A practical guide to model selection, cost & responsible AI use

Every prompt runs a meter.
Most people never see it.

AI feels free because nothing ticks while you type. But every prompt, upload and re-run spends compute — through tokens, usage caps, credits or an API bill. This is a working tool for spending it well.

The phone-credit modelcost per task →
Lighter model
The local call
Prep, proofing, formatting, checklists
Standard model
The mobile call
Client emails, first drafts, summaries
Reasoning model
The international call
Legal, financial, strategy, high-stakes
Note

By Michael Berger. I built this in my own time, as the resource I kept wishing existed — practical, platform-neutral and free. The same principles apply across ChatGPT, Claude, Gemini, Copilot and Grok. Nothing to sign up for, nothing to buy. Take it, use it, share it. If something here is wrong or out of date, tell me.

THE CORE

The 10 commandments

Most people don't need to understand every technical detail of AI to use it well — they need a few simple rules. Everything further down this page exists to back these up: calculators that show the cost, tools to survey your team and draft a policy, and an honest account of what you can actually track.

Start with the outcome, not the tool. Know whether you need a summary, draft, analysis, decision or table before you pick a model.
Don't use the most powerful model by default. Reserve it for genuine reasoning, judgement or high accuracy — not formatting or proofreading.
Use lighter models for preparation. Clarify the task and build the prompt cheaply. Batch related questions into one prompt rather than many turns.
Use stronger models for judgement. Legal, financial, technical and strategic work — where the cost of being wrong is high.
Use lighter models for polish. Proofing, shortening and reformatting don't need premium capacity.
Treat tokens like old phone credit. And remember a long chat resends its whole history every turn — start a fresh chat for a new task. Often the biggest hidden cost.
Don't paste or upload more than is needed. Ask only for the output length you need, and switch off tools you're not using.
Choose the right file format. Clean text beats a PDF when only the words matter; use PDF when layout, signatures or tables do.
Improve bad prompts before trying again. Don't re-run a weak prompt on a heavy model — fix it first.
Use AI generously, but not carelessly. The aim is better AI use, not less. Cut waste, not adoption.
WHY NOW

Tokens are becoming part of the package

The reason this matters more each month: AI budget is shifting from a central IT line to something allocated per person, like a laptop or a phone. Once you have your own allowance, efficiency stops being the company's problem and becomes your productivity ceiling.

It's now a benefit line

A TechCrunch report in March 2026 put the share of tech companies including some AI credit allocation in benefits packages at over 40% — up from under 5% eighteen months earlier. At GTC in March 2026, NVIDIA's Jensen Huang went further, floating engineer token budgets worth roughly half base salary, framing them as a productivity amplifier rather than a perk.

And a cap

The same shift brings limits. Uber exhausted its entire 2026 AI budget by April after Claude Code adoption jumped from 32% to 84% across 5,000 engineers, with individuals generating $500–$2,000 a month — then capped everyone at $1,500 per month per tool. Reported caps elsewhere run from $250 a month at a defence manufacturer to around $2,000 at Workday and Stripe.

The spread is enormous

Ramp's AI Index, covering 70,000+ businesses, put median spend at about $11.38 per employee per month in June 2026 — while the top 1% spent $7,450. An April cut showed a median of $46 with the middle half falling between $3 and $352. Which end you sit at is driven far more by habits than by headcount.

Cheaper tokens, bigger bills

Token prices have fallen more than 90% since 2023, yet total corporate AI spend has roughly doubled since late 2025 — a textbook Jevons paradox. Cheaper units invited far more use, and agentic tools that call a model repeatedly multiplied it again. Waiting for prices to fall is not a cost strategy.

What actually drives the overruns

Ramp names model tier migration — teams moving from a lightweight model to a frontier one for quality reasons, often without anyone noticing — as the single biggest driver of unexpected cost increases. That is precisely what the ten rules on this page are designed to prevent.

What it means for you

If your allowance is capped, every wasted token is one you can't spend on work that matters. Running routine tasks on a premium model in a long thread doesn't just cost the company money — it burns through your budget and throttles you before month end. Discipline buys you more assistant, not less.

Figures as reported mid-2026 and in US dollars. This area moves quickly — treat them as orders of magnitude, not current benchmarks.

WATCH FOR

Common mistakes

Most wasted AI capacity comes down to a handful of recurring habits.

Reaching for the strongest model too early

Don't spend premium reasoning capacity organising messy notes — unless the organising itself needs deep judgement.

Pasting too much material

Long context isn't always better. It raises cost and dilutes focus, and the answer often gets worse, not better.

Uploading PDFs by habit

Useful where layout matters, wasteful where only the text does.

Re-running weak prompts

A bad prompt should be fixed, not repeated on an expensive model. Three heavy re-runs cost more than one good prompt.

Using light models for high-stakes calls

Cost control shouldn't come at the expense of quality where the risk is real. Match the model to the cost of being wrong.

Asking for long answers by default

Output tokens are the expensive ones. Ask for three bullets when three bullets will do.

Letting one chat run forever

Every message resends the whole conversation. Start a fresh chat when you start a new task.

Treating every tool as identical

Different tools and tiers have different strengths, costs and limits. Choose deliberately rather than by habit.

PRACTICAL

Word or PDF?

A Word or text document is like handing someone a typed script. A PDF is often like handing them a photograph of the page — the layout, the headings, the tables, the signatures and sometimes the surrounding clutter too. That extra information is sometimes essential and sometimes just expensive.

Use Word, text or an extract when…Use a PDF when…
Only the wording mattersThe document only exists as a PDF
Formatting isn't importantSignatures or dates matter
You're drafting, editing or summarisingThe layout carries meaning
You already know which section mattersCharts or tables need visual review
The PDF has many irrelevant pagesYou're reviewing evidence exactly as received
Rule

Before uploading a PDF, ask one question: do I need the AI to understand this as a page, or only to understand the words? If it's just the words, send Word, text or a selected extract.

PLAIN ENGLISH

The words people use

You don't need any of this to use AI well, but it helps when reading a bill or an admin dashboard.

TokenThe unit AI reads and bills in — roughly four characters, or about ¾ of a word. "Efficiency matters" is about five tokens. Both what you send and what comes back are counted.
Input vs output tokensWhat you send in, versus what the model writes back. Output is normally the more expensive of the two — which is why asking for shorter answers saves real money.
Context windowHow much the model can hold in mind at once, measured in tokens. A long conversation or a big document fills it up — and everything in it is re-read on every turn.
PromptWhatever you send: the question, the instructions, the pasted material and any files, all together.
Reasoning modelA model that works through a problem in steps before answering. Better at judgement and complex analysis, slower and materially more expensive per task.
InferenceThe act of the model producing an answer. "Inference cost" is simply what it costs to run, as opposed to what it cost to build.
CachingReusing material the model has already processed rather than paying for it again. Helps most when the same long document or instruction is sent repeatedly.
CreditsA vendor's own unit of account sitting on top of tokens. Useful for capping spend; a layer of abstraction away from what a task truly cost.
Seat vs pooled licensingSeat-based charges per person regardless of use. Pooled draws everyone's usage from one shared allowance — so a few heavy users affect everyone else.
Shadow AIWork AI use happening on personal accounts and unsanctioned tools. Invisible to every admin dashboard, and where most data-protection incidents start.
HallucinationOutput that is fluent, confident and wrong. The reason every AI answer needs a human check before anyone relies on it.
WHO WROTE THIS

About

Michael Angelo Berger

I work in portfolio, investments and M&A, which means I spend a lot of time looking at how businesses actually spend money — and increasingly at how they spend it on AI.

This started as a document for my own use. Almost everything written about AI in business was either vendor marketing or abstract strategy, and very little of it answered the practical question people were really asking: which model should I use for this, and what is it costing us? I wrote down what I'd worked out, then kept going until it turned into this.

It's built in my own time and given away deliberately. There's nothing to buy, no email list, and no consultancy behind it. If it's useful to your team, take it.

Connect with me on LinkedIn →

How it's kept honest

Figures are sourced and dated rather than asserted. Where something is an opinion rather than a fact, it says so.

The vendor and pricing sections age quickly, so each page carries a review date and I'd genuinely rather be told when something is out of date than leave it standing.

Corrections have already come from readers. That's the intention.

WHAT PEOPLE SAY

Reviews and feedback

If this has been useful, saying so publicly helps other people decide whether it's worth their time. If something is wrong or out of date, that's even more useful — the vendor tracking in particular ages quickly, and I'd rather hear it from you.