A practical guide to model selection, cost & responsible AI use

Every prompt runs a meter.
Most people never see it.

AI feels free because nothing ticks while you type. But every prompt, upload and re-run spends compute — through tokens, usage caps, credits or an API bill. This is a working tool for spending it well.

The phone-credit modelcost per task →
Lighter model
The local call
Prep, proofing, formatting, checklists
Standard model
The mobile call
Client emails, first drafts, summaries
Reasoning model
The international call
Legal, financial, strategy, high-stakes
Note

I built this in my own time as a practical resource to help us use AI more efficiently and cost-effectively. It's platform-neutral — the same principles apply across ChatGPT, Claude, Gemini, Copilot and Grok — and free to share with colleagues.

THE CORE

The 10 commandments

Most people don't need to understand every technical detail of AI to use it well — they need a few simple rules. Everything further down this page exists to back these up: calculators that show the cost, tools to survey your team and draft a policy, and an honest account of what you can actually track.

Start with the outcome, not the tool. Know whether you need a summary, draft, analysis, decision or table before you pick a model.
Don't use the most powerful model by default. Reserve it for genuine reasoning, judgement or high accuracy — not formatting or proofreading.
Use lighter models for preparation. Clarify the task and build the prompt cheaply. Batch related questions into one prompt rather than many turns.
Use stronger models for judgement. Legal, financial, technical and strategic work — where the cost of being wrong is high.
Use lighter models for polish. Proofing, shortening and reformatting don't need premium capacity.
Treat tokens like old phone credit. And remember a long chat resends its whole history every turn — start a fresh chat for a new task. Often the biggest hidden cost.
Don't paste or upload more than is needed. Ask only for the output length you need, and switch off tools you're not using.
Choose the right file format. Clean text beats a PDF when only the words matter; use PDF when layout, signatures or tables do.
Improve bad prompts before trying again. Don't re-run a weak prompt on a heavy model — fix it first.
Use AI generously, but not carelessly. The aim is better AI use, not less. Cut waste, not adoption.
WHY NOW

Tokens are becoming part of the package

The reason this matters more each month: AI budget is shifting from a central IT line to something allocated per person, like a laptop or a phone. Once you have your own allowance, efficiency stops being the company's problem and becomes your productivity ceiling.

It's now a benefit line

A TechCrunch report in March 2026 put the share of tech companies including some AI credit allocation in benefits packages at over 40% — up from under 5% eighteen months earlier. At GTC in March 2026, NVIDIA's Jensen Huang went further, floating engineer token budgets worth roughly half base salary, framing them as a productivity amplifier rather than a perk.

And a cap

The same shift brings limits. Uber exhausted its entire 2026 AI budget by April after Claude Code adoption jumped from 32% to 84% across 5,000 engineers, with individuals generating $500–$2,000 a month — then capped everyone at $1,500 per month per tool. Reported caps elsewhere run from $250 a month at a defence manufacturer to around $2,000 at Workday and Stripe.

The spread is enormous

Ramp's AI Index, covering 70,000+ businesses, put median spend at about $11.38 per employee per month in June 2026 — while the top 1% spent $7,450. An April cut showed a median of $46 with the middle half falling between $3 and $352. Which end you sit at is driven far more by habits than by headcount.

Cheaper tokens, bigger bills

Token prices have fallen more than 90% since 2023, yet total corporate AI spend has roughly doubled since late 2025 — a textbook Jevons paradox. Cheaper units invited far more use, and agentic tools that call a model repeatedly multiplied it again. Waiting for prices to fall is not a cost strategy.

What actually drives the overruns

Ramp names model tier migration — teams moving from a lightweight model to a frontier one for quality reasons, often without anyone noticing — as the single biggest driver of unexpected cost increases. That is precisely what the ten rules on this page are designed to prevent.

What it means for you

If your allowance is capped, every wasted token is one you can't spend on work that matters. Running routine tasks on a premium model in a long thread doesn't just cost the company money — it burns through your budget and throttles you before month end. Discipline buys you more assistant, not less.

Figures as reported mid-2026 and in US dollars. This area moves quickly — treat them as orders of magnitude, not current benchmarks.

TOOL 01

The waste calculator

Waste doesn't add up — it multiplies. Using a premium model for routine work is expensive; doing it inside a 15-turn thread, twice, asking for long answers each time, compounds that several times over. Set how your team actually works and see the real gap.

Company AI spend & waste estimator

Scale How people actually work
Waste multiplier on routine work
Annual spend as things are
If the same work were run well
Avoidable
Model
Thread
Re-runs
Length

The rates used

Illustrative £ per million tokens, input / output: light 0.20 / 0.80, standard 2.40 / 12, premium 12 / 60. Chosen to reflect the shape of real pricing — output costs roughly 4–5× input, and premium runs roughly 50–60× light — not any one vendor's list price.

Where it stays conservative

Reasoning models also bill hidden "thinking" tokens as output, which can dwarf the visible answer — not counted here. Failed attempts usually stay in the thread too, inflating later turns. Both push the real figure higher.

Where it may overstate

Prompt caching can cut the cost of re-sent history substantially where a tool supports it, and some tools trim old context automatically. On flat per-seat plans this waste shows up as usage caps and throttling rather than an invoice.

TOOL 02

Which model does this task need?

The question is never which brand you're using. It's whether the task needs speed and convenience, or judgement and reasoning. Answer three things and get a recommendation.

Model selector

Recommended
Lighter model
Fast and cheap is all this needs.
TOOL 03

Two ways to do the same job

The saving never comes from avoiding AI — it comes from avoiding bad workflow. This compares the lazy route (one heavy model does everything, twice) against the disciplined three-step route on the same task.

Workflow comparison

Lazy
100
Disciplined
45

The lazy route

The disciplined route

Relative units, not currency — heavy models cost materially more per token than lighter ones, especially on long outputs. The ratio is what matters.

TOOL 04

What does this prompt actually cost?

Tokens are the unit you're billed in — very roughly four characters of English each. Paste anything you'd normally send and see the size of the call before you make it.

Token & cost estimator

Estimated input tokens
0
Nothing pasted yet.
Lighter
Standard
Reasoning

Rates are illustrative placeholders chosen to show the ratio between tiers, not any vendor's real pricing. Token counts are approximate — every model tokenises slightly differently.

THE HIDDEN COST

Why one long chat costs more than ten short ones

This is the cost almost nobody sees. Every time you send a message, the whole conversation so far is sent again — so a chat doesn't cost a steady amount per turn, it compounds. Ten turns in one thread can cost several times ten turns across fresh chats.

Context accumulation

One long thread — cumulative tokens processed Fresh chat per task

One long thread

0

tokens processed in total, because each turn re-reads everything before it.

Fresh chat per task

0

tokens processed — each task carries only its own context.

The difference

more work for the same number of questions.

Assumes roughly 600 tokens of new content per turn. The exact numbers vary; the shape — linear versus compounding — is the point. Caching and context trimming reduce this in some tools, but rarely eliminate it.

TOOL 05

Survey your team

Twelve questions covering model discipline, token habits, data risk and verification. Each person answers and gets a short anonymous code. Whoever is coordinating pastes the codes in below to get a company-wide score. No accounts, no server, no personal data — the code carries only the twelve answers.

AI discipline survey

Answer all twelve to generate your code.

Scaling it

Codes work well up to a few dozen people. Beyond that, collect responses through your usual survey tool, or have a Cloudflare Worker receive them directly — the same twelve questions, stored and averaged automatically. Keep it anonymous either way: a survey that identifies individuals measures how people answer, not how they work.

COPY & USE

Prompt templates

The Prompt-Builder Method in practice: let a lighter model build the prompt, let the stronger model do the thinking, then let a lighter model tidy up. Copy any of these and paste them straight into whichever tool you're using.

BEFORE YOU PASTE

What should never go into an AI tool

Everything else on this page is about spending money well. This part is about not creating a much more expensive problem. Efficiency matters far less than getting this right.

Stop

Don't paste this anywhere
  • Personal data about clients, staff or third parties — names, contact details, health, HR or financial records
  • Anything covered by a confidentiality agreement or legal privilege
  • Credentials, API keys, passwords or access tokens
  • Unreleased financial results, M&A material or price-sensitive information
  • Anything you'd hesitate to forward to someone outside the business

Care

Check before you use it
  • Internal documents that could be de-identified first — strip names and figures, keep the structure
  • Client work where the sector or facts could identify them
  • Anything going into a personal account rather than the company tenancy
  • Third-party material you don't own the rights to
  • Outputs you plan to rely on without checking — always verify before use

Fine

Everyday safe use
  • Public information and published material
  • Your own drafting, notes and rough thinking
  • Anonymised or invented examples that mirror the real problem
  • General questions about law, process or best practice
  • Formatting, structuring and proofreading non-sensitive text

Company account ≠ personal account

Business and enterprise tenancies are normally configured so your inputs aren't used to train models, with retention controls and audit logs. Free and personal accounts often work differently. The same prompt can carry very different risk depending on where it's typed — always use the company tooling for work.

De-identify, don't avoid

You rarely need the real names to get a useful answer. Replace parties with "Company A" and "Supplier B", round the figures, and strip identifiers. You keep the analytical value and remove most of the risk — and the prompt gets shorter and cheaper too.

The output is a draft, not advice

AI output can be confidently wrong. It doesn't know your matter, your client or your obligations. Anything client-facing or decision-bearing needs a human check by someone qualified to give it — that's a professional responsibility no tool transfers.

Please note

This is general good practice, not legal advice, and it doesn't replace our own policies. UK data protection obligations depend on the specifics of what you're handling — if you're unsure about a particular document or dataset, ask our data protection lead or legal team before pasting it anywhere.

RISK

Cyber and intellectual property

Two risks that don't show up on any usage dashboard, and which cost far more than tokens when they land. Both are worth a named owner rather than general awareness.

Cyber exposure

Prompt injection

Hidden instructions inside a document, email or web page that the AI reads and obeys — telling it to ignore its rules or leak what it's seen. The more an AI tool can browse, read files or use connectors, the more this matters. Treat AI-read content as untrusted input.

Agents and connectors

Tools that browse, click, send email or reach into systems widen the blast radius considerably. Grant the narrowest access that works, review what each connector can actually reach, and keep a human approving anything that sends, pays or deletes.

Credential leakage

Keys, tokens and passwords pasted into a prompt for convenience. Assume anything pasted may be logged somewhere. Rotate immediately if it happens — and never paste them in the first place.

Shadow AI

Personal accounts and unsanctioned tools sit outside every admin console, DLP rule and audit log. This is where most incidents originate. Sanctioned tooling that's actually good is the best control — people route around friction.

Better phishing

AI has removed the spelling mistakes and awkward phrasing that staff were trained to spot, and makes convincing voice and video impersonation cheap. Verification procedures for payment and data requests now matter more than spotting a badly written email.

Supply chain

Browser extensions, plugins, MCP servers and AI features bolted onto existing SaaS all process your data. Each is a third party with access. Review them the way you'd review any other processor, and keep a register.

Intellectual property

Who owns the output?

Vendor terms usually assign output rights to the customer, but that isn't the same as the output being protectable. Purely AI-generated material may attract weak or no copyright protection in some jurisdictions — which matters if it's meant to be a defensible asset.

Inbound infringement

Output can reproduce material from training data, particularly for code, brand assets and distinctive style. Anything going into a client deliverable or public campaign deserves a check rather than an assumption.

Your IP going out

Pasting proprietary methods, source code or unpublished material into a tool that may retain or train on it can weaken trade-secret protection. Confidentiality is often maintained by behaving confidentially — casual pasting undermines that argument later.

Client contracts

An increasing number of client agreements now restrict or require disclosure of AI use in delivered work. Check the engagement terms before using AI on client matters — the obligation may already exist and be unmet.

Third-party material

Feeding in licensed content, someone else's documents or paywalled material may breach the licence you hold it under, regardless of what the AI then does with it.

Disclosure and attribution

Decide as a business whether AI-assisted work is disclosed, to whom, and how. Deciding once is far better than each person improvising when a client asks.

Please note

General guidance, not legal advice. IP and data protection positions vary by jurisdiction, by contract and by the specific facts — take proper advice before relying on any of this for a live matter.

TOOL 06

Build a Gen AI policy

Answer a few questions and get a drafted policy you can drop into the staff handbook. It's a starting point written in plain English — not a finished legal document.

Policy generator

Which tools are approved for work use?
Client and personal data
Include clauses on…

Draft policy


      

A drafting aid, not legal advice. Have it reviewed by whoever owns HR policy and by legal before it goes into the handbook, and check it sits consistently with your existing data protection, IT and confidentiality policies.

TOOL 07

The board pack page

AI spend and AI risk are now board-level items in most organisations, but they rarely get reported consistently. This builds a one-page template you can paste into the pack and reuse each quarter.

Board reporting template

Which sections should the page cover?

Board pack template


      
Tip

Report the same handful of measures every quarter rather than a new selection each time. A board can see a trend in three numbers it recognises; it can't see anything useful in twelve it's meeting for the first time.

REALITY CHECK

What you can — and can't — actually track

This is the question every finance and IT team asks, and the honest answer is "it depends what you bought." A subscription hides the meter; an API exposes it. Here's what each setup really lets you see, as of mid-2026.

What you bought Per-user visibility Per-token / true cost By model & tool Programmatic API
ChatGPT Team
seat-based
Aggregate Hidden Limited No
ChatGPT Enterprise
credit-based, 2026 console
Named users Credits Yes Cost API
Claude Team
pooled tokens
Aggregate Pool + caps Some No
Claude Enterprise
Analytics API, 2026
Named users $ per user Model + tool Analytics API
Microsoft 365 Copilot
licence-based
Adoption only Hidden in licence App adoption Graph API
Grok Business
seat-based, $30/user/mo
Console analytics Hidden in seat Limited No
Grok Enterprise
custom pricing, SSO + SCIM
Named users Hidden in contract Limited Audit tools
Direct API
OpenAI / Anthropic / etc.
By key/tag Exact tokens Full detail Full API

Subscriptions hide the meter

Team and licence tiers mostly show adoption — active users, prompt counts, which apps — not what any single task cost. Good for "is anyone using this?", weak for "where is spend going?"

Enterprise adds the dashboard

Through 2026, all three vendors shipped admin consoles with per-user, per-model breakdowns and spend caps. ChatGPT bills Enterprise in credits; Claude Enterprise reports actual dollars per user. Both expose an API.

APIs expose everything

Direct API usage is the only place with true per-token, per-model, per-call cost — down to input, output and cache. Route calls through one gateway and tag by team, or the "one line called OpenAI" hides four different things.

The shadow-AI gap

None of this sees personal accounts. Analysts estimate most AI use in large firms happens on unsanctioned tools. A dashboard only covers what's on the corporate licence — the rest is invisible.

Vendor tooling in this area is changing month to month — treat this as a mid-2026 snapshot and confirm against current admin docs before a procurement decision.

SIDE BY SIDE

How the three stacks meter cost

Same goal — contain and understand spend — three different mental models. Knowing which one you're in tells you how to budget and what to watch.

ChatGPT

Seats, then credits

Team buys seats. Enterprise moved to a credit model with a 2026 Global Admin Console: slice consumption by user, group, project and model, set workspace defaults, cap power users, pull it via a Cost API. Office and agent features now meter in tokens too.

Claude

A utility meter

No seat limits — usage draws from one org token pool, with spend caps at every level. Enterprise analytics report cost by group and by user, showing artifacts, files and tools next to their cost, filtered by your existing org chart. Two APIs expose it.

Copilot

Bundled in the licence

Cost lives inside the M365 licence, so admin reporting is about adoption, not spend: active users, prompts, app usage, agent uptake — via the admin centre, Viva Insights and Graph API. Great for rollout tracking, quiet on per-task cost.

Grok

Seats with a console

Business and Enterprise tiers arrived in January 2026 — Business at $30 per seat, Enterprise on custom pricing. Management runs through the xAI console: seats, consolidated billing, role-based controls and usage analytics, with Enterprise adding SSO, SCIM directory sync, audit tooling and a customer-key data vault. Cost sits inside the seat, so reporting leans toward adoption rather than per-task spend.

Gemini and other Workspace-bundled assistants follow broadly the licence-based pattern described for Copilot — cost inside the licence, reporting focused on adoption — but aren't broken out separately here. Check current admin documentation before relying on any of this for procurement.

PRACTICAL

Word or PDF?

A Word or text document is like handing someone a typed script. A PDF is often like handing them a photograph of the page — the layout, the headings, the tables, the signatures and sometimes the surrounding clutter too. That extra information is sometimes essential and sometimes just expensive.

Use Word, text or an extract when…Use a PDF when…
Only the wording mattersThe document only exists as a PDF
Formatting isn't importantSignatures or dates matter
You're drafting, editing or summarisingThe layout carries meaning
You already know which section mattersCharts or tables need visual review
The PDF has many irrelevant pagesYou're reviewing evidence exactly as received
Rule

Before uploading a PDF, ask one question: do I need the AI to understand this as a page, or only to understand the words? If it's just the words, send Word, text or a selected extract.

WATCH FOR

Common mistakes

Most wasted AI capacity comes down to a handful of recurring habits.

Reaching for the strongest model too early

Don't spend premium reasoning capacity organising messy notes — unless the organising itself needs deep judgement.

Pasting too much material

Long context isn't always better. It raises cost and dilutes focus, and the answer often gets worse, not better.

Uploading PDFs by habit

Useful where layout matters, wasteful where only the text does.

Re-running weak prompts

A bad prompt should be fixed, not repeated on an expensive model. Three heavy re-runs cost more than one good prompt.

Using light models for high-stakes calls

Cost control shouldn't come at the expense of quality where the risk is real. Match the model to the cost of being wrong.

Asking for long answers by default

Output tokens are the expensive ones. Ask for three bullets when three bullets will do.

Letting one chat run forever

Every message resends the whole conversation. Start a fresh chat when you start a new task.

Treating every tool as identical

Different tools and tiers have different strengths, costs and limits. Choose deliberately rather than by habit.

PLAIN ENGLISH

The words people use

You don't need any of this to use AI well, but it helps when reading a bill or an admin dashboard.

TokenThe unit AI reads and bills in — roughly four characters, or about ¾ of a word. "Efficiency matters" is about five tokens. Both what you send and what comes back are counted.
Input vs output tokensWhat you send in, versus what the model writes back. Output is normally the more expensive of the two — which is why asking for shorter answers saves real money.
Context windowHow much the model can hold in mind at once, measured in tokens. A long conversation or a big document fills it up — and everything in it is re-read on every turn.
PromptWhatever you send: the question, the instructions, the pasted material and any files, all together.
Reasoning modelA model that works through a problem in steps before answering. Better at judgement and complex analysis, slower and materially more expensive per task.
InferenceThe act of the model producing an answer. "Inference cost" is simply what it costs to run, as opposed to what it cost to build.
CachingReusing material the model has already processed rather than paying for it again. Helps most when the same long document or instruction is sent repeatedly.
CreditsA vendor's own unit of account sitting on top of tokens. Useful for capping spend; a layer of abstraction away from what a task truly cost.
Seat vs pooled licensingSeat-based charges per person regardless of use. Pooled draws everyone's usage from one shared allowance — so a few heavy users affect everyone else.
Shadow AIWork AI use happening on personal accounts and unsanctioned tools. Invisible to every admin dashboard, and where most data-protection incidents start.
HallucinationOutput that is fluent, confident and wrong. The reason every AI answer needs a human check before anyone relies on it.
GET IN TOUCH

Questions, corrections and feedback

If something here is wrong, out of date, or missing — tell me. The vendor tracking section in particular ages quickly, and I'd rather hear it from you than leave it stale. Happy to send the editable policy template or the Word version of the guide too.