AI feels free because nothing ticks while you type. But every prompt, upload and re-run spends compute — through tokens, usage caps, credits or an API bill. This is a working tool for spending it well.
I built this in my own time as a practical resource to help us use AI more efficiently and cost-effectively. It's platform-neutral — the same principles apply across ChatGPT, Claude, Gemini, Copilot and Grok — and free to share with colleagues.
Most people don't need to understand every technical detail of AI to use it well — they need a few simple rules. Everything further down this page exists to back these up: calculators that show the cost, tools to survey your team and draft a policy, and an honest account of what you can actually track.
The reason this matters more each month: AI budget is shifting from a central IT line to something allocated per person, like a laptop or a phone. Once you have your own allowance, efficiency stops being the company's problem and becomes your productivity ceiling.
A TechCrunch report in March 2026 put the share of tech companies including some AI credit allocation in benefits packages at over 40% — up from under 5% eighteen months earlier. At GTC in March 2026, NVIDIA's Jensen Huang went further, floating engineer token budgets worth roughly half base salary, framing them as a productivity amplifier rather than a perk.
The same shift brings limits. Uber exhausted its entire 2026 AI budget by April after Claude Code adoption jumped from 32% to 84% across 5,000 engineers, with individuals generating $500–$2,000 a month — then capped everyone at $1,500 per month per tool. Reported caps elsewhere run from $250 a month at a defence manufacturer to around $2,000 at Workday and Stripe.
Ramp's AI Index, covering 70,000+ businesses, put median spend at about $11.38 per employee per month in June 2026 — while the top 1% spent $7,450. An April cut showed a median of $46 with the middle half falling between $3 and $352. Which end you sit at is driven far more by habits than by headcount.
Token prices have fallen more than 90% since 2023, yet total corporate AI spend has roughly doubled since late 2025 — a textbook Jevons paradox. Cheaper units invited far more use, and agentic tools that call a model repeatedly multiplied it again. Waiting for prices to fall is not a cost strategy.
Ramp names model tier migration — teams moving from a lightweight model to a frontier one for quality reasons, often without anyone noticing — as the single biggest driver of unexpected cost increases. That is precisely what the ten rules on this page are designed to prevent.
If your allowance is capped, every wasted token is one you can't spend on work that matters. Running routine tasks on a premium model in a long thread doesn't just cost the company money — it burns through your budget and throttles you before month end. Discipline buys you more assistant, not less.
Figures as reported mid-2026 and in US dollars. This area moves quickly — treat them as orders of magnitude, not current benchmarks.
Waste doesn't add up — it multiplies. Using a premium model for routine work is expensive; doing it inside a 15-turn thread, twice, asking for long answers each time, compounds that several times over. Set how your team actually works and see the real gap.
Illustrative £ per million tokens, input / output: light 0.20 / 0.80, standard 2.40 / 12, premium 12 / 60. Chosen to reflect the shape of real pricing — output costs roughly 4–5× input, and premium runs roughly 50–60× light — not any one vendor's list price.
Reasoning models also bill hidden "thinking" tokens as output, which can dwarf the visible answer — not counted here. Failed attempts usually stay in the thread too, inflating later turns. Both push the real figure higher.
Prompt caching can cut the cost of re-sent history substantially where a tool supports it, and some tools trim old context automatically. On flat per-seat plans this waste shows up as usage caps and throttling rather than an invoice.
The question is never which brand you're using. It's whether the task needs speed and convenience, or judgement and reasoning. Answer three things and get a recommendation.
Your input is still messy. Use a lighter model to clean it into a precise prompt, then send that to the model above. Don't spend premium capacity working out what you meant to ask.
The saving never comes from avoiding AI — it comes from avoiding bad workflow. This compares the lazy route (one heavy model does everything, twice) against the disciplined three-step route on the same task.
Relative units, not currency — heavy models cost materially more per token than lighter ones, especially on long outputs. The ratio is what matters.
Tokens are the unit you're billed in — very roughly four characters of English each. Paste anything you'd normally send and see the size of the call before you make it.
Rates are illustrative placeholders chosen to show the ratio between tiers, not any vendor's real pricing. Token counts are approximate — every model tokenises slightly differently.
This is the cost almost nobody sees. Every time you send a message, the whole conversation so far is sent again — so a chat doesn't cost a steady amount per turn, it compounds. Ten turns in one thread can cost several times ten turns across fresh chats.
tokens processed in total, because each turn re-reads everything before it.
tokens processed — each task carries only its own context.
more work for the same number of questions.
Assumes roughly 600 tokens of new content per turn. The exact numbers vary; the shape — linear versus compounding — is the point. Caching and context trimming reduce this in some tools, but rarely eliminate it.
Twelve questions covering model discipline, token habits, data risk and verification. Each person answers and gets a short anonymous code. Whoever is coordinating pastes the codes in below to get a company-wide score. No accounts, no server, no personal data — the code carries only the twelve answers.
Answer all twelve to generate your code.
Codes work well up to a few dozen people. Beyond that, collect responses through your usual survey tool, or have a Cloudflare Worker receive them directly — the same twelve questions, stored and averaged automatically. Keep it anonymous either way: a survey that identifies individuals measures how people answer, not how they work.
The Prompt-Builder Method in practice: let a lighter model build the prompt, let the stronger model do the thinking, then let a lighter model tidy up. Copy any of these and paste them straight into whichever tool you're using.
Everything else on this page is about spending money well. This part is about not creating a much more expensive problem. Efficiency matters far less than getting this right.
Business and enterprise tenancies are normally configured so your inputs aren't used to train models, with retention controls and audit logs. Free and personal accounts often work differently. The same prompt can carry very different risk depending on where it's typed — always use the company tooling for work.
You rarely need the real names to get a useful answer. Replace parties with "Company A" and "Supplier B", round the figures, and strip identifiers. You keep the analytical value and remove most of the risk — and the prompt gets shorter and cheaper too.
AI output can be confidently wrong. It doesn't know your matter, your client or your obligations. Anything client-facing or decision-bearing needs a human check by someone qualified to give it — that's a professional responsibility no tool transfers.
This is general good practice, not legal advice, and it doesn't replace our own policies. UK data protection obligations depend on the specifics of what you're handling — if you're unsure about a particular document or dataset, ask our data protection lead or legal team before pasting it anywhere.
Two risks that don't show up on any usage dashboard, and which cost far more than tokens when they land. Both are worth a named owner rather than general awareness.
Hidden instructions inside a document, email or web page that the AI reads and obeys — telling it to ignore its rules or leak what it's seen. The more an AI tool can browse, read files or use connectors, the more this matters. Treat AI-read content as untrusted input.
Tools that browse, click, send email or reach into systems widen the blast radius considerably. Grant the narrowest access that works, review what each connector can actually reach, and keep a human approving anything that sends, pays or deletes.
Keys, tokens and passwords pasted into a prompt for convenience. Assume anything pasted may be logged somewhere. Rotate immediately if it happens — and never paste them in the first place.
Personal accounts and unsanctioned tools sit outside every admin console, DLP rule and audit log. This is where most incidents originate. Sanctioned tooling that's actually good is the best control — people route around friction.
AI has removed the spelling mistakes and awkward phrasing that staff were trained to spot, and makes convincing voice and video impersonation cheap. Verification procedures for payment and data requests now matter more than spotting a badly written email.
Browser extensions, plugins, MCP servers and AI features bolted onto existing SaaS all process your data. Each is a third party with access. Review them the way you'd review any other processor, and keep a register.
Vendor terms usually assign output rights to the customer, but that isn't the same as the output being protectable. Purely AI-generated material may attract weak or no copyright protection in some jurisdictions — which matters if it's meant to be a defensible asset.
Output can reproduce material from training data, particularly for code, brand assets and distinctive style. Anything going into a client deliverable or public campaign deserves a check rather than an assumption.
Pasting proprietary methods, source code or unpublished material into a tool that may retain or train on it can weaken trade-secret protection. Confidentiality is often maintained by behaving confidentially — casual pasting undermines that argument later.
An increasing number of client agreements now restrict or require disclosure of AI use in delivered work. Check the engagement terms before using AI on client matters — the obligation may already exist and be unmet.
Feeding in licensed content, someone else's documents or paywalled material may breach the licence you hold it under, regardless of what the AI then does with it.
Decide as a business whether AI-assisted work is disclosed, to whom, and how. Deciding once is far better than each person improvising when a client asks.
General guidance, not legal advice. IP and data protection positions vary by jurisdiction, by contract and by the specific facts — take proper advice before relying on any of this for a live matter.
Answer a few questions and get a drafted policy you can drop into the staff handbook. It's a starting point written in plain English — not a finished legal document.
A drafting aid, not legal advice. Have it reviewed by whoever owns HR policy and by legal before it goes into the handbook, and check it sits consistently with your existing data protection, IT and confidentiality policies.
AI spend and AI risk are now board-level items in most organisations, but they rarely get reported consistently. This builds a one-page template you can paste into the pack and reuse each quarter.
Report the same handful of measures every quarter rather than a new selection each time. A board can see a trend in three numbers it recognises; it can't see anything useful in twelve it's meeting for the first time.
This is the question every finance and IT team asks, and the honest answer is "it depends what you bought." A subscription hides the meter; an API exposes it. Here's what each setup really lets you see, as of mid-2026.
| What you bought | Per-user visibility | Per-token / true cost | By model & tool | Programmatic API |
|---|---|---|---|---|
| ChatGPT Team seat-based |
Aggregate | Hidden | Limited | No |
| ChatGPT Enterprise credit-based, 2026 console |
Named users | Credits | Yes | Cost API |
| Claude Team pooled tokens |
Aggregate | Pool + caps | Some | No |
| Claude Enterprise Analytics API, 2026 |
Named users | $ per user | Model + tool | Analytics API |
| Microsoft 365 Copilot licence-based |
Adoption only | Hidden in licence | App adoption | Graph API |
| Grok Business seat-based, $30/user/mo |
Console analytics | Hidden in seat | Limited | No |
| Grok Enterprise custom pricing, SSO + SCIM |
Named users | Hidden in contract | Limited | Audit tools |
| Direct API OpenAI / Anthropic / etc. |
By key/tag | Exact tokens | Full detail | Full API |
Team and licence tiers mostly show adoption — active users, prompt counts, which apps — not what any single task cost. Good for "is anyone using this?", weak for "where is spend going?"
Through 2026, all three vendors shipped admin consoles with per-user, per-model breakdowns and spend caps. ChatGPT bills Enterprise in credits; Claude Enterprise reports actual dollars per user. Both expose an API.
Direct API usage is the only place with true per-token, per-model, per-call cost — down to input, output and cache. Route calls through one gateway and tag by team, or the "one line called OpenAI" hides four different things.
None of this sees personal accounts. Analysts estimate most AI use in large firms happens on unsanctioned tools. A dashboard only covers what's on the corporate licence — the rest is invisible.
Vendor tooling in this area is changing month to month — treat this as a mid-2026 snapshot and confirm against current admin docs before a procurement decision.
Same goal — contain and understand spend — three different mental models. Knowing which one you're in tells you how to budget and what to watch.
Team buys seats. Enterprise moved to a credit model with a 2026 Global Admin Console: slice consumption by user, group, project and model, set workspace defaults, cap power users, pull it via a Cost API. Office and agent features now meter in tokens too.
No seat limits — usage draws from one org token pool, with spend caps at every level. Enterprise analytics report cost by group and by user, showing artifacts, files and tools next to their cost, filtered by your existing org chart. Two APIs expose it.
Cost lives inside the M365 licence, so admin reporting is about adoption, not spend: active users, prompts, app usage, agent uptake — via the admin centre, Viva Insights and Graph API. Great for rollout tracking, quiet on per-task cost.
Business and Enterprise tiers arrived in January 2026 — Business at $30 per seat, Enterprise on custom pricing. Management runs through the xAI console: seats, consolidated billing, role-based controls and usage analytics, with Enterprise adding SSO, SCIM directory sync, audit tooling and a customer-key data vault. Cost sits inside the seat, so reporting leans toward adoption rather than per-task spend.
Gemini and other Workspace-bundled assistants follow broadly the licence-based pattern described for Copilot — cost inside the licence, reporting focused on adoption — but aren't broken out separately here. Check current admin documentation before relying on any of this for procurement.
A Word or text document is like handing someone a typed script. A PDF is often like handing them a photograph of the page — the layout, the headings, the tables, the signatures and sometimes the surrounding clutter too. That extra information is sometimes essential and sometimes just expensive.
| Use Word, text or an extract when… | Use a PDF when… |
|---|---|
| Only the wording matters | The document only exists as a PDF |
| Formatting isn't important | Signatures or dates matter |
| You're drafting, editing or summarising | The layout carries meaning |
| You already know which section matters | Charts or tables need visual review |
| The PDF has many irrelevant pages | You're reviewing evidence exactly as received |
Before uploading a PDF, ask one question: do I need the AI to understand this as a page, or only to understand the words? If it's just the words, send Word, text or a selected extract.
Most wasted AI capacity comes down to a handful of recurring habits.
Don't spend premium reasoning capacity organising messy notes — unless the organising itself needs deep judgement.
Long context isn't always better. It raises cost and dilutes focus, and the answer often gets worse, not better.
Useful where layout matters, wasteful where only the text does.
A bad prompt should be fixed, not repeated on an expensive model. Three heavy re-runs cost more than one good prompt.
Cost control shouldn't come at the expense of quality where the risk is real. Match the model to the cost of being wrong.
Output tokens are the expensive ones. Ask for three bullets when three bullets will do.
Every message resends the whole conversation. Start a fresh chat when you start a new task.
Different tools and tiers have different strengths, costs and limits. Choose deliberately rather than by habit.
You don't need any of this to use AI well, but it helps when reading a bill or an admin dashboard.
If something here is wrong, out of date, or missing — tell me. The vendor tracking section in particular ages quickly, and I'd rather hear it from you than leave it stale. Happy to send the editable policy template or the Word version of the guide too.