The landscape
What your admin console can and cannot actually see, how the vendors differ, and what provider economics suggest about where pricing goes next.
What you can — and can't — actually track
This is the question every finance and IT team asks, and the honest answer is "it depends what you bought." A subscription hides the meter; an API exposes it. Here's what each setup really lets you see, as of mid-2026.
| What you bought | Per-user visibility | Per-token / true cost | By model & tool | Programmatic API |
|---|---|---|---|---|
| ChatGPT Team seat-based |
Aggregate | Hidden | Limited | No |
| ChatGPT Enterprise credit-based, 2026 console |
Named users | Credits | Yes | Cost API |
| Claude Team pooled tokens |
Aggregate | Pool + caps | Some | No |
| Claude Enterprise Analytics API, 2026 |
Named users | $ per user | Model + tool | Analytics API |
| Microsoft 365 Copilot licence-based |
Adoption only | Hidden in licence | App adoption | Graph API |
| Grok Business seat-based, $30/user/mo |
Console analytics | Hidden in seat | Limited | No |
| Grok Enterprise custom pricing, SSO + SCIM |
Named users | Hidden in contract | Limited | Audit tools |
| Direct API OpenAI / Anthropic / etc. |
By key/tag | Exact tokens | Full detail | Full API |
Subscriptions hide the meter
Team and licence tiers mostly show adoption — active users, prompt counts, which apps — not what any single task cost. Good for "is anyone using this?", weak for "where is spend going?"
Enterprise adds the dashboard
Through 2026, all three vendors shipped admin consoles with per-user, per-model breakdowns and spend caps. ChatGPT bills Enterprise in credits; Claude Enterprise reports actual dollars per user. Both expose an API.
APIs expose everything
Direct API usage is the only place with true per-token, per-model, per-call cost — down to input, output and cache. Route calls through one gateway and tag by team, or the "one line called OpenAI" hides four different things.
The shadow-AI gap
None of this sees personal accounts. Analysts estimate most AI use in large firms happens on unsanctioned tools. A dashboard only covers what's on the corporate licence — the rest is invisible.
Vendor tooling in this area is changing month to month — treat this as a mid-2026 snapshot and confirm against current admin docs before a procurement decision.
How the three stacks meter cost
Same goal — contain and understand spend — three different mental models. Knowing which one you're in tells you how to budget and what to watch.
Seats, then credits
Team buys seats. Enterprise moved to a credit model with a 2026 Global Admin Console: slice consumption by user, group, project and model, set workspace defaults, cap power users, pull it via a Cost API. Office and agent features now meter in tokens too.
A utility meter
No seat limits — usage draws from one org token pool, with spend caps at every level. Enterprise analytics report cost by group and by user, showing artifacts, files and tools next to their cost, filtered by your existing org chart. Two APIs expose it.
Bundled in the licence
Cost lives inside the M365 licence, so admin reporting is about adoption, not spend: active users, prompts, app usage, agent uptake — via the admin centre, Viva Insights and Graph API. Great for rollout tracking, quiet on per-task cost.
Seats with a console
Business and Enterprise tiers arrived in January 2026 — Business at $30 per seat, Enterprise on custom pricing. Management runs through the xAI console: seats, consolidated billing, role-based controls and usage analytics, with Enterprise adding SSO, SCIM directory sync, audit tooling and a customer-key data vault. Cost sits inside the seat, so reporting leans toward adoption rather than per-task spend.
Gemini and other Workspace-bundled assistants follow broadly the licence-based pattern described for Copilot — cost inside the licence, reporting focused on adoption — but aren't broken out separately here. Check current admin documentation before relying on any of this for procurement.
What today's prices are really telling you
Everything above treats AI pricing as a fact of life. It isn't — it's a moment in a market that hasn't settled. Understanding what providers charge versus what it costs them to serve you explains why prices move the way they do, and why building the habits on this page now is worth more than waiting for costs to fall.
The providers are not all making money
Anthropic — margin recovered fast
Gross margin moved from roughly −94% in 2024 to about 60% in 2026, driven by inference efficiency rather than price rises — blended token prices have been falling throughout. Revenue per megawatt of compute rose from about $16m to a projected $60m in under a year. Enterprise API mix is the reason: business customers generate several times more revenue per token than consumer users, and their query patterns are cheaper to serve.
OpenAI — scale without margin
Gross margin was around 33% in 2025, improving to roughly 39% in early 2026 — against a $20.9bn operating loss on $13.07bn of revenue. The drag is structural: a very large free consumer tier consuming inference without matching revenue. Internal plans point to cumulative losses near $115bn through 2029 before cash-flow profitability around 2029–30.
Cheaper per query, bigger bills
Inference cost per query has fallen roughly 95% since early 2023. Total spend still rose, because agentic workloads consume many times more tokens per task than a chat message. Whether margins expand depends on whether costs fall faster than demand grows — and so far demand has won.
Figures from press reporting and analyst estimates during 2026, in US dollars. These are private companies; numbers are leaked, estimated or self-reported, and both firms have been reported as missing their own margin targets. Treat as direction of travel, not audited fact.
The margin ladder is being squashed
Classic software earned its valuations on near-zero marginal cost: build once, serve the next customer for almost nothing. AI breaks that in exactly one place — every use spends real compute, so inference lands in cost of goods sold rather than fixed R&D. The result is a visible compression.
Gross margin by business type
Median software gross margin across 342 SaaS and AI-native companies was 80% for full-year 2025, with top-quartile at 86% and usage-only pricing at 62%. ICONIQ's survey of roughly 300 software executives puts average AI product gross margin at 41% in 2024, 45% in 2025 and 52% in 2026. Most analysts expect a structural floor around 60–65% rather than a return to 80%+.
Compression, not collapse
AI margins are improving — 41% to 52% in two years — but the consensus is they settle in the 60–65% band rather than returning to the SaaS 80%. Traditional software adding AI features typically gives up 12–17 points. The old 80% benchmark is unlikely to come back as the industry standard.
What that means for buyers
Vendors facing structurally lower margins have to recover cost somewhere: usage-based pricing, credit systems, tiering and caps. That's precisely the shift described earlier on this page. Expect more metering, not less — and expect the meter to be passed to you.
Why waiting is not a strategy
Per-token prices have fallen sharply and will likely keep falling. Total spend still went up. If cheaper tokens simply invite more use — longer threads, more agent calls, more re-runs — then discipline, not price, is what determines your bill. Efficiency compounds in a way that waiting does not.
Where the value settles: data, not models
Software went from "tech companies" as a single category to fintech, regtech, proptech, healthtech and the rest. AI looks likely to segment the same way — and the dividing line is emerging as proprietary data and workflow depth rather than model access.
Model access is not a moat
Everyone can reach the same frontier models on similar terms, so capability alone stopped being differentiating. Investor language has shifted markedly during 2026 away from thin wrappers and toward workflow ownership, proprietary data and domain depth. One investor went as far as saying the terminal value of software itself is now being questioned.
What defends a position
A Stanford CodeX white paper ranks vertical AI moats in ascending strength: workflow and interface, vertical tooling, built-in compliance, and strongest of all the accumulated data layer. Regulated verticals — legal, healthcare, financial services — score highest, because domain knowledge and compliance are slow to replicate whatever the model can do.
The compounding loop
Use generates data, data improves the product, the better product drives more use. Sectors where the data is both proprietary and directly useful to model performance benefit most. The honest caveat: most software companies discover on inspection that they don't have a data moat — they have a head start, and head starts erode.
This last section is interpretation rather than instruction. The figures are sourced, but the conclusion — that margins converge toward a common band and AI segments by data moat rather than by model — is a reasonable reading of where things are heading, not a certainty. The practical takeaway is narrower and safer: pricing is unsettled, metering is increasing, and the habits described above are what you actually control.