The landscape

What your admin console can and cannot actually see, how the vendors differ, and what provider economics suggest about where pricing goes next.

REALITY CHECK

What you can — and can't — actually track

This is the question every finance and IT team asks, and the honest answer is "it depends what you bought." A subscription hides the meter; an API exposes it. Here's what each setup really lets you see, as of mid-2026.

What you bought Per-user visibility Per-token / true cost By model & tool Programmatic API
ChatGPT Team
seat-based
Aggregate Hidden Limited No
ChatGPT Enterprise
credit-based, 2026 console
Named users Credits Yes Cost API
Claude Team
pooled tokens
Aggregate Pool + caps Some No
Claude Enterprise
Analytics API, 2026
Named users $ per user Model + tool Analytics API
Microsoft 365 Copilot
licence-based
Adoption only Hidden in licence App adoption Graph API
Grok Business
seat-based, $30/user/mo
Console analytics Hidden in seat Limited No
Grok Enterprise
custom pricing, SSO + SCIM
Named users Hidden in contract Limited Audit tools
Direct API
OpenAI / Anthropic / etc.
By key/tag Exact tokens Full detail Full API

Subscriptions hide the meter

Team and licence tiers mostly show adoption — active users, prompt counts, which apps — not what any single task cost. Good for "is anyone using this?", weak for "where is spend going?"

Enterprise adds the dashboard

Through 2026, all three vendors shipped admin consoles with per-user, per-model breakdowns and spend caps. ChatGPT bills Enterprise in credits; Claude Enterprise reports actual dollars per user. Both expose an API.

APIs expose everything

Direct API usage is the only place with true per-token, per-model, per-call cost — down to input, output and cache. Route calls through one gateway and tag by team, or the "one line called OpenAI" hides four different things.

The shadow-AI gap

None of this sees personal accounts. Analysts estimate most AI use in large firms happens on unsanctioned tools. A dashboard only covers what's on the corporate licence — the rest is invisible.

Vendor tooling in this area is changing month to month — treat this as a mid-2026 snapshot and confirm against current admin docs before a procurement decision.

SIDE BY SIDE

How the three stacks meter cost

Same goal — contain and understand spend — three different mental models. Knowing which one you're in tells you how to budget and what to watch.

ChatGPT

Seats, then credits

Team buys seats. Enterprise moved to a credit model with a 2026 Global Admin Console: slice consumption by user, group, project and model, set workspace defaults, cap power users, pull it via a Cost API. Office and agent features now meter in tokens too.

Claude

A utility meter

No seat limits — usage draws from one org token pool, with spend caps at every level. Enterprise analytics report cost by group and by user, showing artifacts, files and tools next to their cost, filtered by your existing org chart. Two APIs expose it.

Copilot

Bundled in the licence

Cost lives inside the M365 licence, so admin reporting is about adoption, not spend: active users, prompts, app usage, agent uptake — via the admin centre, Viva Insights and Graph API. Great for rollout tracking, quiet on per-task cost.

Grok

Seats with a console

Business and Enterprise tiers arrived in January 2026 — Business at $30 per seat, Enterprise on custom pricing. Management runs through the xAI console: seats, consolidated billing, role-based controls and usage analytics, with Enterprise adding SSO, SCIM directory sync, audit tooling and a customer-key data vault. Cost sits inside the seat, so reporting leans toward adoption rather than per-task spend.

Gemini and other Workspace-bundled assistants follow broadly the licence-based pattern described for Copilot — cost inside the licence, reporting focused on adoption — but aren't broken out separately here. Check current admin documentation before relying on any of this for procurement.

ZOOMING OUT

What today's prices are really telling you

Everything above treats AI pricing as a fact of life. It isn't — it's a moment in a market that hasn't settled. Understanding what providers charge versus what it costs them to serve you explains why prices move the way they do, and why building the habits on this page now is worth more than waiting for costs to fall.

The providers are not all making money

Anthropic — margin recovered fast

Gross margin moved from roughly −94% in 2024 to about 60% in 2026, driven by inference efficiency rather than price rises — blended token prices have been falling throughout. Revenue per megawatt of compute rose from about $16m to a projected $60m in under a year. Enterprise API mix is the reason: business customers generate several times more revenue per token than consumer users, and their query patterns are cheaper to serve.

OpenAI — scale without margin

Gross margin was around 33% in 2025, improving to roughly 39% in early 2026 — against a $20.9bn operating loss on $13.07bn of revenue. The drag is structural: a very large free consumer tier consuming inference without matching revenue. Internal plans point to cumulative losses near $115bn through 2029 before cash-flow profitability around 2029–30.

Cheaper per query, bigger bills

Inference cost per query has fallen roughly 95% since early 2023. Total spend still rose, because agentic workloads consume many times more tokens per task than a chat message. Whether margins expand depends on whether costs fall faster than demand grows — and so far demand has won.

Figures from press reporting and analyst estimates during 2026, in US dollars. These are private companies; numbers are leaked, estimated or self-reported, and both firms have been reported as missing their own margin targets. Treat as direction of travel, not audited fact.

The margin ladder is being squashed

Classic software earned its valuations on near-zero marginal cost: build once, serve the next customer for almost nothing. AI breaks that in exactly one place — every use spends real compute, so inference lands in cost of goods sold rather than fixed R&D. The result is a visible compression.

Gross margin by business type

Pure SaaS
80%
Top-quartile SaaS
86%
AI-augmented
~70%
Usage-priced
62%
AI-native (2026)
52%
AI-native (2024)
41%

Median software gross margin across 342 SaaS and AI-native companies was 80% for full-year 2025, with top-quartile at 86% and usage-only pricing at 62%. ICONIQ's survey of roughly 300 software executives puts average AI product gross margin at 41% in 2024, 45% in 2025 and 52% in 2026. Most analysts expect a structural floor around 60–65% rather than a return to 80%+.

Compression, not collapse

AI margins are improving — 41% to 52% in two years — but the consensus is they settle in the 60–65% band rather than returning to the SaaS 80%. Traditional software adding AI features typically gives up 12–17 points. The old 80% benchmark is unlikely to come back as the industry standard.

What that means for buyers

Vendors facing structurally lower margins have to recover cost somewhere: usage-based pricing, credit systems, tiering and caps. That's precisely the shift described earlier on this page. Expect more metering, not less — and expect the meter to be passed to you.

Why waiting is not a strategy

Per-token prices have fallen sharply and will likely keep falling. Total spend still went up. If cheaper tokens simply invite more use — longer threads, more agent calls, more re-runs — then discipline, not price, is what determines your bill. Efficiency compounds in a way that waiting does not.

Where the value settles: data, not models

Software went from "tech companies" as a single category to fintech, regtech, proptech, healthtech and the rest. AI looks likely to segment the same way — and the dividing line is emerging as proprietary data and workflow depth rather than model access.

Model access is not a moat

Everyone can reach the same frontier models on similar terms, so capability alone stopped being differentiating. Investor language has shifted markedly during 2026 away from thin wrappers and toward workflow ownership, proprietary data and domain depth. One investor went as far as saying the terminal value of software itself is now being questioned.

What defends a position

A Stanford CodeX white paper ranks vertical AI moats in ascending strength: workflow and interface, vertical tooling, built-in compliance, and strongest of all the accumulated data layer. Regulated verticals — legal, healthcare, financial services — score highest, because domain knowledge and compliance are slow to replicate whatever the model can do.

The compounding loop

Use generates data, data improves the product, the better product drives more use. Sectors where the data is both proprietary and directly useful to model performance benefit most. The honest caveat: most software companies discover on inspection that they don't have a data moat — they have a head start, and head starts erode.

A view, not a fact

This last section is interpretation rather than instruction. The figures are sourced, but the conclusion — that margins converge toward a common band and AI segments by data moat rather than by model — is a reasonable reading of where things are heading, not a certainty. The practical takeaway is narrower and safer: pricing is unsettled, metering is increasing, and the habits described above are what you actually control.