LAYER 04 / 06
Layer 04 — Models: Training & Inference

Is today's model revenue real?

Token consumption is inflecting upward while the price story split in two: flagship list prices turned back up in 2026 even as last year's models and open weights kept collapsing toward free. The revenue caught between those curves is the most contested number in the stack — and how much of it survives decides where margin lives for the next decade.

The 101

What this layer is

The model layer creates the weights — the frontier labs and open-weight makers that train the intelligence everything above them runs on. Training happens once, at historic capital cost; revenue arrives one served token at a time, so serving economics are this layer's revenue story. Companies that host and serve other people's weights sit a layer down, in Data Centers & Compute — this grid is reserved for the makers.

Two forces squeeze the makers at once. From below, open-weights models keep resetting the price floor for good-enough intelligence. From above, harnesses make models swappable components, holding switching costs near zero. Whatever margin regime the stack lands in, this layer is the hinge.

The defining number

$ per million tokens — the collapse, then the split

Three years of collapsing list prices, then the 2026 fork: flagships repriced upward while last year's models kept deflating.

Price of served intelligence, 2023–2026
$ per million output tokens · published list prices · log scale
$1 / M TOKENS $100 $10 $1 $0.10 2023 2024 2025 2026 FLAGSHIP — TURNS BACK UP LEGACY + OPEN WEIGHTS — STILL FALLING

Hand-curated from published list prices — OpenAI (via independent trackers costgoat.com, pricepertoken.com, aipricing.guru), Anthropic and Google official pricing pages — committed with full citations as data/models-prices.json. List price is not paid price: caching, batch, and enterprise discounts cut effective rates far below these lines. Anthropic notes Claude 4.7+ tokenizers emit ~30% more tokens for the same text, so a flat sticker price there hides an effective increase.

For three years the price of frontier intelligence only fell: GPT-4 launched at $60 per million output tokens in 2023; GPT-5 launched at $10 in August 2025 ($1.25 in / $10 out). Then the curve forked. OpenAI's flagship climbed back to $5 / $30 within roughly eleven months — a 4x input-price reversal — while Anthropic added a $10 / $50 Fable tier above Opus and Google priced Gemini 3 Pro above its predecessor. Meanwhile the trailing edge kept deflating on schedule: legacy tiers took 50–80% cuts and open-weight flagships like GLM-5.2 undercut everything. The frontier stopped deflating; the mid-tier never did.

Both altitudes are true at once. At the unit level, the price of last year's intelligence keeps collapsing — that is the rust line. At the aggregate level, compute is repricing upward: Dwarkesh Patel's July 29, 2026 essay argues lab revenue is growing roughly 10x year over year against compute capacity growing only ~3x, a gap that closes through margins, prices, or reallocation — and the 2026 flagship repricing is what that closing looks like. Cheap tokens and expensive intelligence are the same market, read at different altitudes.

Follow the dollar

Where $1.00 of AI application spend lands

$0.14
Apps + Agents
$0.10
Harness
$0.18
Models
$0.30
Data Centers + Compute
$0.22
Silicon
$0.06
Land + Power + Shell

Models bill roughly $0.18 of the dollar — and most of it flows straight through to the compute layer below, which is why lab margins are the stack's swing variable. The open questions on this page are how much of that $0.18 was billed for work that mattered, and how much of it the layer keeps.

Illustrative split — assumptions in Methodology.

Deep dive

Tokenmaxxing, and the squeeze from both sides

The tokenmaxxing question

Some share of today's token consumption exists because models behave badly: retries, verbose outputs, agents burning tokens in failure loops. That usage books as revenue. Chamath Palihapitiya poses it as the open question of the layer:

"The big open question is how much of the revenue being generated by them today is because of tokenmaxxing and poor model behavior. If it's a lot, then the annualized revenues will diminish meaningfully even as token consumption inflects upwards."

No independent quantification of that share exists yet — it stands as an attributed open question, not a finding. But the mechanism is real: better models mean fewer wasted tokens per unit of work done, which would make this the first layer in history whose revenue partially shrinks as its product improves.

The squeeze from below

Open-weight models — Zhipu's GLM, Moonshot's Kimi, DeepSeek, Qwen — now sit close enough to the frontier that GLM-5.2 serves at $4.40 per million output tokens against comparable closed mid-tier models at $10–12. Chamath argues they are cheap enough that they threaten the U.S. funding capex cycle itself. They are the durable price floor: whatever a frontier lab charges must justify the gap to nearly-free.

The squeeze from above

Harnesses hold enterprise alpha — data, workflows, evals, business rules — outside the model, which keeps switching costs low and model-agnostic. The model underneath becomes a component: swapped for cost, for reasoning, or for policy. A component vendor prices like a component vendor — unless, like Anthropic with Claude Code, the lab climbs into the harness itself.

Revenue quality, contested live

The July 2026 selloff took AI names down 40–60% in a month — yet Gavin Baker's reading of the same tape found no negative quantitative demand metric anywhere: GPU availability, rental pricing, DRAM, and token growth all accelerating. Dwarkesh's essay makes the bull case structural: lab revenue growing ~10x year over year against ~3x compute growth, with blended inference margins moving from roughly 40% in 2025 to over 70% now per SemiAnalysis — and 80%+ claimed, though not verified, for newest-model inference. The error bars stay wide: Anthropic's $100–150B revenue figure for this year is Dwarkesh's extrapolation, against a tracked run-rate near $69–74B in late July. The GLM-5.2 and Kimi K3 releases also shifted measured token mix from high-margin frontier tokens toward low-margin open-source tokens — Baker reads that as margin migrating from the model layer into infrastructure, a story that lives on the Data Centers & Compute page.

Key players

Who's in this layer — and how they differ

OpenAIFrontier lab

Consumer distribution at unmatched scale — the default brand for intelligence, and the lab that led the 2026 flagship repricing.

AnthropicFrontier lab

Enterprise trust and the agentic-coding franchise; its Claude Code harness reaches a layer up, keeping switching costs on its own terms.

Google DeepMindFrontier lab

The full-stack player: own silicon (TPUs), own cloud, own distribution.

MetaOpen weights at scale

Commoditize-the-complement as strategy — frontier-scale weights given away to reset everyone else's price floor.

xAIFrontier lab

Vertically integrated compute buildout — betting speed of scaling beats everything.

Open-weight makersDeepSeek · Qwen/Alibaba · Zhipu GLM · Moonshot Kimi

Frontier-adjacent quality at collapsing cost — the price floor for the whole layer, and the force shifting token mix away from frontier margins.

My read

Some of it, and the market is currently finding out which part. Capability-adjusted prices keep collapsing while frontier stickers quadrupled in eleven months — both are true, and the spread between them is the whole margin question in one number. My open question is how much of today's token consumption is usage inflated by verbose model behavior. Weights alone look like the weakest asset in the stack to me; distribution, cost-to-serve, and the context accumulated above the model decide who keeps the revenue.

— Conner Murphy