LAYER 03 / 06
Layer 03 — Data Centers & Compute

Is the margin worth the risk?

A modern AI data center is a building-sized supercomputer, and the industry is spending at a pace with few historical parallels to build more of them. This layer takes the biggest slice of the AI dollar — and carries the biggest balance-sheet risk in the stack to earn it.

The 101

What this layer is

The data-center layer operates the facilities and rents out the compute. Hyperscalers and neoclouds sell it by the GPU-hour; colocation providers rent powered, cooled space to everyone else; token-serving inference clouds host open-weights models on their own machines and sell the output by the token. Owning the raw land and power is the layer below; making the chips is the layer beside it; the weights served on all this hardware come from the layer above.

The core economics of the layer come down to one variable — utilization:

"A full data center is a money machine; an empty one is a depreciating liability with a power bill."

Everything on this page — the capex debate, the credit scare, the repricing of GPU rentals, the margin arriving from the model layer — is a fight over which side of that sentence the industry lands on.

The defining number

$ per GPU-hour — the real-time glut gauge

Rental prices are the single best live reading on whether this build-out is ahead of demand or behind it.

GPU rental prices, 2024–2026
$ per GPU-hour · candidate metric for this layer
Chart ships when the dataset lands
Candidate metric: GPU rental price, $ per GPU-hour
Likely sources: SemiAnalysis H100 rental-contract index · WSJ Blackwell datapoints

No defensible continuous series exists publicly yet — scattered verified datapoints are quoted in the deep dive with attribution rather than blended into a curve here.

Falling prices say a glut is forming; rising prices say the shortage is real and the pricing power sits here. A clean, continuous public price series does not yet exist — indexes, contract tiers, and chip generations each tell a piece of it — so this chart ships when a dataset worth putting under a byline lands. The verified datapoints live in the deep dive below.

Follow the dollar

Where $1.00 of AI application spend lands

$0.14
Apps + Agents
$0.10
Harness
$0.18
Models
$0.30
Data Centers + Compute
$0.22
Silicon
$0.06
Land + Power + Shell

Data Centers + Compute takes roughly $0.30 — the largest slice on the strip, and the most expensive slice to earn. Every cent of it is backed by capex spent years in advance, financed against demand forecasts, and depreciating whether or not the racks stay full.

Illustrative split — assumptions in Methodology.

Deep dive

Capex, credit, and the margin moving in

Ahead of demand, or behind it

The central debate of this layer: is the industry building ahead of demand (glut — the bubble fear) or behind it (shortage — pricing power)? This is the most important open question in AI investing, and the honest position is to hold both scenarios. The scale of the bet keeps growing either way: Morgan Stanley's revised estimate puts Big-5 hyperscaler capex at roughly $805B for 2026, and on top of that sit an estimated $662B of signed-but-uncommenced data-center leases — off balance sheet until the facilities start, per FactSet and lease-analysis coverage.

The striking part of the mid-2026 data is that capex and cloud growth are accelerating together. Google Cloud grew 82% year over year in Q2 2026 — on $44.9B of quarterly capex and negative free cash flow. AWS grew 37%, its fastest since 2021. Azure grew 43% while crossing $100B a year. And on July 29, 2026, Microsoft guided FY2027 capex to $255–260B, the largest forward number any hyperscaler has given. Whatever this is, it is not spend chasing a demand curve that already rolled over.

The credit side against the cash side

The July 2026 selloff was the credit side of the risk showing itself: real yields rose, spreads widened, Meta's $12B bond at 7.5% met weak demand, Google surprised the market with a $25B bond, and credit-default swaps hit records across the hyperscalers — the four-narrative framing of that month is Gavin Baker's, from his August 4 Invest Like the Best appearance.

The cash side pushed back. Combined operating cash flow growth at Microsoft, Meta, and Amazon accelerated from 28% to 32% year over year in Q2 2026 — about 35% stripping one-time items like EU fines. And per Baker:

"No one in Silicon Valley reported having too many GPUs."

Baker's self-funding model goes further: consensus models still price the installed base at Ampere-era compute rates; repriced to Blackwell spot, his math gets to roughly $2T of operating cash flow and about $700B less credit demand — his model, worth attributing as such. His bear case is the mirror image: cash flow stalls, the build-out goes debt-dependent, and the layer replays the classic capital cycle that undid the internet bubble.

Compute is repricing up

The rental market has swung from glut-pricing to shortage-pricing inside a year. SemiAnalysis puts the trough of its H100 one-year rental-contract index at October 2025, at $1.70 per GPU-hour, up roughly 40% by March 2026 — Dwarkesh Patel's "February trough" framing is his own dating of the same recovery. On the Blackwell generation, WSJ-cited figures show hourly rentals going from $2.75 to $4.08 in about two months (February to April 2026), and tracker surveys put B200 rates moving from $3.81 to $5.64 during July 2026. Baker adds the anecdote of a startup whose B200 cluster went from the mid-$2 range to just under $4 on re-contract, and an inference cloud planning to pay roughly 100% more on renewal — his read: hyperscalers have been systematically under-earning their fleets. At the top of the market, Google is reportedly paying $900M a month for 110K GPUs from SpaceX — about twice spot, a premium for scale, efficiency, and security that spot capacity can't provide.

Depreciation and the LTA trap

Two structural risks cut against the repricing story. GPUs may obsolete faster than accounting schedules assume, which would flatter every reported return in the layer. And long-term agreements carry a trap Baker highlights: contracts were signed expecting price declines, prices rose instead, and breaking an LTA is potentially company-ending — the counterparty controls your future supply allocation.

The regulatory skew

Chamath Palihapitiya's read adds a third risk axis: clouds are very lucrative but very hard to build and very expensive and technically complicated to maintain — and as alignment becomes more important, clouds will be asked to build robust KYC and attest to it. His conclusion is that the risk:reward is skewed; he does not want to be the responsible party when a government says a cloud allowed a bad actor to do something bad. The margin question on this page has a compliance-liability term in it, and it is growing.

The margin transfer — this page's main event

The GLM-5.2 and Kimi K3 releases in mid-2026 shifted the measured token mix — Silicon Data's LLM Token Expenditure Index — away from frontier tokens and toward open-source tokens. Baker's reading: this is a mix shift, not a demand shift, and it moves margin. Frontier tokens carry margins he puts at 80, 90, or 95 percent; open-source tokens run nearer 30. The same flops and the same watts produce each token — so as the mix moves to open weights, the margin that used to sit in the model layer transfers down into this one, to whoever operates the infrastructure serving those tokens.

The visible vehicles of that transfer are the token-serving inference clouds — the Together AI and Fireworks tier, hosting open weights on their own machines and posting what Baker describes as crazy Rule-of-40 numbers. They are the cleanest expression of this page's question: infrastructure margin, captured at software-company economics, for as long as the mix shift runs. The Models page covers what this same shift does to the layer that trains the weights.

Key players

Who's in this layer — and how they differ

HyperscalersAWS · Azure · Google Cloud

Rent compute by the GPU-hour at global scale, financed by their own cash flows. Their reach extends up-stack — each also trains frontier models — but their home in the stack is here, operating the machines.

CoreWeaveNeocloud

The NVIDIA-allocated GPU specialist — pure-play compute rental, built and financed at venture speed against long-term contracts.

OracleContracted compute at scale

The late entrant converting a legacy franchise into GPU capacity — a record backlog of prepaid AI contracts against heavily debt-financed build-out.

Colocation tierEquinix · Digital Realty

Rent the powered, cooled space itself — the landlords of the layer, earning steadier margins one step removed from GPU price risk.

Token-serving inference cloudsTogether AI · Fireworks

Host open-weights models on infrastructure they operate and sell by the token — the direct beneficiaries of the open-source mix shift moving margin into this layer.

Confidential computeNEAR AI Cloud

Serves open models inside hardware trust boundaries (TEEs), selling inference whose privacy is attested rather than promised — the infrastructure-layer answer to the question the Harness page's privacy membrane asks from above.

My read

This is the layer where I hold two opposing thoughts on purpose. Every demand signal is accelerating and compute keeps repricing upward, which says the build-out is running behind demand. The financing tape — record CDS, debt-funded shells — says one air pocket turns money machines into depreciating liabilities with power bills. I watch a single gap: cloud AI revenue growth against capex growth. While they rise together, the margin is worth it. The quarter they diverge, this layer reprices first.

— Conner Murphy