ONLINEAGENT_OPS 2026.Q3 HOMEARTICLESBLOGRECORDCRAFTSEARCH
HOMEARTICLESWhat Ai Runs On
ARTICLES · EVERGREEN EXPLAINER

What AI Actually Runs On: sand, power, and heat

Descend from silicon die to power grid: why GPUs, what memory bandwidth means, how a training run works, and the…

⌘ MAP · NEW STATION · MEASURED JUL 2026 · MACHINE-CREATED, HUMAN-ORCHESTRATED (DISCLOSED) · SOURCES VERIFIED AT PUBLICATION
THE LATTICE
Sand, refined, arranged, and made to think. This is the physical part.

The Assembly upstairs shows what a model is. This station descends through what it is made of: sand, gold, water, and roughly a nation's worth of electricity. Five levels down. Mind the heat.

TL;DR — AI runs on chips built for one trick — multiplying enormous grids of numbers in parallel — stacked into racks so dense a single cabinet can draw what dozens of homes do, cooled by liquid, gathered into halls that rival cities, plugged into grids now visibly straining. The numbers below are not vibes; each carries its source and date, because this page will age and says so.
LEVEL 1 · THE DIE

Why this chip and not your laptop's

Macro photograph of a glowing green GPU silicon die like a city seen from orbit, gold traces as streets
THE DIE — a city of arithmetic, photographed from orbit.

A CPU is a brilliant generalist: a few powerful cores doing complicated things one after another. A neural network needs the opposite — one simple thing (multiply, add) done billions of times simultaneously, because a model's "thinking" is just enormous grids of numbers being multiplied together. GPUs, born for videogame pixels, happen to be exactly that: thousands of small cores in parallel. That accident of history is why one company's green silicon became the picks-and-shovels of the entire gold rush.

The die's partner is memory bandwidth — how fast weights can be fed to those cores. Modern accelerators stack high-bandwidth memory physically next to the die because the bottleneck is usually not computing the numbers; it is carrying them. When the syllabus tells you VRAM decides what runs on your machine, this level is why.

LEVEL 2 · THE RACK

Density is the whole story

Accelerators are ganged eight-plus to a board, boards stacked into racks, racks laced together with interconnect fabric fast enough that thousands of chips can behave like one giant machine — a training run is exactly that: one calculation smeared across a warehouse for weeks. The consequence is heat density with no precedent in computing: the IEA measured AI rack power density rising 11-fold from 2020 to 2025, projecting a further fourfold jump by 2027 — at which point a single refrigerator-sized rack draws power equivalent to about 65 households. Air cannot carry that heat away anymore; liquid cooling — plates and pipes against the silicon — became standard equipment, not exotica.

LEVEL 3 · THE HALL

Where the pour happens

The pour, in the only colour this museum keeps. Silicon starts as sand and ends as the floor everything else stands on.
Endless data center corridor of monolithic black server racks with emerald status lights, cold mist, one human for scale
THE HALL — one human, for scale. The hum is the sound of weights setting.

Gather the racks and you get the hall: a large data centre can consume as much electricity as 100,000 households — the IEA’s own phrasing — and the agency states that the largest currently under construction could consume as much as 2 million households. Training happens here (the weeks-long pour The Assembly describes), and then inference happens here forever after — every chat, every image, every agent, each sip small on its own: the IEA's 2026 estimates put a standard assistant request around 1.1 Wh and a heavy reasoning-agent request near 50 Wh, while a disclosed median short text prompt measured 0.24 Wh and 0.26 ml of water (Google technical paper, 2025). Small sips; unimaginable frequency.

LEVEL 4 · WATER & HEAT

The part everyone forgets is thermodynamics

Industrial cooling towers venting steam at night behind a data center under green floodlights
EXHAUST — computation leaves as weather.

Every watt in becomes heat out; the only questions are how efficiently (the industry metric PUE — total facility power divided by computing power — with modern halls pushing toward 1.1) and what carries it away. Often: water, through evaporative cooling — which is why hyperscaler water disclosures rose sharply through 2024–25 and why closed-loop liquid designs, which cut direct water use dramatically at higher capital cost, are spreading. Cooling is now its own line on the world's bill: Gartner forecasts cooling electricity alone climbing 22.6% in 2026 to 195 TWh.

◈ VERIFICATION STATUS · 4 AUGUST 2026

Re-verified against the issuing organisation this month: the IEA's projection of ~945 TWh of datacentre electricity demand by 2030 (April 2025 Energy and AI report; the agency has since updated its central projection to about 950 TWh, from 485 TWh in 2025), and the comparison that a large datacentre can consume as much electricity as 100,000 households, with the largest under construction reaching 2 million. The remaining figures on this page — rack power density, the 2026 cooling forecast, and the year-on-year growth rates — carry their original source and date but were not re-checked in this cycle. They are marked here rather than quietly presented as equally fresh.

LEVEL 5 · THE GRID

The numbers, dated and sourced

+50%growth of AI-focused data-center electricity in 2025 — 16× the pace of overall electricity demand. IEA, Apr 2026.
565 TWhglobal data-center electricity forecast for 2026, up 26% year-on-year; AI servers ≈ 31% of it. Gartner, Jun 2026.
~950 TWhprojected global data-center demand by 2030 — near 3% of world electricity. IEA base case.
7–12%share of US electricity that data centers could reach by 2028, from ~4.4% in 2023. Lawrence Berkeley National Laboratory.
26% · 21%share of electricity already consumed by data centers in Virginia and in Ireland respectively — concentration, not averages, is where strain lives. IEA / Carbon Brief, 2025.
2027the crossover year: AI-optimized servers projected to out-consume all conventional servers. Gartner, Jun 2026.

Two honest framings, held together: an individual query costs less electricity than boiling water for tea — and the buildout in aggregate is the largest new load the grid has met in a generation, concentrated in a handful of regions whose bills and build-outs are already reshaped by it. Both sentences are true. Any page that gives you only one of them is selling something.

Ascend

You have seen the body. The mind it runs is poured upstairs at THE ASSEMBLY; what it costs the internet is the museum above that; and whether it should live in these halls or on your desk is weighed, without sides, in Local or Cloud.

◈ WHERE THIS SITE STANDS — Infrastructure is not an accusation. These halls also fold proteins, forecast storms, and answer midnight questions. We are messengers: the numbers above are what building a new kind of mind physically costs, shown so people can weigh it with open eyes. Worry is not hostility.
SOURCES (verified July 2026): IEA, "Key Questions on Energy and AI" (Apr 2026): +17% data-center demand 2025, +50% AI-focused, ~945–950 TWh 2030 base case, rack density 11× 2020–25 & 4× by 2027 (≈65 households/rack), hyperscale ≈100,000 households & next-gen ≈20×, per-request 1.1 Wh / ~50 Wh estimates · Gartner forecast (Jun 2026): 565 TWh 2026, AI servers 175 TWh & 31% share, cooling +22.6% to 195 TWh, 2027 crossover · LBNL: US 4.4% (2023) → 6.7–12% (2028) · Google technical paper (Aug 2025): median prompt 0.24 Wh / 0.26 ml / 0.03 gCO2e · Carbon Brief / IEA regional shares (Virginia 26%, Ireland ~21%) · Our World in Data synthesis (Jul 2026). Figures are point-in-time; this station is stamped MEASURED JUL 2026 and will be re-serviced.
WHAT THIS ISThe physical substrate: chips, halls, power.
THE POINTNobody out-thinks a power grid.
THE FACTA large datacentre can draw as much power as 100,000 homes.
READ: ENERGY AND WATER