The Reticle Limit
The marketed 30x generational leap in AI chips collapses to about 18x once you strip out a one-time accounting trick, and the real hardware improvement underneath both numbers is 47 percent. Not a typo.
We're gonna get really geeky with this. I started my life measuring linewidths on reticles, the negatives that print patterns on silicon wafers that power – well everything.
I will not be offended if you skip this one, honestly
Table of Contents
Nvidia's newest gaming chip, the one powering the RTX 5090, is a single piece of silicon roughly 750 square millimeters across, built right up against the edge of what current lithography can physically expose in one shot, a hard ceiling called the reticle limit, somewhere around 850 square millimeters, that no amount of money moves. So when Nvidia needed a bigger brain for the data center, it didn't make a bigger chip. It couldn't. Blackwell's B200 uses a pair of reticle limited compute dies, glued together with a 10 terabyte-per-second connection, because two maximum-size dies was the only way past a wall that a trillion dollars can't buy a way around.
Sit with that for a second, because it's the whole argument in one fact. The last piece in this pairing was about money that gets reshuffled infinitely, debt that vanishes from one balance sheet and reappears somewhere nobody's looking, an architecture built specifically to make scarcity of scrutiny look like abundance of capital. This piece is about the thing underneath all that financial origami that doesn't reshuffle. Silicon has opinions. Physics doesn't take a meeting.
Here's what that physics actually costs, in units you can feel instead of megawatts you can't. A single current-generation AI server rack, the low end of what's shipping right now, draws something like a small residential subdivision's worth of electricity, on the order of twenty-six houses running air conditioning through a Tucson summer, from one rack, in one room. A full-scale campus, the kind Meta is building in Louisiana right now, scales past a million house-equivalents. One company. One site.
Silicon has opinions. Physics doesn't take a meeting.
The Case for Building It Anyway
The stated argument, and it's not a strawman, which is exactly why the numbers in the next section matter.
Give the industry its due before we take it apart. The scaling laws have held, empirically, for a decade: more compute, more data, more parameters, better performance, again and again, and betting against a curve that's held that long is its own kind of risk. Nvidia's own claims for what each generation buys you are, on paper, extraordinary: Blackwell over Hopper is marketed at up to thirty times the inference performance, twenty-five times more performance at the same power draw, and continuous software work has already squeezed another five times lower cost per token out of the same hardware since launch. And there's a real historical counterargument that deserves better than a shrug: the late-1990s fiber buildout looked catastrophically overbuilt right up until it quietly became the backbone the entire internet runs on. Bandwidth nobody could justify in 1999 was dirt cheap and essential by 2005. Maybe GPUs are fiber.
Maybe. Here's what happens when you actually check the marketed numbers against the silicon.
The Reticle Limit
Every headline generational leap has to survive an audit. Most of this one doesn't.
The independent teardown exists, and it's rough. SemiAnalysis, the industry's most trusted analyst shop on exactly this question, pulled apart Nvidia's marketed thirty-times inference gain from Hopper to Blackwell and found that a huge share of it evaporates once you control for the fact that Nvidia was comparing the new chip running at four-bit precision against the old chip running at eight-bit precision, a different numeric format, not the same math done faster. Strip that out and the honest gain is closer to eighteen times, still real, still impressive, not the number in the keynote. Push one step further and ask what the silicon itself improved, independent of any format trick or cherry-picked benchmark: the same analysis puts the iso-power gain, same precision, same watt, at forty-seven percent. Their own words for that number: "hardly anything to write home about." That's the actual generational hardware curve. Not thirty times. Not eighteen. Under one and a half.
Where does the rest of the marketed gain come from, then. Mostly a trick with a floor, not a trend. Going from eight-bit to four-bit numbers roughly doubles raw throughput on paper, because you're pushing half as many bits through the pipe, and you can do that exactly as many times as you have bits left to give up. Blackwell's native four-bit support is most of the "thirty times" story. There isn't a two-bit format waiting in the wings to repeat the trick next generation, not one that keeps a model numerically sane, anyway. You get to cash that particular check once.
The physical reason the free lunch ended over a decade ago is worth naming plainly: Dennard scaling, the old assumption that power density stays roughly flat as transistors shrink, broke down around 2005. Since then, every chip in the industry has had to buy efficiency the hard way, through specialized architecture and precision tricks, not through transistors that magically run cooler as they get smaller. That's the actual, physical reason a GB200 rack now needs something like 120 kilowatts where a general-purpose server rack needs about 12 and even the previous H100 generation topped out air-cooled around 40. Liquid cooling didn't become mandatory because engineers got fancy. It became mandatory because the old way of getting more compute per watt stopped working twenty years ago and nobody's found a physics-level replacement, only better bookkeeping.
And building the thing at all runs through exactly two chokepoints on the planet, neither of them redundant. The dual-die assembly and advanced packaging happens almost entirely inside Taiwan, at TSMC's own fabs, a process serious enough that Nvidia had to redesign part of the chip's own metal routing after early yields on the packaging step came in badly enough to be described, without much hedging, as a yield-killing defect. The high-bandwidth memory bonded into that same package, the component doing most of the heavy lifting and, not coincidentally, the component most prone to failure under sustained load, comes overwhelmingly from South Korea, SK Hynix and Samsung, both. Every dollar of the trillions from the last piece in this pairing eventually has to clear those two countries' fabs. There is no plan B fab sitting somewhere boring and geopolitically stable, waiting to pick up slack. There's Taiwan, there's Korea, and there's hoping nothing happens to either.
Sweaty's Corner doesn't have a marketing budget or an algorithm pushing it into anyone's feed. It has readers who share it. If this one landed, sending it to someone who'd get something out of it — or subscribing so the next one finds you directly — genuinely moves the needle.
What One Rack Actually Costs
Most people can't tell a Slack server from an AI campus, which is a problem, because they're not the same animal.
A traditional enterprise data center, the kind quietly running Salesforce or Teams somewhere nobody's ever thought about, typically pulls under 100 megawatts total and moves in slow, predictable, daily rhythms, the kind of load a grid operator can plan around without losing sleep. An AI campus is planned in gigawatts, ten to a hundred times larger on the same "data center" label, and it doesn't behave like a big version of the boring kind. Grid engineers describe AI training loads as synchronized, high-amplitude swings, capable of dumping or demanding over a thousand megawatts almost instantaneously as a training run starts or stalls, the kind of volatility that genuinely worries the people whose job is keeping the lights on for everyone else, not just the campus itself.
A lot of these sites don't even wait for the grid. Meta is putting a 200-megawatt natural gas plant directly on-site at one facility and three more gas plants, three billion dollars worth, at another, specifically because the interconnection queue can't move fast enough. A neighbor to a Slack data center gets a quiet building. A neighbor to an AI campus potentially gets a merchant power plant next door, running around the clock, plus the water draw that comes with cooling something this dense.
None of this is theoretical anymore, and it's stopped being a niche complaint. Data center backlash is showing up as a live variable in real 2026 races, crossing party lines in a way most issues this cycle don't, which tells you something about how visceral the actual experience of living next to one of these things has become, regardless of which side of the aisle you're standing on when you complain about it.
Nobody's Lived Through the Warranty Period
The honest answer to "how long does this hardware actually last" is that nobody knows yet, and the accounting is proceeding as if somebody does.
The best real reliability data available comes from Meta's own disclosed training runs: a sixteen-thousand-plus GPU cluster logged over four hundred unexpected hardware failures across fifty-four days, which independent analysts have used to back into a real mean-time-between-failure figure, somewhere around five and a half to six years per GPU. That's a genuine number, not marketing, and it's a reasonably solid one. High-bandwidth memory shows up as the single largest or second-largest category of hardware failure across multiple independently reported clusters, which should sound familiar: it's the same component that took the brunt of the wear during the crypto mining boom, a hardware generation ago, on a completely different chip, running a completely different workload. Sustained load finds the same weak point again and again, apparently regardless of what you're actually computing.
Here's the honest limit of what that number can tell you, though. H100s only reached real volume in 2023, which means the oldest fleets in existence right now are about three years old. Nobody has five-year field data on a chip generation that's only existed for three. The industry's own accounting assumes something like a six-year useful life for this hardware. That assumption has never once been tested against an actual six-year-old fleet, because one doesn't exist yet. The number everyone's building financial models on is an extrapolation dressed up as a fact.
The last piece in this pairing described money that gets financed against a promise, moved somewhere it won't get scrutinized, backed by revenue that's often another company's promise right back. This piece is the other half of the same bet: hardware whose true multi-year degradation curve nobody has actually observed yet, whose marketed generational improvement is substantially an accounting trick with a floor rather than a repeatable trend, manufactured almost entirely through two countries' worth of infrastructure that has nothing to do with artificial intelligence and everything to do with ordinary seismic risk and ordinary geopolitics. The debt can be restructured. The revenue assumptions can be renegotiated. The reticle limit cannot. Somewhere underneath several trillion dollars of financial engineering is a physical object that gets hot, needs water, and doesn't care what anybody's balance sheet says about it.
Like what you've read, but not able to switch to paid? Why not buy me a coffee instead! All money goes to help defray the expense of running the site.