More on the Data Centers that Will Kill Us All


A lot of this is what I consider being from the Department of Obviousness Department, which I normally try to avoid, but sometimes, the timing is such that I appear to have powers above and beyond those of an ordinary ex-nerd thinking about the future. There was a time, once, when I was arrogant enough to reply – when someone asked me to predict the future – “I make the future, I don’t predict it.” Which was literally true but also JD Vance-level stupid.

Item #1: [arstechnica]

DeepSeek, the Chinese startup developing large language models that are competitive with those from US companies like OpenAI and Anthropic, is planning to enter the silicon business, according to Reuters.

Citing three people familiar with the matter, Reuters writes that DeepSeek has been working on a move into silicon for about a year. It has been meeting with potential partners in the hardware and silicon space and has been hiring engineers for the project.

The focus is on data center chips for inference, not training, and the goal is likely to reduce reliance on both Huawei and Nvidia.

Nvidia is the chipmaker for most AI companies in North America and Europe, but a United States export ban has prevented the company from achieving a similar presence in China. Huawei controls about half of the data center chip market there, and DeepSeek isn’t the only one trying to enter; Chinese tech giants like Alibaba and Baidu have been making moves, too.

Well, yeah. I constantly gnash my teeth in frustration at the way the US sits back, like a dumb and happy techbro, confident in its control of the market. Oh, it’s probably right enough, but every time it pushes the wrong part of the balloon, the other parts squish out uncontrollably. This is not a new move for the US; I first encountered ITAR (international traffic in arms regulation) back when I was working on crypto projects in the early 1990s. The US desired to control the global availability of high quality encryption, by regulating the export of chips that could run parallelized to attack (or build) Feistel-net cyphers like DES. Nobody at NSA or State Department or NIST believed that any of the export regime would slow people who were willing to break the rules but it sure was a pain for law-abiding researchers. I am developing a new principle – which is that government export controls always hit any target except the one they were designed for. (insert F-35 assembly in Turkey stories) In this case, I am mostly rooting for the Chinese, who are now more capitalist than the USA. I hope that they build a chip that can both process large AI trees concurrently, and mine for cryptocurrency. Those are not the same problem but if you’re making custom silicon, why not add a few more circuits; it’s waffer theen <- pronounced as formerly funny John Cleese.

The reason everyone is using Nvidia is twofold: they’re the most multi-core programmable chips on the consumer market, and ‘cuda. ‘Cuda is a programming language for talking to all the Nvidia CPUs on a chip, serializing inputs and outputs, downloading programs etc. ‘Cuda also handles bus-mastering memory and has powerful synchronization locking and queueing, so that a game can asynchronously calculate object trajectories without interrupting the CPU. Other than that, there’s no rocket science that makes AI love Nvidia, and the AI researchers would be perfectly happy with something that ran better and cost less. You can see where this is going: the Chinese develop a chip, start producing it with one of the government-influenced businesses, and the US freaks out and bans the chip from all US machines and networks, citing vague but terrifying backdoors. Because, apparently, Nvidia hasn’t thought of that …? Joking aside, the Chinese today are more pure capitalists than the US is. Witness a case in point: US networking companies created an alleged backdoor in Huawei networking gear and got it banned on critical and corporate networks. Do you think that Cisco Systems’ wholly-owned subsidiaries in congress were thinking of the competitive infrastructure landscape when they passed those directives? Not for a second! We have the best congress money can buy.

Anyhow, the big power-gobblers on a GPU are the screen RAM, the RAM cache for object models, and the multiple cores to drive it all. I forget how many cores my GPU is running, but it’s a lot and, uh, it gets things done. If someone produces a board with a bunch of RAM and many cores and a usable interface that is not tied up in non-disclosures, AI researchers will link their code – -define NONCUDA -lnoncuda and not only will the result be faster, cheaper, and it will gobble less electricity.

There has been considerable discussion about whether or not we are in a bubble. Of course we are! We are always in a bubble of something-or-other since US Steel, Ford Motors and Pullman rail cars stock got overvalued based on the cost of labor. The entire history of the US has been a sequence of bubbles: the Oklahoma land rush bubble, the railway bubble, the steam power bubble, the civil war weapons bubble, the cotton bubble, the slave bubble. But those are all examples of Great American Rip-Offs, where people paid ridiculous prices for growth stock, based on the obvious growth of an industry. The problem, like with the railways, is that AI data centers are something you can get the shareholders to invest in; remember it’s just the company diluting its stock to raise $1bn and build a new data center – but what that also does is raises the value of early insiders’ shares to catastrophic heights. When the “Lamborghinis and cocaine” guys start to glom onto the businesses, they’ll pretty quickly lead them astray. I know those guys. Those are the guys that’ll tell Facebook to buy some stupid company for a ton of money in order to move the stock (which they happen to hold) and the whole thing gets written down in the loss column 3 years later. Anyhow, yes, we are in a bubble. Companies like OpenAI are stuck because they have convinced the market that they need more money for GPUs and data centers – therefore they do a stock offering to the public for more GPUs and data centers, and raise $2bn. Then, the insiders liquidate a few of their shares, buy Lamborghinis and coke, and what the heck, may as well build a data center while we are at it! What happens if the cost balance implodes? Do you remember 2008? A lot of people got out over their skis, and face-planted. When people talk about a “bubble” – ignore them, they’re just rich guys wondering about the fate of other rich guys. The workers will get fucked either way, because capitalism needs a cheap work-force that is desperate, to threaten the expensive work-force that is comfortable. [No, I am not a Marxist but Marx’ understanding of the flaws in capitalism was profound.]

[Krea, the view from within a bubble, looking into a data center.]

Your Nvidia stock is good as gold for a long time because there are a lot of gamers and Intel hasn’t managed to cough up a good integrated chip with CPU and graphics together. They have been consistently out-marketed. Some day, they may implode but I doubt it, because that would take down Oracle, Microsoft, Cisco, and the rest of the US tech market. So, if the Chinese come up with something great, they’ll get buried in legislation and evaluations, etc. Where there might be problems is if a company that is ostensibly allied, (e.g.: Qualcomm) comes out with an AI chip and then the US will have to put on its stripper shoes and dance really hard. If you look at the history of the US stock market, the public gets a fleecing pretty much every 14 years. It’s not a hard-defined cycle but I think it has something to do with investors, and forgetting, and the memory effect of cocaine. Or something? [Actually, cocaine boosts my memory to uncomfortable levels, so much that I discontinued it long ago]

Item #2: [nvidia]

In another step forward, Equinix opened in January a dedicated facility to pursue advances in energy efficiency. One part of that work focuses on liquid cooling.

Born in the mainframe era, liquid cooling is maturing in the age of AI. It’s now widely used inside the world’s fastest supercomputers in a modern form called direct-chip cooling.

Liquid cooling is the next step in accelerated computing for NVIDIA GPUs that already deliver up to 20x better energy efficiency on AI inference and high performance computing jobs than CPUs.

When they say “born in the mainframe era” they mean “this is how real computers have always worked” – the CRAYs, IBM mainframes, CDC Cybers, etc. Another way of looking at it is that “room temperature” computing is a relatively new phenomenon which did not really exist! If you put an 80386-powered desktop in someone’s cubicle it’s not unbearably hot, or phenominally loud, but it’s a 400-watt toaster oven worth of heat and it’s only going into the office-space air conditioning. What we are looking at is incredibly lazy computing: the AI guys have taken basic computing and just scaled it insanely with software, without thinking of where they could improve it.

[It’s not a hot tub, it’s the output of a head-exchanger powered by a volcano. Iceland rules.]

In terms of cooling, closed loop is the obvious answer. There are more subtle answers, too, but most of them build on the closed loop concept of capturing “waste” energy. Also, if you think about it, a closed loop cooling system brings all of its members toward a sort of average temperature. A closed loop does not need municipal water to cool it – you run a big heat-exchanger and bury it in a lake, or underground, or offshore, or build an avocado farm on top of it. The icelanders have figured this out: they have a thing they call “the blue lagoon” which is the heat output of a geothermal plant that they charge people to swim around in the winter: win, win, win. Don’t tell me we’re not as smart as the icelanders, damn it. OK, most of us aren’t but some of us are.

In separate tests, both Equinix and NVIDIA found a data center using liquid cooling could run the same workloads as an air-cooled facility while using about 30 percent less energy. NVIDIA estimates the liquid-cooled data center could hit 1.15 PUE, far below 1.6 for its air-cooled cousin.

Liquid-cooled data centers can pack twice as much computing into the same space, too. That’s because the A100 GPUs use just one PCIe slot; air-cooled A100 GPUs fill two.

In raw Marcus-terms, liquid cooled data centers appear to be twice as fucking smart as air-cooled data centers. Because we’ve known about liquid cooling since the 70s. Even advanced PC case-modders know about liquid cooling. [many of them won’t run antifreeze in them because it’s conductive and will cause a big no-no if it leaks. So, instead they run water and their CPU chiller blocks become coral reefs full of obscure forms of algae.] I say, run vodka in the damn things, and confusion to the French!

Item #3: optimizing code

There are some indications that AI engines aren’t very well-coded. I don’t have references here to cite, but I saw one researcher who had actually done some cache effectiveness analysis on the AI’s keyword->vector cache and determined that it was ineffective. For example (making up things here) a typical query didn’t hit the same words often enough that caching the mappings did anything but eat a few gb of memory over time. When I saw about that, I was horrified. It appeared that some Computer Science PhDs had possibly been working on the code – and it’s stereotypically amateurish. A professional programmer (at least, any one that ever worked for me) would know “you do not pre-optimize your code.” The fact that someone just decided, apparently without measuring cache hits/misses versus updates, just put a big cache into production code – horrifying. Back when I managed development, you didn’t just add a feature, you documented an analysis of what it would do, and the benefits of adding it. When I see features being added uncomprehendingly, I assume I’m looking at Bullshit Code(tm) that can be reduced by 20-30% and improved by 100-500%. That’s just my industry experience, I could be wrong. But my experience also says, “I doubt it.” If code flaws like that are slipping into production, I can pretty much predict the software team is a “scrum” or some such bullshit, more focused on their Lamborghini than writing tight code.

That brings me to another philosophical point: the singularity. The premise of the singularity is that AIs will code better AIs and eventually they will tight-loop improvements until humans no longer comprehend what is going on. Well, not if the current generation of latte-swilling motherfuckers are teaching the AIs how to code. There is an interesting problem there, some of my friends and I have noodled about, which would be to start a business rating the code quality of available code, before AI ingests it. “Everything Paul Vixie writes has buffer overruns” etc. There will be no singularity if AIs are depending on the code of the last generation of latte-swilling motherfuckers who brought us PHP and JSON. It just won’t happen. Maybe we’ll get a reverse singularity, where AIs start acting as dumb as coder-bros and we have to nuke the whole mess from orbit, to be sure.

So, all in all, I think the bubble will burst for financial reasons of un-sustainability, and the technology bubble will continue forward. I’ve noticed that fads in computing often liberate previous incarnations to be art, instead of commerce. A lot of artists are complaining (fair enough, but their dying complaints are pro forma) that AI is collapsing the commercial value of illustration. Like illustration collapsed the commercial value of painting, and photography collapsed the commercial value of illustration. Each of these collapses is a blessing and a curse – mediocre artists who were making their money doing commercial illustration, are now liberated to pursue their skills as pure art – because nobody is paying for the old stuff. “But what if they can’t make it as pure artists?” you ask. Well, if they couldn’t make it as pure artists, they were kidding themselves about their commercial skills, too. Those folks will be unhappy but they were going to be unhappy no matter what. I’m sorry.

There are so many tropes tangled up in the anti-data-center memeology that it’s hard for me to sort out. The opponents come off, to me, like irrational anti-vaxxers: they are throwing everything at the wall and hoping something sticks. “Data centers are sucking all our electricity!” OK but what if the code was more efficient by a factor of 10, and the GPUs were replaced with cards that ran less power and did the same work? “Data centers are still bad because they are consuming all our water!” Sure I hear you but what if they were closed-loop and used the extra heat for growing yummy marijuana plants? “AI are using stolen data!” Well, you need to understand a bit more about machine learning, and study the differences between human learning and machine learning and ask yourself if reading a textbook on math is stealing from a mathematician or not. Fortunately for us all, runaway capitalism is taking the reins and what will happen is what will happen. I hope the AI turn out to be non-profitable, and are liberated as art. I am beginning to see some signs of that, but also signs of great transformation, so I am unsure where it all will go. The good news is that the technology will remain, because innovation is irreversible. I am often reminded of the AI in Brin’s The Postman – a genius-level gentle AI that was produced just as human civilization collapsed, which died because of lack of super-cooling for its chips.

Ironically, it probably could have rescued humanity if the power hadn’t failed.

------ divider ------

As usual there is various stuff going on in my life that I have stopped talking about. My crafting work interests mostly me, and explaining it outside of myself is time-consuming and uninteresting. I don’t want to just post pictures of blades without their stories, but I don’t want to write their stories, either.

My mom died a month or so ago, and my last few weeks have been me and dad doing a memory tour (for him) retracing my childhood memories. I am pretty sure he enjoyed refreshing them. For me, since Paris was experiencing a supremely unlikely heat-wave (predicted by science) it was especially difficult. Holding a handful of what used to be my mother, and throwing it away, was a surreal experience. Fuck dementia. And, please, dementia, stay away from me.

I’ve been insanely busy with new stuff, some of which you might enjoy some of which you might not. I think that is what has held me back from blogging. I keep trying to think what is interesting, but I fundamentally realize that nothing is, any more. I have some fun AI art stuff and other things – painful beautiful things – I don’t know if I want to summon up the energy to write about. We’ll see. Stay tuned or fuck off, as you see fit.

@Pierce: I see your thoughtful comments about F-35s and the ongoing disaster there. I’ve been tempted to write about that, (concurrency and the radar is one, but also just plain flight-worthiness is not there) but it’s so painful to contemplate my tax money going for those bullshit aircraft which are only good at making short-range bombing runs against defenseless civilians. It makes me unusually sick.

Comments

  1. C Sue says

    Aw, so sorry to hear about your mom. My one memory of her, is her letting me play her harpsichord. :>

  2. Pierce R. Butler says

    My condolences too – it gets harder to navigate life when a major landmark disappears.

    … chips for inference…

    They’ve come a long way conceptually, to design for high-level abstractions out of basic binary math. Why not go straight for “profundity” or even “wisdom”?

    The entire history of the US has been a sequence of bubbles: the Oklahoma land rush bubble …

    Going back even earlier, until just after the Revolution, the trans-Allegheny land bubble once King George’s treaties with those pesky Indians could be torn up and Ohio went on the market. Robert Morris, primary financier of the Revolution (if you don’t count Louis XVI), who reshaped the US government to get his loans paid back (unlike Louis), lost his shirt (again) and died in poverty when that one popped.

    … you run a big heat-exchanger and bury it in a lake… or offshore…

    Once again Florida has served the nation by offering a cautionary example. A big power station did exactly that, correctly calculating their waste heat would be easily absorbed by the Gulf of Mexico (as it was known at the time). And it was. Until they had to shut down for repairs one nippy winter, and a big herd of manatees which otherwise would’ve migrated southwards all froze to death. (I know: California techbros never heard of manatees…)

    As for the F-35, maybe the AF can buy some Ford F-350s, chisel off the zeros and epoxy on wings and an afterburner. After all, according to DHS/ICE, motor vehicles are quite easily weaponized.

  3. JM says

    Back when I managed development, you didn’t just add a feature, you documented an analysis of what it would do, and the benefits of adding it. When I see features being added uncomprehendingly, I assume I’m looking at Bullshit Code(tm) that can be reduced by 20-30% and improved by 100-500%.

    The LLM companies don’t know how their LLM’s work internally, it’s all emergent properties of big node networks. The idea of LLMs itself was just an accident, somebody took one of the older much smaller AIs and raised it’s node cap 100x and this created an AI that appears many times smarter then old AIs for reasons that are not entirely clear.
    One side effect of this is obscurity of internal function is that there is no way to rigorously determine what a change will do ahead of time. They just have to put forward something that the team agrees looks like a good idea and try it.
    At the same time the AI companies are in computer market startup mode. They are all trying to build up enough market share they can use that market share to their advantage. The goal going along with this is to get big enough fast enough that the IPO looks good before people start asking questions about long term profitability. Because of this sprint for market share they rush stuff into production so they can constantly claim improvements.

  4. Reginald Selkirk says

    There are so many tropes tangled up in the anti-data-center memeology that it’s hard for me to sort out. The opponents come off, to me, like irrational anti-vaxxers: they are throwing everything at the wall and hoping something sticks. “Data centers are sucking all our electricity!” OK but what if …

    That what if is doing a lot of heavy lifting. The problems with AI data centers are very real: their power use, their water use, their noise, they way they are sucking resources out of every other part of the economy…
    What if the rich fucks who see this as their opportunity to get richer by exploiting common resources held off until some of your what if improvements came through?

    And now the rich fucks are playing the patriotism card: China, Russia and Others Seek to Inflame Debate Over A.I. Data Centers

    So unfriendly foreign powers wish to divide our country further- so what? Does that make any of the well-documented problems with AI data centers untrue? (Hint: the answer is two letters long and starts with ‘n’.)

  5. says

    JM@#3:
    One side effect of this is obscurity of internal function is that there is no way to rigorously determine what a change will do ahead of time. They just have to put forward something that the team agrees looks like a good idea and try it.

    This ought to be reasonably possible to regression test. But I’m hardly arguing for any functional changes to the code – just do basic performance analysis and optimization. For example, every cache I ever coded into a production application had some kind of cachestats( ) function that dumped hits, misses, chain depth, hot updates, cold updates, etc. If the cache wasn’t serving 80-90% of the reads, it wasn’t doing very well. If my hash function wasn’t performing well, I knew that, too. I knew what parts of my search were getting called a lot, and whether memory allocation was a problem, or not. Etc. If someone went ahead and added a cache to a process that didn’t benefit by it, that tells me a tremendous amount about the code-base: it’s written by CS101 students who haven’t learned not to pre-optimize, yet. You don’t start just slamming hash functions onto code until you know it needs it, and whether a hash is the right thing, or a b-tree, or a bloom filter, etc. [Also, with the amount of memory available to modern systems, it’s a fair question whether cache is needed, versus just letting the systems paging algorithm take care of things if there’s any locality of reference.] At the scales that we are talking about, a small performance improvement in the critical path can have a mind-blowing effect. There is a lot of stuff I simply have not tried to understand about these implementations but it does look to me like most of the work is on training sets rather than engine implementation. If we’re not at the point where anyone cares, we’re going to be over-scaling in order to make up for poor code.

    At the same time the AI companies are in computer market startup mode. They are all trying to build up enough market share they can use that market share to their advantage. The goal going along with this is to get big enough fast enough that the IPO looks good before people start asking questions about long term profitability. Because of this sprint for market share they rush stuff into production so they can constantly claim improvements.

    That’s a neat summary of being in a start-up. “We’ll fix it later or we’ll go out of business.” That’s one of those efficiencies that capitalist competition encourages. And, of course, it never gets fixed in either case.

    One of my friends used to call the early 90s (after the netscape IPO) “the age of the endless beta-test” and I think he was right. Netscape conclusively showed that a half-product could be worth a brazillian dollars.

  6. says

    Reginald Selkirk@#4:
    The problems with AI data centers are very real: their power use, their water use, their noise, they way they are sucking resources out of every other part of the economy

    True.
    In the 1980s I owned and drove a Honda CRx HF. 50mpg highway. But the techniques used to build it were not widely adopted, the current options are hybrids that are barely as efficient as a big-block Chevy. What didn’t happen? What didn’t happen was the customers didn’t turn out to be ready for it, and still aren’t. The same sort of thing seems to be the case with all of the energy efficiency and whatnot. It’s capitalism, damn it. The data centers represent a huge form of pump-and-dump. That’s one reason that things like the Chinese code releases are so scary to the capitalists: what if someone comes along with an AI engine that runs in CPU time and doesn’t need a GPU? Or that uses the built-in GPU that is on every Intel processor? [That’s an old dream of mine: could an Intel CPU be fooled into performing GPU operations even while there is a GPU card installed? It’s a spare billion instructions per second and a memory cache…]

  7. Dunc says

    Sorry to hear about your mom. Mortality kinda sucks.

    Even advanced PC case-modders know about liquid cooling.

    Yeah, I was building liquid cooled PCs back in the glory days of overclocking, maybe 15-20 years ago now… People were even building multi-stage rigs, where you cool your primary loop with a TEC chiller, and then run a second liquid cooling loop to cool the hot side of that through a big radiator. Condensation management becomes an interesting problem once you start doing that sort of thing though. Still, it all has to get dumped to background one way or another, and all of the clever tricks you can employ also put more heat into the system in total, so you can pretty quickly hit a point of diminishing returns. Cooling into the ocean is a good trick if you can manage it, but it does kinda limit where you can build… (Also it has some surprising failure modes.) Cooling into the ground is limited by thermal conductivity, and in a lot of places that’s not actually very good unless you’ve got a lot of moisture down there. (This can be a serious design consideration even for domestic ground-source heat pumps, which are dealing with really quite modest amounts of heat.) One of the most effective and flexible options is evaporative cooling – sucks in a domestic environment, but it’s great at industrial scale – but that’s when you start actually using up a lot of water. (OK, it all rains back down somewhere eventually…)

    OK but what if the code was more efficient by a factor of 10, and the GPUs were replaced with cards that ran less power and did the same work?

    And as the saying goes, if my grandmother had wheels, she’d be a bike. Sure, in a different world, where this technology was being built in a different way, for different reasons, by different people, then things would be different. Meanwhile, in the real world in which we actually live right now, people are turning obsolete jet engines into gas turbines to power these things. [IEEE]

    The data centers represent a huge form of pump-and-dump.

    This is the fundamental problem. It’s not about technology any more, it’s about leveraged real-estate deals. There’s no incentive to make anything more efficient, because the only thing that matters right now is NUMBER GO UP! Making things more efficient does not make the number go up. What does the number measure? Who cares! Number is number. Big number better than small number. NUMBER MUST GO UP!

    Companies like OpenAI are stuck because they have convinced the market that they need more money for GPUs and data centers – therefore they do a stock offering to the public for more GPUs and data centers, and raise $2bn.

    Dude, you’re off by two orders of magnitude. OpenAI’s last private funding round was for $122bn, and Oracle have issued $100bn of debt to build data centers for them*. They’re targeting an IPO valuation of one trillion dollars! [Insert Dr Evil gif]. A mere $2bn doesn’t even get you a cup of coffee in this market.

    *And those bonds are going to need to be paid, whether the data centers actually get built or not, and whether OpenAI can actually pay for them or not…

  8. rrutis1 says

    Marcus @6.
    You are right about the capitalism ruining opportunities for the energy use reductions you are talking about. I work in energy efficiency projects and unless mandated the money people balk at anything that does not have a 3 year or less ROI. I have seen stakeholders reject projects that would save 40% of their utility bills and make the building/process/whatever work better. So frustrating.

    I am also very sorry to hear about your mom. Maybe if you write some of your stories about her here it will help you process the loss.

Leave a Reply