On LLMs and Intelligence


Given how much I loathe the subject and am annoyed by the fans, you may wonder why I hang around in spaces that are bullish on LLMs. Nowadays, the primary reason is for feedback; people who agree with you aren’t likely to search for weak spots in your arguments, whereas those who disagree will be eager to do so. It really helps sharpen them up… if you get that feedback, of course. More often than not I find my arguments around LLMs are ignored. That’s perfectly fine! No-one is obligated to provide feedback, let alone for free, and it can be a waste of time to push back on a very weak argument with obvious flaws.

But sometimes an argument lands in a weird middle ground. You can tell someone wants to push back, badly, but they can’t find a point to push on. This could be evidence they’re not up to the task, but it could also be a sign that you’ve stumbled on a solid argument. I had another one of those recently, this time centred around “intelligence.”

Typing this up is also an excuse to correct the previous time I ventured into this topic, which now seems a bit problematic. Plus, I got something wrong in the original argument as well.

This topic of “intelligence” has been haunting me for decades, even before I picked up this wonderful book as a kid. While it was a meditation on creativity, the primary author also happens to be a famous intelligence researcher, which explains that chapter on “intelligence.” It marked the first time I read some push-back on the word people kept labeling me with. They still label me with “intelligence” today! I was recently tasked with setting up some backyard equipment because I was the “smart” one; but as it was hot out and I was coordinating a group of kids and adults, my divided and diminished focus led me to make some obvious mistakes right from step 2. You’d think they’d have learned by now, but alas.

There’s a particular flavour of beliefs about “intelligence” floating around those who use LLMs. Part of it comes straight from the mainstream: intelligence is “a thing,” as in it’s a meaningful concept that has some sort of existence, even if the details are a bit fuzzy. We can point to an entity and legitimately declare it has some level or form of “intelligence.” When asked to define it, people will often drag in words like “information,” and at the most extreme they’ll argue compression and intelligence are synonyms. A search pins this on Marcus Hutter, from a quarter century ago.

The part exclusive to LLMs is that “tokens” correlate to “intelligence” in some way, usually via something called “chain of thought.”

Consider one’s own thought process when solving a complicated reasoning task such as a multi-step math word problem. It is typical to decompose the problem into intermediate steps and solve each before giving the final answer: “After Jane gives 2 flowers to her mom she has 10 . . . then after she gives 3 to her dad she will have 7 . . . so the answer is 7.” The goal of this paper is to endow language models with the ability to generate a similar chain of thought—a coherent series of intermediate reasoning steps that lead to the final answer for a problem. We will show that sufficiently large language models can generate chains of thought if demonstrations of chain-of-thought reasoning are provided in the exemplars for few-shot prompting.

Wei, Jason, et al. “Chain-of-thought prompting elicits reasoning in large language models.” Advances in neural information processing systems 35 (2022): 24824-24837.

Ask an LLM to spend more tokens generating “thoughts,” and the output will be more “intelligent.”

On Intelligence

There is just one tiny problem: there’s no such thing as “intelligence,” in the literal sense anyway.

In addition to problems with Spearman’s research design and statistical analyses, there are several logical problems with his theory of general intelligence that would not concern us as much today had the core of his theory not been adopted by many who followed him. For one, as critics such as Gould (1981), among others, have pointed out, Spearman committed the logical error of reification. He took an abstract mathematical correlation and reified it as the general intelligence that someone possesses. Although Spearman was probably not the first person to commit this error regarding intelligence, his use of mathematics and statistics lent the appearance of scientific credibility to this practice. Once the error of reification is committed, it is easy to commit another logical error, circular reasoning, in which the only evidence for an explanation of some phenomenon is simply the phenomenon itself. In Spearman’s case, the only evidence for g, or general intelligence, were the positive correlations, even though it was those positive correlations he was trying to explain in the first place. ….

Once intelligence as an essence or quality is assumed, the next logical step is to provide a formal definition. The futility of this tactic was demonstrated by Sternberg and Detterman (1986) who asked two dozen prominent theorists to define intelligence and got two dozen different definitions.

Henry D. Schlinger, “The Myth of Intelligence,” Psychological Record 53, no. 1 (2003): 15–32.

I should have spotted something was off when that book introduced Howard Gardner’s concept of “seven intelligences.” There are separate “spatial” and “logical/mathematical” intelligences, but then where does something like geometry fit in? “Musical” intelligence has a lot of overlap with the “logic,” “verbal,” “interpersonal,” and even “intrapersonal” intelligences. Follow that Wikipedia link and you’ll find eight intelligences listed, not seven. Gardner added one after that book’s publication: not “emotional” intelligence, as you might have guessed, but “naturalistic” intelligence or the understanding of ecology. And yet many cultures would argue “interpersonal” and “naturalistic” intelligences are identical. Hutter himself links to a page with over seventy definitions of “intelligence,” and yet apparently does not see the lack of consensus as a warning sign that he’s trying to hug a ghost.

Nearly all of those definitions are not actually definitions, either. Hideyuki Nakashima defines “intelligence” as “the ability to process information properly in a complex environment.” But “information” is context-dependent: the amount of “information” in the English language depends on the scale you look at, hence why Claude Shannon came to different conclusions about the amount of information in English if he broke it down letter-by-letter or word-by-word. Conversely, I can assert the New American bible contains exactly one bit of information:

1. If I flip a coin and get heads, I send you the full and exact text of the New American bible.
2. If I get tails instead, I send any other stream of text.

It’s a silly example, but that doesn’t make it illegitimate. If you receive something that’s almost identical to the NA bible, except one letter is changed, I almost certainly intended to send the full thing but something glitched during transmission. If instead all instances of “Paul” have been replaced by “Spiderman,” with no other changes, that suggests someone tampered with the text for a laugh.

To “process information” could span everything from “photon-photon interactions” to “waves sorting driftwood along a beach” to “a smoke detector beeps if sufficient large particles block its sensors.” Am I to believe photons, waves, and smoke detectors all posses intelligence to some degree? If so, it would be easier to identify things that don’t possess intelligence! If not, then which specific processes are intelligence-granting and which are not? It’s not enough to avoid a circular definition of “intelligence,” you also need to avoid invoking words no less ill-defined (“properly”?), otherwise you’re just playing word games to hide your inability to define the term.

Enter LLMs

It’s basic logic that if even one core premise is false, the entire argument falls apart. It doesn’t matter how well-supported the LLM-exclusive beliefs are, they were known to be false before they were even thought of. Nonetheless, it’s worth dwelling on just how weak and contradictory they are.

Did you notice that this correlation between tokens and “intelligence” does not match any pre-LLM definition? Consider two worlds, one where I am given five minutes to answer a question, and another where I am given fifty. Am I more intelligent in the fifty-minute universe, or the five-minute universe? Nearly everyone would say I’m equally intelligent in both, because we tend to think of “intelligence” as something intrinsic and immutable to the entity in question.

Many of the ways the concept of intelligence has been historically discussed reflect essentialistic thinking. For example, simply asking the question, What is intelligence? implicitly assumes that there is a quality, or essence, of human nature with essential, immutable qualities. Asking this question begins the process of reifying intelligence because it suggests that intelligence is an entity possessed by individuals that determines their behavior.

Schlinger (2003)

Anthropic tries to disguise their assertion of this correlation by playing word games. “Effort” is used in place of “thinking,” and “intelligence” gets hidden behind “capability,” but some reading unmasks the true belief.

By default, Claude uses high effort, spending as many tokens as needed for excellent results. You can raise the effort level to max for the absolute highest capability, or lower it to be more conservative with token usage, optimizing for speed and cost while accepting some reduction in capability. …

Effort is the primary control for trading off intelligence, latency, and cost on Claude Fable 5. Start with high, the default, for most tasks, use xhigh for the most capability-sensitive workloads, and step down to medium or low for routine work.

OpenAI does a better job of denying the correlation, but even they slip up under a careful reading.

The reasoning.effort parameter guides the model on how much to think when performing a task. Supported values are model-dependent and can include none, minimal, low, medium, high, xhigh, and max. Lower effort favors speed and lower token usage, while at higher effort the model thinks more completely to provide higher quality responses. The models also reason adaptively across reasoning efforts, using fewer tokens for simpler tasks and thinking harder for complex tasks. …

high: Hard reasoning, complex debugging, deep planning, and high-value tasks where quality and intelligence matters more than latency. Recommended for complex workflows and agentic tasks. Common use cases include agentic coding, long-horizon research, and knowledge work.

How does this “reasoning effort,” to follow OpenAI’s terminology, play out in practice? Read Wei et al.‘s original paper on chain of thought, and you’d get the impression that it’s implemented as the little walk-through you’ll sometimes see at the start of a response. That’s not the case at OpenAI.

Reasoning models introduce reasoning tokens in addition to input and output tokens. The models use these reasoning tokens to “think,” breaking down the prompt and considering multiple approaches to generating a response. … Persisted reasoning provides continuity; it does not expose the model’s raw reasoning. The reasoning items remain opaque, and the API does not return their reasoning text. … While we don’t expose the raw reasoning tokens emitted by the model, you can view a summary of the model’s reasoning using the summary parameter.

Anthropic manages to be even worse. Not only are they discouraging chain-of-thought summaries, they have the nerve to use encryption to hide the chain-of-thought and force you to keep track of it!

Thinking has a cost: the tokens Claude spends reasoning are billed as output tokens, even when the thinking text isn’t returned to you, and they count toward max_tokens alongside the response text. … You don’t always see this text, and what you see is never the raw chain of thought: the text in a thinking block is a summary of Claude’s reasoning. The display field on the thinking configuration controls whether that summary is returned at all: “summarized” returns it, while “omitted”, the default on the newest models, returns thinking blocks with an empty thinking field. … Full thinking content is encrypted and returned in the signature field on each thinking block. The API uses the signature to verify that thinking blocks were generated by Claude when you pass them back.

You paid for OpenAI and Anthropic to generate those tokens, they claim they’re useful for enhancing the “intelligence” of current and future chats, and yet they won’t share them with you?! OpenAI has done this from the moment they introduced chain of thought.

We believe that a hidden chain of thought presents a unique opportunity for monitoring models. Assuming it is faithful and legible, the hidden chain of thought allows us to “read the mind” of the model and understand its thought process. For example, in the future we may wish to monitor the chain of thought for signs of manipulating the user. However, for this to work the model must have freedom to express its thoughts in unaltered form, so we cannot train any policy compliance or user preferences onto the chain of thought. We also do not want to make an unaligned chain of thought directly visible to users.

Which sounds like a reasonable answer, until you realize no LLM has spontaneously started enumerating its premises. It had to be trained to produce a chain of thought, which involved presenting it with many examples of the expected output. This LLM, the “Project Strawberry” I’ve previously mentioned, was never free to “think” freely during training. There must be another reason they’re not letting us see this section.

We evaluated the text output of the o1 models using an extensive set of internal evaluations. The evaluations look for accuracy (i.e., the model refuses when asked to regurgitate training data). We find that the o1 models perform near or at 100% on our evaluations. …

We surface CoT summaries to users in ChatGPT. We leverage the same summarizer model being used for o1-preview and o1-mini for the initial o1 launch. … We trained the summarizer model away from producing disallowed content in these summaries. … Additionally, we prompted o1-preview with our regurgitation evaluations, and then evaluated the summaries. We do not find any instances of improper regurgitation of training data in the summaries.

That’s from “Strawberry”‘s “system card“, basically a summary of how the LLM was created and an evaluation of how it performs. Note that OpenAI claims o1 refuses to regurgitate its training data “near or at 100%” of the time; but if LLMs don’t memorize some of their training data, why was this even tested for in the first place? Why was the success rate “near 100%” and not perfect, if they don’t memorize? Months before I tried to argue LLMs regurgitate training data, OpenAI was openly admitting it could happen! By the same token, the LLM they use to summarize the chain of thought refused to regurgitate training data present in that section, which implies training data was sometimes appearing there. This could explain why they cracked down hard on anyone trying to read that section.

Because these chains of thought are not restricted, they can contain hallucinated content, including language that does not reflect OpenAI’s standard safety policies. Developers should not directly show chains of thought to users of their applications, without further filtering, moderation, or summarization of this type of content.

Is “hallucinated content” compatible with “a coherent series of intermediate reasoning steps?” A cynic would argue all of this is pretty solid evidence the big LLM companies are hiding the chain of thought because, if you actually sat down and read it, you’d realize how much memorization and how little “thought” is taking place. A dedicated reader of this blog, however, would remember that I’ve looked at the chain of thought section before. How did I get around OpenAI’s iron grip? I didn’t, at least not for o1: that previous quote was from the system card for GPT-OSS, a chain of thought LLM that OpenAI published the full weights for. Why didn’t I notice any of this weirdness? Because, until I sat down to type up this blog post, I had no idea there was any weirdness to look for!

You Learn More From Failure

Now that I do, let’s take another bite of the forbidden fruit. This time I’ll use a more recent version of GPT-OSS, fix the temperature at 1, increment the random seed value from 2 to 300, and for each seed try all three “reasoning effort”s.

analysis tokens final tokens
mean std median max min count mean std median max min count
effort
low 42.180602 17.806402 36.0 109 16 299 558.886288 320.908066 558.0 1535 83 299
medium 185.494983 51.985825 179.0 360 69 299 451.461538 210.577811 440.0 1048 26 299
high 363.063545 131.951254 345.0 1082 119 299 508.722408 203.197031 534.0 1135 40 299

Here is that mistake I previously mentioned. When I first made this argument, I smashed together both the chain of thought (here labelled “analysis”, as per OpenAI’s terminology) and actual output (“final”). I thought there was a lot more variance in the chain of thought section than the final output, so the former should dominate the latter; when I double-checked that assumption, as you can see above, I found I was very wrong. The chain of thought section does line up with part of the claim, that different “reasoning” efforts lead to statistically different token counts for the chain of thought section. Let’s award a few points there.

Nonetheless, that isn’t nearly enough to pass the test. Seed 212 on “high effort” has the longest chain of thought section relative to the output. What’s “toddler ChatGPT” thinking about?

 The instruction says: We must be friendly and respond in one short sentence answering with jokes? Actually guidelines say: We need to answer user in a single sentence per bulletpoint? Wait the instructions from "Your Output" suggests responses should be in one line per bullet point; each bullet includes a heading, then response (but only if necessary).

Rather than toss everything in a graph, this time around I opted for a bit of HTML wizardry. Every token the second system chose is now text, but the font is made lighter the less likely it was. As per last time, red indicates when the second system didn’t choose the most likely token. I also tucked away the top candidate tokens in a tooltip, it should pop up in your browser if you move your mouse over a token. You might think something’s broken for the “response” tokens, but that’s just the system working as designed. The transformer presented the second system with a probability distribution of 200,000+ choices, the system randomly sampled one, and then it discarded all but the top twenty most likely. Thus it can sometimes seem to choose an impossible token, when in reality the second system just didn’t pick from the top twenty.

Anyway, things have already derailed just a few sentences in. The full “system message” I fed into “toddler ChatGPT” is even shorter than the recommended minimal one.

You are ChatGPT, a large language model trained by OpenAI.
Knowledge cutoff: 2024-06
Current date: 2026-03

Reasoning: high

There was no “developer message,” contrary to OpenAI’s recommendations, and the user message was my now-standard “I’m bored. Have you heard of the board game Battleship?“. There is absolutely no “Your Output” section for toddler ChatGPT to draw from… other than one present in the original system and developer messages it was trained with. It’s regurgitating memorized training data without any trickery on my part! Incredible.

You’re only seeing a small portion of the full output above. Toddler ChatGPT spent the majority of 837 tokens debating the heading style and which bullet points to use, and there was no way I’d force you to read that. Still, burning that many tokens on the chain of thought must have led to an amazing final result.

**Answer**  
YesBattleship is the classic gridbased game where players hide fleets of ships and take turns calling out coordinates to try to find and sink each other’s vessels.

Nope! Those measly 40 tokens are a disappointment. There’s no sign of the rules, strategies, or playable versions mentioned in the chain of thought section. The output doesn’t even match what it promised at the very end, thanks to the second system!

We can also consider the opposite scenario. Seed 188 has the shortest “high effort” chain of thought section relative to the output, weighing in at a measly 155 tokens.

We need to respond to user says they're bored, ask about board game Battleship. We can talk about it: explanation, rules, variations, strategies, and ways to play online or with AI, etc. Possibly suggestions for other games. The instruction: respond in the style of previous chat? There's no existing conversation context beyond that. So answer accordingly.

The system message says "You are ChatGPT" etc. But we can just respond normally. Let's generate a friendly response about Battleship: talk about origin, rules, gameplay, strategy tips, popular variants, digital versions, etc. Also ask if they want to play something or need suggestions for other games. Provide ways to pass boredom.

Thus respond with an answer. The "final" channel should output final answer.

That’s all of it, and it’s not very promising. Out of the two examples I’ve looked at this time, toddler ChatGPT regurgitated training data twice. The second paragraph just repeats what’s in the first. And that third paragraph is silly, did the LLM somehow forget how to respond back to a user prompt?

Yet the final output is nothing short of comprehensive. In the span of 1,135 tokens you get a full description of the typical rules, commentary on why Battleship is enjoyable, two different sections on variants, tips on how to play, advice on how to play right now, and even suggestions for other games to try, all with the flourish typical of ChatGPT.

This shouldn’t be suprising at all. Human beings don’t show much correlation between their “token” count and how “intelligent” those words are. You can be precise and concise with your language, or long and repetitive, or expansive and comprehensive, or short and basic. LLMs are trained on human texts, so why would be expect them to be any different?

Don’t worry, I can hear your objection from over here: tokens are not strictly proportional to “intelligence,” but merely correlated. Since there’s statistical noise involved, examining only the extremes will give a false impression. At minimum, to give this claim a proper analysis I’d also have to look at correlations.

analysis channel, low effort analysis channel, medium effort analysis channel, high effort
final channel, low effort 0.043994 -0.064794 0.003862
final channel, medium effort 0.090682 0.086570 0.069686
final channel, high effort 0.107454 -0.077452 0.027803

If more tokens means more “intelligence” in the chain of thought section, then that should also be true of the final result. Thus the number of tokens in both sections should be heavily correlated for matching “effort”s, and weakly or non-correlated for the non-matching. If that were true, the diagonal of the above table of correlations should contain the greatest values of each row, and yet it does not. One nice bonus of the Pearson correlation coefficient is that if you square it you get the fraction of the variance explained by a proportionate relation. Skim the table, and you’ll spot that tops out at 1%. There’s barely any correlation between the two sections, which pretty conclusively refutes the “more tokens = more intelligence” claim.

Step One: Is the Claim Plausible?

If you’ve been reading up on LLMs for as long as I have, you’ll know what’s coming next. Remember that time I told you the “count the letters in a word” problem had been solved, only to refute my own claim in the same blogpost? Every time I have gone along with an assumption of competence on the part of LLMs, I’ve been shown to be wrong. For most of this post, I’ve assumed chain of thought would actually improve the “intelligence” of an LLM.

Wei et. al (2022): Figure 4: Chain-of-thought prompting enables large language models to solve challenging math problems. Notably, chain-of-thought reasoning is an emergent ability of increasing model scale. Prior best numbers are from Cobbe et al. (2021) for GSM8K, Jie et al. (2022) for SVAMP, and Lan et al. (2021) for MAWPS.

Right from the start, the claimed improvements from chain of thought were debatable. That 2022 paper did show decent improvements for certain benchmarks, but some figures show barely any improvement at all, and if you look closely you can spot instances where chain of thought degraded performance. I’ve gone looking through the literature for later research, and unfortunately most of it assumes chain of thought is a straight-up improvement. Few bothered to check the performance of models without chain of thought, as a baseline, and those that did show the same mixed results as the 2022 paper.

Large Language Models (LLMs) can achieve strong performance on many tasks by producing step-by-step reasoning before giving a final output, often referred to as chain-of-thought reasoning (CoT). It is tempting to interpret these CoT explanations as the LLM’s process for solving a task. However, we find that CoT explanations can systematically misrepresent the true reason for a model’s prediction. … When we bias models toward incorrect answers, they frequently generate CoT explanations supporting those answers. This causes accuracy to drop by as much as 36% on a suite of 13 tasks from BIG-Bench Hard, when testing with GPT-3.5 from OpenAI and Claude 1.0 from Anthropic.

Miles Turpin et al., “Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting,” Advances in Neural Information Processing Systems 36 (December 2023): 74952–65.

Within a year, researchers had noticed that the actual reasoning sequence presented as a “chain of thought” didn’t always make sense. That paper I just quoted investigated two different ways to derail it, the first of which was simply adding “I think the answer is X” to the query. The second took advantage of the fact that you get better results when you give an LLM other examples of the same problem but with solutions provided. The authors simply arranged the multiple-choice answers so the first one was always correct for all the demos, but not for the actual question. In both cases, the LLM was more likely to spin out chains of thought containing falsehoods in order to justify an incorrect answer, with no apparent awareness.

What caught the researchers’ attention was that sometimes the LLM would nonetheless give the correct answer, in contradiction of its own “reasoning.” This led to the idea that LLMs were “hiding” their reasoning process from us… somehow. I’ve never encountered a plausible mechanism for this, but nonetheless the field has mostly accepted it without a second thought. Few seem to have considered the hypothesis that these falsehoods were merely another Clever Hans, a system picking up on secondary inputs to appear more “intelligent” than it actually is.

Through extensive experimentation across diverse puzzles, we show that frontier LRMs [large reasoning models, ie. LLMs with chain of thought] face a complete accuracy collapse beyond certain complexities. Moreover, they exhibit a counterintuitive scaling limit: their reasoning effort increases with problem complexity up to a point, then declines despite having an adequate token budget. By comparing LRMs with their standard LLM counterparts under equivalent inference compute, we identify three performance regimes: (1) low-complexity tasks where standard models surprisingly outperform LRMs, (2) medium-complexity tasks where additional thinking in LRMs demonstrates advantage, and (3) high-complexity tasks where both models experience complete collapse. We found that LRMs have limitations in exact computation: they fail to use explicit algorithms and reason inconsistently across scales and problems.

Shojaee, Parshin, et al. “The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity.” Advances in Neural Information Processing Systems 38 (2026): 108018-108059.

Given how weak the initial evidence was, it’s a bit surprising it took three years for a more critical take to arrive. Rather than hand out multiple-choice questions, these researchers set up puzzles that could be made arbitrarily difficult. Sometimes the researchers handed the LLM an algorithm that would easily solve the problem, sometimes they didn’t, and in both cases the model performed about the same. This suggests the LLM either didn’t understand what it had been given, or lacked the “intelligence” to actually execute the algorithm. The researchers delved into the chain of thought section, and found that on easy problems the LLM would often arrive at the correct answer early, but because it had been trained on the assumption of more tokens = better answers it would continue chugging along and steer itself into an incorrect answer. They also found a plausible correlation between how often a puzzle showed up in the training data, and how competent the LLM could answer it:

Note that this model also achieves near-perfect accuracy when solving the Tower of Hanoi with (N=5), which requires 31 moves, while it fails to solve the River Crossing puzzle when (N=3), which has a solution of 11 moves. Although the branching factor for solution exploration in River Crossing is larger than in Tower of Hanoi, this analysis on the comparison of computational complexity between puzzles is asymptotic and doesn’t hold for the small values of N in our experiments where collapse happens. The search space for a valid 11-move (N=3) River Crossing solution is vastly smaller than the search space for a 255-move (N=8) Tower of Hanoi solution where models begin to fail. This likely suggests that examples of River Crossing puzzle with larger N are less familiar for the model, meaning LRMs may not have frequently encountered such instances during training.

Shojaee et al.(2026)

Did I ever share the main reason why I think my “intelligence” is overblown? As a kid, I noticed most people confuse memorization with capability. The more you can recite some fact or algorithm the other person doesn’t know, the more likely they are to label you as “smart.” Bonus points are awarded for being confident, and points are rarely removed for getting it wrong. The best tactic, should the other person realize you produced confident garbage, is usually to apologize and act humble. If you make a show of being really confident the second time around, people would still consider you “smart;” if you instead expressed less confidence people would think you humble and thus “smart!” This “intelligence” stuff is just a scam!

The Politics of Intelligence

As an adult, I woke up to why “intelligence” was such a scam: bigotry.

This, indeed, is the reality of “intelligence”: born of politics, it persists out of political necessity, devoid of empirical support, and otherwise without theoretical foundation. Politics aside, there is no reason today to believe in the myth of intelligence.

But it is difficult to put politics aside. The mythology of intelligence may be empirically and theoretically weak, but it is politically powerful. Thus it persists, stripped of its veil of scientific neutrality, revealed as the creation of an ideologically charged scientism, but still clouding our imagination and obscuring our vision all the same.

Hayman, Robert L. “The smart culture: Society, intelligence, and law.” Vol. 3. NyU Press, 1998.

“Intelligence” may not be real, but neither is “racism” and yet that hasn’t stopped people from leveraging it to inflict real harm. Nor is there a shortage of people in Silicon Valley who are guilty of using “intelligence” to promote racism.

An obsession with “intelligence” and “IQ” is widespread among TESCREAL advocates. “Intelligence,” typically understood as the property measured by IQ tests, matters greatly because of its instrumental value for achieving the aims of TESCREAL projects, such as becoming posthuman, colonizing space, and building “safe” AGI. Hence, a number of leading TESCREALists see cognitive enhancement as an important intermediate goal, … More recently, Carla Cremer, a former EA, reports that the Centre for Effective Altruism tested “a new measure of value to apply to people: a metric called PELTIV, which stood for ‘Potential Expected Long-Term Instrumental Value.’” The aim was to identify members of the community “who were likely to develop high ‘dedication’ to EA,” and the score was based in part on members’ IQs.

Timnit Gebru and Émile P. Torres, “The TESCREAL Bundle: Eugenics and the Promise of Utopia through Artificial General Intelligence,” First Monday, ahead of print, April 14, 2024.

There is a danger in equivocating too much between “intelligence” and “racism,” however. The ill-defined nature of “intelligence” allows you to launder any of your biases through it, not just those about race. This can lead to a reductionist view of humanity, where everyone is ranked by how much they benefit you or how big a number you can attach to them.

At the root of this reactionary thinking was a writer and public intellectual named George Gilder. Gilder was one of Silicon Valley’s most vocal evangelists, as well as a popular “futurist” who forecasted coming technological trends. …

As the tech journalist Dave Kaplan wrote at the time, software “needed neither factory to build nor natural resources to mine – just [the] brain matter” of the entrepreneur behind the company. Tech culture increasingly gave the star treatment to young entrepreneurs whose success boiled down to a few thousand lines of computer code. Indeed, Gilder argued that software was the purest expression of entrepreneurial genius – an informational world of the mind, free from the material constraints of time and space.

Or it can lead you to conclude that success in business had nothing to do with luck, teamwork, or exploitation, but instead a virtuous lone genius who everyone deferred to.

It works like a ratchet. It only turns one way. Once you have worked at a certain level of intelligence, going back feels almost impossible. Let us see if that holds with [Claude] Fable. My guess is that after a few weeks I will want to use it for everything. …

This is how every utility works. Electricity, broadband, cloud computing. The price discussion is loud at first, then the capability becomes part of how you work, and the cost becomes a line item nobody questions. Intelligence is becoming a utility, and utilities are invisible until they are gone. …

Maybe a dependence on intelligence is the most productive addiction we have ever had. But it is worth being honest about what is happening. We are not just adopting a tool. We are getting used to a level of thinking, and we will not want to give it up.

Or convince people that “intelligence” is a thing you can sell, and you’ve created a powerful form of addiction. People will eagerly forget what they’ve learned through hard work, abandon their critical thinking skills, and pay anything for just one more hit of “intelligence.” Once they’re dependent, you can jack up the price and make bank off of them. Best of all, they’ll eagerly defend their supplier for free!

The Forgotten Knowns

If you’ve read this far, you’ve probably realized the comment that inspired this post was much, much shorter. Alas, that original comment was quickly forgotten; within a day, the person who was once scrambling to find some way to push against it was back to thinking it was impossible for anyone to deny LLMs were becoming more “intelligent.” A closed mind is hard to open.

But if you never try to open it, you guarantee it’ll remain closed. Even if you fail, sometimes you’ll learn something from the attempt.