@theo: With LLMs, we like to think of "smart" and "dumb" as one axis (because we think of humans this way). I'd like to argue …
Summary
The author argues against viewing LLM intelligence on a single axis, proposing separate axes for 'smart' and 'dumb' traits, with examples like Gemini, Astra, and Fable 5.1 illustrating how models can exhibit both simultaneously.
View Cached Full Text
Cached at: 09/27/26, 09:15 AM
With LLMs, we like to think of “smart” and “dumb” as one axis (because we think of humans this way).
I’d like to argue against this framing. Instead, try to think of “smart” and “dumb” as two different axes. A model can be incredibly smart AND dumb at the same time.
For a good example of this, look at a Gemini model. They are incredibly smart: you can bench them and see the capability. The amount of knowledge Google bakes into their models is incredible. Yet when you ask it to do work, the amount of stupid things the models will do is similarly incredible. They are “smart” and “dumb” at the same time.
Similarly, look at a model like Fable 5.1. It is not quite as smart as Astra, but it is significantly less dumb. Astra is still the smartest model available today, but it is also one of the dumbest, regularly doing things that make literally no sense whatsoever.
Finding the right balance of both “smart” and “dumb” is a challenge that everyone needs to figure out, from users to researchers. If we stop thinking of it as either-or and instead realize both can exist together, understanding model behavior gets much easier.
Similar Articles
Fast, Slow, and Tool-augmented Thinking for LLMs: A Review
This paper presents a taxonomy of LLM reasoning strategies along two orthogonal axes: fast vs. slow thinking and internal vs. external knowledge, and surveys recent adaptive reasoning methods.
Why your local LLM feels dumber than it is
The article explores why locally run large language models might seem less intelligent, addressing potential performance or perception issues.
@DailyDoseOfDS_: A harnessed LLM agent, clearly explained! Most people picture this as a model with tools bolted on. The real architectu…
Explains the inverted architecture of a harnessed LLM agent, where intelligence is externalized into memory, skills, and protocols around a thin model core, with mediators governing interactions.
@paulg: Using "think" to describe what an LLM does reminds me of the 16th century, when astronomers didn't really believe in th…
The tweet draws an analogy between the historical use of the heliocentric model for calculations despite initial disbelief and the current use of the word 'think' for LLMs, suggesting a gradual acceptance of AI cognition.
Seeing how differently people prompt LLMs is funny
The author humorously contrasts his precise prompting style for a local GLM 5.3 Flash model with his brother's berating approach to GPT-6-Astra for coding tasks, highlighting different LLM interaction styles.