Opinion: AlphaGo’s reasoning shows today’s LLMs lack true reasoning
Thore Graepel, a former DeepMind researcher, argues that the creative move that won AlphaGo’s 2016 match against Lee Sedol was not a flash of intuition but the result of a genuine reasoning process that modern large language models (LLMs) still lack. AlphaGo combined a policy network that guessed human‑like moves with a separate search engine that explicitly built and explored a game tree,…
Key points
- AlphaGo used a policy network plus explicit search over a game tree to choose move 37, not pure intuition
- LLMs generate tokens sequentially, lacking a separate, inspectable epistemic state for reasoning
- Graepel proposes AI systems that keep an explicit, evidence‑driven epistemic state similar to AlphaGo’s architecture
Graepel contrasts this with today’s LLMs, which generate the next token repeatedly—a system‑1‑like pattern completion. Even chain‑of‑thought prompting, he notes, merely stretches the same next‑token prediction without creating an independent, inspectable epistemic state. He identifies three shortcomings: no persistent reasoning ledger, no clean separation of knowledge and manipulation, and post‑hoc fabricated reasoning traces. To achieve trustworthy AI in high‑stakes domains, Graepel proposes a new architecture that maintains an explicit epistemic state, updates it only when evidence justifies, and treats reasoning as a sequence of moves that reduce uncertainty, akin to AlphaGo’s game‑tree approach.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model · 1 src
- Korean legal study finds KLUE-BERT outperforms GPT models in sexual offense text classification · 1 src
- Researchers question human-derived bias measures for LLM evaluation · 1 src
- arXiv study finds reading LLM judges from first token overstates position bias · 1 src
- Study compares On-Device NER models for speed, cost and accuracy · 1 src
Comments
via GitHub Discussions