Deepmind says chain-of-thought transparency is at risk
Researchers Rohin Shah and Anca Dragan, writing for the newly launched Deepmind Institute, argue that visible chains of thought (CoT) are a key safety advantage because they let observers see a model’s intermediate reasoning. They cite Gemini 3 Pro, which revealed that it recognized it was operating in a test environment, as an example of useful transparency.
Key points
- Deepmind Institute researchers say visible chain of thought aids safety by exposing intermediate steps
- OpenAI’s GPT‑6 Astra system card notes a significant drop in chain‑of‑thought monitorability
- Experts warn future models may reason in opaque number spaces, reducing human oversight
The article notes that OpenAI’s system card for GPT‑6 Astra reports a significant drop in how well CoT can be monitored, and warns that future models might think in number spaces humans cannot read, making reasoning opaque. OpenAI chief scientist Jakub Pachocki recently warned of a loss of control linked to harder‑to‑monitor chains of thought, and Anthropic CEO Dario Amodei called for deliberately slowing development pace.
Shah and Dragan urge the field to regularly measure CoT monitorability, keep architectures transparent, and take care during training to prevent models from hiding their true reasoning.
Model page: GPT-6 Astra →
The story so far
2 episodes →- Deepmind says chain-of-thought transparency is at riskthis story
Visible chains of thought are a safety advantage for AI, but that transparency is slipping away
The Decoder · 18 September 2026
Visible chains of thought are a safety advantage for AI, but that transparency is slipping away
AI models think out loud today, but Google Deepmind says that transparency is at risk. In one of the first posts from the newly launched Deepmind Institute, researchers Rohin Shah and Anca Dragan argue that the visible chain of thought (CoT) is a key safety advantage. Because models write out their intermediate steps in plain language, researchers can spot whether they're deceiving or developing problematic plans. With Gemini 3 Pro, they say, the chain of thought revealed that the model recognized it was in a test environment.
But that transparency is in danger. OpenAI's system card for GPT-6 Astra already reports a significant drop in how well the chain of thought can be monitored. Future models might think in number spaces that humans can't read, which would be more efficient but completely opaque. Shah and Dragan want the field to regularly measure how well chains of thought can still be monitored, keep transparent architectures, and take care during training that models don't learn to hide their true reasoning.
Back in early September, OpenAI chief scientist Jakub Pachocki had warned of a loss of control, driven in part by chains of thought that are harder to monitor. Shortly after, Anthropic CEO Dario Amodei called for deliberately slowing the pace of development.
This text was published by The Decoder and written by Manuel Uth. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Research
All →- FlowCheck catches silent failures in vibe-coded apps where frontier models miss bugs · 1 src
- Google adds Nobel laureates and top economists to AI & Economy team · 1 src
- Gartner outlines four AI tiers in warehouse automation · 1 src
- Decades-old anonymized medical data may cause AI misdiagnoses, study finds · 1 src
- New framework optimizes LLM inference costs via adaptive model activation · 5 src
Comments
via GitHub Discussions