Microsoft study finds AI models score 27% on long-term decision tasks
Microsoft researchers published a paper showing AI systems struggle with long-term decision-making. In a simulated year-long task requiring interconnected choices, models reached only 27% of human performance. The best-performing setup, Qwen3.7-Max with Hermes, still lagged far behind human benchmarks, highlighting persistent reliability gaps in extended planning.
Key points
- Microsoft’s study shows AI models achieve 27% of human performance in long-term decision tasks over a year
- Qwen3.7-Max with Hermes was the top-performing setup but still underperformed human benchmarks significantly
- Market observers speculate findings may impact Anthropic’s competitive position ahead of November 2026
The findings, shared by Rohan Paul on social media, raise questions about industry progress. While unnamed market observers link the results to Anthropic’s November 2026 timeline for a top-tier model, the study itself does not name competitors or speculate on pricing. Google and Meta’s upcoming model releases could further shape perceptions of AI capabilities in this area.
Microsoft study reveals AI struggles with long-term decision-making tasks
cryptobriefing.com · 30 September 2026
Loading the full article…
This text was published by cryptobriefing.com and written by Estefano Gomez. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Study finds frontier AI models outperform junior accountants on specific tasks · 1 src
- Researchers propose heterogeneity metric for LLM team selection · 1 src
- Researchers compose frozen RWKV and Pythia models via shared latent adapter · 1 src
- ContextAdapt evaluates LLMs on value alignment across medicine, law, finance, and national security · 1 src
- Research shows prompt framing shifts LLM political stance · 1 src
Comments
via GitHub Discussions