{"version":1,"type":"story","url":"https://digestai.news/story/microsoft-study-finds-ai-models-score-27-on-long-term-decision-tasks","json":"https://digestai.news/story/microsoft-study-finds-ai-models-score-27-on-long-term-decision-tasks.json","markdown":"https://digestai.news/story/microsoft-study-finds-ai-models-score-27-on-long-term-decision-tasks.md","slug":"microsoft-study-finds-ai-models-score-27-on-long-term-decision-tasks","headline":"Microsoft study finds AI models score 27% on long-term decision tasks","summary":"Microsoft researchers published a paper showing AI systems struggle with long-term decision-making. In a simulated year-long task requiring interconnected choices, models reached only 27% of human performance. The best-performing setup, Qwen3.7-Max with Hermes, still lagged far behind human benchmarks, highlighting persistent reliability gaps in extended planning.\n\nThe findings, shared by Rohan Paul on social media, raise questions about industry progress. While unnamed market observers link the results to Anthropic’s November 2026 timeline for a top-tier model, the study itself does not name competitors or speculate on pricing. Google and Meta’s upcoming model releases could further shape perceptions of AI capabilities in this area.","keyPoints":["Microsoft’s study shows AI models achieve 27% of human performance in long-term decision tasks over a year","Qwen3.7-Max with Hermes was the top-performing setup but still underperformed human benchmarks significantly","Market observers speculate findings may impact Anthropic’s competitive position ahead of November 2026"],"whyItMatters":"The study underscores a critical weakness in AI’s ability to handle sustained, self-correcting decision-making, a core requirement for real-world applications like finance, healthcare, and logistics.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["Microsoft","Anthropic","Google","Meta"],"models":["Qwen3.7-Max","Hermes"],"people":["Rohan Paul","Dario Amodei","Sundar Pichai"]},"firstPublishedAt":"2026-09-30T17:00:00Z","updatedAt":"2026-09-30T17:00:00Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"cryptobriefing.com","title":"Microsoft study reveals AI struggles with long-term decision-making tasks","url":"https://cryptobriefing.com/microsoft-study-reveals-ai-struggles-with-long-term-decision-making-tasks","publishedAt":"2026-09-30T17:00:00Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Microsoft study finds AI models score 27% on long-term decision tasks\", 30 September 2026, https://digestai.news/story/microsoft-study-finds-ai-models-score-27-on-long-term-decision-tasks","publisher":"Digest AI","title":"Microsoft study finds AI models score 27% on long-term decision tasks","datePublished":"2026-09-30T17:00:00Z","url":"https://digestai.news/story/microsoft-study-finds-ai-models-score-27-on-long-term-decision-tasks"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}