AI Forecasting Trials Reveal Model Limits
The series follows a journalist testing various AI assistants on their ability to handle numbers in proposal drafts and make forecasts in flawed scenarios. After the latest episode, it shows that AI models still struggle with numerical reasoning and unreliable predictions.
-
Author tests four AI assistants on flawed forecasting scenarios
A researcher hid four common pitfalls in a forecasting task—leakage, reporting delays, promotion effects, and structural breaks—and asked four AI assistants to solve it. The assistants included…
1 source -
AI models differ in how they handle numbers in proposal drafts
A user compared drafts generated by Claude, ChatGPT, and Gemini for the same business proposal. Each model produced distinct errors in numerical data, such as plausible but unverifiable figures. The…
1 source