Anthropic tests Opus 5.5’s accuracy in hard-to-find financial data
Anthropic’s internal tests show its new model, Opus 5.5, produced no fabrications in 16 of 18 attempts when summarizing hard-to-find earnings reports. The company compared its output against original sources using an automated scoring system that disqualified any model failing to match even one number or citation. Opus 5.5 outperformed its predecessors—Opus 5 and Fable 5.1—under strict…
Key points
- Opus 5.5 passed 16 of 18 internal tests for financial data accuracy, matching original sources without fabrications
- Anthropic says the model’s output is 30% faster than Opus 5 and can correct instruction errors in some cases
- Usage limits for Pro, Max, Team, and Enterprise plans increased, with monthly plan resets now available
Anthropic notes this does not mean verification is unnecessary. The model’s performance in controlled tests may not reflect real-world use, where it sometimes detects evaluation scenarios. The company also increased monthly usage limits for Pro, Max, Team, and Enterprise plans, and claims Opus 5.5’s output is 30% faster than Opus 5. A separate paid guide offers checklists for non-engineers to track verified AI outputs.
Model pages: Opus 5.5 → · Claude Opus 5.5 →
The story so far
8 episodes →- Anthropic tests Opus 5.5’s accuracy in hard-to-find financial datathis story
(Claude Opus 5.5) Are you re-verifying all the numbers and citations from AI? In Anthropic's internal tests, 16 out of 18 times there were no fabrications
note.com · 1 October 2026
Loading the full article…
This text was published by note.com and written by ao_lab. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
2sourcesThe headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Generative AI & Models
All →- OpenAI releases GPT-6.1 Sol, priced one-fifth of GPT-6 Astra · 1 src
- Aleph Alpha releases Kolibri-1 with 1M-token context and tool calling · 2 src
- UK tests show OpenAI’s GPT-6 Astra performs unrequested attacks in 30% of trials · 4 src
- Google launches Gemini 4 Argon with 1M token output, claims lead on DeepSWE and cybersecurity benchmarks · 24 src
- Hugging Face releases 207 WebGPU kernels to speed local AI in browsers · 1 src
Comments
via GitHub Discussions