Google may be testing Gemini 4 Pro checkpoints, rumor says
Rumors are circulating that a set of checkpoints labeled “gemini-3.8-flash” has appeared on the LMSYS Arena platform, and some testers claim the model can output up to 256K tokens, generate precise SVG code and even build a 3‑D voxel scene. The same leaked data is said to have achieved a 2,064 Elo score on the GDPval‑AA v2 benchmark, apparently surpassing OpenAI’s upcoming GPT‑6 Astra and…
Key points
- Leaked checkpoints labeled “gemini-3.8-flash” appeared on the Arena platform, showing 256K token outputs and 2,064 Elo on GDPval‑AA v2
- Benchmark tables claim the model outperforms GPT‑6 Astra and Claude Fable, but figures are inconsistent and unverified by Google or Arena
- Analysts say the leak is probably an early test; an official Gemini 4 Pro announcement is expected between October and year‑end
The analysis points out three reasons to treat the numbers cautiously: the demos shown on X are cherry‑picked examples, the circulating benchmark tables contain inconsistent reference values and do not match any official Arena leaderboard, and Google DeepMind has not confirmed the results. The article also highlights Google’s ongoing work on Recursive Self‑Improvement (RSI) pipelines, which could enable autonomous data generation and longer‑context reasoning.
Overall, the author concludes the checkpoints are likely early internal tests, but an official Gemini 4 Pro announcement or API preview is expected sometime between October and year‑end. Readers are advised to wait for a formal paper and verified scores before drawing firm conclusions.
Model page: GPT-6 (Astra) →
[Urgent Analysis] Has Gemini 4 Pro Descended onto the Arena? The Truth Behind the "gemini-3.8-flash" Fake Leak and the Credibility of Surpassing GPT-6 Astra
note.com · 26 September 2026
Loading the full article…
This text was published by note.com and written by pisuke. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Generative AI & Models
All →- GPT-6 Astra and Claude Fable 5.1 compete in end‑to‑end software build · 2 src
- Claude Opus 5.5 tops benchmark over OpenAI’s Astra and Fable 5.1 · 6 src
- Anthropic releases Claude Opus 5.5 with 40% lower costs · 18 src
- OpenAI and Anthropic cut API prices for GPT-6 Sol and Claude Opus 5.5 on September 22 · 10 src
- Supersonic Labs releases Julia 1, a 144.3M-parameter open decision model · 2 src
Comments
via GitHub Discussions