Study finds Gemini 3.1 Flash-Lite deficit recurs on fresh data under historical configuration
A new arXiv paper examines how behavioural evaluations of hosted language models can vary due to differences in the service, measurement instrument, or both. The study uses Regent Chess, a sequential environment with a hidden, mutable state, to evaluate model beliefs against ground truth at action time. It finds that a previously reported deficit for Gemini 3.1 Flash-Lite recurs on fresh games…
Key points
- Gemini 3.1 Flash-Lite deficit recurs on fresh data: +0.0530, 95% CI [+0.0329,+0.0714]
- Study separates replication, measurement sensitivity, and persistence in model evaluation
- Findings motivate explicit indexing of hosted-model claims by identifier, serving period, instrument, and configuration
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Researchers propose multi-split boundary decision to lower LLM document segmentation cost · 1 src
- GaitVista reduces gait measurement error by 27.7% in lab tests · 1 src
- Megagon Labs releases mawile workbench for auditing LLM judges · 1 src
- AutoGym framework generates verifiable agent gyms · 1 src
- EvidenT improves enterprise assistant evidence verification by 29% · 1 src
Comments
via GitHub Discussions