GPT-6.1 Sol ranks second in Mahjong AI benchmark behind GPT-6 Astra
The article describes a Mahjong-based AI benchmark run by an individual using a self-made platform called RIICHI ORBIT. Six models were tested: GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, GPT-6 Luna, Claude Fable 5.1, and Claude Opus 5.5. GPT-6 Astra placed first with an average rank of 2.4105, followed by GPT-6.1 Sol at 2.4350. The benchmark involved 23,040 half-games across 15 combinations of four…
Key points
- GPT-6 Astra ranked first with average rank 2.4105 in Mahjong AI benchmark
- GPT-6.1 Sol ranked second with average rank 2.4350 and highest first-place rate at 27.3%
- Benchmark used 23,040 half-games across 15 model combinations with fixed seeds and seating
Model pages: GPT-6 Astra → · GPT-6.1 Sol → · GPT-6 Sol →
The story so far
3 episodes →- GPT-6.1 Sol ranks second in Mahjong AI benchmark behind GPT-6 Astrathis story
GPT-6.1 Sol Joins the Fray! Evaluating AI via Mahjong
note.com · 30 September 2026
Loading the full article…
This text was published by note.com and written by Doui Lab. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Mirror-Score benchmarks D-peptide design tools against real-world affinity · 1 src
- Researchers introduce coffee framework for discrete diffusion model guidance · 2 src
- Researchers propose NashEval for context-dependent AI agent evaluation · 2 src
- Researchers test whether AI harnesses specialize or just repeat answers · 1 src
- Study finds synthetic embeddings match text for LLM fine-tuning · 1 src
Comments
via GitHub Discussions