{"version":1,"type":"story","url":"https://digestai.news/story/gpt-6-1-sol-ranks-second-in-mahjong-ai-benchmark-behind-gpt-6-astra","json":"https://digestai.news/story/gpt-6-1-sol-ranks-second-in-mahjong-ai-benchmark-behind-gpt-6-astra.json","markdown":"https://digestai.news/story/gpt-6-1-sol-ranks-second-in-mahjong-ai-benchmark-behind-gpt-6-astra.md","slug":"gpt-6-1-sol-ranks-second-in-mahjong-ai-benchmark-behind-gpt-6-astra","headline":"GPT-6.1 Sol ranks second in Mahjong AI benchmark behind GPT-6 Astra","summary":"The article describes a Mahjong-based AI benchmark run by an individual using a self-made platform called RIICHI ORBIT. Six models were tested: GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, GPT-6 Luna, Claude Fable 5.1, and Claude Opus 5.5. GPT-6 Astra placed first with an average rank of 2.4105, followed by GPT-6.1 Sol at 2.4350. The benchmark involved 23,040 half-games across 15 combinations of four models selected from six, using fixed seeds and seating orders to reduce bias. Each model had 30 minutes to implement and refine their Mahjong AI, with unlimited practice matches allowed. The author notes that GPT-6.1 Sol achieved the highest first-place rate at 27.3%, despite ranking second overall. The discussion suggests the benchmark effectively measures comprehensive AI capability in imperfect-information settings, though results may not align with official model rankings.","keyPoints":["GPT-6 Astra ranked first with average rank 2.4105 in Mahjong AI benchmark","GPT-6.1 Sol ranked second with average rank 2.4350 and highest first-place rate at 27.3%","Benchmark used 23,040 half-games across 15 model combinations with fixed seeds and seating"],"whyItMatters":"The benchmark highlights how imperfect-information games like Mahjong can reveal nuanced AI capabilities in judgment, adaptation, and strategic trade-offs, offering a complementary view to standard LLM evaluations.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["OpenAI","Anthropic"],"models":["GPT-6 Astra","GPT-6.1 Sol","GPT-6 Sol","GPT-6 Luna","Claude Fable 5.1","Claude Opus 5.5"],"people":[]},"firstPublishedAt":"2026-09-30T00:23:00Z","updatedAt":"2026-09-30T00:23:00Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"note.com","title":"GPT-6.1 Sol Joins the Fray! Evaluating AI via Mahjong","url":"https://note.com/doui_lab/n/n0eb198d80d28?hl=en","publishedAt":"2026-09-30T00:23:00Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":{"title":"AI Model Wars GPT-6 Claude and the Market","url":"https://digestai.news/thread/openai-and-anthropic-costs-drive-firms-toward-open-weight-ai-report-says","storyCount":3},"cite":{"text":"Digest AI, \"GPT-6.1 Sol ranks second in Mahjong AI benchmark behind GPT-6 Astra\", 30 September 2026, https://digestai.news/story/gpt-6-1-sol-ranks-second-in-mahjong-ai-benchmark-behind-gpt-6-astra","publisher":"Digest AI","title":"GPT-6.1 Sol ranks second in Mahjong AI benchmark behind GPT-6 Astra","datePublished":"2026-09-30T00:23:00Z","url":"https://digestai.news/story/gpt-6-1-sol-ranks-second-in-mahjong-ai-benchmark-behind-gpt-6-astra"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}