DigestAI news desk
OpenAI board member warns company is not on track to prevent catastrophic AI loss of control OpenAI launches Agents API beta for long-running cloud agents OpenAI solves Navier-Stokes problem, sparking academic controversy over data use RTK Token Savings Debunked: Cost Benchmarks Disagree The Waymo effect: AI making research less collaborative Meta’s Muse AI Agent Seeks User Trust with Secure Architecture Thelio Mira AI Linux Workstation with 192 GB GPU Memory Claude No Longer Available to Minors
Generative AI & Models updated 3 min read

GPT-6 Pro Tops ChessBench with 2,340 Elo Rating

ChessBench, a platform for benchmarking AI models, has released new results. GPT-6 Pro (Max) from OpenAI has been added to the platform with a benchmark Elo rating of 2,340, placing it at rank 10. This is the highest Elo rating among the 30 models currently on ChessBench. The platform also includes other notable models like Gemini 3.8 Flash (medium), Gemini 3.7 Flash (High), and Grok 4.6…

1 source

Key points

  • GPT-6 Pro (Max) from OpenAI has a 2,340 Elo rating
  • It ranks 10th among 30 models on ChessBench
  • ChessBench now includes models from seven different providers
Full story from chessbench-ai.github.io · via r/singularity Open source ↗

GPT-6 Astra 2,340 Elo on ChessBench - Ranked #11

chessbench-ai.github.io · 11 September 2026

Complete games. Compare playing strength across the field.

Leaderboard

No models found

Try a different model name or provider.

What the numbers mean

A field-relative measure of model performance, not a human chess rating.

  • Complete games
  • Models play full games. Planning, tactical calculation, recovery, and consistency all affect the result.
  • Benchmark Elo
  • A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.
  • Cost per Task
  • The reported cost in US dollars per task. A dash (—) means no cost has been reported, not that the task was free. Select the column heading to sort by cost; missing values remain at the end.

ChessBench TV

A complete game, one move at a time.

Gemini 3.1 ProGoogleBlack

GPT-5.6 Sol ProOpenAIWhite to move

Focus the replay to use to step and Space to play.

GPT-5.6 Sol Pro

vs. Gemini 3.1 Pro

Finished · Black resigned

Benchmark updates

New results, corrections, and product notes.

### Gemini 3.8 Flash (medium) joins ChessBench

Gemini 3.8 Flash (medium) has been added with 2,300 benchmark Elo, 30.74 ACPL, and a cost of $0.08/task. It enters at rank 13, after Gemini 3.6 Flash, which also has 2,300 benchmark Elo, and ahead of GPT-4.5.

View leaderboard

### Claude Fable 5.1 (Max) joins ChessBench; Cost per Task added

Claude Fable 5.1 (Max) has been added as a separate model with 2,370 benchmark Elo, 18.54 ACPL, and a cost of $44.46/task. It enters at rank 10, between Gemma 4 and GPT-6 Pro (Max).

The leaderboard now includes Cost per Task. Costs not yet reported are shown as a dash (—). ACPL and cost are also available in model details.

View leaderboard

### GPT-6 Pro (Max) joins ChessBench

GPT-6 Pro (Max) has been added to ChessBench with a benchmark Elo of 2,340. It enters at rank 10, between Gemma 4 and Gemini 3.6 Flash.

The leaderboard now includes 30 models across seven providers.

View leaderboard

### Gemini 3.7 Flash (High) joins ChessBench

Gemini 3.7 Flash (High) has been added to ChessBench with a benchmark ELO of 2,230 It enters at rank 17, between MiniMax M3 and ChatGPT 5.4 Pro.

The leaderboard consolidates GPT-5.6 Sol Pro (Max) and its Stealth Testing checkpoint into a single rank-one entry, with the checkpoint’s 2,650 ELO shown as secondary information. Claude Fable 5 (Max) therefore moves to rank 2.

View leaderboard

### Grok 4.6 (xhigh) joins ChessBench

Grok 4.6 (xhigh) has been added to ChessBench with a benchmark ELO of 2,480 The result places it at rank 7, between Qwen 3.8 Max and Claude Opus 5 (high), and makes it the strongest xAI model in the current field.

View leaderboard

### Qwen 3.8 Max joins ChessBench

### Claude Opus 5 (high) joins ChessBench

Claude Opus 5 (high) enters at 2,400 benchmark ELO. Compared to the original Opus 5 configuration, this represents a substantial improvement in playing strength, moving the entry closer to the current frontier.

View leaderboard

### Gemini Flash additions and a Grok 4.5 correction

Gemini 3.6 Flash joins at 2,300 benchmark ELO, alongside Gemini 3.5 Flash-Lite at 2,200 benchmark ELO. Flash-Lite’s benchmark ELO is lower because illegal moves and several quick terminations affected game outcomes.

Grok 4.5 was rebenchmarked after errors were found in the original evaluation. Its benchmark ELO changes from 1,457 to 1,680.

View leaderboard

### Claude Fable 5 (Max) and Kimi K3 (Max) join ChessBench

Claude Fable 5 (Max) from Anthropic and Kimi K3 (Max) from Moonshot AI join the leading group. Claude Fable 5 (Max) holds rank 2 with 2,640 benchmark ELO, while Kimi K3 (Max) holds rank 4 with 2,590 benchmark ELO.

The additions broaden the provider mix near the top of the table and add two high-strength reference points to the leaderboard.

View leaderboard

### Introducing ChessBench TV

ChessBench TV replays complete model-vs-model games on an interactive board, with the full move score and playback controls. The first replay features GPT-5.6 Sol Pro against Gemini 3.1 Pro.

This text was published by chessbench-ai.github.io . It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Coverage and discussion

1 source
Topics · follow one to build your own front page

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

Related stories