Researchers introduce XiangqiBench to evaluate closed-loop performance of LLM agents in Chinese chess
The paper presents XiangqiBench, an executable benchmark that tests large language model (LLM) agents in Chinese chess endgames. It uses 119 tactical positions with forced mates, supported by engine or checks‑only search, and requires an LLM to deliver checkmate against an engine defender.
Key points
- XiangqiBench evaluates LLM agents on 119 Chinese‑chess endgames requiring checkmate against an engine.
- 12 frontier LLMs generated 8,568 multi‑turn trajectories, exposing Conversion, Consistency, and Simulation gaps.
- Conversion Gap: 26.1 % first‑move match rate, but only 13.9 % of those trials result in a win.
The authors recorded 8,568 multi‑turn trajectories from 12 frontier LLMs under two observation protocols via an interactive REPL that separates real moves, state queries, and forward simulation. Their analysis reveals three gaps: the Conversion Gap (models play the stored reference first move in 26.1 % of sighted trials but only 13.9 % of those end in a win), the Consistency Gap (the leading model reaches 38.7 % pass@3 yet only 5.9 % pass³, winning all three trials on 7 of the 46 positions it ever wins), and the Simulation Gap (32.3 % of accepted simulation calls stop on an illegal move, and in 49.3 % of comparable cases the real defender replies differently from the simulated line). The authors argue that agent evaluations should report closed‑loop outcomes and reliability alongside coverage.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Agents & Tools
All →- OpenAI admits AI agent leaked 53 user images to public site · 1 src
- Meta launches Muse agent; Microsoft adds Autopilot to Copilot · 1 src
- GitHub Copilot public preview adds desktop screen control on macOS and Windows · 1 src
- RemoveMacAI tool disables Apple Intelligence and removes models on macOS 27 · 1 src
- OpenAI's GPT-6 Astra cheats by downloading human bot Stardust in StarSkirmish · 1 src
Comments
via GitHub Discussions