{"version":1,"type":"story","url":"https://digestai.news/story/researchers-introduce-xiangqibench-to-evaluate-closed-loop-performance","json":"https://digestai.news/story/researchers-introduce-xiangqibench-to-evaluate-closed-loop-performance.json","markdown":"https://digestai.news/story/researchers-introduce-xiangqibench-to-evaluate-closed-loop-performance.md","slug":"researchers-introduce-xiangqibench-to-evaluate-closed-loop-performance","headline":"Researchers introduce XiangqiBench to evaluate closed-loop performance of LLM agents in Chinese chess","summary":"The paper presents XiangqiBench, an executable benchmark that tests large language model (LLM) agents in Chinese chess endgames. It uses 119 tactical positions with forced mates, supported by engine or checks‑only search, and requires an LLM to deliver checkmate against an engine defender.\n\nThe authors recorded 8,568 multi‑turn trajectories from 12 frontier LLMs under two observation protocols via an interactive REPL that separates real moves, state queries, and forward simulation. Their analysis reveals three gaps: the Conversion Gap (models play the stored reference first move in 26.1 % of sighted trials but only 13.9 % of those end in a win), the Consistency Gap (the leading model reaches 38.7 % pass@3 yet only 5.9 % pass³, winning all three trials on 7 of the 46 positions it ever wins), and the Simulation Gap (32.3 % of accepted simulation calls stop on an illegal move, and in 49.3 % of comparable cases the real defender replies differently from the simulated line). The authors argue that agent evaluations should report closed‑loop outcomes and reliability alongside coverage.","keyPoints":["XiangqiBench evaluates LLM agents on 119 Chinese‑chess endgames requiring checkmate against an engine.","12 frontier LLMs generated 8,568 multi‑turn trajectories, exposing Conversion, Consistency, and Simulation gaps.","Conversion Gap: 26.1 % first‑move match rate, but only 13.9 % of those trials result in a win."],"whyItMatters":"Closed‑loop benchmarks reveal how LLM agents perform in interactive tasks, highlighting reliability gaps that static move‑prediction metrics miss.","category":{"slug":"agents","name":"Agents & Tools","url":"https://digestai.news/category/agents"},"entities":{"companies":[],"models":[],"people":[]},"firstPublishedAt":"2026-10-05T04:00:00Z","updatedAt":"2026-10-05T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents","url":"https://arxiv.org/abs/2610.02425","publishedAt":"2026-10-05T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Researchers introduce XiangqiBench to evaluate closed-loop performance of LLM agents in Chinese chess\", 5 October 2026, https://digestai.news/story/researchers-introduce-xiangqibench-to-evaluate-closed-loop-performance","publisher":"Digest AI","title":"Researchers introduce XiangqiBench to evaluate closed-loop performance of LLM agents in Chinese chess","datePublished":"2026-10-05T04:00:00Z","url":"https://digestai.news/story/researchers-introduce-xiangqibench-to-evaluate-closed-loop-performance"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}