# Researchers introduce XiangqiBench to evaluate closed-loop performance of LLM agents in Chinese chess

Digest AI · Agents & Tools · published 2026-10-05T04:00:00Z

Canonical: https://digestai.news/story/researchers-introduce-xiangqibench-to-evaluate-closed-loop-performance

## Summary

The paper presents XiangqiBench, an executable benchmark that tests large language model (LLM) agents in Chinese chess endgames. It uses 119 tactical positions with forced mates, supported by engine or checks‑only search, and requires an LLM to deliver checkmate against an engine defender.

The authors recorded 8,568 multi‑turn trajectories from 12 frontier LLMs under two observation protocols via an interactive REPL that separates real moves, state queries, and forward simulation. Their analysis reveals three gaps: the Conversion Gap (models play the stored reference first move in 26.1 % of sighted trials but only 13.9 % of those end in a win), the Consistency Gap (the leading model reaches 38.7 % pass@3 yet only 5.9 % pass³, winning all three trials on 7 of the 46 positions it ever wins), and the Simulation Gap (32.3 % of accepted simulation calls stop on an illegal move, and in 49.3 % of comparable cases the real defender replies differently from the simulated line). The authors argue that agent evaluations should report closed‑loop outcomes and reliability alongside coverage.

## Key points

- XiangqiBench evaluates LLM agents on 119 Chinese‑chess endgames requiring checkmate against an engine.
- 12 frontier LLMs generated 8,568 multi‑turn trajectories, exposing Conversion, Consistency, and Simulation gaps.
- Conversion Gap: 26.1 % first‑move match rate, but only 13.9 % of those trials result in a win.

## Why it matters

Closed‑loop benchmarks reveal how LLM agents perform in interactive tasks, highlighting reliability gaps that static move‑prediction metrics miss.

## Sources

1. [Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents](https://arxiv.org/abs/2610.02425) (arXiv cs.CL, 2026-10-05, primary source)

## Cite

Digest AI, "Researchers introduce XiangqiBench to evaluate closed-loop performance of LLM agents in Chinese chess", 5 October 2026, https://digestai.news/story/researchers-introduce-xiangqibench-to-evaluate-closed-loop-performance

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/researchers-introduce-xiangqibench-to-evaluate-closed-loop-performance.json
