DigestAI news desk

Cut through the AI noise.

Agents & Tools

Researchers introduce XiangqiBench to evaluate closed-loop performance of LLM agents in Chinese chess

The paper presents XiangqiBench, an executable benchmark that tests large language model (LLM) agents in Chinese chess endgames. It uses 119 tactical positions with forced mates, supported by engine or checks‑only search, and requires an LLM to deliver checkmate against an engine defender.

1 source primary source

Key points

  • XiangqiBench evaluates LLM agents on 119 Chinese‑chess endgames requiring checkmate against an engine.
  • 12 frontier LLMs generated 8,568 multi‑turn trajectories, exposing Conversion, Consistency, and Simulation gaps.
  • Conversion Gap: 26.1 % first‑move match rate, but only 13.9 % of those trials result in a win.

The authors recorded 8,568 multi‑turn trajectories from 12 frontier LLMs under two observation protocols via an interactive REPL that separates real moves, state queries, and forward simulation. Their analysis reveals three gaps: the Conversion Gap (models play the stored reference first move in 26.1 % of sighted trials but only 13.9 % of those end in a win), the Consistency Gap (the leading model reaches 38.7 % pass@3 yet only 5.9 % pass³, winning all three trials on 7 of the 46 positions it ever wins), and the Simulation Gap (32.3 % of accepted simulation calls stop on an illegal move, and in 49.3 % of comparable cases the real defender replies differently from the simulated line). The authors argue that agent evaluations should report closed‑loop outcomes and reliability alongside coverage.

Read the original at arXiv cs.CL · by Yekun Chai, Qiwei Peng, Haoyi Xiong primary sourceOpen source ↗

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Agents & Tools

All →

Related stories