DigestAI news desk

Cut through the AI noise.

Research

Researchers release ArgGYM benchmark for testing defeasible reasoning in AI models

Researchers introduced ArgGYM, a new benchmark designed to evaluate structured defeasible reasoning in AI models. The framework breaks down reasoning into twelve tasks and uses a symbolic argumentation engine to verify outputs. It includes a frozen dataset of 1,440 verified instances across fifteen configurations, with two argument preference orderings and two set orderings. The benchmark…

1 source primary source

Key points

  • ArgGYM benchmark tests structured defeasible reasoning across 12 tasks with 1,440 verified instances
  • Uses symbolic argumentation engine for formal state evaluation and dynamic instance generation
  • Frontier models recover partial answers but struggle with longer dependencies and complex structures

Frontier and open-weight models exhibit distinct reasoning patterns, recovering partial answers without solving tasks entirely. Performance drops in later configurations with longer dependencies and complex structures. The authors release the benchmark, generators, and verifiers for reproducibility and reinforcement learning with verifiable rewards.

Read the original at arXiv cs.AI · by \.Ibrahim Ethem Deveci, Funda Tan \c{C}al{\i}k, Bar{\i}\c{s} Deniz Sa\u{g}lam, Duygu Ataman primary sourceOpen source ↗

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories