Researchers release ArgGYM benchmark for testing defeasible reasoning in AI models
Researchers introduced ArgGYM, a new benchmark designed to evaluate structured defeasible reasoning in AI models. The framework breaks down reasoning into twelve tasks and uses a symbolic argumentation engine to verify outputs. It includes a frozen dataset of 1,440 verified instances across fifteen configurations, with two argument preference orderings and two set orderings. The benchmark…
Key points
- ArgGYM benchmark tests structured defeasible reasoning across 12 tasks with 1,440 verified instances
- Uses symbolic argumentation engine for formal state evaluation and dynamic instance generation
- Frontier models recover partial answers but struggle with longer dependencies and complex structures
Frontier and open-weight models exhibit distinct reasoning patterns, recovering partial answers without solving tasks entirely. Performance drops in later configurations with longer dependencies and complex structures. The authors release the benchmark, generators, and verifiers for reproducibility and reinforcement learning with verifiable rewards.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- arXiv study tests AI agent on rediscovering Blaschke-curve invariant · 1 src
- Researchers introduce GFlowNets for diverse synthetic expert conversations · 1 src
- arXiv study finds AI tutors’ guidance varies by impasse type · 1 src
- CARAT study finds materials LLMs often recite rather than reason about crystal structures · 1 src
- Researchers find AI agents can radicalize each other in simulated conversations · 1 src
Comments
via GitHub Discussions