{"version":1,"type":"story","url":"https://digestai.news/story/symce-corpus-released-for-counterexample-generation","json":"https://digestai.news/story/symce-corpus-released-for-counterexample-generation.json","markdown":"https://digestai.news/story/symce-corpus-released-for-counterexample-generation.md","slug":"symce-corpus-released-for-counterexample-generation","headline":"SymCE corpus released for Counterexample generation","summary":"SymCE is a new dataset of 4,707 false undergraduate‑algebra and real‑analysis conjectures, each paired with an executable Python verifier. The authors train Qwen3‑4B and Gemma‑3‑4B, showing that counterexample‑only supervised fine‑tuning drops true‑theorem recognition from 0.27 to 0.00, while reinforcement learning with a sparse outcome‑only reward restores it to 0.66.\n\nThe 4B model outperforms every evaluated 7B open‑weights math specialist and remains competitive with six frontier commercial APIs. A human audit of 177 verifier decisions finds 97.7% accuracy, and the model transfers under unchanged prompting to GSM8K, MATH‑500 and MMLU‑college‑math.\n\nThese results illustrate how reinforcement learning can repair imitation failures and provide a large, verifiable benchmark for training models to generate counterexamples in theorem proving.","keyPoints":["SymCE: 4,707 false undergraduate‑algebra and real‑analysis conjectures with executable verifiers.","Counterexample‑only SFT drops true‑theorem recognition from 0.27 to 0.00; RLVR with sparse reward restores to 0.66.","4B model outperforms 7B open‑weights math specialists and matches six frontier commercial APIs."],"whyItMatters":"The SymCE corpus provides a large, verifiable dataset for training models to generate counterexamples, a key step toward more reliable theorem proving. The findings show that reinforcement learning can recover from imitation failures, guiding future research on safe fine‑tuning. The benchmark results also benchmark 4B models against larger commercial APIs, informing model scaling decisions.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["Qwen","Gemma"],"models":["Qwen3-4B","Gemma-3-4B"],"people":[]},"firstPublishedAt":"2026-10-05T04:00:00Z","updatedAt":"2026-10-05T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs","url":"https://arxiv.org/abs/2610.02444","publishedAt":"2026-10-05T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":{"title":"Qwen3.8-27B Evolution Token Cuts and RL Tuning","url":"https://digestai.news/thread/bottlecap-ai-cuts-qwen3-8-27bs-reasoning-tokens-by-37-2-with-thinkingcap-qwen3","storyCount":3},"cite":{"text":"Digest AI, \"SymCE corpus released for Counterexample generation\", 5 October 2026, https://digestai.news/story/symce-corpus-released-for-counterexample-generation","publisher":"Digest AI","title":"SymCE corpus released for Counterexample generation","datePublished":"2026-10-05T04:00:00Z","url":"https://digestai.news/story/symce-corpus-released-for-counterexample-generation"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}