arXiv study finds PRM-Pruned Fragment Grafting shows no benefit in reasoning tasks
A new paper on arXiv examines PRM-Pruned Fragment Grafting (PPFG), an inference-time method that extracts high-scoring prefixes from pruned chains and grafts them into sibling chains. Researchers tested PPFG on Qwen2.5-7B-Instruct with Math-Shepherd across the full MATH500 dataset (n=500, three seeds) and found no statistically significant improvement over independent parallel chain-of-thought…
Key points
- PPFG shows no benefit over independent parallel-CoT on Qwen2.5-7B-Instruct with Math-Shepherd on MATH500 (n=500, three seeds)
- Only 14% of PPFG injections targeted genuinely struggling chains, rest were redundant or ineffective
- Random targeting matches independent CoT at 2.4x the firing rate, with no population-level compensation
The study analyzed 322 stagnation-rule injection events and found only 14% targeted genuinely struggling chains. The rest were applied to already successful chains, near-completion chains, or those on a flat PRM plateau. Even with random targeting, PPFG did not outperform independent CoT. The findings replicate across three base LMs, six benchmarks, and a second PRM, with equivalence testing confirming no meaningful gain. The paper concludes that PPFG’s inertness is not heuristic-specific and contributes an equivalence-testing template for future research.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model · 1 src
- Korean legal study finds KLUE-BERT outperforms GPT models in sexual offense text classification · 1 src
- Researchers question human-derived bias measures for LLM evaluation · 1 src
- arXiv study finds reading LLM judges from first token overstates position bias · 1 src
- Study compares On-Device NER models for speed, cost and accuracy · 1 src
Comments
via GitHub Discussions