Study finds decomposed-to-composed asymmetry in RL‑post‑trained language models
The arXiv paper proposes a dependency‑graph framework that formalizes compositional reasoning in language models, defining three increasingly complex levels of compositionality. Using deterministic data‑structure tasks that expose clear compositional structure, the authors evaluate reinforcement‑learning post‑training methods.
Key points
- Authors introduce a dependency‑graph framework with three compositionality levels.
- Experiments show decomposed‑skill training fails to transfer to composed tasks, but the reverse transfers better.
- Pilot study on tool‑calling benchmarks suggests the asymmetry may extend to real‑world settings.
Their experiments reveal a consistent decomposed‑to‑composed asymmetry: training on decomposed skills does not reliably transfer to composed tasks, while training on composed tasks more readily transfers back to decomposed skills. The paper offers a theoretical explanation for this pattern and tests compositional generalization under length extrapolation, structural distribution shift, and transfer to tasks requiring unseen skills. A pilot study on real‑world tool‑calling benchmarks provides preliminary evidence that the same asymmetry can appear in practical applications.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Research
All →- New framework optimizes LLM inference costs via adaptive model activation · 4 src
- Qwen3.5-4B outperforms larger LLMs on new user-side conflict benchmark · 1 src
- Neo-Classic benchmark evaluates linguistic-aesthetic reasoning in Classical Chinese poetry · 1 src
- Study finds trust and friction issues in major generative AI app reviews · 1 src
- Study finds PCA can detect stylistic axes in LLM activations without training · 1 src
Comments
via GitHub Discussions