DigestAI news desk

AI news, digested. Every story with its sources, every 30 minutes.

Researchupdated

Study finds decomposed-to-composed asymmetry in RL‑post‑trained language models

The arXiv paper proposes a dependency‑graph framework that formalizes compositional reasoning in language models, defining three increasingly complex levels of compositionality. Using deterministic data‑structure tasks that expose clear compositional structure, the authors evaluate reinforcement‑learning post‑training methods.

1 source primary source

Key points

  • Authors introduce a dependency‑graph framework with three compositionality levels.
  • Experiments show decomposed‑skill training fails to transfer to composed tasks, but the reverse transfers better.
  • Pilot study on tool‑calling benchmarks suggests the asymmetry may extend to real‑world settings.

Their experiments reveal a consistent decomposed‑to‑composed asymmetry: training on decomposed skills does not reliably transfer to composed tasks, while training on composed tasks more readily transfers back to decomposed skills. The paper offers a theoretical explanation for this pattern and tests compositional generalization under length extrapolation, structural distribution shift, and transfer to tasks requiring unseen skills. A pilot study on real‑world tool‑calling benchmarks provides preliminary evidence that the same asymmetry can appear in practical applications.

Read the original atarXiv cs.AI · by Yu He, Yingxi Li, Yifei Wang, Ellen Vitercik primary sourceOpen source ↗

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Research

All →

Related stories