DEEPO improves hallucination in multimodal large language models
DEEPO is a new reinforcement‑learning enhancement for multimodal large language models that targets hallucination. The authors identify two weak points in the correction chain: high‑entropy queries collapse advantage at rollout, and confident‑but‑wrong tokens become gradient‑invisible at optimization. They propose a dual‑stage approach that combines signal‑variance regularization with gradient…
Key points
- DEEPO combines signal variance regularization and gradient preconditioning to reduce hallucination in multimodal large language models.
- The method improves over GRPO individually and shows a statistically significant +4.0$ improvement on VideoMMMU with 95% CI [1.1, 6.9].
- DEEPO addresses two weak points: high‑entropy queries collapse advantage and confident‑but‑wrong tokens become gradient‑invisible.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Researchers release COILD corpus for Indian language machine translation · 1 src
- Research finds script knowledge in LLMs emerges only in final layers · 1 src
- Researchers test how language models handle numerical formats in word problems · 2 src
- PTC-Bias improves speech LLM accuracy with phoneme-level bias correction · 1 src
- Researchers benchmark how LLMs handle political character attacks · 1 src
Comments
via GitHub Discussions