DigestAI news desk

Cut through the AI noise.

Research

DEEPO improves hallucination in multimodal large language models

DEEPO is a new reinforcement‑learning enhancement for multimodal large language models that targets hallucination. The authors identify two weak points in the correction chain: high‑entropy queries collapse advantage at rollout, and confident‑but‑wrong tokens become gradient‑invisible at optimization. They propose a dual‑stage approach that combines signal‑variance regularization with gradient…

1 source primary source

Key points

  • DEEPO combines signal variance regularization and gradient preconditioning to reduce hallucination in multimodal large language models.
  • The method improves over GRPO individually and shows a statistically significant +4.0$ improvement on VideoMMMU with 95% CI [1.1, 6.9].
  • DEEPO addresses two weak points: high‑entropy queries collapse advantage and confident‑but‑wrong tokens become gradient‑invisible.
Read the original at arXiv cs.AI · by Yingxuan Zhuang, Miao Pan, Wangjie Gan, Jingxiao Yang, Fan Wang, Weiming Liu, Cheng Tan, Xuhong Zhang, Jintao Chen primary sourceOpen source ↗
Topics · follow one to build your own front page
DEEPOGRPOVideoMMMU

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories