{"version":1,"type":"story","url":"https://digestai.news/story/deepo-improves-hallucination-in-multimodal-large-language-models","json":"https://digestai.news/story/deepo-improves-hallucination-in-multimodal-large-language-models.json","markdown":"https://digestai.news/story/deepo-improves-hallucination-in-multimodal-large-language-models.md","slug":"deepo-improves-hallucination-in-multimodal-large-language-models","headline":"DEEPO improves hallucination in multimodal large language models","summary":"DEEPO is a new reinforcement‑learning enhancement for multimodal large language models that targets hallucination. The authors identify two weak points in the correction chain: high‑entropy queries collapse advantage at rollout, and confident‑but‑wrong tokens become gradient‑invisible at optimization. They propose a dual‑stage approach that combines signal‑variance regularization with gradient preconditioning. Semantic‑entropy‑triggered expert prefixes inject grounded continuations on high‑uncertainty queries, while Renyi preconditioning counteracts logit‑level saturation to reach confident errors. The authors report that both branches improve over GRPO individually and that their interaction is statistically significant on VideoMMMU, the most complex long‑horizon task in their suite, with a +4.0$ improvement and 95% CI [1.1, 6.9]. DEEPO is said to reduce hallucination while preserving accuracy and training stability.","keyPoints":["DEEPO combines signal variance regularization and gradient preconditioning to reduce hallucination in multimodal large language models.","The method improves over GRPO individually and shows a statistically significant +4.0$ improvement on VideoMMMU with 95% CI [1.1, 6.9].","DEEPO addresses two weak points: high‑entropy queries collapse advantage and confident‑but‑wrong tokens become gradient‑invisible."],"whyItMatters":"Reducing hallucination improves reliability of multimodal LLMs, benefiting applications that rely on accurate multimodal reasoning such as AI assistants, content generation, and decision support systems.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":["DEEPO","GRPO","VideoMMMU"],"people":[]},"firstPublishedAt":"2026-09-25T04:00:00Z","updatedAt":"2026-09-25T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.AI","title":"DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs","url":"https://arxiv.org/abs/2609.28570","publishedAt":"2026-09-25T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"DEEPO improves hallucination in multimodal large language models\", 25 September 2026, https://digestai.news/story/deepo-improves-hallucination-in-multimodal-large-language-models","publisher":"Digest AI","title":"DEEPO improves hallucination in multimodal large language models","datePublished":"2026-09-25T04:00:00Z","url":"https://digestai.news/story/deepo-improves-hallucination-in-multimodal-large-language-models"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}