{"version":1,"type":"story","url":"https://digestai.news/story/researchers-unveil-ovis-embedding-a-universal-omni-modal-embedding-fam","json":"https://digestai.news/story/researchers-unveil-ovis-embedding-a-universal-omni-modal-embedding-fam.json","markdown":"https://digestai.news/story/researchers-unveil-ovis-embedding-a-universal-omni-modal-embedding-fam.md","slug":"researchers-unveil-ovis-embedding-a-universal-omni-modal-embedding-fam","headline":"Researchers unveil Ovis-Embedding, a universal omni-modal embedding family","summary":"A team of researchers announced Ovis-Embedding, an omni‑modal embedding family that encodes text, images, video and audio in a single representation space. The system builds on a pretrained Qwen‑omni backbone, which is adapted through low‑rank contrastive training rather than separate modality towers.\n\nThe authors describe three main advances: native omni‑modal initialization using Qwen‑omni, a data‑centric training corpus that spans all four modalities with homogeneous‑source sampling for task‑consistent batches, and embedding‑specific optimization that combines focal loss with similarity‑based embedding distillation. At inference, low‑rank feature decomposition yields compact embeddings with flexible dimensionality and minimal loss. Empirical results show state‑of‑the‑art performance on the MMEB‑v3, MMEB‑v2, MVEB, MAEB and RTEB benchmarks, demonstrating the model’s ability to support any‑to‑any retrieval across modalities.","keyPoints":["Ovis-Embedding uses a shared Qwen‑omni backbone with low‑rank contrastive training.","Training data covers text, images, video and audio, using homogeneous‑source sampling and focal loss.","Achieves state‑of‑the‑art scores on MMEB‑v3, MMEB‑v2, MVEB, MAEB and RTEB benchmarks."],"whyItMatters":"A unified omni‑modal embedding reduces the need for separate models, enabling more efficient cross‑modal search and retrieval for AI applications.","category":{"slug":"models","name":"Generative AI & Models","url":"https://digestai.news/category/models"},"entities":{"companies":[],"models":["Ovis-Embedding","Qwen-omni"],"people":[]},"firstPublishedAt":"2026-09-23T04:00:00Z","updatedAt":"2026-09-23T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.AI","title":"Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings","url":"https://arxiv.org/abs/2609.25165","publishedAt":"2026-09-23T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Researchers unveil Ovis-Embedding, a universal omni-modal embedding family\", 23 September 2026, https://digestai.news/story/researchers-unveil-ovis-embedding-a-universal-omni-modal-embedding-fam","publisher":"Digest AI","title":"Researchers unveil Ovis-Embedding, a universal omni-modal embedding family","datePublished":"2026-09-23T04:00:00Z","url":"https://digestai.news/story/researchers-unveil-ovis-embedding-a-universal-omni-modal-embedding-fam"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}