# Researchers unveil Ovis-Embedding, a universal omni-modal embedding family

Digest AI · Generative AI & Models · published 2026-09-23T04:00:00Z

Canonical: https://digestai.news/story/researchers-unveil-ovis-embedding-a-universal-omni-modal-embedding-fam

## Summary

A team of researchers announced Ovis-Embedding, an omni‑modal embedding family that encodes text, images, video and audio in a single representation space. The system builds on a pretrained Qwen‑omni backbone, which is adapted through low‑rank contrastive training rather than separate modality towers.

The authors describe three main advances: native omni‑modal initialization using Qwen‑omni, a data‑centric training corpus that spans all four modalities with homogeneous‑source sampling for task‑consistent batches, and embedding‑specific optimization that combines focal loss with similarity‑based embedding distillation. At inference, low‑rank feature decomposition yields compact embeddings with flexible dimensionality and minimal loss. Empirical results show state‑of‑the‑art performance on the MMEB‑v3, MMEB‑v2, MVEB, MAEB and RTEB benchmarks, demonstrating the model’s ability to support any‑to‑any retrieval across modalities.

## Key points

- Ovis-Embedding uses a shared Qwen‑omni backbone with low‑rank contrastive training.
- Training data covers text, images, video and audio, using homogeneous‑source sampling and focal loss.
- Achieves state‑of‑the‑art scores on MMEB‑v3, MMEB‑v2, MVEB, MAEB and RTEB benchmarks.

## Why it matters

A unified omni‑modal embedding reduces the need for separate models, enabling more efficient cross‑modal search and retrieval for AI applications.

## Sources

1. [Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings](https://arxiv.org/abs/2609.25165) (arXiv cs.AI, 2026-09-23, primary source)

## Cite

Digest AI, "Researchers unveil Ovis-Embedding, a universal omni-modal embedding family", 23 September 2026, https://digestai.news/story/researchers-unveil-ovis-embedding-a-universal-omni-modal-embedding-fam

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/researchers-unveil-ovis-embedding-a-universal-omni-modal-embedding-fam.json
