DigestAI news desk

Cut through the AI noise.

Generative AI & Models

Researchers unveil Ovis-Embedding, a universal omni-modal embedding family

A team of researchers announced Ovis-Embedding, an omni‑modal embedding family that encodes text, images, video and audio in a single representation space. The system builds on a pretrained Qwen‑omni backbone, which is adapted through low‑rank contrastive training rather than separate modality towers.

1 source primary source

Key points

  • Ovis-Embedding uses a shared Qwen‑omni backbone with low‑rank contrastive training.
  • Training data covers text, images, video and audio, using homogeneous‑source sampling and focal loss.
  • Achieves state‑of‑the‑art scores on MMEB‑v3, MMEB‑v2, MVEB, MAEB and RTEB benchmarks.

The authors describe three main advances: native omni‑modal initialization using Qwen‑omni, a data‑centric training corpus that spans all four modalities with homogeneous‑source sampling for task‑consistent batches, and embedding‑specific optimization that combines focal loss with similarity‑based embedding distillation. At inference, low‑rank feature decomposition yields compact embeddings with flexible dimensionality and minimal loss. Empirical results show state‑of‑the‑art performance on the MMEB‑v3, MMEB‑v2, MVEB, MAEB and RTEB benchmarks, demonstrating the model’s ability to support any‑to‑any retrieval across modalities.

Read the original at arXiv cs.AI · by Embedding Team primary sourceOpen source ↗
Topics · follow one to build your own front page
Ovis-EmbeddingQwen-omni

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Generative AI & Models

All →

Related stories