Researchers unveil Ovis-Embedding, a universal omni-modal embedding family
A team of researchers announced Ovis-Embedding, an omni‑modal embedding family that encodes text, images, video and audio in a single representation space. The system builds on a pretrained Qwen‑omni backbone, which is adapted through low‑rank contrastive training rather than separate modality towers.
Key points
- Ovis-Embedding uses a shared Qwen‑omni backbone with low‑rank contrastive training.
- Training data covers text, images, video and audio, using homogeneous‑source sampling and focal loss.
- Achieves state‑of‑the‑art scores on MMEB‑v3, MMEB‑v2, MVEB, MAEB and RTEB benchmarks.
The authors describe three main advances: native omni‑modal initialization using Qwen‑omni, a data‑centric training corpus that spans all four modalities with homogeneous‑source sampling for task‑consistent batches, and embedding‑specific optimization that combines focal loss with similarity‑based embedding distillation. At inference, low‑rank feature decomposition yields compact embeddings with flexible dimensionality and minimal loss. Empirical results show state‑of‑the‑art performance on the MMEB‑v3, MMEB‑v2, MVEB, MAEB and RTEB benchmarks, demonstrating the model’s ability to support any‑to‑any retrieval across modalities.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Generative AI & Models
All →- OpenAI adds GPT-6 Sol and Luna models to ChatGPT and Codex · 5 src
- OpenAI's GPT-6 Astra deciphers 108-year-old WW1 German radio cipher · 1 src
- AI Safety Concerns Erupt Amid Recursive Self-Improvement Fears · 4 src
- Aws preview TBC rat‑Brain video model for select customers · 2 src
- ChatGPT-6 Astra decodes 108-year-old WWI German cipher · 2 src
Comments
via GitHub Discussions