Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings
Ovis-Embedding is presented as a state-of-the-art omni-modal embedding family built on native integration of text, image, video, and audio. Rather than assembling separate modality towers, it uses a shared multimodal backbone to encode different modalities in a common representation space, aiming at universal any-to-any retrieval.
The report highlights three key advances. First, for native omni-modal initialization, it adopts a pretrained Qwen-omni model as the embedding backbone and adapts it through contrastive training with low-rank initialization. Second, for data-centric omni-modal training, it constructs a broad, high-quality corpus spanning text, images, video, audio, and interleaved multimodal data; to improve data efficiency, it introduces homogeneous-source sampling to form task-consistent batches with informative in-batch negatives.
Third, for embedding-specific training and inference optimization, it uses focal loss to emphasize hard examples and similarity-based Embedding Distillation to transfer fine-grained similarity structure from complementary experts. At inference time, low-rank feature decomposition enables compact embeddings with flexible dimensionality and minimal performance loss.
Empirical evaluations show that the Ovis-Embedding family achieves state-of-the-art performance on MMEB-v3, MMEB-v2, MVEB, MAEB, and RTEB, demonstrating effectiveness across text, image, video, and audio modalities. The authors say these results highlight the potential of unified omni-modal training to overcome modality fragmentation and advance universal embedding models for any-to-any retrieval.