Model Releases Hugging Face Blog

NeoMME: an efficient Multimodal-native and Multilingual Encoder

NeoMMEmultimodal encoderdocument retrieval

Many recent visual document retrievers are adapted from pretrained generative visual language models (VLMs), which rely on a separate vision encoder, a projector, and a causal decoder. NeoMME argues that for non-generative tasks like retrieval and classification, this architecture adds unnecessary parameter and compute overhead. Instead, NeoMME is a multimodal foundation encoder built from one bidirectional Transformer that processes both text tokens and raw image patches, with no pretrained vision tower, text encoder, or text decoder. It is trained entirely from scratch using a masked discrete-diffusion objective, available in 260M and 800M parameter variants. Text inputs use factorized token embeddings, while images are divided into patches and fed into the same computational path as text.

For visual document retrieval, NeoMME was fine-tuned using ColPali's page-image approach. The resulting NeoMME-Retriever produces both dense and late-interaction embeddings in a single forward pass. On the ViDoRe v3 benchmark, both model sizes sit on the Pareto frontier for nDCG@10 versus model size. At a matched 2048×2048 image input size on an NVIDIA L40S GPU, the 260M model encodes about 51 pages per second—roughly twice the throughput of ColModernVBERT.

The authors also introduce two efficiency techniques: hierarchical token pooling and asymmetric quantization. These reduce late-interaction index storage from about 1.5 MB per page to just 6 kB per page (255× smaller) while preserving more than 95% of baseline nDCG@10. NeoMME is available in Hugging Face Transformers, and all model checkpoints are released under the Apache 2.0 license. The release includes a model collection, a technical report, and a visual RAG demo.

Read original →

← Back to home