Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
The blog post introduces a new addition to Sentence Transformers v6.0: MultiVectorEncoder, a fourth model type for ColBERT-style late-interaction retrieval. The motivation is that dense embedding models compress an entire text into a single fixed-size vector, forcing all information – rare entities, exact identifiers, or multiple query requirements – to compete for limited space. For example, a query for "green sofa with wooden legs and rounded cushions" gets blended into one point, so a green sofa with the wrong legs can appear similar.
A multi-vector model runs the same transformer but, instead of pooling token embeddings into one vector, projects each token embedding down to a small dimension (classically 128) and keeps them all, so a 9-token document becomes a 9x128 matrix. Query-document scoring uses the MaxSim operator, preserving token-level matching information that a single vector averages away. This usually yields stronger retrieval at the cost of a larger index.
The new encoder loads any PyLate checkpoint, Stanford-NLP ColBERT checkpoint, and colpali-engine models for visual document retrieval, all through the same familiar API used for dense, sparse, and reranker models. It is described as state of the art for visual document retrieval, where a text query is matched directly against page images with no OCR step. The blog also covers audio and video retrieval, interpretability, token pooling, and inference speedups.
Practical guidance in the post includes loading the various checkpoint formats, encoding and scoring, plugging the models into a search stack, running them on page images, and keeping the index affordable. The instructions run on a plain `pip install -U sentence-transformers`.