Product Updates Hugging Face Blog

Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers

sentence-transformersmulti-vectorcolbertretrieval

Sentence Transformers is a Python library for training and using embedding and reranker models for retrieval augmented generation, semantic search, and semantic textual similarity. Its v6.0 update introduces a fourth model type, MultiVectorEncoder, which enables ColBERT-style late interaction retrieval, along with a complete training approach for multi-vector models. The blogpost demonstrates how to finetune such a model and also how to train strong multi-vector models from scratch.

A dense embedding model compresses a whole text into a single vector, losing fine-grained details. A multi-vector model instead keeps one small vector per token and scores query-document pairs using the MaxSim operator, where each query token finds its best-matching document token and the scores are summed. This token-level matching preserves fine-grained signals and typically yields stronger retrieval, at the cost of a larger index. The training pipeline covers the model, datasets, loss functions, training arguments, evaluators, and the trainer class, with practical examples for each component.

The author finetuned the multi-vector-encoder/mLateOn-medical model in 14.5 hours on a single RTX 3090. Evaluation shows that this model outperforms every general-purpose retrieval model tested on a medical retrieval evaluation, including dense, sparse, lexical, and multi-vector models. The blogpost also includes a table of contents covering multi-vector model concepts, dataset handling (Hugging Face Hub and local), loss functions, training arguments, evaluators, callbacks, multi-dataset training, evaluation, and index optimization, with links to a companion blogpost about using multi-vector models and to prior posts on finetuning dense, sparse, and reranker models.

Read original →

← Back to home