Skip to content
TrackPodcasts
scienceOct 4, 202523:21pending

ModernVBERT: Towards Smaller Visual Document Retrievers

About this episode

🤗 Upvotes: 24 | cs.IR

Authors:
Paul Teiletche, Quentin Macé, Max Conti, Antonio Loison, Gautier Viaud, Pierre Colombo, Manuel Faysse

Title:
ModernVBERT: Towards Smaller Visual Document Retrievers

Arxiv:
http://arxiv.org/abs/2510.01149v1

Abstract:
Multimodal embedding models are gaining prevalence, notably for document retrieval as efficient alternatives to text-only pipelines. These models are typically built by finetuning large vision-language decoders (VLMs) with contrastive losses on text-image pairs. In this work, we show that, while cost-efficient, this repurposing approach often bottlenecks retrieval performance. Through controlled experiments, we establish a principled recipe for improving visual document retrieval models. We notably measure the impact of attention masking, image resolution, modality alignment data regimes, and late interaction centered contrastive objectives which emerge as central performance factors. Building on these insights, we release ModernVBERT, a compact 250M-parameter vision-language encoder that outperforms models up to 10 times larger when finetuned on document retrieval tasks. Models and code are made available at https://huggingface.co/ModernVBERT.

Get every episode summarized

Each time Daily Paper Cast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

No transcript yet

This episode has not been transcribed. Request it and it moves to the front of the queue.

ModernVBERT: Towards Smaller Visual Document Retrievers

Daily Paper Cast

0:00
23:21

More episodes

More from Daily Paper Cast

View all episodes →