The top open-weight visual document retrieval models in October 2026 are EVIE from Tencent and VultronRetriever from Vultr. Both are Apache 2.0 and both are ColQwen-style retrievers built on Qwen3.5. Tencent's card puts EVIE-8B first on ViDoRe V3 at 66.24, with EVIE-4.5B second at 65.70. VultronRetriever Prime scores 64.26 and uses 320-dimension vectors, which keeps the index small. webAI's ColVec1.1 and NVIDIA's Nemotron ColEmbed V2 score in the same range but are licensed for non-commercial use.
These models retrieve document pages as images. They embed a rendered page directly, so tables, charts and layout are searchable without OCR. This page covers that category only. For text retrieval see open-weight embedding models and open-weight rerankers. For OCR see OCR and document AI models.
ViDoRe V3 Open-Weight Comparison
| Model | Developer | Parameters | License | Released | ViDoRe V3 mean nDCG@10 and source | Hardware to run |
|---|---|---|---|---|---|---|
| EVIE-8B | Tencent | 8.4B | Apache 2.0 | September 2026 | 66.24 (Tencent card) | Not stated on card |
| EVIE-4.5B | Tencent | 4.5B | Apache 2.0 | September 2026 | 65.70 (Tencent card) | Index of 3.81 GiB per million pages with token compression |
| webAI-ColVec1.1-8b | webAI | 8.4B | webAI Non-Commercial License v1.0 | July 2026 | 64.95 (webAI card) | Not stated on card |
| VultronRetriever Prime | Vultr | 8.4B | Apache 2.0 | May 2026 | 64.26 (Vultr card, official MTEB run) | About 17 GB in bf16, one GPU |
| VultronRetriever Core | Vultr | 4.5B | Apache 2.0 | June 2026 | 63.57 (Vultr card) | About 9 GB in bf16 |
| Nemotron ColEmbed VL 8B V2 | NVIDIA | 8.7B | CC BY-NC 4.0 | January 2026 | 63.42 (leaderboard, as quoted by Vultr and webAI) | NVIDIA GPU |
| tomoro-colqwen3-embed-8b | Tomoro AI | 8B | Apache 2.0 | November 2025 | 61.59 (leaderboard, as quoted by Vultr and webAI) | Not stated |
| colqwen3.5-4.5B-v3 | athrael-soju | 4.5B | Apache 2.0 | March 2026 | 61.46 (model card) | About 8.7 GB memory |
| jina-embeddings-v4 | Jina AI | 3.8B | Qwen Research License | May 2025 | 57.52 (leaderboard, as quoted by Vultr) | Not stated |
| VultronRetriever Flash | Vultr | 0.85B | Apache 2.0 | June 2026 | 56.16 (Vultr card) | Not stated |
| ColQwen2.5 v0.2 | ViDoRe team | 3B | MIT | January 2025 | 51.90 (leaderboard, as quoted by Vultr) | Not stated |
| ColPali v1.3 | ViDoRe team | 3B | MIT | November 2024 | Not reported | Not stated |
ViDoRe V3 is the current version of the benchmark behind the ViDoRe leaderboard. The headline score is mean nDCG@10 over ten tasks: eight public and two private tasks scored by the maintainers. The lead changed at least three times in 2026. NVIDIA's paper put Nemotron ColEmbed V2 first on February 3 at 63.42. Vultr's card shows VultronRetriever first in early July, webAI's card puts ColVec1.1 ahead by the end of July, and Tencent's card claims the top two places from September.
Which to Pick for Which Job
| Job | Pick | Reason |
|---|---|---|
| Highest accuracy, commercial use | EVIE-8B | Top reported V3 score, Apache 2.0 |
| Large collections where index size matters | VultronRetriever Prime or EVIE-4.5B | 320-dimension vectors on Vultron. Token compression to 32 vectors per page on EVIE-4.5B |
| Mid-size model | VultronRetriever Core or EVIE-4.5B | About 4.5B parameters each. Vultron Core is about 9 GB in bf16 |
| Small footprint | VultronRetriever Flash or ColSmol-500M | Under 1B parameters. ColSmol-500M is the most downloaded model in the category on Hugging Face |
| One vector per page in a standard vector database | Qwen3-VL-Embedding-8B | Single-vector multimodal embedding, Apache 2.0. Scores 83.3 on the visual document part of MMEB-V2 per Qwen |
| Mostly plain text documents | A text embedding model | Vultr's own card notes a single-vector text embedder is cheaper for text-only search |
Licences: Open Weight Is Not Always Open Source
Permissive. EVIE, VultronRetriever, the Tomoro ColQwen3 models, colqwen3.5-4.5B-v3 and Qwen3-VL-Embedding are Apache 2.0. The original ColPali and ColQwen2.5 adapters are MIT, on top of base models with their own terms.
Non-commercial. Nemotron ColEmbed VL 8B V2 is CC BY-NC 4.0, and NVIDIA's card says it is for non-commercial and research use. webAI ColVec1 and ColVec1.1 use a webAI non-commercial licence. jina-embeddings-v4 falls under the Qwen Research License. Its card says an earlier CC BY-NC label was an error.
Two of the top six models in the table cannot be deployed commercially without a separate agreement, so the leaderboard order and the usable order differ.
Caveats
- Top scores are developer-reported. We could not load the live leaderboard, so the EVIE figures are from Tencent's card. EVIE results are present in the public MTEB results repository. Tencent says its paper is still to come.
- The benchmark is hard. No model exceeds 67 on V3, while the same models score above 90 on ViDoRe V1. Expect misses on multi-hop and chart-heavy questions.
- Multi-vector indexes are large. Late-interaction models store hundreds of vectors per page. Vector dimension ranges from 128 to 4096 across this table, which changes storage cost by more than an order of magnitude.
- A reranker can matter more than the retriever. On the ViDoRe V3 pipeline results listed by benchmarklist.com, jina-embeddings-v4 with a text reranker scored 0.631 nDCG@5 on English, ahead of Nemotron ColEmbed VL 8B V2 alone at 0.620.
Brand Visibility Implications
Document retrievers sit inside enterprise search and RAG products. What they return is not observable from outside, and no brand can track its presence in someone else's private index. Two things are worth noting. First, the model choice itself is now often made by asking an assistant, so retriever developers and RAG vendors depend on assistants knowing the current leaderboard and licences. Second, these models read pages as images. Reports, slides and PDFs that carry key facts only in charts are now retrievable, which rewards clear labels and legible figures in published documents.
Methodology
This page is compiled from published sources, not from Presenc AI measurements. Parameter counts, licences and repository dates come from each model's Hugging Face card and the Hugging Face model API, read on October 1, 2026. Where a developer gives no release date, the month shown is the month the repository was created. Scores are quoted from the party that ran them and are labelled as such: developer model cards and papers, the GIFT-Eval leaderboard result files, the ViDoRe leaderboard as quoted on model cards, the openpi and Isaac GR00T repositories, and independent studies such as the Domyn guard model benchmark. Developer-reported scores use each developer's own test setup and are not directly comparable across rows. "Not reported" means we found no published score on that benchmark. Rankings in these categories change monthly. Status as of October 1, 2026.
How Presenc AI Helps
Presenc AI tracks which retrieval models, vector databases and RAG tools AI assistants recommend and how they describe them. Vendors can see whether they appear when engineers ask an assistant how to build document search.