Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQA

Perplexity AI has released pplx-embed-v2-late, a new pair of multimodal embedding models designed to handle text, images, and rendered PDF pages within…

By Vane October 8, 2026 3 min read

Perplexity AI has released pplx-embed-v2-late, a new pair of multimodal embedding models designed to handle text, images, and rendered PDF pages within a single vector space. The release includes two versions: a 0.6B model intended for fast, low-cost queries on edge devices, and a 9B model focused on maximum retrieval quality. Both are available on Hugging Face under an MIT license, though a hosted API endpoint is currently planned rather than live.

Performance and Benchmarks

The 9B model achieves a score of 92.4% on MADQA, an agentic PDF question-answering benchmark. This result beats Mixedbread’s retriever (88.9%) but trails Mixedbread Agentic Search (93.4%). The 0.6B model scores 90.1% on the same task. On domain-specific text retrieval across 72 tasks, the 9B model reaches 81.3% nDCG@10, leading all tested models by 1.6 percentage points, while the 0.6B version sits at 78.0%. Both models improve upon the previous best for Q2D-Web recall.

Image retrieval shows a wider gap. The 9B model scores 65.2% on ViDoRe v3, but Tencent’s EVIE model still outperforms it. The 0.6B model hits 62.3% on that metric and 61.2% on ViDoRe v3 Markdown. While the Markdown score is the lowest reported for the suite, it remains the second-best result on that specific benchmark. BrowseComp+ accuracy is listed at 64.0% for the 9B model, which is 4.9 percentage points higher than the next ColBERT model.

A notable finding is that mixing the models yields better results than using identical sizes. Querying a 9B index with the 0.6B model scored 63.5% on ViDoRe v3 image retrieval. This beats the 62.3% achieved when using the 0.6B model for both indexing and querying, at the same query cost.

Technical Specifications

Unlike dense models that compress a document into a single vector, this system stores a 128-dimensional vector for every token. Retrieval uses MaxSim, where each query token finds its best match in the document, and those maximums are summed. Perplexity distilled both models from an 18B teacher using LEAF-style token-level training to create a shared embedding space. Pages are encoded as images, removing the need for a separate OCR step.

The 0.6B version uses roughly 340 million active parameters for images and runs on laptops or small GPUs, requiring about 1.2 GB of memory in bf16 format. The 9B model requires 16 to 18 GB of memory and is suited for datacentre environments. Both require sentence-transformers version 6.0.0 or higher and transformers version 5.4.0 or higher.

Storage is a consideration because the system keeps one vector per token. This means the index size grows as the document length increases. The authors note that all scores are self-reported and a full technical report is not yet available.

How It Compares

The new models differ significantly from competitors like NVIDIA’s nemotron-colembed-vl-8b-v2 and Google Gemini Embedding 2. While Gemini supports text, image, video, audio, and PDF inputs, pplx-embed-v2-late currently handles text, images, and page renders. A key advantage is the shared embedding space across the two sizes, allowing a 0.6B query to search a 9B index effectively. The MIT license allows commercial use, unlike the proprietary Google model or the CC-BY-NC-4.0 license on the NVIDIA model.

What it means

Developers can now deploy a lightweight query encoder on edge devices while maintaining a high-quality index in the cloud. The ability to mix the smaller model for queries with the larger model for storage offers a practical balance between speed and accuracy without needing to retrain or switch systems. However, teams should account for the linear growth of storage requirements and verify the self-reported benchmark results independently.

Scroll to Top