Perplexity Research and turbopuffer have released pplx-embed-v2-context-9b-preview, a contextual embedding model designed for RAG pipelines. It embeds each chunk while keeping the full document in view. The core change is the training signal: the model learns to retrieve the answer alongside the context needed to verify it, rather than targeting a single ‘gold passage.’
In this article
The weights are available on Hugging Face under the MIT license for self-hosted use. Deployment requires transformers version 5.4.0 or higher with trust_remote_code set to True. The model is not yet on the Perplexity API, and the release notes warn that weights and interfaces may change without maintaining backward compatibility.
Why the gold passage falls short
RAG systems split long documents into chunks. A single chunk often depends on an entity, heading, or definition stated elsewhere in the text. Contextual models address this using late chunking. The document is encoded in one pass, then pooled per chunk.
Training, however, usually marks one gold chunk per query. Every other chunk becomes a negative, including the sentences that make the answer checkable. Perplexity lists three further problems. Binary labels provide a coarse signal. LLM annotation costs grow linearly with dataset size. Labels are also tied to one specific chunking strategy.
How the training works
The teacher is Perplexity’s query-aware context compression model. It reads the query and document together and scores every token.
- Chunk relevance: the mean of the top n token scores inside each chunk.
- Soft target: a temperature-scaled softmax over chunks in the positive document. Chunks in other documents get zero.
- Distillation loss: forward KL divergence between teacher and student distributions.
- Document loss: InfoNCE, where a document scores as its best chunk, inspired by ColBERT’s MaxSim.
Each batch samples a random chunking strategy. Chunks are separated by a learned <|chunk_sep|> token and mean-pooled. The teacher runs only during training, so inference adds no latency or storage.
The model starts from an in-house 9B ColBERT retrieval model. A linear projection outputs 2048 dimensions. Matryoshka training also supports 1024 dimensions. Quantization-aware training enables native int8 embeddings. The release is a model soup of several checkpoints. Training used roughly 430 datasets covering over 50 languages, with no ConTEB data.
Interactive explainer
How pplx-embed-v2-context retrieves answers and their evidence
An interactive walkthrough of Perplexity’s contextual embedding preview: why isolated chunks fail, how teacher distillation replaces the single gold passage, and what the reported numbers mean.
Three lease files share the sentence “Monthly rent is …”. Only one belongs to 5 Park Avenue. Switch modes and run retrieval.
Illustrative example modeled on the lease scenario in Perplexity’s post. Scores are for explanation only, not model outputs.
A context compression model acts as teacher. It scores every token for the query. Those scores are pooled per chunk (mean of the top n tokens) and turned into a soft target, instead of a one-hot gold label.
Temperature0.25
Change the boundaries: the same token scores re-aggregate without re-annotation. That is the “flexible chunk boundaries” property Perplexity describes. Token scores here are illustrative.
context-bench (2,099 queries, 38,894 documents, 2,458,072 sentence chunks, exhaustive ranking). Numbers below are as reported by Perplexity at K = 10.
Voyage values are computed as Perplexity’s figure minus the stated gap (14.4 and 5.0 points). Other Voyage metrics appear only in Perplexity’s chart and are not shown here.
Contextual embeddings store one vector per chunk, same as a normal chunk index. Cost depends on vector size. Perplexity reports that 1024-dim int8 (1 KB) slightly exceeds voyage-context-4 at 2048-dim float32 (8 KB) on its chunk-retrieval suite.
Bytes = dimensions x bytes per value. Sensitivity is mean nDCG@10 across 74 MTEB tasks, as reported by Perplexity.
What it means
For developers building retrieval systems, the shift is practical. The model allows you to treat the entire document as the source of truth rather than forcing a rigid split. This reduces the need to manually tune chunk sizes to capture context. It also lowers annotation costs because the system learns from token-level relevance rather than requiring perfect single-sentence labels. Storage remains efficient, with int8 quantization keeping vector sizes low.




