What did Perplexity publish, and when?

The dated announcement is the 7 October 2026 research post. It presents pplx-embed-v2-late as the late-interaction sibling of pplx-embed-v2-context, which the same post says shipped the week before. Simple HTTP clients hit Cloudflare on that URL; the text we checked on 9 October is the article body opened through a browser-class fetch. We are dating the news to that post, not to an inferred Hub upload.

The two public checkpoints are perplexity-ai/pplx-embed-v2-late-0.6b and perplexity-ai/pplx-embed-v2-late-9b. Unauthenticated Hub API calls on 9 October returned gated=false, private=false, library_name sentence-transformers, and lastModified 2026-10-08T14:34:03Z for both. The createdAt stamps are earlier — 20 July 2026 for the 9B repo and 3 August 2026 for the 0.6B repo — so the repositories existed before the dated blog. That is a Hub metadata fact, not a second launch.

This is a retriever release, not a chat model. It sits nearer Google’s EmbeddingGemma 2 than a workplace agent: vectors for search, not answers. EmbeddingGemma 2 is a single-vector multimodal encoder. Perplexity’s pair keeps one vector per token.

How does late interaction differ from one vector?

A dense embedder compresses a query or a document into one vector and scores with a dot product. That is cheap enough for approximate nearest-neighbour search, and it is the setup in our text-embeddings explainer. Perplexity’s post repeats the usual cost: long or visually dense pages have to fit into that one vector, and some score patterns are impossible no matter how you train.

Late-interaction, ColBERT-style models still encode the document offline, but they keep a vector for each retained token and score with MaxSim: each query token takes its best document-token match, then those maxima are summed. The post says pplx-embed-v2-late emits 128-dimensional vectors and uses that scoring rule. Storage and candidate scoring then grow with document length. That is a different serving problem from one-vector nearest-neighbour search.

What can you actually download and run?

Both cards say the models are multimodal late-interaction retrievers for text, images and visual documents, built on Qwen3.5 with bidirectional attention. They load with sentence-transformers >= 6.0.0 and transformers >= 5.4.0 through MultiVectorEncoder. The usage block encodes queries and documents separately, then calls similarity. That is first-stage retrieval, not a complete RAG pipeline.

The cards are explicit about a batching limit: encode text-only and image-only batches in separate calls. “Mixed text+image inputs are not supported.” They also warn that PyLate inserts query and document markers at the second position; this export expects those markers first. Neither repository had a LICENSE file when we requested /raw/main/LICENSE on 9 October (HTTP 404). The YAML on both READMEs, and the Hub cardData.license field, say mit.

The blog’s “Getting started” paragraph says both sizes are on Hugging Face and that Perplexity “will progressively roll out support for late-interaction, dense, and contextual embeddings on the Perplexity API Platform.” That is a planned API, not a claim the late-interaction pair is a public hosted endpoint today. A technical report is promised later this year. The 18B teacher used for distillation is not released.

Why index with 9B and query with 0.6B?

The interoperability claim is the reason to care about two sizes. Perplexity trains an 18B teacher contrastively, then distils 9B and 0.6B students with a token-level LEAF-style objective that aligns each student’s per-token vector with the teacher’s. Because both students share that space, the post lists four layouts: 9B on both sides; 0.6B on both sides; 9B for the index and 0.6B for live queries; and a local–cloud split that encodes private documents or queries with 0.6B and compares them to a hosted 9B index.

The 0.6B card lists 340 million active parameters. The blog’s more detailed count starts from Qwen3.5-0.8B, prunes the 24-layer text tower to 12 layers, and arrives at 594 million total parameters, of which about 240 million are active for text encoding and 340 million for image encoding. The 9B card lists 7.4 billion active parameters. Those are Perplexity’s parameter tables, not an independent count of the safetensors.

The post’s own comparison for the asymmetric layout — 0.6B queries against a 9B-built index — is a 1.6 point nDCG@10 gain on a 72-task domain-specific suite and 63.5% versus 62.3% on public ViDoRe v3 images. Symmetric 9B remains higher. That is a deployment option for teams that can afford a one-off 9B index pass but not 9B at query time. It is not a substitute for measuring recall on your own files, the same warning we give for vector search.

What do the vendor scores actually cover?

On the public ViDoRe v3 subset the cards print 62.3% / 61.2% nDCG@10 for the 0.6B model (image / Markdown) and 65.2% / 64.7% for the 9B model. The blog says the 0.6B image score beats other vision-language baselines they plotted and trails some larger vision-only ColBERT systems, while using 128-dimensional tokens instead of 2,048 or 4,096. Those comparisons are Perplexity’s. Visual-document search here means scoring a text query against a rendered page, which is a different pipeline from OCR-then-parse document AI.

For text, the post reports 81.3% and 78.0% average nDCG@10 on 72 domain-specific tasks it says were held out of training, and Recall@1000 of 74.8% / 73.6% on Q2D-Web Combined against a previous best of 69.3%. Q2D-Web is Perplexity’s own web-scale set: about 70,000 agent-reformulated production queries over 190 million documents. Combined judgements mix citations, the company’s production ranker, and LLM labels. That is useful context and a home-field risk. PPLX-Q2I is likewise built from Perplexity image-search logs.

Agentic tables use other models as the agent: GPT-OSS-120B on BrowseComp+, where the 9B retriever is reported at 64.0% answer accuracy, and Gemini 3.5 Flash on MADQA, where the 9B / 0.6B retrievers are reported at 92.4% and 90.1% answer accuracy over 800 PDFs. Those scores measure a retriever-plus-agent stack the authors chose. We did not rerun ViDoRe, Q2D-Web or MADQA.

What should readers not assume?

“State of the art” in the post is Perplexity’s reading of the baselines it plotted. The training mix is described as 186 million query–document pairs from 594 datasets in 46 languages, with eval-associated sets excluded. That exclusion is the authors’ contamination control; it does not make the public numbers independent. A late-interaction index is also larger than a single-vector index of the same corpus, which matters before you replace a production OpenDocRouter-style parser swap or a dense embedder.

The MIT mark is the Hub card field. There is no LICENSE file in either repo we requested. Qwen3.5 lineage is stated on the cards; we are not offering a licence-compatibility opinion. The API Platform sentence is forward-looking. createdAt on the Hub is not the announcement date.

This is an evidence review of pages opened on 9 October. It is not a first-hand retrieval run and not a claim that pplx-embed-v2-late is generally better than EmbeddingGemma 2, Nemotron ColEmbed or Voyage outside Perplexity’s tables.

Common questions

Can I call these models on Perplexity’s API today?

Not according to the 7 October post. The weights are on Hugging Face. The company says late-interaction support on its API Platform will roll out with dense and contextual embeddings. We did not find a public late-interaction endpoint in the pages we opened.

Did anyone outside Perplexity rerun the 65.2% ViDoRe score?

Not in the pages we opened. The cards and the blog print Perplexity’s public-subset nDCG@10. We did not download weights or score ViDoRe.

If I already have a single-vector index, do I have to rebuild it?

Yes, if you want these models. Late interaction stores one 128-dimensional vector per document token, not one vector per document. A dense index is not a drop-in MaxSim index. The 0.6B and 9B pair can share an index with each other because they share a space; they do not share a space with EmbeddingGemma 2 or a generic dense model.

THE TAKEAWAY

What to remember

Open the 7 October post for the shared-space layouts and the 0.6B parameter breakdown, and the Hugging Face cards for the Sentence Transformers snippet and the mixed-batch limit. Keep ViDoRe and Q2D-Web in the vendor column until you score your own pages.

Sources & further reading

  1. Multimodal embeddings beyond a single vector ↗
  2. perplexity-ai/pplx-embed-v2-late-0.6b model card ↗
  3. perplexity-ai/pplx-embed-v2-late-9b model card ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories