Perplexity and turbopuffer release pplx-embed-v2-context-9b-preview embedding model
Perplexity Research and turbopuffer have launched a preview of their new contextual embedding model, pplx-embed-v2-context-9b-preview. The 9‑billion‑parameter model is available as self‑hosted weights on Hugging Face under an MIT license, but it is not yet accessible through the Perplexity API. Loading the model requires transformers 5.4.0 with trustremotecode=True, and the model card warns that…
Key points
- Perplexity and turbopuffer released pplx-embed-v2-context-9b-preview, a 9B contextual embedding model
- Model achieves 45.5% answer recall and 40.6% evidence recall at K=10 on context‑bench
- Weights are MIT‑licensed on Hugging Face; not yet available via Perplexity API
The model improves retrieval‑augmented generation by training a token‑level teacher that scores every token for a query‑document pair, producing soft targets instead of a single “gold” chunk. It encodes an entire document in one pass, then pools per chunk, allowing retrieval of answers together with the evidence needed to verify them. Benchmarks on the private context‑bench (2,099 queries, 38,894 documents, 2,458,072 chunks) show 45.5% answer recall and 40.6% evidence recall at K = 10, a 14.4‑point and 5.0‑point lead over the competing voyage‑context‑4 model respectively. The model stores 1024‑dim int8 vectors (≈1 KB each) and supports 2048‑dim float32 vectors.
Model page: pplx-embed-v2-context-9b-preview →
The story so far
2 episodes →- Perplexity and turbopuffer release pplx-embed-v2-context-9b-preview embedding modelthis story
Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence
MarkTechPost · 1 October 2026
Loading the full article…
This text was published by MarkTechPost and written by Asif Razzaq. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Generative AI & Models
All →- OpenAI releases GPT-6.1 Sol, says it nearly matches Astra on agentic coding at one-fifth the price · 31 src
- Google rolls out Gemini 4 Argon to trusted cyber defenders first · 8 src
- Google faces internal doubts over Gemini 4 coding performance ahead of launch · 2 src
- Gemini Notebook isolates AI interactions for testing · 1 src
- OpenAI scraps GPT-6.1 Astra release over safety concerns · 61 src
Comments
via GitHub Discussions