{"version":1,"type":"story","url":"https://digestai.news/story/perplexity-and-turbopuffer-release-pplx-embed-v2-context-9b-preview-em","json":"https://digestai.news/story/perplexity-and-turbopuffer-release-pplx-embed-v2-context-9b-preview-em.json","markdown":"https://digestai.news/story/perplexity-and-turbopuffer-release-pplx-embed-v2-context-9b-preview-em.md","slug":"perplexity-and-turbopuffer-release-pplx-embed-v2-context-9b-preview-em","headline":"Perplexity and turbopuffer release pplx-embed-v2-context-9b-preview embedding model","summary":"Perplexity Research and turbopuffer have launched a preview of their new contextual embedding model, **pplx-embed-v2-context-9b-preview**. The 9‑billion‑parameter model is available as self‑hosted weights on Hugging Face under an MIT license, but it is not yet accessible through the Perplexity API. Loading the model requires *transformers* 5.4.0 with `trustremotecode=True`, and the model card warns that the weights and interface may change without backward compatibility.\n\nThe model improves retrieval‑augmented generation by training a token‑level teacher that scores every token for a query‑document pair, producing soft targets instead of a single “gold” chunk. It encodes an entire document in one pass, then pools per chunk, allowing retrieval of answers together with the evidence needed to verify them. Benchmarks on the private *context‑bench* (2,099 queries, 38,894 documents, 2,458,072 chunks) show **45.5% answer recall** and **40.6% evidence recall** at K = 10, a 14.4‑point and 5.0‑point lead over the competing voyage‑context‑4 model respectively. The model stores 1024‑dim int8 vectors (≈1 KB each) and supports 2048‑dim float32 vectors.","keyPoints":["Perplexity and turbopuffer released pplx-embed-v2-context-9b-preview, a 9B contextual embedding model","Model achieves 45.5% answer recall and 40.6% evidence recall at K=10 on context‑bench","Weights are MIT‑licensed on Hugging Face; not yet available via Perplexity API"],"whyItMatters":"The model’s ability to retrieve answers with supporting evidence could raise the reliability of RAG systems, reducing reliance on single gold passages and improving downstream AI applications.","category":{"slug":"models","name":"Generative AI & Models","url":"https://digestai.news/category/models"},"entities":{"companies":["Perplexity Research","turbopuffer"],"models":["pplx-embed-v2-context-9b-preview"],"people":["Asif Razzaq"]},"firstPublishedAt":"2026-10-01T03:23:39Z","updatedAt":"2026-10-01T03:23:39Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"MarkTechPost","title":"Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence","url":"https://marktechpost.com/2026/09/30/perplexity-releases-pplx-embed-v2-context-9b-preview-a-contextual-embedding-model-that-retrieves-answers-and-their-supporting-evidence","publishedAt":"2026-10-01T03:23:39Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":{"title":"AI Embedding Race Heats Up","url":"https://digestai.news/thread/linkup-research-releases-sparseup-149m-parameter-sparse-embedding-model","storyCount":2},"cite":{"text":"Digest AI, \"Perplexity and turbopuffer release pplx-embed-v2-context-9b-preview embedding model\", 1 October 2026, https://digestai.news/story/perplexity-and-turbopuffer-release-pplx-embed-v2-context-9b-preview-em","publisher":"Digest AI","title":"Perplexity and turbopuffer release pplx-embed-v2-context-9b-preview embedding model","datePublished":"2026-10-01T03:23:39Z","url":"https://digestai.news/story/perplexity-and-turbopuffer-release-pplx-embed-v2-context-9b-preview-em"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}