{"version":1,"type":"story","url":"https://digestai.news/story/prime-intellect-launches-prime-inference-for-serving-open-models","json":"https://digestai.news/story/prime-intellect-launches-prime-inference-for-serving-open-models.json","markdown":"https://digestai.news/story/prime-intellect-launches-prime-inference-for-serving-open-models.md","slug":"prime-intellect-launches-prime-inference-for-serving-open-models","headline":"Prime Intellect launches Prime Inference for serving open models","summary":"Prime Intellect has launched Prime Inference, a serving platform for frontier open-source models. It offers serverless endpoints and reserved capacity on the company's own GPUs across multiple datacenters. Before public release, the platform processed nearly a trillion tokens per day internally, used for RL rollouts, synthetic data generation, evaluations, and long-running coding agents. Prime Inference is the serving layer of Prime Intellect's open training stack, designed to close the loop by feeding production traces back into training.\n\nThe platform supports OpenAI-compatible endpoints at https://api.pinference.ai/api/v1 and runs on NVIDIA Blackwell hardware, with Vera Rubin listed as coming soon. Prime reports its GLM-5.3 endpoint ranks among the fastest on OpenRouter, with near-zero tool-call error rate and 100% uptime since launch. The serving stack combines NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer, built with Inferact and NVIDIA, and includes prefill/decode disaggregation, cache-aware routing, and NVFP4 KV compression to improve latency and capacity.\n\nKey technical results include a 1:4 prefill/decode ratio serving 66 sessions per prefill group at 101 tokens per second per user, NVFP4 KV cache increasing cached tokens per decoder from 1.09M to 1.63M, and nearly 40% lower p90 inter-token latency in tests. Prime Intellect also contributed a structural-tag builder to Dynamo to improve tool-call reliability for agents. The company notes batch inference and 1-click dedicated deploys are next on the roadmap.","keyPoints":["Prime Intellect launched Prime Inference, a serving platform for open models with serverless and reserved modes.","Before launch, it processed nearly a trillion tokens per day internally for RL rollouts, synthetic data, evaluations, and coding agents.","GLM-5.3 on GB200 NVL72 achieved 101 tokens per second per user with a 1:4 prefill/decode ratio and 66 sessions per prefill group."],"whyItMatters":"Prime Inference provides a high-performance, open-model serving option that could reduce reliance on proprietary APIs, especially for agentic workloads requiring low latency and high throughput.","category":{"slug":"agents","name":"Agents & Tools","url":"https://digestai.news/category/agents"},"entities":{"companies":["Prime Intellect","NVIDIA","OpenRouter","ComputePrices"],"models":["GLM-5.3"],"people":["Michal Sutter"]},"firstPublishedAt":"2026-10-03T05:37:18Z","updatedAt":"2026-10-03T05:37:18Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"MarkTechPost","title":"Prime Intellect Launches Prime Inference: Serverless and Reserved Serving for Frontier Open Models","url":"https://marktechpost.com/2026/10/02/prime-intellect-launches-prime-inference-serverless-and-reserved-serving-for-frontier-open-models","publishedAt":"2026-10-03T05:37:18Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Prime Intellect launches Prime Inference for serving open models\", 3 October 2026, https://digestai.news/story/prime-intellect-launches-prime-inference-for-serving-open-models","publisher":"Digest AI","title":"Prime Intellect launches Prime Inference for serving open models","datePublished":"2026-10-03T05:37:18Z","url":"https://digestai.news/story/prime-intellect-launches-prime-inference-for-serving-open-models"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}