# Prime Intellect launches Prime Inference for serving open models

Digest AI · Agents & Tools · published 2026-10-03T05:37:18Z

Canonical: https://digestai.news/story/prime-intellect-launches-prime-inference-for-serving-open-models

## Summary

Prime Intellect has launched Prime Inference, a serving platform for frontier open-source models. It offers serverless endpoints and reserved capacity on the company's own GPUs across multiple datacenters. Before public release, the platform processed nearly a trillion tokens per day internally, used for RL rollouts, synthetic data generation, evaluations, and long-running coding agents. Prime Inference is the serving layer of Prime Intellect's open training stack, designed to close the loop by feeding production traces back into training.

The platform supports OpenAI-compatible endpoints at https://api.pinference.ai/api/v1 and runs on NVIDIA Blackwell hardware, with Vera Rubin listed as coming soon. Prime reports its GLM-5.3 endpoint ranks among the fastest on OpenRouter, with near-zero tool-call error rate and 100% uptime since launch. The serving stack combines NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer, built with Inferact and NVIDIA, and includes prefill/decode disaggregation, cache-aware routing, and NVFP4 KV compression to improve latency and capacity.

Key technical results include a 1:4 prefill/decode ratio serving 66 sessions per prefill group at 101 tokens per second per user, NVFP4 KV cache increasing cached tokens per decoder from 1.09M to 1.63M, and nearly 40% lower p90 inter-token latency in tests. Prime Intellect also contributed a structural-tag builder to Dynamo to improve tool-call reliability for agents. The company notes batch inference and 1-click dedicated deploys are next on the roadmap.

## Key points

- Prime Intellect launched Prime Inference, a serving platform for open models with serverless and reserved modes.
- Before launch, it processed nearly a trillion tokens per day internally for RL rollouts, synthetic data, evaluations, and coding agents.
- GLM-5.3 on GB200 NVL72 achieved 101 tokens per second per user with a 1:4 prefill/decode ratio and 66 sessions per prefill group.

## Why it matters

Prime Inference provides a high-performance, open-model serving option that could reduce reliance on proprietary APIs, especially for agentic workloads requiring low latency and high throughput.

## Sources

1. [Prime Intellect Launches Prime Inference: Serverless and Reserved Serving for Frontier Open Models](https://marktechpost.com/2026/10/02/prime-intellect-launches-prime-inference-serverless-and-reserved-serving-for-frontier-open-models) (MarkTechPost, 2026-10-03)

## Cite

Digest AI, "Prime Intellect launches Prime Inference for serving open models", 3 October 2026, https://digestai.news/story/prime-intellect-launches-prime-inference-for-serving-open-models

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/prime-intellect-launches-prime-inference-for-serving-open-models.json
