Sarvam AI launches Saaras V4 speech-to-text model for 22 Indian languages and global English
Sarvam AI released Saaras V4, a speech-to-text model supporting all 22 scheduled Indian languages plus global English accents. The encoder-decoder system uses a 3B-parameter hybrid state-space language model (Sarvam-3B) trained in-house to process audio embeddings into transcripts. It claims state-of-the-art accuracy across languages, including noisy audio, with vendor-reported benchmarks…
Key points
- Saaras V4 supports 22 Indian languages and global English accents in one model, with vendor-reported state-of-the-art accuracy
- Keyterm prompting allows up to 50 terms to bias transcription, and five output modes let users choose transcription style
- API pricing is ₹30/hour for real-time transcription, with streaming support and sub-150ms latency
Key features include five output modes (transcription, verbatim, codemix, transliteration, and translation), keyterm prompting (up to 50 terms), and streaming capabilities with sub-150ms time-to-first-token. Pricing starts at ₹30 per hour for real-time transcription, with optional diarization at ₹45/hour. The model is available via Sarvam’s API today, though self-hosting documentation still references Saaras v3. No independent validation of the benchmarks has been published yet.
Model page: Saaras V4 →
The story so far
2 episodes →- Sarvam AI launches Saaras V4 speech-to-text model for 22 Indian languages and global Englishthis story
Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English
MarkTechPost · 26 September 2026
Loading the full article…
This text was published by MarkTechPost and written by Asif Razzaq. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Generative AI & Models
All →- OpenAI releases GPT-6 Sol with 7.7% improvement and half API cost · 7 src
- Anthropic releases Claude Opus 5.5 with 40% lower costs · 16 src
- Supersonic Labs releases Julia 1, a 144.3M-parameter open decision model · 1 src
- Claude Opus 5.5 tops benchmark over OpenAI’s Astra and Fable 5.1 · 3 src
- OpenAI releases GPT-6 Sol and Luna for cheaper coding and high-volume tasks · 5 src
Comments
via GitHub Discussions