# Sarvam AI launches Saaras V4 speech-to-text model for 22 Indian languages and global English

Digest AI · Generative AI & Models · published 2026-09-26T21:56:23Z

Canonical: https://digestai.news/story/sarvam-ai-launches-saaras-v4-speech-to-text-model-for-22-indian-langua

## Summary

Sarvam AI released Saaras V4, a speech-to-text model supporting all 22 scheduled Indian languages plus global English accents. The encoder-decoder system uses a 3B-parameter hybrid state-space language model (Sarvam-3B) trained in-house to process audio embeddings into transcripts. It claims state-of-the-art accuracy across languages, including noisy audio, with vendor-reported benchmarks showing lower word error rates (WER) than competitors like Deepgram Nova-3 and GPT-4o Transcribe in some cases.

Key features include five output modes (transcription, verbatim, codemix, transliteration, and translation), keyterm prompting (up to 50 terms), and streaming capabilities with sub-150ms time-to-first-token. Pricing starts at ₹30 per hour for real-time transcription, with optional diarization at ₹45/hour. The model is available via Sarvam’s API today, though self-hosting documentation still references Saaras v3. No independent validation of the benchmarks has been published yet.

## Key points

- Saaras V4 supports 22 Indian languages and global English accents in one model, with vendor-reported state-of-the-art accuracy
- Keyterm prompting allows up to 50 terms to bias transcription, and five output modes let users choose transcription style
- API pricing is ₹30/hour for real-time transcription, with streaming support and sub-150ms latency

## Why it matters

This model fills a critical gap for multilingual speech recognition in India, where English and regional languages coexist. Its low-latency streaming and multi-mode outputs could simplify workflows for call centers, media, and accessibility tools, though vendor claims await independent verification.

## Sources

1. [Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English](https://marktechpost.com/2026/09/26/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english) (MarkTechPost, 2026-09-26)

Part of the developing story: [Rapid Advances in Multilingual Speech Recognition Technology](https://digestai.news/thread/xai-releases-grok-voice-transcribe-2-0-speech-to-text-model) (2 stories)

## Cite

Digest AI, "Sarvam AI launches Saaras V4 speech-to-text model for 22 Indian languages and global English", 26 September 2026, https://digestai.news/story/sarvam-ai-launches-saaras-v4-speech-to-text-model-for-22-indian-langua

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/sarvam-ai-launches-saaras-v4-speech-to-text-model-for-22-indian-langua.json
