{"version":1,"type":"story","url":"https://digestai.news/story/xai-releases-grok-voice-transcribe-2-0-speech-to-text-model","json":"https://digestai.news/story/xai-releases-grok-voice-transcribe-2-0-speech-to-text-model.json","markdown":"https://digestai.news/story/xai-releases-grok-voice-transcribe-2-0-speech-to-text-model.md","slug":"xai-releases-grok-voice-transcribe-2-0-speech-to-text-model","headline":"xAI releases Grok Voice Transcribe 2.0 speech-to-text model","summary":"On September 18, 2026, xAI introduced Grok Voice Transcribe 2.0, its newest speech‑to‑text model. The service keeps batch pricing at $0.10 per hour of audio and $0.20 per hour for streaming, and xAI says the model is twice as accurate as the previous version. It was trained on live, noisy, multilingual audio from diverse environments and is claimed to be among the most accurate transcription models for real‑world use.\n\nThe company reports that Grok Voice Transcribe 2.0 tops a public leaderboard of 32 streaming models and outperforms its predecessor on four internal test sets, including an 8 kHz English telephony set. Multilingual performance shows the biggest gain, with word error rate dropping from 20.6 percent to 6.8 percent on a short‑phrase multilingual benchmark. The model adds features such as word‑level timestamps, speaker diarization, up to eight‑channel transcription, and automatic language detection.\n\nAtlassian evaluated the model for Loom and now uses it to transcribe every Loom video. Existing Speech‑to‑Text API integrations receive the accuracy boost without code changes, and the older version will be deprecated in the coming weeks.","keyPoints":["xAI launched Grok Voice Transcribe 2.0 on Sep 18 2026, pricing $0.10 per hour batch, $0.20 per hour streaming.","The model claims twice the accuracy of version 1.0 and leads a 32‑model leaderboard, cutting word error rate from 20.6% to 6.8% on multilingual phrases.","Atlassian evaluated the model for Loom, adopting it to transcribe all Loom videos, with integration requiring no code changes."],"whyItMatters":"A more accurate, multilingual transcription model at low cost can improve customer‑support, video captioning, and voice‑agent workflows, giving developers and businesses a stronger AI tool for real‑time and batch audio processing.","category":{"slug":"models","name":"Generative AI & Models","url":"https://digestai.news/category/models"},"entities":{"companies":["xAI","Tesla","Atlassian","Loom"],"models":["Grok Voice Transcribe 2.0","Grok Voice Transcribe 1.0","Gemini 3.5 Transcribe","ElevenLabs Scribe v2","Deepgram Nova-3","Whisper Large v3"],"people":["Sanchan Saxena"]},"firstPublishedAt":"2026-09-18T17:56:01Z","updatedAt":"2026-09-18T17:56:01Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"Unite.AI","title":"xAI Releases Grok Voice Transcribe 2.0 Speech-to-Text Model","url":"https://unite.ai/xai-releases-grok-voice-transcribe-2-0-speech-to-text-model","publishedAt":"2026-09-18T17:56:01Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"xAI releases Grok Voice Transcribe 2.0 speech-to-text model\", 18 September 2026, https://digestai.news/story/xai-releases-grok-voice-transcribe-2-0-speech-to-text-model","publisher":"Digest AI","title":"xAI releases Grok Voice Transcribe 2.0 speech-to-text model","datePublished":"2026-09-18T17:56:01Z","url":"https://digestai.news/story/xai-releases-grok-voice-transcribe-2-0-speech-to-text-model"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}