2 hrs ago
Sarvam AI Launches Saaras V4 for Indian-Language Voice Applications
Sarvam AI has made a new computer model called Saaras V4.
It listens to speech and turns it into useful text.
The model understands English and 22 Indian languages.
It is designed to handle different accents, dialects, background noise and mixtures of languages.
It can write down speech, keep the exact words, change writing systems or translate the speech.
It can also work quickly while someone is speaking.
Sarvam AI says it was tested on English and Indian-language speech datasets.
Developers can use it through an API and software tools.
Saaras V4 supports English and 22 Indian languages, including code-mixed speech, accents, dialects and noisy audio.
The model offers five modes: verbatim, transcription, code-mixed text, transliteration and translation.
Sarvam AI reported a language-identification error rate of 5.22% across 22 languages and 2.9% across 10 major languages.
The model provides low-latency WebSocket streaming, with a reported time-to-first-token below 150 milliseconds.
Saaras V4 is available through Sarvam AI’s API, Python and Node.js SDKs, and several developer integrations.
- Who
- Sarvam AI, a Bengaluru-based company.
- What
- The company launched Saaras V4, a multilingual automatic speech recognition model with five speech-processing modes.
- Where
- The model is intended for voice applications in India and is available through the Sarvam AI API.
- When
- Sarvam AI announced it on 24 August.
- Why
- The company aims to expand voice AI use across Indian languages and support real-world, multilingual audio conditions.
Key facts
- Language support
- English and 22 Indian languages
- Speech modes
- Verbatim, transcribe, codemix, translit and translate
- Model architecture
- An audio encoder paired with a 3-billion-parameter hybrid state-space language model
- Language-identification error rate
- 5.22% across 22 languages; 2.9% across the 10 most widely spoken Indian languages
- Streaming speed
- Reported time-to-first-token of less than 150 milliseconds
- Audio processing
- Supports real-time WebSocket streaming and long-form audio; Sarvam AI says multi-minute recordings can be processed in about one second
- Developer access
- Available through the Sarvam AI API, Python and Node.js SDKs, Vercel AI SDK, LiveKit Agents and Pipecat Agents










