Saaras V4 uses a 3-billion-parameter decoder trained from scratch, giving Sarvam full control over how the model handles Indian speech.
The model produces five transcript formats from one audio file, removing the need for separate tools at each stage.
Keyterm prompting lets developers guide recognition toward brand names and technical terms, a feature aimed squarely at real business use.
A phone call in India rarely stays in one language for long. A customer might open in Hindi, switch to English mid-sentence, and close with a regional word that has no clean translation. Most speech recognition tools were not built for this kind of switching. They expect one language at a time, spoken clearly, without interruption.
Sarvam AI, an Indian startup working on language technology, has released Saaras V4 to close that gap. The model reads code-mixed speech, handles background noise, and adjusts to regional accents without extra tuning. That focus places it in a different category from ASR tools designed mainly for English.
Saaras V4 runs on two main parts: an audio encoder and a language decoder. Sarvam says the decoder, a 3-billion-parameter hybrid state-space model, was trained from the ground up rather than adapted from an existing system. That choice gives the company direct say over how the model learns Indian sounds and speech patterns, rather than inheriting the habits of a model built for other languages first.
The result shows up in the numbers. Sarvam reports state-of-the-art performance across all 22 Indian languages the model supports, including several with limited digital speech data available for training.
Source link







