Bengaluru, India – The Indian Institute of Science (IISc) SPIRE Lab has officially launched SraVaani, a groundbreaking open-source speech artificial intelligence model. Released on August 13, 2026, the SraVaani model is designed to recognize speech across 65 diverse Indian languages and dialects. This initiative aims to bridge significant language gaps in AI technology, potentially benefiting over 25 crore (250 million) people whose native languages often lack robust digital support.
Developing Multilingual AI for India
The SraVaani model, developed by IISc's SPIRE Lab in partnership with ARTPARK and with support from Google, represents a major step towards inclusive AI for India. It covers 20 scheduled Indian languages and an additional 45 regional languages and dialects. This broad coverage includes languages such as Garo, Angika, Chakma, Kokborok, Tulu, Bundeli, and Bajjika, many of which existing speech recognition systems do not officially support. The system can convert spoken words into text across 10 different scripts and features automatic language identification, eliminating the need for users to manually select their language.[whalesbook+3]
The development of SraVaani relied on the extensive Vaani dataset, a large-scale project that collected over 31,000 hours of spontaneous speech. This data came from 156,000 individuals across 165 districts in 28 Indian states. The focus on natural, real-world speech, rather than rehearsed audio, helped the model perform significantly better on less common languages.[timesofindia+1]
Bridging Language Gaps with Advanced Technology
SraVaani is built upon a FastConformer architecture and was trained using a three-stage pipeline. This process included self-supervised pretraining on the Vaani corpus and an innovative audio-image representation alignment stage. This multimodal approach helps the speech encoder learn richer representations by connecting visual context with spoken content, especially improving recognition for low-resource languages.
In performance tests, SraVaani showed impressive results, particularly for underserved languages. For instance, it achieved a 9.5% word error rate on the Garo language. This marks a substantial improvement compared to other evaluated systems, which often recorded error rates as high as 69.4% for the same language. Alower error rate means the software makes fewer mistakes when converting speech to text. The model is freely available on the Hugging Face platform under an MIT license, allowing developers and startups to use, modify, and build new applications.[arxiv+4]
The Vision Behind Project Vaani
Professor Prasanta Kumar Ghosh, a professor at IISc and the principal investigator of Project Vaani, highlighted the core mission behind the initiative. "When we began Project Vaani four years ago, the aim was simple: that voice AI should work for every Indian, not only for those whose languages already had the resources behind them," Ghosh said. This vision underscores the commitment to inclusive language technology.[timesofindia]
The release of SraVaani is a key step towards building "Sovereign AI" in India. This concept emphasizes developing technology specifically tailored for local needs, rather than solely relying on global models that may not fully grasp regional linguistic nuances. By offering automated language identification and removing the need for manual language tagging, SraVaani simplifies the integration of voice features into various applications without requiring extensive technical expertise from developers.[whalesbook]
Future Prospects and Impact
SraVaani holds immense potential for various sectors, including banking, e-governance, and e-commerce. Reaching users in their native languages is crucial for growth strategies in these areas. Beyond these, the model can enhance accessibility in education, healthcare, and other public services, enabling more citizens to interact with digital platforms in their preferred language. The model's pan-India coverage spans 19 languages from the Northeast, 16 from eastern India, nine from the west, eight from the north, six from the south, and five from central India, alongside English and Sanskrit.[whalesbook]
The open-source nature of SraVaani fosters innovation within the Indian technology ecosystem. It lowers costs and technical barriers for companies aiming to build localized Indian applications, promoting the development of AI solutions that truly serve India's diverse linguistic landscape. This initiative is expected to drive further research and development in speech AI, creating a more inclusive digital future for all Indians.[thehindu]




