Current Affairs
science techUPSCState PCSSSCIBPSRRBCDS

SraVaani 1.0 — IISc Releases India's Broadest-Coverage Multilingual Speech Recognition Model: 65 Indian Languages and Dialects

17 August 2026 13 min read 61 IISc / ARTPARK / arXiv
Why in news

The SPIRE Lab at the Indian Institute of Science (IISc), Bengaluru, in collaboration with ARTPARK, released SraVaani-1.0 on August 13, 2026 — the broadest-coverage Automatic Speech Recognition (ASR) model for Indian languages reported to date, covering 65 languages and dialects (20 scheduled + 45 regional). The model achieves 9.5% word error rate on Garo vs 69.4% for the next-best system. It is open-source (MIT licence) on Hugging Face. Note: it is NOT India's first multilingual ASR model — AI4Bharat's IndicConformer (22 scheduled languages) predates it.

At a glance

Why in News

IISc's SPIRE Lab released SraVaani-1.0 (August 13, 2026) — the broadest-coverage Indic ASR model: 65 Indian languages and dialects, including 45 regional languages beyond the 22 Eighth-Schedule languages. MIT open-source licence.

What Changed

SraVaani-1.0 achieves 9.5% WER on Garo (vs 69.4% for the next-best system). Covers 65 languages in 10 scripts. ~430M parameters; ~900 MB. Available on Hugging Face.

Key Correction

NOT India's first multilingual ASR model — AI4Bharat's IndicConformer (IIT Madras + MeitY) covering all 22 scheduled languages predates it. SraVaani's claim: broadest coverage (65 vs 22 languages).

Significance

Enables voice-based digital services in low-resource tribal and regional languages. Supports Digital India, IndiaAI Mission, and constitutional obligations to linguistic minorities (Articles 350A, 350B, 8th Schedule).

Timeline

1909
IISc Founded
Indian Institute of Science, Bengaluru — India's premier research university; established with support from Jamsetji Tata and Maharaja of Mysore
2020
IN-SPACe / ARTPARK
AI & Robotics Technology Park (ARTPARK) established at IISc as industry-academia interface
Pre-2026
IndicConformer released
AI4Bharat (IIT Madras + MeitY) releases IndicConformer — first comprehensive government-backed Indic ASR system; covers all 22 scheduled languages
Pre-2026
Project Vaani launched
IISc-ARTPARK initiative to build India's largest multilingual speech dataset; 31,255 hours, 105 languages — pretraining corpus for SraVaani
Aug 12, 2026
arXiv preprint
SraVaani-1.0 paper published on arXiv
Aug 13, 2026
SraVaani-1.0 Released
Model released publicly on Hugging Face under MIT licence; 65 languages, 10 scripts, 430M parameters, 900 MB

Why in News

On August 13, 2026, the SPIRE Lab (Speech, Information and Recognition Engineering Lab) at the Indian Institute of Science (IISc), Bengaluru, in collaboration with ARTPARK (Artificial Intelligence and Robotics Technology Park at IISc), released SraVaani-1.0 — the broadest-coverage Automatic Speech Recognition (ASR) model for Indian languages reported to date. The model covers 65 Indian languages and dialects, comprising 20 of the 22 constitutionally scheduled languages and 45 additional regional languages and dialects. An arXiv preprint was published on August 12, 2026, and the model is publicly available under the MIT open-source licence on Hugging Face.

Important factual correction: SraVaani-1.0 is NOT India's first multilingual speech recognition model. AI4Bharat — a joint initiative of IIT Madras and the Ministry of Electronics and Information Technology (MeitY) — had already released IndicConformer, covering all 22 scheduled Indian languages. SraVaani's achievement is breadth of coverage — 65 languages, including low-resource tribal and regional languages that no previously published system had addressed — making it the broadest-coverage Indic ASR model to date.

Background

India's Linguistic Diversity

India is one of the most linguistically diverse nations on earth. The Census of 2011 recorded 121 major languages and over 19,500 distinct mother tongues. The Eighth Schedule of the Constitution currently recognises 22 official (scheduled) languages. However, hundreds of additional languages — spoken by tribal communities, hill populations, and regional groups — remain outside formal constitutional recognition yet serve as the daily medium of communication for millions of citizens.

This diversity poses a fundamental challenge for digital governance. Most digital systems — search engines, voice assistants, government portals — are built for English and, at best, a small number of major Indian languages. For the hundreds of millions of Indians who are not literate in English, voice-based interfaces in their own languages are a prerequisite for meaningful participation in the digital economy and governance.

Speech Recognition in India: Prior Developments

The government's National Language Technology Mission (NLTM), housed under MeitY, funded research into Indian-language tools including ASR systems. AI4Bharat (IIT Madras, supported by MeitY) released IndicConformer, the most comprehensive government-backed Indic ASR system, covering all 22 Eighth-Schedule languages. SraVaani-1.0 extends that coverage to 65 languages — adding Garo, Angami, Ao, Chakma, Kokborok, Tulu, Bundeli, Bajjika, Bhili, Gondi, Garhwali, Wancho, and others spoken by tribal and hill communities historically underserved by technology.

Current Developments

SraVaani-1.0: Technical Details

Built on the FastConformer architecture (~430 million parameters), SraVaani uses a hybrid TDT-CTC decoder and FP16 (16-bit floating-point) quantization, resulting in a deployed model size of approximately 900 megabytes — feasible to run on modern hardware without specialised infrastructure.

Training proceeded in three stages: (1) 31,255 hours of unlabelled audio from Project Vaani (105 languages) for self-supervised pretraining; (2) audio-image alignment using 11 million pairs; and (3) fine-tuning on 31,270 hours of labelled audio across 65 languages. The standard accuracy metric — Word Error Rate (WER) — is the percentage of words incorrectly transcribed. SraVaani achieves a WER of 9.5% on Garo versus 69.4% for the next-best evaluated system, demonstrating particular strength for low-resource languages where prior systems were effectively unusable.

Key Facts

ParameterDetail
Model nameSraVaani-1.0
Release dateAugust 13, 2026
arXiv preprintAugust 12, 2026
Developing institutionSPIRE Lab, IISc Bengaluru + ARTPARK
SupporterGoogle (no named government scheme)
Languages covered65 (20 scheduled + 45 regional/dialectal)
Scripts supported10
ArchitectureFastConformer (~430M parameters); hybrid TDT-CTC decoder
Model size~900 MB (FP16 quantised)
Training data31,255 hours unlabelled (Project Vaani) + 31,270 hours labelled (65 languages)
Best performance (Garo)9.5% WER vs 69.4% for next-best evaluated system
LicenceMIT (open-source); available on Hugging Face
Prior artIndicConformer (AI4Bharat, IIT Madras + MeitY) — 22 scheduled languages
Correct claimBroadest-coverage Indic ASR model to date (NOT "first multilingual")

Constitutional Provisions

  • Article 343 — Official Language of the Union: Hindi in the Devanagari script. SraVaani's Hindi coverage aligns with this priority while extending reach far beyond it.
  • Article 344 — Official Language Commission: recommends progressive use of Hindi for Union official purposes. ASR tools in Hindi and scheduled languages support this transition.
  • Article 350A — Directs States and local authorities to provide adequate facilities for instruction in the mother tongue at the primary stage of education for children of linguistic minority groups. Voice-based educational tools built on ASR models like SraVaani can assist practical implementation in tribal and remote areas.
  • Article 350B — Provides for a Special Officer for Linguistic Minorities appointed by the President to investigate safeguards. ASR tools for minority languages provide a technological complement to these constitutional protections.
  • Eighth Schedule — Lists the 22 official scheduled languages of India. SraVaani covers 20 of these 22 plus 45 additional regional languages and dialects.

Legal Framework

  • Official Languages Act, 1963 — Governs use of Hindi and English for official Union purposes. Government deployments of ASR in official functions must accommodate the multilingual requirements this Act establishes.
  • Digital Personal Data Protection Act, 2023 (DPDPA) — Voice data is personal data under Indian law. Applications collecting, processing, or storing users' voice recordings must comply with DPDPA requirements: lawful purpose, informed consent, data minimisation, purpose limitation, and security safeguards. Since SraVaani is open-source, third-party deployers bear independent DPDPA compliance responsibility.
  • Information Technology Act, 2000 — Overarching legislation governing electronic data, cybersecurity, and digital intermediaries. AI-based voice processing systems must comply with IT Act rules.

Institutional Framework

  • IISc Bengaluru — India's premier research university; established 1909 with support from Jamsetji Tata and the Maharaja of Mysore. SPIRE Lab is a speech technology research group within IISc.
  • ARTPARK (AI and Robotics Technology Park) — industry-academia interface body at IISc; co-developed Project Vaani (India's largest multilingual speech dataset used to pretrain SraVaani).
  • AI4Bharat (IIT Madras + MeitY) — established government-backed centre for Indic AI research; prior art with IndicConformer (22 scheduled languages); used as direct baseline in SraVaani authors' paper.
  • MeitY — Ministry of Electronics and Information Technology; nodal ministry for digital governance, AI policy, language technology, IndiaAI Mission, BharatGen, NLTM.
  • NLTM (National Language Technology Mission) — MeitY initiative to develop translation, transcription, and ASR tools for all 22 scheduled languages.

Science and Technology Dimensions

How Automatic Speech Recognition Works

ASR converts an audio signal into text through three steps: (1) the audio waveform is transformed into a numerical representation (spectrogram); (2) a deep learning model maps these representations to probable sequences of sounds and words; (3) a decoder selects the most likely text output. The FastConformer architecture simultaneously attends to short phoneme-level patterns and long sentence-level patterns — dramatically improving accuracy over older Hidden Markov Model (HMM) approaches that dominated the field before 2015.

Word Error Rate as a Metric

WER = (substitutions + insertions + deletions) / total reference words. A WER of 9.5% means approximately 90–91 words in every 100 are correctly transcribed. A WER of 69.4% means roughly only 31 words in 100 are correct — a level of inaccuracy that makes a system impractical for real-world use.

Open-Source AI and Sovereign Capability

The MIT licence allows any developer, researcher, or government agency to download, modify, and deploy SraVaani-1.0 without licensing fees. This reduces dependence on proprietary foreign ASR systems for processing the voice data of Indian citizens in Indian languages and enables Indian startups, state governments, and civil society to build voice-enabled applications without prohibitive technology costs.

Economic and Social Dimensions

  • Rural and tribal digital inclusion: Over 60% of India's population lives in rural areas where English literacy may be limited. Voice interfaces in local languages enable access to banking (PM Jan Dhan), healthcare (Ayushman Bharat), and agricultural advisory services (e-NAM) without requiring text literacy.
  • Frontline workers: ASHAs and Anganwadi workers in tribal areas can use voice-based data-entry tools to file reports in their own language, reducing administrative burden.
  • Government service delivery: Integration with UMANG, DigiLocker, and e-Gram Swaraj would allow citizens to navigate government services through voice commands in their mother tongue.
  • Language preservation: For endangered languages such as Wancho (Arunachal Pradesh), an ASR model creates a technological record complementing academic and governmental language preservation efforts.

Challenges

  • Data scarcity for low-resource languages: Sustained, community-inclusive data collection is required to improve and maintain performance over time for the 65 covered languages.
  • Speaker diversity: Accents, regional dialects, age-related patterns, and recording conditions vary widely within a single language; a model trained on a limited speaker pool may underperform for speakers outside that group.
  • Code-switching: Indian speakers routinely mix languages within a sentence (Hindi-English, Tamil-English). SraVaani is designed as monolingual-per-utterance and is not optimised for code-switched speech.
  • Absence of government funding for this specific model: SraVaani was supported by Google. Long-term maintenance and expansion require either continued private support or government funding, which has not been announced.
  • Data privacy and community trust: Any voice application built on SraVaani collects personal voice data. Building trust among tribal and rural communities, ensuring DPDPA compliance, and preventing misuse are significant governance challenges.

Government Initiatives

  • Digital India — flagship programme for digital governance and connectivity. Voice interfaces in regional languages represent the next frontier for inclusion.
  • IndiaAI Mission (2024) — ₹10,371 crore budget; builds India's AI compute infrastructure, curated datasets, and startup ecosystem. Open-source models like SraVaani align with its democratisation goals.
  • BharatGen — MeitY initiative to develop large AI models for Indian languages and cultural contexts. Complements SraVaani's open-source ASR capability.
  • National Language Technology Mission (NLTM) — builds tools for all 22 scheduled languages; institutional and financial support for Indic NLP research including ASR.
  • National AI Portal (ai.gov.in) — government's central platform for sharing AI resources; models like SraVaani can be catalogued here for maximum accessibility.

Way Forward

  • MeitY, Ministry of Tribal Affairs, and Ministry of Panchayati Raj should jointly evaluate integrating SraVaani-1.0 into UMANG, DigiLocker, and e-Gram Swaraj for voice-based service delivery for tribal and rural populations.
  • NLTM should fund data collection and model fine-tuning for the most underserved of the 65 covered languages, particularly languages of the northeastern states and central Indian tribal belts.
  • A privacy impact assessment and community data governance protocol should be developed before government-scale deployment, ensuring DPDPA 2023 compliance.
  • AI4Bharat and IISc SPIRE Lab should explore formal collaboration to pool datasets, avoid duplication, and work towards a unified, government-endorsed Indic ASR standard.

Possible Mains Questions

  1. "India's linguistic diversity is simultaneously a democratic imperative and a technological challenge for digital governance." Examine with reference to recent advances in Automatic Speech Recognition for Indian languages, particularly SraVaani-1.0. (GS Paper III / GS Paper II)
  2. "What are the constitutional and legal obligations of the Indian state towards its linguistic minorities? Evaluate the role that language technology tools can play in fulfilling these obligations, and identify governance gaps that remain." (GS Paper II)

Possible Prelims MCQs

  1. SraVaani-1.0, released in August 2026, is best described as:

    • (A) India's first multilingual speech recognition model
    • (B) The broadest-coverage Automatic Speech Recognition model for Indian languages reported to date
    • (C) A government-funded model covering all 22 scheduled languages
    • (D) A text-to-speech model developed by IIT Madras

    Answer: (B). SraVaani is the broadest-coverage Indic ASR model by language count (65), not the first. Developed by IISc, not IIT Madras. Supported by Google, not a government scheme. IndicConformer (AI4Bharat) predates it.

  2. Which institution developed SraVaani-1.0?

    • (A) AI4Bharat, IIT Madras
    • (B) C-DAC, Pune
    • (C) SPIRE Lab, IISc Bengaluru
    • (D) TCS Research, Hyderabad

    Answer: (C).

  3. Which Article directs states to provide facilities for instruction in the mother tongue at the primary stage for children of linguistic minority groups?

    • (A) Article 343
    • (B) Article 344
    • (C) Article 350A
    • (D) Article 350B

    Answer: (C) Article 350A. Article 350B provides for a Special Officer for Linguistic Minorities.

  4. Word Error Rate (WER) is defined as:

    • (A) The percentage of sentences transcribed with no errors
    • (B) The ratio of incorrectly transcribed words (substitutions + insertions + deletions) to total reference words
    • (C) The number of languages a model can process per second
    • (D) The ratio of the model's output file size to the audio input duration

    Answer: (B). Lower WER = higher accuracy.

  5. Project Vaani, which provided the pretraining dataset for SraVaani-1.0, is an initiative of:

    • (A) MeitY and NASSCOM
    • (B) IISc and ARTPARK
    • (C) IIT Madras and Google DeepMind
    • (D) ISRO and C-DAC

    Answer: (B) IISc and ARTPARK.

Essay Dimensions

  1. "Language, technology, and power: the availability or absence of language technology determines which linguistic communities can participate in the digital economy and governance."
  2. "Open-source AI as public infrastructure: foundational AI models for languages spoken by tens of millions of Indian citizens may warrant treatment as public goods."
  3. "The limits of technological inclusion: when technology reaches a community in its own language but was designed without community participation, the result may be a new form of technological paternalism."
  4. "Sovereignty and strategic dependency: India's most advanced language AI research is supported by a foreign corporation — what does this mean for data sovereignty and strategic autonomy?"
  5. "The census and the algorithm: decisions about which languages receive technological resources involve trade-offs between speaker population, political representation, linguistic endangerment, and financial viability."

Interview Questions

  1. SraVaani-1.0 covers 65 Indian languages. India has over 19,500 mother tongues. How should the government prioritise which languages receive ASR support?
  2. The model is supported by Google, not a government scheme. What are the risks and benefits of private-sector funding for foundational language AI that will eventually serve public governance functions?
  3. Article 350A mandates mother-tongue instruction at the primary level. How can speech recognition technology practically support implementation for children in the tribal areas of Chhattisgarh or Jharkhand?
  4. How does the DPDPA 2023 apply to a government agency that wishes to deploy SraVaani to accept voice inputs from citizens applying for welfare benefits? What specific safeguards would you design?
  5. SraVaani achieves a WER of 9.5% on Garo. A farmer using a crop disease advisory service cannot tolerate errors of that magnitude. How would you bridge the gap between benchmark accuracy and the reliability required for life-affecting decisions?

FAQ

Q1. Is SraVaani-1.0 India's first multilingual speech recognition system?
No. AI4Bharat (IIT Madras + MeitY) released IndicConformer — covering all 22 constitutionally scheduled languages — before SraVaani. SraVaani's distinction is coverage breadth: 65 languages, including 45 regional languages and dialects beyond the Eighth Schedule, making it the broadest-coverage Indic ASR model to date.
Q2. Can SraVaani-1.0 be used commercially without cost?
Yes. The MIT licence permits commercial use, modification, and redistribution without royalty payments. Organisations deploying it in products that collect voice data remain independently responsible for compliance with applicable Indian law, including the DPDPA 2023.
Q3. Which languages does SraVaani cover that IndicConformer did not?
SraVaani adds 45 regional languages and dialects beyond the 22 scheduled languages. These include Garo, Angami, Ao, Chakma, Kokborok, Tulu, Bundeli, Bajjika, Bhili, Gondi, Garhwali, and Wancho — spoken predominantly by tribal and hill communities.
Q4. What is ARTPARK, and what role did it play?
ARTPARK (AI and Robotics Technology Park) is an industry-academia interface body at IISc Bengaluru. It co-developed Project Vaani — the IISc-ARTPARK initiative to build India's largest multilingual speech dataset — whose 31,255 hours of unlabelled audio formed the pretraining corpus for SraVaani-1.0.

Further Reading

Constitutional provisions

Article 343

Official Language of the Union: Hindi in Devanagari script; English retained for official purposes. SraVaani's Hindi coverage aligns with this priority.

Article 344

Official Language Commission to recommend progressive use of Hindi. Speech recognition tools in Hindi and scheduled languages support this transition.

Article 350A

Directs States to provide adequate facilities for instruction in mother tongue at primary stage for children of linguistic minority groups. Voice-based educational tools built on ASR can assist implementation.

Article 350B

Special Officer for Linguistic Minorities appointed by the President to investigate safeguards. ASR tools for minority languages provide technological complement to constitutional protections.

Eighth Schedule

Lists 22 official scheduled languages of India. SraVaani covers 20 of these 22 + 45 additional regional languages and dialects.

Relevant Acts & Judgments

Acts
Digital Personal Data Protection Act, 2023 (DPDPA)
Voice data constitutes personal data. Any application collecting user voice recordings must comply with DPDPA: lawful purpose, informed consent, data minimisation, purpose limitation. SraVaani is open-source; deployers bear independent DPDPA compliance responsibility.
Official Languages Act, 1963
Governs use of Hindi and English for Union official purposes. Government deployments of speech recognition in official functions must accommodate multilingual requirements.
Information Technology Act, 2000
Overarching legislation governing electronic data, cybersecurity, and digital intermediaries. AI-based voice processing systems must comply with IT Act rules.
Key distinction: Don't confuse SraVaani (IISc + ARTPARK; 65 languages; broadest coverage; Google-supported; not government-funded) with IndicConformer (AI4Bharat; IIT Madras + MeitY; 22 scheduled languages; government-backed). SraVaani extends coverage; it does NOT replace IndicConformer as the government's official Indic ASR system.
GS-IIIScience and TechnologyAISpeech RecognitionIIScLinguistic DiversityDigital IndiaIndiaAI Mission8th ScheduleArticle 350ALanguage TechnologyNLTM

0 Comments

Sign in to join the discussion.

SraVaani IISc Multilingual Speech Recognition 65 Languages 2026 | UPSC | UPSC.wiki