4 hrs ago
India’s Sovereign AI Faces Test Beyond Twenty-Two Language Coverage
India is building its own computer programs that can work in many Indian languages.
One study found that some advanced programs could not properly recognise important Khasi and Garo names and words.
Those programs received a score of zero on that test.
A smaller tool made with help from the local community scored very highly.
This shows that speaking grammatically is not the same as understanding a culture.
The national effort has funded many models and created a large shared computing system.
It has also collected more than 15,000 datasets, but not all of them contain deep cultural or linguistic material.
The article says communities should help create and check the data so the technology respects their meanings.
A study presented in the Association for Computational Linguistics record found tested language models scored zero on a Khasi and Garo culturally focused recognition task.
The models’ tokenisers also corrupted Khasi diacritics and a Garo morpheme marker in up to half of tested cases.
A community-built regional tagger scored 0.964 on the same task, highlighting the importance of local linguistic expertise.
Under the IndiaAI Mission, the government has funded 20 indigenous model proposals and assembled a subsidised pool of tens of thousands of processors.
The article calls for language- and region-specific scorecards, community review, error reporting, and stronger consent and ownership practices for linguistic data.
- Who
- Badal Nyalang, regional language researchers, community linguists, the IndiaAI Mission, and BharatGen are central to the discussion.
- What
- Research exposed severe weaknesses in language-model recognition of culturally significant Khasi and Garo terms, while a community-built tagger performed strongly.
- Where
- The research concerns Northeast Indian languages, including Khasi and Garo, while the national model effort operates across India.
- When
- The study appears in this year’s Association for Computational Linguistics record; the article also discusses current IndiaAI Mission developments.
- Why
- The gap reflects limited and uneven linguistic data, tokenisation problems, and insufficient community-led evaluation and governance.
Coverage and Scale
Depth and Community Fidelity
How to expand language capability
Coverage and Scale
Engineers argue that broad coverage should come first because modern models can transfer structure from data-rich languages to languages with less data, with depth improving as corpora grow.
Depth and Community Fidelity
Critics argue that transfer can impose the assumptions of English and Hindi, producing fluent text that misses local registers, names, cultural meanings, and dialectal distinctions.
How to measure progress
Coverage and Scale
A model covering all 22 scheduled languages is presented as an important technical and political achievement for India’s sovereign-AI effort.
Depth and Community Fidelity
The article argues that language counts, parameter totals, processor numbers, and dataset counts do not show whether models understand or distort particular communities.
How to build language datasets
Coverage and Scale
AIKosh is described as a useful national platform for pooling and reusing Indian-language datasets instead of rebuilding them separately.
Depth and Community Fidelity
The article says annotated cultural material requires slower community work, including consent, ownership decisions, expert checking, and public mechanisms for correcting errors.
Key facts
- Study result
- Tested models recorded an F1 score of zero on the reported culturally focused recognition task.
- Regional tagger result
- A purpose-built regional tagger scored 0.964 on the same ground.
- Tokenisation problem
- Khasi diacritics and a Garo morpheme marker were corrupted in up to half of the tested cases.
- Indigenous proposals funded
- The government selected 20 model proposals from more than 500 applications.
- Param2
- BharatGen has released Param2, a 17-billion-parameter model it says generates text across all 22 scheduled languages.
- Shared computing
- The IndiaAI Mission assembled a common pool of tens of thousands of subsidised processors.
- AIKosh datasets
- AIKosh grew from a few hundred datasets at launch to more than 15,000.









