4 hrs ago

India’s Sovereign AI Faces Test Beyond Twenty-Two Language Coverage

India’s Sovereign AI Faces Test Beyond Twenty-Two Language Coverage
22 languages, score 0 · thestatesman.com

India is building its own computer programs that can work in many Indian languages.

One study found that some advanced programs could not properly recognise important Khasi and Garo names and words.

Those programs received a score of zero on that test.

A smaller tool made with help from the local community scored very highly.

This shows that speaking grammatically is not the same as understanding a culture.

The national effort has funded many models and created a large shared computing system.

It has also collected more than 15,000 datasets, but not all of them contain deep cultural or linguistic material.

The article says communities should help create and check the data so the technology respects their meanings.

Key facts

Study result
Tested models recorded an F1 score of zero on the reported culturally focused recognition task.
Regional tagger result
A purpose-built regional tagger scored 0.964 on the same ground.
Tokenisation problem
Khasi diacritics and a Garo morpheme marker were corrupted in up to half of the tested cases.
Indigenous proposals funded
The government selected 20 model proposals from more than 500 applications.
Param2
BharatGen has released Param2, a 17-billion-parameter model it says generates text across all 22 scheduled languages.
Shared computing
The IndiaAI Mission assembled a common pool of tens of thousands of subsidised processors.
AIKosh datasets
AIKosh grew from a few hundred datasets at launch to more than 15,000.

Sources

Related news