AI Sovereignty: From Language Coverage to Cultural Representation
AI Sovereignty: From Language Coverage to Cultural Representation

Context: India’s Multilingual AI Push
- India is developing indigenous AI models under the IndiaAI Mission, aiming to build AI capabilities suited to Indian languages and contexts.
- A major objective is to expand AI access across the 22 languages in the Eighth Schedule of the Constitution.
- BharatGen, an IIT Bombay-led consortium, is developing multilingual foundation models capable of processing Indian languages.
- The IndiaAI Mission has created a subsidised national compute ecosystem, with tens of thousands of processors made available to support AI development.
- The article highlights an important distinction: language coverage does not automatically mean linguistic or cultural understanding.
The Problem: Breadth vs Depth
- A model may generate grammatically correct text in a language while remaining weak at understanding its cultural context.
- Testing cited in the article found major problems in recognising Khasi names, festivals, clan terminology and linguistic markers.
- Diacritics and morphemes carrying meaning can be incorrectly processed, potentially distorting names and cultural references.
- Regional or community-built language tools performed better when trained using terminology developed and validated by the communities themselves.
- This demonstrates that AI performance cannot be assessed only through:
- Number of languages supported.
- Size of datasets.
- Number of model parameters.
- Computing capacity.
- Fluency ≠ fidelity: A model may produce convincing sentences while misunderstanding the people, cultural references and meanings embedded in a language.
Why Data Quality and Community Participation Matter?
- Indian languages differ significantly in their availability of high-quality training data.
- Data-rich languages can transfer linguistic structures to data-poor languages through multilingual model training, but this can also transfer biases and blind spots.
- Smaller and underrepresented languages may therefore require directly curated linguistic resources, rather than relying primarily on transfer from English or other data-rich languages.
- The article highlights AIKosh, a public platform intended to provide datasets for Indian-language AI development.
- However, a large number of datasets does not necessarily constitute a sufficiently rich corpus.
- Important resources include:
- Transcribed speech.
- Community-validated named entities.
- Dialectal material.
- Cultural and linguistic annotations.
- Building such datasets requires participation from linguists, scholars and language communities.
AI Sovereignty and Data Governance
- AI sovereignty requires more than domestic computing infrastructure or indigenous foundation models.
- A genuinely sovereign AI ecosystem must also preserve the meaning, context and cultural knowledge embedded in Indian languages.
- Creating national-scale language corpora raises questions about:
- Who owns the recordings and linguistic material?
- Whether communities have consented to their use.
- How the data will be stored and accessed.
- Who benefits from models trained on community-generated material.
- Careless collection can reproduce cultural distortions at scale.
- Therefore, linguistic AI development requires ethical data governance, community consent and transparent provenance.
Way Forward: Measuring Linguistic Depth
- Publicly funded AI models should disclose performance by language and region, rather than presenting aggregate multilingual scores alone.
- Models should be tested before deployment by linguists and community scholars, alongside technical evaluators.
- Establish public mechanisms through which communities can report and correct errors involving:
- Names
- Festivals
- Cultural terms
- Sacred or sensitive expressions
- Build high-quality, community-validated corpora for underrepresented languages instead of relying excessively on cross-lingual transfer.
- AIKosh and similar public data platforms can support a more open ecosystem if accompanied by strong standards for data quality, consent and provenance.
- India's AI sovereignty should therefore be measured not simply by how many languages an AI model can generate, but by how accurately and respectfully it understands the communities and civilisations represented by those languages
