Anantam IASPost · 2 June 2026

Sovereign AI and BharatGen: India’s Bid for Indigenous Foundation Models (UPSC Science & Tech)

Study Notes · General Studies · Governance · GS III · Science & Tech

Sovereign AI is a country's capability to build and control its own artificial intelligence — compute, data, foundation models and talent — without leaning on foreign providers. BharatGen and the IndiaAI Mission are India's first serious attempts at it. Here is what they actually are, where they sit, and what to watch.

For most of the last three years, every Indian who typed a prompt into a chatbot was, in a quiet way, borrowing. The model answering them was trained abroad, on data scraped mostly from English-language web pages, running on graphics chips bought from a single American company, hosted on cloud servers owned by three foreign firms. It worked beautifully — and it left India dependent on other people’s infrastructure for a technology the government has decided is as strategic as electricity or telecom. Sovereign AI is the response to that unease: the idea that a nation should be able to build and control its own artificial intelligence end to end, rather than rent it.

BharatGen is the clearest expression of that ambition. Launched in late 2024 as India’s first government-funded effort to build a homegrown, multilingual, multimodal large language model, it is a bet that the country can train its own foundation models — on Indian languages, Indian data and increasingly Indian compute — instead of waiting for Silicon Valley to get around to Bhojpuri. Sitting above it is the much larger IndiaAI Mission, a roughly ₹10,372-crore programme meant to supply the GPUs, datasets, skills and money that any such effort needs. For UPSC, this is a near-perfect GS3 science-and-technology story with a strong GS2 governance edge, and it is only getting more important. So it is worth understanding properly, not as a slogan but as a system.

What “Sovereign AI” Actually Means

Let’s pin the term down before it floats away, because “sovereign AI” is now used to mean almost anything. At its core it describes a country’s capability to build, own and control its own AI stack without depending on foreign providers for the parts that matter. And that stack has four pillars, each of which India has historically been weak in.

The first is compute — the raw processing power that trains and runs models, supplied by specialised chips called GPUs (graphics processing units) housed in vast data centres. Frontier models are trained on tens of thousands of these chips for months at a stretch, and the high-end ones come overwhelmingly from one company, Nvidia, in the United States. The second pillar is data: a foundation model is only as good as what it learns from, and a model trained mostly on English web text will be fluent in California and clumsy in Chhattisgarh. The third is the foundation models themselves — the large, general-purpose neural networks (the “GPT” in ChatGPT stands for one) that everything else is built on top of. And the fourth is talent: the researchers and engineers who can actually design, train and align these systems, a pool India exports more of than it retains.

Why does owning all four matter enough to spend public money on? Four reasons keep recurring in government statements. There is strategic autonomy — the same logic that drove India’s space and nuclear programmes, that you do not want a critical capability sitting behind a foreign export-control decision or a sudden change in another company’s pricing. There is data localisation and security: if Indian citizens’ records, government documents and defence-adjacent information are processed by models running on foreign soil, that is a sovereignty problem, not just a commercial one. There is linguistic and cultural representation — most global models simply do not understand India’s languages, scripts and contexts well, because they were not built to. And there is plain economic value: AI is expected to add hundreds of billions of dollars to India’s economy this decade, and a country that only consumes AI captures far less of that than one that produces it. Sovereign AI is the attempt to move India from the consumer column to the producer column.

Diagram of the four pillars of sovereign AI — compute, data, foundation models and talent — that a nation must control to be AI-independent
Sovereign AI is not one thing but a stack of four — and India has been thin on every layer.
Infographic listing the seven pillars of the IndiaAI Mission from Compute to Safe and Trusted AI
The IndiaAI Mission is the scaffolding meant to supply what sovereign AI needs: chips, data, skills and capital.

BharatGen: India’s First Government-Funded Foundation Model

So what exactly is BharatGen? It is India’s first government-funded initiative to build indigenous, multimodal, multilingual large language models, and it was unveiled on 30 September 2024. The word “multimodal” matters here — it means the system is designed to handle not just text but speech and images too, so it can read a scanned land record, listen to a farmer’s question in Marathi and reply in the same tongue. “Multilingual” is the whole point: BharatGen exists to serve India’s languages, not to bolt them on as an afterthought to an English model.

The institutional plumbing is worth getting right, because UPSC loves the chain of command. BharatGen runs under the National Mission on Interdisciplinary Cyber-Physical Systems (NM-ICPS), which sits in the Department of Science and Technology (DST) — so this is the science ministry’s project, distinct from the electronics ministry’s IndiaAI Mission, a distinction students routinely muddle. It is led by IIT Bombay, with the heavy lifting done by the TIH Foundation for IoT and IoE (a “Technology Innovation Hub” set up under NM-ICPS), and a consortium of partners including IIT Madras, IIT Kanpur, IIT Hyderabad, IIT Mandi, IIIT Hyderabad and IIM Indore. As the DST described it at launch, the goal is to “revolutionise public service delivery and boost citizen engagement” by building foundational models in language, speech and computer vision — with the first phase targeting roughly a two-year build.

And it has started shipping, which is the part that lifts this above press-release territory. BharatGen has released a family of models under the “Param” name. Param-1 is a 2.9-billion-parameter foundational text model trained on around 7.5 trillion tokens, with about a third of that data being Indian content — built bilingually on Hindi and English with a deliberate focus on fact-rich material. There is also Patram, billed as India’s first document-vision model, a roughly seven-billion-parameter system that can read complex Indian document layouts in Indian languages and English, the sort of thing that could one day parse a stack of handwritten government forms. The roadmap has pushed toward coverage of all 22 scheduled languages and open-source releases so that others can build on the work — the open release being a quiet but important policy choice, since it spreads the capability rather than locking it inside one lab. (“Parameters,” by the way, are the adjustable weights inside a model; more parameters usually means more capability but also far more compute to train, which is exactly the tension this whole story turns on.)

The IndiaAI Mission: The Scaffolding Around the Models

BharatGen would be a heroic but lonely effort without the IndiaAI Mission underneath it. Approved by the Union Cabinet in March 2024 with an outlay of about ₹10,371.92 crore over five years, and run by the Ministry of Electronics and Information Technology (MeitY), the mission is the country’s attempt to build the whole environment a homegrown AI industry needs, rather than fund one model at a time. Its tagline — “Making AI in India and Making AI Work for India” — captures both halves: build the capability, then point it at Indian problems.

The mission is organised around seven pillars, and they map almost exactly onto the gaps in the sovereign-AI stack. IndiaAI Compute Capacity was the headline: a plan to assemble more than 10,000 GPUs through a public-private partnership and offer them to startups and researchers at heavily subsidised rates, so that a small Indian team need not pay frontier-lab prices for the chips it cannot otherwise afford. IndiaAI Innovation Centre is the pillar dedicated to developing indigenous large multimodal models and domain-specific foundation models — the home, in policy terms, for exactly the kind of work BharatGen and others are doing. IndiaAI Datasets Platform, branded AIKosh, gathers and cleans high-quality non-personal datasets for ethical use, because models need fuel. Then come Application Development (solving problem statements thrown up by ministries and states), FutureSkills (expanding AI courses and supporting thousands of students and scholars), Startup Financing (patient, risk-taking capital for deep-tech founders), and Safe & Trusted AI (the responsible-AI and governance pillar). Take them together and you have compute, data, models, talent, money and guardrails — the sovereign-AI checklist, pillar for pillar.

The numbers have since outrun the original plan, which is the rare good kind of slippage. Against the initial 10,000-GPU target, the government has reported assembling a far larger pool — figures of around 38,000 GPUs have been cited as the mission scaled — offered at subsidised rates as low as ₹65 per GPU-hour, a fraction of commercial cloud pricing. That common-compute model is the mission’s most distinctive idea: rather than every startup building or renting its own cluster, the state aggregates compute and rents it out cheaply, the way it once built shared industrial estates. Whether the demand, the data and the talent can keep pace with all that silicon is the open question — but the silicon, at least, is showing up.

The Push for Indigenous Models — and Who Is Building Them

The mission’s most-watched move has been to stop merely funding infrastructure and start commissioning actual sovereign foundation models. Under the IndiaAI Mission, the government invited proposals to build India’s own foundation models from scratch — and in 2025 it began selecting teams, handing them dedicated subsidised compute and, in effect, a national mandate. The headline selection came in April 2025, when the startup Sarvam AI was chosen to build a sovereign LLM, with a proposal aiming for a large model (reported in the 120-billion-parameter range) trained with a meaningful share of Indian-language data. Sarvam has since released models such as Sarvam-M, a 24-billion-parameter system tuned for Indian languages and reasoning, and pitched itself as a “full-stack” sovereign platform spanning speech, translation and conversational agents across 22 Indian languages. Several other teams — among the dozen-odd selected, names like Soket AI, Gnani AI and Gan AI have appeared — were picked to build models of varying sizes and specialisms.

It helps to be honest about how crowded and how fragile this field is, because UPSC rewards nuance over cheerleading. Beyond the government-backed efforts sit private bets like Krutrim, the Ola founder Bhavish Aggarwal’s venture, which became India’s first AI unicorn in January 2024 and open-sourced models such as Krutrim-2, a 12-billion-parameter multilingual system. But Krutrim’s trajectory is itself a cautionary tale: by late 2025 it was reported to be pivoting from building frontier models toward cloud services, a reminder that training large models is brutally expensive and that several Indian players may end up consuming compute rather than producing frontier intelligence. So the picture is not “India has built a rival to GPT.” It is more accurate to say India now has a portfolio — a government science-ministry effort (BharatGen), a flagship MeitY-funded sovereign LLM (Sarvam), and a scatter of private and open-source models (Krutrim and others) — all of them small to mid-sized by global standards, all of them betting that Indian-language depth and low-cost compute can carve out a defensible niche even if they never match the largest American or Chinese frontier models head-on.

Why It Matters — and Where It Could Fail

The strongest case for all this is linguistic inclusion, and it is not abstract. India recognises 22 scheduled languages and hundreds of mother tongues, and a citizen who speaks only Maithili or Santali gains nothing from a model fluent in English alone. Tie indigenous foundation models to Bhashini, the government’s national language-translation platform, and you get a plausible route to AI that answers a pension query in Odia, explains a crop advisory in Telugu or reads a court order aloud in Tamil. That is governance, agriculture and health delivery at population scale — the kind of “AI for public good” that justifies spending tax money rather than leaving it to the market, which will always serve the largest, richest language first. Add strategic autonomy (not being hostage to a foreign export-control decision on chips) and the prize is real.

But the obstacles are equally real, and a balanced answer names them. The deepest is compute and chip dependence: India is buying GPUs, not making them, so “sovereign” AI still runs on imported silicon, and a serious indigenous chip remains years away. There is the cost of frontier training — a single large model can cost tens of millions of dollars in compute alone, which is why even well-funded Indian players are mid-sized, and why some are quietly retreating to cloud services. There is data quality: high-grade, clean, rights-cleared Indian-language text is genuinely scarce, and a model fed thin data will stay thin. There is the talent drain, with India’s best AI researchers often working abroad. There is energy, since data centres are voracious power consumers in a country still expanding its grid. And there is governance and safety: the Digital Personal Data Protection (DPDP) Act of 2023 — whose detailed Rules were notified in 2025 — governs how personal data may be used to train models, and in November 2025 MeitY released India AI Governance Guidelines built around principles it calls the “Seven Sutras,” choosing a light-touch, pro-innovation approach (using existing laws and a new AI Governance Group rather than an EU-style omnibus AI Act). Whether that balance protects citizens without strangling the very industry it is trying to grow is the live debate to watch. So the honest verdict: India has, for the first time, the money, the chips and the political will to attempt sovereign AI — but ambition is not the same as arrival, and the next three years will decide which it turns out to be.

For Your Mains Answer

This topic is built for GS Paper 3, under “developments in science and technology” and “indigenisation of technology and developing new technology,” and it carries a strong GS Paper 2 governance angle whenever a question touches digital public infrastructure, e-governance or India’s regulatory approach to emerging tech. It is also prime Essay material on technology, self-reliance and development. The examiner is not testing whether you can spell “BharatGen” — they are testing whether you can explain why a country builds its own AI and weigh the trade-offs.

How to Build the Answer

Open by defining sovereign AI as a four-pillar capability — compute, data, models, talent — rather than a buzzword. Then ground it: BharatGen under DST/NM-ICPS led by IIT Bombay, sitting inside the wider IndiaAI Mission under MeitY. Move to significance (linguistic inclusion via Bhashini, strategic autonomy, governance delivery), then to challenges (chip dependence, training cost, data, talent, energy, regulation). Close with a measured line on where India actually stands. That arc — define, locate, justify, critique, conclude — works for almost any “indigenous technology” question.

Common Mistakes to Avoid

Don’t conflate BharatGen (DST, science ministry) with the IndiaAI Mission (MeitY, electronics ministry) — keeping the two ministries straight signals real reading. Don’t claim India has “built a ChatGPT rival”; the honest framing is mid-sized, language-focused models. Don’t drown the answer in parameter counts; one or two anchors are plenty. And don’t present sovereign AI as cost-free — the chip and energy dependence is the point that earns marks.

A Compact Answer Spine

Sovereign AI = control of compute + data + models + talent → why it matters (autonomy, localisation, language, economy) → BharatGen (DST/NM-ICPS, IIT Bombay, multimodal multilingual, Param models) → IndiaAI Mission (~₹10,372 cr, seven pillars, 10,000+ GPUs scaled up, subsidised common compute) → indigenous-model push (Sarvam selected 2025, others) → significance (Bhashini, 22 languages, governance) → challenges (chips, cost, data, talent, energy, DPDP Act + 2025 AI Governance Guidelines) → balanced conclusion.

Diagram or Flowchart Idea

Draw a four-layer pyramid for the sovereign-AI stack — Talent at the apex, then Models (BharatGen, Sarvam), then Data (AIKosh, Bhashini), then Compute (GPUs, data centres) as the base — and label the whole thing “what a country must own to be AI-sovereign.” A clean stack diagram like this earns marks fast and is quick to reproduce under time pressure.

A Balanced-Conclusion Line

“India now has, for the first time, the capital, the compute and the political will to attempt sovereign AI; turning that scaffolding into frontier-grade, trustworthy models — without importing the chips, exporting the talent or over-regulating the industry — is the test of the next three years.”

How to Use Data Without Cramming

Carry just a handful of anchors and deploy them precisely: the IndiaAI Mission outlay (~₹10,372 crore over five years), the GPU target (10,000+, scaled toward ~38,000), BharatGen’s launch (30 September 2024, DST/NM-ICPS, IIT Bombay), and one governance marker (DPDP Act 2023 plus the 2025 AI Governance Guidelines). Attribute them in prose — “as MeitY framed the IndiaAI Mission” — and you read like someone who has followed the policy, not memorised a list.

FAQ

What is sovereign AI in simple terms? Sovereign AI is a country’s ability to build, own and control its own artificial-intelligence capability — the compute (GPUs and data centres), the data, the foundation models and the talent — without depending on foreign providers for the parts that matter most. It matters for strategic autonomy, data security, representation of local languages and capturing the economic value of AI at home rather than abroad.

What is BharatGen and who runs it? BharatGen is India’s first government-funded initiative to build indigenous, multimodal, multilingual large language models, launched on 30 September 2024. It runs under the National Mission on Interdisciplinary Cyber-Physical Systems (NM-ICPS) of the Department of Science and Technology, is led by IIT Bombay through the TIH Foundation, and works with partners including IIT Madras, IIT Kanpur and IIIT Hyderabad. It has released models in the “Param” family aimed at India’s languages.

How is BharatGen different from the IndiaAI Mission? They sit in different ministries. BharatGen is a specific model-building project under the Department of Science and Technology. The IndiaAI Mission is the much larger ₹10,372-crore programme under the Ministry of Electronics and Information Technology that supplies the surrounding ecosystem — subsidised GPUs, datasets (AIKosh), skills, startup funding and an Innovation Centre — and has separately selected startups such as Sarvam AI to build sovereign foundation models.

What are the main challenges to India’s sovereign-AI ambition? The biggest is compute and chip dependence — India buys GPUs rather than making them. Others include the enormous cost of training frontier models, a shortage of clean Indian-language data, talent migrating abroad, the heavy energy demand of data centres, and the need to regulate AI (via the DPDP Act 2023 and the 2025 India AI Governance Guidelines) without choking the industry it is trying to grow.

Practice Questions

Prelims MCQs

  1. With reference to BharatGen, consider the following statements:
    (a) It is led by IIT Bombay
    (b) It functions under the National Mission on Interdisciplinary Cyber-Physical Systems of the Department of Science and Technology
    (c) It is designed as a multimodal, multilingual large language model
    (d) It was launched in 2022.
    Which of the statements is/are correct?
    Answer: (a),
    (b) and (c)
    — BharatGen is an IIT Bombay-led, DST/NM-ICPS multimodal multilingual LLM launched on 30 September 2024, not in 2022.
  2. The IndiaAI Mission, approved by the Union Cabinet in March 2024, is implemented primarily by which ministry, and what is its approximate outlay?
    (a) Ministry of Science and Technology; ₹5,000 crore
    (b) Ministry of Electronics and Information Technology; about ₹10,372 crore
    (c) Ministry of Finance; about ₹20,000 crore
    (d) NITI Aayog; ₹1,000 crore.
    Answer: (b) — The IndiaAI Mission is run by MeitY with an outlay of about ₹10,371.92 crore over five years.
  3. Which of the following is NOT one of the seven pillars of the IndiaAI Mission?
    (a) IndiaAI Compute Capacity
    (b) IndiaAI Datasets Platform (AIKosh)
    (c) IndiaAI Innovation Centre
    (d) IndiaAI Sovereign Wealth Fund.
    Answer: (d) — The seven pillars are Compute, Innovation Centre, Datasets Platform (AIKosh), Application Development, FutureSkills, Startup Financing and Safe & Trusted AI; there is no “Sovereign Wealth Fund” pillar.
  4. “Bhashini,” frequently linked to India’s indigenous-AI efforts, is best described as:
    (a) a domestic GPU-manufacturing programme
    (b) a national language-translation and digital-inclusion platform
    (c) a data-localisation law
    (d) a semiconductor fabrication plant.
    Answer: (b) — Bhashini is India’s national platform for language translation and digital inclusion across Indian languages, complementing indigenous language models.
  5. Consider the following with reference to India’s AI governance:
    (a) The Digital Personal Data Protection Act was enacted in 2023
    (b) Its detailed Rules were notified in 2025
    (c) MeitY released India AI Governance Guidelines built around “Seven Sutras” in 2025
    (d) India enacted a binding EU-style omnibus AI Act in 2025. Which is/are correct?
    Answer: (a),
    (b) and (c)
    — India chose a light-touch, principles-based approach (the “Seven Sutras”) rather than an EU-style omnibus AI Act.

Mains Practice Questions

  1. “Sovereign AI is to the 2020s what the space and nuclear programmes were to an earlier India.” Critically examine the rationale for building indigenous foundation models, and assess India’s readiness across compute, data, models and talent. (15 marks, 250 words)
  2. Discuss the institutional architecture of India’s AI push, distinguishing the roles of BharatGen, the IndiaAI Mission and platforms such as Bhashini. How do these initiatives together advance the goal of self-reliance in artificial intelligence? (15 marks, 250 words)
  3. Indigenous large language models are often justified on grounds of linguistic inclusion and governance delivery. Evaluate this claim with reference to India’s 22 scheduled languages and the delivery of public services. (10 marks, 150 words)
  4. “India’s sovereign-AI ambition still runs on imported silicon.” Examine the challenge of compute and chip dependence, and suggest a realistic strategy for reducing it over the next decade. (15 marks, 250 words)
  5. Compare India’s light-touch, principles-based approach to AI governance — the DPDP Act 2023 and the 2025 AI Governance Guidelines — with more prescriptive global models. What are the trade-offs between innovation and citizen protection? (10 marks, 150 words)