Sarvam AI: IndiaAI Mission Pick, Sarvam 105B and the Sovereign AI Debate
Sarvam AI explained: the Bengaluru startup building India's sovereign LLMs under the IndiaAI Mission, its models and the debate around them.
Sarvam AI is a Bengaluru startup building large language models for Indian languages, and in April 2025 it was cleared to build India’s sovereign language model under the IndiaAI Mission. In February and March 2026 it released two models trained from scratch in India, Sarvam 30B and Sarvam 105B, and opened an assistant called Indus on the larger one. For anyone tracking India’s AI policy, Sarvam is the clearest test case of what “sovereign AI” means in practice.
Most readers carry two wrong ideas about it. The first is that Sarvam is a government body, when it is a private company (its legal name is Axonwise Private Limited) that receives public compute and money under a competitive scheme. The second is that “sovereign” means India no longer depends on anyone. This note sets out what the government actually picked, what the company has shipped and where the honest limits of the sovereignty claim lie.
What is Sarvam AI?
Sarvam AI is a private Indian company that builds foundation models, the large general-purpose AI models that other applications are built on, with a focus on Indian languages and voice. It is one of 12 teams chosen under the IndiaAI Mission’s foundation-model track.
| Fact | Detail |
|---|---|
| What it is | Private AI company; legal entity Axonwise Private Limited |
| Headquarters | Indiranagar, Bengaluru, Karnataka |
| Co-founders | Vivek Raghavan and Pratyush Kumar |
| First platform launch | 13 August 2024, with the open-source Sarvam 2B model |
| Government selection | April 2025, to build India’s sovereign LLM ecosystem under the IndiaAI Mission |
| Public support | Rs 246.72 crore in financial and compute support (PIB, February 2026) |
| Flagship models | Sarvam 30B and Sarvam 105B, open-sourced on 6 March 2026 |
| Speech and document models | Bulbul (text to speech), Saaras (speech to text), Sarvam Vision (documents) |
| Consumer product | Indus assistant, limited beta from 20 February 2026 |
The Rs 246.72 crore figure is the one to fix in memory. Set it against the Mission’s total outlay of Rs 10,371.92 crore and Sarvam’s share is a little over 2 percent of the whole programme, which tells you the Mission is spreading its bets.
Why did the IndiaAI Mission pick Sarvam?
The short answer is that the Mission wanted Indian firms to train base models from scratch, and Sarvam had already done it at small scale. To see why that mattered, you need the Mission’s design first.
The Union Cabinet approved the IndiaAI Mission in March 2024 with an outlay of Rs 10,371.92 crore over five years, run by MeitY through its IndiaAI division. The PIB’s December 2025 explainer lists seven pillars:
- Compute: a shared pool of GPUs (graphics processing units, the chips that do the parallel arithmetic AI training needs), which crossed 38,000 at a subsidised rate of about Rs 65 per hour.
- Application Development: AI tools for India-specific problems in health, farming, climate and governance.
- AIKosh: a national platform for datasets and models.
- Foundation Models: India’s own large models built on Indian data and languages.
- FutureSkills: fellowships and training for AI talent.
- Startup Financing: funding support for AI startups.
- Safe and Trusted AI: tools and frameworks for responsible use.
The Foundation Models pillar drew over 500 proposals. In two phases the Mission selected 12 teams: Sarvam AI, Soket AI, Gnani AI, Gan AI, Avaatar AI, the IIT Bombay consortium BharatGen, Zenteiq, Gen Loop, Intellihealth, Shodh AI, Fractal Analytics and Tech Mahindra Maker’s Lab. Sarvam’s approval came in April 2025.
What Sarvam had built before selection
By the time it applied, Sarvam had a track record. Its August 2024 launch introduced Sarvam 2B, which the company described as the first LLM trained from scratch by an Indian company on Indian compute, using about 4 trillion tokens across 10 Indian languages. A token is a chunk of text, roughly a word or part of a word, so “4 trillion tokens” is a measure of how much reading the model did during training.
That phrase “from scratch” is the whole secret of the selection. Plenty of Indian teams had taken a foreign open model and fine-tuned it on Hindi or Tamil text. Very few had built the base model themselves, and that is the capability the Mission was paying for.
What models has Sarvam released?
Sarvam’s output falls into two families: text models that reason and converse, and speech and document models that make the text models usable by people who don’t type in English. Treat them separately, because questions on “Sarvam’s models” usually mean the first family while government releases dwell on the second.
The text models
The headline release came on 6 March 2026, when Sarvam open-sourced Sarvam 30B and Sarvam 105B. The numbers are parameters, the adjustable values inside a model; more parameters generally means more capacity. Both are reasoning models trained in India on compute provided under the IndiaAI Mission, and both use a Mixture-of-Experts design.
Mixture-of-Experts is easier than it sounds. Picture a hospital with 128 specialists where each patient sees only the few doctors relevant to the complaint. The model has 128 expert sub-networks, and for each token only a handful are switched on. The hospital picture is good for the cost logic but it breaks in one place: the “experts” don’t hold neat subjects like cardiology, they learn whatever split of the work helps training. What the design buys is a large total parameter count without paying the full compute cost on every word.
The details the company published are worth keeping:
- Sarvam 30B is tuned for real-time use and powers Sarvam’s conversational agent platform; it was pre-trained on about 16 trillion tokens.
- Sarvam 105B is the larger reasoning model, pre-trained on about 12 trillion tokens, and it powers the Indus assistant.
- Weights are open: anyone can download them from AIKosh and Hugging Face and run them locally.
The speech and document models
The PIB’s February 2026 explainer leads with these, and for public services they matter more:
- Bulbul turns text into speech in 11 Indian languages with 39 voices.
- Saaras turns speech into text across all 22 scheduled languages, including 8 kHz telephone audio and code-mixed speech (a Hinglish sentence, for example).
- Sarvam Vision reads documents in 22-plus Indian languages, including mixed scripts and handwriting.
Why does telephone audio get its own line? Because a farmer calling a helpline on a basic phone produces exactly that low-quality 8 kHz signal, and a model trained only on clean studio audio fails on it.
Where is Sarvam’s work being used?
The government’s case for Sarvam rests on public-service use, and the February 2026 PIB release names three projects.
- UIDAI: a custom generative AI stack running inside UIDAI’s own on-premise systems to support Aadhaar services with voice interaction, fraud alerts and support in 10 Indian languages.
- Odisha: the state government and Sarvam are setting up a 50 MW Sovereign AI Capacity Hub, with use cases in mining, industrial safety and Odia-language skilling.
- Tamil Nadu: the state government and IIT Madras are developing Digital Sangam, described as a sovereign AI research park anchored by a 20 MW data centre.
The UIDAI detail is the most telling. Aadhaar data cannot be shipped to a foreign cloud API, so a model that runs on hardware the government controls is the only kind that fits. That single requirement explains much of the sovereign AI argument better than any slogan does.
Sarvam’s consumer face is Indus, opened in limited beta on 20 February 2026 during the AI Impact Summit week. Sarvam itself says the 105B model behind Indus is significantly smaller than the models running global chat apps, which is an unusually candid admission from a company about its own product.
Is Sarvam AI really sovereign? The debate
Sarvam is sovereign in the sense that matters most for public data, and not sovereign in the sense that would end all foreign dependence. Both halves need to be on the page.
The case for calling it sovereign is strong on four counts:
- Training happened in India, on Indian compute, with data pipelines the company built itself.
- Language coverage targets the 22 scheduled languages that global models handle unevenly.
- Deployment can happen on government-controlled hardware, as the UIDAI project shows.
- Open weights mean the model can’t be switched off or repriced by a foreign vendor.
Sarvam’s own framing is careful here. Its May 2025 statement described the goal as strategic autonomy rather than decoupling from the rest of the world, and promised that models trained under the Mission would be open-sourced under permissive licences.
The limits are just as real. The GPUs in the national compute pool are imported chips, which is why the AI story is tied to the India Semiconductor Mission. Scale is a second limit: by Sarvam’s own description, its largest model is smaller than the frontier systems it is compared with. A third concern is that public money and subsidised compute are flowing to a private firm, and open-sourcing is the main condition that keeps the benefit public.
One episode shows why the word “sovereign” needs care. In May 2025 Sarvam released Sarvam-M, a reasoning model built by post-training Mistral Small, a French open model with 24 billion parameters. It was a legitimate research release, but it was not a from-scratch Indian base model. The 30B and 105B models of 2026 are. When a model is called Indian, the first question to ask is whether it was trained from scratch or adapted from someone else’s base.
The balanced view is that India has moved from adapting foreign models to training its own at mid-size, which is real progress, while the hardware and the frontier remain out of reach for now. Neither the triumphant version nor the dismissive version survives contact with the facts.
Sarvam AI today
As of October 2026, Sarvam AI is a funded participant in the IndiaAI Mission with open-weight models in public use. The dated milestones run like this:
- April 2025: approved to build the sovereign LLM ecosystem.
- February 2026: Sarvam’s models launched at the AI Impact Summit 2026 in New Delhi, alongside models from BharatGen, Gnani and Soket, according to a PIB reply of 13 March 2026.
- 6 March 2026: Sarvam 30B and 105B open-sourced.
The same March 2026 reply said the compute pool stood at more than 38,000 GPUs through 14 service providers, with 20,000 more being added. The December 2025 status of the IndiaAI Mission gives the earlier figures if you want the trend. Sarvam has said it intends to train larger models next, including ones specialised for coding and multimodal conversation; until those ship, treat them as plans.
How to study Sarvam AI for exams
Sarvam sits in GS Paper III under science and technology, indigenisation of technology and IT, with a governance link in GS Paper II through digital public services. The broader frame is artificial intelligence in India and the mechanics of large language models, so read those first and treat Sarvam as the worked example.
The theme is tested through AI policy and applications, not through company trivia. Mains 2026 GS Paper III asked “What is agentic Artificial Intelligence (AI)? Explain its working. Describe its applications with suitable examples. Discuss the advantages, risks and challenges associated with agentic AI systems.” A Sarvam example fits naturally into the applications and risks parts.
Mains 2023 GS Paper III asked candidates to “Introduce the concept of Artificial Intelligence (AI). How does AI help clinical diagnosis? Do you perceive any threat to privacy of the individual in the use of AI in healthcare?”, where the on-premise UIDAI model is a ready illustration of privacy-safe deployment. On the Prelims side, 192 of the 1,403 questions in the Prelims question bank are Science and Technology, and AI schemes are a natural target.
Revision facts:
- IndiaAI Mission: approved March 2024, Rs 10,371.92 crore over five years, seven pillars.
- 12 foundation-model teams selected from over 500 proposals; Sarvam approved in April 2025.
- Sarvam’s support: Rs 246.72 crore in financial and compute support.
- Sarvam 30B and 105B: Mixture-of-Experts, 128 experts, trained from scratch in India, open-sourced 6 March 2026.
- Bulbul is text to speech; Saaras is speech to text (all 22 scheduled languages).
- Compute pool: 38,000-plus GPUs at about Rs 65 per hour.
- Indus is the consumer assistant on the 105B model.
The common confusions are easy to separate once named. Sarvam is a private company; BharatGen is a government-funded consortium led by IIT Bombay; Bhashini is MeitY’s language-translation platform, not a foundation-model builder. And Bulbul and Saaras get swapped constantly: Bulbul, the songbird, speaks; Saaras listens.
| Initiative | Who runs it | What it does |
|---|---|---|
| Sarvam AI | Private company, Bengaluru | Foundation models (30B, 105B), speech and document models |
| BharatGen | IIT Bombay-led consortium, government-funded | Multimodal Indian-language models under the IndiaAI Mission |
| Bhashini | MeitY | Translation and speech platform for Indian languages |
| AIKosh | IndiaAI Mission | Public platform hosting datasets and models, including Sarvam’s |
For your preparation, Sarvam is less a fact to memorise than an argument to own. If you can explain why “trained from scratch” matters, why an on-premise model suits Aadhaar and where the imported-chip limit bites, you can answer almost any question on sovereign AI with a concrete Indian example instead of a generic paragraph.
Frequently Asked Questions
What is Sarvam AI?
Sarvam AI is a Bengaluru-based private company that builds large language models and speech models for Indian languages. Its legal entity is Axonwise Private Limited, and its co-founders are Vivek Raghavan and Pratyush Kumar. It is one of 12 teams building foundation models under the IndiaAI Mission.
Is Sarvam AI a government company?
No. Sarvam AI is a private startup. It receives financial and compute support of Rs 246.72 crore under the IndiaAI Mission, a competitive government scheme, and models trained under the Mission are to be released as open source.
When was Sarvam AI selected under the IndiaAI Mission?
Sarvam AI received approval in April 2025 to build India’s sovereign large language model ecosystem. It is one of 12 foundation-model teams chosen from more than 500 proposals.
What are Sarvam 30B and Sarvam 105B?
They are reasoning models that Sarvam trained from scratch in India and open-sourced on 6 March 2026. Both use a Mixture-of-Experts design with 128 expert sub-networks. The 30B model is tuned for real-time conversation, and the 105B model powers the Indus assistant.
What is the difference between Bulbul and Saaras?
Bulbul is Sarvam’s text-to-speech model, available in 11 Indian languages with 39 voices. Saaras is its speech-to-text model and covers all 22 scheduled languages, including telephone-quality audio and code-mixed speech.
Why is Sarvam AI called a sovereign AI model?
Its main models were trained in India, on Indian compute, with in-house data pipelines, and their weights are open so they can run on hardware the government controls. The claim has limits, because the GPUs used for training are imported and the models are smaller than global frontier systems.
What is Indus by Sarvam AI?
Indus is Sarvam’s AI assistant, opened in limited beta on 20 February 2026. It runs on the Sarvam 105B model and is aimed at reasoning tasks and Indian-context use.
How is Sarvam AI different from Bhashini?
Sarvam AI is a private company that trains foundation models. Bhashini is a MeitY platform that offers translation and speech tools across Indian languages. The two can work together, but one builds models while the other is a government service layer.
Practice Questions
Prelims
1. Consider the following statements about the IndiaAI Mission: 1. It was approved by the Union Cabinet in March 2024 with an outlay of over Rs 10,000 crore. 2. Its Foundation Models pillar selected a single company, Sarvam AI, to build all of India’s indigenous models. Which of the statements given above is/are correct?
- (a) 1 only
- (b) 2 only
- (c) Both 1 and 2
- (d) Neither 1 nor 2
Answer: (a) The outlay is Rs 10,371.92 crore; 12 teams were selected under the Foundation Models pillar, not one.
2. Consider the following statements about Sarvam AI’s models: 1. Saaras is a speech-to-text model covering all 22 scheduled languages. 2. Bulbul is a document-understanding model for handwritten text. Which of the statements given above is/are correct?
- (a) 1 only
- (b) 2 only
- (c) Both 1 and 2
- (d) Neither 1 nor 2
Answer: (a) Bulbul is a text-to-speech model; document understanding is handled by Sarvam Vision.
3. In the context of AI models such as Sarvam 105B, the term ‘Mixture-of-Experts’ refers to:
- (a) A panel of human reviewers who check every answer the model gives
- (b) An architecture in which only a few of many sub-networks are activated for each token
- (c) A method of combining several companies’ models into one product
- (d) A licensing rule that requires expert approval before weights are released
Answer: (b) Mixture-of-Experts routes each token to a small subset of expert sub-networks, keeping compute per token low.
4. Which one of the following is hosted by the IndiaAI Mission as a national platform for datasets and AI models?
- (a) Bhashini
- (b) AIKosh
- (c) AIRAWAT
- (d) iGOT Karmayogi
Answer: (b) AIKosh is the IndiaAI datasets and models platform; Sarvam’s 30B and 105B weights are available on it.
5. Consider the following pairs: 1. BharatGen: IIT Bombay-led consortium. 2. Sarvam AI: Bengaluru-based private company. 3. Bhashini: Foundation model built by Sarvam AI. How many of the pairs given above are correctly matched?
- (a) Only one
- (b) Only two
- (c) All three
- (d) None
Answer: (b) Bhashini is MeitY’s language-translation platform, not a Sarvam model.
Mains
- What do you understand by a ‘sovereign’ foundation model? Using the example of Sarvam AI, examine how far India’s AI efforts have achieved technological sovereignty. (15 marks, 250 words)
- What is agentic Artificial Intelligence (AI)? Explain its working. Describe its applications with suitable examples. Discuss the advantages, risks and challenges associated with agentic AI systems. (15 marks, 250 words) Previous year: Mains 2026, GS Paper III.
- Discuss the role of shared compute infrastructure under the IndiaAI Mission in lowering entry barriers for Indian AI startups. What are its limitations? (10 marks, 150 words)
- Indian-language speech and document models can widen access to public services. Examine with examples, and identify the privacy safeguards such deployments need. (10 marks, 150 words)
- Public funds and subsidised compute are being given to private firms to build national AI models. Critically evaluate this approach to building digital public goods in India. (15 marks, 250 words)