The Genome India Project (GIP) is India’s flagship whole-genome sequencing mission, launched in 2020 by the Department of Biotechnology to sequence 10,000 Indians and build a national reference genome that reflects the country’s population diversity. The 10,000-genome target was met in early 2024, with 10,074 individuals sequenced, and in 2025 the dataset was opened to researchers through the Indian Biological Data Centre — the cohort analysis was published in Nature Genetics in April 2025.
The reference human genome that geneticists have used for two decades is built on a small number of donors of largely European ancestry. The drugs, the diagnostic tests, the genetic-risk panels, the polygenic scores, and the rare-disease databases that flow from it are all calibrated against that reference. For most of the world, including most of India’s population, the reference is a poor fit. Variants common in Indian sub-populations are absent. Variants present in the reference but rare in India are over-represented. The result is a quietly biased genomic infrastructure that under-serves the people who carry the largest share of human genetic diversity.
The Genome India Project is India’s response. It is a flagship mission of the Department of Biotechnology, launched in 2020, that has set out to sequence the whole genomes of 10,000 individuals representing the country’s ethnic, linguistic, and tribal diversity, build a national reference genome, and create a publicly accessible dataset for medical, agricultural, and population research. The first phase reached its 10,000-genome target in 2024, and the data is now hosted at the Indian Biological Data Centre.
For UPSC purposes, the Genome India Project intersects biotechnology, the Indian health system, data sovereignty, ethics in human research, and the broader story of India’s positioning in the genomics era. This article walks through the project’s design, the institutional architecture, the scientific outputs, and the policy questions it raises.
Quick Facts on the Genome India Project

The Genome India Project, GIP, is a flagship genomics initiative of the Department of Biotechnology under the Ministry of Science and Technology. It was approved by the Government of India in 2020 with a target of sequencing the whole genomes of 10,000 individuals representing India’s diverse populations.
The coordinating institution is the Centre for Brain Research at the Indian Institute of Science, Bangalore. The project consortium brings together about 20 academic and research institutions across the country, including the National Institute of Biomedical Genomics in Kalyani, the Centre for Cellular and Molecular Biology in Hyderabad, the Indian Institute of Technology Madras, and several institutes of the Indian Council of Medical Research.
Sample collection covered 99 distinct ethnic and linguistic populations across India, with explicit oversampling from tribal communities and under-represented language groups. Samples were collected with informed consent under approved ethics protocols. Whole-genome sequencing at typical depth of 30 times provides the raw data, which is then assembled, annotated, and deposited.
The target dataset is a national reference genome that captures the variants common in Indian sub-populations and that can be used as a baseline for future medical, population-genetic, and rare-disease research. The completed dataset is held at the Indian Biological Data Centre in Faridabad, India’s first institutionally managed life-sciences data repository, modelled on international counterparts such as the European Bioinformatics Institute and the United States National Center for Biotechnology Information.
What the Genome India Project Actually Is
A whole-genome sequence is the complete readout of an individual’s three billion base pairs of DNA. From that sequence, researchers can identify single-nucleotide variants, insertions and deletions, structural variants, copy-number changes, and patterns of population history. Aggregating thousands of genomes from a defined population produces a frequency map of variants that lets researchers separate common harmless variation from rare disease-causing changes.
The Genome India Project’s design has three layers. The first is sample collection from individuals in well-defined ethnic and linguistic populations across India. The second is whole-genome sequencing at the participating laboratories, with a standard pipeline for read alignment and variant calling. The third is data integration and analysis at the Centre for Brain Research and at the Indian Biological Data Centre, producing the population-level reference and the annotated variant catalogue.
The scientific output is a tiered set of resources. There is a population-frequency map of variants found across the 99 populations sampled. There is a list of medically relevant variants whose frequencies in Indian groups differ markedly from the reference, including pharmacogenomic variants that affect drug response. There are leads on rare-disease causal variants in specific Indian sub-populations. And there is a set of foundational data for studying population history, migration, and admixture in the South Asian region.
The Genome India Project is not the only genomic effort in India. The IndiGen programme of the Council of Scientific and Industrial Research, launched in 2019, sequenced about 1,000 individuals as a pilot to demonstrate domestic capability. The Genome India Project absorbs the lessons from IndiGen and scales the cohort by an order of magnitude.
Background and Historical Context
India’s genomic story predates the Genome India Project by a long way. The Centre for Cellular and Molecular Biology in Hyderabad, the National Institute of Biomedical Genomics in Kalyani, and several genetics laboratories at the All India Institute of Medical Sciences and the Indian Institutes of Science Education and Research had been studying Indian populations for decades, including the genetics of malaria resistance, sickle-cell disease, and inherited disorders concentrated in particular caste, tribal, and regional groups.
The international genome era began with the Human Genome Project, which produced the first reference assembly in 2003 at a cost of about three billion dollars and over a decade of work. By the late 2010s, falling sequencing costs, with whole-genome sequencing approaching a few hundred dollars per sample, had made population-scale projects financially feasible. The UK’s 100,000 Genomes Project, launched in 2013 and completed in 2018, became the leading model. China launched its own large-scale sequencing programmes. The All of Us programme in the United States targeted a million participants.
India’s response was twofold. The Council of Scientific and Industrial Research launched IndiGen in 2019 with the Centre for Cellular and Molecular Biology and the CSIR Institute of Genomics and Integrative Biology in Delhi. IndiGen sequenced about 1,000 Indians and demonstrated that domestic infrastructure, ethics frameworks, and analytics pipelines could deliver a population-genomic dataset. The Department of Biotechnology then sanctioned the Genome India Project in 2020 as the flagship mission, a tenfold scaling of IndiGen’s cohort with a deliberate emphasis on ethnic and linguistic diversity.
The first sequencing began in 2020. The COVID-19 pandemic delayed several rounds of fieldwork. By early 2024, the consortium reported that the 10,000-genome target had been achieved. The Government of India formally announced the completion at a ceremony at the Indian Institute of Science. The Indian Biological Data Centre was inaugurated in late 2022 and now hosts the dataset alongside other large biological data repositories.
Key Features of the Project
Population coverage is the project’s signature feature. The 99 ethnic and linguistic populations sampled span all four major language families spoken in India: Indo-European, Dravidian, Austroasiatic, and Tibeto-Burman. Tribal groups, including those in the Andaman and Nicobar Islands, central Indian forest belts, and the Northeast, were sampled in proportion higher than their share of the national population, to capture the deepest ancestral lineages and the most distinctive genetic variation.
Whole-genome sequencing rather than exome or array genotyping is the technical choice. Whole-genome sequencing reads the entire three-billion-base-pair sequence rather than just the protein-coding regions. It is more expensive but captures regulatory variants, structural variants, and non-coding regions that exome sequencing misses. The data quality at 30-fold depth is sufficient for most clinical and research applications.
Distributed sequencing and centralised analysis is the operational architecture. Sequencing was done at multiple institutional laboratories with standard reagents and pipelines, and the resulting data was uploaded to a central computing facility. The architecture leverages the network of biotechnology and genomics centres that the Department of Biotechnology has built up over two decades.
The Indian Biological Data Centre is the long-term custodian. Located at the Regional Centre for Biotechnology in Faridabad, IBDC hosts the Genome India dataset and provides controlled access for qualified researchers under data-sharing agreements. It is India’s parallel to international biological data repositories and is intended to become the default destination for human-genomics data generated in the country.
Why the Genome India Project Matters

Precision medicine is the most direct downstream application. Indian populations carry pharmacogenomic variants that differ in frequency from European populations, affecting drug response in ways that the standard international evidence base does not capture. CYP2D6, CYP2C9, and other drug-metabolising enzymes have variants that change the effective dose of common cardiovascular and psychiatric medications. A reference dataset built on Indian genomes lets clinicians and pharmaceutical regulators calibrate dosing guidelines for the Indian population.
Rare-disease diagnosis is the second application. Many genetic disorders are rare globally but concentrated in specific Indian populations because of historical patterns of endogamy. Identifying the causal variants in those populations is much faster when the reference dataset includes thousands of genomes from the same group. The work on autosomal-recessive disorders in particular benefits.
Agricultural and conservation applications are the third strand. The genomic infrastructure built for the Genome India Project also supports parallel programmes on Indian crop varieties and livestock breeds. The methodologies of population-scale sequencing transfer cleanly between human, animal, and plant applications, and the same data centres host the relevant datasets.
Population history and demographic research is the fourth. Indian populations preserve some of the deepest ancestral lineages and the richest genetic diversity in the world. The Genome India dataset, properly curated, becomes a global resource for understanding human migration, admixture, and adaptation, complementing similar national efforts in Europe, China, and the Americas.
A fifth, less-discussed reason is data sovereignty. Indian biological samples and genomic data have historically flowed to international laboratories, often as part of collaborations that left the originating institutions with limited downstream access. A national project hosted at a domestic data centre changes that equation, both for ethical reasons and for the long-term scientific economy.
Detailed Analysis: The 10,000-Genome Cohort
The cohort design balanced statistical power, geographic coverage, and budget. With 10,000 samples drawn from 99 populations, the project averages about 100 individuals per population, with larger samples for populations of high public health interest and smaller samples for very small tribal groups. This gives reasonable statistical power for detecting variants down to allele frequencies of around one per cent, with rarer variants detectable in aggregate.
Sample types and consent were managed through a standardised ethics framework. Local institutional ethics committees approved protocols for each participating site. Consent was obtained in the local language, with explicit clauses on data sharing, secondary use, and the right to withdraw. The DNA samples themselves were obtained from peripheral blood, processed at the local institutional laboratory, and forwarded for sequencing.
Sequencing pipelines used standard short-read platforms, primarily Illumina, supplemented by long-read platforms for selected samples to capture structural variation. Read alignment to the GRCh38 human reference assembly, variant calling using GATK best practices, and joint variant analysis across the cohort produced the integrated dataset.
Quality control included standard population-genomics filters: removing samples with high contamination or low depth, excluding variants with poor mapping quality, and flagging known genotyping errors. The dataset has been described in peer-reviewed publications by the consortium, with population-frequency tables, structural-variant maps, and pharmacogenomic loci as the headline outputs.
Comparative View: Genome India vs IndiGen vs International Programmes
IndiGen, the CSIR pilot, sequenced about 1,000 Indians from a smaller and more curated cohort, primarily focusing on the urban and accessible populations near the participating institutes in Delhi and Hyderabad. IndiGen’s value was as a proof of concept; its data informed pharmacogenomic and rare-disease analyses but did not capture the full ethnic diversity of the country.
The Genome India Project scales the cohort tenfold, spreads sample collection across the country, and emphasises ethnic and tribal diversity. The two projects are complementary; IndiGen data has been integrated into broader analyses, and the methodologies developed for IndiGen informed the Genome India pipeline.
The UK’s 100,000 Genomes Project, completed in 2018, focused on patients with rare diseases and cancer, generating clinically interpretable data alongside a population reference. The All of Us programme in the United States targets a million participants with broad health-data linkage, including environmental exposure and lifestyle factors. China’s national genomics programmes have produced substantial cohorts but with less open-access policy. Each of these programmes has features that the Genome India Project has selectively borrowed, particularly the ethics framework and the data-access model.
A key differentiator for Genome India is ethnic and linguistic diversity. The 99 populations sampled in India represent a level of within-country genetic variation that no European or East Asian programme can match. This makes the dataset uniquely valuable for understanding human variation and uniquely complex to manage in terms of consent and benefit-sharing.
Challenges and Ethical Considerations

Data privacy is the most immediate concern. Whole-genome sequences are intrinsically identifying. Even with names and direct identifiers removed, a genome can be re-linked to an individual through ancestry-database matching. Strong access controls, audit trails, and clear data-sharing agreements are essential. The Indian Biological Data Centre operates a controlled-access model with researcher vetting and data-use agreements.
Informed consent in vulnerable populations is the second concern. Tribal and isolated populations are particularly susceptible to consent procedures that do not fully translate the implications of genomic research. The Genome India Project has used local-language consent forms and community-level engagement, but the long-term ethics of using such data, particularly for commercial or pharmaceutical applications, remain a live debate.
Benefit-sharing is the third concern. The communities whose genomic data drives drug discovery or pharmacogenomic guidelines should benefit from the resulting medical advances. India’s existing legal framework, including the Biological Diversity Act, addresses some aspects of benefit-sharing, but the application to large-scale human-genomics projects is still evolving.
Data interpretation in clinical use is the fourth concern. Population-level frequency data is one input into clinical decisions. Translating it into pharmacogenomic guidelines, rare-disease diagnostics, and risk-stratification tools requires further work and validation in clinical trials. India’s clinical-genomics infrastructure is improving but still lags the genomic data infrastructure.
A fifth issue is sustainability and scaling. The next phase of the Genome India Project, with a target potentially in the hundreds of thousands of genomes, will require larger sequencing capacity, expanded data infrastructure, and continuing investment. Whether this is funded as a continuation of the current flagship or as a new programme is a policy question for the Department of Biotechnology and the broader science establishment.
Prelims Pointers
The Genome India Project was launched in 2020. The target is sequencing 10,000 whole genomes representing 99 populations across India. The nodal agency is the Department of Biotechnology, Ministry of Science and Technology. The coordinator is the Centre for Brain Research at the Indian Institute of Science, Bangalore. The dataset is held at the Indian Biological Data Centre in Faridabad. About 20 institutions participated in the consortium. The IndiGen project, run by the Council of Scientific and Industrial Research, sequenced about 1,000 Indians beginning in 2019 as a pilot. Whole-genome sequencing at 30-fold depth is the standard. The 10,000-genome target was reached in 2024.
Mains Practice Questions
- The Genome India Project aims to build a national reference genome by sequencing 10,000 Indians from 99 populations. Discuss the scientific rationale, the institutional architecture, and the policy implications of this flagship initiative. (250 words)
- Genomic data is intrinsically identifying and intrinsically valuable. Examine the ethical and legal challenges of large-scale human-genomics projects in India and the safeguards that the Genome India Project has put in place. (250 words)
- Compare the Genome India Project with international counterparts such as the UK’s 100,000 Genomes Project and the United States All of Us programme. What are the relative strengths of the Indian programme, and where does it need to invest further? (250 words)
Way Forward
The first task is making the data usable. A genomic dataset is not, on its own, a medical resource. It needs interpretation tools, pharmacogenomic guidelines, rare-disease diagnostic panels, and integration into clinical workflows. Building those layers, and making them accessible to public hospitals and primary-care clinicians, is the bridge from research output to public health outcome.
A second task is expanding the cohort. Indian populations are vast and diverse enough that 10,000 genomes, while a transformative dataset, is still a small fraction of what the country’s medical research and clinical genomics will need. A second-phase programme reaching one hundred thousand or a million genomes, with deeper clinical phenotyping, would put India among the global leaders in population genomics.
A third task is governance. A national framework for human-genomics data, covering consent, access, commercial use, benefit-sharing, and international collaboration, would fill a current gap. The Department of Biotechnology, the Indian Council of Medical Research, the Ministry of Health, and the Department of Science and Technology need a standing coordination mechanism.
A fourth task is industrial uptake. The pharmaceutical, diagnostic, and agricultural genomics sectors should be able to license or use Genome India data under appropriate safeguards. A clear pathway, with oversight and benefit-sharing, will accelerate the translation of the dataset into products.
The Genome India Project is one of the largest human-genomics undertakings ever launched in a developing country. It has met its initial cohort target and produced a national reference genome. The next decade is where the value of that resource will be tested, in the clinic, in agricultural breeding programmes, in conservation, and in the country’s wider biotechnology economy.
Frequently Asked Questions
What is the Genome India Project?
The Genome India Project, GIP, is a flagship genomics initiative of the Department of Biotechnology, launched in 2020, with the goal of sequencing the whole genomes of 10,000 individuals representing India’s diverse ethnic, linguistic, and tribal populations to build a national reference genome.
Who is the nodal agency for the Genome India Project?
The nodal agency is the Department of Biotechnology under the Ministry of Science and Technology. The coordinating institution is the Centre for Brain Research at the Indian Institute of Science, Bangalore. About twenty research institutions participate in the consortium.
How many populations were sampled?
The project sampled 99 distinct ethnic and linguistic populations across India, including Indo-European, Dravidian, Austroasiatic, and Tibeto-Burman language groups, with deliberate oversampling from tribal and under-represented communities.
Where is the Genome India dataset stored?
The dataset is held at the Indian Biological Data Centre, IBDC, in Faridabad, located at the Regional Centre for Biotechnology. IBDC is India’s first institutionally managed life-sciences data repository and operates a controlled-access model for researchers.
What is the difference between Genome India and IndiGen?
IndiGen is a pilot project of the Council of Scientific and Industrial Research, launched in 2019, that sequenced about 1,000 Indians to demonstrate domestic capability. The Genome India Project is the flagship initiative of the Department of Biotechnology, launched in 2020, scaling the cohort to 10,000 individuals with much broader population coverage.
What is whole-genome sequencing?
Whole-genome sequencing is the technique of reading an individual’s complete three-billion-base-pair DNA sequence. It captures variation in protein-coding genes, regulatory regions, and structural elements, providing more comprehensive data than exome sequencing or array-based genotyping.
Why does India need its own reference genome?
The international reference genome is built largely on European-ancestry donors and does not adequately represent variants found in Indian populations. An India-specific reference allows more accurate calibration of pharmacogenomic guidelines, rare-disease diagnostics, and risk-stratification tools for the Indian population.
What are the applications of the Genome India dataset?
Applications include precision medicine that tailors drug doses to genetic background, faster diagnosis of rare genetic diseases concentrated in specific Indian populations, agricultural genomics through shared methodology, conservation genomics for Indian wildlife, and population-history research on migration and admixture.
What ethical safeguards does the Genome India Project use?
The project uses local-language informed consent, ethics-committee approvals at every participating institution, controlled-access data sharing through the Indian Biological Data Centre, and consent provisions covering secondary use and the right to withdraw. Benefit-sharing and data-privacy frameworks are evolving alongside the project.
Has the 10,000-genome target been reached?
Yes. The consortium reported reaching the 10,000-genome target in 2024, and the Government of India formally announced the completion. The dataset is now hosted at the Indian Biological Data Centre, and the next phase of the project, potentially scaling to a much larger cohort with deeper clinical phenotyping, is under discussion.
Tell Google you want more of this.
Add Anantam IAS as a preferred sourceOne tap, and this site shows up more often in your own Top Stories, AI Overviews and AI Mode. Remove it any time.