South Asians are missing from global health databases: why this matters and what needs to change


Globally, more than one in 10 adults now live with diabetes. If you have South Asian roots, that risk is even higher, and it hits earlier than it does in many other populations. In India alone, the number of people with diabetes is projected to reach 125 million by 2045

Yet, when scientists try to understand why diseases such as diabetes and cardiovascular disease affect South Asians differently, they often have to rely on genetic data drawn from European populations.

Advances in artificial intelligence and machine learning are allowing scientists to mine vast amounts of genomic and health data to detect disease earlier, predict risk, monitor patients and tailor treatments to individuals. But at the heart of this changing landscape lies an old, constant problem – the data used to build these tools lack diversity.

The missing factor

Integrated biobanks such as the U.K. Biobank, which combine participants’ genomic information with electronic health records, environmental exposures and lifestyle data, have transformed biomedical research. These repositories have accelerated drug development, informed clinical guidelines and helped shape public health policy across the world.

But South Asians remain largely absent from these datasets. The NHGRI-EBI GWAS Catalogue, an online database of human genome-wide association studies, shows that between 2005 and 2025, more than 86% of participants in these studies were of European ancestry, while South Asians accounted for less than 1%.

“More than 20% of the world is being neglected in multi-modal data integration,” said Bhramar Mukherjee, senior associate dean of public health data science and data equity, Yale School of Public Health, United States. “It denies [them] the human right and opportunity to attain the maximal possible health”.

The pattern repeats in newer tools, too. A recent study published in Cell Genomics reviewed more than 13,500 samples across three major single-cell resources—the Human Cell Atlas, the Human Tumour Atlas Network and the PsychAD Consortium and found found a “striking, pervasive European overrepresentation and underrepresentation of Asian and Latino individuals.”  

“These single-cell atlases are becoming the reference maps for biology and medicine, and they are increasingly used to train the AI models that will shape future research and care,” said Kuan-lin Huang, senior author of the study, who is an associate professor of genetics and genomic sciences, and AI in human health, at the Icahn School of Medicine, United States.

The data gap is a health gap

These data points are more than just statistics. South Asians face higher rates of type 2 diabetes, cardiovascular disease and asthma than people of European ancestry, which means the tools built on European-heavy data are less accurate for the population that needs them most.

Take polygenic risk scores, which combine the effects of many genetic variants associated with a disease to estimate a person’s overall genetic risk. A 2023 study found that polygenic risk scores for multiple sclerosis were less accurate when applied to South Asian populations. 

“Most predictions about how variants affect gene expression or cell function are inferred from European datasets, and we don’t know which of those predictions hold in South Asians. This limits our ability to understand disease mechanisms and identify drug targets relevant to South Asian populations,” said Shweta Ramdas, a geneticist based in Bengaluru.

Even the well-established measures of disease risk can vary between populations. Genetic traits such as G6PD deficiency, which can cause a type of anaemia, vary considerably across South Asia, with some ethnic groups in Pakistan and Afghanistan carrying the trait at much higher rates than others. 

Another study from Sri Lanka found that cardiometabolic risk did not fit into a single metabolic syndrome profile. Even within the same population, men and women showed distinct patterns of obesity, blood sugar, cholesterol and blood pressure.

“Diagnostic thresholds, risk scores and prediction models developed predominantly from European populations should be validated and, where necessary, recalibrated using South Asian data. Locally generated evidence is essential for equitable and accurate health care,” said Athula Sumathipala, director, Institute for Research and Development in Health and Social Care, Sri Lanka.

South Asia is not one population

South Asia constitutes one of the most diverse human populations in the world, shaped by thousands of years of migration, cultural diversity, endogamy and consanguineous marriages. Much of the existing research does not reflect this diversity. South Asians, Southeast Asians, West Asians and other Asian populations are often lumped together, obscuring important differences between them.

India alone illustrates how much diversity can disappear when populations are treated as a single group. The GenomeIndia Project, launched in 2020 to capture the country’s genetic diversity and build a reference database for Indian populations, has already identified more than 40 million genetic variants unique to the Indian population.

“South Asia, and India in particular, cannot realistically be treated as one genetic block. We need to include distinct endogamous and tribal groups, not just a few urban cohorts,” said Dr. Ramdas. “A lot of these harmful variants aren’t seen anywhere else”.

Why is the data missing?

Research funding, institutions, registries, biobanks and large population cohorts have historically been built and sustained where the money already was, leaving low- and middle-income countries with inadequate laboratory infrastructure, biobanking facilities and trained personnel to run comparable studies at scale.

“The global health landscape remains deeply unequal. While over 90% of the world’s potential years of life lost occurred in low- and middle-income countries (LMICs), only about 10% of global health research funding addressed their health needs,” said Dr. Sumathipala. “This imbalance goes beyond money, shaping whose problems are studied, whose questions are prioritised and whose evidence informs health policy and practice”.

For most South Asian countries, genomic research can be difficult to prioritise given more immediate and pressing public health needs such as infectious diseases, maternal and child health and non-communicable diseases.

But the region has made strides in the health landscape. Sustained investments in genomics have enabled several large prospective cohorts and population datasets, including GenomeIndia, Phenome India, Longevity India, the Sri Lankan Twin Registry Biobank and the Pakistan Genome Resource. But these independent cohorts and biobanks are mostly focused on individual diseases or specific populations, and often use different systems for collecting and storing data, which makes it difficult to bring them together for large genetic studies. India, for example, has several sizeable cohorts, but no harmonised system yet exists that lets researchers within and across borders work across them easily.

Towards building a beginning

A recent perspective in the Lancet Regional Health – Southeast Asia, authored by scientists across India, Pakistan, Bangladesh and Sri Lanka, argues that the region risks being excluded from the genomic revolution unless it builds an infrastructure itself.

The authors propose building greater regional collaboration between existing biobanks and cohorts, while ensuring that South Asian researchers and institutions retain a meaningful role in how their data are used.They propose to build a system in which existing datasets can speak to each other, populations that have historically been overlooked can be included, and the researchers generating the data can share in the scientific benefits.

This is vital because if the underlying data continues to stay skewed, the AI models and clinical tools built on top of it will reproduce and repeat those biases—only at a much larger scale, 

“It may be late but still not too late,” said Dr. Sumathipala.

(Rupsy Khurana is science communication and outreach lead at the National Centre for Biological Science, Bengaluru. khurana.rupsy@gmail.com)

  • Related Posts

    New studies pursue the ‘perfect’ blend for coffee and health

    While many Indians have a fondness for coffee, almost everyone in the U.S. and Europe drinks coffee, and more of it, every day. | Photo Credit: Mike Kenneally/Unsplash Every morning,…

    Continue reading
    How Gaganyaan’s thermal protection system will survive re-entry | Explained

    Compared to rockets, re-entering modules have a unique and unforgiving set of challenges | Photo Credit: PTI The atmospheric phase of all space missions is challenging for both ascending rockets…

    Continue reading