When AI that Works in Delhi Fails in Sonipat - Ashoka University

Other links:

Other links:

When AI that Works in Delhi Fails in Sonipat

Suvrankar Datta, Faculty Fellow, Koita Centre for Digital Health - Ashoka, shares why we need to rethink how we evaluate health AI and start asking the right questions.

AI in healthcare and wellness is the coolest thing to work on right now, with thousands of investors and longevity enthusiasts pouring in billions of dollars in recent years to ‘disrupt’ and capture a piece of a potentially trillion-dollar market.I will start by sharing an incident. What I am sharing with you is not just an isolated example but a pattern we keep seeing repeatedly now across institutions and geographies, and it captures a major problem in health AI today that I believe deserves far more attention than it currently receives.

A few years back, a team of brilliant engineers and doctors in a private South Delhi hospital spent two years building an AI system to read chest X-rays. The system performed brilliantly on their test data with 95% accuracy, clean ROC curves, a paper accepted at a reputable venue, and investors were understandably excited. The government had also shown interest. Everyone agreed: scale it up, take it to the people who need it most!

So the system was deployed at a primary health centre in rural Sonipat. The X-ray machine there was an older model. The images were grainier, often with suboptimal positioning. The patient population was fundamentally different and had a much different burden of tuberculosis, occupational lung diseases, and the conditions the model had been designed and trained to detect. And the radiologist who was supposed to oversee the AI’s outputs and catch its mistakes? There was no radiologist. That was the whole point of the deployment: the AI radiologist (like AI doctors being announced by startups in SF every week in 2026) was supposed to fill that gap.

Within weeks, the system was confidently generating reports that no one on the ground had the training or infrastructure to verify. Some of those reports were wrong. Nobody knew which ones. And nobody had set up a mechanism to find out. The poor results were discovered by the clinical research team once they decided to do a QC (Quality control) of the output reports! They were so bad and so wrong that the entire project had to be shut down!

This is not a completely hypothetical scenario constructed for rhetorical effect, but a pattern that we are, quite literally, sleepwalking into… across India and across much of the Global South, who are becoming users of models developed on datasets which have a meagre representation of their own population. And the reason it keeps happening comes down to something that sounds deceptively simple but is, in practice, enormously consequential: how we evaluate AI before, during, and after deployment in the real world.

The Problem with ‘It Works’

When someone tells you that an AI model ‘works’, the first question you should ask is: for whom? The second should be: where, when, and under what conditions? And the third, which is often the most important and the most neglected: who decided that it works, and by what criteria? Most health AI systems today are validated in much the same way we might grade a final degree examination – tested once, under controlled conditions, on a fixed dataset, against a predefined set of metrics, and then declared “ready”. Researchers call this static evaluation, and this is the most common way of evaluating AI today – at least in places where they are getting evaluated. You measure metrics like accuracy, sensitivity, specificity, and perhaps area under the curve. If you are smart and can afford a research team, you write a paper. You obtain some form of regulatory clearance. And then you deploy, with the implicit assumption that the validation you performed at one point in time, in one clinical context, will continue to hold across other times, other populations, and other care settings.

But healthcare and our patient population do not stand still, and that is an assumption which is deeply flawed but often made by tech companies that do not understand the nuances of healthcare in the real world. Patient populations shift with migration, with seasonal disease patterns, and with epidemiological transitions that unfold over years. Equipment degrades or gets replaced with devices that produce subtly different image characteristics. Clinical protocols evolve as guidelines are updated and as the staff on the ground changes.

The junior MBBS doctor using an AI triage tool on a night shift in August, in a district hospital running on a skeleton crew, is operating in a fundamentally different clinical context than the senior radiologist who validated the same tool in a well-equipped conference room at a hospital in a metro city in January.

A model that was “accurate” at the time of testing can silently degrade over the weeks and months that follow (a phenomenon that the machine learning literature refers to as performance drift), and nobody may notice, because nobody is checking anymore after some time! The certification has happened, and the regulatory box has been ticked.

Approval, in this context, is often just a decision made at one moment in time, but it is treated as if it guarantees long-term safety. This can be misleading. Instead of truly ensuring safety, it may sometimes become more about meeting bureaucratic requirements while appearing to follow scientific standards.

What Changes When You Actually Look

Suppose your chest X-ray AI model was trained overwhelmingly on data from large urban tertiary care hospitals and perhaps a few private hospital chains in metro cities. The images are high-resolution, acquired on well-maintained machines, and the diagnostic labels were annotated by specialist radiologists with years of training. The patient demographics in the training set skew towards a certain age range, a certain distribution of comorbidities, and a certain socioeconomic profile that is characteristic of the population that accesses tertiary care in Indian cities.

Now when you deploy that same model in a primary health centre in a state such as Jharkhand, where the disease prevalence is different and the imaging infrastructure is a generation older, or let us assume a district hospital in rural Assam, where the clinical presentation of common conditions may be altered by high rates of malnutrition and co-infections, in each of these settings, the data distribution and clinical context has shifted.

The patient population has shifted, and in some cases, the very definition of “ground truth” – the standard against which the AI’s predictions would be judged has shifted, because the conditions being diagnosed, their prevalence, and the manner in which they present are all different from what the model encountered during training and validation. This is what researchers mean by the concept of contextual clinical validity: the recognition that a model validated in one setting cannot simply be assumed to perform safely or effectively in another.

But it is a principle that almost no current policy framework for health AI takes seriously enough. Today, we regulate AI systems as though they were pharmaceuticals (test once, certify and release) when in reality they behave more like living systems that interact with their environment in complex and sometimes unpredictable ways, and whose behaviour can change even when the model itself has not been modified.

The Equity Question You Cannot Ignore

Consider a model that reports 94% accuracy overall. That headline number sounds reassuring, even impressive. But when you perform what we call stratified evaluation (breaking down the performance across demographic, geographic, and clinical subgroups), you might discover that the model achieves 97% accuracy for male patients aged 40 to 60 from urban settings, but only 78% accuracy for women over 65 from rural areas. Or that it systematically underperforms for patients with certain comorbidities, or for communities that speak languages underrepresented in the clinical notes on which the model was trained, or for populations from specific ethnic or socioeconomic backgrounds that were simply not well-represented in the training data.

That 94% is an average. And in healthcare, averages can obscure exactly the kind of harm that falls disproportionately on the people who are already most underserved. If an AI system is being used to “triage” patients, deciding who gets escalated to a specialist and who gets sent home with reassurance, then a 20-percentage-point accuracy gap between demographic groups is not a “limitation” to be noted quietly in the appendix of a research paper. It is a form of structural harm, reproduced and scaled through computational means. It means that the communities with the least access to qualified healthcare professionals are also getting the least reliable AI.

Equity, therefore, cannot be treated as an afterthought or an optional add-on in AI evaluation. It must be at the centre of the evaluation framework itself. The question is not merely “does this model work for most?” but “does it work for the people who need it the most, and does it work in the places where it will actually be deployed?” If your evaluation framework does not mandate stratified performance reporting (across gender, age, geography, socioeconomic status, language, and clinical subgroups) then it is not, in any meaningful sense, an evaluation framework.

From One-Time Certification to Continuous Oversight

So, what would a better evaluation actually look like in practice? It begins with a fundamental shift in how we think about the relationship between evaluation and deployment. Instead of treating evaluation as a gate or a checkpoint that a system passes through once before being released into the world, we need to reconceptualise it as a continuous process that runs for the entire lifecycle of a deployed system, and that is embedded within the clinical workflows where the system operates.

This means, first, real-time monitoring: not annual audits conducted retrospectively, but dashboards and feedback mechanisms that track model performance continuously against incoming clinical data. Such systems should flag when accuracy drops below acceptable thresholds, when the demographic or clinical composition of the patient population changes in ways that may affect model performance, or when the system begins to be used for clinical tasks that it was never designed or validated for.

It means, second, periodic recalibration: models should be regularly retested against fresh, locally relevant data that reflects the current clinical reality of the deployment site. A system that was deployed in 2024 and has never been reassessed against local data by 2026 should not be trusted uncritically. It must be actively questioned, and its outputs treated with appropriate caution.

And it means, third, structured feedback from the frontlines. The clinicians who use these tools every day – the MBBS doctors in primary health centres, the nurses in community health centres, the technicians who operate the imaging equipment – are often the first to sense when something is off, when the AI’s suggestions stop aligning with what they see in the patient in front of them. Evaluation systems need to be designed to capture that experiential knowledge systematically, and to treat it as data that is just as important as the model’s log files and performance metrics.

This is what dynamic, continuous evaluation means in practice. Not a one-time examination that grants a permanent certificate, but an ongoing, iterative conversation between the AI system, the clinical environment in which it operates, and the human beings (both clinicians and patients) whose lives and well-being depend on it.

The question today is not whether AI will transform healthcare. It will, and in many ways it already is. The question, the one that matters far more and that remains dangerously unresolved, is whether we will build the evaluation infrastructure that makes that transformation safe, equitable, and accountable, or whether we will keep certifying systems in Delhi and hoping, without evidence, that they work in Sonipat.

The future of health AI in India and the world will surely not be decided solely in research laboratories or corporate boardrooms. It will be decided by whether we, as a generation of researchers, clinicians, policymakers, and engaged citizens, choose to ask the harder questions and to insist that the answers meet the standard that patients deserve.

The author is a radiologist and currently a Simons-Ashoka Early Career Fellow leading a healthcare AI research lab at the Koita Centre for Digital Health, Ashoka University. Views expressed are personal.

Sticky Button