Over the past year, in my capacity reviewing research and data-sharing proposals for medical colleges, I have watched a particular kind of document arrive with increasing frequency. A health-tech company — well-funded, well-spoken, slides in the right shade of blue — proposes a "collaborative research partnership." The institution will "digitise and structure" its patient records. In exchange, the company will use the de-identified data to improve a diagnostic model. There is usually a data-sharing agreement. There is usually a paragraph about future co-authorship. There is rarely a paragraph about who owns the finished model, who profits from it, or whether the originating institution will ever see it again except as a customer…!
I have watched capable, well-meaning colleagues prepare to sign such documents. Not out of carelessness. Out of the ordinary, unglamorous scarcity that defines under-resourced Indian medical colleges: little in-house AI or data-governance expertise, ethics committees stretched thin across a hundred other responsibilities, and a genuine, understandable hunger for anything that looks like partnership with people who have resources the institution does not. This is not naivety in the way the word usually implies. It is a rational decision made inside a system that has never given these institutions the leverage to negotiate better terms.
The more accurate word for what is happening across underfunded Indian institutions right now may not be exploitation of ignorance. It may be extraction, dressed as collaboration, aimed precisely at the places least equipped to price what they are giving away.
I have argued elsewhere that AI is unlikely to make Indian healthcare cheaper overall. I want to go further here, because "overall" hides the more important story, which is divergence. India does not have one healthcare system into which a new technology arrives. It has several, stacked on top of each other and barely on speaking terms:
five-star corporate chains with the capital to license the best diagnostic tools and the fee-for-service incentive to deploy them liberally;
mid-tier private nursing homes chasing whatever technology promises differentiation; and
government and rural teaching hospitals running on numbers that would be a rounding error in the first category's annual report.
The corporate tier will likely use AI to add cost — more flagged findings, more confirmatory scans, more billable next steps, absorbed easily by a payment structure built to convert efficiency gains into revenue. That is a real problem, but one experienced mostly by patients with insurance or savings to absorb it.
The trajectory that concerns me more runs backwards through the second tier: not AI arriving to add cost to care already delivered, but AI arriving to extract value from the one asset resource-poor institutions hold in surplus — patient volume, diagnostic diversity, and the density of undocumented clinical experience that a data-hungry model needs, and that a smaller, more curated corporate caseload often cannot supply as cheaply.
Indian medicine has lived this pattern before, in a different currency. Government and rural hospitals have long been where clinical trials recruit fastest, where postgraduate theses find their sample sizes, where research eventually cited worldwide gets its raw material — often with informed consent that is technically valid and substantively thin, given by patients navigating illness, poverty, and a language barrier from the consent form itself.
The incentive structures now forming around AI training data appear to be reproducing that shape, faster, and with even less institutional memory of what fair benefit-sharing should look like. A model trained substantially on the pathology load of underfunded hospitals — tuberculosis, rheumatic heart disease, malnutrition-complicated illness, the diagnostic long-tail a well-resourced urban hospital sees less of — becomes commercially valuable precisely because of that data…!
The finished, licensed product is then typically priced for institutions that can pay: the corporate chains, again. The originating institution is, at best, offered a discounted subscription to a refined version of its own patients' data. At worst, it is offered nothing beyond the original proposal's vague promise of "future collaboration."
This is not a hypothetical harm sitting downstream. Ethics committees at under-resourced colleges across the country are, right now, evaluating proposals of exactly this kind — framed as capacity-building, assessed by reviewers without the specific expertise to price data as an asset, approved by people acting in good faith under real institutional pressure. This is not a failure of individual judgment. It is what happens when an entire tier of institutions is asked to negotiate data value without ever having been given the tools, precedent, or leverage to do so…
I want to be fair to the genuine counter-example, because I raised it before and it deserves to stand: an AI triage tool that lets a technician in a primary health centre flag a probable TB pattern on a chest film, before it ever reaches a radiologist three hundred kilometres away, is not a cost driver — it may be the only realistic way that film gets read at all. I still believe that. But I no longer think it is separable from the extraction question.
Ask where that triage model’s training data came from. Ask whether the institutions that supplied the chest films that taught it to recognise TB were compensated, credited, or consulted about how the finished tool would be priced and to whom it would be sold. Often, the honest answer is that they were not asked, because nobody thought to ask, because the data-sharing agreement that made the model possible was signed the way the one on my desk nearly was — quickly, hopefully, by people doing their genuine best inside a system that has never taught them to say no on commercially informed terms.
If this trajectory holds, the disparity AI produces in Indian healthcare will not simply be the familiar one — corporate hospitals get better tools, government hospitals get worse ones.
It will be stranger and more self-defeating than that: government and rural institutions will have supplied the raw material that made the tools possible, received little of the financial return, and then be expected — by policy, by accreditation bodies, by the general cultural pressure to look “AI-enabled” — to adopt the finished products anyway, at prices set by a market their own data helped create.
The return on investment for the institutions doing the actual data-generating work will be poor almost by design, because the value was priced and captured somewhere else in the chain, by people who never treated a single one of the patients whose records built the model…!
The proposals keep arriving — reasonable-sounding, well-intentioned on both sides, evaluated by good people without the specific expertise or leverage the moment requires.
Institution by institution, we are generating the data, signing away the data, and quietly reinforcing a system that will use our own openness against the very patients and institutions that made it possible.
Nobody in any of these rooms intends harm. That is precisely what allows the pattern to survive. We are building, with good faith and very little malice, the exact asymmetry we will spend the next decade trying to explain to ourselves.
Note: examples in this piece are composite and illustrative, drawn from patterns visible across multiple institutions rather than any single proposal or review.
Thanks for reading! This post is public so feel free to share it.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.