RSS Amplifier

Intersections by The Utopia Studio · Apr 20, 2026

The Best Surgeon in the World Can Only Be in One Room

0
Sign in to vote or save

Ollie Graham-Yooll, Amyn Haji, Karan Pinto · Intersections by The Utopia Studio

On a Thursday morning in an endoscopy suite at King’s College Hospital, Amyn Dr. Haji is dissecting a polyp out of a patient’s colon that fifteen years ago would have been treated by cutting the colon out instead. The polyp is a laterally spreading tumour, roughly six centimetres across, carpeting a section of the sigmoid wall the way moss covers a stone. He’s using a technique called endoscopic submucosal dissection — a needle-thin knife pushed through a flexible scope, lifting the lesion with saline, then slicing it free from the healthy tissue beneath, one millimetre at a time. An ESD on a lesion this size takes most endoscopists two to three hours. Dr. Haji does it in under 60 minutes. He has done this operation more than a thousand times.

Somewhere behind him, a camera is recording. Somewhere in the building, a trainee is watching a live feed. Somewhere in the hospital’s archives, a case note will be typed, filed, and, within six months, effectively lost — indexed by date, not by what happened.

Here is what is not being recorded. Not the video. The video is fine. What isn’t being recorded is the reasoning underneath the video. Why Dr. Haji lifted the lesion at this corner and not that one. Why he paused for seven seconds before the next cut. Why, when the submucosal plane dropped away in the middle of the procedure, he switched his injection solution — and why that switch will be invisible on the video unless you already know what to look for. Why he doesn’t cut here. Why he cut there. Why, when the anaesthetist says something about CO₂ absorption, he steps his pace down a fraction without announcing it.

That reasoning is the operation. The cuts are the easy part. The cuts are what a good general surgeon can learn in a year. The reasoning is what takes twenty years to build, and it is almost entirely unwritten.

Dr. Haji trained under Professor Shin-ei Kudo at Showa University in Yokohama, where pit-pattern classification — the visual taxonomy that lets an endoscopist read a lesion’s malignancy from its surface pattern before cutting — was invented in the early 1990s. He trained under Professor Haruhiro Inoue, who in 2008 pioneered peroral endoscopic myotomy, POEM, for the treatment of achalasia: a procedure that, until a single surgeon in a single hospital in Yokohama decided otherwise, required a full chest incision. Kudo’s pit-pattern classification took two decades to cross the Pacific. Inoue’s POEM took a decade to become standard of care. The knowledge exists. The transmission mechanism is the slowest part of modern medicine.

Dr. Haji wrote more than a hundred papers. He runs King’s Live Endoscopy, an annual event where surgeons from around the world watch him operate via live broadcast. He is directly responsible for training a generation of British and European colorectal endoscopists in techniques that Western centres were telling their fellows, as recently as 2015, could not be safely adopted outside Japan.

He’s fifty-one. Even if he operates until he’s seventy, the arithmetic is unforgiving. A lifetime of surgical knowledge dies with the surgeon unless something structural happens between now and then to preserve it.

He’s watched it happen before. Every master does.

There is a useful exercise if you work in medical education or surgical robotics and you want to understand the shape of the problem. Sit in on a teaching ward round. Count the number of times a senior surgeon says something like, “You can see here that…” or “The thing to notice is…” or “In my experience when you get this appearance, you want to…” Now count how many of those statements get written down anywhere a future trainee could find them. The answer is approximately zero.

This is not a failure of diligence. This is the architecture of the apprenticeship.

Surgical training in the West still runs on a model designed in 1889. William Halsted, at Johns Hopkins, formalised a system in which young surgeons learn by watching, then assisting, then doing under supervision, then finally operating unsupervised — passing, as Halsted put it, from observer to assistant to surgeon-in-charge. The method works. It produces excellent surgeons. It also scales linearly with the number of senior surgeons willing to teach and the number of junior surgeons in their operating room, which is to say it does not scale at all.

The cost of that architecture is legible in the learning curves. A surgeon acquiring proficiency in laparoscopic colorectal surgery requires somewhere between 50 and 152 cases before complication rates drop to expert baseline, depending on the definition and the case mix [Fact]. Anastomotic leak rates in the first 200 cases for a new laparoscopic colorectal surgeon run 6–9 percent, falling below 2 percent only after sustained case volume. Endoscopic submucosal dissection — Dr. Haji’s specialty — requires approximately 80 procedures before an endoscopist can reliably achieve the dissection speed, en bloc resection rate, and perforation avoidance that define expert performance [Fact]. Forty procedures just to stop punching holes through the colon wall.

A surgeon who trains one fellow a year transmits her expertise to perhaps twenty people in a career. At the population level that is a rounding error. The NHS diagnoses a colorectal cancer patient every fifteen minutes. Fewer than two hundred endoscopists in the UK are trained in complex ESD. The supply of expertise is not keeping up with demand, and the pipeline cannot be accelerated without compromising the patient on whom the trainee is learning.

Meanwhile, the data that would compress the learning curve sits in three places. On hospital video servers, unsearchable. In case notes filed by date and patient identifier, not by clinical feature. And in the head of the surgeon who did the case, a few terabytes of pattern recognition with no export format.

Miss the capture window and it’s gone. The institution writes a card. Everyone claps. A successor picks up the baton and starts the twenty-year climb from somewhere near the bottom.

There is a specific type of surgical knowledge the field has spent a decade trying to automate, and it is worth understanding the ceiling the automation has hit, because it tells you where the actual opportunity is.

Look at polyp detection.

A colonoscopy performed by an average endoscopist misses somewhere between 20 and 30 percent of adenomas — precancerous polyps. That miss rate is the source of a meaningful fraction of post-colonoscopy colorectal cancers, the ones that appear in the three to five years between a “clear” screening and a diagnosis. For a decade, the field has known that computer-aided detection systems — CADe — can improve adenoma detection rate by around 20 percent in head-to-head trials, with a 55 percent decrease in adenoma miss rate when pooled across dozens of randomised studies [Fact]. The pooled adenoma detection rate in the CADe group of a major 44-trial meta-analysis was 44.7 percent, versus 36.7 percent in the standard colonoscopy group. These are not marginal gains. These are meaningful numbers of cancers prevented.

So why is CADe not yet standard of care in every hospital in the world?

Part of the answer is procurement cycles and variable reimbursement. The underlying technical ceiling is that detection is not the hard part of endoscopy. Detection is the foothill. The mountain is what comes next — characterisation of the lesion, decision to resect, choice of technique, execution, management of complication. That entire sequence is where the cost, the risk, and the patient outcome live. CADe answers “is there something there.” It does not answer “what is it, should I take it out, and how.”

Systems that pattern-match against Kudo’s pit-pattern classification can get partway to characterisation. Systems that track instrument kinematics — the path a needle takes through tissue — can score dexterity. Intuitive Surgical, the company that manufactures the da Vinci robotic system, has accumulated an extraordinary video corpus. Their SurgVU dataset contains more than 840 hours of robotic surgery footage, sampled at 60 frames per second, labelled for tool use and surgical step — roughly 18 million annotated images [Fact]. Research groups at ETH Zurich, MIT, and NVIDIA are training video-language foundation models on progressively larger surgical corpora, using frameworks like VidLPRO and GenSurg+ to pair video with transcript-derived captions [Fact].

This is meaningful work. It is not the ceiling.

The ceiling is that almost all of this data is observational. It shows what happened. It does not show why. The master surgeon who stepped back for seven seconds before the next cut did so for a reason, and until the reason is captured alongside the video, the foundation model trained on it will learn the choreography without learning the judgment.

You end up with a system that can recognise an ESD but cannot perform one.

The kinematics are visible. The decision tree that produced the kinematics is not.

What is missing is the expert annotation layer. The running commentary. Here’s why I’m doing this, here’s what I’m looking for, here’s what would make me stop, here’s what this reminds me of. That annotation layer exists — in the heads of Dr. Haji, Kudo, Inoue, and a few hundred other world-class endoscopists globally. It has simply never been systematically captured.

Amyn Dr. Haji has, by his own count, somewhere north of four thousand hours of operative video sitting on hospital servers, going back fifteen years. He has case notes for every single one of those procedures. He remembers, with startling specificity, the difficult ones. Which tells you something: the mind of a master surgeon is not a filing cabinet. It is a retrieval system. Given the right cue — a lesion appearance, a patient history, a scope angle — he can surface the relevant priors from a fifteen-year career in roughly the time it takes to make a decision mid-procedure. Three to five seconds. That retrieval, more than the hands, is what makes him one of the best in the world.

Now imagine you could capture that retrieval system. Not the video alone. Not the notes alone. The reasoning that connects the two.

This is the insight behind what Mentix is building.

The initial concept, which is where the company is live today, is narrower than the eventual vision but already interesting on its own terms. You take a single surgeon — Dr. Haji, as the founder-operator pilot — and you build a system capable of ingesting the raw material of a career: operative videos, case notes, the published papers, the teaching lectures, the recorded broadcasts from King’s Live Endoscopy. You structure it. You annotate it, in partnership with the surgeon himself, layering the why onto the what. You build an interface that lets a trainee, mid-procedure, query the corpus in natural language.

“I’m looking at a 30mm LST-granular in the rectum with a type IV pit pattern near the anorectal junction. What would you do?”

And the system responds not with a textbook answer but with Dr. Haji’s answer — grounded in a specific case he did in 2019, or a variant he discussed in a paper in 2022, or a teaching moment he recorded in 2024.

That is a personalised AI. It is a specific surgeon, abstracted from the biological constraint that he can only be in one operating room at a time.

If you accept that such a thing is technically feasible — and a decade of progress in multimodal foundation models, expert retrieval systems, and surgical video analysis suggests it is — then the second-order implications become the interesting part.

A personalised AI of one expert surgeon is a tool. A network of personalised AIs, covering every major surgical specialty and every major procedural variant, is an infrastructure. It is the first time in the history of the field that expertise can be queried without the expert being physically present. It is the first time that a 28-year-old registrar in Doha can have, in her ear, the reasoning of a specific Japanese endoscopist she will almost certainly never meet. It is the first time that the Halsted apprenticeship model gets genuinely unbundled — not replaced, but extended, because the senior surgeon’s judgment can be carried into rooms the senior surgeon will never enter.

And that infrastructure, once it exists, is the data layer the next generation of surgical robotics has been missing.

Here is the robotics parallel, which matters because it shows where the real scale of this opportunity lives.

In the last three years, the frontier of robotics has moved from specialised, hand-engineered systems to generalist foundation models — large neural networks trained on enormous, heterogeneous corpora of physical interaction data. Google DeepMind’s RT-2. NVIDIA’s ORBIT-Surgical. Physical Intelligence’s π-series. These models learn motor skills the way a language model learns grammar — by consuming huge volumes of example data and extracting the underlying patterns. The key lesson from the last decade of AI is that data quality and diversity beat architecture [Fact]. The teams that win are the teams with the best corpora.

Surgical robotics has a corpus problem.

The existing datasets — the SurgVU data, the academic collections, the fragments of kinematic logs released from da Vinci installations — are large but narrow. They skew heavily toward training tasks on porcine models, toward early-career surgeons, toward procedures chosen for their pedagogical cleanliness rather than their clinical complexity. The corpus required to train a robot to perform a complex ESD on a live human colon does not currently exist in a structured, queryable form. It exists as raw video on hospital servers, untagged, unannotated, and bound by the data governance of fifty different health systems.

And the missing layer — the one that would transform that video from an observational record into training data — is exactly the layer Mentix is building. The annotations. The reasoning. The expert commentary. The why.

Think about it structurally. Intuitive’s SurgVU dataset is, at 840 hours, smaller than the body of operative video sitting on Dr. Haji’s King’s College servers for a single subspecialty. Multiply that by every subspecialty, every senior endoscopist and surgeon, every teaching hospital globally. The raw footage exists. What is missing is the connection between the raw footage and the reasoning of the operator. Once that connection is made — once you have a systematic way to extract expert judgment and pair it with the corresponding video — you have something the robotics field does not yet have: a structured dataset of surgical decision-making at expert grade.

This is the ImageNet moment for surgery.

The analogy is deliberate. In 2009, Fei-Fei Li and collaborators released ImageNet — fourteen million hand-labelled images, organised by a taxonomy of real-world objects. The dataset was not conceptually novel. What was novel was the scale, the labelling, and the commitment to making it open to researchers. Within three years, ImageNet had catalysed the deep learning revolution. Not because the neural network architectures were new — most of them dated to the 1980s — but because, for the first time, there was enough labelled data for the architectures to learn on. The corpus was the unlock.

Surgical AI is sitting at the equivalent point. The architectures exist. The compute exists. The raw video exists. What is missing is the expert-annotated, structured, reasoning-laden corpus. Whoever builds it becomes the substrate on which the next generation of surgical AI — training tools, decision support, and eventually autonomous surgical robotics — will be trained.

This is not a small prize. Intuitive Surgical alone is a $180 billion company, built on a single robotic platform that requires a human surgeon in the loop. The companies that build the successor platforms will be built on data they do not currently have. Mentix is positioning to be the supplier.

Now engage the hardest version of the counter-argument, because a piece like this is only credible if it does.

Surgical knowledge cannot be extracted because it is fundamentally tacit. It lives in the hands. It lives in the pattern recognition that a surgeon cannot quite put into words. Attempts to codify it have been tried for decades and produced textbooks that no trainee reads.

This is the strongest objection. It is also incomplete, and it’s worth being precise about why.

The objection assumes the extraction problem is linguistic — that the only way to capture tacit knowledge is to get the expert to write it down or speak it aloud, and that this process loses the essential thing. It also assumes the extraction has to be complete. Neither assumption holds.

Two things have changed in the last five years that should update the prior.

The first is multimodal capture. A senior surgeon narrating mid-procedure, with video of the operation running alongside, kinematic data from the scope, the audio of the operating room, the patient’s vital signs, and the pathology result a week later — that is not a linguistic record. That is a dense, time-aligned dataset in which the reasoning is grounded in the physical reality that produced it. The commentary does not need to be complete. It needs to be paired with what is on the screen at the moment it is spoken. Modern foundation models are very good at learning from aligned multimodal data. They are much worse at learning from text alone.

The second is retrieval-augmented generation. The AI systems being deployed in medicine today do not need to memorise the corpus. They need to retrieve the right fragment at the right time. A trainee mid-procedure does not need a generic summary of colorectal ESD. She needs the three cases in Dr. Haji’s career that most resembled the lesion on her screen, the reasoning he applied to each, and the outcomes. That is a retrieval problem, and the tools to solve it at scale exist now.

Tacit knowledge resists extraction into text. It yields to extraction into structured, multimodal, retrievable datasets. The technology to build those datasets did not exist five years ago. It does now. That’s the update.

The historical parallel worth holding in mind is Go. For decades, the orthodox view among Go professionals was that top-level play required an intuition that could not be taught outside of a decade of direct apprenticeship under a master. The 2016 AlphaGo result did not disprove the existence of that intuition. It showed that the intuition could be induced in a system that had consumed enough high-quality game data. The extraction wasn’t linguistic. It was behavioural. The surgical equivalent is where we are heading.

Mentix’s roadmap, stated plainly, has three phases, each of which is independently valuable and none of which depends on the later phases being right.

Phase 0 - The Mentor Platform, a foundational platform allowing surgeons to connect with mentees and effectively train and transfer knowledge, onboarding them to Mentix and allowing them to explore deeper, long-term knowledge sharing. This phase is live.

Phase One: The Personalised Surgeon AI.

Pilot with a single expert — Dr. Haji — operating as both founder and subject. Build the capture, annotation, and retrieval infrastructure on one career, proving that a queryable personalised surgical AI is technically feasible and clinically useful. This phase is in R&D.

Phase Two: The Training Platform
Deploy the personalised AI into NHS trusts, Royal College training programmes, and MedTech partner organisations, starting with the institutions where Dr. Haji already has teaching relationships. The unit of sale is not the model; it is the training license. A training programme director at a teaching hospital signs a contract that gives her fellows access to the Dr. Haji corpus, with the retrieval interface embedded in their training workflow. She measures success by time-to-proficiency — the compression of the learning curve that the field has been waiting for since Halsted.

Phase Three: The Multi-Surgeon Corpus
The multi-surgeon corpus and the data layer for surgical AI broadly. By the time ten or twenty senior surgeons across specialties are on the platform, Mentix holds a structured, expert-annotated surgical corpus with no true competitor. That corpus becomes the substrate for three distinct businesses: continued training platform expansion, decision-support tools for practicing surgeons, and — the largest prize — a licensed data layer for the robotics platforms that will need exactly this data to move from teleoperation toward semi-autonomy. That last market is the one Intuitive, Medtronic, CMR, Asensus, and half a dozen well-funded startups are all racing toward, and the one that will be structurally unreachable without an expert-grade surgical dataset.

The moat across all three phases is the same. Every expert onboarded makes the next one cheaper to onboard. The annotation methodology, built with Dr. Haji, gets refined with the second expert and polished with the fifth. The retrieval infrastructure gets more useful with every additional case. The trainees using the platform generate engagement data that feeds back into which annotations are most valuable. And — crucially — exclusive relationships with a small number of world-class surgeons per specialty, early, create a corpus that a competitor with more capital but a later start cannot easily reconstruct.

This is Bloomberg’s moat. Not the terminal. The settlement record. The data that accumulates as a byproduct of the service and that, over a decade, becomes the thing the industry cannot work without.

There is a version of this piece that ends on the commercial arithmetic. The surgical simulation market is around $665 million today and projected to approach $1 billion by 2033 [[Assumption]](market reports vary by source; order of magnitude is consistent). Surgical robotics is meaningfully larger — roughly $7 billion, growing at mid-teens. Adjacent markets — decision support, post-operative quality assurance, medical AI broadly — extend the surface into the tens of billions.

That version is real. It is the smaller version.

The larger version is that, for the first time in modern medicine, there is a credible path to decoupling surgical expertise from the physical presence of the expert. A 32-year-old surgeon in Dar es Salaam, operating on a complex colorectal lesion, should have access to the same pattern recognition, the same priors, the same decision support as a surgeon down the corridor from Dr. Haji at King’s. She does not yet. The gap between her outcomes and his trainees’ is not a function of her capability. It is a function of her geography and the slowness of the transmission mechanism that determines what she has access to.

That gap is what Mentix is built to close. Not immediately. Not for every procedure. But the architecture of the closure — a queryable corpus of expert surgical judgment, built in partnership with the experts themselves, delivered as infrastructure — is the first serious attempt at a structural fix.

Dr. Haji has watched the knowledge leave the room before. Every master has. Every master has also known, somewhere behind the frustration, that it did not have to. The video always existed. The notes always existed. The reasoning always existed. The missing piece was a way to connect them.

That piece exists now.

Interested in working with Dr. Haji or collaborating with Mentix?

Dr. Amyn Haji is the clinical lead and founder of Mentix. If you are a surgical trainer, NHS trust, MedTech company, or research institution interested in the Mentix platform, reach out directly.

Contact Mentix

Mentix is a founding fellow company at Utopia Studio, operating within our Co-Build at The Utopia Studio.

Further reading: Dr. Haji et al. — ESD outcomes study (Springer Surgical Endoscopy, 2026)

No posts

Read the original on intersectionsutopiastudio.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.