RSS Amplifier

Partial Agonism · Feb 12, 2026

Lab Automation is Bodybuilding Hard

0
Sign in to vote or save

William H. Bragg · Partial Agonism

I had this article in drafts for a while and last week I finally got around to finishing it up, only to be brutally substack-mogged by Owl Posting who dropped his 8.4k word opus on lab automation a few days ago:

That piece is comprehensive and great, you should go read it. If you read one substack about lab automation this week read that one. This may be of interest if you’re curious about the nitty gritty of why things “often aren’t worth automating”.

From “Hard is Not Defensible” by Alex Compton (2018):

There are two kinds of hardness: Math hard – the hardness of working something out; and Bodybuilding hard – the intrinsic hardness of doing certain things.

To paraphrase, elaborate, and somewhat butcher the original point to fit my own thesis: Math hard is problems where the next step in solving the problem isn’t clear. It’s problems where all the resources in the world might not actually solve it if you can’t come up with the right idea. The example in the piece is Pythagoras’ theorem: could you derive it if you’d never heard of it and weren’t even sure it existed?

This is contrasted with Bodybuilding hard: problems that take resources, time, money, personnel etc, but where the next steps are pretty clear. How do you add an inch to your biceps? Well, you do curls until failure 3-4x per week, you eat enough protein, you obtain and take anabolic steroids. It’s not easy, but it’s difficult in a very distinct way to deriving Pythagoras’ theorem cold.

That specific framing — that there are 2 kinds of hard — clicked immediately for me when I thought about it in the context of my field: lab automation.

First to define it, lab automation is taking the things done by humans in a life science wet lab and getting a robot to do them. The vast, vast majority of this boils down to 3 fundamental processes: manipulating liquids (moving, heating, cooling, mixing), manipulating containers (putting plates in stacks, moving plates between robots, putting plates into readers or PCR machines, shaking, putting lids on and off), and getting a readout of some sort (fluorescence, absorbance, OD, image etc). Nearly everything can be reduced to some variant of these processes, with fractally expanding details and nuance as you zoom in.

Science is math hard. You're almost by definition working at the edge of human knowledge. You don't actually know what the result should be. If something doesn't give the result you hypothesised, you don't know if your model is wrong or if one of a million things you didn't consider is interfering. Cells are complicated, animals and humans even more so. Is your assay on plaques actually relevant for Alzheimer's disease? Is a small molecule cure even possible? Sure, there may be literature on these questions, but does it apply to what you're doing?1

Lab automation on the other hand, is bodybuilding hard. The process to automate a workflow done by a bench scientist is actually very straightforward. You take each step of the process, and figure out how to get a robot to do that step. Then you stick it all together. That's basically it.

To make this concrete, let’s walk through something simple. Here’s a toy example: cloning genes out of a bacterium into an expression plasmid, and transforming that into E. coli2. For a bench scientist, step by step, this would look like:

  1. Streaking out the bacteria

  2. Picking a colony into water

  3. Amplifying the gene by PCR with your colony as a template (after you’ve designed and ordered the primers)

  4. Purifying the PCR product

  5. Ligating the purified gene into the expression vector

  6. Heat shock transformation into E. coli

  7. Recovery in rich media

  8. Plating onto selective LB agar

  9. Picking a few colonies

  10. PCR of those to confirm the gene is there

  11. Culturing for protein expression or whatever experimental work you want to do

Workflows like this absolutely can and have been automated to very high degrees, but something this simple still requires an astonishing amount of equipment and thinking in detail about steps and sub-steps that a technician could do in their sleep.

Let’s start at step one: streaking out bacteria. You have your bacteria as a frozen glycerol stock3, and you normally would streak it out by touching an inoculation loop onto the surface of the glycerol, and then rubbing it onto an agar plate to get isolated bacterial colonies. Right away we’re at something that isn’t trivial for a 4-axis liquid handling robot with no sensors to do. The process of streaking when a human does it involves a lot of tactile feedback to do nicely (scratching the surface of the frozen glycerol, pressing gently onto the agar).

So we have a few options here depending on your setup.

We can skip streaking and just grab some of the glycerol to go directly into our PCR. Cool. Are we sure it’s not contaminated with something? We’re skipping a process control step, which may or may not be formalised but “this strain looks/smells/grows funny” is absolutely something a lab technician will note. OK so maybe we are confident in the sterility and purity of our strains. Are they in a format a robot can handle? If they’re in Eppendorf tubes, we could load them into racks. But wait, did we check the lids don’t get in the way?

Eppendorf Tubes 1.5ml | NHBS Wildlife Survey & Monitoring
Perfectly optimized to be flicked open with a human thumb, a nightmare for a robot to deal with.


Shit, OK we get a rack with special slots to hold the lids open. Did we make sure we keep them cold enough that the glycerol doesn’t lose viability through freeze-thaw cycles? Our new rack doesn’t fit in a cooling block, dammit. OK we can 3D print something, whatever. Is the robot aspirating or just touching the top of the glycerol? If it’s aspirating, we’re comfortable letting it melt each time? Or if we’re just touching the top, we know all our glycerol stocks are the same volume so we don’t miss any? You know glycerol doesn’t freeze as a nice flat surface in a small tube, right?

OK forget all that nonsense. We're just going to buy a specific robot for streaking and picking.

Are we sure we’re going to have the throughput to justify dropping $100-200k USD plus service costs? We know we have a technician or scientist who will stop doing whatever their job is at the moment and invest the time to learn how to operate it? Does it run OK if we only need to do 10 strains and not 96+? Does that smaller batch take longer with the setup time on a robot than a technician would take to do the whole thing (in which case the robot won’t get used)?

To be clear, there are normally pretty clear answers to every one of these questions. There are solutions to all these little problems that crop up along the way. None of them require particularly novel or inventive thinking. Every solution will take time, money and personnel to implement, test and troubleshoot though.

Most of the steps along a process will have some level of issue like the colony streaking example to think about. There will be fewer issues for bread and butter liquid handling, but it’s rare that there isn’t some fiddliness to hammer out.

Once you automate each step and start sticking everything together, you'll hit similar issues at pretty much all the integration steps. In addition, the likelihood of a workflow running end to end without fail when you have a sequence of steps each with X% failure rate drops off fast as you add steps, and the numbers get pretty brutal really quickly once the instrument count starts to stack up.

Now the real questions: how long is sorting all this out going to take? How much will it cost in capex and service contracts? Alright, so how many samples do you run per week? How long are you going to do that for before you come with a better assay or change targets or the department gets laid off? Which gets to the heart of all this: will lab automation really pay for itself?

I am short-term bearish and long-term incredibly bullish on AI for lab automation. No I will not define my timelines, which makes them unfalsifiable.

Many of the things we need to solve for lab automation are the things we need to solve for robotics generally: autonomous robots with sufficient sensory feedback and visual acuity. My error bars on the timelines to this are massive and a lot of it comes down to “how much do LLMs help with general robotics”.

The arms and pipettes used in lab automation currently can follow instructions with pretty good precision, but they are nearly all exclusively deterministic instructions in the form of “move the axes to coordinates X, Y, Z, and then activate the aspirator at such and such a speed”.

The upshot of this is that truly cutting-edge AI-first biotech companies, with huge funding rounds whose entire model relies on loops of, for example, protein synthesis and screening, still have incredibly brittle physical systems. Every time they take out the plate reader to look at something or do maintenance, it needs to be clicked back into the system without more than 1 mm of deviation, lest the arm miss the plate loading step and the run fails.

Recently a Ginkgo x OpenAI collaboration got a fair bit of coverage in the lab automation space. Ginkgo has been touting their “autonomous” lab for a while now, and to their credit they have automated and integrated a bunch of units along with what I have heard is a decent controlling software, in a very modular system.

They gave GPT-5 access to running experiments on one of their autonomous labs and got some perfectly nice results.

The key thing though, is they had already sunk years and god knows how many FTEs into getting the autonomous lab to work. The “hard” part (I have to assume) was not the GPT-5 integration. The hard part was almost certainly troubleshooting problem number 589: “arm doesn’t successfully pull out the freezer rack due to frost” or problem 4568: “trigger from plate reader completion comes before the plate is actually ejected meaning our arm tries to pick up the plate before it’s ready.”

The challenge with lab automation currently as I see it isn’t that we have automated systems that we need good hypotheses for. It’s that the automated systems still need lots of bodybuilding-hard development to get running properly.

1

Of course, there are chunks of lab science — or processes like "writing and publishing a paper" or "getting pre-clinical safety data on a drug before human trials" — that fall more into the bodybuilding hard category where obvious, expensive and/or time consuming tasks just need to get done.

2

Now you’d almost certainly just buy synthetic genes from Twist or something, but let’s imagine for some reason we want to clone them instead.

3

A tube of bacterial culture mixed with glycerol to keep it in suspended animation so it's available to wake up whenever you need it.

No posts

Read the original on partialagonism.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.