For the past two years, pharmaceutical AI has followed the same script as the rest of machine learning.
Bigger models.
More parameters.
More GPUs.
The assumption was simple: scale would eventually translate into better science.
But the release of Insilico Medicine and Liquid AI’s new model suggests something else may matter more.
Architecture.
Their new model, LFM2-2.6B-MMAI, is a 2.6 billion parameter scientific foundation model designed specifically for drug discovery workflows. It runs entirely on private pharmaceutical infrastructure.
No cloud dependency.
No external APIs.
No exposure of proprietary compound libraries.
And despite being roughly one-tenth the size of comparable models, it performs competitively across multiple discovery tasks.
That combination is unusual.
Because in pharmaceutical AI, privacy and performance rarely coexist.
Most AI models used in drug discovery today fall into one of two camps.
General frontier models
These models are powerful. But they run on external infrastructure. That means sending proprietary data to third-party servers. For pharmaceutical companies, this is often a non-starter.
Compound libraries, assay data, and target hypotheses represent billions in intellectual property.
No one wants that leaving the building.
Specialized point models
These live on internal infrastructure and solve specific tasks. A model for ADMET prediction. Another for docking. Another for retrosynthesis.
But stitching together dozens of narrow tools creates fragile workflows.
Every handoff becomes a failure point.
The industry has quietly been searching for something in between.
A general scientific model that is small enough to run internally.
LFM2-2.6B-MMAI is built on Liquid AI’s LFM architecture, which does not follow the standard transformer blueprint.
Instead, it draws inspiration from dynamical systems and signal processing.
This is not just a stylistic difference.
The goal is to make models more compute-efficient without sacrificing reasoning ability.
According to Ramin Hasani, efficient architecture design may matter more than raw parameter count when building models for scientific applications.
In other words, scaling alone may not be the path forward.
The model was trained on roughly 120 billion tokens of pharmaceutical data.
But the more interesting piece is the training framework.
Insilico developed a system called MMAI Gym, a benchmarking and training environment containing over 1,000 pharmaceutical tasks across two major tracks:
Chemical Superintelligence (CSI)
Medicinal chemistry reasoning, molecular design, optimization.
Biology and Clinical Superintelligence (BSI)
Target biology, pathway reasoning, and clinical planning.
Instead of optimizing for a single benchmark, the model learns across the entire drug discovery pipeline.
Property prediction.
Molecular optimization.
Affinity estimation.
Retrosynthesis planning.
The idea is to train the model to reason across the discovery loop rather than within isolated tasks.
The model’s performance is notable for one reason.
It consistently competes with systems far larger.
On property prediction tasks involving pharmacokinetics and toxicology, LFM2-2.6B-MMAI outperformed TxGemmaon 13 of 22 benchmarks, despite TxGemma having more than ten times the parameters.
On molecular optimization benchmarks, it achieved success rates up to 98.8 percent on MuMO-Instruct.
On affinity prediction tasks, evaluated against 2.5 million experimental measurements across 689 protein targets, it produced stronger correlation scores than several frontier models.
Notably including GPT-5.1, Claude Opus 4.5, and Grok-4.1.
None of these systems were designed specifically for drug discovery.
But they are the models researchers often reach for.
A 2.6B parameter model changes deployment economics.
Models of this size can run on private GPU clusters inside pharmaceutical companies.
That means:
• No external data transfer
• Lower inference cost
• Faster iteration cycles
This matters for tasks like ADMET screening or lead optimization, where thousands of predictions may be generated daily.
Latency and privacy both matter.
This release hints at a broader trend.
For a while, the dominant belief in AI was that scale solves everything.
But scientific domains behave differently.
Training data is scarce.
Benchmarks are complex.
Reasoning often matters more than pattern recognition.
In these environments, architecture and training design may matter more than sheer size.
If a 2.6B parameter model can outperform models ten times larger on pharmaceutical benchmarks, it suggests something important.
The next generation of scientific AI might not come from bigger models.
It may come from smarter ones.
For Alex Zhavoronkov (CEO of Insilico), the goal is straightforward.
Compress discovery timelines.
Insilico has been building an AI-driven platform that spans target discovery, molecular design, and clinical development. Opening MMAI Gym and releasing models trained on it expands that platform beyond the company itself.
It turns their internal infrastructure into something closer to industry infrastructure.
For the broader ecosystem, the signal is subtle but clear.
Pharmaceutical AI may not be headed toward trillion-parameter models running in hyperscale clouds.
It may instead converge on efficient, domain-specific systems that live inside research organizations.
Quietly embedded in the workflows that determine whether a drug candidate survives long enough to reach patients.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.