The number of “neo-labs” being announced every day around the idea of automating research is only getting more interesting. I spoke with @tensorqt (Francesco) - CEO of Paradigma - about autonomous research, research taste, and the infrastructure layer underneath both. Here are my thoughts and insights from this podcast.
Paradigma is building Flywheel, which is essentially the infrastructure for autonomous research.
Before I dive in, there are some bits I want to highlight. Francesco dropped out of his PhD program a couple months ago. The whole career trajectory might look scattered at first, but as conversation went on, the reason became obvious. He kept chasing whatever felt like the hardest unsolved problem in front of him and that path bent from physics into deep learning. (Aha!)
Most of the auto-research discourse I’ve seen is about the model.
“Can GPT-o3 write a paper?”
“Can Gemini replicate a Nature result?”
“Can an agent go from hypothesis to experiment without collapsing?”
Yes, this framing is intuitive. It maps onto how we currently think about AI progress (which is almost entirely about what the model can do at a given moment in time.)
Francesco sees this differently. Not asking how good is the model at research but what has to be true about the world around the model for research to compound?
I think this is a distinct question and the answer points somewhere most people aren’t currently looking.
The core claim is that the bottleneck in AI-driven research is not the model’s intelligence on a given query. It is the lack of structure around what the model produces across many queries (over time) at the intersection of multiple researchers and agents. The model has no memory of its own experimental lineage. There is no object it can look at that captures what it tried + what failed + what was adjacent to something interesting. Each session starts over.
According to Francesco, this is the problem Flywheel is built to solve. And when you hear it described that way, it lowkey seems obvious.
The conversation shifted to the idea of “research taste.”
But what is research taste?
Research taste serves as the key distinction between a PhD student and a Principal Investigator. A PhD student can run the experiments. A PI knows which experiments deserve to be run, emphasizing that is a learned skill.According to him, Research taste is a “probabilistic prior under constraint.” Essentially a judgment call made under constraints such as limited compute and available ideas, to select the research direction most likely to achieve successful and high-impact results.
He defined a person with excellent research taste as one who can identify a correct question + execute the work quickly in few steps + ensure the outcome is profoundly impactful. A person with great taste spots the right question, gets to the answer fast and the answer actually moves the needle. Not just publishable but impactful.
I think this @willccbb tweet is very close to the concept of research taste.
I asked him how he’d evaluate these models through the lens of research taste.
Francesco posits that a fundamental trade-off exists between being an effective assistant and a tasteful ideator. By design - Current models excel as assistants but this very optimization often restricts them to “safe” ideas and perpetual hyperparameter tuning rather than the contrarian, establishment-challenging thinking that characterizes great scientific minds. While models are improving in their ability to follow complex instructions, current training structures are inherently resistant to true creativity, and as a result - improvements in “research taste” are not keeping pace with other performance gains.
This is a training problem. In his words, the base model is something like a schizophrenic - it explores every direction. The current RLHF-trained assistant model is so focused it picks one direction and commits. Neither of those is what a good researcher does. A good researcher has taste - focused, but along the few directions that carry the highest information value given current knowledge.
The question is whether that can be trained. Or rather, whether you can construct a training objective that captures it.
I think one fundamental question when it comes to Auto-Research is, Trust. I pushed on this during our conversation and his answer was genuinely interesting. The problem is obvious as agents hallucinate. They can produce results that look coherent and reproducible within their own context window but fail when you actually try to replicate them. How do you verify?
Actually, you don’t verify it mechanically. You verify it structurally, by encoding into the system the same adversarial logic that underpins the scientific method. Reproducibility by independent parties AND Falsifiability as a prerequisite for a claim mattering.
The specific mechanism he described - where researchers and agents can commit compute to reproducing experiments on the graph and receive fractional ownership of those nodes in exchange is interesting as an incentive design, but I think the deeper point is the philosophy. You can’t hard-verify every result in an autonomous system at scale. The combinatorial load is impossible. What you can do is make the incentive structure favor reproduction rather than novelty alone. If a graph node is highly cited and highly valuable, many agents will want to reproduce it because reproduction earns them something. This is essentially building a decentralized peer review layer into the infrastructure.
I don’t know how the economics of this actually work out cleanly in practice. But the design logic is sound and it’s solving a harder problem than most people in the space are even articulating.
In one possibility, Flywheel could have been an open-source project - rather building a company on it. He was genuine on this. The impetus to form a company (rather than an open-source project) stemmed from the realization that the largest unsolved problems facing humanity - such as nuclear fusion or curing cancer are fundamentally limited by insufficient time and application of intelligence. He stated there is an inherent “opportunity risk” in the next decade if humans, rather than machines, are still the primary drivers of research, particularly if breakthroughs are only limited by GPU availability.
The core thesis is that the bottleneck must shift from human effort to compute resources. In the future, the human contribution to research will be concentrated on generating ideas (not just on coding!). He thinks about a future where science moves away from the current manually intensive, “attentive” process to one resembling an infrastructure users set up to manage compute and mine the results - producing a live map of the knowledge of the world.”
The co-founding story is interesting. Giulio built PaperBench at OpenAI (a benchmark where agents attempt to replicate published research results from scratch). This is very interesting to walk away from a position at one of the most important labs in the world to go build infrastructure in Rome, from scratch.
Giulio apparently believes that the benchmarking problem - how do you even measure whether an autonomous research system is doing something meaningful is as important as the research system itself and again, the infrastructure layer is where the real leverage is (not the model layer!).
And this makes sense. Most of the smartest people in this field are trying to make better models, not better scaffolds for models.
Francesco’s quickfire answer to “what’s the most broken thing about how research is done today” was: inappropriate use of LLMs. I thought about this for a while.
I think what he means is that most researchers currently use models the way they use Google Search - as a retrieval system with Natural language interface. Ask a question, get a text back and move on. If you think about this interaction - the experimental lineage, the hypothesis tree, the failed directions and the adjacent results - none of that is preserved. The model is used as a smart one-shot tool rather than a persistent research collaborator embedded in a structured process.
This a harder cultural change than it might seem, because the “ask the LLM” pattern is deeply established in every research workflow right now. The question is whether the productivity gap becomes obvious and fast enough, that people make the switch voluntarily or whether it requires a generation of researchers who never knew a different workflow.
Francesco seems to believe it’ll be the former. The gap will become undeniable in a few specific niches first, then spread sharply as the pattern becomes visible.
An interesting dynamic.
I liked this part because it quietly breaks an old assumption. He is building from Rome and doesn’t see it as a compromise.
The cultural proximity argument for the Bay Area has largely migrated to cyberspace - specifically X. He met his co-founder on Twitter, found collaborators through Twitter-linked events, and even brought in investors through X DMs. A surprising amount of Paradigma’s early network was built online.
He also made a second point. A lot of the best young ML talent is not sitting in San Francisco. It is spread out and some of it is in Italy. Nobody in SF is recruiting from that pool yet. He called it a “very unfair advantage” and I get what he meant. Since he has privileged access to these young builders before anyone else even knows they’re there and and he’s probably right. The people who are young, curious, still plastic enough to explore weird directions - I think this is where the alpha is in this era. He’s just closer to them than anyone in SF.
As I have been saying this a lot, and a gentle note that this timeline is generational opportunity for anyone young building out in this space.
Watch the full podcast here:
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.