By Sean Hill
The figure above does not exist — and the harder problem is that for most real figures, the evidence underneath is almost as hard to reach.
The figure above is not mine. It was posted on X on April 24 by Ruslan Rust, produced with a single prompt to the new ChatGPT image model. There were no mice. No 5xFAD cohort, no behavioral assay, no western blots, no Thioflavin-S sections, no Morris water maze. The “data” was generated, in the strictly technical sense of that word, by a chatbot.
Rust posted it as a warning, with the line: “Can water intake prevent Alzheimer’s disease? No. This is fully AI-generated… but the data below could easily pass as real.” He is right. It could.
It looks familiar because every panel is borrowing the rhythm of a real paper somewhere. Error bars the right size. Asterisks placed where you would expect them. A caption that reads like a caption I might have written myself. I have spent a long time looking at figures like this, and the visual signature here is now indistinguishable from the real thing. If an editor pushed this to me as a reviewer at the end of a long week, I am not confident I would catch it.
A few months ago, I would have caught it. The labels would have been malformed. Stained sections would have melted, the way generative models used to melt anatomical detail. Gel bands would have been biologically or physically implausible. The image would have given the game away.
That window has closed, and it closed faster than most of us expected.
The instinct, which I share, is to push back with detection — watermarks, image forensics, C2PA-style provenance baked into pixel data, classifiers trained to spot generated content. We should keep building those tools, and the major image-generation vendors should keep adding signatures. None of it can be the foundation. Generators improve faster than detectors. They always have. If the integrity of the scientific record depends on detection staying ahead of generation forever, the record is already in trouble.
Rust’s deeper point is the one to take seriously. AI-generated figures did not create a weakness in science. They are exposing one.
Most published research cannot actually be checked, because the data is not there to check.
When the data is not accessible, the quality of the science cannot be assessed. It can only be taken on faith.
The numbers, for anyone who has not lived inside this literature, are difficult to believe.
Gabelica and colleagues, writing in J Clin Epidemiol in 2022, contacted 1,792 corresponding authors of papers carrying “data available upon request” statements. They received data from 6.8% of them. Tedersoo and colleagues, in Scientific Data in 2021, found that across nine disciplines in Nature and Science, only about 42% of datasets with availability statements could actually be obtained from authors upon request. After that, decay accelerates: Vines and colleagues, in Current Biology in 2014, measured a 17% annual decline in data availability. Combined with Tedersoo’s at-publication baseline, that rate implies roughly 6% of the original data is still accessible a decade after publication. Freedman, Cockburn and Simcoe, in PLOS Biology in 2015, put the cost of irreproducibility in U.S. preclinical research alone at roughly $28 billion a year. The European Commission has put the annual cost of non-FAIR research data across the EU in the tens of billions of euros.
When I cite these numbers in workshops, the response is almost always the same. People nod. They already knew. They had been quietly assuming it was somebody else’s data that had vanished, not theirs.
These failures share a structural cause. The scientific record was designed to display findings. The evidence behind those findings has always been treated as supplementary.
A figure shows a claim. A plot summarizes data. A caption explains what the reader is looking at. None of them carries the experiment behind it: the dataset, the provenance, the quality checks, the analytical choices, the boundaries of reuse. That part lives elsewhere — if it lives at all. On someone’s institutional drive. In a folder named final_v3 that nobody has opened in three years. In an inbox belonging to a corresponding author who left for industry in 2019.
Synthetic figures are not a new category of misconduct disconnected from this. They exploit the same gap as the reproducibility crisis: the space between a scientific claim and the evidence needed to verify it. When verification is expensive and fabrication is cheap, trust gets brittle.
The PDF is good at what it was designed to do. It presents an argument. It fixes text, tables and figures into a stable, readable form. It is portable, familiar, and still essential for the kind of close reading we call peer review.
It does not carry raw measurements, variable definitions, instrument settings, processing steps, quality checks, schemas, provenance, governance constraints or version history in any form that people and machines can reliably use. Those things are flattened into prose, tables, and figures, which is a lossy compression. The article becomes the scholarly record. The data becomes supporting material. Supporting slides into supplementary. Supplementary slides into wherever.
Repositories help, and in many cases they are the only reason any data survives at all. Deposit alone is not enough. Sharing typically happens at the very end of a project, after the analysis is finished and the human context has already begun to drift away from the files. A zip with a README satisfies the letter of a policy without producing reusable scientific evidence — which is exactly why “available upon request” continues, in practice, to mean “not actually available.”
This is partly an infrastructure problem. Mostly, though, it is an incentives problem. We reward the research paper. We cite the narrative. We count the article. We rarely give equivalent credit to the dataset, the documentation, the cleaning, the governance, the quality control, and the long, unrewarded years of stewardship that make any of it reusable.
The reproducibility crisis, fragile data availability, synthetic figures and weak data stewardship all point to the same thing: findings are easy to read; the data behind them is hard to reach.
A better PDF will not solve this. Neither will another submission checklist. Another scholarly object has to join the research paper.
The dataset has to be treated as a first-class scholarly artefact, with the apparatus we already give to a paper. That means more than uploading a CSV alongside the manuscript. It means a structured, citable, peer-reviewable, reusable object that carries the context required to understand it.
That context is not abstract. It is the things any honest researcher already knows they would need from somebody else’s dataset before they would be willing to use it:
Provenance — who collected the data, when, where, on what instrument, under what protocol, with what calibrations.
Methods — collection procedures, processing pipelines, analytical choices, software versions, parameter settings.
Structure — variables, units, types, allowable values and relationships, in a machine-actionable form.
Quality — missingness, anomalies, completeness, validation checks, known limitations and uncertainty.
Governance — consent, access conditions, licensing, sovereignty constraints, responsible-use terms.
Reuse boundaries — what the data can support, what it cannot, what would be inappropriate or unsafe to infer from it.
Persistence — durable identifiers, durable hosting, indexing, a credible plan for the data to remain available.
This is what I mean by contextualized data. It is the difference between a number and an observation. A number is a string. An observation is a measurement that someone, somewhere, can trust enough to use.
Without that context, what we have is files. Not evidence.
When data is structured and accompanied this way, several long-standing problems start to look more tractable, often at the same time.
Reproducibility becomes less heroic. A reviewer, a funder, a competitor or an AI agent can load the dataset, read the metadata, rerun the analysis and test the claim — instead of reconstructing the project from email threads, supplementary PDFs and whatever the original PhD student happens to remember.
Cumulative science becomes realistic. Datasets connect across studies, instruments, cohorts and domains because their structure and meaning are made explicit, rather than encoded in the tribal knowledge of each individual lab.
Fabrication becomes harder. A plausible image is now easy to make; a coherent, provenance-rich, machine-actionable dataset that survives independent integrity checks and reanalysis is much harder to fake. There are more places where a forgery has to be internally consistent, and more places where it can fall apart.
Trust becomes less dependent on reputation. A claim backed by an inspectable dataset can be challenged, reused and built upon by anyone — not only by people who happen to know, and trust, the corresponding author personally.
Access becomes fairer. Whether a researcher in Lagos or São Paulo can verify or reuse a result should not depend on whether the corresponding author at MIT still answers email.
None of this requires a miracle. Most of the technology already exists — MLCommons Croissant for ML-ready packaging, Zenodo for persistent open deposit, OpenAIRE and Google Dataset Search for indexing, schema standards across most major domains. What has been missing is the commitment to treat datasets with the seriousness we have long reserved for papers.
That means review. Credit. Persistent identifiers. Indexing. Hosting. Machine-readable structure. Clear governance. And workflows that make all of these things normal, instead of exceptional.
There is a wave of “AI scientists” being built right now — FutureHouse, Sakana’s AI Scientist, the various agentic chemistry and biology pipelines, the autonomous research demos that come out roughly monthly. Some of the engineering is genuinely impressive. Some of the demos are striking.
The deeper claim — that these systems can reason over the scientific literature in any robust sense — is much weaker than the marketing suggests.
Most of these systems are reading text. They extract statements from papers — “X reduces Y,” “treatment A outperforms B,” “this method generalizes” — and then reason over the statements as if the statements were evidence. They are not.
A statement in a paper is a compressed claim. It carries the weight of the data, the sample size, the protocol, the controls, the missingness, the analytical choices and the limits the original authors put on the result. Strip that context away and you are left with a sentence. The agent can cite it, recombine it, summarize it and re-assert it. It cannot assess it.
If the data behind the claim is not contextualized, machine-actionable and inspectable, no amount of orchestration on top of the literature recovers what was lost. The agent reasons over a flattened version of science. It can produce fluent, plausible, internally consistent results. It cannot reliably tell whether those results are correct.
In some narrow domains, on some narrow questions, this happens to work. The literature is uniform, the constraints are tight, the variance is low. That is not robustness. That is luck.
This is the input-side mirror of the synthetic-figure problem. Synthetic figures show that fabrication on the output side is becoming cheap. “AI scientists” reading PDFs show that automated reasoning over the literature alone is structurally limited on the input side. Both fail in the same place: the polished representation is overweighted, and the underlying evidence is hard to reach.
If AI is going to genuinely accelerate science, rather than amplify the weaknesses we already have, the substrate has to change. Agents need to operate on the evidence underneath the claims, with provenance, methods, limitations and governance attached. Otherwise we are scaling up the same fragile foundation we already know is breaking.
When a reader replied to Rust’s thread asking whether reviewers couldn’t simply request the original data behind the figure, Rust was direct:
Yes and no. I think it’s also relatively easy to fabricate raw data that reproduce the same mean ± SD. Sharing raw images with metadata or uncropped Western blots, would definitely help, but I’m not sure we currently have the infrastructure to properly control and verify all of that. I’m also not sure whether, in 1–2 years, even these types of data could be convincingly fabricated.
He is right, and this is the harder version of the problem. Generative models are not stuck at the figure layer. Soon they will produce convincing underlying datasets — synthetic CSVs, fabricated time-series, simulated raw measurements at the pixel level. Some of this is already possible.
This is a real escalation, and it is also a different kind of problem from a fake figure, in a way that helps the verifier.
A figure has to look right under one constraint — visual plausibility. A dataset published with its methods, provenance, instrument metadata, statistical structure and analytical pipeline has to look right under thousands of constraints at once, all of which have to agree with each other under the stated method. Faking any one of them is easy. Faking all of them in mutual agreement, at the resolution where automated checks operate, is a categorically harder problem — and an AI verifier scales at least as fast as an AI forger.
This is a different arms race from the one we are currently losing on figures. We can plausibly stay ahead of this one, provided the data is published in a form that makes deeper checking possible. Why this asymmetry favors the verifier in practice — with concrete examples from instrument signatures, cross-modal correlations and statistical fingerprints — is the subject of a follow-up post.
The original FAIR principles, set out by Wilkinson and colleagues in Scientific Data in 2016, are the right starting point. Data should be Findable, Accessible, Interoperable and Reusable. Almost a decade on, FAIR is necessary but no longer sufficient. Data now needs to be usable not only by humans reading papers but by computational systems that need structure, semantics, provenance and computable governance. And it needs to be verifiable, not just declared.
FAIR² is an open specification that adds two requirements to FAIR that I think are now non-negotiable: AI-readiness, and responsible, verifiable governance. The “²” is those two additions. The specification is openly maintained and any publisher or platform can implement it.
Frontiers FAIR² Data Management is the service that implements FAIR² for researchers — built and operated by Senscience in partnership with Frontiers. This is what we have been building.
In practice, treating the dataset as a scholarly object in its own right means the service does the structural work for the researcher. That work is handled by Clara, the service’s AI curator. Clara validates the dataset (missingness, inconsistencies, format issues, PII screening), structures it to specification, and then drafts, packages and builds each of the outputs below. The researcher reviews and approves Clara’s work before submission to peer review. Concretely:
A peer-reviewed Data Article, drafted by Clara, with its own DOI, indexed in Scopus and Web of Science. The review focuses on the data itself: its quality, variables, methods, provenance, limitations and reuse potential. This is not the same review as the research paper, and it does not replace it.
A custom-built interactive data portal for every dataset, built by Clara — with data explorers, visualizations, and an AI chat that lets a reader, reviewer or agent ask questions about the dataset in plain language and run analyses without ever downloading the files. This is what reviewers actually use to interrogate the data, and it remains live for the lifetime of the publication, hosted in the cloud — so the dataset is no longer dependent on a former student’s hard drive, an author’s expired email, or a lab’s funding cycle.
A Jupyter notebook with worked analysis code, prepared by Clara and shipped alongside the dataset. The reader does not have to reconstruct the pipeline from a Methods section; they can open the notebook, rerun every cell and see how each result was produced. Reproducibility stops being an aspiration written into prose and starts being a button anyone can press.
Machine-actionable packaging in MLCommons Croissant, so the dataset can be loaded into modern AI workflows in one line of code instead of being a download that someone has to spend a week parsing.
A FAIR² Certificate that documents adherence to FAIR principles, AI-readiness and responsible governance, in a form that funders and institutions can audit programmatically rather than by reading a PDF.
Repository deposit in Zenodo (operated by CERN), indexed in OpenAIRE and Google Dataset Search — open infrastructure independent of any single publisher or platform.
This will not detect every fake figure. It changes the burden of trust.
A synthetic image can look convincing. A fabricated claim should not survive simply because the image is convincing. It should have to survive contact with the data object behind it: the provenance, the structure, the methods, the quality checks, the governance, the review, and the possibility of independent reanalysis by anyone who opens the notebook and reruns the cells.
That will not eliminate misconduct. Nothing will. It raises the cost of deception and lowers the cost of verification.
It moves the trust anchor to where it should always have been — underneath the figure, in the evidence itself.
Software went through a version of this transition about fifteen years ago. A package without a public repository, a license, version history, documentation or tests is treated with suspicion now. The culture changed first. Code without context became unacceptable. Git, semantic versioning, package indexes, signed releases — the infrastructure followed the cultural commitment, not the other way around.
Science is overdue for the same shift, for data.
The reproducibility crisis told us the foundation was thin. Generative AI is now showing us, in a way that is much harder to ignore, that the surface layer — the figure, the polished PDF, the visual impression of rigour — was never structurally load-bearing in the first place.
Better detection of fake figures will not save a record where the evidence underneath is already inaccessible. The evidence itself has to be made inspectable.
Make the dataset citable. Carry its context with it. Ship the analysis alongside it. Review it. Preserve it. Structure it for machines as well as humans. Give it governance that can be understood and audited. Treat it as scientific output in its own right, alongside the research paper.
Then the figure becomes what it should always have been: one visualization of evidence that lives somewhere others can actually check.
The mice in the figure above do not exist. For too much of the published literature, we cannot answer the more important question with any confidence: does the evidence behind the figure exist, in a form anyone — human or machine — can actually verify?
We need a scientific record where, for every paper, the answer is yes. That is the work in front of us. Frontiers FAIR² Data Management is one concrete way to do it. The ecosystem of tools building on the open FAIR² specification is what carries this forward.
FAIR² is an open specification. Frontiers FAIR² Data Management, built and operated by Senscience, is the implementation described here. The Data Article — peer-reviewed, indexed and assigned its own DOI — is the scholarly record of the dataset itself, complementing the research paper.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.