RSS Amplifier

The Variant · Jul 18, 2026

The Most Misleading Sentence in Scientific Publishing

0
Sign in to vote or save

Kuan-lin Huang, PhD · The Variant

Should an NIH-funded paper still be allowed to end with this sentence?

Data are available upon request from the corresponding author.

There may be no repository. No persistent identifier. No machine-readable metadata. No clear access criteria. No response deadline. No durable steward after the graduate student, postdoc, or principal investigator leaves the institution.

The prospective user sends an email and waits. I have tried this many times myself.

Perhaps the request is answered. Perhaps another form is required. Perhaps the original analyst is gone. Perhaps the files can no longer be located. Perhaps the request is quietly judged insufficiently “reasonable.”

The phrase may be appropriate when ethical, legal, privacy, consent, or contractual restrictions prevent open release. Controlled access is not the enemy of data sharing. For sensitive human data, it is often essential.

But “upon reasonable request” should not be the default substitute for depositing data that could responsibly be shared.

Our solution to credit and empower data sharing at theSindex.org

Share

The NIH Data Management and Sharing Policy requires investigators generating scientific data to submit a plan, budget for data management and sharing, and comply with the approved plan. These policies matter.

But policy primarily answers:

Did you make a plan to share the data?

It does not fully answer:

Did the shared data actually enable anyone else to do better science?

A dataset can technically exist while remaining nearly impossible to find, understand, access, or reuse.

Conversely, a carefully constructed cohort, atlas, reference database, benchmark, or molecular resource may power hundreds of downstream studies. Yet, the researchers who assembled, documented, curated, and supported it may receive far less academic recognition than the scientists who publish the resulting analyses.

That is an incentive failure.

Academia has mature systems for counting papers, citations, grants, and principal-investigator roles. We have much weaker systems for recognizing the scientific infrastructure that makes future research possible.

Our finalist entry in the NIH Data Sharing Index—or S-Index—Challenge begins with a simple premise:

A published dataset is worth the science it empowers.

We built theSindex.org as an open platform for measuring the downstream network impact of biomedical datasets and recognizing the researchers, institutions, and funders behind them.

Its core metric, DataRank, starts by identifying papers whose principal contribution is a shared dataset. It then examines the citation network around each resource.

Rather than treating every citation as identical, DataRank considers both:

  1. The probability that a given citation of a data resource paper is due to data reuse.

  2. The downstream influence of the papers that cite it, adjusted for how broadly those papers distribute their references.

Self-citations are excluded from the propagated network signal. The resulting score is converted into a percentile so that a dataset can be compared with other resources in the indexed corpus. The current platform covers a demonstration corpus of NIH-funded biomedical data papers and can calculate a score for any paper with a DOI, although network estimates should improve as the corpus expands.

The goal is not to create another closed ranking system.

The scores, citation networks, methodological components, CSV exports, and APIs are open so that others can inspect, challenge, reproduce, and build upon them.

A dataset may be highly cited but poorly documented.

Another may be exceptionally well curated but too new to have accumulated downstream use.

That is why we treat two questions separately.

DataRank asks:

What scientific activity did this dataset enable?

Our FAIR Agent asks:

How findable, accessible, interoperable, and reusable is the resource today, and what could its creators improve?

The FAIR Agent examines a paper and its sharing practices, then provides concrete recommendations such as depositing data in an appropriate repository, adding a persistent identifier, specifying a license, linking accessions directly from the publication, or improving documentation.

Importantly, the FAIR assessment is not secretly folded into DataRank. One measures observed downstream influence; the other provides prospective coaching on sharing quality.

The S-Index should not become merely another number scientists feel pressured to maximize.

Used well, it could become infrastructure for changing behavior.

Funders could identify data investments that generated unusually large downstream returns.

Institutions could recognize researchers who create widely reused scientific resources.

Promotion committees could credit data stewardship as a substantive research contribution.

Journals could distinguish between data that are nominally “available” and data that are persistently, responsibly, and practically reusable.

Researchers could receive guidance before publication—when deficiencies can still be corrected—rather than discovering years later that no one could reuse their work.

And early-career scientists who perform the often-invisible labor of organizing, documenting, validating, and supporting datasets could finally have evidence of their contribution.

At theSindex, we would like to open the discussion.

A few questions for researchers, editors, funders, data stewards, and open-science advocates:

Should journals continue accepting “available upon request” when an appropriate repository exists?

Should a data-sharing metric reward openness, downstream reuse, reproducibility, FAIRness, or some combination of this?

How should we measure the value of controlled-access human datasets without penalizing responsible privacy protections?

How should credit be distributed among the many people who create and maintain a dataset?

What safeguards would prevent an S-Index from becoming another metric that rewards scale, seniority, or strategic gaming?

Would you use a measure such as DataRank in funding, promotion, institutional reporting, or journal evaluation? What evidence would you need before trusting it?

Our position is not that we have solved every aspect of data-sharing measurement.

It is that the current alternative—requiring sharing while leaving its scientific value largely invisible—is no longer good enough.

Explore the current finalist platform at theSindex.org, test a DOI, inspect the methodology, and tell us where the approach succeeds or fails.

What would a fair, useful, and difficult-to-game S-Index look like to you?

#OpenScience #DataSharing #SIndex #NIH #FAIRData #ResearchImpact #Reproducibility

No posts

Read the original on drkuan.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.