BioMetadataAudit

Public genomic metadata you can trust — and trace.

OpenBioData recovers missing or inconsistent sample metadata in NCBI records by tracing every field back to its source paper — with a citation and a confidence score you can check yourself.

81% field recovery vs 19% from NCBI alone Every field cited Confidence scored 0–100
EXAMPLE RECORD — ILLUSTRATIVE BioSample
✕ As deposited
host
isolation_source
collection_dateunknown
geo_loc_name
✓ Recovered
hostHomo sapiens92
Source: linked publication, Table 2 · direct statement
isolation_sourcenasopharyngeal swab87
Source: linked publication, Methods · direct statement
collection_date2021-0374
Source: supplementary table, cross-checked against 2 citing papers
geo_loc_nameVietnam: Ho Chi Minh City81
Source: linked publication, sample table · confirmed across sources
Hover or tab into a recovered field to see its source →
The problem

The record exists. It just can't be trusted.

NCBI records are often missing exactly the fields that matter for real research — even when the answer is sitting in plain text in the paper that deposited the data.

01

A researcher needs 1,000+ public samples to study a disease population. He assumes the public record is accurate — that's the default assumption everyone starts with.

02

He opens the database. Country: blank. Sample type: wrong. Collection date: missing. So he reads the original papers by hand, one at a time, for weeks — before analysis even starts.

03

Run against his own hand-curated, trusted dataset, the same checks surface systematic errors — samples mislabeled by origin, era, or species. If a careful lab's own curation has this, the public commons has it too.

How it works

One pipeline. Two directions.

The same four-step process works on public records by default — or point it at your own internal data and it becomes an audit layer.

01

Trace public + internal sources

Give it an accession or a paper link. It finds the source publication and, if you connect one, your own internal records too — public sources are the default, internal is opt-in.

02

Cross-check

It doesn't stop at the original paper — it finds every publication that cites or reuses the same sample and checks them all against each other for agreement.

03

Recover & verify

Missing fields get filled in, existing fields get checked against the evidence, and every value ships with a citation and a confidence score — never a black box.

04

Validate internal data

Point the same pipeline at your own sample metadata instead of NCBI's, and recovery becomes auditing — surfacing exactly where your internal records disagree with the evidence.

Proof

Real records. Real gaps found.

81%
Fields recovered (vs. 19% from NCBI alone) — PRJNA976261 benchmark
20+
Weekly active researchers
4
Continents represented
0
Original NCBI records modified — every output is traceable and reversible
Academic curation lab

A metadata-curation lab at a major public health graduate school is using OpenBioData to speed up manual curation and flag likely errors before they enter their reference dataset.

Public health surveillance

A genomic surveillance program is running its antimicrobial-resistance sample set through the tool to recover context missing from the original NCBI deposits.

AI bioinformatics platform

An AI-native bioinformatics platform is piloting OpenBioData as the curated dataset layer feeding its own analysis agents.

Who it's for

Open for researchers. Private for teams.

Same engine, two ways to use it — depending on whether your data is public or yours alone.

Individual researchers

Free & open source

MIT-licensed, self-hostable, and free to try — built for anyone tracing public NCBI records.

  • 10 samples without an account, 30 signed in
  • Confidence score and source citation on every field
  • Self-host with your own API key — nothing sent to us
  • Open pipeline: source-fetching and scoring logic is auditable on GitHub
Try it free →
Teams & companies

Private validation, at your scale

For teams whose product or research depends on data quality they can't fully verify by hand.

  • Cross-check your own internal metadata, not just public records
  • Private deployment — your data stays yours
  • Schema-aligned output for your existing pipelines
  • Volume pricing and a scoped pilot before any commitment
Talk to us →
Team

Built by people who ran into this problem firsthand.

Founder & CEO

Vy Khanh Phung

B.S. Computer Science & Biochemistry, Dickinson College. Ran into public metadata quality problems firsthand during a research internship at Oxford, which became the starting point for OpenBioData.

CTO & Co-founder

Gowtham Gopalakrishnan

M.S. Data Science, University of Arizona. AI/ML engineer with hands-on research in computational drug repurposing and epidemic modeling — architects and leads engineering for BioMetadataAudit.

Get started

See what's missing in your data.

Try it free on public records, or talk to us about validating your own.