How these numbers are made
The Matflow Benchmark Hub is trustworthy only if its methods are public and its gates are executable, not policy slides. This page is that contract: the split protocols, calibration and gating rules the runners and routes actually enforce — and how to re-run any published number yourself.
Every value is classified by how it was produced
A number's class decides what it may be used for — and the class travels with the number through the API, the UI, exports, reports, and training filters. Each class also carries an evidence-level ceiling (E0 demo … E5 certified measurement); a card may claim a level below its class ceiling but never above it, and the runtime clamps an over-claim instead of trusting it.
| Class | Meaning | May it enter production decisions or training? |
|---|---|---|
| MEASURED | Raw experimental / instrument result linked to a sample, method, and quality state. | Yes, after the defined quality/review gate; the measurement is E5 at the ceiling. |
| COMPUTED | Reproducible external/open solver run with exact inputs, engine version, and artifacts. | Yes, with fidelity and limitation labels (E4 ceiling). |
| PREDICTED | Output of a registered statistical/ML model with uncertainty and an applicability-domain result. | As a reviewed recommendation, never as a measured fact (E3 ceiling). |
| EXTRACTED | Provisional value derived from a document/image/source. | Only after citation-linked human review (E2 ceiling). |
| HYPOTHESIS | LLM/generative/agent proposal or mechanistic interpretation. | No; a proposal until independently evidenced (E1 ceiling). |
| DEMO | Synthetic, heuristic, stub, or educational output. | No; excluded from production datasets and decisions (E0). |
Engine cards extend the class with the specifics: which engine ran, its version, license, citation, and its known limitations — e.g. the M3GNet card documents the PBE-level lattice bias, and the battery analytic fallback is labelled as such. A skeptic can reach the source of any number from its card instead of taking the badge on trust.
Random is diagnostic. Extrapolative and cluster are the hard questions.
The hub publishes three questions for one dataset, because publishing only the first flatters the model: can it interpolate among formulations like the ones I have; can it rank things better than my best; can it say anything useful about chemistry it has never seen.
Random holdout (seeded)
The classic benchmark split: a seeded 80/20 shuffle. Every held-out row sits between training rows, so the model interpolates. Useful as a diagnostic and for reproducing published random-split results — never the headline on its own.
Extrapolative (top-performer) holdout
Rows are ranked by the mean of their standardised targets; the model trains on the lowest-performing 80% and is scored only on the top slice it has never seen. This is the question a screening tool is actually asked — "can it rank things better than my current best?" — and it is much harder. On this split the holdout is a narrow band of top performers, so its spread is tiny and R² can read as catastrophic even when the model comfortably beats the no-model alternative. That is why the headline on extrapolative runs is skill vs no model, not R².
Cluster (leave-one-cluster-out) holdout
KMeans clusters the standardised inputs; an entire cluster — a whole region of chemistry the model has never seen — is held out, with the cluster closest to the requested fraction chosen so the holdout stays comparable. If the data is too small or too clustered to split honestly, the runner refuses and records a failure rather than silently falling back to a random split under a cluster label.
Grouped / leave-one-group-out
Used by the base-model harness: whole groups (catalyst family, composition cluster, lot) are held out so no group appears in both train and test. This is the honest answer to "does it work on new materials, not just new rows?" and is reported as the headline for that program.
Temporal (time-forward)
Also from the base-model harness: the earliest rows train, the latest rows test, for time-on-stream, cycle-life and campaign data. It never shuffles time.
Baselines on the identical split
Every split is scored against simple models fitted on exactly the same train/holdout rows: the training-mean predictor, ridge regression, and an untuned single random forest, plus the harness baseline set where applicable. The comparison is published even when a baseline wins.
Calibration and applicability
Conformal interval coverage, out-of-domain/extrapolation flags, abstention behaviour and per-domain coverage are reported alongside every metric, because a single aggregate score can hide the domain a user actually cares about.
Intervals with stated, audited coverage
The predictor calibrates a conformal half-width per target from the absolute out-of-fold residuals of the same cross-validation the metrics already compute: the finite-sample-corrected order statistic at the requested level (90% by default), not a silent sample quantile. Small datasets (under 10 rows) return no interval rather than an unstable one. Coverage is reported on those calibration residuals and explicitly labelled in-sample — never passed off as a held-out estimate.
Where a held-out number exists, it is reported as such: the live benchmark computes coverage from the pipeline's own prediction intervals on rows the model never trained on, and the public dataset audit builds a 90%-quantile band from training residuals and reports the fraction of held-out rows inside it. A per-domain coverage audit additionally splits residuals by feature quartile/level and lists every domain that falls below the nominal level, so an overall 90% cannot hide an under-covered temperature range or support material.
The ten-minute floor, published first
Every suite run is scored against deliberately untuned baselines fitted on the identical train/holdout split: ridge regression (the "is this even non-linear?" check) and a single random forest (the "does the ensemble beat one decent tree model?" check). The table reports each baseline's RMSE next to the ensemble's and marks the rows that beat it, because anyone evaluating the platform will ask that question and answering it first is worth more than winning it. The predictor's own default ensemble exists in this context: a regularised linear model is opt-in precisely because it outperformed the tree-heavy default on the extrapolative split of the bundled dry-reforming dataset, and that result is published instead of buried.
What the generator comparison measures — and what it does not
The validation study compares generators on distribution shape (ks-complement for continuous columns, total-variation complement for categorical), column-pair correlation similarity, and an ML-efficacy indistinguishability score from a classifier trained to tell real from synthetic rows — the harder that discrimination, the more faithful the synthesis. Privacy metrics (new-row synthesis, distance-to-closest-record protection, overfitting protection) and an anomaly score on the real rows run alongside. These metrics quantify statistical fidelity; they do not prove scientific validity. A high score means the synthetic rows are usable for modelling, never that they are experimental evidence.
Official tasks, official folds — and a DEMO evaluator labelled as one
The MatBench table runs Matflow's lightweight composition baseline against the official MatBench v0.1 tasks and the official 5-fold splits, bundled verbatim from the upstream validation file. Datasets are downloaded once from the official host and cached; the baseline is a five-member gradient-boosted-tree ensemble over seven pymatgen-derived composition features, scored with the official metrics (MAE for regression, ROC-AUC for classification) so the numbers sit next to the published leaderboard under identical folds. Only composition-input tasks are enabled in v1; the leaderboard is linked, never hardcoded, and tasks with no cached dataset are reported as not run.
The interactive evaluator is a DEMO and every response says so: it trains a real sklearn estimator on a deterministic synthetic dataset (or relabels it under the requested published model name, which the service displays only as a truthful "(demo surrogate)" name). It is never comparable to the published leaderboard, and the public route is disabled by default on hosted instances. The genuine result is separate and untouched: GET /api/evaluation/matbench/real serves only the locally computed gold reports from benchmarks/run_real_matbench.py, with evidence class COMPUTED and real offline data; tasks with no report return an honest empty state, not a substitute number.
Severson LFP: bundled split, published baselines, labelled data modes
The battery benchmark uses the bundled 124-cell Severson LFP split (41 / 43 / 40 cells) and the published Table 1 baselines. The model is an elastic net on log10(cycle life) from the standardised Severson feature block — the same family as the paper. Three data modes are possible, and the response states which ran: a supplied per-cell feature CSV (COMPUTED, directly comparable to the paper), fetching the Q(V) curves from the MIT reference repository on first use (COMPUTED, network required once), or a clearly labelled synthetic DEMO feature set so the page can render. Results are cached for 24 hours and carry a reproducibility SHA-256 over task, data mode, model and scores; the public table shows the evidence class on the Matflow row itself.
Real engines, legacy tables, and what the errors mean
The engine benchmark compares the real M3GNet (MatPES-PBE) potential against the legacy surrogate table and experimental lattice constants for six prototype materials, and spot-checks PyBaMM (SPM, Chen2020) against the in-house analytic SPM on an NMC811/graphite 1C discharge. The legacy table showed near-zero error only because it contained the answer; the real engine carries genuine physics, and the residual errors are the potential's own PBE-level bias (for example the layered c-axis), documented rather than corrected away.
Versioned fixtures and executable gates
Versioned fixtures
Every suite is a versioned fixture set with a manifest recording suite, release and provenance (backend/gold/*, backend/corpora/*), so the same code plus the same fixtures yields the same report.
Production code only
Runners evaluate the shipped services, not a special evaluation build. CF-DataMap scores services.column_mapper.infer_column_roles; CF-Instrument scores the characterization parsers; CF-Atlas scores services.atlas_service.extract_catalyst_records_from_text.
Explicit gate semantics
CF-DataMap enforces its stop-ship condition in the runner (unsafe input↔output conversions must be 0.0); CF-Instrument enforces a 1.0 pass rate over its 16 checks. CF-Agent, CF-Sci-Record and the base-model program publish pass rates / holdout metrics that CI tests assert. CF-Atlas is deliberately not a pass/fail gate — it publishes real field recall and the count of records needing review.
Honest headlines
A gate that fails stays visible with its real number. CF-Atlas publishes its actual field recall and routes uncertain fields to human review rather than rounding up; a gold report that fails its serve-time hash check is never selected as a best run.
CI-reproducible artifacts
Each runner writes a public JSON report and names its regeneration command on the serving endpoint. Tests import and execute the runners (backend/tests/test_gold_suite.py, test_wave3c_instrument_evals_agent.py, test_wave3d_agent_gold_claims.py, test_wave3f_outbox_cost_gate_sci.py, test_wave3_gold_atlas_split.py), so a deterministic gate regression fails the test job and the build.
The six suites: CF-DataMap (column-role mapping against 10 fixtures; unsafe input↔output conversion rate must be 0.0) · CF-Instrument (XRD/BET/TPR parsers against exact synthetic ground truth; all 16 checks must pass) · CF-Agent (10 tool-layer policy cases, including confirm-before-execute) · CF-Sci-Record (campaign/record round-trips across a fresh database) · CF-Atlas (literature extraction, published honestly at its real field recall with the review-queue count) · CF-Materials-Base (base-model program: random vs grouped splits and a warm-start comparison with true holdout metrics, evidence E2). Every report is public and regenerable from its named command.
Every leaderboard row is a bundle, not a claim
Each live and suite run carries its dataset identity, seed, split kind, trained/holdout row counts, per-target metrics, prediction artifacts and log, downloadable as a full reproducibility bundle (.zip) that includes the dataset, the seeded train/holdout CSVs, the column configuration and the metrics JSON — re-runnable with python -m benchmarks.benchmark_runner. When a run directory is transient, the served bundle is rebuilt deterministically from the committed published record and dataset, and says so in its README. Locally computed MatBench and JARVIS gold reports carry a serve-time SHA-256 over their canonical payload; a missing or mismatched hash marks the report verified: false / status: hash_mismatch and it is never selected as a best run. No number is hardcoded on the hub: every figure is served from a report or artifact, and the public claim registry ↗ registers every externally quoted claim against a dated evidence reference.
From a rendered value back to its artifact
- Read the row's evidence class and engine card first. A DEMO value is excluded from decisions and training by policy; an analytic fallback names itself.
- For a suite or live result, open the source endpoint named on the section (all public reports are served from
/api/public/benchmarks/*) and copy the run id, split kind, seed and row counts. - For a live/suite run, download the reproducibility bundle and re-run the runner with the same seed and split; the bundle contains the exact dataset, split and configuration.
- For a gold suite, run the regeneration command shown on the endpoint (for example
python -m benchmarks.gold_runner) and compare to the served report. - For MatBench and JARVIS rows, check the report's
report_sha256against its canonical payload; a hash mismatch means the served row is not verified. - For community submissions, remember
verifiedmeans "scored by Matflow against the official task format", not "a good model"; the delta vs the best gold is reported as-is with no better/worse claim.
Working through this page alongside /science is enough to reconstruct how any figure on the hub was produced and to reproduce it on your own hardware.
The claims we make publicly — and the evidence for each
Hold us to these rules
Every number on the hub can be re-run from its artifact. If one can't, that's a bug — report it and the gate will catch it too.