Matflow Benchmark Hub

Measured against reality

Read-only scores served from the public /api reports — nothing hardcoded, empty states instead of fallback numbers. Every table names its source; the methodology documents the splits, gates and bundles, and /science covers the models.

Zone 1 — Results hub (read-only)

Validation study

Model leaderboard

Held-out regression. Ensemble balances accuracy and stability.

ModelMSE ↓R² ↑Verdict
Validation study

Synthetic data fidelity

Statistical similarity of generated rows to experiments — distribution shape, category balance and column-pair correlation.

Validation study

Laboratory case study

Four Ni-hydrotalcite catalysts. Predicted vs measured. Sources below.

CatalystCalcination °CCH4 exp %CH4 AI %CH4 errorCO2 exp %CO2 AI %CO2 error

Sources: ACS Omega 2025 ↗ Gas Sci Eng 2025 ↗.

Validation study

External validation datasets

Four independent sets: sparse, small-data, rig, operational.

Validation study

Novel catalyst candidates

Top ten optimization outputs for prioritization.

#CatalystScoreCH4 conv %CO2 conv %H2/COTemp °CNi wt%SA m²/g
Datasets

Benchmark datasets — download and verify

Local copies with licenses. D-Epox-Cat-ML CC-BY-4.0 Zenodo 15750387.

CC-BY-4.0 · Zenodo 15750387

D-Epox-Cat-ML — cyclohexene epoxidation literature compilation

207 rows · yield target · used in the benchmark suite above

Bundled with Matflow

Ni-hydrotalcite dry reforming of methane (demo dataset)

51 experimental rows · 3 targets (CH4 conversion, CO2 conversion, H2/CO)

Download .csv ↓
Standard tasks

MatBench standard tasks

Official v0.1 tasks and folds. MAE / ROC-AUC. The Matflow column is the composition baseline, not the DEMO evaluator in Zone 2. Protocol: /benchmarks/methodology.

Battery RUL

Battery cycle-life prediction (Severson LFP)

124-cell Severson LFP (doi:10.1038/s41560-019-0356-8). 41 / 43 / 40 split. First-100-cycle elastic-net. Baselines are the published Severson Table 1; the Matflow row is DEMO until real features are supplied. Protocol: /benchmarks/methodology.

Real engines vs. legacy surrogates

What the physics engines actually deliver

M3GNet (matgl) and PyBaMM vs legacy surrogate tables and experiment. The old table showed ~0% error because it contained the answers; the real engine's residual error is its own PBE-level bias. Protocol: /benchmarks/methodology.

Regression gates

Gold suites — the regression gates

Six versioned fixtures with public reports and stated gate semantics — two enforce stop-ship in the runner, the rest assert their pass rates in CI. Full reports linked. Gate protocol: /benchmarks/methodology.

Fixtures + runners in backend/benchmarks/. CI fails on gate regression. /science.

Download the validation dataset

Run the numbers yourself

Ni-hydrotalcite CSV.

Zone 2 — Run it yourself (interactive) · runners, uploads, submissions, leaderboards, orgs

Collapsed by default. All controls stay mounted and functional when closed.

Your data

Benchmark Your Own Dataset

CSV audit. 3 splits. AutoML baselines. Conformal calibration.

Runner

MatBench runner

Sign-in required. First run downloads datasets.

Runner

Battery-RUL runner

Sign-in required. This button runs the labelled synthetic DEMO feature set so the page can render; a COMPUTED number requires a per-cell feature CSV or fetch mode through the API.

MLIP leaderboard

Machine-learned interatomic potentials — Matbench Discovery style

Model × task MAE / R² / F1. Live endpoints with fallback. Pairs with /viz. Methods: /science.

MLIP leaderboardComputed· E3loading…
modelMAE ↓R² ↑F1 ↑E_hull MAE ↓paramssource
No rows for task “e_form”.
Hull-error plot
MAE (x) vs E-above-hull MAE (y) · eV/atom
No hull errors for this task.
Matbench Discovery ranks stability classification (F1 on E-above-hull ≤ 0) above raw formation-energy MAE — the diagonal marks equal error on both.
Community

Research organizations

Opt-in public profiles. Member counts listed.

Participate

Evaluation & submissions

DEMO evaluator on synthetic data (disabled by default on hosted instances). Submissions are scored against the official task format; the verified leaderboards are the gold reports above.

Run an evaluation (DEMO)

POST /api/benchmarks/evaluate · synthetic task data + baseline estimators · DEMO-labelled, not comparable to the published leaderboard · disabled by default on hosted instances

Submit your model to MatBench

POST /api/evaluation/matbench/submit · scored against the official task format

Requires a signed-in session — anonymous submissions are rejected with 401. The scored result lands on the public submissions table below.

MatBench tasks

GET /api/evaluation/matbench/tasks · official task definitions with published model references

Scored submissions

GET /api/evaluation/matbench/submissions · every scored (predictions, targets) pair

Standardized leaderboards

GET /api/benchmarks/leaderboards · published suite leaderboards across Matbench, JARVIS and Catalysis

Validate your own materials

Bring experimental rows. Pipeline returns scored candidates.