Measured against reality
Read-only scores served from the public /api reports — nothing hardcoded, empty states instead of fallback numbers. Every table names its source; the methodology documents the splits, gates and bundles, and /science covers the models.
Zone 1 — Results hub (read-only)
Model leaderboard
Held-out regression. Ensemble balances accuracy and stability.
| Model | MSE ↓ | R² ↑ | Verdict |
|---|
Synthetic data fidelity
Statistical similarity of generated rows to experiments — distribution shape, category balance and column-pair correlation.
Laboratory case study
Four Ni-hydrotalcite catalysts. Predicted vs measured. Sources below.
| Catalyst | Calcination °C | CH4 exp % | CH4 AI % | CH4 error | CO2 exp % | CO2 AI % | CO2 error |
|---|
Sources: ACS Omega 2025 ↗ Gas Sci Eng 2025 ↗.
External validation datasets
Four independent sets: sparse, small-data, rig, operational.
Novel catalyst candidates
Top ten optimization outputs for prioritization.
| # | Catalyst | Score | CH4 conv % | CO2 conv % | H2/CO | Temp °C | Ni wt% | SA m²/g |
|---|
Benchmark datasets — download and verify
Local copies with licenses. D-Epox-Cat-ML CC-BY-4.0 Zenodo 15750387.
D-Epox-Cat-ML — cyclohexene epoxidation literature compilation
207 rows · yield target · used in the benchmark suite above
Ni-hydrotalcite dry reforming of methane (demo dataset)
51 experimental rows · 3 targets (CH4 conversion, CO2 conversion, H2/CO)
MatBench standard tasks
Official v0.1 tasks and folds. MAE / ROC-AUC. The Matflow column is the composition baseline, not the DEMO evaluator in Zone 2. Protocol: /benchmarks/methodology.
Battery cycle-life prediction (Severson LFP)
124-cell Severson LFP (doi:10.1038/s41560-019-0356-8). 41 / 43 / 40 split. First-100-cycle elastic-net. Baselines are the published Severson Table 1; the Matflow row is DEMO until real features are supplied. Protocol: /benchmarks/methodology.
What the physics engines actually deliver
M3GNet (matgl) and PyBaMM vs legacy surrogate tables and experiment. The old table showed ~0% error because it contained the answers; the real engine's residual error is its own PBE-level bias. Protocol: /benchmarks/methodology.
Gold suites — the regression gates
Six versioned fixtures with public reports and stated gate semantics — two enforce stop-ship in the runner, the rest assert their pass rates in CI. Full reports linked. Gate protocol: /benchmarks/methodology.
Fixtures + runners in backend/benchmarks/. CI fails on gate regression. /science.
Run the numbers yourself
Ni-hydrotalcite CSV.
Zone 2 — Run it yourself (interactive) · runners, uploads, submissions, leaderboards, orgs
Collapsed by default. All controls stay mounted and functional when closed.
Benchmark Your Own Dataset
CSV audit. 3 splits. AutoML baselines. Conformal calibration.
MatBench runner
Sign-in required. First run downloads datasets.
Battery-RUL runner
Sign-in required. This button runs the labelled synthetic DEMO feature set so the page can render; a COMPUTED number requires a per-cell feature CSV or fetch mode through the API.
Machine-learned interatomic potentials — Matbench Discovery style
Model × task MAE / R² / F1. Live endpoints with fallback. Pairs with /viz. Methods: /science.
| model | MAE ↓ | R² ↑ | F1 ↑ | E_hull MAE ↓ | params | source |
|---|---|---|---|---|---|---|
| No rows for task “e_form”. | ||||||
Research organizations
Opt-in public profiles. Member counts listed.
Evaluation & submissions
DEMO evaluator on synthetic data (disabled by default on hosted instances). Submissions are scored against the official task format; the verified leaderboards are the gold reports above.
Run an evaluation (DEMO)
POST /api/benchmarks/evaluate · synthetic task data + baseline estimators · DEMO-labelled, not comparable to the published leaderboard · disabled by default on hosted instances
Submit your model to MatBench
POST /api/evaluation/matbench/submit · scored against the official task format
Requires a signed-in session — anonymous submissions are rejected with 401. The scored result lands on the public submissions table below.
MatBench tasks
GET /api/evaluation/matbench/tasks · official task definitions with published model references
Scored submissions
GET /api/evaluation/matbench/submissions · every scored (predictions, targets) pair
Standardized leaderboards
GET /api/benchmarks/leaderboards · published suite leaderboards across Matbench, JARVIS and Catalysis
Validate your own materials
Bring experimental rows. Pipeline returns scored candidates.