How Matflow actually works
A methodology-level account of every stage — written against the shipped implementation, with the evidence class, engine card and public benchmark behind every class of number.
Overview
Experimental R&D is data-scarce: a typical campaign — catalysis, formulation, battery, polymer — produces tens, not thousands, of rows. Matflow's pipeline exists to make that scarcity workable: generate statistically faithful synthetic rows, prove the synthesis is trustworthy before using it, train a model that can explain itself, and then search that model's response surface for candidates worth testing in the lab — ranking them by performance and cost. Every number the pipeline produces carries an evidence class (measured, computed, predicted, extracted, hypothesis, demo) and opens to an engine card naming exactly which engine produced it. The same discipline drives the public benchmark hub, whose split protocols, pass gates and bundles are documented at /benchmarks/methodology.
5-Stage End-to-End Discovery & Scale-Up Graph
Combines 5 tree & linear regressors with physics transform layers (inverse, log, logit). Computes SHAP waterfalls and conformal confidence bounds.
Data Synthesis
Matflow's Data Synthesis engine offers three modes:
- Matflow Neural — a conditional-GAN-style mode designed for tabular data with mixed continuous/categorical columns and multi-modal distributions. The strongest default for larger, more complex datasets.
- Matflow Hybrid — combines a GAN-style network with a copula transform, often more stable than Matflow Neural on smaller samples.
- Matflow Statistical — a fully statistical (non-neural) mode that fits marginal distributions and a correlation structure directly. Fastest, and the most defensible choice when rows number in the tens rather than hundreds.
Mode selection is data-size aware: below 500 usable rows the engine runs the statistical (Gaussian-copula) path regardless of the requested mode, and logs that it did. Generated rows then pass a deterministic feasibility pass, in this order: an under-covering column is first mapped onto the requested [min, max] with a monotonic affine stretch (a positive affine map preserves Pearson and Spearman correlations, so only the marginal is rescaled); values are then clipped to the requested bounds; non-negative physical quantities (pressure, flow, time, loading, energy, …) are floored at zero, Celsius temperatures at −273.15 °C and pH at 0–14; and detected composition groups are projected onto their 100% (or 1.0) simplex so multi-component formulations stay consistent. Candidate batches are deduplicated and re-sampled until the requested row count is reached, and discrete columns are rounded before deduplication so near-integer duplicates cannot slip through.
Quality evaluation
Synthetic data is never trusted by assumption. Every batch can be scored with:
- Quality Report — the SDMetrics quality report plus per-column detail: a ks-complement ("Distribution Match Score") for continuous columns, a total-variation complement ("Category Match Score") for categorical columns, mean/std statistic similarity, and a column-pair correlation similarity. A Diagnostic Report adds data-validity and data-structure findings.
- ML efficacy — an Indistinguishability Score from a logistic detector: 1.0 means a classifier cannot tell real from synthetic rows apart; 0.0 means they are trivially separable.
- Dimensionality reduction — a PCA projection (linear) and, when
umap-learnis installed, a UMAP projection of real vs. synthetic rows, saved as interactive JSON and static plots, so divergence is visible rather than only scored. - Privacy — the Privacy Shield trio from SDMetrics: new-row synthesis (what share of synthetic rows are genuinely novel rather than near-duplicates), distance-to-closest-record baseline protection, and an overfitting-protection check. The overfitting check needs a real train/validation split, so it is skipped with a stated reason on very small datasets.
- Outlier/anomaly detection — the Matflow Anomaly Detector scores the real rows with two unsupervised PyOD detectors (ECOD, an empirical-CDF method, and COPOD, a copula-based method) at a fixed 10% contamination prior, returns the worst rows first with the score from each detector, and states when the data is too small (fewer than 2 usable numeric columns or 10 rows) to score.
These scores are diagnostics, not automatic gates: the evaluation report is what tells a scientist whether the augmentation is safe to model on. The public validation study reports synthetic-data fidelity with the same family of metrics; see /benchmarks.
Predictive modeling
The predictor is the Matflow Predictive Ensemble — a weighted voting ensemble, not a single model, because no one algorithm is reliably best across the small, heterogeneous datasets this platform sees. The default five members are an auto-tuned XGBoost-style gradient-booster, a random forest, CatBoost and LightGBM boosters, and partial least-squares regression (the linear member), combined at weights 3 : 2 : 1 : 2 : 2. A ridge-regression member is available as an explicit opt-in because the public benchmark suite found plain ridge beating the default ensemble on the extrapolative split of the bundled dry-reforming dataset — a result the platform publishes rather than hides. Every member is pinned to single-threaded execution internally; letting any one of them multithread while the ensemble itself is cross-validated causes severe CPU oversubscription (a training run that should take under two minutes can otherwise take twenty).
Feature engineering ahead of training includes:
- A physics-derived transform layer (inverse, log, sqrt, square) whose breadth scales with data size — none below 15 training rows, log+square below 40, all four above — plus a logit transform for targets whose name and observed values look like a bounded percentage, with each target's observed [lo, hi] range stored for the inverse transform.
- Up to three partial-least-squares latent components from the transform layer as additional features.
- Automated recursive feature elimination under a sample-aware budget: never more than ~n/3 or 15 selected features, whichever is smaller.
- Optional Matflow composition descriptors (132 Magpie elemental descriptors) for any input column holding a chemical formula, plus structure-aware atomistic descriptors when those inputs are configured.
Explainability is not an afterthought: tree-SHAP attributions are computed on the fitted ensemble (the PLS member is excluded because it has no tree path, and the tree weights are renormalized and documented as such), then rolled up from physics/categorical/descriptor sub-features to one bar per input the user configured; model-agnostic permutation importance runs alongside. Uncertainty is reported two ways. Conformal prediction intervals use a cross-conformal half-width: the finite-sample-corrected order statistic of the absolute out-of-fold residuals from the same cross-validation the metrics already compute (so honest intervals cost no extra training time), with coverage reported on those calibration residuals — not passed off as a held-out estimate — and a per-domain coverage audit that lists feature bins/levels where the interval under-covers. Second, an ECOD extrapolation flag marks predictions whose inputs fall outside the training envelope, because the model cannot tell you itself when it is guessing. Classification targets instead get per-target soft-voting classifiers with per-class probabilities and a confidence threshold; an interval around a class label would be meaningless. Each run also writes a model card tying the metrics to the dataset hash, configuration and runtime package versions.
Quantum & active-site physics
Active-site physics enters the tabular pipeline as descriptors, not as isolated demos. Matflow's composition descriptor pipeline computes Hammer–Nørskov-style d-band-center trends, and the fast estimators add composition-rule adsorption energetics for candidate surfaces; the descriptor-based catalyst screener ranks candidates on that pipeline; and the analytic Sabatier two-leg volcano model plots a distinct descriptor and optimum per reaction, built on the published linear-scaling and volcano literature (Bligaard et al.; Medford et al.) and labelled PREDICTED because it is a trend model, not a measurement. Where a real quantum engine or MLIP is installed, the same quantities can be recomputed at higher fidelity, and the engine card states which tier actually ran. The ledger below shows how the evidence classes render side by side.
Multi-objective optimization
Two search strategies are available. Matflow Adaptive Search is surrogate-guided: on small datasets (under ~30 rows) it searches locally around the elite training points, and on larger ones it runs a global surrogate-guided search whose objective is the same ensemble used for final scoring. Matflow Evolutionary Search runs a genuine multi-objective genetic algorithm — NSGA-II by default, with NSGA-III, MOEA/D and SMS-EMOA selectable — over the same surrogate. Candidates are ranked by TOPSIS against the configured objective weights, and Pareto dominance is computed directly so the result is an actual trade-off front, not a single "best" answer. Compositional constraints (a set of columns that must sum to 100%, essential for formulations such as catalyst compositions) and per-feature bounds are enforced during candidate generation, not filtered out afterward. Both searches find good fronts, not provably optimal ones.
TEA
The cost-aware optimizer, TEA, prices every candidate from five modules feeding one annualized unit cost per kilogram: raw-material quantities (which can be linked per recipe to composition columns), equipment sizing, CapEx factors over a stated project life, OpEx (utilities, labour, maintenance) and spent-material end-of-life economics — a metal-recovery credit in the catalysis pack. Equipment sizing uses the standard six-tenths power-law rule with a default 0.6 scaling exponent (cost ∝ (actual_size / base_size)^0.6) when actual and base sizes are supplied. Where the OpenPyTEA package is installed, equipment costing can delegate to it and the response states which engine was used. These are screening-grade engineering estimates over the stated boundary, not vendor quotes.
Screening, design of experiments & active learning
Virtual screening ranks a candidate library by structural similarity — Morgan fingerprints with Tanimoto similarity via the Matflow Cheminformatics Engine — to a reference molecule, with an optional SMARTS substructure filter; unparseable molecules are reported honestly rather than silently dropped. Design of Experiments generates full-factorial and fractional-factorial, Box–Behnken, central-composite, Latin Hypercube and Sobol designs — the inverse question to prediction: what should be measured next, not what a given formulation will do. Circumscribed central-composite designs deliberately place axial ("star") points outside the factor range to estimate curvature, and Sobol sequences are most uniform at power-of-two sample counts; the returned design notes both properties. Active learning closes the loop: the Matflow Active Learning Engine uses a BoTorch Gaussian-process model (qLogEI expected improvement by default; qNEHVI hypervolume improvement for multiple targets) to propose the next experiment batch from an optimization job's history, falling back to a deterministic heuristic or committee surrogate when the full engine is not installed — the response always states which engine actually ran. Strategies include expected improvement, UCB, probability of improvement, uncertainty sampling, query-by-committee, and cost-aware value-per-cost ranking.
Atomistic & quantum engines
Beyond the tabular pipeline, Matflow runs a tiered physical-sciences stack and states on every result which tier actually executed. Fast estimators cover composition-based formation energy, d-band center, adsorption energetics and band gap, plus a real ASE EMT relaxation for parameterised metals (a semi-empirical potential, labelled as such); HPC submissions render and queue Quantum ESPRESSO Slurm jobs via SSH when HPC_SSH_USER and HPC_CLUSTER_HOST are configured, and refuse honestly rather than fake a periodic DFT result when no cluster exists. Machine-learned interatomic potentials are probed live — MACE, CHGNet, MatterSim, SevenNet, ORB, FAIR-Chem (OMat24), M3GNet, ALIGNN or EMT, whichever is installed — and drive relaxation, MLIP molecular dynamics, equation-of-state fits and potential benchmarks; MLIP output is PREDICTED, never COMPUTED, and a missing model returns no number, with an install hint instead. Molecular dynamics is rendered server-side as real GROMACS (.mdp presets, five-file decks, RMSD/RMSF/radius-of-gyration/ H-bond trajectory analysis), OpenMM (platforms, runs, seeded replicates; falls back to an ASE EMT minimisation) and LAMMPS (NPT decks, vacancy workflows) input decks with Slurm templates. Phonons use a native Phonopy port for displacements, force constants, band structures, DOS, thermal properties and mode-animation data, with forces from the unified MLIP driver (default EMT) or a caller-supplied DFT force set; imaginary modes are reported, never hidden, and without forces only the displacement dataset is returned rather than zero-filled band/DOS fields. A composition heuristic is labelled as such; a real engine run is COMPUTED.
Battery physics
Battery work has two layers. The physics layer runs real PyBaMM continuum models (SPM by default; DFN is switchable in the service configuration; published parameter sets such as Chen2020 NMC111, Ai2020 LFP and Sulzer2019 LCO) wherever PyBaMM is installed, and falls back to in-house reduced-order SPM-style discharge and equivalent-circuit EIS models otherwise — both deterministic, both labelled by engine card. The depth layer fits semi-empirical capacity fade (power-law with an Arrhenius temperature term, 1800 K scale) to cycler history, reports remaining-useful-life at a stated end-of-life threshold with a split-conformal 90% prediction interval on the log-loss axis, and screens electrolyte mixtures by viscosity/density/conductivity with honest group-contribution and Arrhenius-mixing fallbacks (explicitly not a trained QSPR). The /benchmarks hub additionally runs the Severson LFP cycle-life benchmark against the published Table 1 baselines, labelled DEMO until real feature data is supplied.
Characterization & enrichment
Measured instrument files are parsed into numbers, not screenshots: XRD (background, peaks, FWHM, Scherrer crystallite size, Bragg d-spacing, phase matching), BET (multi-point surface area plus a real BJH pore-size distribution), TPR/TPD (multi-peak Gaussian deconvolution, H2 consumption), chemisorption (dispersion, active-site density, TOF normalisation), DSC (Tg, Tm, enthalpy of fusion, crystallinity) and TGA (onset/peak, residual mass), plus XPS/FTIR/Raman baseline and peak fitting with functional-group assignments. The characterisation narrations ground an LLM in those computed numbers — the model interprets, it does not invent metrics — and every result carries the parser's evidence class (MEASURED for live instrument data, EXTRACTED for vendor-file extraction) with the producing engine card. The enrichment layer parses free text and notes into material/formula, process, property, value and units (LLM when configured, deterministic regex fallback otherwise) and adds element-level descriptors for composition columns. LLM interpretation can be wrong; treat it as a reading aid over the deterministic numbers, not the other way around.
Literature intelligence
Matflow Atlas is the literature layer. It searches OpenAlex for scholarly works (single queries with a server-side contact, no bulk scraping), ingests a DOI-cited record, and extracts structured catalyst/material rows from text with value-level fields. Extraction is EXTRACTED evidence (E2): every stored record keeps its source DOI, and records land in a human review queue as pending_review until a person accepts or rejects them — only reviewed records may be treated as data. The CF-Atlas Gold v0 benchmark scores the production extractor on versioned source paragraphs and publishes its real field recall and count of records needing review; extraction at v0 is imperfect by design and the benchmark quantifies it rather than hiding it. Reviewed records can be exported to datasets, and they feed descriptor screening and campaign planning. The report is public on the benchmark hub.
Sustainability (LCA)
The LCA studio prices the environmental side of a composition: cradle-to-gate global warming potential (kg CO2eq / kg), cumulative energy demand (MJ / kg), waste E-factor (kg waste / kg product) and a 0–100 ESG sustainability score, with LCA flowsheets, BioSTEAM flowsheet simulation (delegated when the package is installed) and marginal-abatement-cost / carbon-sensitivity analysis. These are screening-grade COMPUTED estimates over the stated boundary, not a certified ISO 14040 study: the licensed ecoinvent database is never bundled, and when an operator supplies one — directly or through a Brightway install — the response reports which database was actually used.
Molecular design
For molecular and catalyst work, the platform connects several design engines under one evidence discipline. Docking prepares real receptors and boxes, runs single, flexible-residue and ensemble docking via AutoDock Vina or Smina when installed, and rescores and clusters poses; without a binary every scoring path is labelled DEMO. Free-energy perturbation builds atom-mapped alchemical networks (star, fully connected or minimal-spanning) for relative binding energetics, with a real OpenFE/OpenMM sanity path when both are installed and a validation-only DEMO otherwise. Chemical-space projection/diversity tools place ligand sets in context. Generative chemistry samples and diffuses candidates (GFlowNet-style sequence sampling and crystal-diffusion-lite are HYPOTHESIS, E1) and scores them with reward functions over RDKit QED, synthetic-accessibility, docking-proxy and MLIP-energy-proxy terms, alongside QSAR train/predict and an OPERA-style ADME/Tox panel (a rule-based analogue when the weights are absent). Reaction networks estimate elementary-step barriers with Brønsted–Evans–Polanyi scaling, render Sabatier volcano curves, and plan retrosynthetic routes with estimated yields, route confidence and a benchtop protocol. Real engine output is COMPUTED; analytic or heuristic output is PREDICTED / HYPOTHESIS / DEMO and says so.
Campaigns & the digital thread
Campaigns are the primary research object: an objective with constraints, linking the datasets, jobs, reports and candidates that belong to it. Around them, the ELN gives samples, containers, microplates/wells and aliquots first-class records with versioned experiment templates; instrument results attach to the originating sample as EXTRACTED rows. The LIMS layer tracks physical metadata, creates child aliquots that deduct from the parent, derives multi-source samples with mass-weighted concentration, checks volume conservation, and returns a lineage DAG whose edges carry relationship, transfer volume and dilution factor — the chain of custody from a result back to its original lot. Microplate layouts (96/384) support serial dilution series and export directly as Opentrons protocols. Integrity is technical, not claimed: records are hashed into a tamper-evident chain (canonical SHA-256 always; HMAC-SHA256 signatures when ELN_SIGNING_KEY is set), and e-signatures are immutable, hash-chained rows that can be independently re-verified. This is explicitly not a validated regulatory engine — responses carry regulatory_claim: false and scope integrity-draft. Treat it as a tamper-evident ledger, not as ISO-17025 / GLP / 21 CFR Part 11 certification.
Self-driving laboratory
The SDL loop is: plan round → generate protocol → ingest result → decide whether to advance. Planning reuses the Active Learning engine (exploit by expected improvement or top value, explore by maximum uncertainty, cold-start with space-filling designs); protocol generation reuses the Opentrons OT-2/Flex generator (API v2.14) with injected safety bounds and AST validation, and only real liquid-volume parameters become dispense actions. Ingestion applies a deterministic physical-plausibility gate per target (conversion/selectivity/yield 0–100%, temperature −50–2000 °C, viscosity strictly positive, pH 0–14, …): a value inside the bounds is recorded as MEASURED and appended to the training set; an absurd value is flagged for human approval and never silently folded into the surrogate. Planned candidates are PREDICTED, and every mutation is written to the audit log. The edge layer registers hardware nodes, streams telemetry, treats a node as offline after 90 s without a heartbeat, validates readings against per-node safety envelopes, auto-trips a software E-stop on critical breach, and requires an authenticated operator with a verified checklist and justification to reset — recorded in a hash-chained sign-off ledger. Action risk is classified LOW/MEDIUM/HIGH, with medium/high actions held behind an approve/reject gate.
Collaboration & review
Research teams share a room rather than an export. Rooms track participants, colours, avatars, live cursors and the active view over server-sent events; a shared candidate spreadsheet can be loaded, locked, tagged and shortlisted, and a synchronized 3D viewer keeps camera and selected atoms in lockstep. Decisions move through an explicit review state machine — DRAFT → IN_REVIEW → APPROVED / REJECTED → SENT_TO_LAB — with idempotent per-user votes (approve / reject / abstain) and quorum auto-approval at two distinct voters. Every sign-off is an immutable SHA-256 hash-chained certificate that can be re-verified end to end, and shared Pareto-filter state and trade-off weights are broadcast to the room so everyone is looking at the same front.
The Pilot layer
Matflow Pilot is deliberately outside the scientific loop but bound by it: it can only call the same capabilities the UI exposes (69 in the live catalog), each carrying a risk tier. Reads, navigation and drafts run freely; writes, deletes, cancellations, real compute spend and billing require an explicit Approve click backed by a two-phase, HMAC-signed, single-use token with a ten-minute TTL that is bound to its user and rejects replay, tampering or cross-user use. Every compute path admits credits first, so a zero balance stops a run before it spends. A post-turn critic checks that the assistant never narrates a completed action it did not actually take — the check runs once, followed by exactly one bounded retry, and then a deterministic honesty note if the claim persists. Drives operate pages visibly (fill the real form, call the real submit, scroll the result into view), and every number the assistant reports keeps its evidence class and engine card.
Visualization workbench
Structures and optimization runs are inspected, not just tabulated. The pluggable 3D viewer speaks four backends — Mol* (trajectories, docking poses, density maps), NGL (cartoon / surface / contacts with a trajectory scrubber), 3Dmol.js (bundled lightweight WebGL) and backend-rendered PyMOL stills for machines without WebGL — with automatic fallback and an evidence badge on every view. The crystal toolkit adds supercell expansion, Miller-slice highlighting with analytic d-spacing, XRD stick overlays and Brillouin-zone paths. The 2D sketcher round-trips SMILES / MOL / RXN with SMARTS highlighting. Optimizer runs open in an Optuna-style dashboard (Pareto front, parameter importance, timeline, parallel coordinates), MLIP quality is read off the leaderboard with hull-error plots, phonons render as bands + DOS with mode animation, and 4D-STEM data opens in a py4DSTEM-lite diffraction viewer (virtual BF/ADF, orientation maps). See the live workbench at /viz, the sketcher at /studio/recipe, screening at /studio/screening, and the MLIP board at /benchmarks.
Limitations, stated plainly
- Small-data caveats. Every stage above is built to cope with tens of rows, but "cope with" is not "perform as well as with thousands." Confidence intervals widen and extrapolation flags fire more often exactly when data is scarcest.
- Synthetic data is not a substitute for real experiments. It preserves statistical structure well enough to train a model, but it cannot manufacture information that was never measured in the first place.
- Extrapolation risk. The ensemble is a strong interpolator within the training envelope and an honest, flagged guesser outside it — the extrapolation detector exists specifically because the model itself cannot tell you when it is guessing. Benchmark methodology publishes both the random-split and target-extrapolative results for this reason.
- Both searches are heuristics. Matflow Evolutionary Search and Matflow Adaptive Search find good fronts, not provably optimal ones; neither guarantees the global Pareto front has been found, only a good approximation of it.
- High-fidelity engines are optional. Real PyBaMM, GROMACS, Vina, phonopy, QM and MLIP runs require the relevant package, binary or cluster. When one is absent the platform returns a labelled analytic fallback or an explicit DEMO with
value: null— never a fabricated number — and the engine card says which path ran. - TEA and LCA are screening-grade. They are order-of-magnitude engineering and emission estimates over a stated boundary, not supplier quotes and not a certified ISO 14040 study. Licensed databases such as ecoinvent are never bundled.
- Extraction and narration are imperfect. Literature extraction publishes its real field recall and routes uncertain records through a human review queue; LLM characterisation narration interprets computed peaks but can misread them. Deterministic numbers remain the ground truth.
- Benchmark numbers are instance- and protocol-specific. The hub renders only what this deployment generated, with the split, seed and evidence class attached; an empty state means no result yet, not a default score. The audit path for any published number is documented at /benchmarks/methodology.
- Lab validation is the final word. Predictions, synthesized data and ranked candidates prioritize which experiment to run next; they do not replace running it.
- Your data stays yours. Models are trained per account and are never merged into a shared model; results are exportable.
References
The validation datasets and laboratory case study are documented in the platform's public benchmarks and evidence reports; the underlying methods used within the platform are drawn from the peer-reviewed work listed below. Links resolve through the DOI system.
- "High-throughput computational screening of cathode materials for Li-O2 battery." Computational Materials Science, 197, 110592 (2021). doi:10.1016/j.commatsci.2021.110592 ↗
- "Visualising multi-dimensional structure/property relationships with machine learning." Journal of Physics: Materials, 2(3), 034003 (2019). doi:10.1088/2515-7639/ab0faa ↗
- "Statistical Analysis and Discovery of Heterogeneous Catalysts Based on Machine Learning from Diverse Published Data." ChemCatChem, 11(18), 4537–4547 (2019). Source of the CADS oxidative-coupling-of-methane benchmark dataset. doi:10.1002/cctc.201900971 ↗
- "On the Quality of Synthetic Generated Tabular Data: Evaluation Metrics and Best Practices." Mathematics, 11(15), 3278 (2023). doi:10.3390/math11153278 ↗
- "PyBaMM: Python Battery Mathematical Modelling." Journal of Open Source Software, 4(40), 1654 (2019). Underlying the continuum battery models where PyBaMM is installed. doi:10.21105/joss.01654 ↗
- "Data-driven prediction of battery cycle life before capacity degradation." Nature Energy, 4, 383–391 (2019). Source of the Severson LFP cycle-life benchmark and its published Table 1 baselines. doi:10.1038/s41560-019-0356-8 ↗