Stage Guides

Prediction — predictive modeling

What it does

Trains the Matflow Predictive Ensemble — five model families (tree-based and kernel-based regressors) — to predict material or chemical performance from input conditions. The ensemble exists because no single algorithm is reliably best across the small, heterogeneous datasets this platform sees.

Configuring a run

  • Dataset — real, synthetic, or a mix. Training on augmented data is a core pattern.
  • Target column — the outcome to predict (e.g. conversion, selectivity, yield).
  • Feature columns — everything else, minus any columns you explicitly exclude.
  • Composition descriptors — optional: for any column holding a chemical formula, 132 descriptors are computed automatically.

Feature engineering

  • Physics-derived transforms (inverse, log, sqrt, square) applied per feature shape, plus a logit transform for bounded-percentage targets.
  • Latent components from the transform layer as additional features.
  • Automated feature selection to control dimensionality.

Explainability

Every trained ensemble comes with Global Impact (summary, dependence, waterfall),Local Impact per prediction, and permutation importance. You see which inputs drive the outcome — not a black box.

Uncertainty

Two signals, always: conformal prediction intervals (90th-percentile out-of-fold residuals from the same cross-validation the metrics use) and an extrapolation flagon every prediction. A number the model is guessing at looks different from one it has real support for. Treat flagged predictions as hypotheses, not results.

Go from reading to running

The quickstart tutorial walks the first synthesis, evaluation and prediction on a demo dataset — each step matches a real platform page.