Trains the Matflow Predictive Ensemble — five model families (tree-based and kernel-based regressors) — to predict material or chemical performance from input conditions. The ensemble exists because no single algorithm is reliably best across the small, heterogeneous datasets this platform sees.
Configuring a run
Dataset — real, synthetic, or a mix. Training on augmented data is a core pattern.
Target column — the outcome to predict (e.g. conversion, selectivity, yield).
Feature columns — everything else, minus any columns you explicitly exclude.
Composition descriptors — optional: for any column holding a chemical formula, 132 descriptors are computed automatically.
Feature engineering
Physics-derived transforms (inverse, log, sqrt, square) applied per feature shape, plus a logit transform for bounded-percentage targets.
Latent components from the transform layer as additional features.
Automated feature selection to control dimensionality.
Explainability
Every trained ensemble comes with Global Impact (summary, dependence, waterfall),Local Impact per prediction, and permutation importance. You see which inputs drive the outcome — not a black box.
Uncertainty
Two signals, always: conformal prediction intervals (90th-percentile out-of-fold residuals from the same cross-validation the metrics use) and an extrapolation flagon every prediction. A number the model is guessing at looks different from one it has real support for. Treat flagged predictions as hypotheses, not results.