The Matflow Predictive Ensemble votes across five members — three gradient-boosted tree models (XGBoost, LightGBM, CatBoost), a random forest, and a partial least squares (PLS) linear model — with an optional ridge member added where it improves the held-out split (for example on extrapolative splits, where ridge outperformed the default). Each member is pinned to single-threaded execution internally; letting any one multithread while the ensemble cross-validates causes severe CPU oversubscription (a two-minute training run can become twenty).
Feature transforms
Physics-derived transforms (inverse, log, sqrt, square) per feature shape.
Logit transform for bounded-percentage targets.
Latent components from the transform layer as additional features.
Automated feature selection for dimensionality control.
Optional composition descriptors (132 columns) from chemical formulas.
How uncertainty is estimated
Conformal prediction intervals come from the 90th-percentile out-of-fold residual of the same cross-validation the metrics already compute — so honest intervals cost no extra training time. The extrapolation flag comes from an unsupervised anomaly detector on the fitted feature space: inputs outside the training envelope are flagged, because the model cannot tell you itself when it is guessing.