Glossary
Glossary of terms
Core concepts
- Job — one execution of one stage (data synthesis, evaluation, prediction, optimization) with its own progress, status and results.
- Pipeline — multiple stages chained so each stage's output becomes the next stage's input.
- Synthetic data — rows generated by a model to statistically resemble your real experiments without copying them.
- Ensemble — a model made of several member models whose predictions are combined; more robust than any single member.
Machine learning
- GAN — generative adversarial network: a generator competes with a discriminator until the generated rows are indistinguishable from real ones.
- Copula — a statistical structure that models the dependence between columns separately from their individual distributions.
- Cross-validation — repeatedly training on subsets and testing on the rest, so reported metrics aren't a lucky split.
- Conformal interval — a prediction range with a coverage guarantee, computed without distributional assumptions.
- Extrapolation — predicting outside the region the training data covers; flagged because it is inherently less reliable.
- Recursive Feature Elimination — iteratively removing the least informative features to keep models small and interpretable.
Optimization
- Pareto front — the set of candidates where improving one objective worsens another; the trade-off surface of the search.
- TOPSIS — a ranking method scoring each candidate by distance to the ideal and anti-ideal solutions.
- Multi-objective genetic search — an evolution-based algorithm that maintains population diversity while converging on the front.
- Compositional constraint — a requirement that a set of columns sums to a fixed value (typically 100%).
Economics
- TEA — techno-economic analysis: costing a process or product including capital, operating and raw-material expenses.
- CapEx / OpEx — capital expenditure (equipment, plant) and operating expenditure (utilities, labor, materials).
- Six-tenths rule — equipment cost scaling: cost ∝ (size/base)^0.6, the standard chemical-industry scale-up approximation.