Data Preparation

Column roles & common mistakes

What each role means

  • Feature — an input the model learns from (composition, conditions).
  • Target — the outcome to predict (conversion, selectivity, yield).
  • Identifier / metadata — excluded from modeling (sample IDs, dates, notes).
  • Composition — columns that hold parts of a formulation, often constrained to sum to 100%.

How roles are detected

The platform auto-detects column types on upload — categorical columns are recognized by low cardinality, continuous columns by numeric spread. Review the detection on the Data Synthesis and Prediction pages; a mislabeled column type is the most common cause of weak results.

Common mistakes

  • Leaving the experiment ID as a feature — the model memorizes the ID instead of learning chemistry.
  • A target with no variance (a column where every value is identical) — detection or prediction will fail; pick a target that varies across rows.
  • Highly correlated features (e.g. both weight % and mole % of the same metal) — the ensemble copes, but interpretation gets harder.
  • Missing values left as empty strings in a continuous column — the type detector may classify it as categorical.

Go from reading to running

The quickstart tutorial walks the first synthesis, evaluation and prediction on a demo dataset — each step matches a real platform page.