Stage Guides

Data Synthesis — synthetic data generation

What it does

Turns a handful of experimental runs into a statistically faithful dataset. The engine learns the joint distribution of your real experiments — not just per-column shapes, but the non-linear correlations between conditions and outcomes.

Choosing a mode

  • Matflow Neural — conditional-GAN-style mode for mixed continuous/categorical columns and multi-modal distributions. Strongest default for larger, more complex datasets.
  • Matflow Hybrid — GAN-style network combined with a copula transform; often more stable on small samples.
  • Matflow Statistical — fits marginals and a correlation structure directly. Fastest and the most defensible when rows number in the tens.

Inputs and settings

  • Dataset — your uploaded or demo dataset.
  • Rows to generate — how many synthetic rows you want (within tier limits).
  • Column types — verified automatically; override if the detector mislabeled a column.
  • Seed — set a fixed seed for reproducible runs.

Range clipping

Generated rows are clipped to the physically observed range of each column — a generator can otherwise extrapolate a percentage column past 0–100%. The clipping is applied after generation; constraints enforced natively are on the roadmap.

After generation

A synthetic dataset appears in your Datasets list, marked with source "synthetic". Run Evaluation on it before using it for modeling — that is the step that decides whether the augmentation is trustworthy.

Go from reading to running

The quickstart tutorial walks the first synthesis, evaluation and prediction on a demo dataset — each step matches a real platform page.