Stage Guides

Evaluation — quality scoring

What it does

Synthetic data is never trusted by assumption. Evaluation stress-tests every generated batch with 20+ metrics, so you know — not hope — that the augmentation matches reality.

The five report families

  • Quality Report — Distribution Match Score for continuous columns, Category Match Score for categorical, plus column-pair correlation similarity.
  • ML Efficacy — an Indistinguishability Score measuring how easily a classifier can tell real rows from synthetic ones. The harder that is, the more faithful the data synthesis.
  • Dimensionality reduction — Linear Map and Structure Map projections of real vs. synthetic rows, so divergence is visible rather than just scored.
  • Privacy Shield — record-level distance to the closest real row, an overfitting guard, and an exact-duplicate check. Relevant the moment synthetic data is shared outside your lab.
  • Anomaly detection — unsupervised, threshold-free detection of the worst-fitting rows, flagged for manual review.

Reading the headline scores

  • Overall Quality Score — the summary: how close the synthetic distribution is to the real one.
  • Correlation Similarity — whether relationships between columns survived generation.
  • Distribution Match Score — per-column distributional accuracy.

What to do with the numbers

High quality + high correlation similarity means the synthetic data is a safe stand-in for the real set in modeling. A low Indistinguishability Score with a weak quality report means the generator struggled — switch modes (try Statistical on small data), reduce the number of rows generated, or fix mislabeled column types before regenerating.

Go from reading to running

The quickstart tutorial walks the first synthesis, evaluation and prediction on a demo dataset — each step matches a real platform page.