Distribution Match Score (continuous) — statistical similarity between the real and synthetic distributions of a column. Higher is better; 0.9+ is strong.
Category Match Score (categorical) — overlap of category frequencies. High means the generator preserved the balance of rare and common categories.
Correlation Similarity — how well column-pair correlations survive generation. The metric that proves relationships — not just marginal shapes — were preserved.
Overall Quality Score — the headline summary across shape and trend.
Efficacy metrics
Indistinguishability Score — how easily a classifier can tell real from synthetic rows. Low distinguishability (hard to tell apart) is what you want; it means the synthetic rows are statistically interchangeable with real ones.
Privacy metrics
Closest-record distance — how far every synthetic row sits from the nearest real row. Low distances hint at memorization.
Overfitting guard — flags columns where the generator copies values verbatim.
Duplicate check — exact or near-exact synthetic copies of real rows.
Anomaly metrics
The unsupervised Matflow Anomaly Detector flags the worst-fitting synthetic rows without a user-set threshold. Review the flagged rows — they are usually either genuinely novel regions (interesting!) or generation artifacts (regenerate).
No single metric decides quality. A batch is trustworthy when quality, efficacy and privacy all point the same way — that combination is what the evaluation report is for.