Synthetic Data

Synthetic data is training or evaluation data generated by models or programs rather than collected from people. Uses: instruction and reasoning traces from a stronger teacher (distillation), verified solutions for RLVR, simulated environments and conversations for agents, augmentation where real data is private or scarce (data-privacy-in-ai), and test cases for evals.

Risk: model collapse

Training repeatedly on model-generated output can degrade later models, losing the tails of the original distribution (“The Curse of Recursion”, arXiv 2023-05-27; the journal version, Shumailov et al., “AI models collapse when trained on recursively generated data”, Nature 631, 755-759, 2024-07-24). Mitigations: keep real data in the mix, filter and verify generated data, use checkable domains such as code and maths.

Motivation in 2026: high-quality human text is finite, which feeds the scaling-laws debate. See also snorkel-ai for programmatic data labelling and ai-slop for the low-quality public-web side effect.

Related: large-language-model, reasoning-models.

Sources