Synthetic Data
Synthetic data is training or evaluation data generated by models or programs rather than collected from people. Uses: instruction and reasoning traces from a stronger teacher (distillation), verified solutions for RLVR, simulated environments and conversations for agents, augmentation where real data is private or scarce (data-privacy-in-ai), and test cases for evals.
Risk: model collapse
Training repeatedly on model-generated output can degrade later models, losing the tails of the original distribution (“The Curse of Recursion”, arXiv 2023-05-27; the journal version, Shumailov et al., “AI models collapse when trained on recursively generated data”, Nature 631, 755-759, 2024-07-24). Mitigations: keep real data in the mix, filter and verify generated data, use checkable domains such as code and maths.
Motivation in 2026: high-quality human text is finite, which feeds the scaling-laws debate. See also snorkel-ai for programmatic data labelling and ai-slop for the low-quality public-web side effect.
Related: large-language-model, reasoning-models.
Sources
- https://arxiv.org/abs/2305.17493 (The Curse of Recursion, v1 2023-05-27, re-checked 2026-10-07)
- https://doi.org/10.1038/s41586-024-07566-y (Nature 2024 version, metadata via Crossref API 2026-10-07)
- https://en.wikipedia.org/wiki/Synthetic_data (secondary)