Model Collapse in Recursive Synthetic Data Training: Mechanisms, Evaluation, and Mitigation Strategies
Main Article Content
Keywords
model collapse, recursive synthetic data training, synthetic data, distribution shift, data governance
Abstract
Generative models increasingly produce content that can enter later training corpora, creating recursive synthetic-data feedback loops. This review examines when such loops cause model collapse and when synthetic data remain useful. It synthesizes theoreti cal, empirical, and methodological evidence across language, diffusion, evaluation, mitigation, and governance, organized into four dimensions: training regime, degradation mechanism, outcome, and intervention point. Collapse is most consistently associate d with repeated replacement of independent observations with unverified outputs, resulting in tail loss, declining diversity, cumulative error, and bias amplification; recursive diffusion studies also report a shift toward memorization. However, stable or beneficial outcomes occur when independent data or accumulated history are retained and generated samples are corrected, filtered, reweighted, or verified. Findings vary across model families, sampling policies, synthetic ratios, recursive depths, and metr ics, limiting direct comparison and large-scale generalization. Priorities include standardized multigeneration benchmarks, multidimensional longitudinal evaluation, recovery research, realistic multimodel ecosystems, and provenance-aware governance. Overall, model collapse is conditional on the workflow rather than an inevitable consequence of synthetic data.
References
- [1] Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. “Self-Instruct: Aligning Language Models with Self-Generated Instructions.” In Proceedings of ACL 2023, 13484–13508. https://doi.org/10.18653/v1/2023.acl-long.754.
- [2] Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D. Goodman. “STaR: Bootstrapping Reasoning With Reasoning.” In Advances in Neural Information Processing Systems 35 (2022): 15476–15488.
- [3] Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. “MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.” In The Twelfth International Conference on Learning Representations (ICLR), 2024.
- [4] Shekoofeh Azizi, Simon Kornblith, Chitwan Saharia, Mohammad Norouzi, and David J. Fleet. “Synthetic Data from Diffusion Models Improves ImageNet Classification.” Transactions on Machine Learning Research, 2023.
- [5] Zhen Tan, Dawei Li, Song Wang, Alimohammad Beigi, Bohan Jiang, Amrita Bhattacharjee, Mansooreh Karami, Jundong Li, Lu Cheng, and Huan Liu. “Large Language Models for Data Annotation and Synthesis: A Survey.” In Proceedings of EMNLP 2024, 930–957. https://doi.org/10.18653/v1/2024.emnlp-main.54.
- [6] Zhuoyan Li, Hangxiao Zhu, Zhuoran Lu, and Ming Yin. “Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations.” In Proceedings of EMNLP 2023, 10443–10461. https://doi.org/10.18653/v1/2023.emnlp-main.647.
- [7] Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. “AI models collapse when trained on recursively generated data.” Nature 631 (2024): 755–759. https://doi.org/10.1038/s41586-024-07566-y.
- [8] Sina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun, Hossein Babaei, Daniel LeJeune, Ali Siahkoohi, and Richard G. Baraniuk. “Self-Consuming Generative Models Go MAD.” In The Twelfth International Conference on Learning Representations (ICLR), 2024.
- [9] Elvis Dohmatob, Yunzhen Feng, Pu Yang, François Charton, and Julia Kempe. “A Tale of Tails: Model Collapse as a Change of Scaling Laws.” In Proceedings of the 41st International Conference on Machine Learning, PMLR 235 (2024): 11165–11197.
- [10] Elvis Dohmatob, Yunzhen Feng, and Julia Kempe. “Model Collapse Demystified: The Case of Regression.” Advances in Neural Information Processing Systems 37 (2024): 46979–47013. https://doi.org/10.52202/079017-1490.
- [11] Yanzhu Guo, Guokan Shang, Michalis Vazirgiannis, and Chloé Clavel. “The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text.” In Findings of the Association for Computational Linguistics: NAACL 2024, 3589–3604. https://doi.org/10.18653/v1/2024.findings- naacl.228.
- [12] Sierra Wyllie, Ilia Shumailov, and Nicolas Papernot. “Fairness Feedback Loops: Training on Synthetic Data Amplifies Bias.” In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, 2113–2147. https://doi.org/10.1145/3630106.3659029.
- [13] Lianghe Shi, Meng Wu, Huijie Zhang, Zekai Zhang, Molei Tao, and Qing Qu. “A Closer Look at Model Collapse: From a Generalization-to-Memorization Perspective.” Advances in Neural Information Processing Systems 38 (NeurIPS 2025), Main Conference Track, 2025.
- [14] Quentin Bertrand, Avishek (Joey) Bose, Alexandre Duplessis, Marco Jiralerspong, and Gauthier Gidel. “On the Stability of Iterative Retraining of Generative Models on their own Data.” In The Twelfth International Conference on Learning Representations (ICLR), 2024.
- [15] Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Tomasz Korbak, Henry Sleight, Rajashree Agrawal, John Hughes, Dhruv Bhandarkar Pai, Andrey Gromov, Dan Roberts, Diyi Yang, David L. Donoho, and Sanmi Koyejo. “Is Model Collapse Inevitabl e? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data.” In First Conference on Language Modeling (COLM), 2024.
- [16] Nate Gillman, Michael Freeman, Daksh Aggarwal, Chia-Hong Hsu, Calvin Luo, Yonglong Tian, and Chen Sun. “Self-Correcting Self-Consuming Loops for Generative Model Training.” In Proceedings of the 41st International Conference on Machine Learning, PMLR 235 (2024): 15646–15677.
- [17] Yunzhen Feng, Elvis Dohmatob, Pu Yang, François Charton, and Julia Kempe. “Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification.” In The Thirteenth International Conference on Learning Representations (ICLR), 2025.
- [18] George Drayson, Emine Yilmaz, and Vasileios Lampos. “Machine-generated text detection prevents language model collapse.” In Proceedings of EMNLP 2025, 29657–29673. https://doi.org/10.18653/v1/2025.emnlp-main.1506.
- [19] Elvis Dohmatob, Yunzhen Feng, Arjun Subramonian, and Julia Kempe. “Strong Model Collapse.” In The Thirteenth International Conference on Learning Representations (ICLR), 2025.
- [20] Shi Fu, Sen Zhang, Yingjie Wang, Xinmei Tian, and Dacheng Tao. “Towards Theoretical Understandings of Self-Consuming Generative Models.” In Proceedings of the 41st International Conference on Machine Learning, PMLR 235 (2024): 14228–14255.
- [21] Shi Fu, Yingjie Wang, Yuzhu Chen, Xinmei Tian, and Dacheng Tao. “A Theoretical Perspective: How to Prevent Model Collapse in Self-consuming Training Loops.” In The Thirteenth International Conference on Learning Representations (ICLR), 2025.
- [22] Aymane El Firdoussi, Mohamed El Amine Seddik, Soufiane Hayou, Reda Alami, Ahmed Alzubaidi, and Hakim Hacid. “Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory.” In The Thirteenth International Conference on Learning Representations (ICLR), 2025.
- [23] Mohamed El Amine Seddik, Suei-Wen Chen, Soufiane Hayou, Pierre Youssef, and Merouane Abdelkader Debbah. “How bad is training on synthetic data? A statistical analysis of language model collapse.” In First Conference on Language Modeling (COLM), 2024.
- [24] Feiyang Kang, Newsha Ardalani, Michael Kuchnik, Youssef Emad, Mostafa Elhoushi, Shubhabrata Sengupta, Shang-Wen Li, Ramya Raghavendra, Ruoxi Jia, and Carole-Jean Wu. “Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, a nd Pitfalls.” In Proceedings of EMNLP 2025, 10739–10758. https://doi.org/10.18653/v1/2025.emnlp-main.544.
- [25] Ryuichiro Hataya, Han Bao, and Hiromi Arai. “Will Large-scale Generative Models Corrupt Future Datasets?” In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023, 20555–20565. https://doi.org/10.1109/ICCV51070.2023.01879.
- [26] Xiaodan Xing, Fadong Shi, Jiahao Huang, Yinzhe Wu, Yang Nan, Sheng Zhang, Yingying Fang, Michael Roberts, Carola-Bibiane Schönlieb, Javier Del Ser, and Guang Yang. “On the caveats of AI autophagy.” Nature Machine Intelligence 7 (2025): 172–180. https://doi.org/10.1038/s42256-025-00984-1.
- [27] Boris Van Breugel, Zhaozhi Qian, and Mihaela Van Der Schaar. “Synthetic Data, Real Errors: How (Not) to Publish and Use Synthetic Data.” In Proceedings of the 40th International Conference on Machine Learning, PMLR 202 (2023): 34793–34808.
- [28] Ahmed M. Alaa, Boris van Breugel, Evgeny Saveliev, and Mihaela van der Schaar. “How Faithful is your Synthetic Data? Sample-level Metrics for Evaluating and Auditing Generative Models.” In Proceedings of the 39th International Conference on Machine Learning, PMLR 162 (2022): 290–306.
- [29] Mehdi S. M. Sajjadi, Olivier Bachem, Mario Lucic, Olivier Bousquet, and Sylvain Gelly. “Assessing Generative Models via Precision and Recall.” Advances in Neural Information Processing Systems 31 (2018): 5228–5237.
- [30] Vishakh Padmakumar and He He. “Does Writing with Language Models Reduce Content Diversity?” In The Twelfth International Conference on Learning Representations (ICLR), 2024.
- [31] Muhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi, and Jaejun Yoo. “Reliable Fidelity and Diversity Metrics for Generative Models.” In Proceedings of the 37th International Conference on Machine Learning, PMLR 119 (2020): 7176–7185.
- [32] Emily M. Bender and Batya Friedman. “Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science.” Transactions of the Association for Computational Linguistics 6 (2018): 587–604. https://doi.org/10.1162/tacl_a_00041.
- [33] Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. “Model Cards for Model Reporting.” In Proceedings of the Conference on Fairness, Accountability, and Trans parency, 2019, 220–229. https://doi.org/10.1145/3287560.3287596.
- [34] Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. “Datasheets for Datasets.” Communications of the ACM 64, no. 12 (2021): 86–92. https://doi.org/10.1145/3458723.
- [35] John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. “A Watermark for Large Language Models.” In Proceedings of the 40th International Conference on Machine Learning, PMLR 202 (2023): 17061–17084.
- [36] Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, and Chelsea Finn. “DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature.” In Proceedings of the 40th International Conference on Machine Learning, PMLR 202 (2023): 24950–24962.
- [37] Sumanth Dathathri, Abigail See, Sumedh Ghaisas, Po-Sen Huang, Rob McAdam, Johannes Welbl, Vandana Bachani, Alex Kaskasoli, Robert Stanforth, Tatiana Matejovicova, Jamie Hayes, Nidhi Vyas, Majd Al Merey, Jonah Brown-Cohen, Rudy Bunel, Borja Balle, Taylan Cemgil, Zahra Ahmed, Kitty Stacpoole, Ilia Shumailov, Ciprian Baetu, Sven Gowal, Demis Hassabis, and Pushmeet Kohli. “Scalable watermarking for identifying large language model outputs.” Nature 634 (2024): 818–823. https://doi.org/10.1038/s41586-024-08025-4.
- [38] Mauro Giuffrè and Dennis L. Shung. “Harnessing the power of synthetic data in healthcare: innovation, application, and privacy.” npj Digital Medicine 6 (2023): 186. https://doi.org/10.1038/s41746-023-00927- 3.
- [39] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. “GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium.” In Advances in Neural Information Processing Systems 30, 2017, 6626–6637.
- [40] Tuomas Kynkäänniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila. “Improved Precision and Recall Metric for Assessing Generative Models.” Advances in Neural Information Processing Systems 32 (2019): 3927–3936.
- [41] Theresa Stadler, Bristena Oprisanu, and Carmela Troncoso. “Synthetic Data—Anonymisation Groundhog Day.” In 31st USENIX Security Symposium, 2022, 1451–1468.
- [42] Joshua Snoke, Gillian M. Raab, Beata Nowok, Chris Dibben, and Aleksandra Slavković. “General and Specific Utility Measures for Synthetic Data.” Journal of the Royal Statistical Society: Series A 181, no. 3 (2018): 663–688. https://doi.org/10.1111/rssa.12358.
