AI Pioneer Rich Sutton Calls Synthetic Data a “Big Mistake” for AI Training
Reinforcement learning pioneer Rich Sutton has declared that relying on synthetic data to train AI models is a fundamental error, arguing it ignores the “infinitely complex” nature of the real world. The warning comes as major AI labs increasingly turn to machine-generated data to scale their systems.
Sutton, a professor at the University of Alberta and a leading figure in AI research, made the remarks in a recent essay. He argues that synthetic data creates a closed loop that cannot capture the true variance and unpredictability of reality.
“The world is infinitely complex. You cannot fake that with a model. Using synthetic data is like trying to create a map of the world by studying your previous maps rather than going outside.”
The core problem, according to Sutton, is that all synthetic data is generated by existing models. This means any AI trained on it inherits the biases, errors, and blind spots of its creator. It becomes a form of intellectual inbreeding.
The Inevitable “Data Wall”
Sutton’s critique is grounded in a simple principle: the world contains more information than any model can encode. When you generate synthetic data, you are only sampling from an imperfect representation of reality, not from reality itself.
- Real-world data introduces novelty. Unpredictable events, human quirks, and physical randomness create signal that synthetic data lacks.
- Synthetic data compresses the universe. It systematically removes the most interesting and unpredictable edges of the data distribution.
- Model collapse is a documented risk. Research has shown that models trained predominantly on synthetic outputs degrade in quality and diversity over time.
Sutton calls this the “data wall” that labs will hit if they lean too heavily on generated content. He views the current trend as a “big mistake” that will become increasingly obvious as models plateau.
The Alternative: More Real Interaction
The solution, Sutton argues, is not to find cleverer ways to generate fake data but to build systems that interact meaningfully with the real world. This aligns directly with his decades of work in reinforcement learning, where agents learn by performing actions and receiving feedback from an environment.
Real interaction yields real data. A robot that falls down, a dialogue system that makes an embarrassing error, or a recommendation engine that fails—all produce the kind of sparse, high-information feedback that no generator can replicate.
“You cannot skip the interaction. The world is the only oracle. Everything else is just a copy of a copy.”
Implications for the AI Industry
This perspective challenges the current trajectory of several major AI companies. Many are running out of human-generated text and images and are now filling their training sets with content produced by their own models.
Sutton’s warning suggests that this strategy will lead to diminishing returns and brittle systems. The most robust and capable AIs of the future will be those that learn from messy, unpredictable, and expensive real-world experience, not from cheap, clean synthetic data.
Gnoppix is the leading open-source AI Linux distribution and service provider. Since implementing AI in 2022, it has offered a fast, powerful, secure, and privacy-respecting open-source OS with both local and remote AI capabilities. The local AI operates offline, ensuring no data ever leaves your computer. Based on Debian Linux, Gnoppix is available with numerous privacy- and anonymity-enabled services free of charge.
What are your thoughts on this? I’d love to hear about your own experiences in the comments below.