The Data Is Running Out, and AI Has to Learn From Itself
What happens when high-quality data runs out and models begin learning from their own outputs?
Check whether Synthetic data and OpenAI develop a skill or merely hide a lack of understanding; the assessment must check whether the model supports learning or merely takes over the decision and assessment.
What to watch for
Key takeaways
The practical meaning of “The first stage of large-model development followed a simple logic” is that the logic of collecting as much human text, imagery, and code as possible has hit its limit: books, websites, and open archives are already used, while new high-quality data appears more slowly than compute grows.
The “Synthetic data and their impact on AI and the future” scene leads to a working conclusion: a strong model can devise a problem, verify the solution, and create millions of training examples, but when it reproduces its own mistake the next system accepts it as truth — and synthetic data turns into a closed loop.
The discussion of “How important new data is for creating AGI” yields a practical test: for AGI as a self-improving system what matters is not volume but specific data about how models are trained; whatever an independent system creates is synthetic by definition — the task is to make it high quality, while today's synthetics still hallucinate.
The decision in “Autonomous NPCs from NVIDIA” depends on one criterion: at CES 2025 NVIDIA showed NPCs that no longer follow a script but develop on their own — leave the game for a day and the character has lived a life of its own; the open question is what data all of this learns from.
The “Use cases for AI in medicine” topic becomes clearer once this point is included: from dental scans the model gave the host more detail than three doctors, and a German breast-cancer study of 460,000 patients showed diagnostic accuracy improving by almost twenty percent — AI's value appears in tandem with the physician.
In the context of “The release of an autonomous AI agent from OpenAI,” this criterion applies: Operator is an autonomous agent for complex tasks such as writing code and booking trips with minimal human involvement; expectations for the first quarter are low, but by year's end agents will start visibly changing workflows — all three major players are betting on it.
The working conclusion from “Planning travel with AI” is that the agent's promise quickly runs into Kayak, Aviasales, and Airbnb — if a specialized service searches better and more transparently, a general-purpose chat creates no value merely because it can talk; the main risk is getting used to one intermediary as the only window.
What this episode is about
Large models have already consumed books, websites, and open archives. The next stage is being built on synthetic data—examples created by other models. That can accelerate progress, but it can also trap the system inside its own mistakes. Against this backdrop, the market is testing agents, medical use cases, and travel tools, where quality can be compared directly with conventional services.
The first stage of large-model development followed a simple logic: collect as much human text, imagery, and code as possible. But that resource is not infinite.
Books, websites, and open archives have already been used, while new high-quality data is being created more slowly than compute is growing. Companies are therefore training models more and more often on synthetic examples—data generated by AI itself.
This is not necessarily a dead end. A strong model can devise a difficult problem, verify the solution, and create millions of useful training examples. The problem begins when it reproduces its own mistake and the next system accepts that mistake as truth. If the quality of the original generator is not controlled, synthetic data quickly becomes a closed loop.
That is why Ilya Sutskever and other researchers are looking for new fundamental approaches, while DeepSeek shows that progress may come from more than simply adding still more data. Restrictions in the United States may slow local companies, but they will not stop developers in countries with different rules. The race is becoming regulatory as well as scientific.
Practical products reveal the real boundary. In medicine, AI can improve diagnostic accuracy when it works alongside a physician and uses verified data. In travel, the promise of an agent quickly runs into Kayak, Aviasales, or Airbnb: if a specialized service searches better and more transparently, a general-purpose chat does not create value merely because it can hold a conversation.
The main risk is becoming accustomed to a single intermediary. Just as Amazon trains people not to compare prices, an AI agent can become the only window through which someone plans trips, buys services, and receives recommendations. Data quality is therefore not an abstract research problem. It determines which options the system shows at all—and which decisions the user eventually stops checking.
Data quality determines which options a system can surface and which decisions users may stop checking. Synthetic training data therefore changes not only model performance, but also the range of reality the model can represent.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 87 segments: 53 identified, 4 mixed, 19 probable, and 11 unresolved.
Read transcript on a separate page
Loading…