Skip to content
Artificial intelligence · OpenAI · GoogleEpisode 041 · 19 January 2025 · 41:57

The Data Is Running Out, and AI Has to Learn From Itself

Central question

What happens when high-quality data runs out and models begin learning from their own outputs?

What you take away

Check whether Synthetic data and OpenAI develop a skill or merely hide a lack of understanding; the assessment must check whether the model supports learning or merely takes over the decision and assessment.

Main threads

What to watch for

1Compare “The first stage of large-model development followed a simple logic” with “How important new data is for creating AGI”: they provide different criteria for judging the same issue.
2Test the conclusion from “Planning travel with AI” in your own use case—what actually changes in the process and what remains a promise.
3Before choosing a product or approach, record the constraint identified in “Autonomous NPCs from NVIDIA”.
4Define the owner of the outcome and the quality metric for the situation described in “Use cases for AI in medicine”.
Signals to track afterwards
→Watch for actions by Google and OpenAI that confirm or challenge the episode’s central claims.
→Compare new launches and policy changes with “Planning travel with AI”: have access, quality, price, or constraints changed?
→Check whether the scenario in “Use cases for AI in medicine” becomes repeatable practice rather than a one-off demonstration.
Most useful for
AI usersCompaniesSecurity specialistsExecutives and managersProfessionalsPeople planning their careers

Key takeaways

00:00The first stage of large-model development followed a simple logic: collect as much human text, imagery, and code as possible

The practical meaning of “The first stage of large-model development followed a simple logic” is that the logic of collecting as much human text, imagery, and code as possible has hit its limit: books, websites, and open archives are already used, while new high-quality data appears more slowly than compute grows.

02:19The practical meaning of the issue: synthetic data and their impact on AI and

The “Synthetic data and their impact on AI and the future” scene leads to a working conclusion: a strong model can devise a problem, verify the solution, and create millions of training examples, but when it reproduces its own mistake the next system accepts it as truth — and synthetic data turns into a closed loop.

08:25This is not necessarily a dead end

The discussion of “How important new data is for creating AGI” yields a practical test: for AGI as a self-improving system what matters is not volume but specific data about how models are trained; whatever an independent system creates is synthetic by definition — the task is to make it high quality, while today's synthetics still hallucinate.

11:58What determines the outcome: autonomous NPCs from NVIDIA

The decision in “Autonomous NPCs from NVIDIA” depends on one criterion: at CES 2025 NVIDIA showed NPCs that no longer follow a script but develop on their own — leave the game for a day and the character has lived a life of its own; the open question is what data all of this learns from.

13:48Why an announcement is not enough: use cases for AI in medicine

The “Use cases for AI in medicine” topic becomes clearer once this point is included: from dental scans the model gave the host more detail than three doctors, and a German breast-cancer study of 460,000 patients showed diagnostic accuracy improving by almost twenty percent — AI's value appears in tandem with the physician.

15:43The market tests it through use: the release of an autonomous AI agent from OpenAI

In the context of “The release of an autonomous AI agent from OpenAI,” this criterion applies: Operator is an autonomous agent for complex tasks such as writing code and booking trips with minimal human involvement; expectations for the first quarter are low, but by year's end agents will start visibly changing workflows — all three major players are betting on it.

19:18That is why Ilya Sutskever and other researchers are looking for new fundamental approaches, while DeepSeek shows that progress may come from more than simply adding still more data

The working conclusion from “Planning travel with AI” is that the agent's promise quickly runs into Kayak, Aviasales, and Airbnb — if a specialized service searches better and more transparently, a general-purpose chat creates no value merely because it can talk; the main risk is getting used to one intermediary as the only window.

What this episode is about

Large models have already consumed books, websites, and open archives. The next stage is being built on synthetic data—examples created by other models. That can accelerate progress, but it can also trap the system inside its own mistakes. Against this backdrop, the market is testing agents, medical use cases, and travel tools, where quality can be compared directly with conventional services.

The first stage of large-model development followed a simple logic: collect as much human text, imagery, and code as possible. But that resource is not infinite.

Books, websites, and open archives have already been used, while new high-quality data is being created more slowly than compute is growing. Companies are therefore training models more and more often on synthetic examples—data generated by AI itself.

This is not necessarily a dead end. A strong model can devise a difficult problem, verify the solution, and create millions of useful training examples. The problem begins when it reproduces its own mistake and the next system accepts that mistake as truth. If the quality of the original generator is not controlled, synthetic data quickly becomes a closed loop.

That is why Ilya Sutskever and other researchers are looking for new fundamental approaches, while DeepSeek shows that progress may come from more than simply adding still more data. Restrictions in the United States may slow local companies, but they will not stop developers in countries with different rules. The race is becoming regulatory as well as scientific.

Practical products reveal the real boundary. In medicine, AI can improve diagnostic accuracy when it works alongside a physician and uses verified data. In travel, the promise of an agent quickly runs into Kayak, Aviasales, or Airbnb: if a specialized service searches better and more transparently, a general-purpose chat does not create value merely because it can hold a conversation.

The main risk is becoming accustomed to a single intermediary. Just as Amazon trains people not to compare prices, an AI agent can become the only window through which someone plans trips, buys services, and receives recommendations. Data quality is therefore not an abstract research problem. It determines which options the system shows at all—and which decisions the user eventually stops checking.

Data quality determines which options a system can surface and which decisions users may stop checking. Synthetic training data therefore changes not only model performance, but also the range of reality the model can represent.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 87 segments: 53 identified, 4 mixed, 19 probable, and 11 unresolved.

Read transcript on a separate page

Loading…