Skip to content
ChatGPT · OpenAI · GeminiEpisode 098 · 22 February 2026 · 53:00

The World of AI Agents Has Already Arrived: A Model Makes Discoveries, Hires People, and Opens New Paths for Deception

Central question

What changes when an AI agent can make discoveries, hire people, and create new methods of deception at the same time?

What you take away

Evaluate agents by both utility and blast radius: who sets the goal, which tools are available, how the result is verified, and who is accountable for deception.

Main threads

What to watch for

1Compare “GPT Codex Spark - superspeed model” with “Claude and millions of context tokens — how the user shapes the result”: they provide different criteria for judging the same issue.
2Test the conclusion from “China stuns with its city automation” in your own use case—what actually changes in the process and what remains a promise.
3Before choosing a product or approach, record the constraint identified in “The main AI-agent releases”.
4Define the owner of the outcome and the quality metric for the situation described in “ChatGPT 5.2 PRO”.
Signals to track afterwards
Watch for actions by Lockdown Mode and Anthropic that confirm or challenge the episode’s central claims.
Compare new launches and policy changes with “China stuns with its city automation”: have access, quality, price, or constraints changed?
Check whether the scenario in “ChatGPT 5.2 PRO” becomes repeatable practice rather than a one-off demonstration.
Most useful for
Executives and managersAI usersEntrepreneursProcess ownersDevelopersProduct teams

Key takeaways

00:00The World of AI Agents Has Already Arrived: A Model Makes Discoveries, Hires People, and Opens New Paths for Deception

The discussion of “The World of AI Agents Has Already Arrived: A Model Makes Discoveries, Hires People, and Opens” yields a practical test: speed grows through models, context, and tools, but the longer the chain of actions, the easier it is to substitute data, redirect an agent, and get a confident but wrong result.

00:48Model speed is not growing because of one magical feature

The decision in “GPT Codex Spark - superspeed model” depends on one criterion: speed grew not because of one magical feature — architecture, inference, context handling, and tool-calling all improved, so an agent does in minutes what recently required separate services and manual steps.

04:02What determines the outcome: heavy thinking and super speed of responses

The practical meaning of “Heavy thinking and super speed of responses” is that heavy reasoning and high speed are useful only when the result is reliable and can be checked, not merely produced fast.

07:11Why an announcement is not enough: what drives the growth in model speed

The boundary of the “What drives the growth in model speed” case is defined by this point: the gain comes from architecture, inference, context handling, and tools working together, not from a single announced feature.

14:30The market tests it through use: ChatGPT 5.2 Pro

The boundary of the “ChatGPT 5.2 PRO” case is defined by this point: heavier reasoning helps with large projects, but volume alone does not guarantee understanding: a model can lose an important detail inside a million tokens just as a person can inside an enormous folder.

16:40ChatGPT 5.2 Pro provides heavier reasoning, while Claude works with up to a million tokens of context

In the context of “Claude and millions of context tokens — how the user shapes the result,” this criterion applies: a million tokens lets you load a whole project, but the user still has to point the model at what matters, because volume by itself is not understanding.

20:14Who owns the outcome: Lockdown Mode and AI safety

The “Lockdown Mode and AI safety” issue should be assessed with one constraint in mind: safety modes matter because the radius of possible harm grows as an agent acts, so access control has to be part of the product rather than a setting added at the end.

22:06What changes in real work: AI makes scientific discoveries

The practical meaning of “AI makes scientific discoveries” is that the model proposes a hypothesis and speeds up testing alternatives, but authorship stays shared: the person poses the question, judges novelty, and answers for the proof.

34:59Agents are already entering the real world and can hire a person for a physical task

In the context of “China stuns with its city automation,” this criterion applies: agents are already entering the real world and can hire a person for a physical task, and once the digital layer connects to cameras, transport, and services, an error stops being local.

50:00Scams and data substitution become natural attacks

The “The main AI-agent releases” topic becomes clearer once this point is included: scams and data substitution become a natural attack: a changed page, email, or instruction can be taken as part of the task, so the key skill is controlling sources and authority, not the launch itself.

What this episode is about

ChatGPT 5.2 Pro, Claude’s million-token context, scientific papers, and agents in the real world show a new scale. Speed grows through models, context, and tools. But the longer the chain becomes, the easier it is to substitute data, redirect an agent, and obtain a confident but wrong result.

Model speed is not growing because of one magical feature. Architecture, inference, use of context, and the ability to call tools all improve. A modern agent can therefore complete in minutes work that recently required separate services and manual transitions.

ChatGPT 5.2 Pro provides heavier reasoning, while Claude works with up to a million tokens of context. For the user, this means the ability to upload a large project or archive. Volume alone does not guarantee understanding, however: a model can lose an important detail inside a million tokens just as a person can inside an enormous folder.

Scientific discoveries made with AI show the strongest use case. A model proposes a hypothesis, helps a researcher test alternatives, and accelerates the path. Authorship remains shared because the person formulates the question, evaluates novelty, and answers for the proof.

Agents are already entering the real world and can hire a person for a physical task. Chinese urban automation demonstrates how quickly the digital layer connects to cameras, transportation, and services. In such an environment, an error stops being local.

Scams and data substitution become natural attacks. If an attacker changes a page, email, or instruction, an agent may accept it as part of the task and act on its own.

Codex, Qwen, and other agent releases are expanding capability faster than users are learning to protect it. The world of agents is already here, but its central skill is not launching them—it is controlling sources and authority.

The world of agents is already here, but the decisive skill is not launching another agent. It is controlling sources, permissions, and the chain of actions that follows.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 85 segments: 56 identified, 2 mixed, 19 probable, and 8 unresolved.

Loading…