Skip to content
OpenAI · ChatGPT · Deep ResearchEpisode 068 · 27 July 2025 · 47:31

The ChatGPT Agent Promises to Act for You—and Immediately Runs Into Control, Marketing, and Safety

What to watch for

1Compare “The ChatGPT Agent Promises to Act for You—and Immediately Runs Into Control, Marketing, and Safety” with “USA, EU, China: three approaches to AI and ideology management”: they provide different criteria for judging the same issue.
2Test the conclusion from “Code and algorithms control: EU/France cases vs xAI/Meta/Google” in your own use case—what actually changes in the process and what remains a promise.
3Before choosing a product or approach, record the constraint identified in “AI at the International Mathematical Olympiad: 'gold' — how is it possible - part 2/2”.
4Define the owner of the outcome and the quality metric for the situation described in “Anthropic: risks of AI abuse and MENA investors”.
Signals to track afterwards
→Watch for actions by Anthropic and Apple that confirm or challenge the episode’s central claims.
→Compare new launches and policy changes with “Code and algorithms control: EU/France cases vs xAI/Meta/Google”: have access, quality, price, or constraints changed?
→Check whether the scenario in “Anthropic: risks of AI abuse and MENA investors” becomes repeatable practice rather than a one-off demonstration.
Most useful for
EntrepreneursAI usersProduct teamsExecutives and managersProcess ownersDevelopers

Key takeaways

00:00An AI agent sounds the next obvious step: instead of telling a person where to click, it opens a site, gathers information, and completes the task itself

The discussion of “The ChatGPT Agent Promises to Act for You—and Immediately Runs Into Control, Marketing, and Safety” yields a practical test: in a presentation, this looks natural. In a test, there are delays, wrong turns, restrictions, and situations where doing the action manually would be easier. The gap between promise and product begins exactly here.

03:00Why an announcement is not enough: the 'AI without ideology' order — what it means

The boundary of the “New law: AI without ideology, what does it mean?” case is defined by this point: the White House put out a plan of dozens of points, including federal contracts only with “ideologically neutral” models; it sounds progressive, but with no definition and no benchmark for bias it reads more like a political slogan.

07:04The ChatGPT agent combines a browser, reasoning, and tools

For the “USA, EU, China: three approaches to AI and ideology” scene, the decisive point is this: the US more often leaves room for companies, Europe demands transparency and restricts practices, and China ties model development tightly to state control — the same “AI” is regulated under three different logics.

08:47Regulators see this market differently

The “Code and algorithms control: EU/France cases vs xAI/Meta/Google” topic becomes clearer once this point is included: the United States more often leaves room for companies, Europe demands transparency and restricts practices, and China connects model development to state control. Disputes around X, Meta, and Google show that this is not only about code safety. Marketing, access to data, and an algorithm’s influence on society are also being controlled.

11:32Who owns the outcome: restrictions on AI marketing

In the context of “Restrictions on AI marketing,” this criterion applies: California wants to require a chatbot to remind users at the start of a session and at least every three hours that it is not human, and to label AI content; there is sense in it — the line between real and generated is already blurring — but the panel split on whether it is freedom or a restriction.

17:25What changes in real work: OpenAI's ChatGPT agent — promised vs delivered in tests

In the context of “OpenAI's ChatGPT agent: promised vs delivered in tests,” this criterion applies: the agent is not one model but a construct that decides on its own when to run Deep Research, reasoning, or connect a service; the first impression is strong, but the real boundary runs along the agent's permissions, the visibility of its actions, its action log, and the ability to stop it.

19:55Why context matters more than one metric: the 'go order by photo' agent case — why it's not daily-use

The boundary of the “The 'go order by photo' agent case: why it's not daily-use” case is defined by this point: the “photograph a dish, go order it” demo has been shown for two years as a fun trick rather than an everyday scenario; a flashy picture is no substitute for the repeatable usefulness that keeps a product in daily use.

22:58How the issue moves from news to product: what ChatGPT agents are

For the “What ChatGPT agents are” scene, the decisive point is this: unlike Deep Research, which only digs into sources, the agent can click and interact with sites and interfaces; but every new capability needs access to accounts, files, and payments, so usefulness grows together with the need for control.

33:06Against the background of everyday failures, model results at the International Mathematical Olympiad are especially impressive

The practical meaning of “AI at the International Mathematical Olympiad: 'gold' — how is it possible” is that “gold” on a formalized problem does not mean a universal agent — an Olympiad gives clear conditions and a verifiable answer, while an internet task is made of ambiguous pages, authentication, and human rules; a model can reason brilliantly and still navigate a real process poorly.

44:45XAI, OpenAI, and Google are now trying to control the market narrative, while Anthropic separately warns about abuse and access risks

The “Anthropic: risks of AI abuse and MENA investors” topic becomes clearer once this point is included: the winner will not be the first company to call its system an agent. It will be the product that shows its actions, allows the process to be stopped, limits authority, and honestly admits when a human should take control.

What this episode is about

OpenAI is demonstrating a browser agent, models are winning gold at the Mathematical Olympiad, and the United States, Europe, and China are choosing different rules. A live test separates autonomy from a polished demo: an agent is useful only when it is clear what it is doing, which data it can see, and who is responsible for an error.

An AI agent sounds like the next obvious step: instead of telling a person where to click, it opens a site, gathers information, and completes the task itself. In a presentation, this looks natural. In a test, there are delays, wrong turns, restrictions, and situations where doing the action manually would be easier. The gap between promise and product begins exactly here.

The ChatGPT agent combines a browser, reasoning, and tools. It can plan a sequence of steps, but every new capability requires access to accounts, files, and payments. An ordinary chat error remains text. An agent’s error can submit a form, change data, or expose information. Utility therefore grows together with the need for control.

Regulators see this market differently. The United States more often leaves room for companies, Europe demands transparency and restricts practices, and China connects model development to state control. Disputes around X, Meta, and Google show that this is not only about code safety. Marketing, access to data, and an algorithm’s influence on society are also being controlled.

Against the background of everyday failures, model results at the International Mathematical Olympiad are especially impressive. But “gold” on a formalized problem does not mean a universal agent. An Olympiad provides clear conditions and a verifiable answer; an internet task consists of ambiguous pages, authentication, and human rules. A model can be outstanding at reasoning and still navigate a real process poorly.

xAI, OpenAI, and Google are now trying to control the market narrative, while Anthropic separately warns about abuse and access risks. The winner will not be the first company to call its system an agent. It will be the product that shows its actions, allows the process to be stopped, limits authority, and honestly admits when a human should take control.

The winning agent will not be the first system to claim autonomy. It will be the product that shows its actions, limits its authority, can be stopped, and clearly hands control back to a person when needed.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 97 segments: 64 identified, 1 mixed, 15 probable, and 17 unresolved.

Read transcript on a separate page

Loading…