Skip to content
Apple · Anthropic · OpenAIEpisode 030 · 3 November 2024 · 41:22

Claude Can Already Control a Computer, While Apple Intelligence Still Gets in the User's Way

Central question

Why can Claude already control a computer while Apple Intelligence can still get in the way of a simple user task?

What you take away

Determine which work can safely be entrusted to Apple and Claude before granting real permissions; the assessment must set permissions, boundaries, stop conditions, and ownership of the outcome before automation begins.

Main threads

What to watch for

1Compare “Anthropic's Claude learned to control a PC” with “What for? What are the use cases for AI controlling a PC?”: they provide different criteria for judging the same issue.
2Test the conclusion from “Apple has started rolling out AI features” in your own use case—what actually changes in the process and what remains a promise.
3Before choosing a product or approach, record the constraint identified in “Microsoft is putting every model into Copilot: who wins the product race?”.
4Define the owner of the outcome and the quality metric for the situation described in “What for? What are the use cases for AI controlling a PC?”.
Signals to track afterwards
Watch for actions by Anthropic and Apple that confirm or challenge the episode’s central claims.
Compare new launches and policy changes with “Apple has started rolling out AI features”: have access, quality, price, or constraints changed?
Check whether the scenario in “What for? What are the use cases for AI controlling a PC?” becomes repeatable practice rather than a one-off demonstration.
Most useful for
EntrepreneursAI usersProduct teamsExecutives and managersInvestorsStrategy teams

Key takeaways

00:00The market tests it through use: ToTheMoon — a podcast about the world of

The “ToTheMoon — a podcast about the world of modern technology” topic becomes clearer once this point is included: the episode is built on a contrast — Anthropic's Claude already clicks buttons in a browser like a person, while Apple Intelligence features show up at the wrong moment, so the assistant race will be decided by the quality of actions and interfaces rather than raw model power.

01:35The boundary between value and constraint: why Joe Biden imposed new AI restrictions

The discussion of “Why Joe Biden imposed new AI restrictions at the end of his term, right before the U.S. election” yields a practical test: the national-security memorandum with AI restrictions was signed at the very end of the term, and the hosts read it as political timing — pass the unpopular measure now so it does not land on the next candidate.

05:15Who owns the outcome: aI in the defence industry

The discussion of “AI in the defence industry” yields a practical test: the memorandum includes a clause on attracting foreign AI specialists so they do not work for competitors, and private companies are required to give the state access to their models for analysis.

08:13Claude Computer Use looks more important than another chatbot because the model gains the ability to act inside an interface: open a browser, find information, fill out a form, and move between windows

The “Anthropic's Claude learned to control a PC” scene leads to a working conclusion: it does not have to wait for a dedicated API from every service—it repeats the user's path. That makes the agent more universal, but also slower and more dangerous: an error now ends not in incorrect text, but in an incorrect action.

14:07The real use cases are obvious

The boundary of the “What for? What are the use cases for AI controlling a PC?” case is defined by this point: the system can gather data from websites, process email, transfer information into a spreadsheet, or perform a routine operation in old enterprise software. But reliability has to be far higher than for an ordinary answer. If the agent clicks the wrong place once, sends an email to the wrong person, or confirms a purchase, the time savings disappear quickly.

18:13Apple Intelligence demonstrates the opposite problem

The decision in “Apple has started rolling out AI features” depends on one criterion: email-editing functions and suggestions are already built into the system, but they can appear without being requested and fail to explain what they offer; a good AI interface either solves the task at the right moment or stays out of the way.

19:34The practical meaning of the issue: voice models: what is the main problem of

For the “Voice models: what is the main problem of voice recognition and voice communication with the AI” scene, the decisive point is this: speech recognition is already high quality, yet a conversation with an assistant collapses into clarifications and mistakes that send you back to the screen — voice does not replace the interface yet.

39:14Microsoft is therefore betting not on one winner, but on Copilot as a layer over different models

The “Microsoft is putting every model into Copilot: who wins the product race?” scene leads to a working conclusion: for developers and content creators, that may be stronger than a standalone chatbot: the system chooses the appropriate tool and remains inside the workflow. The race will be won not by the company that first taught AI to click a button, but by the one that makes the action predictable, understandable, and safe.

What this episode is about

Anthropic demonstrated a model that opens a browser and clicks buttons like a person. Apple is adding AI directly to email and the operating system, but the functions appear at the wrong time and for unclear reasons. The contrast shows that the future of assistants will be decided not by model power, but by the quality of their actions and interfaces.

Claude Computer Use looks more important than another chatbot because the model gains the ability to act inside an interface: open a browser, find information, fill out a form, and move between windows. It does not have to wait for a dedicated API from every service—it repeats the user's path.

That makes the agent more universal, but also slower and more dangerous: an error now ends not in incorrect text, but in an incorrect action.

The real use cases are obvious. The system can gather data from websites, process email, transfer information into a spreadsheet, or perform a routine operation in old enterprise software.

But reliability has to be far higher than for an ordinary answer. If the agent clicks the wrong place once, sends an email to the wrong person, or confirms a purchase, the time savings disappear quickly.

Apple Intelligence demonstrates the opposite problem. The company has already integrated email-editing functions and suggestions into the system, but they can appear without being requested and fail to explain what they are offering. Users should not have to guess what a new button means. A good AI interface either solves a task at the right moment or stays out of the way.

Voice models still run into the same boundary. High-quality speech recognition is useful, but a conversation with an assistant often collapses into clarifications, mistakes, and the need to look at the screen again. Search inside ChatGPT also has no value by itself when email, cars, or office applications are already receiving their own embedded models.

Microsoft is therefore betting not on one winner, but on Copilot as a layer over different models. For developers and content creators, that may be stronger than a standalone chatbot: the system chooses the appropriate tool and remains inside the workflow.

The race will be won not by the company that first taught AI to click a button, but by the one that makes the action predictable, understandable, and safe.

An agent’s usefulness grows with the radius of possible harm, so access control is part of the product rather than a separate security setting.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 64 segments: 31 identified, 1 mixed, 13 probable, and 19 unresolved.

Loading…