Claude Can Already Control a Computer, While Apple Intelligence Still Gets in the User's Way
Why can Claude already control a computer while Apple Intelligence can still get in the way of a simple user task?
Determine which work can safely be entrusted to Apple and Claude before granting real permissions; the assessment must set permissions, boundaries, stop conditions, and ownership of the outcome before automation begins.
What to watch for
Key takeaways
The “ToTheMoon — a podcast about the world of modern technology” topic becomes clearer once this point is included: the episode is built on a contrast — Anthropic's Claude already clicks buttons in a browser like a person, while Apple Intelligence features show up at the wrong moment, so the assistant race will be decided by the quality of actions and interfaces rather than raw model power.
The discussion of “Why Joe Biden imposed new AI restrictions at the end of his term, right before the U.S. election” yields a practical test: the national-security memorandum with AI restrictions was signed at the very end of the term, and the hosts read it as political timing — pass the unpopular measure now so it does not land on the next candidate.
The discussion of “AI in the defence industry” yields a practical test: the memorandum includes a clause on attracting foreign AI specialists so they do not work for competitors, and private companies are required to give the state access to their models for analysis.
The “Anthropic's Claude learned to control a PC” scene leads to a working conclusion: it does not have to wait for a dedicated API from every service—it repeats the user's path. That makes the agent more universal, but also slower and more dangerous: an error now ends not in incorrect text, but in an incorrect action.
The boundary of the “What for? What are the use cases for AI controlling a PC?” case is defined by this point: the system can gather data from websites, process email, transfer information into a spreadsheet, or perform a routine operation in old enterprise software. But reliability has to be far higher than for an ordinary answer. If the agent clicks the wrong place once, sends an email to the wrong person, or confirms a purchase, the time savings disappear quickly.
The decision in “Apple has started rolling out AI features” depends on one criterion: email-editing functions and suggestions are already built into the system, but they can appear without being requested and fail to explain what they offer; a good AI interface either solves the task at the right moment or stays out of the way.
For the “Voice models: what is the main problem of voice recognition and voice communication with the AI” scene, the decisive point is this: speech recognition is already high quality, yet a conversation with an assistant collapses into clarifications and mistakes that send you back to the screen — voice does not replace the interface yet.
The “Microsoft is putting every model into Copilot: who wins the product race?” scene leads to a working conclusion: for developers and content creators, that may be stronger than a standalone chatbot: the system chooses the appropriate tool and remains inside the workflow. The race will be won not by the company that first taught AI to click a button, but by the one that makes the action predictable, understandable, and safe.
What this episode is about
Anthropic demonstrated a model that opens a browser and clicks buttons like a person. Apple is adding AI directly to email and the operating system, but the functions appear at the wrong time and for unclear reasons. The contrast shows that the future of assistants will be decided not by model power, but by the quality of their actions and interfaces.
Claude Computer Use looks more important than another chatbot because the model gains the ability to act inside an interface: open a browser, find information, fill out a form, and move between windows. It does not have to wait for a dedicated API from every service—it repeats the user's path.
That makes the agent more universal, but also slower and more dangerous: an error now ends not in incorrect text, but in an incorrect action.
The real use cases are obvious. The system can gather data from websites, process email, transfer information into a spreadsheet, or perform a routine operation in old enterprise software.
But reliability has to be far higher than for an ordinary answer. If the agent clicks the wrong place once, sends an email to the wrong person, or confirms a purchase, the time savings disappear quickly.
Apple Intelligence demonstrates the opposite problem. The company has already integrated email-editing functions and suggestions into the system, but they can appear without being requested and fail to explain what they are offering. Users should not have to guess what a new button means. A good AI interface either solves a task at the right moment or stays out of the way.
Voice models still run into the same boundary. High-quality speech recognition is useful, but a conversation with an assistant often collapses into clarifications, mistakes, and the need to look at the screen again. Search inside ChatGPT also has no value by itself when email, cars, or office applications are already receiving their own embedded models.
Microsoft is therefore betting not on one winner, but on Copilot as a layer over different models. For developers and content creators, that may be stronger than a standalone chatbot: the system chooses the appropriate tool and remains inside the workflow.
The race will be won not by the company that first taught AI to click a button, but by the one that makes the action predictable, understandable, and safe.
An agent’s usefulness grows with the radius of possible harm, so access control is part of the product rather than a separate security setting.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 64 segments: 31 identified, 1 mixed, 13 probable, and 19 unresolved.
Loading…