Skip to content
OpenAI · Microsoft · GoogleEpisode 034 · 1 December 2024 · 46:19

AI Agents Promise to Replace Employees, but They Still Cannot Replace a Bad Interface

Central question

Why do AI agents promise to replace employees while still being unable to compensate for a bad interface and an unclear process?

What you take away

Determine which work can safely be entrusted to AI agents and Salesforce before granting real permissions; the assessment must set permissions, boundaries, stop conditions, and ownership of the outcome before automation begins.

Main threads

What to watch for

1Compare “What AI agents are?” with “Start-ups with I. agents lazy, part 2/2”: they provide different criteria for judging the same issue.
2Test the conclusion from “Call the man. Why are the current AI agents so stupid” in your own use case—what actually changes in the process and what remains a promise.
3Before choosing a product or approach, record the constraint identified in “Didn't all the AI studs fail?”.
4Define the owner of the outcome and the quality metric for the situation described in “What would the new level AI agent look, part 1/3”.
Signals to track afterwards
Watch for actions by Amazon and Anthropic that confirm or challenge the episode’s central claims.
Compare new launches and policy changes with “Call the man. Why are the current AI agents so stupid”: have access, quality, price, or constraints changed?
Check whether the scenario in “What would the new level AI agent look, part 1/3” becomes repeatable practice rather than a one-off demonstration.
Most useful for
EntrepreneursInvestorsDevelopersCompany leadersProcess ownersStrategy teams

Key takeaways

01:20What changes in real work: what is the AI agents? When will I

The “What is the AI agents? When will I replace people?” topic becomes clearer once this point is included: the practical boundary is defined by the agent’s permissions, the visibility of its actions, its action log, and the ability to stop execution.

04:38The word ‘agent’ has become a convenient way to sell the future

In the context of “What AI agents are?,” this criterion applies: the practical boundary is defined by the agent’s permissions, the visibility of its actions, its action log, and the ability to stop execution.

07:31How the issue moves from news to product: what kind of life-saving instruments do you use

The decision in “What kind of life-saving instruments do you use?” depends on one criterion: the important signal is not one funding number: the next round, available runway, and closure rate show whether a company can survive the new cost of capital.

10:31Large companies already offer agent builders, but businesses face a simple question: why pay a separate startup for a complicated wrapper if OpenAI, Anthropic, Microsoft, or Google will add a similar feature to familiar…

In the context of “Start-ups with I. agents lazy, part 2/2,” this criterion applies: the practical boundary is defined by the agent’s permissions, the visibility of its actions, its action log, and the ability to stop execution.

12:57Where the promise meets reality: what would the new level AI agent look

In the context of “What would the new level AI agent look, part 1/3,” this criterion applies: the practical boundary is defined by the agent’s permissions, the visibility of its actions, its action log, and the ability to stop execution.

19:32What determines the outcome: why did all the AI studs fail

The boundary of the “Why did all the AI studs fail?” case is defined by this point: the important signal is not one funding number: the next round, available runway, and closure rate show whether a company can survive the new cost of capital.

22:07Why an announcement is not enough: will AIT.T.A. replace

The practical meaning of “Will AIT.T.A. replace?” is that the important signal is not one funding number: the next round, available runway, and closure rate show whether a company can survive the new cost of capital.

33:32Users still often want to press ‘talk to a person.’ An agent confuses cities, fails to understand local names, gets stuck on an exception, and cannot explain why it made a choice

The working conclusion from “Call the man. Why are the current AI agents so stupid” is that ’ An agent confuses cities, fails to understand local names, gets stuck on an exception, and cannot explain why it made a choice. In enterprise work this is critical: one unusual case can cost more than all of the promised savings.

34:38This does not mean every AI startup is doomed

The boundary of the “Didn't all the AI studs fail?” case is defined by this point: this section clarifies the mechanism behind the topic and preserves a constraint that would otherwise be lost in an overly simple conclusion.

What this episode is about

Salesforce, Microsoft, IBM, OpenAI, and dozens of startups sell enterprise agents, yet mass adoption is almost nonexistent. The reason is not a shortage of presentations: an agent must connect many systems, understand business context, and act reliably in places where an ordinary chat interface already makes regular mistakes.

The word ‘agent’ has become a convenient way to sell the future. If a chatbot answers a question, an agent is supposed to open systems, gather data, make a decision, and perform the work on its own. But the models we criticize for poor answers do not become more reliable simply because they have been given access to buttons and corporate data.

Large companies already offer agent builders, but businesses face a simple question: why pay a separate startup for a complicated wrapper if OpenAI, Anthropic, Microsoft, or Google will add a similar feature to familiar software? A young company has to glue together several models, integrations, and databases—and then take responsibility for errors across the entire chain.

A real agent begins not with autonomy, but with context. An interior designer does not need a bot that generates a beautiful room. They need a system that can ask questions, find actual products, account for dimensions, budget, and the client's taste, and return several coherent alternatives. The market still rarely does even that hybrid process—search, Pinterest, generation, and dialogue—well.

Users still often want to press ‘talk to a person.’ An agent confuses cities, fails to understand local names, gets stuck on an exception, and cannot explain why it made a choice. In enterprise work this is critical: one unusual case can cost more than all of the promised savings.

This does not mean every AI startup is doomed. Glean proved the value of a new kind of enterprise search, and Perplexity made search easier to distribute.

The winners solve a narrow pain better than the platform and build data, distribution, or habit before the platform catches up. But selling a ‘universal digital employee’ before it can perform one job consistently is a direct path to disappointment.

The winners solve a narrow pain better than the platform and build data, distribution, or habit before the platform catches up. As a result, but selling a ‘universal digital employee’ before it can perform one job consistently is a direct path to disappointment.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 52 segments: 39 identified, 4 mixed, 7 marked with ✓, and 2 unresolved.

Loading…