AI Agents Promise to Replace Employees, but They Still Cannot Replace a Bad Interface
Why do AI agents promise to replace employees while still being unable to compensate for a bad interface and an unclear process?
Determine which work can safely be entrusted to AI agents and Salesforce before granting real permissions; the assessment must set permissions, boundaries, stop conditions, and ownership of the outcome before automation begins.
What to watch for
Key takeaways
The “What are AI agents? When will AI replace people?” topic becomes clearer once this point is included: if a chatbot answers a question, an agent must open systems, gather data, make a decision, and do the work itself — but the models we criticize for poor answers do not become more reliable just because they were given access to buttons and corporate data.
In the context of “What kinds of AI agents are there?,” this criterion applies: a wave of startups builds support and sales agents, but the economics do not add up — there is no reason to pay another twenty dollars per employee on top of Salesforce, such a wrapper is easy to assemble yourself, and these systems still work like primitive chats.
The decision in “What tools do you use to simplify your life?” depends on one criterion: the hosts ask viewers whether they have seen systems that genuinely replace an employee rather than work as one-offs — write your examples in the comments, and the hosts promise to review them in future episodes.
In the context of “How AI-agent startups screw up — part 2/2,” this criterion applies: a young company has to glue together several models, integrations, and databases — and then take responsibility for errors across the entire chain.
In the context of “What a next-level AI agent will look like — part 1/3,” this criterion applies: a real agent begins not with autonomy but with context — an interior designer needs a system that asks questions, finds actual products, accounts for dimensions, budget, and the client's taste; the market still rarely does even that hybrid of search, Pinterest, generation, and dialogue well.
The boundary of the “Why did all the AI startups fail?” case is defined by this point: the first wave of generative-AI startups is too weak — Anthropic and OpenAI are effectively labs attached to Amazon and Microsoft, the ex-Salesforce CEO's startup raised over a billion and a half on revenue under twenty million, and the real process will start when the big five enter the market and M&A begins.
The practical meaning of “Will AI replace Tatyana?” is that at the basic level — layouts, selecting and counting materials — agents will do interiors better than a layperson, and a wave like “Pinterest interiors” will repeat; but high-end interiors remain art and client contact, so Tatyana does not expect to be fully replaced.
The working conclusion from “‘Call a human.’ Why today's AI agents are so dumb” is that an agent confuses cities, fails to understand local names, gets stuck on an exception, and cannot explain why it made a choice. In enterprise work this is critical: one unusual case can cost more than all of the promised savings.
The boundary of the “NOT all AI startups failed?” case is defined by this point: Glean proved the value of a new kind of enterprise search, and Perplexity made search easier to distribute — the winners solve a narrow pain better than the platform and build data, distribution, or habit before the platform catches up.
What this episode is about
Salesforce, Microsoft, IBM, OpenAI, and dozens of startups sell enterprise agents, yet mass adoption is almost nonexistent. The reason is not a shortage of presentations: an agent must connect many systems, understand business context, and act reliably in places where an ordinary chat interface already makes regular mistakes.
The word ‘agent’ has become a convenient way to sell the future. If a chatbot answers a question, an agent is supposed to open systems, gather data, make a decision, and perform the work on its own. But the models we criticize for poor answers do not become more reliable simply because they have been given access to buttons and corporate data.
Large companies already offer agent builders, but businesses face a simple question: why pay a separate startup for a complicated wrapper if OpenAI, Anthropic, Microsoft, or Google will add a similar feature to familiar software? A young company has to glue together several models, integrations, and databases—and then take responsibility for errors across the entire chain.
A real agent begins not with autonomy, but with context. An interior designer does not need a bot that generates a beautiful room. They need a system that can ask questions, find actual products, account for dimensions, budget, and the client's taste, and return several coherent alternatives. The market still rarely does even that hybrid process—search, Pinterest, generation, and dialogue—well.
Users still often want to press ‘talk to a person.’ An agent confuses cities, fails to understand local names, gets stuck on an exception, and cannot explain why it made a choice. In enterprise work this is critical: one unusual case can cost more than all of the promised savings.
This does not mean every AI startup is doomed. Glean proved the value of a new kind of enterprise search, and Perplexity made search easier to distribute.
The winners solve a narrow pain better than the platform and build data, distribution, or habit before the platform catches up. But selling a ‘universal digital employee’ before it can perform one job consistently is a direct path to disappointment.
The winners solve a narrow pain better than the platform and build data, distribution, or habit before the platform catches up. As a result, but selling a ‘universal digital employee’ before it can perform one job consistently is a direct path to disappointment.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 52 segments: 39 identified, 4 mixed, 7 probable, and 2 unresolved.
Read transcript on a separate page
Loading…