Skip to content
Google · OpenAI · GoEpisode 057 · 11 May 2025 · 42:21

Voice Assistants Still Make Mistakes, While Scammers Already Use AI at Full Scale

Central question

Why do voice assistants still make basic mistakes while fraudsters already use generative AI at full scale?

What you take away

Compare the maturity of legitimate voice assistants and fraudulent uses through accuracy, speed of adaptation, data access, and the user’s ability to opt out.

Main threads

What to watch for

1Compare “A class-action lawsuit over Apple Intelligence” with “On Alexa Plus: pros and cons”: they provide different criteria for judging the same issue.
2Test the conclusion from “An explosion of AI-powered cybercrime” in your own use case—what actually changes in the process and what remains a promise.
3Before choosing a product or approach, record the constraint identified in “The UAE: drafting federal laws with AI”.
4Define the owner of the outcome and the quality metric for the situation described in “Why speed matters more than repeat-founder experience”.
Signals to track afterwards
Watch for actions by Apple and Google that confirm or challenge the episode’s central claims.
Compare new launches and policy changes with “An explosion of AI-powered cybercrime”: have access, quality, price, or constraints changed?
Check whether the scenario in “Why speed matters more than repeat-founder experience” becomes repeatable practice rather than a one-off demonstration.
Most useful for
AI usersProduct teamsEntrepreneursExecutives and managersContent creatorsDesigners

Key takeaways

00:00The market tests it through use: in this issue

The discussion of “In this episode” yields a practical test: while big companies refine voice interfaces, cybercriminals already use generation with no quality requirement — they only need to fool a small percentage of people; the episode is built on that contrast.

04:00The boundary between value and constraint: gemini Metriki

The boundary of the “Gemini's metrics” case is defined by this point: pretty Gemini metrics mean little if the model hallucinates on real queries; the value of a voice and text assistant is decided not by a number in a report but by accuracy and predictability on the user's tasks.

05:55Apple Intelligence promised features that were delayed or did not work as expected

The practical meaning of “A class-action lawsuit over Apple Intelligence” is that Apple Intelligence promised features that were delayed or work differently, and the class action matters not for the amount but for the principle — the company sold a device through future AI while the user paid now; the Mail summary is useful but does not replace the promised intelligent Siri.

07:00Amazon describes the Alexa Plus problem honestly: the accuracy of complex multi-step commands may remain at a level that is unacceptable for a household assistant

The discussion of “On Alexa Plus: pros and cons” yields a practical test: Amazon honestly admits the accuracy of complex multi-step commands may stay unacceptable for a household assistant — if the agent sometimes orders the wrong thing or fails to act, trust is lost entirely.

10:59Why context matters more than one metric: using ChatGPT's voice mode

For the “Using ChatGPT's voice mode” scene, the decisive point is this: ChatGPT's voice mode sounds more natural than Alexa but still requires verification — pleasant speech does not guarantee a correct action, and the person still has to check the result.

14:07How the issue moves from news to product: transpa Decree: AI‐Chunglasm K‐12 throughout the country

The discussion of “Trump's executive order: AI courses for K-12 nationwide” yields a practical test: an order to introduce AI courses in schools changes access to learning at scale, but the real value will come not from the order itself but from what and how children actually learn and who is accountable for the programs' quality.

16:00Scammers do not need that level of reliability

The practical meaning of “An explosion of AI-powered cybercrime” is that scammers do not need reliability — they generate thousands of emails, voices, and messages, and a one-percent hit rate is enough; that is why cybercrime captures the benefit before the ordinary user does.

20:52The UAE is going to the opposite extreme by officially using AI to analyze and draft legislation

The “The UAE: drafting federal laws with AI” issue should be assessed with one constraint in mind: a system can find contradictions in laws faster, but political choice and responsibility cannot be delegated to it — the higher the level of the decision, the more important it is to see the underlying data and a human approval.

22:13What determines the outcome: startap in defence: 21-year-old faunder attracted +$185 million

In the context of “A defense startup: a 21-year-old founder raised +$185 million,” this criterion applies: a large round for a young founder shows a bet on execution speed, but the business's durability will be decided not by the amount and age but by repeatable revenue, customer retention, and reaching the next round of funding.

23:55OpenAI and Anthropic assembled a rare concentration of talent before the mass boom

The “Why speed matters more than repeat-founder experience” issue should be assessed with one constraint in mind: OpenAI and Anthropic assembled a rare concentration of talent before the boom, which new startups find hard to reproduce, so execution speed matters more than a repeat founder's biography; but devices — glasses, pendants, voice assistants — add continuous data collection, and the market matures only when it shows accuracy, data-training rules, and the ability to opt out rather than a promise.

What this episode is about

Apple faces a lawsuit over unfulfilled promises, Amazon acknowledges low accuracy in multi-step agents, and ChatGPT and Meta are testing voice. While large companies refine their interfaces, cybercriminals use generation without any quality requirement—they only need to fool a small percentage of people.

Apple Intelligence promised features that were delayed or did not work as expected. The class-action lawsuit matters not because of the amount, but because of the principle: the company sold a device through future AI, while the user paid in the present. Summaries in Mail are useful, but they do not replace the promised intelligent Siri.

Amazon describes the Alexa Plus problem honestly: the accuracy of complex multi-step commands may remain at a level that is unacceptable for a household assistant. If an agent sometimes orders the wrong thing or fails to carry out an action, the person stops trusting it altogether. ChatGPT’s voice mode sounds more natural, but it also requires verification.

Scammers do not need that level of reliability. They can generate thousands of emails, voices, and messages; a one-percent success rate is enough. Cybercrime therefore captures the benefit before the ordinary user does—the criminal does not have to build a stable product or maintain a reputation.

The UAE is going to the opposite extreme by officially using AI to analyze and draft legislation. A system can find contradictions more quickly, but political choice and responsibility cannot be delegated to it. The higher the level of the decision, the more important it is to see the underlying data and a human approval.

OpenAI and Anthropic assembled a rare concentration of talent before the mass boom. New startups find that hard to reproduce, so execution speed matters more than a repeat founder’s biography.

Devices—glasses, pendants, and voice assistants—also add continuous data collection. The market matures only when it shows not the promise, but the accuracy, the rules for training on data, and the ability to opt out.

The voice-assistant market will mature only when accuracy, training-data rules, and the ability to opt out matter more than promises.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 108 segments: 71 identified, 3 mixed, 13 probable, and 21 unresolved.

Loading…