Skip to content
McDonald’s · Artificial intelligence · QuickBooksEpisode 146 · 7 August 2026 · 32:27

Will AI Solve Your Task—or Make You Lose Money? How to Tell in Advance

Central question

How can you determine before major spending whether AI will solve your specific task, what error rate is acceptable, and whether manual work with a model is more economical than a dedicated automation system?

What you take away

Gain a framework for evaluating an AI project: task value, cost of error, quality on real data, need for human intervention, and the economics of the full process.

Main threads

What to watch for

1Before development, define acceptable error through consequences: what happens if the system is wrong and a person does not notice?
2Run an end-to-end test on real data and edge cases, not only on a convenient demonstration sample.
3Count human interventions, corrections, retries, and support costs; they are part of the system outcome.
4Compare a dedicated AI product with a manual process enhanced by ChatGPT or Codex; automation is not always the best next step.
Signals to track afterwards
The share of operations AI completes without intervention, not only average accuracy.
The cost of handling exceptions and errors after deployment.
The pace of change in base models relative to the development cycle of a dedicated product.
Cases where a general model removes the need for a familiar software intermediary.
Most useful for
companies planning AI automationoperations and finance leadersfounders of AI productsproduct managers and developerspeople evaluating the return on an AI project

Key takeaways

00:00Not Every Task Should Be Automated

The first question is not which model to choose, but whether the task has enough value, longevity, and acceptable error for a dedicated system.

01:35A General Model Can Solve the Task Without a New Product

In the financial experiment, Codex received source documents and helped assemble the needed result within hours. This changed the question of whether a separate application was needed at all.

05:36Familiar Software Can Become an Unnecessary Layer

If AI works directly with bank transactions and tax documents, the value of an intermediary accounting system has to be proven again through control, reliability, and required functions.

10:18Trust Depends on the Consequences of Error

A creative task can tolerate many unsuccessful variants. A financial document, tax calculation, or customer order requires a different threshold and a different form of review.

16:38Demo Accuracy Is Not the Same as a Working System

The McDonald’s voice project could look convincing on most orders, while the remaining errors continually brought employees back into the process.

22:03Eighty-Five Percent Can Be Both Success and Failure

One accuracy number says little without context. You need to know which 15% of cases are lost, what correction costs, and how often a critical scenario occurs.

25:17Real-World Exceptions Consume the Economics of Automation

Noise, accents, similar names, and complex orders turn a small statistical error into a persistent operational burden.

29:47The Decision Must Be Based on the Whole Process

The model should not be evaluated in isolation. The full path matters: input data, review, exceptions, accountability, and cost per completed operation.

What this episode is about

One AI project can remove an unnecessary financial layer in a matter of hours. Another can spend years failing to reach the quality required for real work. A comparison between a personal Codex experiment and McDonald’s voice-ordering system shows how to evaluate automation before making a large investment.

First Ask Whether the Task Should Be Automated at All
Before building a large AI system, it is important to ask a simple question: should this task be automated at all right now? In some cases, manual work with ChatGPT already solves the problem faster and more cheaply than a separate product. In others, the technology changes so quickly that a complex system becomes obsolete before deployment is complete. The first step is therefore not choosing a model, but evaluating the value of the task, the acceptable error rate, and the likely lifetime of the solution.

Financial Accounting Showed How a Software Layer Can Disappear
In a personal experiment, Codex received bank transactions, tax documents, and other financial data from a company. Within hours, the system matched the materials, prepared a report, and helped verify balances. That raised a direct question: why keep QuickBooks as a separate layer if AI can work with the primary data and assemble the required view for a specific task? The answer depends on reliable verification, but the status of the familiar software product has already changed.

The McDonald’s Case Shows the Cost of the Missing Percentage Points
McDonald’s spent years testing voice-based order taking. In ten Chicago restaurants, the system produced correct orders about 85 percent of the time, while later estimates in Illinois were closer to 80 percent. That looks respectable in a demonstration. For a chain where one complex order can involve accents, noise, similar product names, and many items, the remaining 15 to 20 percent means constant employee intervention, errors, and customer frustration. The company wanted to approach 95 percent, but the project was eventually shut down.

The Acceptable Error Rate Is Defined by the Task, Not the Model
In a creative task, a system may be wrong most of the time and still produce useful options. In a financial report, tax document, or customer order, even a few percent of errors may be unacceptable. Comparing AI projects through one accuracy number is therefore meaningless. The user has to define what an error actually means, who will detect it, how much correction costs, and what happens if it is missed. Only then can the business decide whether to work manually with a model, build a full automation, or wait for the next technological step.

An AI project succeeds not when the model shows an attractive accuracy number, but when the entire process reaches the required quality more cheaply and reliably than the alternative. Sometimes the answer is automation; sometimes it is a person using a general model well.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 47 segments: 47 identified, 0 mixed, 0 marked with ✓, and 0 unresolved.

Loading…