ChatGPT vs. Claude: Which One Actually Works for the Person?
What separates AI that genuinely understands the user from a system that merely produces a plausible answer quickly?
Learn how to evaluate voice and design AI tools in real work: whether they ask clarifying questions, use memory at the right moment, explain their criteria, and let the user correct the course.
What to watch for
Key takeaways
ChatGPT Voice sounds natural, but in the coffee-shop test it began recommending places without first learning what mattered to the user. A strong interface cannot compensate for a poorly framed task.
The system can already listen, respond, and search at the same time. The central problem has shifted from technical capability to transparency: where the recommendation came from and which assumptions produced it.
Once prompted directly, ChatGPT used prior conversations and moved much closer to the user’s real preferences. The memory existed, but it was not integrated into the initial decision.
For complex work, a system that stops and checks the goal is more useful than one that instantly produces a polished but incorrect result.
Before creating the design system, Claude asked about the audience, style, constraints, and degree of freedom. Those questions clarified the task itself rather than merely improving the visual output.
GPT-5.6 Sol can reason deeply, but latency, overcomplication, and the need to stop the process manually reduce its practical value.
When AI does not establish the user’s intent, it can replace that intent with goals embedded in the product and the developer’s business model. User control begins with the ability to inspect and change the criteria.
What this episode is about
The new ChatGPT Voice sounds almost human, while Claude Design knows when to stop and ask questions first. Through two real tests, Alexander Volchek examines why the quality of AI depends not only on model power, but also on whether the system tries to understand the person using it.
A Strong Voice Does Not Yet Mean Understanding
The new ChatGPT voice mode makes a strong first impression: it listens better, responds more naturally, and genuinely resembles a human conversation partner. But a real coffee-shop search exposed a more important problem. The system began recommending places without first learning what mattered to me: coffee quality, atmosphere, convenience, food, or a quiet place to work. It relied on generic criteria and admitted only after several direct questions that it had made assumptions instead of learning about the person.
Memory Exists, but the System Still Has to Use It
ChatGPT stores a large history of previous conversations, and in a repeated test it moved closer to my real preferences: it first suggested The Coffee Movement and Cafe Trieste, then adjusted the recommendation toward Revel. The result was surprisingly close to what I would actually choose. But the system reached it only after being pushed to consult its memory and explain its criteria. That shows the difference between possessing data about a person and knowing how to use it at the right moment.
Claude Design Started with Questions, Not an Answer
The contrast appeared in my work with Claude Design. When I gave it the task of creating a design system, the tool did not immediately begin drawing a structure or proposing colors. It stopped and asked ten to fifteen substantive questions: which audiences mattered, which styles fit, what was wrong with the current website, whether multiple color systems were needed, and where the system could make the decision itself. The value was not in the number of questions, but in the way they forced the task to become more precise.
The Real Race Is to Be Useful to a Specific Person
The first day with GPT-5.6 Sol added another layer to this picture. The model asks good questions when prompted to do so, but it works slowly, sometimes overcomplicates simple tasks, and may need to be forced to stop reasoning and produce a result. Claude Design has shortcomings as well: awkward handoffs, token-limit constraints, and too few follow-up questions after the first stage. The choice between these systems therefore cannot be reduced to a single benchmark. The more important question is which one genuinely adapts to the person, asks about goals, and helps make a decision—and which one simply executes its own script as quickly as possible.
The best AI assistant is not the one that starts talking or generating fastest, but the one that recognizes missing context, asks the right questions, and remains under the user’s control.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 38 segments: 38 identified, 0 mixed, 0 marked with ✓, and 0 unresolved.
Loading…