Artificial intelligence has already learned to speak almost like a person. But can it understand who you are and what you actually need? I tested two major updates: the new ChatGPT Voice and Claude Design. One system immediately began giving me answers after learning almost nothing about me. The other stopped and first asked dozens of precise questions. Why do they behave so differently? Which system is genuinely built around the person, and whose goals does the other actually serve—yours, or those of the company that created it?
Today we will examine what these two updates reveal about the real future of AI and how to work with artificial intelligence so that it truly helps you.
You are watching ToTheMoon: technology news and insights from Silicon Valley. We now release these special episodes several times a week. I want to explore this question through two newly released systems. We have mentioned Claude Design repeatedly, and I want to explain why it matters. Some people may think design is a narrow subject and the system is unnecessary, but the underlying idea is extremely important. I do not usually devote an episode to such a small topic. The second story is much larger: ChatGPT’s new voice model, or more accurately, a new way of communicating with ChatGPT by voice.
In OpenAI’s demonstration, a user says: “I am knitting a sweater for my grandson, but I do not want the needles to be too large. Do you think larger needles would make it slightly baggy? Although baggy clothing is popular now, isn’t it?” The assistant answers, “Of course. I think so too.” The user then asks it to verify several dates: Édouard-Léon Scott de Martinville’s phonautograph in 1857, Edison’s tinfoil phonograph in 1865, and Berliner’s gramophone in 1887. While the assistant checks them, the user asks whether there are delays at 16th Street BART station.
This is not merely a voice API for programmers to configure. It is meant to be your new voice assistant. We therefore need to look closely at how it works and what it carries inside it. Why did I open with the idea that people do not know how to listen to one another and do not truly learn who another person is? If you examine the goals of a person or a company, you should understand that when someone meets you, shows no interest in who you are, and asks nothing about what you do, they probably already have a plan for you.
Today I began testing the new ChatGPT Live Voice model. It is a fundamental update, and we discussed it in the podcast a few days ago. The remarkable claim is that the model can listen and speak at the same time. OpenAI also suggested that it could perform many different operations in parallel while talking with you. I gave it a very simple test: could it continue searching while we spoke? I asked it to recommend a coffee shop in Los Altos, the Silicon Valley city where I live.
Around us are small cities and areas such as Los Altos, Palo Alto, Mountain View, Stanford, and Menlo Park. I asked for one specific coffee shop where I would enjoy having coffee, told it to answer quickly, and then asked it to study the options more carefully in parallel using different criteria and what it already knew about me. ChatGPT stores a very large history about me. It started searching and listing coffee shops, while I kept trying to determine whether any work was continuing in parallel.
The longer we spoke, the clearer it became that nothing was happening in the background. I repeatedly asked whether it could work in parallel. It kept saying that it could not, although the screen appeared to show a search while it was talking. I do not think the speaking model itself fully understands how the overall system works. The voice has improved dramatically. It listens differently, adapts fairly well, and sounds very much like a person. But honestly, it reminded me of employees at work—employees who want to contribute something
to a company before understanding what the company actually needs. I then gave it a San Francisco example: “I am going to the area where Chinatown, the Italian neighborhood, and downtown meet. Find two coffee shops there that I am most likely to enjoy. First, give me a quick answer. Then, while we discuss other subjects, conduct a deeper search and return with the results.” The assistant replied: “Understood. I will generate a quick answer now. For coffee quality, I would choose The Coffee Movement in Chinatown, at 1030 Washington Street.
It is a small specialty coffee shop, but the cup quality is top-tier. For atmosphere, Cafe Trieste in North Beach, at 601 Vallejo Street. It is less about cutting-edge specialty coffee and more about an authentic Italian-American atmosphere. It is a historic, iconic place. I cannot continue researching in the background, but we can return to this subject later and do a detailed analysis.” I interrupted: “On what basis did you choose these two cafés?” It answered: “First, I looked at their reputation for coffee quality among knowledgeable people and publications.
Second, atmosphere and historical significance. Third, convenience for the Chinatown—” I stopped it again: “How did you decide that I need any of that? How did you decide that coffee quality or historical significance matters to me? Where did you get those criteria?” It admitted that it did not know my precise priorities and had made silent assumptions. It asked whether I cared most about coffee, food, cuisine, or atmosphere. I reminded it that it had a memory of my previous chats and asked why it could not look there.
After checking, it said it should have relied on that information instead of guessing. It claimed that I usually favor a calm, high-quality atmosphere, convenience, and a certain level of service rather than a “hipster cup at any cost.” On that basis it decided The Coffee Movement was probably not my format, Cafe Trieste was closer but still imperfect, and it would shift the recommendation toward Revel on Columbus Avenue while keeping Trieste. Revel, it said, was calmer, more modern, cleaner in presentation, and more convenient for sitting down.
I asked where it was and whether it was closer than Trieste to the intersection of the Italian district, Chinatown, and downtown. It gave the address as 200 Columbus Avenue, in the southern part of North Beach and closer to downtown. Trieste is deeper inside North Beach and feels more Italian. This second San Francisco attempt was interesting.
The system first proposed a coffee shop in Chinatown and then Cafe Trieste. I genuinely like Cafe Trieste; in terms of atmosphere it may be my original top choice in San Francisco. Then it proposed Revel, which is my second favorite in that area and perhaps the most sensible choice there. I was surprised. Either something changed in the few days between my first recording and this repeat, or it finally used feedback I had given earlier. Its use of memory was clearly better this time.
This is an important subject. Several years ago I said that the model that learns to ask a person questions will win. A model that is not interested in people may also win by becoming some extraordinary giant,
and ChatGPT’s development may be moving in that direction. But the model that wins the broad consumer market, among large numbers of people, will be the one that asks questions. Early in ChatGPT’s history, OpenAI introduced Deep Research. Some viewers may not remember it, especially those who were not early users. Its purpose was to conduct a large-scale search, mainly on the internet but also across files and other sources. At the beginning of a research task it would ask three or four questions.
The mechanism contained a standard set of clarifying questions, and OpenAI experimented with different ways of asking them. In my coffee-shop test, however, the system immediately began making recommendations. I then asked it to use my preferences, and it proposed more places. I asked for deeper research, and it proposed still more. It made errors and odd assumptions. At some point I asked it what criteria it was using and whether it had examined who I was through my previous chats.
It responded with nonsense—generic phrases about universal preferences and organized spaces, vague marketing language that meant almost nothing. The longer we spoke, the more obvious the problem became. If it did not know me and could not retrieve the information from my previous chats, why did it not simply ask what mattered to me? I tried to make the point through my history of hotel searches.
I travel often and use a fairly standard method to evaluate hotels. I compare them across many parameters and score them on a hundred-point scale, where one hundred represents an Aman hotel, the famous ultra-expensive chain whose rooms can average around six thousand dollars per night. The purpose is to give the model a reference point for quality. It cannot award one hundred points to a property that is nowhere near that standard; using Aman as the benchmark forces it to break quality down into separate dimensions.
In theory, the system’s memory should have shown it this recurring method. When I asked what it knew about me, it listed some basic facts that were largely correct. We have already made an episode about this, and I am running a large project that analyzes all of my chats. I discuss that project on the channel and will show it when the implementation is ready. Yet the model could tell me nothing meaningful about my coffee preferences because, in practice, it did not know me. I told it: “If you do not know me, why do you not simply ask what interests me?” It agreed to ask, but after I answered, it still did not clarify what I meant in any depth.
It used an interesting phrase: “I will do everything without guessing, exactly as you said.”
But I had just explained that some judgment and informed inference would be necessary. Here we have a new voice system that appears extraordinary, but the crucial question is what concept it contains. It embodies OpenAI’s concept. OpenAI is building a system for its own objective: creating some kind of extraordinary computer. We discussed that in last Friday’s AGI episode and in an earlier episode with Max Grigoriev about the missions of companies such as Anthropic and OpenAI.
Max visited me today, and we agreed to make an episode about terminology in AI—not simply to define words, but to show the reasoning models behind them. ChatGPT, in this framing, treats people partly as experimental subjects. Any normal assistant designed around the user should have asked me questions—not a fixed questionnaire, but questions generated from my own history. I explicitly told it from the beginning to rely on that history, so it could have examined substantial detail.
On the same day, I began testing Claude Design. Ilnar mentioned it in our podcast in early July, and Tanya said she had looked at it but did not understand where to use it. I had a concrete case. I am building a number of projects in parallel as AI-native software, meaning complex software created entirely by artificial intelligence, not merely a collection of agents. One task is to create a production system that supports all of my commercial projects and can generate different websites and portals.
I am testing it across several of my own businesses and some partner businesses. In one company, I said that we lacked a precise design system for creating landing pages and sites, and I asked Claude Design to develop one. I asked Claude Design to work on that system because ChatGPT had already proposed some reasonably good ideas about color schemes and the structure of the site, but it focused mainly on structure. I wondered what Claude Design would do differently. I returned about an hour later and saw that it had not moved.
Before continuing, it had decided to ask me ten or fifteen questions. Those were some of the better questions I have received from an AI system in quite a long time—not only well structured and comprehensive, but written in human language. Anyone who uses Codex or Claude Code has seen systems ask extremely technical questions, invent a discussion that never happened, or use bizarre jargon. These systems often speak in a strange internal slang. A modern voice assistant should begin by adapting itself to me: what voice do I prefer, what pace is comfortable, should it address me formally or informally, and what communication style should it use?
The new ChatGPT model did none of that, and I consider it a major design failure. When a system shows that it was made for you and is willing to configure itself around you, that is incredibly powerful. Claude Design did this by asking questions. It asked how important it was to support different color schemes, especially because ChatGPT had proposed several and I had shared that material as a starting point. It asked which kinds of users and design styles mattered most, and what I considered bad about the existing site.
That was a genuinely interesting question, and I answered it. Almost every set of answer options included “I will decide myself,” leaving room for the system to make the professional decision. That is excellent. Claude Code usually does not do this, and it frustrates me. When you work with a system professionally, you often want it to say: “I came to consult you, but I am the professional here and will make the final decision.” GPT-5.6 Sol, which had just begun rolling out the day before I recorded this, also asked me some decent
questions—but only after I explicitly told it to ask them. After roughly twenty-four hours of work with GPT-5.6 Sol, using many five-hour windows, sessions, and a great many tokens, I found that its questions were more substantial than those of GPT-5.5 High. But I am not very satisfied with the model. The main reason is speed: it is unrealistically slow and appears to complicate tasks that do not require complexity. Simple tasks that earlier systems completed quickly now take a very long time.
Several times I had to stop it and say, “You must finish the task right now,” after which it somehow completed the work immediately. That looked extremely strange, almost as though the system were tuned to consume tokens or prolong the process. Anthropic’s Fable now seems to behave similarly during development work. When Fable first appeared a month earlier, I was in Santa Barbara and used it intensively. It was much faster and produced better work. Today I again had to move many tasks to Opus 4.8 because Fable seemed to be slowing down.
It is also constrained by usage limits, so you do not know whether a quick audit is worth burning expensive tokens. Many of my routine audits can start costing an additional fifty, one hundred, two hundred, three hundred, or four hundred dollars each, which becomes tens of thousands per month.
Claude Design asked strong questions and then went to work. What still bothered me was that after completing a stage, it did not return with a new set of questions. I had to ask whether it had comments, suggestions, or ideas. Only then did it continue offering material. Ideally, it should have done that automatically—not merely asking whether it could continue somewhere, but deepening its understanding of the task and drawing me further into the creation process. It has many tools available for that.
The product also has significant usability problems. It can export materials in multiple formats and pass work onward to Claude Code, but the interface makes those capabilities far from obvious. Even with those shortcomings, Claude Design was interesting enough to share. Long-time ToTheMoon viewers know that one of my current projects is moving my entire large website to a new CMS. I mentioned it in Sunday’s episode. Today Codex parsed the whole site in about thirty minutes using GPT-5.5 High and very few tokens.
The result was an archive of roughly 530 to 550 megabytes. Just before recording, I sent the archive to ChatGPT so that it could think through the design and style structure, as well as the possible architecture of the portal and website. I have many ideas of my own that still need to become a proper specification. I also wanted to give the archive to Claude Design, but a few minutes before filming I exhausted my limits and a new usage window began, so I decided to let the system wait and return to it in a few hours.
Claude Design consumes the ordinary Claude token allowance.
Its willingness to ask questions produced the strongest reaction in me. How important is it for artificial intelligence to ask questions? And how important is it in human life for people who communicate with one another to know how to listen and learn who the other person really is? Is there not an enormous problem in our society—something I also discuss on my other channel about personal development—that people do not want to hear or understand others? This becomes a crucial test for AI: is the system doing something for the person in front of it, for you, or is it serving goals and tasks of its own that are often invisible to us?
I explored that question in the AGI episode. Share your conclusions, approaches, and tests. I would like to know how GPT-5.6 Sol, Claude Design, and ChatGPT Voice have worked for you. We will see one another in the next episodes.