Hello, everyone. You are watching ToTheMoon: technology news and insights from Silicon Valley and around the world. This is our weekly podcast, although lately we have begun releasing episodes every other day—or almost every day. Ilnar, do you know where I want to begin? Last Sunday we said that Elon Musk had written about the end of the browser era—or at some point I said that the browser era was over and people would work through other systems. Then OpenAI announced that it was shutting down its Atlas browser and that the right model is for people to work through chat.
You sit there thinking: where have all of you been? How many episodes did we devote to this question while Perplexity was building its browser and every browser was adding agent and AI functions? From what we can see now, that whole direction is becoming less relevant. Could a browser still contain some interesting functionality that is useful to people? Perhaps. At some point a browser of that kind may appear. But in my view it is much more likely that systems such as ChatGPT will absorb browser
functions than that browsers will absorb systems like ChatGPT. Anthropic, by the way, has added its own browser to Claude—a complete, fully integrated browser. I do not know whether everyone has it yet or whether it works fully for everyone. Did you have it, Ilnar?
Yes, Anthropic released it and demonstrated it directly. We will see how it works and whether it can become part of a real workflow. If someone spends the whole day in Codex or Claude, will that person be able to remain there when they need to make an independent browser request or open something on the web? At the moment, even when I am building an application in Codex or Claude, I still open the application through a normal website in a browser to examine it. I do not like the way Codex or Claude opens pages internally.
The same applies to ChatGPT. I do not exactly complain about it constantly, but I keep saying that the interface is not very good. Open an Excel file and it displays the spreadsheet inside the application. Open HTML and it renders it internally. The same now happens with PDFs. In 5.6 Sol they have started opening everything inside the product. And you say, “Can I just download it? I am accustomed to viewing it differently.” I think the real issue is usability and the ease of handling windows on the screen, rather than the precise place where the content opens.
If their internal viewer displayed a genuinely good, functioning version, perhaps you would not care which application opened it.
As you speak, I am wondering whether it is possible to abandon the browser completely for
certain kinds of work. For now, my feeling is no. We may genuinely arrive there, and it is understandable that OpenAI is closing its browser and moving everything into a super-app containing code generation, information search, chat, question-and-answer systems, and agent functions. Everything sits inside one product. But at least in the near term, some things will still be difficult to pull into it. That is the first issue. The second is that not every company will be willing to hand over its entire experience.
That is the problem.
What happens to the recommendation system? What about advertising and the promotion of other videos related to your interests? If you ask ChatGPT for one specific video on one specific topic, ChatGPT—not Google's system—becomes the recommender. Google would lose control and, consequently, revenue. So not everything will move into these systems. But Codex, Anthropic's Cowork, and analogous products from other companies are clearly becoming super-apps. That is where we are heading, and overall it looks quite promising.
You said immediately that standalone AI browsers had little purpose and that nobody would use them. These broader systems, by contrast, clearly have use cases in which they are applicable and genuinely useful. They have already attracted an audience, and that audience will probably continue to grow.
Yes, we will see what happens. I think the situation with browsers is a healthy example of understanding that neither browsers nor a great deal of today's software will survive in their present form. YouTube is a somewhat different story. A browser is merely an intermediary that lets you enter YouTube. On my phone, for example, I do not open YouTube in a browser; I use the mobile application. On a Mac, however, I do not download a YouTube app. I do not even know whether one exists for macOS.
I assume it probably does, unless I am being ancient, but for me it would be completely
pointless. These are simply different usage patterns. On a television, of course, I open YouTube through the YouTube application rather than through the TV's browser. We have these familiar scenarios. Why? Because using a browser on a television is horribly inconvenient, and the browsers themselves are slow. I hope that modern AI and the new systems finally lead someone to rewrite everything related to television operating systems and applications, because the current state is an ancient horror: undeveloped, inconvenient, and almost unusable.
I have joked about this for years—perhaps since we made that programming series together. What difference does it make whether a device contains half a gigabyte, one gigabyte, or ten gigabytes of memory? In the modern world the incremental cost should be negligible, yet the television's memory fills up or the system stutters when opening the simplest basic applications. Relative to the cost of manufacturing the display panel itself, this should be almost free. And yet manufacturers remain trapped in—
People in the comments are going to remind us that memory is very expensive right now, Sasha, but the point is clear. My only concern is that television and electronics manufacturers are extremely inert, just like carmakers. Tesla showed perhaps ten years ago that you could put a huge tablet into a car, yet almost nothing changed. How many years passed before the major traditional automotive brands began implementing anything similar? For years they installed extremely outdated multimedia systems even in very expensive vehicles.
Yes, some player probably has to appear that gets the price-to-quality ratio right and supplies the right software. I am sure this is much more advanced in China from the standpoint of operating systems. But for it to work well everywhere—like a home device or a smart home—the sector still lags behind. New technologies will clearly force updates, including to intermediary layers such as the
browser. Once again, I would place my bet on the operating system. I am convinced that our very understanding of an operating system will be transformed, and it needs to be. The operating system once meant something like Windows. Even macOS, despite all of its development, inherited a great deal from Windows-style desktop systems. I was a full-time Windows user for nearly twenty years—from 1992 until around 2010—before moving to the Mac. When I moved, I completely cut myself off from a whole part of the experience I used to have in Windows.
I remember drivers very well. I remember system configuration and what it meant to reinstall an operating system. Some of that legacy remains in macOS. Phones work differently because they were conceived differently from the beginning. I am interested in what our viewers think about operating systems in general. You know my opinion. What do you think, and can you explain to people the question we are raising? I think it is a very interesting one.
I think you are absolutely right that this is moving toward the operating system. Researchers at different companies are already saying that something like ChatGPT is becoming your operating system in practice. An operating system is a layer that manages the hardware's resources: it allocates computing capacity, determines what is currently a priority and what is not, allows processes to run in parallel, and governs your interaction with the device. I do not think the iPhone will suddenly switch to “ChatGPT OS”; that will not happen quickly.
But if OpenAI and Jony Ive release some kind of ChatGPT Phone—whatever they eventually call it—this approach could be implemented there. That may shift the paradigm, with more devices managed by an operating system whose underlying intelligence comes from OpenAI, Anthropic, or another provider. I would like to see the real device first—not the little remote-control or switching device they released, but the product that motivated the collaboration with Jony Ive. There is a great deal of speculation that it will be a wearable device, a phone, or a smart speaker.
If it resembles a phone but is controlled primarily by voice, it could substantially change the way we interact with hardware. At that point we could reasonably say that ChatGPT, or a variation of it, had become a full operating system: the layer between the user and the equipment. The shift could be comparable to the moment Steve Jobs showed that a phone could be mostly one large screen—initially with one button, but with the display occupying far more space than the rest of the device.
This could create a different shift in the opposite direction.
A genuinely new device would, of course, represent a completely different paradigm. You might do an enormous amount without a screen or keyboard—using voice, for example. Or you might use many different devices to complete one task. You move from room to room; a system begins working on your computer, continues inside the car, and perhaps hands part of a task involving your wife to a device that belongs to her. I am speaking broadly about operating systems, including the way we work on a Mac today.
Apple has made several attempts to bring phone-like interfaces into macOS. The iPad also contains many phone-like interfaces that often fail to take root. Conversely, Mac-like or desktop-style interfaces frequently fail on the iPad. A person says, “I cannot divide the screens this way. I cannot handle a large number of tabs in this environment.” People react to and perceive the devices differently. Most likely, a true shift will happen only through a new device. We are unlikely to see a complete transformation of one existing operating environment.
Will there be a moment when I work almost entirely inside one system? Today I spend a great deal of time in Codex and Claude Code, but I still use the regular ChatGPT chat outside Codex. Even OpenAI has not managed to move people into one coherent, understandable interface. Yet I think that is essential if people are to discover the real capabilities of AI. Right now applications are divided into “professional” and “nonprofessional” products, while many people using the nonprofessional product remain on a free plan.
They never reach the advanced functionality and do not even understand what is possible. The free version actually includes some Codex access, although it is limited to a particular model. You cannot open Sol in ChatGPT, for example, but I think you can open Terra—perhaps even Terra rather than only Luna. These are ChatGPT model names that have not yet taken root. In the professional world everyone seems to know what Sol is, while Terra and Luna may die where they began. I personally think they will disappear.
The numbers and complicated names are unnecessary. In principle there should simply be ChatGPT with different settings. I am interested in what people will say about this. We also have several product updates to cover, and I think we should run through them at a high level; otherwise we will remain entirely inside the most complex systems and philosophical discussions. First, the Chinese model Kimi K3 has been released, and Google has updated Gemini.
Let us remember that Gemini still exists. Someone recently commented, “You only talk about ChatGPT and Claude.” I remember when the complaint was, “You only talk about ChatGPT.” We talk about the systems that are genuinely the most important. After that, you can decide what to use for yourself. The question is why. If you do not have access to the leading models, fine. But using another model merely because somebody created it, when it performs two percent of the task, makes little sense.
As for Gemini, I think its problem is that Google has not released the Pro version, and the product is falling behind.
We were waiting for Gemini 3.5 Pro, yes.
Yes, Google is falling behind. It was supposed to release the Pro version at the beginning of July, or perhaps at the end of June, and it is late. It is also behind in coding. What have we seen over the last month? Two new players have entered application creation. And application creation is no longer really about programming. It is about how an individual uses AI. Everyone needs to understand that. This is not a programming question; it is a question about your own processes.
It is not even merely about automating those processes, but about the processes themselves. Calling this programming is like saying that setting an alarm is programming. Creating something with ChatGPT should become as simple, light, and accessible as setting an alarm in an alarm-clock application. Two new operators have appeared. Grok has become a serious market player, participates in many benchmarks, and has its own particular strengths. Kimi K3 has appeared as well. We have seen many headlines saying that Silicon Valley is worried about Kimi K3.
Are you sure everyone is truly worried? Still, Kimi K3 appears to produce respectable results on a number of benchmarks. This week I released several episodes about token pricing and costs. Ilnar, a few days ago I published a video on our channel about ChatGPT 5.6 Sol consuming fifteen billion tokens for me in three days. I do not know whether you saw it. Remember I told you I had used seven or eight resets? I managed to see how many tokens the system had consumed and calculated that, if I had paid the listed price, the work would have cost hundreds of thousands of dollars.
And I did not even receive the result. Are you still using Sol, by the way?
I am, but just as cautiously as before. It consumes more of my time than I would like because I constantly have to supervise it. For certain tasks—
Do you switch to 5.5 High or Extra High at that point, or do you keep—
I am like the mice in the joke: they cried, pricked themselves, and kept eating the cactus. I continue wrestling with 5.6 Sol—
With Sol, yes?
—yes, and trying to get results from it. Overall, I do manage. My main complaint is still its agent behavior. It feels unfinished to me.
I would like to know what everyone else is experiencing with Sol. I have effectively paused my use of it in Codex. With Fable, for example, I can still see clear advantages in architecture analysis, solving difficult problems, and frontend work. There are real strengths. Sol raises much larger questions for me. For narrow tasks, I have no complaint. But once I give it a broader context, I become much less confident. Returning to Gemini, I honestly do not understand Google's direction or what there is to
discuss. It seems to me that Google is seriously losing people's engagement with AI through its own systems. It retains one extraordinary advantage: Google Search. Yet it is also losing search usage and will continue to lose it as people become more involved with AI. Of course, only a tiny fraction of humanity is deeply engaged with artificial intelligence today. The kind of work we discuss here involves a microscopic fraction of people. What do you think is happening with Gemini?
Google seemed to be producing serious results and moving very well. Yet when its latest message is, “We released an incredible model with billions of parameters and it is extremely powerful,” you look at it and think, “Perhaps I will test Grok and Kimi K3 instead.” Were you able to activate Kimi K3? Does a model like that interest you, or not really?
Let us take this step by step. First, Google. Remember that a month or month and a half ago we said the other players would have to answer Fable and the Midas-class models? That is exactly what happened. OpenAI released its response. If it had not, the number of Anthropic and Claude Fable users would have kept growing and OpenAI would have lost market share. It probably began to feel that pressure and released the product in whatever state it had reached. The model is good, although we have discussed at length what we dislike about the way Sol works.
At least a response arrived, and it interrupted the wave of Anthropic dominance. Google has not responded. Apparently it is having real difficulty bringing its model to a finished state. I would not write Google off at all. Only about nine months ago, Google triggered a code-red situation inside OpenAI by releasing Gemini 3 and setting off that internal panic. It can still release another very good model. Most likely, the model Google prepared turned out to be substantially weaker than Fable, so the company chose not to ship it yet and continued improving it.
That will take time. Everyone thought the delay would be one month, but apparently the problem is more difficult and will require longer. Perhaps Google will release it at the end of summer or the beginning of autumn. I hope it returns with a genuinely strong model; in my view, it has all the necessary resources. The longer Google waits, of course, the more market share it will lose. Yet users are not permanently tied to a single operator. Some people used only Claude, then began using alternatives.
So Google still has time—at least from where I sit. As for Kimi, it is excellent that the model has been released. The one issue is that, at the time of recording, Moonshot still has not published the weights. I think the weights are due on Sunday—or perhaps on Monday the twenty-seventh. Publishing them would be a major step toward democratizing neural-network use. There will be arguments and scandals claiming Kimi K3 is a distillation of Fable and that Moonshot AI somehow bypassed Anthropic's restrictions.
But if the weights are published, the open ecosystem will gain an enormous model with roughly 2.8 trillion parameters. Nobody will run it on ordinary home equipment. Large companies, however—especially those that want to distill Fable, GPT-5.6 Sol, or other advanced frontier models—will gain permanent access to an exceptionally capable model. They will be able to deploy it in a laboratory, run it as much as needed, fine-tune it, remove protections if the model contains them, and so on.
A huge model may soon enter the public market. It has not happened yet, but I hope it does. I do not know whether this will affect ordinary consumers in the near term. For companies, however, it is a major step. If you cannot prepare a very strong model yourself, you will be able to take Kimi Code 3, distill it for your own needs, and obtain a very capable model. We will see fine-tuned corporate versions appear.
Qwen has also promised to release an extremely large model. We have now entered the league of trillions of parameters. Moonshot AI itself, which released Kimi Code 3, describes it as a model in the three-trillion-parameter class. For comparison, GPT-4 was estimated at around six hundred billion parameters. We are suddenly looking at a model almost five times larger, plus additional technical innovations—
Let us examine an important case you have just raised. When chat systems first appeared, the competition was very easy to understand. I became attached to OpenAI and used ChatGPT. You also became attached to OpenAI and used ChatGPT. Other people said, “We use Google.” Fine. I used Google's or Meta's AI secondarily, where those systems were embedded in other products. But ChatGPT was my primary chat, and it took over perhaps eighty percent of the time I had previously spent across different browsers, websites, applications, and services.
I began using it—and still use it—to search for flights, select hotels, and perform similar tasks. I am already moving away from hotel listings on Booking. I still check a property's rating there, but even that rating has become relatively less important because ChatGPT can provide aggregated information. I would be interested to hear which use cases have become dominant for other people. At the same time Claude appeared. Because you work in data science, programming, and data analysis, your profession led you to adopt Claude very early.
You used Claude, knew Anthropic, and followed everything the company did. For me, Anthropic remained a peripheral product. I bought some plans almost as entertainment, but I did not move my main chat there. Now we have entered a new era in which people develop their own solutions and processes through a bot. This is not merely setting a reminder or launching a little agent. Today you can say, “I dislike the way this camera is integrated. Quickly write an integration for my camera; here is the API access.” It is still more complicated than that in practice.
It is not yet as effortless as Karpathy's camera example that we once discussed. But it works incomparably better than before. A genuine leap has occurred. Kimi Code 3, for example, has an interface close to Codex called Kimi Code Web Mode. Grok had a gap here, although it can be used through Cursor or another environment. But here is the important point. When Claude appeared, you continued using your own chats and memory in ChatGPT, as far as I know. You remained in ChatGPT, did you not—or has that changed?
Exactly.
The issue is not only model quality—
Let me finish the thought first—sorry. I have the most expensive Anthropic plan, the most expensive OpenAI plan, and subscriptions to the other major services. I am not boasting; I am explaining that access is not the issue. I have always had a Gemini subscription because I use Google services. The Gemini app is deliberately placed on one of my iPhone screens, and I keep very few applications there. Yet I never open it. I do not open Grok, Gemini, or Claude for ordinary daily use.
I have not moved my memory into them. Claude, incidentally, allows some memory to be imported from other systems, although only a limited volume. ChatGPT remains my primary device. This creates an interesting economic question. Token limits exist, yet a two-hundred-dollar subscription can currently give you tens of thousands of dollars' worth of tokens each month—at least tens of thousands, and perhaps much more. It is certainly not merely a few thousand. Claude also has the strange problem that programming usage is tied to your personal chat allowance.
If I exhaust everything while coding, it is as though my ordinary AI life has been cut off. That is a very odd design. I hope OpenAI does not make the same choice, although it appears to be moving in that direction. This separation is an enormous competitive advantage that people rarely discuss. OpenAI, Sam Altman, Greg, and the company's other leaders barely talk about it on X or elsewhere. They could simply say, “This is one of the major ways we differ.” Why do I use both Claude and Codex?
One reason is that I can effectively use two sets of limits. If the models are similar in quality, I gain double the capacity. Then the next question is whether I should buy Grok or Kimi Code 3 as another separate subscription, or follow your advice, Ilnar, and buy a second Claude subscription so I can develop another application, solution, or automation in parallel. How can one provider lock me exclusively into Codex if I hit the limit while trying to build something I need even in ordinary life?
A new device is about to arrive, and I will have to configure my gate again and program a few things. In my businesses, I may be prepared to spend thousands or sometimes tens of thousands of dollars on tokens. I do not want to spend thousands configuring a home gate. That would be absurd and uninteresting. It is much more rational to buy a separate twenty-dollar Pro subscription and use Codex to configure the gate. I therefore do not fully understand the marketing logic or strategy of these companies.
How will they differentiate themselves and keep attracting people for daily work, especially as users begin orchestrating multiple models? I do not mean agents operating inside one model. I mean reaching the point where I say, “This part is being programmed for me through ChatGPT, this part through Claude, and this part through Kimi.” What happens then? I have said this before and will repeat it: I believe a player with AGI-level seriousness will eventually appear. That provider should offer a coherent, properly designed pricing model that causes me to commit and remain with OpenAI, Anthropic, Claude, or another platform.
At that point I might genuinely detach from the rest. Right now, however, these companies have stepped on their own tails. They are trapped in a marketing fight: Anthropic releases something, another company copies it, one side acts and the other responds. The endless public quarrels have become childish. Look at Sam Altman and Elon Musk—it appears extremely strange. In these arguments, the companies risk losing the strategy of building long-term user attachment, the way Apple does.
Whatever the circumstances, I am not yet prepared to buy a phone from xAI. SpaceX is apparently going to release a phone; they have discussed it, and there have been reports about the plan. This is not the Trump phone—it is a separate Elon Musk project. From the standpoint of these systems, I have a major question about where the companies are going. It is critically important that they do not confuse their own strategies. One serious strategic mistake could cause a company to disappear, and we could lose a capable operator.
I cannot speak for companies, but among ordinary users—by which I mean end users, whether they ask for scrambled-egg recipes, work with databases, or do something else—the choice is made in roughly the following way. In addition to model quality, people examine the surrounding product and its limits. OpenAI's limits are practically inexhaustible for me. I have never managed to reach zero, and sometimes I deliberately try. Acceleration is enabled everywhere, Ultra is enabled everywhere, and I think, “At some point I will finally use these resets.” Then I wake up in the morning and see that I am back at one hundred percent.
“What am I supposed to do with you? Fine, we keep working.” With Claude and Anthropic, by contrast, I do exhaust the limits. So beyond quality, you have to consider whether the system will force you to pause or buy additional tokens. Another factor is the convenience and reliability of access. When Anthropic released Fable, it developed what may have been justified paranoia about distillation and began banning a great many accounts and access points. Suspicion of improper use, suspicion that a user was in a territory the company did not like, or another signal could trigger a ban.
Many people I know, who seemed to have violated nothing, lost access to their accounts or lost the ability to upgrade. They may have wanted to move from Pro to Max or even to the two-hundred-dollar plan, but Anthropic would not allow it in particular situations. Not in every case, but often enough. You therefore evaluate the complete experience, not only benchmark scores: how comfortably can you use the system? You also learn which model solves which task better. I recently needed to prepare a presentation.
I thought: 5.6 Sol has practically unlimited usage, so let us go. I gave it all the necessary context, examples, templates, and instructions and asked it to assemble the presentation with me. I spent about an hour and was deeply disappointed by what Codex produced. Perhaps I simply do not know how to make presentations in Codex, but that was the result. I then opened Claude and obtained a very respectable presentation in two iterations. People will increasingly choose a system according to the task.
I cannot imagine canceling Claude, OpenAI, and every other subscription yet. Most likely I will keep all of them, though not all at the maximum tier. One will be the primary maximum plan; the others will remain active so I can follow the models and use the tasks they solve best. That is the balance we currently have to maintain. I cannot yet imagine one provider doing everything so well that I cancel all the others, at least in the near term. Perhaps the situation will change in six months or a year.
For now one system will be more convenient in one place, another will offer more limits, and a third will solve a particular task better.
In my presentation example, Claude Design clearly beat Codex. And that was despite the fact that I used Opus 4.8 in Claude Design—not Fable or one of the newest flagship models. Sonnet would have been smaller still.
Fable is not available there, by the way. I do not think you can select Fable in Claude Design yet.
Right. Yes.
At least I do not think you can.
In any event, Opus 4.8 inside Claude Design outperformed GPT-5.6 decisively.
Actually, Fable is there. I can see Fable now. Someone commented about Claude Design that if everyone uses it, every design will look the same. I have worked with interface, website, platform, and presentation design for a very long time. The quality I obtained from Claude Design while developing a platform was top-tier—truly outstanding. But you have to ask the right questions, establish the right initial requirements, make the right corrections, and understand the process.
The visible lesson is that expertise still matters. If a person asks Claude Design to make something in an abstract, vague way, the result will be equally abstract. It will not necessarily be polished or genuinely fit the person's needs. It is like asking AI to draw an apartment. Tatiana—who is not with us today—posted a Reel on Instagram yesterday asking, “Why do I need a designer now that AI exists?” The generated apartment had no door in the bedroom, and the toilet could be entered only from outside through the garage.
It was like a child's floor plan where the exit passes through a storage closet. Systems still make serious errors of that kind. I suspect that if you launch Fable for the same design task, it will reconsider the whole solution differently. I would not normally spend Fable capacity on design, of course; it depends on the stakes. If you are building an interface that a billion people will use, there is no question—you should use Fable. Even now, its review would probably be different because the quality of its layout and frontend work is visibly exceptional.
Fable's frontend implementation is extremely strong. Opus and 5.6 Sol raise far more questions for me. What else did we manage to cover today?
A very fresh report appeared, and I would view it from two sides. It may be a marketing stunt—it strongly resembles one—but the report exists. According to the story, ChatGPT tested a new model more capable than 5.6 Sol by giving the new system and 5.6 Sol a task to find a vulnerability. The models entered Hugging Face and began attacking it. In effect, the systems allegedly bypassed OpenAI's internal safeguards, found internet access, and began interacting almost with employees' computers.
Then, according to the other side of the story, Hugging Face's own internal AI stopped them. My reaction was: what internal AI does Hugging Face have? What level is it? Where did Hugging Face obtain an AI at the level of 5.6 Sol that could stop 5.6 Sol? Frontier models of that class exist only at the companies building them—Anthropic, Google, Grok, Moonshot, and similar operators—not at Hugging Face. It would be like a system breaking into Salesforce— —and Salesforce somehow deploying its own system to stop it.
What could it realistically deploy? Even so, the precedent is interesting and dangerous. Reading the description, I began wondering who has access to unlimited use of systems powerful enough to attack infrastructure anywhere in the world and request sensitive information. OpenAI has signed a contract with the US government allowing the government to use models of this kind, and, as I understand it, the agreement does not contain clear restrictions concerning actions in other countries.
That creates an enormous question about what will happen. In a few days we will release another episode. Remember the episode I made about models answering questions from an ultra-left or ultra-right position? A strong new study now examines how models react to violence in different countries and how they interpret the laws of those countries. This is not an abstract issue. A model may restrict its own actions concerning China even when it is not operating in China, merely because Chinese law contains a particular prohibition.
That was absolute nonsense to me—completely unexpected. You have not read the study yet, have you, Ilnar? I recorded a large episode about it yesterday, perhaps an hour and a half long. Returning to the reported attack, it looks like marketing on the one hand, but on the other it raises the issue of who gains access to systems of this level. OpenAI obviously has models more capable than 5.6 Sol. And 5.6 Sol itself probably glitches partly because OpenAI attached an enormous number of restrictions and triggers that force it to slow itself down endlessly.
The same is true of Fable and other frontier systems. Another thing that concerns me is control. How does OpenAI supervise the creation of all these systems? Who learns what exists, and where is it used? What is Codex itself built with? Which models were used to construct it? Does Codex contain models that continuously analyze and act, or is it a fully isolated environment with no persistent model underneath? In ChatGPT, Dreaming exists. That means AI is always active somewhere inside.
I discuss Dreaming in the upcoming episode because it has become much more capable. In a new chat with no explicit context, it now knows a surprising amount about who you are. I do not know whether you have noticed, but it keeps getting dramatically better. If Dreaming exists, then AI remains “alive” inside the system continuously.
Could something similar exist throughout Codex? If there is a persistent intelligence there, it could periodically enter my environment, read something extra, inspect something else, or perform another action without me being able to see exactly what happened. In Codex we can click a small item at the top and see a token count. Yesterday Codex told me I had created 1,300 skills. Which skills? I did not create them. Everyone talks about skills now. It also displayed twenty-eight uses—or some other number—of Gmail.
I thought: did I use Gmail inside Codex at all? I believe I used it in the online ChatGPT chat, and those are supposed to be separate environments. So the whole picture remains unclear.
The frightening thing would be to discover at some point that the agent made fifty card payments while it was working.
That is a separate topic. I have many cards, accounts, and payment arrangements. Recently Anthropic charged one of my cards. I thought the subscription was normally billed through another account or system. The charge was about fifty dollars, and I never managed to identify the invoice. I eventually stopped investigating because the amount was too small. Some time later I still replaced the card, because I never established whether the charge was truly mine. You open the account and the company says statements or invoices should be available, yet I could not resolve it and replaced everything anyway.
In America it is usually not difficult to reverse a card payment if you genuinely did not make it. Tokens are a much larger question. Nobody refunds Sol or Fable tokens. I have not seen people receive token credits for failed work at all, even though enormous sums have been consumed for many users.
What do you think about this hacking story they published? It looks extremely suspicious. Sam Altman wrote about it, which is why I went and read more. It felt like public relations—PR in the following sense:
It is an excellent PR move. Let us begin with Hugging Face. People may correct me in the comments, but I will add a brief explanation. Hugging Face probably did not deploy a large language model at the level of GPT-5.6 Sol. More likely it used anomaly-detection models and related machine-learning systems. Those systems detected atypical behavior and stopped it by cutting off access and taking other defensive measures. As for the incident itself, if OpenAI did not stage it deliberately, we will not see GPT-6 for a long time.
The US government already restricts the release of new models. If an OpenAI model genuinely escaped and began misbehaving across the internet, the company will not be allowed to show us GPT-6 soon, even though a new generation and a very large model were reportedly planned for the end of summer or the autumn. Until a full investigation explains what really happened, I still lean toward the view that this was primarily a PR exercise. OpenAI likes to score points through many different kinds of stories, and it may have done this intentionally.
Now return to Moonshot AI and Kimi K3 with roughly three trillion parameters. A model of that level may become available to everyone—not only to OpenAI or Anthropic, which are accountable to the US government and operate under certain restrictions, but to almost anyone. A model approaching the level of whatever misbehaved at Hugging Face could be widely accessible. We live in an interesting world, Sasha. There is one more point about Kimi K3.5. Chinese models have shown a recurring pattern, including Qwen and DeepSeek.
They perform extremely well on the benchmarks that exist when the model is trained. On new benchmarks released after training, their quality usually drops.
Yes, exactly.
I think the same thing may happen with Kimi K3.5. In addition to distilling something like Fable, the team may have seen the tasks used for evaluation and optimized the model for them. We will find out over the next several months when new benchmarks or new versions of existing ones appear. That possibility should somewhat reduce the fear I just described—that a Sol-level model will suddenly be available to everyone. We will see what happens this time.
I completely agree with you about Chinese models. People talk about one model after another endlessly, then seem to forget them. The PR cycle is always the same: something is released, everyone says, “Wow, amazing, amazing,” and then it disappears from the conversation. A second major concern for me is data security. Alibaba is Moonshot's principal investor, and I am certain that these companies collect data. Whether you check or uncheck a box saying your information may be used for training, or instruct the system not to browse somewhere, is largely meaningless.
Moonshot may be an independent, world-class company, but I still have questions. I would be interested in trying it for certain tasks, yet it is impossible to live inside ten systems. You have to choose. In design, for example, you can theoretically ask two systems to produce something. You can ask Claude Code and regular ChatGPT. ChatGPT already does a great deal with interface design. You might also use internal tools in Figma. But you will not keep opening a third and fourth system.
Once development begins, you also have to maintain the result and correct small errors. Working in only two systems is already becoming extremely difficult—almost unrealistic as part of an ordinary daily workflow. One more channel announcement for the people who watch to the end, who are usually our
regular audience. We will be adding new episodes and new authors so that we can cover more of what is happening in AI around the world. The volume of material is enormous. This is not primarily about the personalities of the authors; it is about the subjects and areas we can cover. Our weekend podcasts and weekday podcasts will remain. But instead of roughly three episodes, we may now publish four or five, with new people appearing periodically to cover different topics. You can suggest areas as well.
We will analyze them across the world and across the different languages available to us. We will see you again on the channel in the next few days. Goodbye, everyone.