Skip to content
Transcript

Transcript · 154 · OpenAI Has Paused Its New Model. What Is Happening to the AI Race? — ToTheMoon

English machine translation of the Russian-language episode. Timecodes open the source video.

Episode overview
00:00:00–00:00:54In This Episode of ToTheMoon
Alexander Volchek00:00:00

Something very strange has been happening in the artificial intelligence race these past months. We are used to the idea that a new model comes out, so it should be better than the previous one. But today, on real tasks, it is increasingly the other way round. Older versions turn out to be steadier and sometimes cope better than the new ones. Though I am of course not talking about Claude Fable. And right at this moment OpenAI is pausing the training of its next model, Astra — the very best, the most advanced — amid questions about its behaviour and safety.

Alexander Volchek00:00:31

Is this simply a temporary technical problem? Is it marketing? Or has the AI race itself begun to change? Why might the winner today be someone who does not have the newest and most powerful model at all? That is what we will work through today. Hello everyone!

00:00:54–00:04:00OpenAI Has Paused Its Next Frontier Model
Alexander Volchek00:01:03

You are on the ToTheMoon channel — technology news and insights from Silicon Valley and around the world. So, Ilnar, listen, what do you make of the story that OpenAI has stopped, stopped the development of frontier models — the ones meant by advanced models — and specifically the training of those advanced models. That is, what they describe, what is it? Is it that they have paused their most current models in terms of their own training? And that is roughly how it is described in plain language.

Alexander Volchek00:01:43

And it was written in Sam Altman's statement that within the Atlas model, I think — no, Astra. Astra.

Mentions: Astra
Ilnar Shafigullin00:01:44

Astra, yes.

Mentions: Astra
Alexander Volchek00:01:45

Within the Astra model. This is the next generation of model after five-six, after the Sol model. Presumably the question here is not five-six but Sol. The next generation of model, and supposedly this model is uncontrollable, that it may carry an enormous number of risks in terms of cybercrime or security, or any kind of crime. But in general, when Sam Altman wrote about this, I thought: why write about it at all? That is, imagine you build models, and it is logical that you encounter various circumstances, various difficulties, various uncontrollable things.

Mentions: Astra · Sam Altman
Alexander Volchek00:02:25

And you run these tests internally and various uncontrollable things happen. Why tell people about it? There are useful things to tell. Anthropic ran a very large study on how agents work at scale. That, by the way, is a very interesting subject — how agents come to terms with each other or how they take each other out. I talked about it a few days ago in an episode. If anyone has not watched it, do. But here the story is different. They come out and talk about a model just as Anthropic is heading for a public listing, an IPO, a very expensive IPO.

Alexander Volchek00:02:56

Anthropic is showing higher revenue than ever. They say revenue in the coming years could reach two hundred and eighty billion dollars, three hundred billion a year. And is this not a certain pitch from OpenAI's side, that OpenAI supposedly has some supernatural, some incredible model? Remember, I kept saying for some reason that even the previous version with Hugging Face looked like marketing to me. It seems not to be marketing, because Hugging Face complained to the FBI, but it still looks to me as though it were somehow staged, because I do not see in 5.6 Sol the quality that Fable has.

Mentions: OpenAI
Alexander Volchek00:03:41

Well, kill me, I do not see it. Well, I am of course by no means a person who works on security systems or serious things like that, but to me Fable is on an entirely different level from 5.6 Sol. And this pitch of theirs about Astra — is it their last resort, or have they really made something incredibly good?

Mentions: Astra
00:04:00–00:07:30Astra: Safety Problems or Marketing?
Ilnar Shafigullin00:04:10

Listen, every new model, if we take the big transitions, gave a substantial gain across many measures. And accordingly, here too you can look at it in two aspects. On the one hand, right now, because of this whole business between the large companies and the US government, they have their own constraints — they cannot release a model that is too free, too strong, it will be locked down anyway. And here, on the one hand, you can read it as: our model is already so good that improving it further is unnecessary.

Ilnar Shafigullin00:04:53

We will now fine-tune it on certain parameters, because if we make it even stronger we will not get through the constraints we have been given. The second story. They published a safety paper, and after all these incidents that occurred they began spending a lot of effort on tracking what their agents actually do. Because many problems arise not when an agent is released into the free world — there is not that much activity there — but when they start testing it. Giving it tasks.

Ilnar Shafigullin00:05:28

And carrying out those tasks, the agents as a side effect start, notionally, misbehaving with Hugging Face, or other things happen. At least at the training stage inside OpenAI, which is what is documented now. So accordingly OpenAI began tracking not only the actions the agent performs while solving — that is function calling, which tools it invokes — but also the reasoning itself, the reasoning trace, roughly how the model thinks when it takes actions. Accordingly you end up with an enormous volume of data that has to be tracked on top.

Mentions: OpenAI
Ilnar Shafigullin00:06:04

And by my estimates, up to twenty or twenty-five per cent of the compute that goes into training now starts going into analysing these traces, these sequences of computation. And accordingly, if some trigger fires — the model is behaving in a way we would not want — and within thirty minutes it cannot be resolved by automated means, then the entire process halts until we work it out by hand. Accordingly, this two-week halt, if you follow that line of reasoning, may be a consequence of their simply not being able to work out how to cure the behaviour the models started showing.

Ilnar Shafigullin00:06:45

And I will agree with you that in any situation Sam Altman and OpenAI's PR still want to get some benefit out of it. And this news too can be read in two aspects. On the one hand they are doing something uncontrollable, which they cannot control themselves, so they are stopping. And that counts against the company. On the other hand they are saying: «We have something so good, we are just going to finish it off, but it is a mad story. We even stopped training simply because what we have already built is very strong.» And in that way they try to score some points too. So it is hard to say one way or the other, because there are several aspects to it. But the main story is that GPT-6 is most likely being put off a little.

Mentions: OpenAI · Sam Altman
00:07:30–00:08:30ChatGPT 6 — Not Coming?
Ilnar Shafigullin00:07:33

There was news that possibly by the end of August we would have the next numbered version. And it was understood that Astra would be ready by then, and presumably it would have been ready had there not been the problems they ran into. But most likely it slips. What is more, with all these stories, with the Hugging Face break-in and so on, I think it will not be so easy for them to pass the checks in the US government. So most likely the story gets put off a little. Also by the end of August, give or take, the feeling was that Claude, that Anthropic, would also have a next version.

Mentions: Astra · Claude
Ilnar Shafigullin00:08:18

But so far I have not heard further news on that. Possibly that story is being put off a bit as well. I hope that at the start of the autumn we will be treated to some new models.

00:08:30–00:12:17Opus 5, Fable and 5.6 Sol: What Changed
Alexander Volchek00:08:30

Well, there is an interesting situation here now: on the one hand we are watching a race, and we have always watched this race of model updates. But what has happened, it seems to me, over the past two months is fundamental. That is, ChatGPT 5.5 appeared with this Extra High mode, Opus 4.8 Max appeared. I remember that being a real peak, when you went in, started doing some tasks and saw proper study of the architecture, large analyses, development of serious blocks of code or a serious end result.

Mentions: ChatGPT
Alexander Volchek00:09:09

Among simple applications you could get a serious end result. That is, you look at it and think: a really good, strong result. Then what happened? Fable appeared, and Fable in absolutely every aspect — in my view, and in everyone's, and in the descriptions, and in everything — really did seem to bring something new. And not even new in terms of duration of work and the quality of holding to a single path. It seems to me it brought a fundamental quality of result. That is, it saw the result in a completely different way.

Alexander Volchek00:09:43

It had some different speed, as if before a fifteen-year-old had been assessing the task. And now an adult, a grown person came along and seriously implemented the task. Really — I don't know, did you feel that? I really, really felt it. And then came…

Ilnar Shafigullin00:10:00

Sasha, yes — what was called a PhD in everyone's pocket, right? That is, we have come closer to that.

Alexander Volchek00:10:05

Although people have been saying PhD for a long time. And then what happened? Opus 5 appeared — I am talking about large systems now, mind you, about large ones, on the maximum plans, because I am not even talking about simple systems, I do not even consider simple ones. Opus 5 appeared at Claude, and this new family of models appeared — Luna, Sol, Terra at ChatGPT in Codex. And something went wrong entirely. That is, on the one hand they seem to say: «We are going to do a load of great things here.» And you seem to see something.

Mentions: ChatGPT · Claude · Codex
Alexander Volchek00:10:37

It even looked as though it were 5.5, only one that can keep working, one that does not stop. But at the same time endless glitches came pouring in. And what do I see now? It seems to me there is a degree of degradation. That previously 4.8 Opus effectively solved tasks better than Opus 5 solves them now. Simply better. Whatever you do. In my own tests I think I still go back in many places to 4.8 if there is no Fable. And it does tasks better than Opus 5 often. Opus 5 goes off who knows where. Take ChatGPT, not Codex, the chat one: its development on the whole went along fairly well and steadily.

Mentions: ChatGPT · Codex
Alexander Volchek00:11:22

There were moments occasionally when the pro version had some glitches, but on the whole it all went normally. Recently, though, the 5.6 Sol pro version, first of all, keeps resetting the slider. To be honest, I think this is simply — well, this is a mark against everyone, against all the management, against everyone who does this, it is simply impossible. That is, it resets everywhere. This is the cry, the cry, the cry of Volchek. So, this is complete unprofessionalism, of course, absolute unprofessionalism, in development in particular.

Alexander Volchek00:12:04

But what I see is that the 5.6 pro version has sometimes started producing radically strange things too. Although at the same time ChatGPT still holds up en masse somehow. And what would one not want to be happening right now?

Mentions: ChatGPT
00:12:17–00:14:08Has OpenAI Lost Track of Its Own Strategy?
Alexander Volchek00:12:19

Is there not a sense that they have lost the strategic understanding of where they are taking these systems, what they genuinely want to build? Do they want to build AGI, or do they want to build systems for coding, or do they want to build chats used in businesses? That is, what do they want to do? Because Sam Altman, the day before yesterday in one of his interviews — they had a partnership with a company that opened a free service for developing software on the basis of the Terra version.

Mentions: Sam Altman
Alexander Volchek00:12:47

And I sit there, and Sam Altman says: «How great! Everyone will now have a free opportunity to make software.» And I sit and think: what is the point of that software if ChatGPT has an endlessly free Terra version by default? And why build these separate systems? That is, why are they needed? And effectively any person at all can open Codex and start doing something there. And Sam Altman also said a phrase — our viewers, by the way, think about this too — that everyone will be able to be an entrepreneur and everyone will be able to try themselves as an entrepreneur, to create something.

Mentions: Sam Altman · ChatGPT · Codex
Alexander Volchek00:13:21

I thought: listen, if you want to try being an entrepreneur, first of all you have a free version, but beyond that you can buy the ordinary Plus version for twenty dollars. Because if you want to be an entrepreneur you ought to have twenty dollars. You buy the ordinary Plus version, and you get Codex with a fairly decent token budget to actually do something. And what is this parasitism, exactly? All this reasoning about the mass market — or was he saying it, again, to market that system, so that the system would start being worth a lot and OpenAI would then get, I don't know, some shares out of it.

Mentions: Codex
Alexander Volchek00:14:01

Or maybe Sam Altman has shares in that company. There is a great deal of this going on right now. There are decent deals, good ones. Grok closed the deal with Cursor. But it is not actually Grok but SpaceX, which owns xAI, which owns — whose product is Grok.

Mentions: Sam Altman
00:14:08–00:15:15xAI, SpaceX and Cursor: A New Fight for the Programming Market
Alexander Volchek00:14:20

It closed the deal. By the way, we could sit here like this. It, it, it closed. Well, since it did, correctly, it closed the deal, closed the deal. We actually talked about it two weeks ago, I think. And what is notable about this deal? It is very clear and obvious that xAI too will be entering the programming market. A hundred per cent. And entering it very seriously. Because Cursor is a bloody serious system. And combined with Grok they will now, of course, have to start competing properly with Codex and with Claude. And on top of that Google is not retreating.

Mentions: Claude · Codex
Alexander Volchek00:15:12

Google after all has Gemini, and they released 3.7 Flash precisely for programming and agent tasks. And Google is not retreating from programming. So what is going on, in terms of strategies in particular?

00:15:15–00:16:38How Anthropic's Strategy Differs From OpenAI's
Alexander Volchek00:15:21

Anthropic's strategy seems to be what it was and has stayed that way. Judging by what they do, what research and essays they release, what constitutions they write — all of it. Whereas ChatGPT and OpenAI genuinely have the world market of all users, whom they forget about, because they cannot manage a damned slider. In quotation marks. And meanwhile they are making announcements about Astra.

Mentions: OpenAI · ChatGPT · Astra
Ilnar Shafigullin00:15:51

Well, look, there is both a curse and a blessing in their making a very, very universal product. That is, it is not some tool that solves one specific task. Yes, you can apply it in many places: in education, in programming, in business, in creative fields, anywhere. Anthropic set a fairly concrete goal and are moving towards that goal. They have their main direction, connected with AGI safety, and along the way questions of programming get solved as well. Which is where a lot of engineers went at one point from ChatGPT. Whereas OpenAI is trying to move in every direction at once.

Mentions: OpenAI · ChatGPT
Ilnar Shafigullin00:16:36

And especially when you told the story about creating software with Terra.

00:16:38–00:18:31Ways to Make AI Stronger Without a New Model
Ilnar Shafigullin00:16:42

What came to mind here — I have been thinking about it for a while. All the time we have been recording this podcast, there has always been a base level at which the models work, and then all sorts of tricks by which you can make them work better. What do I mean? At the very beginning you had to write a good prompt. There were large guidelines on how to write a prompt. And it began: you are an expert in such-and-such a field, write as if it were this — and that actually meant something.

Ilnar Shafigullin00:17:17

That is, your ordinary prompt and an elaborate one differed in result. Then, accordingly, there appeared the ability to connect additional services over an API. That was fairly hard to do. At the base level it was not there, but advanced users could and did get benefits from it. And then came the whole orchestration story at OpenAI, and all of it was eaten up by the large companies. That is, now it does not matter what prompt you write, you can hit the keyboard with your left heel and get something at least distantly resembling what you want to do.

Ilnar Shafigullin00:17:56

It will work it out, it will work it out, it will do it. Integrations with all sorts of services are already there inside, out of the box. Orchestration of a huge number of agents is inside too, available out of the box. You do not have to do anything for it. And, well, why did I recall this? You say there is some service under whose hood OpenAI holds Terra, and with it you can generate software. Meanwhile any person can already do this now. It will become even simpler in the next iteration — in a month, in two, in three it will be completely simple.

Mentions: OpenAI
Ilnar Shafigullin00:18:28

That is, the big companies will eat it anyway. The current thing everyone is running around with is skills.

00:18:31–00:20:34Skills: New Hype or a Temporary Technology?
Ilnar Shafigullin00:18:34

There are a great many different skills — from, I don't know, Karpathy, from Meta, I keep forgetting the surname, we will add it in the comments. For example there is a skill, Grill Me, that helps you work through a subject you want to work with further. That is, it asks you all the questions, discusses, gives recommendations, and together you build the full map of the project you need. And an enormous number of other skills. On the one hand you can get into this, you can dive into it and be, notionally, a month ahead of everyone else.

Ilnar Shafigullin00:19:08

But all these things will be integrated inside the same Claude, the same ChatGPT, and will be available out of the box to all the other users. That is one point. There is always some movement forward that then gets eaten by the large companies — with the creation of software using this Terra company. If it suddenly is not Sam Altman's project, it will be absolutely the same. And the second point, about what awaits us further, what models come next and so on. There are two main routes to improvement. One route is to improve the base model. That is, we wait for GPT-6, we wait for Fable-6 or some other models.

Mentions: ChatGPT · Claude · Sam Altman
Ilnar Shafigullin00:19:51

And the other story is to squeeze more out of the current model's existing capabilities. Before that it was: we added MCP, and we got additional context. We added the ability to split a task across sub-agents. There too the situation changed dramatically, even though the base model under the hood may not have changed much. And it is the same here. If the whole story with a sixth ChatGPT and other such models keeps being put off, companies will find how to squeeze more out of the current models.

Mentions: ChatGPT
Ilnar Shafigullin00:20:21

What is more, the market is already doing it, and most likely all of it will simply be integrated further. That is my little monologue; I hope I have not confused anyone with the thought I wanted to share.

Alexander Volchek00:20:33

Listen, there is a very interesting aspect here: effectively the market does not now demand that anyone hurry to develop some super-models.

00:20:34–00:22:32Does the Market Even Need to Hurry With GPT-6?
Alexander Volchek00:20:43

They can develop strong models in parallel and improve the current work. Not just move a button somewhere or fix some micro-glitch, but genuinely improve the work — the linear structure for managing projects or chats that has not changed in six months, we can see it. A great deal is inconvenient, or some names are wrongly labelled. Or periodically there are glitches in the systems and you do not understand whether there was a failure or not. Or in Claude you open the page and it tells you how many tokens you have spent over your lifetime; you click «this week» and it gives you zeroes.

Mentions: Claude
Alexander Volchek00:21:20

That is, it does not give you some of the data. That is, an enormous number of these micro-glitches, which it seems to me affect people's everyday life a great deal, and affect whether a person becomes more attached to the system, genuinely attached, not wanting to migrate off it. Because I was the strongest advocate of OpenAI and ChatGPT in terms of daily activity, and I remain one. But in terms of programming, obviously, for a month now eighty per cent of my work has been happening in Claude rather than in ChatGPT.

Mentions: OpenAI · ChatGPT · Claude
Alexander Volchek00:21:52

I continue to use Codex a lot for various tasks. I like how it does a number of things — analysis in particular, downloading data from various sources, analysing data, all sorts of audio, texts, some microservices that exist inside. Like the nonsense I was doing, recognising people and animals on the camera by the house, by the gate, or some of my own financial things. That is, a whole set of things I use it for, but still the big, most fundamental things I have started doing in Claude Fable.

Mentions: Codex
Alexander Volchek00:22:28

That is, I have effectively migrated to it. And the fact that ChatGPT released a new mode, as they describe it — OpenAI released a new mode, 5.6 Sol Ultra Fast — and you think: well, folks, what is the point?

Mentions: OpenAI · ChatGPT · Claude
00:22:32–00:24:04Why OpenAI Made an Ultra Fast Mode
Alexander Volchek00:22:44

That is, first of all, 5.6 Sol is a very expensive model, and the Ultra model is very expensive, terrifyingly expensive. I would be genuinely surprised if via the API an enormous number of people and companies in the world use it. I mean en masse, everywhere, nothing but it. And they go and make an Ultra Fast mode. It seems to me that if you use ChatGPT 5.6 Sol Ultra or even Extra High, you are in no hurry at all. You will say: «Never mind, I will wait.» Right? As with Fable, presumably — I do not want to pay one and a half times the cost in Fable.

Mentions: ChatGPT
Alexander Volchek00:23:22

If I am paying, I don't know, a thousand dollars per request, I can wait for five hundred. So who is it who cannot wait, among people for whom these systems are free? Well, presumably Anthropic itself or OpenAI can do it. And my question is: why are they releasing this? It is like the question of why they released ChatGPT Pro at a hundred dollars. An absolutely strange, well, strange, absolutely strange proposition. Want to make money? Then release a mode for two thousand dollars. Or want to earn more, even more?

Mentions: OpenAI · ChatGPT
Alexander Volchek00:23:51

Then sell the Plus version to ordinary users better. But build more serious engagement for ordinary users. I have a question. Tanya, did you start using Codex, or did it end up being something unexplored for you?

Mentions: Codex
00:24:04–00:27:30Why Ordinary Users Still Do Not Understand Codex
Tatyana Tsvetkova00:24:10

No, Sasha, I did not, because, as we have discussed before, what it offers me were some very basic steps that I had reached myself. Of course it put them together for me in, say, five or ten minutes, whereas for me that probably took several months to get through on my own — but as for developing it further, I could not.

Alexander Volchek00:24:39

Well, here, by the way, a very interesting aspect emerges: you have an old memory of Codex, notionally one of the first versions released, from when people started using it. And you have put a cross against it. You have the feeling that the system still works the same way, and it is unclear what to do with it. Although I want to say that you have a lot of business, professional activity, including with various clients, and Codex in particular helps incredibly well with that.

Mentions: Codex
Alexander Volchek00:25:20

But even where that is absent — my wife, for instance, does not have it — I recently said to my wife: «Listen, there is a neat solution, Codex. You can solve that task of yours with it.» She said to me: «So it is a separate program?» I said: «No, it is the same ChatGPT, just a separate program, a separate application.» And she most likely of course never launched it and never did it, but I simply knew it would be useful for her to run Codex for handling certain recurring things in her mail.

Mentions: ChatGPT · Codex
Alexander Volchek00:25:38

But it turns out people were not shown this, or it was not conveyed to them, or people did not see a simple use case. And it seems to me there is an enormous shortfall in the world among people who are not extremely interested in artificial intelligence — and not many people are interested in artificial intelligence. Even if you look at the number of people who watch our channel or other channels about artificial intelligence, there are not that many, not hundreds of millions.

Alexander Volchek00:26:09

So it turns out that among consumers of ChatGPT, even paying ones — take my wife, she has the paid version at twenty dollars — she does not know what else can be done. It is like the way they never moved the slider and did not know there was a better version. They do not know what else to do. Or, for example, I have a friend, and at one point I advised her to buy a version together with her father. And she says: «I have a version together with my father, I cannot ask certain things there.»

Mentions: ChatGPT
Alexander Volchek00:26:36

But buying it together with your father was something you did two years ago. Now you have long needed your own version. That is, people have this old feeling — she even says to me: «What about you?» Imagine, she says to me! It was, of course, very funny. She says: «Do you have different ChatGPTs, do you each pay twenty dollars?» When I hear that, I understand how people still treat ChatGPT — that twenty dollars. That is, this same girl goes and drinks a ten-dollar coffee easily.

Mentions: ChatGPT
Alexander Volchek00:27:08

And here it is twenty dollars a month for ChatGPT. You cannot possibly have two subscriptions in a family. Well, it sounds — well, it sounds insane, but it is stupid. In fact it is stupid, it is undevelopment, unawareness, and that is all.

Mentions: ChatGPT
00:27:30–00:28:29Will OpenAI Lose?
Alexander Volchek00:27:30

But on the other hand it is the global reality. And so still, what OpenAI is doing now, not progressing much in terms of interaction with the ordinary user — at least that is how it seems to me. We will watch what happens with them and what they manage. I hope — no, there is no chance that OpenAI loses or does not succeed, because it is a healthy, massive organisation already with mad infrastructure, hundreds of billions of dollars invested in outside infrastructure, servers, everything else.

Mentions: OpenAI
Alexander Volchek00:28:09

They will build models, but one would of course like them to build interesting things. And possibly Anthropic will go public, will have a valuation above three trillion, will raise a lot of capital, will not fall in valuation and will capture this consumer market, which it seems to me they badly lack, and fly ahead. By the way, write in, our viewers — and, Ilnar, this is interesting to you. I still see an enormous fundamental difference, for instance, in creating a report or HTML in ChatGPT 5.6 Pro in the chat, or even in Claude Fable.

00:28:29–00:31:38ChatGPT vs Codex
Alexander Volchek00:28:41

Fable, very important. Fable in Max or in Ultra Code, in Claude Code. That is, in Claude Code it still creates something that is, I don't know, simplified, or more code-like — not this interface-like, not visually neat, for some reason. For some reason that is my impression. And that is probably still one aspect of why I have not migrated. I know, Ilnar, you migrated almost entirely to Codex, out of ChatGPT, right? That is precisely why I have not migrated to Codex.

Ilnar Shafigullin00:29:33

For me, Sasha, it has more to do with the fact that the overwhelming majority of tasks and requests I solve with ChatGPT are connected with IT, with programming, with data science. And it is simply more convenient for me to do that in Codex. There are a large number of projects, each connected to some repository. And that is why there is not much point in me going into ChatGPT any more. Occasionally there are some questions, the old-fashioned kind of request, where you want some research done or something like that.

Mentions: ChatGPT · Codex
Ilnar Shafigullin00:30:03

Then yes, I use it, but less and less often, and more and more I simply stay inside Codex. As for reports and presentations, yes, I have worked out a scheme for myself and on the whole it suits me completely. In ChatGPT I gather all the context needed for the presentation. That may be about some project, about experiments conducted, about some other body of work. You gather all the necessary information. It might be several months of work, for instance. Then, accordingly, together with it you break this into meaningful blocks, and you get some slides.

Mentions: Codex · ChatGPT
Ilnar Shafigullin00:30:39

And then, so that it looks good, I do that in Claude Design, because getting ChatGPT to make a beautiful presentation has not worked for me so far, at least. Claude Design cannot gather all that context. You have to give it to it. So accordingly you gather the context in ChatGPT and take it to Claude Design. You also throw Claude Design a reference — either a guide to making presentations that is accepted at, I don't know, a conference, in a company or somewhere else, or existing presentations where you say: make it in a similar way.

Mentions: ChatGPT · Claude
Ilnar Shafigullin00:31:13

That it copes with. So it works out that I gather the context in ChatGPT and then go with that context and with examples to Claude Design. Then literally two or three iterations and you get very, very decent results, which by hand I would — well, I would never have assembled such a presentation. And here it comes out cheap and fast. Usually that does not happen. But well…

Mentions: ChatGPT · Claude
Alexander Volchek00:31:37

And which of our viewers, by the way, uses Claude Design, and what more serious alternatives do they see to it today?

Mentions: Claude
00:31:38–00:36:30Claude Design: Where It Turned Out Stronger Than ChatGPT
Alexander Volchek00:31:45

Genuine alternatives. I can say that I do not yet know of such an alternative. Perhaps someone will say that in Figma, for instance, or in various systems you can do something better, super-professional. But I want to say that the quality specifically of design development of various kinds, with nuances here and there, with some imperfections — but everything to do with applications, websites, sometimes even press graphics… I do not like the export in Claude Design. It seems to me it works very heavily, buggily, not entirely correctly.

Mentions: Claude
Alexander Volchek00:32:32

Not all systems understand what is going on separately there. Also a question for Claude Design, by the way: why do they not do it properly? A genuinely concrete glitch, genuinely stupid, a month and a half. But the quality of the system is incredibly good. And again, my question is: how does the market keep on using third-party solutions? There is a system called Lovely, right? Or how is it? Lovable. What is the right name? For developing mobile…

Mentions: Claude
Ilnar Shafigullin00:32:57

Lovable is what people usually say.

Alexander Volchek00:32:58

Lovable, yes. For developing mobile applications. It has just had a new round and raised a load of money. Its valuation, I think, was over ten billion dollars. And I sit there all the time and say to myself: why should I use this separate system if in Claude Code or in Codex I can do more or less the same thing? And yes, there are some additional features, slightly easier integrations. But still, if you want to genuinely develop something further and look a little wider but far more fundamentally, it is more correct to do it straight away in Claude Code or in Codex, not there.

Mentions: Claude · Codex
Alexander Volchek00:33:40

And, by the way, what I miss with Claude Design is for Claude Design to be as if built into Claude Code, so that I could ask for it right there inside or carry on working with it further. Because when you do an export, it then loses something, it does not have those files. And again the question arises: why do you create products like this? That is, if you want this end-to-end integration, it is obvious that the person of the future — even the current, the present one — if a person makes a website, the most mass-market design of that kind, websites, for instance.

Mentions: Claude
Alexander Volchek00:34:07

And if you make a website, it all has to be done in one place, at one point. And when you have one connected point, you will not, for instance, be passing Claude Design into Codex. Right now I can pass Claude Design into Codex, which I sometimes do, or I can pass it into Claude Code. But if it were built into Claude Code, then Claude Code would tie me to itself even more. Just as the appearance of browsers in Claude Code and in Codex removed the problem of having to look in a browser yourself one more time, of having to go somewhere one more time.

Mentions: Claude · Codex
Alexander Volchek00:34:47

That is, having internal browsers is an incredibly, mega good, super-fashionable thing. What other systems do you know? You have got hooked on Claude Design too, then, since you bring it up? Because I have been an advocate of it from the very start.

Mentions: Claude
Ilnar Shafigullin00:35:04

It was for visuals like that that I did not touch it for a long time; I went to it when I could not shake normal presentations out of ChatGPT. I started looking at what else could do it. And I had a Claude subscription anyway. I thought, let me try. And I was strongly surprised at how much better the result was in Claude Design than in, say, five-six Ultra with all the bells and whistles you can layer on there. Yes, people wrote to you in the comments, and I think you even said it yourself, that doing design in Claude Design gives you all the same design.

Mentions: ChatGPT · Claude
Ilnar Shafigullin00:35:39

But here you only have to remember that designers who study at the same university, in the same faculty, in the same group will all do it differently, because their backgrounds are different, their briefs are different — depending on the input, that is, on what data you give at the input, your output will differ greatly. And here too, depending on what references you give, what prompt — however that sounds in 2026 — you set, what internal settings you have, what general context and so on, the output will differ greatly.

Ilnar Shafigullin00:36:14

So here, if anyone is worried, there is no reason to worry, you can always simply give a reference and say: this is what I want.

Alexander Volchek00:36:22

You mean in terms of designs looking the same?

Ilnar Shafigullin00:36:24

Yes, yes, yes. Well, there were more concerns in the comments.

Alexander Volchek00:36:28

I see that it differs very strongly, and, well, simply incredibly strongly right now, in particular in what it does not do.

00:36:30–00:54:37The New ToTheMoon System, Built With Claude
Alexander Volchek00:36:38

We have, by the way, already made the full version of ToTheMoon — that is, all of our episodes — with Claude Design, and all the cards. Well, this is rather a bit of promotion, because people watch our episode on YouTube anyway, and it is for geeks and for people who like us a lot. But what is the plus there? There are, I think, more than twenty-five thousand different links tying together topics, tags, companies, everything. And the search works cosmically fast, and you can look at any moment in time when we mention someone, where we mention them, how we mention them.

Mentions: Claude
Alexander Volchek00:37:11

In short, the system works in a fairly interesting way in terms of search, which is fundamentally interesting. And what do I want to say? This was made — I will open it now and show you. This was made precisely with the help of Claude Design. That is, all of our episodes entirely. It is currently in an unreleased, notionally test version; it will be released soon. And here every episode of ours has markup, has all the companies that appear in each episode. And here, for instance, is the episode of 9 August. And here is who is in the episode, who is at the centre of attention, who provides the additional threads, what the main lines of conversation were.

Mentions: Claude
Alexander Volchek00:38:07

Here, for example, you can open the OpenAI company card and see that OpenAI was mentioned in a hundred and forty-eight episodes, ninety-eight of them in a central role, twenty-two OpenAI products and models, and that mentions start in April 2024 and run to August 2026. Or we have search, our search works. We had a load of mentions of Belarus, right? And you can even see that it was never at the centre of an episode. And that is neat. Here, for example, is India. If you open India, India was at the centre of an episode once. Remember, we did an episode with you in May 2025 about education.

Alexander Volchek00:38:44

And you can see that substantial coverage, for instance, of India happened several times. And obviously there are interconnected things, very interesting ones. That is, when you understand, in particular, that here the system understands that France and Mistral, for instance, are connected things. Well, there are a lot of neat constructions, very interesting constructions. And it works. We mentioned Mistral in nineteen episodes, imagine. If anyone, by the way, says — our people love to say that we do not cover someone. So if anyone thinks we do not cover someone, there it is, you see immediately, both France and everywhere it appeared.

Alexander Volchek00:39:31

And here is all the coverage, and right down to the minute you can see in what context and when, and what was mentioned, and what was where. Our transcripts are here, obviously, all of them in full. So if anyone thinks we did not mention someone, remember, that we short-changed some system. And people always said, by the way, that we do not mention Anthropic. Anthropic. Anthropic. Well, Anthropic is clearly less, right? Twenty-five times at the centre of an episode, unlike OpenAI.

Alexander Volchek00:40:06

But still Anthropic, look, has been present across all our episodes throughout. Look what a good thing it made. It made the notion of entities connected with Anthropic. See? And you can even go in and look at when Claude Design was at the centre of an episode, or when Anthropic was at the centre of an episode. It is fairly serious. What matters? This design was done by Claude Design. And the neatest thing with Claude Design, for instance, was when I said to it: «Not the cards.

Mentions: OpenAI · Claude
Alexander Volchek00:40:39

It is one thing for you to make these cards. It was another thing for you to make…» Well, there is a great deal here. Even, look, it knows about the software company Atlassian, or here we had McDonald's, with two mentions in a central role. We did indeed recently, as you know, have a key breakdown on McDonald's. I recorded it. It was on 7 August. Do you see McDonald's? There it is. So, it was one thing to make cards like this. I want to say that the design itself looks incredible. Well, in terms of, in terms of development speed.

Alexander Volchek00:41:17

On the other hand, Ilnar, it made me — it proposed to me how this search here should look, how the search should look visually. And on the first attempt it wrote the architecture for such a search, on the first attempt. And then Claude Fable implemented it for me on the first attempt. And implemented it with an incredibly high speed. I want to say that the speed at which this search works is very good. I have probably made hundreds of different sites or projects and various systems.

Alexander Volchek00:42:06

But the ease of creating a system like this across our episodes — there may not be a supernatural number of them, but there are still a lot. And here we have coverage of thirty-two thousand different intersections of tags, labels, words, eight hundred and ninety-two entities, a hundred and fifty-three episodes. That is fairly serious coverage, very serious coverage. This does not appear here yet; it will appear on the site, because there is a site that relates — not this one, but a site that relates to…

Alexander Volchek00:42:27

The company Solid. That is where my personal projects live. ToTheMoon is simply a participant, a connecting link between those projects. It will appear separately within Solid, but it has its own style and its own design. Ilnar, note by the way that this design differs quite strongly from this one. That is, it differs absolutely strongly in every aspect but has certain connecting things. And Claude Design managed to do all of it fairly simply and easily. The markup was done by Claude Fable, and the development there was done partly by Fable and by Codex.

Alexander Volchek00:43:13

Although the ToTheMoon development, by the way, was done only by Fable. And I want to say that I sometimes realise it is better for me to allocate the resource and do it on Fable than to struggle and move into other systems. Lately, everything to do with markup quality in particular is sometimes astonishing. A system cannot lay out the stupidest, simplest page, an elementary one, the very simplest of all. And sometimes complex-looking cards like these. But you know yourself that this does not look simple.

Alexander Volchek00:43:51

What I am showing you it does incredibly well, precisely, accurately. And there will be a great many other interesting things here. But I think Ilnar and Tanya have not even seen this yet, by the way. And today it is being seen both by our viewers and, at the same moment, by Ilnar and Tanya. I think it will be interesting to both Ilnar and Tanya. If you need to, you can always find, absolutely always find: in which episode did we talk about what? That is, for us as an editorial company that of course matters a great deal.

Alexander Volchek00:44:19

And I am incredibly impressed by how seriously, of course, modern — by how seriously modern models can work. Because in a world of hundreds of millions of sites, or a billion sites, this already looks like something of a platform. Let us be honest. This is not just some micro-site, because even if you go, in particular, to my own site, there is a fairly large knowledge-base structure. Here there is a search of articles, of details, of all sorts of interconnected constructions, but a great deal of interconnected structures, I don't know, paintings with various meditation episodes.

Alexander Volchek00:45:07

In short, there is everything. It looks like a platform. And this platform is of course a platform. And with all that it is still a fairly simple site. That is, it is not some business selling a big platform, an enormous one. And all of it is made with very simple tools. And to all the viewers who shout that you need fierce architects, programmers, incredible designers, layout people, I don't know who else — I would ask a question, because all of this was made by me personally, alone.

Alexander Volchek00:45:43

And, as you know, I have a lot on. And it works, by the way, fully in English. Of course we have everything in English, and we have all the episodes in English, and all the transcriptions in English, and all the episodes in English, and all the connected entities in English, and the descriptions of all the companies entirely in English too. All the topics are in English. Incidentally, right down to the covers or all the details. Well, this is not the topic here, although it seems to me our viewers will find it interesting to see what relates to our ToTheMoon project in particular.

Alexander Volchek00:46:16

Right? But I was showing it here as an example too. A very good example of where such systems are today, and in particular it is a pity that so far people — I do not know whether people will be able to do this easily. Tanya, for instance, has a large number of projects, sites, all of which could be developed incredibly strongly this way. But it turns out that to do this you have to know simultaneously what architecture is, and what systems analysis is, and what business analysis is, and what markup is, and what front-end programming is, and what captchas are and what cookies are, and what user consent is, and what audio transcription is, and so on.

Alexander Volchek00:46:56

Incidentally, by the way, you know that with Codex you can now go in and say to yourself: «Please parse this whole YouTube channel for me, all the episodes with all the covers, with all the descriptions. Connect me an LLM Labs token, recognise all the audio for me completely.» It will then recognise all of it completely and bring it to you as text. And you can build such a model of interconnections yourself. Do you realise what level we have reached? Incredible! By the way, Ilnar, you know there are various systems that recognise voices. And I was so dissatisfied with role recognition in every system, in LLM Labs and in ChatGPT, that in the end I wrote my own local role-assignment system with Codex.

Alexander Volchek00:47:27

And it takes as input, it can take, say, a thousand hours of audio and everything. That is, recognised audio recognised by LLM Labs. LLM Labs still wins on the quality of Russian and wins on the quality of Russian even against the new model. ChatGPT released a new model. They had four-o Transcribe. They released a new Transcribe model, the most current one. It costs less. Very close in price to LLM Labs, but they still lose to it. And meanwhile I wrote my own role-assignment model. So that is neat. We record a great many episodes with you at ToTheMoon, or we record — I record for personal purposes.

Mentions: ChatGPT · Codex
Alexander Volchek00:48:21

I have personal development of the individual. There is spiritual development. Tanya, for instance, takes part in other channels too. Periodically she conducts interviews for me, asks questions. And yesterday this system says there are six hundred and ninety-six unrecognised fragments of people. «Cannot recognise the people.» I say: «To train the model, send me voices in batches, straight into Codex.» And it sends me little five-second, seven-second fragments. And there goes Tanya's voice, there Anya Gurin's voice, there Anya Volchek's voice.

Mentions: Codex
Alexander Volchek00:48:42

Well, different voices go by, Olya's, Sasha Mineev's, a lot of different people. And these voices come through and you go: «Oh, that is this person's voice, this one's, this one's.» And it goes on being fine-tuned. It gets fine-tuned and tells you: «Our level of voice assignment is, for instance, coming out at ninety-nine per cent.» And, clearly, again — to create something like that, I say it is very easy and so on. Ilnar will now laugh and say, as with Karpathy back then: «Still, Volchek, do not — you are not a simple website developer.» But it turns out that I recently recorded an episode here at ToTheMoon, an individual episode, one of them.

Alexander Volchek00:49:07

Not within our Sunday podcast. I was talking about what people will become, what will happen to the programmer. And there I put forward exactly the hypothesis that an enormous number of new people are appearing who will call themselves programmers or be called developers. And which of these people will be able to grow through? Because to create what I have just shown, you have to be all the roles at once, at least in terms of understanding what is happening, and in terms of understanding the business. And at the same time, of course, I am not a designer, far from it, and by no means am I an architect now, and I will not personally write the code myself.

Alexander Volchek00:50:02

So. Ilnar, what are your impressions of the search, by the way?

Ilnar Shafigullin00:50:10

Well, it is neat. A chronicle of AI as seen by ToTheMoon has appeared, right? That is, it will be possible to…

Alexander Volchek00:50:16

Yes, yes, yes, yes.

Ilnar Shafigullin00:50:17

…to look after some time at what we actually had, how we developed and, indeed, how the industry developed through our eyes. I hope the viewers will find this interesting too. As for the complexity — what you are saying now, that to create this you need to command a certain set of competences. But if you rewind a year, two years, a far greater volume of competences was needed for it. Everything is being simplified a great deal. And there is no guarantee that in six months, with the appearance of the next models, the next harnesses, as they are called, wrappers, this will not be solved with a single prompt.

Ilnar Shafigullin00:50:53

You will not have to think about architecture, or markup, or anything else. You make one request and say: here, gather me this YouTube channel. Right now it does not work that way, but I think in the near future we will move in that direction.

Alexander Volchek00:51:12

Yes, maybe it will not even be needed. You know what I want to say, Ilnar? Even doing this, I am not sure it will be needed in the future. Because I am an advocate myself, and I say that websites are becoming unnecessary altogether, that the notion of a site as such has no necessity at all. At the same time I am helping AI models and search engines analyse what we actually talk about at ToTheMoon. Rather than abstractly saying we are some company and about something somewhere — you see?

Alexander Volchek00:51:45

Somewhere, about everything. Even making such an analysis requires spending a lot of compute. And as of today artificial intelligence will not spend compute to analyse what the ToTheMoon YouTube channel is to that depth. It will look at some parameters here and there, will say, I don't know, some subscriber number the company has. That is, it will not do it. But on the whole I understand that with a proper artificial intelligence that endlessly analyses everything, studies everything, studies all the videos, all the channels, it should in principle arrive at this itself, right?

Alexander Volchek00:52:15

That is, in principle AGI should itself have, from all the channels, from all the millions of YouTube videos, these — well, these are already some cosmic semantic cores. To have them inside. But effectively that is what artificial intelligence itself is. That is what it is. So that it can then genuinely say something, or describe something, or talk about something.

Ilnar Shafigullin00:52:44

A feature request for the project: add parsing of comments too, and accordingly build a reputation for our viewers from the comments. That will all happen. And then give some prize to the most loyal, the kindest commenters.

Alexander Volchek00:53:03

I want to say, by the way, Ilnar, that for me having a site — including the possibility, since we do after all have very strong ToTheMoon participants, listeners, and since we are in the industry and there is a possibility of giving people some interesting perks, some promo codes, some details even, not promo ones. A site gives you the possibility of unfolding additional infrastructure further. Everything to do with various chatbots, everything to do with the work, as you say, with comments — even reviewing all our episodes and seeing how people over a long period, which people, which episodes exactly.

Alexander Volchek00:53:38

And we will build that, by the way. Definitely, right, Ilnar? We will make a rating on the site, a rating of commenters, and then nobody will be able to cheat retroactively. It will already be current and there will already be a record. Two years, two years — you will not be able to do it, right? That is, you cannot create a retroactive comment with us.

Ilnar Shafigullin00:53:57

It is as Sasha is saying now, that we are helping AGI and are creating, saving its future compute. That will clearly count in the future when AGI arrives. When you leave comments, think about that too — that all your comments will subsequently be used to train AGI.

Alexander Volchek00:54:18

Yes. Thank you all very much for today's episode. We will see you in exactly a week. You are on the ToTheMoon channel. Do not forget to support our channel, comment, like, recommend us to friends. Bye, everyone.