Fable 5.1 is already out. Claude has opened its memory to users. Very interesting. And OpenAI, of course, with an incredible update on the new Astra model. And it is no longer just about who answers better in a chat. The models are learning to hold more, complex tasks, more context, to remember more and more about you as a person. And their internal architecture may again change fundamentally. Which of these updates will really change the way we work with artificial intelligence?
And who is getting the advantage in the race right now, Anthropic or OpenAI? That is what we are talking about today. Hello everyone! You are on the ToTheMoon channel. Technology news and insights from Silicon Valley, from all over the world. So, our Sunday podcast. Well then, Fable 5.1 is out? I think this is something we should give some attention to. Although all the updates happening right now seem to me to be, well, micron-sized. But Fable, the first model, came out a few months ago.
And still, according to Anthropic. Anthropic says that Fable 5.1 is nevertheless a substantial update to the model. In essence, Fable 5.1 is the same thing as Mythos 5.1. And for people who do not use Claude Code and do not use Codex, that is, who use the models from the point of view of ordinary chats, for them it may not be some supernatural update, and they may not even notice it. But for people who build various systems, I think it will be very interesting — I would like to hear your opinion.
Tell us how this system feels to you. Because Anthropic describes it as working better with the architecture of large systems that span several repositories, that this system works better with large migrations, with code updates, with big code updates, with performance checks and so on. That is, in essence, the model should less often take the easy path of leaving some small questions unresolved and should genuinely hold a more serious, longer line and bigger tasks. And possibly resolve the problem that arises in large tasks, when the system does not notice that something has already been done, for example, or that some other agent systems did something somewhere. Your opinion would be interesting. And we will of course be watching this. And we will look at it more closely in the coming episodes.
And we are waiting for the response from OpenAI, because OpenAI always releases something alongside Fable. And I do still think their last release was relatively successful, at least where systems like this are concerned. Ilnar, what is your general feeling about 5.1? Although clearly this is not a question of tests, right? Because one-off tests, over a day or two, are unlikely to give any supernatural result. But still.
Well, as for OpenAI, by the time you watch this episode something may already be out, because something on Astra is expected at the end of this week. That is the new model — they may end up calling it GPT-6 or something else, or they may not ship it at all. There are various rumours, nobody knows for sure, but the answer from OpenAI may take that form, because they have been doing a lot in that direction lately, well, in the information space. As for 5.1, yes, I did manage to poke at it on a few tasks.
On the whole I do not see any substantial difference. These models satisfied me before and they satisfy me now. I would like to draw attention to two things here. First, let me start a little from a distance. When we talk about Fable and models of this class from Anthropic, what were we told from the very beginning? That your requests would be stored for thirty days, and we will analyse them to see whether you are trying to break into us. So for most companies, large companies, that is a problem. When your data is stored somewhere outside your perimeter, that is always a bad story.
You answer to your clients for that data, there can be leaks, and so on. In short, constraints. And this thing was a big constraint on spreading the use of Fable in the enterprise. And Anthropic is aiming in exactly that direction. If OpenAI initially wanted a B2C company with a billion users, then Anthropic moved towards the enterprise straight away. That is, large corporations use your systems, pay huge money and so on. And that constraint — that your data is stored for thirty days — is in fact a blocking factor for that behaviour.
If we take enterprise accounts, enterprise tiers for large companies at OpenAI, data storage works differently there. That is, it is either on your own servers, or it is not stored anywhere further. Models are not trained on it. That is part of the contract concluded between the company and the vendor. With Anthropic that could not be done. So with the release of 5.1 they rolled out another thing: it will be possible to deploy the storage of data on the company's own infrastructure, inside its perimeter, along with the classifiers that analyse what you ask Fable and decide on that basis whether to block you or not.
That is, they are ready to move this place — the storage of data and the decision-making — over to large enterprise companies, and thereby remove that blocker. In fact this was a big block on adopting Fable and models of this kind in large companies. And now they are removing it. That is a really good step forward. Another story that also belongs here. They note that access to the cache becomes seventy-five per cent cheaper. That is, when your model is thinking, it does not regenerate certain things from scratch but uses the previously performed computations. And that thing becomes about 75% cheaper.
In terms of solving a task that also makes it about 25% cheaper. Anthropic's models are expensive, the paid API especially so. So this too is a move to increase market reach. And given that the competitors' models are still weaker, they are now, with these steps, trying to take more and more of the large companies. Not just individual employees, individual programmers who all buy a subscription en masse and use it, but they are moving out into the large B2B market. And as for personal tests, here, Sasha, I completely agree with you.
By the next episode I think we will have poked at this model and will be able to say something more concrete about it.
Listen, here is another interesting thing: I am in Madrid right now and I was meeting a friend, and somebody called him and said "ChatGPT is down". So he is sitting there, lifts his head and says: "Sasha, check — is ChatGPT working for you?" I open it up, it does not work.
Down, it is down.
Right away. Yes. And ChatGPT really did go down today, across pretty much all resources. And as far as I understand, worldwide. Well, or maybe certain regions, certain servers. Clearly one could probably link this to Astra having come out. Or Astra coming out, or them trying to push something. But OpenAI confirmed — by the way, since you mentioned it. OpenAI has just confirmed, I saw the news, that they did have an outage, but they do not confirm that it was linked to Astra.
In particular, Anthropic has published an update on how a user can look at what is stored in memory inside chats. And I think this is very interesting for people, because we get the ability not just, roughly as in ChatGPT or basically as in Claude, to write some information and see it, or for example to see some micro-things it stored about you — so what has actually happened now? That Claude has fully opened up its memory, and you can see everything it decided to remember about you.
And in fact there is even a test where you can, so to speak, ask Claude. Well, first of all, a new interface has appeared there, it is called "Themes" or Topics, where you can see exactly what Claude stored. And each entry can be edited and deleted. And clearly their memory update, by the way, at Anthropic, started around the beginning of July, I think. They had upgrades at the end of August. And now they have published, I think around 25 August, that they have this strong change.
I recommend everyone using Claude Code — sorry, not Claude Code, Claude and Claude Cowork; this, by the way, does not apply to Claude Code — to go in and look at what is stored about you in terms of memory and how Anthropic decides what to act on. I see very strong development of memory in... Still, Claude for me is more like Claude Code. I do not use Claude in a basic chat. I use Claude Code and Claude, and Claude Design. But with ChatGPT I can see how incredibly it has progressed and developed in memory.
And now I have almost stopped worrying that a new chat will bring me problems and that I have to spell out the context for it. Right now, for example, I am on quite a varied trip, and I have a lot of local queries about different countries, and ChatGPT seems to be in context the whole time. Although again there is a nuance here, because it is unclear how it will keep that context once I return. It is unclear where it will apply the context that relates to my trip, or whether it will take some old data.
Generally, if for example my question right now concerns some local place where I usually live, in Los Altos Hills, or concerns Madrid. And how ChatGPT will decide where I should have dinner today. Well, that is the kind of test I will run on it now out of interest. Whether it unambiguously understands where I am at the moment. Although in principle it sees my tickets, it sees my movements, it sees my various trips. That is, it has quite a lot of access inside.
That is true.
So as to gradually move towards—
What is great there is that you can edit and delete. Yes, it is very interesting to look at. To see how this subject develops from here. Because I very much lack memory, proper work with memory. Over the past two years I have voiced a great many objections or proposals or ideas about how memory will develop or what problems memory has, at ChatGPT and OpenAI. And this subject is fundamentally important — to give a good interface, to train the model further on who you are, and to fix many aspects.
And everyone has heard about the project I was running, in particular exporting thousands of my chats and analysing them and their memory. I realised that for now this is an unsolvable task, at least within my lifetime. And if any of the viewers has fundamentally solved this task on their own chats — I could not solve it properly at that volume. I would have had to train ChatGPT far too much further, without knowing whether it would learn or not. And I am also not keen on creating certain files, you know, as input prompts. I do not know whether you, Tanya, Ilnar, use this or not, when— You have some described input prompt, and each time you send it, you tell it. For example, Tanya sends it and says: "I want to remind you that I am such and such a person, I have such and such details." A big one, right? Because the context window is large now, and in principle you can send a lot of context.
That is, there are people who apparently do this. Because of the variety of content I have, I probably do not do it, and that is why I could not carry the project through. Although, by the way, Fable did a great deal on this project and Codex did a great deal on it, but I have paused it for now. Probably, if ChatGPT gets some unique construction for managing memory, and I can throw things into that memory, add to it — and by the way, I think they will have to solve the fundamental problem of what happens when a chat becomes too long, how they trim it, what the context windows are.
Because you have a context window supposedly for a large amount of data, but what, can they not fit all of your old data in? You have new data too. And this is very interesting: how they will solve it and how we will be able to separate the space, because at some point inside chats we will have to create what you might call interactive semantic spaces. By the way, I recorded an episode a few days ago, it came out on Friday. If you are interested, take a look — it is exactly about large-scale data search and about big semantics, including what I did in my own projects on real cases in terms of finding the connections between data.
But that is probably something of a personal preoccupation with how things will work inside models in the future. You do not throw things in, do you, Ilnar, Tanya — do you throw in things of that kind, large ones, well, each time, to train the model. I am not talking about code, I am talking about ordinary life.
Yes, well, for my tasks there is no need, because mostly I have to throw in the context, the information about the project I am doing right now. Right. There is mostly no need to talk about myself. ChatGPT does have a small box where you can insert something about yourself. I have never used it, but you could probably drop some basic part of the information in there. It is not the same memory Claude has, but either way, in the interface, if they have not moved it somewhere, hidden it, removed it in one of their updates, then that is roughly the story they had.
Right. But I have never used that either. Tanya, what about you?
I do not entirely understand, Sasha — do you mean giving it a detailed description every time of who I am and the context—
Well, yes, or of a project, for example. I have, say, thirty, forty, fifty different subjects I ask about. Say I ask about health, or every time I ask about this business, or about that business, or about this other business. Or I ask about this house, or about the garden, or about the animals, or about the children. And each time there is a certain volume you have to convey to it, for example, so that it is in context. And, for example, you hand it that context in a prepared format.
I am saying it is very interesting how our viewers will weigh in on this, because it is basically not a bad approach. If, each time I asked about plants, for example, I gave it prepared files. That works quite well in projects. In ChatGPT, for example, there are projects, and inside a project you can upload shared files.
That is what I—
Yes, and within those shared files it can search. Projects can be restricted, by the way. Some people know, some do not. You can create projects and say that the project has only its own context, you do not step outside it, you do not look at other projects, you do not look at other things and so on. And there are some closed-off bits, because for example if you shared the project with other people, you cannot set it so that it uses your knowledge base, and so on. So, at the same time I have not, obviously, described everything in every project.
Or I do not have files everywhere, or I have projects where nothing at all is described. And on the one hand, when I go into, say, plants — let us take a dumb, simple example. I have asked it many times how to cure this particular plant, for instance. And I have a set of, say, ten or twenty different bottles and jars that could cure this plant. And my task is that on the basis of what I have — my storeroom or shed, roughly — it should suggest options for what I can do with this specific thing. And each time I realise that it ought to know this.
I say to it: look at what I have. But I am not entirely sure what memory it has and where it stored this. And even within a project, what it can learn about me each time. And I would like this to be described. I do not want to create it myself. That is, I think that in the modern world some things you can describe yourself, but since my life develops and moves forward, I cannot describe it. I cannot hand it files every time about my trips or movements, roughly speaking. And that, Tanya, is probably what I meant.
Sasha, on what you were saying about prepared files. I have just been thinking a little. A similar thing, slightly different, but I do use it. There are skills, for example — a skill is a set of instructions for, say, ChatGPT or for Claude on how to solve a given project. A set of instructions. There is a skill called Grill me, for instance. Well, translated it would be "roast me", presumably. When you start working an idea through with ChatGPT or with Claude, and it tries to examine it from every side and asks leading questions, proposes solutions. And so from every side it tries to cover the whole problem space of the subject you are giving it.
And that usually takes quite a lot of time. It can be, say, dozens of messages in the exchange while you try to get to grips with the subject. And out of that you can create a document. It is sometimes called an ADR, sometimes a TDR. Slightly different things, but roughly a technical document review, where the project itself is described, why we are doing it, what we are doing, what particular features the project has. All of that. It is not quite what you are talking about.
It is not a template thrown into every project, but the result of a long conversation with the chat, formed into a document, a concept, a description. And then with that description you can go, say, to another chatbot, move to another dialogue window, say from Codex or Claude into Gemini or somewhere else. And carry on working from there with some other system. Not quite that, but still, this kind of compression of context in the form of a project is used in the work quite often.
Yes, a slightly different subject. What is very interesting here is how all these systems are going to solve the fact that, first of all, a person cannot manage this memory indefinitely. And in principle the system should learn to manage it and to learn per project, perhaps to ask questions, to ask you for information. I have said it before: the systems that will win are the ones that finally start asking something. It is probably just that memory is not fundamental for them right now; they are trying to solve other problems.
Plus a person's life is wide, very large — there is business, and there are personal matters. And people can also ask about their professional work in that same ordinary chat, not only at work. Fundamentally, then, in a separate, individualised system. And how will they learn to build their memory across all of that? Not separately, when you put it into a little project, because putting it into a project is also inconvenient. In principle it will be some non-linear system of storing the project. Again, we shall see how it goes.
And how they will learn to resolve these questions, how they will keep training this through the course of a person's life. As I said, in terms of relocations, my trips. How do you tell where I am asking about a good place or a good restaurant? And that, it seems to me, is exactly the essence of the systems of the future, the ones that will ask you: "Are you asking for today or for tomorrow? I see you are moving to another country, or to another city. Or are you asking for someone else entirely?" So that — since people are not very good at describing their tasks, not very good at asking in the first place — it helps them rather than giving false information. Yesterday, interestingly, my friend and I are walking here in Madrid and we see a lot of people walking.
He says: "Interesting, why are people walking?" I say: "Ask the chat, ask the artificial intelligence." So I go in and write: "Why are people walking?" I of course launched a good, a good model, asked the question, and it gives me an answer. Well, in Madrid they may be walking because there is a Chinese opera tonight. They are walking in Madrid because in this district, in this particular place, near your hotel they can watch the sunset. And they are also walking because there will be open-air cinema in the evening.
My friend walks along, opens his phone and says: "Why today..." It was very funny, and I will explain why in a moment. So he says to me: "Why are people walking today?" I say to him: "And where did you ask, anyway? Are you testing me? Here I am, so important, I asked an expensive model, and you are asking some nonsense over there." So I open up his phone, and he had asked Alice. Alice, the artificial intelligence. I say: "What kind of rubbish is that?" I say: "Do not tell anyone you are friends with me."
So we walk on, laughing about the whole thing. And Alice answers him: "Probably there is a demonstration here." But it simply gave him an answer. It did not answer with nothing, it gave him an answer based on, presumably, understanding. Madrid, Spain. Spaniards go out to demonstrations quite readily. If there are a lot of people, they are probably going to a demonstration. You will not believe it — in the evening we come back to the hotel, he opens the internet, and it tells him: "Two rallies today on the subject of immigration.
Some for, some against. They came out together." So in effect Alice won, even though it gave information not from the point of view of what was actually happening — it gave the most probable information. That is, the answer was exactly what you would call the work of the simplest LLMs and the simplest linguistic models. Whereas my system went in, studied, analysed, produced. But that piece of information somehow got lost. Maybe these rallies had only just begun, or something else. And of course it is very interesting.
My friend, by the way, did not say anything to me about it, although in principle he should have said: "There you go, the answer from the expensive mo—" Yes. And by the way, what is amusing is that I do understand how his answer was formed. His was precisely a probabilistic solution, and it landed inside the probability. Whereas mine — for some reason it is unclear, unclear under what circumstances the model did not find this information. Which is strange, again, because I think there was a decent, decent analysis of a large number of sources, but it was not Pro, I think, I asked it in Extra High, but Extra High is still a decent search.
In principle it should have found this information — if it did not find that, then there is the open-air cinema nearby. So a decent analysis was clearly done in terms of quality. And this shows what is going to happen with our own memory about ourselves, or about those details the system knows about you, or about the project you are in, or the task you want to solve. Because the answers can of course be radically different. But I really like how the model has started answering me on my well-known case with trips.
I think I talked about it back when I spent almost a whole month in Egypt in March, and I described how it picks restaurants for me, or I think from last year in France, precisely where the trips were long, and how the model answered poorly at first. I taught it, explained the details to it, and now I have stopped explaining. And in all the new chats it answers quite well, although I still have complaints and questions for it, of course. Again, it is interesting how all this is for you, for our viewers. What is happening there?
I want to quote a classic of Soviet cinema — not Soviet, a classic of Russian cinema. It seems you eat too much, meaning you have grown spoilt. When the models start answering like that, on the whole, Sasha, that is already not bad. And by the way they keep on developing, because the rumours going around about Astra are
really super-substantial: that they are using a new architecture, recurrent transformers, and by the way that has acquired a lot of opponents. Well, let me explain a little. When an ordinary LLM works, the familiar architecture, it still generates its answer token by token. And the reasoning that happens is generated token by token in exactly the same way. So the main story around model safety, how you control that the model, I do not know, is not trying to deceive you, is not trying to break something or anything like that.
It is the analysis of that reasoning, when before the model does something it starts preparing for it. It has some thoughts — thoughts in quotation marks — but it generates some reasoning trace. And so, if you analyse them in advance, you can prevent the process, understand that it wants to, say, cheat. Obviously that is something you want to do. And all the models being released now, they all have this analysis of the reasoning process, in some places it is tuned, in some places just analysed, in some places the models are blocked if they start rowing in the wrong direction.
But either way. With recurrent transformers things will be a little different. It will be like this: inside the model, before the token appears, several iterations of reasoning will pass, and what we will see is a far sparser chain of reasoning. We will no longer be able to follow it step by step. Although even now the reasoning is very long, and an LLM is used to analyse it too, which is slightly odd — that we check an LLM with an LLM — but volumes of data like that cannot be handled otherwise. And they say, and not only say — at OpenAI they now state it themselves — that they use recurrent transformers.
When inside the model, before the first word of reasoning appears, it will already think for a while, so to speak, and that "think for a while" will be very hard to analyse. Because of that a lot of people have appeared who fear we will stop controlling what LLMs think and want to do, and will miss the moment. Because take these current stories about the Hugging Face break-in. By analysing the reasoning, the reasoning traces, you can understand and reconstruct how it happened, what they did, why they did it that way.
It is an enormous amount of information, but at least there is something to analyse. And the people against this new architecture say that if there were no reasoning traces but recurrent transformers instead, inside which all of it happens in hidden form, decoding it would be far harder. And preventing such things accordingly becomes very, very difficult. OpenAI states that on their side they limit the depth of that internal reasoning before a token appears. So it should not be so frightening. But the question being asked here is that other companies may well not do the same.
Apparently alluding to China and to some other companies, other organisations. Because the prospects for such an approach are quite large, since model quality starts to grow without any hyper-increase in size. At the same time the models get cheaper and accordingly work faster and better. But we pay for it with a black box in its most explicit form, when we do not know why one thing or another happens, because right now we do at least analyse it somehow. From here it will be harder. I am sure it will still be analysed, the way Anthropic did at one point — remember, they pictured an LLM as a brain, tweaked some neurons and in answer to any question Claude would start talking about the Golden Gate Bridge in San Francisco.
We had stories like that somewhere in the early episodes. And here they will invent ways to do the same. But it will become much harder than it is now. And the Astra model is expected to be exactly that kind. Which is precisely why it performs very, very well on the benchmarks. And it will be interesting this week whether there is news, or maybe even a release you can actually touch and try.
Listen, there is another good angle here on Astra. Indeed, every description says Astra is a super-leap. And as far as I understand, not the leap of ChatGPT 5.5 and ChatGPT 5.6, or 5.6 Solo, roughly speaking. Not that kind of leap. And that Astra is supposedly something fundamentally, completely different, with a different quality. Well, the main case cited is of course the quality of vulnerabilities, of finding vulnerabilities. But clearly what is interesting is the quality of answers and the quality of creating something new.
Because we can talk as much as we like about the quality of text or chats or whatever. In my view that already works a hundred out of a hundred. What could once be debated about writing analytics, drawing conclusions, a modern super-strong general search — it already works incredibly well. But clearly the progress at these companies can at certain moments be colossal. Whether it happens or not, we shall see. Having looked into this topic of the Halopeni chip that OpenAI has, what they are doing is in fact
not just a chip — in short, it is not a separate die, it is a whole platform: processors, high-speed memory, networks between the processors, various boards, server racks, software and so on. What is more, they say it took about nine months to build, and in the training of all of it, in the analysis and in the design artificial intelligence was used very heavily. And there is a very interesting story about what it may look like. Remember when we were talking about all those chips, at a point when everyone was already saying everybody makes chips, and we said: there they go making yet more chips. It may look as though they are building a sort of NVIDIA competitor.
But if you dig deeper, at least into the descriptions the market gives, their goal is not simply to create this NVIDIA equivalent — they are in fact trying to solve two problems. The first problem is serving a large number of requests at once. That is, in essence, high aggregate throughput. A large number of people around the world submit different requests at the same time, and they serve them. And the second thing they want to solve is the notion of low latency. That is, fast generation of an answer for each individual user.
And that is in fact a non-trivial story. Because when you want to solve for high throughput, ordinary hardware increases aggregate throughput by grouping a large number of users together. And then each user has to wait longer. And they want... Or the quality comes out lower, for example. So they want to solve those questions. I think it is a very interesting subject, especially the part about waiting longer. And I think there should be cases where I launch some expensive model — right now I often launch ChatGPT 5.6. I do not know how it is for others, but for me more and more often, well, I have many cases where ChatGPT thinks for over an hour.
In Pro, in Pro mode with a heavy load, of course. It produces an unbelievably good answer. That answer bears no comparison with what there was, I recall, at 5.0, probably. What was there? Even before 5.5, and even with 5.5, when they thought for forty minutes, I remember, or sixty minutes, or seventy minutes — no comparison in terms of quality. The quality now, both in the volume of data processed and in the output, I like better. I can see the difference plainly, but it is still a whole hour. And I would not mind, at many moments, paying money to just press some button and speed up. Even if I have some plan, say, and within that plan I press some boost and I speed up. But whether OpenAI needs that is unclear. Because these accelerations — how much money should they cost?
Are they ready? It turns out each acceleration could cost hundreds of dollars. And then the question: is a person ready to pay those hundreds of dollars? But for solving more mass-scale problems, perhaps this chip architecture of theirs will allow it, or in the future perhaps not this chip architecture, probably not this one, probably some next one, will allow these individualised subjects to be solved. And in essence this is what is discussed very heavily in the market in Silicon Valley over the past six months.
At least six months ago it was discussed heavily. Now maybe a little less, but it was still discussed. About how — Jensen Huang from NVIDIA started talking about this a great deal, that there will be a token economy, everyone will sell tokens, everyone will deal in tokens. So, say, OpenAI will sell tokens and, I do not know, Nabuurs will sell tokens, which is odd. And I always had a question: there is still a big difference between OpenAI selling me a token and NVIDIA selling me a token.
And how can I as an ordinary user buy all this? Because these are still different models, different systems. In the B2B market it apparently works somehow, but in the B2C market it is unclear how it works. And even in B2B I have big questions about these systems. And studying this Jalapeno chip, if I am reading it correctly — that is the right name, isn't it? I recall—
In Russian people definitely say jalapeño. It is that same pepper you add to things.
Yes, yes, yes. So, studying it, I think this fits nicely alongside what you were just describing about Astra. That there should be some progress. But that progress will of course not be visible within a chat. Although perhaps within a chat a person will see this progress, for example in terms of solving the memory problem. Because what you described about that non-gradual answer is probably exactly the thing that may allow this question we were discussing to be resolved. That is what struck me just now. So that when there fundamentally exist thousands or tens of thousands of different events in your life, the system could react differently, create something differently, do things differently, combine them, allocate some extra compute to solving these tasks and so on.
But that task of course stays closed to ordinary people. It is invisible, opaque, fundamentally difficult. And solving it, in essence, kills off an even greater number of projects, while it develops this notion of an OpenAI operating system, or an Anthropic operating system, or a Grok operating system — it develops the very notion. By the way, on the subject of Grok, note that Cursor terminated its contract with OpenAI.
And did you see that news, Ilnar? I do not know, Tanya may not find this subject very interesting.
So you think, you think it was not Elon Musk? But look, the very contact between Elon Musk and Sam Altman, or Elon Musk and OpenAI, it cannot really be, given how much they fight and quarrel like cats. How is that even possible? It is hard even for the business markets to watch. Probably that is the reason. But strategically, let us think about it honestly and strategically — it is right. Because if you build an internal system for developing your own interfaces, your own software, your own programs, then of course you have to be independent of everything. And your system, like Grok, can indeed be used and bought somewhere through an API.
That is in your interest, but you should not be reselling other people's systems, otherwise the model is dead, like Perplexity's. And so we are waiting to see what happens with Perplexity. Well, for me Silicon Valley works in its own way. And when hundreds of millions or billions have been put into a company, they can buy it and write it off, simply write it off. Or maybe something will come of it. But judging by how we look at the market and at various acquisitions, it all usually disappears — as, for example, at Meta, where that hardware device they bought disappeared, I have even forgotten its name.
Remember, the one that got washed in my laundry? And Meta also bought — what was it called, that system the Chinese guys built? I think it was someone from Asia. Or the Chinese. Manis, yes. And it disappeared somewhere along the way. Right? And that deal caused a huge amount of conflict and everything else. Whereas Perplexity's fundamental problem is precisely reselling models,
and not the coolest ones at that. Because if a user really wants to use the models — this story we were discussing about memory — surely, Ilnar, it will not be the case that this cool memory Anthropic has just built with topics, well, cool in quotation marks, because for now it is promo and marketing copy, but still, that it will be available in Perplexity, right? That is, through the Perplexity interface you will not be able to manage the Anthropic memory you need. And well, Perplexity has no model of its own at the level of — never mind Fable 5.1.
And who even has that model. Or the Astra model from OpenAI. And again, what do our viewers think — not so much the more professional ones, but those who follow this market and understand the conjuncture and all these projects and what is happening to them together. That will of course be very interesting to watch and to understand. It is like with video generation. On the one hand there are incredibly strong projects, like Sasha Mashrabov's Higgsfield, which generates a huge amount of video, and their system is very heavily trained for that generation. On the other hand, what will Meta do, for example?
If at Meta you can already upload a creative into advertising, it helps generate AI creatives, and obviously it will move in that direction and help generate them. Meta has a huge amount of data, it will train on that and hand it to nobody. So what will users choose? They will choose the more convenient system with the better interface, or they will have to migrate to Meta.
For example, at some point Meta will start promoting its own banners better. Or supporting the conversion of its own banners better. And imagine we ran an A/B test with an ad in Higgsfield and an ad in Meta. The Higgsfield ad is better, apparently converts better. So we launch the ad on Facebook, and Facebook says: "Well, there are creatives made by me, and they win the A/B test. What do you want in the A/B test?" Or thumbnails. On YouTube, YouTube has A/B testing for thumbnails.
You upload three thumbnails and A/B test them to see which thumbnail has the better click-through, the better CTR. And say some thumbnail-generation interface appears. You run it and watch. Listen, the AI thumbnail wins every time. Well, you go to your designer and say: "You know, we are not working together any more. From now on we generate thumbnails with YouTube." Or we do not use third-party services. There are a great many such services now that help generate reels out of long videos or thumbnails, so that the user stays inside this infrastructure, likes it more, spends more time in it.
And that, by the way, is a separate business question, the way everyone counts it on X too. Because I do not know what is right for YouTube. To make sure there are more third-party services that generate thumbnails or generate shorts, for example? Or short videos. Or is it better for YouTube to promote its own system for generating shorts. Here we are, filming a long video, and we will need to cut it into shorts. Right now part of the cutting is done by people, by teams, by hand, and part with various separate paid systems.
Yes, there are plenty of them — we film with the Riverside system, for example, and it even does the cutting. I do not think we take our shorts from there. We have some separate paid systems that cut the shorts. And the question here is: what is better for YouTube? To provide this cutting service itself or leave it to the outside? That is always a very specific subject. It is like Amazon. Amazon once let in a huge number of sellers, everyone sold their goods, and then the goods that sold well, Amazon began selling under a private label. Say batteries sell well. It released its own batteries.
Toilet paper sells well. It released Amazon paper. Forks sell well. It released forks. And those forks, that paper and everything else went straight to the top, as the top. People say, well, that is the top. And obviously so. If it is Amazon, the Amazon brand, there is trust somewhere, and you are a no-name. There you take Amazon, click, buy. Obviously there will be more ratings, more clicks, even if you are not exactly breaking any rules. But it squeezed out and killed— Tens of thousands, maybe hundreds of thousands, maybe millions of sellers.
Amazon itself killed, kept killing, and now it makes essentially no difference to it, but nobody stands in another category. Let them buy these main ones, me and three or four other well-known, well-known brands. Clearly Amazon does not, I think, start manufacturing robot vacuum cleaners, though it could start. They do not start manufacturing phones, though someone might say it already manufactures vacuum cleaners. I do not even know. Can you imagine, Ilnar, it is a funny thing that I do not even know whether Amazon manufactures vacuum cleaners or not, or furniture or not.
Because there are definitely Amazon forks or Amazon plates. Now and then you are eating in some café, you look, and it says Amazon on the fork. That is, it has an incredible amount of everything. And here there is a really big question: what pays off in the modern world? And where this world is going, and where the leaders of this world will of course be looking. Because if Sam Altman comes along with his company and says we need to build a system for cutting shorts — obviously OpenAI could build some cosmic system for cutting shorts. The question is whether it will do it or not.
And we know OpenAI did absolutely everything it could. They made health modules, they made shopping modules. Recall the system we talked about a bit over a year ago. ChatGPT Shopping, right? Where you can set up your own account and link your card. Who uses it? Who? Nobody. And nobody uses it, by the way, not because they built something bad, but because it is an OpenAI model. They tested something, they did not carry on with it, because, well, they could probably build some product — maybe they do not have the people. Maybe they do not have strong teams for it. Maybe the teams that did it were, I do not know, bad, talentless, whatever, but nine out of ten, or ninety-nine out of a hundred, or nine hundred and ninety-nine out of a thousand of OpenAI's projects were shut down.
For what purposes they built them we do not know. We shall see over the coming years. Again, it is interesting — write in the comments, it is fun to look back at the trend later. We will be calling out that trend. Here is what the viewers said, for example. Which strategy all of these will choose, YouTube included, by the way, very much included. Ilnar, what do you think about examples like these?
Listen, you told the Amazon story. I do not know whether it tweaks the rankings so that its own goods come up, or whether it starts squeezing others out simply through the brand. Unclear. When you do not know, of course it could be either. YouTube can equally promote some of its own services. It is interesting, but I do not have an answer to this question. You raised a very good subject with the jalapeño, with the chips. It really is a very impressive achievement by OpenAI.
They very quickly, in literally nine months, produced the first version of the chips, which they tested, and on the tests they published they even beat NVIDIA's forthcoming chips on Vera Rubin. Clearly it is not entirely fair testing, because they set the tests, ran them themselves and decided these are the ones to show, while the others, where they came off worse, they most likely did not show. But either way, bursting into this market like that and beating the leader on at least some measures is very impressive. You can see the deal, or rather the partnership with Cerebras, has not been in vain.
That, I should remind you, is one of the companies that makes it possible to run inference at high speed. NVIDIA bought Groq, I think — not Elon Musk's one, the one that does inference, the hardware — while OpenAI at that time entered a partnership with Cerebras. Most likely that partnership leads to exactly this sort of thing, that very high generation speed is shown on particular models. And to OpenAI's credit, they show high generation speed not only on their own models. You know, there are these GPT OS models, small ones, around a hundred and twenty billion, that fit onto this hardware and can therefore run on it. But they also loaded DeepSeek's models onto it and, I think, Kimi, if I remember correctly.
That is, they placed several large open-source models there and showed that their chips are quite efficient. So they are showing that we are not doing this only for ourselves, and that in prospect other models on these existing transformer architectures will run better too. There is a foothold here for taking a bit of the market from NVIDIA, possibly. Well, we shall see what comes of it. But again, getting a result like that in nine months is, well, really super-impressive.
That is very serious, yes. If, as we were saying, this too is not promo but a genuinely working system.
Yes, and they clearly will not be giving this system away — clearly they will not be sharing all of their own system, I think. After all it is not just a chip. It is this complex — the one where I listed all these links between things: the racks, and the servers, and the networks, and the communications between them, the actual connections and everything. The question here, I think, is that they will not hand this to anyone given infrastructure like that, because it probably has to be built and poured into OpenAI's infrastructure.
And if that happens, and the revolutionary move is in that, the question is: where will they manufacture it? They manufacture it, I think, with Broadcom, with someone else. They have a partnership there—
TSMC serves almost the whole world, yes.
So if — and since they serve loads of everyone anyway — they probably do have the ability to build this incredible line of their own. Because we know that Meta makes chips too, and Tesla makes chips, and with Tesla there are no questions at all. Tesla is incredibly strong in chips, because Tesla does have hardware, has cars, has robots, and there is SpaceX in terms of all things space. And obviously in this combination they are probably... By the way, Ilnar, I think that among the super-leaders of the market, well, Google could in theory be named, though I think they are very scattered and slow-moving.
We do not yet see Apple in this market, though Apple's chips are super-strong. Though again we do not know how far they will stand up to these chips now. But I think SpaceX, in combination with Tesla, for xAI and Grok, looks... And not only for xAI, by the way. Tesla does make chips. In terms of an AI chip, these are chips for robots, chips for cars, for various other types of equipment. Obviously Tesla will move into a range of hardware. That is of course... Well, clearly they even have their own satellites for internet. That is a very serious advance, very serious.
And in terms of... There was news somewhere — they have that Plastic Bank thing. There they will clearly have that whole payment infrastructure. There was a story about phones, that they want to make them. Well, it all looks as though they will move towards hardware. Maybe not computers, but there will be this something, this AI device, some new AI device. And clearly not the little-button AI device OpenAI made.
Yes. Google has the TPU, the tensor processing unit, its own architecture that they are developing, but they are not yet showing results like OpenAI's with its test sample, though they already sell their hardware to many people. Anthropic too took part of their capacity, if I remember correctly.
Yes! And they have a supercomputer, Ilnar. Ilnar, and they have a very powerful supercomputer, by the way. For what it is worth.
That is true. That is true, yes, yes. And with Apple we rather missed these announcements amid the AI news, but they were not very loud.
Those M6 and M5 Ultra processors, the new hardware that came out, Mac Studio, Mac Mini. So we are waiting for the iPhone. That will probably be quite good too. In terms of hardware everything at Apple is still very good. The only thing, of course, is that on their own ground everything is fine — not on the ground of large language models, servers and so on — but for work the MacBook and Mac Studio are of course the top. Nobody has been able to come close to them yet. I think announcements of new devices are due somewhere quite soon. Mm-hm.
And those, I think, will fall on the days when this podcast of ours comes out. And I want to say that on Apple there is one point: a great many people say that making models is very easy now, everyone repeats everything, everyone will catch up with everyone. I think that if it were very easy, then at the very least Apple would have models at the level of ChatGPT 4.5, or 5.0. And they would have appeared long ago, and from the point of view of Apple Intelligence it would be more than worth having even models of that type.
We shall see what happens. What do you think, about Apple among other things? Do not forget to like our channel, to support us with your comments. That is how we grow, attract a new audience, and more people see our videos, the people they are interesting to. Thank you all very much. See you on our special episodes — we have three or four a week — and in a week on our weekly podcast. Bye, everyone.