OpenAI has released Astra, and today it is GPT-6. It is the first OpenAI model the company has placed at the critical capability level in cybersecurity. Today we will work out why there is so much talk around it, how it fundamentally differs from the previous models, where the real leap happened and what that could change for people who use artificial intelligence in their work every day. For you specifically. That is today's subject. And separately, by the way, we will look at what in all of this is genuinely important, from Astra's own point of view, and what for now remains just pretty numbers in OpenAI's tests.
Let's start. Hello everyone! We are on the ToTheMoon channel. So, the long-awaited OpenAI model Astra is out. Just a few days ago on our podcast with Ilnar and Tanya we were saying that Fable 5.1 had come out and that Astra was expected. And it has arrived. Today is a separate episode devoted to Astra. In general, over the last six months I have not been a great fan of devoting an episode to one individual model. Well, I would probably do it for Fable, of course. And we have done a few episodes like that, but Astra at OpenAI does have a unique feature.
That is what we will talk about today. And, of course, it deserves serious attention. What made Astra interesting to me was looking at everything that is going on. First, on the one hand, OpenAI talked about Astra a great deal, saying that it had, unconfirmed, broken into Hugging Face together with 5.6 Sol, that it was Astra whose development as the frontier model was paused, relative to all the others, and that it is Astra that is the new, new special model which does not work like all the previous ones. And we did not even know for sure whether Astra would be GPT-6.
And today it is GPT-6. Right? It is not GPT-5.6 with some internal branches like Sol and Luna, is it? It is not GPT-5.5 and everything else that came before it. It is a different line of OpenAI models altogether. And yet there was a big difference between GPT three, four and five. And from what I see in the description, Astra is of fundamental importance. And most importantly, yes, there is a comparison. I think I will start with that, to understand why this model deserves attention and why you should not treat it as something ordinary. Because there is a lot of basic reasoning along the lines of: what difference does it make which model came out, the Chinese will copy everything anyway.
What difference does it make which model came out. Someone will now distil the models and launch training against them. These models will be replicated, right? And in particular, when the Fable model came out from Anthropic, that was Fable 5. Now it is already Fable 5.1, and the update was quite significant according to their descriptions. OpenAI released 5.6. It was clear that this was not Fable's level, although on the whole they talked about it a lot. And then even in China Kimi Moonshot releases the Kimi model, K3, and says it is at the level of Fable and at the level of 5.6 Sol, right?
Even their PR service wrote to me and said this model is at the level of 5.6 Sol and Fable. But in the weeks since Kimi K3 came out, or it is probably a month by now, I have not seen such an explosion in worldwide use of this model that it really showed the level of Fable's work, just as 5.6 from OpenAI did not show it either. And it seems to me that the Astra model is still a breakthrough, from OpenAI in particular, something completely new. And together with their new infrastructure line, everything to do with the Halpenia chips and around all the servers and software, OpenAI may once again come out into the leading position relative to Anthropic, right? Because at some point, let's say, they drew level.
So, on to the indicators, the parameters worth paying attention to. There is a parameter called "scientific tasks with programs and a terminal". And here GPT-6 Astra suddenly shows sixty-four percent completion in internal tests, against the GPT-5.6 Sol model. Sixty-four against twenty-two is, in essence, a fundamental leap. And I want to say right away that I had a certain indignation towards 5.6 Sol. In recent months I even moved more towards building systems with Claude. Although, of course, I remain a complete adept of using ChatGPT every day, and 5.6 Sol Pro, or Extra High. But a leap like this really gets your attention.
Or there is a parameter where the leap is especially noticeable for Astra, and that is why I am telling you about it today. That is the automation of professional work. Forty-one percent. Against eighteen percent, where Astra is better than the previous models at understanding when it can make a reasonable decision on its own and when it needs to check with the user. So this is a very big indicator, a big difference in independence, in understanding what is going on. Although this understanding story does not mean that Sol asked more often what to do; rather, we will hope, together with you, because we will be able to speak about Astra's effectiveness probably in September and October, we will certainly be discussing it then, together with everyone here. And the whole market will be talking about it. We hope it is really so, because 5.6 Sol, in terms of its own reasoning, I think, lost quite seriously. Just as I do not like what happened with Opus five, right?
Anthropic had an update, there was Opus four point eight, Max. And then they released Opus, Opus five. Opus five, and it works in Ultra Code, I believe. That is sort of the maximum mode. So you must always understand that these models have different modes, right? Just like 5.6 Sol. It had many different modes. And, for example, even in the paid version extra thinking was available. The Pro version had Extra High and Pro. And in essence, today these are fundamentally different capabilities, different outputs.
Even for simple requests. It is hard to say whether the Extra High or the Pro model does better at helping you choose a restaurant for the evening. One can wonder about that. But for serious tasks the progress is, of course, obvious. Where else is there a leap in terms of exactly this checking with the user? In OpenAI's internal tests there is a notion called "complex programming tasks through the terminal". Right? Not scientific tasks, just complex tasks. And there it is fifty-seven percent against thirty-seven.
Twenty percentage points relative to thirty-seven is more than half. Also quite a serious move. Right? There is the notion of understanding computer interface elements. Ninety-two versus seventy-six percent. There is another parameter: false statements in the internal test. Right? There, the lower the value, the better. It dropped almost to zero for Astra. Just four percent against the twelve that it was. So. Astra is positioned from the start, if you look at the biggest win, it is positioned in agentic work, in programming.
Somewhere they also consider the science area. But we will approach it fundamentally broadly, from the point of view of ordinary human life. I always say that ordinary human life is far more important and complex than scientific life, than any supernatural programming. Far more important. That is, arranging and managing a person's life is far harder. Right? And to me, talking about some agentic work in relation to a person's life is laughable. Right? Because in a person's life hundreds or thousands of different processes are fundamentally happening at once, including parallel and multifaceted ones.
And that is why ordinary human life is far more complex. And when we see more serious results in that same agentic work or in programming, it is some movement towards these systems genuinely helping all of us in our own lives. And helping not just with silly things, with simple things. I remember how the GPT era began three years ago. And that whole story. ChatGPT will write a joke, ChatGPT will write a letter, ChatGPT will write a poem, will write a song. Very abstract things from the point of view of a person's life. But when ChatGPT genuinely understands who you are, right?
And judging by Astra's descriptions, at some point Astra will probably be able to do exactly that. And I hope memory will improve there. If you did not watch our episode last Sunday, Ilnar and I went through memory in a lot of detail, including the fact that Anthropic released an update about memory. Theirs is the memory of what Anthropic knows about you. Watch that episode. Ah. There are many different features there. Right? So, clearly, if you assume that you use ChatGPT in the style of "find me the errors in this letter", "where is such-and-such street" or "what is the population of this country".
Well, the difference from GPT-5.6 will probably be small for you. Although it seems to me there are also fundamental changes happening in Astra in terms of what this model actually is. And I would not answer that abstractly either. Right? Just as in the Fable model, you can see a different quality of work. Clearly the programming market, the market for building all these systems, the venture market, everyone is looking at what these systems cost, what they spend, what tokens cost.
They say complicated words to people. Here, still, the model is fundamentally new in its architecture, and this may change the way, and the possibilities, of your work together with it. Right? Even if we talk about the versatility of the chat. Right? As of today that task is relatively solved. But overall. In the programming interface Astra supports a context of up to a million tokens, and the answer there is a hundred and twenty-eight thousand tokens.
Again, these are just some ordinary numbers. Because, on the one hand, we can say abstractly that this lets you give Astra a large number of documents, big ones, I do not know, load millions of some letters into it, or tens of thousands of letters. Millions is questionable. Load a big codebase. On the other hand, it is abstract, again, because when I read about all these details, I said, well, conditionally, while studying this, I wrote ChatGPT a question. Listen, every time a new model has been released over the last year, it is said that this model is better at agentic work than the previous one.
And you stand there thinking: so what is the trick? So what is it better at? And if, for example, Astra basically claims that this model is better at filling in forms or works better with the calendar or with mail. Well, all the previous models also said they work better with the calendar, they work better with mail or they are better at creating and editing various documents and presentations. If you look at OpenAI's presentation of Astra, they have a video, there are some things you watch and you do not understand what the fundamental difference is, right? Well, this system lets you draw a rocket. Well, apparently it draws a cooler rocket, or it lets you draw some 3D world.
But who exactly needs to draw these 3D worlds? And again, how much does this move the system towards that same AGI.
That is, if they say AGI is about a system being above human level, then, well, I made an episode about that. I hope our editors will show a card, because I went through all these terms, AGI, ASI, quite recently. And I talked about the singularity of artificial intelligence. And OpenAI, by the way, has claimed it already has AGI. Or even, in part, Sam Altman said a month or so ago that in essence the technological singularity has already been reached. Although I only relatively believe in reaching the singularity, because if there is a technological singularity, then who can control that technological singularity?
In theory, such a system cannot be controlled. But still. About Astra OpenAI says that Astra is better than the previous models at understanding when it can once again make a reasonable decision. Yes. That is an interesting aspect of this wording, again. Why is it interesting to me? Because I believe it would be great if the current models, and especially those of the last three years, asked the user a question, asked about the user themselves, about where they are. And even when offering them something, looked at why it is being offered at all.
For example, OpenAI periodically sends me emails offering me scenarios which, conditionally, could be offered to a beginner in artificial intelligence. Or they write to me: "Use voice. It is our new feature." And you think: "Guys, I have already used your voice a huge number of times. You can see that. Why are you offering me this?" And that is where you ask yourself this question. That is, these interfaces you are building, do they work with the help of your artificial intelligence, or are you still building solutions for us yourselves, ones that are not artificial intelligence. And we are, after all, in a kind of vacuum.
As I explain to many professionals, to you and to us as well, to everyone among the viewers of the ToTheMoon channel, there is a big difference. For example, you work as a marketer or some middle manager in a company, even in a small startup. And this company gave you some interface for working with agents. That is, the company said, for example: "We have connected an AI bot in our system. It could be in Notion, it could be, I do not know, in Slack, it could be in Telegram, it could be in a home-grown ERP or a CRM system.
We have connected a bot, and this bot is artificial intelligence, and it can answer any of your questions about all kinds of information." And it seems to people that when they work and interact through this bot, which then works with some software, and that software possibly works with some other software, and then somewhere there some model works, or possibly no model works at all and this part is simply programmed. It seems to them that this is the same thing as if I had connected directly to different systems in Codex or to different systems in Claude Code.
I will try to make a separate episode about this, by the way. I think this topic is very relevant. And recently, by the way, I filmed a similar story where I asked you what you see coming. Will there be something like Codex as an operating system with no interfaces, or will there be interfaces? Once again, what do you see in this story? By the way, write in parallel what you think about Astra while I keep telling you the rest. But do not forget to support our channel. It really helps our promotion. And any like and comment of yours helps promotion a great deal.
So it turns out that a person who works through this agent fundamentally does not understand the difference at all. If he works from inside Codex, the system is completely different. He has no idea how these systems are built. And, in essence, he loses the highlight of working directly in ChatGPT, or inside Anthropic, for example. Just as a person loses that highlight if he works in platform aggregators such as Perplexity. Or a person loses that highlight if he essentially uses some third-party solutions to process documents rather than processing those documents directly in systems such as Astra. Again, someone may argue and say no, that is not so important.
I do not remember, maybe someone in the comments will point it out for our viewers, or perhaps our editors will. Whether Fable is available in Anthropic's free mode or not. Or, say, 5.6 Sol Pro. 5.6 Sol, by the way, I think is unavailable. Either Luna or Terra is available there, one of them, but Sol is not. And if these models are unavailable to you, I think even Opus five, Max is certainly unavailable, or Ultra Code, in the ordinary modes. You do not have access to them. That is, you have access to some kind of architecture.
You are often in isolated environments compared to people who work in the pro models. Again, the discussion here could be very long. We can quickly write down what is going to happen. So, as of today Astra is available, judging by all the information, at least at the moment this episode was filmed, and at the moment it is published a few days later. We will see exactly what it will be. Available in the paid versions: in Plus, in Pro, in Business, in Enterprise. Over the next few days it will clearly be rolled out in different countries. Most likely it will be the US first and then all the other countries.
Or maybe not. Let's see what happens, right? Somewhere today there is a description with a clarification that Astra is so far available only in the Pro version, not included in Plus. Again, we will see. I recommend it. And there, by the way, even in Pro, in the Pro version at two hundred dollars, it says there is a limit of two hundred messages a week in the ordinary chat. But Codex and OpenAI's Work are counted separately from the chat, I remind you. But I will not give you any guarantees here about how these models are available, because everything can change every day. That is, even if you have the simplest subscription or cannot pay, it makes sense to hear about these models in terms of their differences.
Let me remind you once again that Astra differs in the architecture of its systems. And why is there so much talk around it? Because Astra is the first OpenAI model that the company has placed at the critical level, at the critical level, if one can put it that way. That is the critical level of cybersecurity capability, in the trials. And what did they say? That in trials without public restrictions Astra could find unknown vulnerabilities and create working ways of exploiting them.
And therefore, I will remind you, the same happened with Fable. And that is why the public version got stricter restrictions, clearly. Additional monitoring of the whole sequence of actions. Not for nothing did OpenAI recently introduce control over how its models work inside businesses, in the business version of ChatGPT. And began, among other things, saving your chats in isolated environments. Today I would not assume at all that any chats in which you want to isolate something will be guaranteed one hundred percent to disappear from everywhere. That is, as of today I think one cannot speak of one hundred percent confidentiality in such systems, in any system at all.
I am not even talking about Chinese models, and of course that includes OpenAI and Anthropic. So clearly, at its core, once again, Astra is not a smarter chat. Although in their reasoning, once again, some systems will tell you that this is a new digital executor or some new agentic environment. I do not regard any of this as an agent. I do not think at all that modern models such as Fable or Astra are agents, right? This is what can genuinely be called artificial intelligence, by the way. Not that it is artificial intelligence.
I always say that real artificial intelligence exists in films, like Oblivion, conditionally. My favourite example. Even real artificial intelligence is not what is in the film "Ready Player One". But we have not even reached that level yet. And still, Astra is closer to artificial intelligence. And in this case let's understand artificial intelligence not as a person, not as an agent and not as some single procedure, but as a kind of world or space in which you are and which does something. I think that is the most descriptive story for models like Astra, or possibly models like Fable.
Let's see once again how far Astra has surpassed Fable, because according to OpenAI Astra has substantially surpassed Fable. Today I also want to list for you, since Astra does have limitations, where you can think about what could be fundamental growth for you with Astra, right? That is, today I wanted to show you that Astra is still not just a model for comparison.
Like, bigger, cooler at programming, or cooler at working with files, or better at filling in questionnaires for you. I wanted to show you, fundamentally, that Astra could become the beginning of systems that will understand you more fundamentally as a person. Let's see. Perhaps that is my assumption. It will move in that direction anyway. But in parallel I still want to show you, for your life, where you use models professionally. Many of our people find this very interesting. For those who use models every day, where Astra can really show a leap, so that we are not in the space of some purely abstract things. So: the automation of complex work in programs.
That is, many people automate various processes, including in accounting, in sales, marketing, editorial work, in management, in analytics, finance. Well, in very many things. There, there is a very strong leap, more than a twofold increase. It is clearly visible. In terms of the result. Yes, if we talk about the result. And clearly there is the story for the people who work in Codex. There the leap shown is quite strong. That is, if you want to develop various solutions, both solutions with interfaces, that is, conditionally, you created some website or some program that people use, and solutions where you simply wrote some code that constantly recalculates something, goes somewhere, checks something.
The model should work significantly better. And today the leap there is from thirty-seven to fifty-seven percent. That is, a really fast gain sits inside. There are leaps that are quite noticeable in everything to do with programming. There, for example, database migrations. Well, that seems to me a rather abstract story, more applicable to strong techies. But there is a topic where you may see the progress. There is the notion of understanding interface elements on the screen, when you ask the system to keep going somewhere in parallel, constantly, not once, to order something for you, as in OpenAI's favourite examples of ordering sushi, but to do specific things, go to certain sites, gather different information every day; then before this, the previous system's understanding of interfaces was seventy-five percent, and it became ninety-two, right?
That is, in essence, this system presses the wrong thing and gets lost in the interface less often. Although ninety-two is still not one hundred percent. But a person, I suppose I can say, can make a mistake here too. Still, such an error, eight percent, can cost you a great deal, right? If, for example, you are filling in some questionnaire, OpenAI had an example: writing and correcting a contract in the interface, right there in Word or whatever they did it in. Or filling in a form, it was a tax form. They showed how a tax form was filled in. And you think, well, making a mistake in a tax form with an eight percent probability, well, that is not allowed, right? You must be guaranteed that your probability of error is, I do not know, below one percent.
I would probably say zero point one percent. Right? That is, either the system has to recheck everything several times and reach that level. But an eight percent error is, for instance, decent work. Where there is a very strong leap is scientific work through terminal programs.
That is the most practical leap. And, of course, what scientific work is is somewhat of an abstraction. But you can, first of all, study Astra's description if you need to. That work, of course, involves large datasets. Complex systems, yes, a complex understanding of systems. And remember that video I released recently about how Anthropic did a very cool study on agents working simultaneously and whether agents agree among themselves or not, and how badly they worked with each other.
And that is exactly where this theme is: when you start changing something, how well does the system understand the complex context it is in. It seems to me that this parameter, scientific work through terminal programs, is very much about that. Finding the right information. In a very large context it was seventy-three percent and became ninety-six percent.
Far less chance of forgetting an important detail among hundreds of documents. And in fact that is a critical story. Although I want to say that 5.6 Sol, the Pro mode, seems to already work with an incredible volume of data. I see it on extremely complex tables with probably millions of different parameters. Very decent processing quality. But complex web interface search itself is almost unchanged. Web search, right? Not web interface, web search. It had ninety percent quality and it stays at ninety percent quality.
That is, in essence, we are today in an era when ChatGPT with a good version already searches superbly. Do we need to reach one hundred percent? It is not even clear what one hundred percent means in this measurement. I think seventy or eighty percent is already great. Right? Because what is one hundred? And then there are the indices. So that we understand there is no particular difference.
For example, the index, the general intelligence index, has stayed practically the same, at around sixty percent. Design creation. For example, if any of you are involved in that; Tanya and I will even be discussing it. Astra has had practically no effect on it, and practically nothing has changed in terms of work for programmers with ordinary code. That is, Astra has not become radically smarter in all matters. For simple things nothing much has changed. But when we talk about the automation of complex work in programs, that leap is truly cardinal, striking, right?
Or the questions I told you about from the very beginning. Where less needs to be found out from the person about what actually happened. What do you think about Astra? What do you think about Astra and Fable? What do you think about the Chinese systems? What lies ahead, what and where will affect this? See you in our new episodes on the ToTheMoon channel. Bye, everyone.