GPT-6 Astra — A New Leap for AI? Where It Has Become 2–3 Times Stronger
Where does Astra really leap ahead of GPT-5.6 and Fable, where does the new architecture remain pretty numbers in OpenAI's tests — and what does that change for someone who uses the models every day?
The reader will see the numbers from OpenAI's internal tests of Astra: 64% against 22% on scientific tasks with a terminal, 41% against 18% on automating professional work, 92% against 76% on understanding an interface, 4% false statements instead of 12%. They will understand why the host sees Astra not as a "smarter chat" but as the beginning of systems that understand a person, and why he does not call it an agent. They will learn what the critical cybersecurity level means and which restrictions follow from it. And they get a sober frame: for simple requests the difference from 5.6 will be small, while web search, the general intelligence index and ordinary code have barely moved.
What to watch for
Key takeaways
OpenAI has released Astra, and it is already GPT-6: not the internal 5.6 branches like Sol and Luna and not 5.5, but a different line. The difference is of the same order as between GPT three, four and five, and for the first time the company has placed one of its models at the critical capability level in cybersecurity.
Sceptics say any model will be distilled and replicated. Moonshot's Kimi K3 claimed the level of Fable and 5.6 Sol — even their PR service wrote to the host about it — yet a month later there has been no explosion of worldwide use, just as OpenAI's 5.6 never showed Fable's level.
In the host's view OpenAI and Anthropic recently drew level, although OpenAI keeps a large lead in daily users. Astra, together with the company's own infrastructure line — the Jalapeño chips, servers and software — could put it back in the leading position.
In OpenAI's internal tests against GPT-5.6 Sol, Astra completes 64% of scientific tasks with programs and a terminal against 22% — the host calls it a fundamental leap. He admits 5.6 Sol disappointed him and in recent months he built systems mostly on Claude while remaining a daily ChatGPT user.
41% against 18%: Astra is better than the previous models at knowing when it can make a reasonable decision on its own and when it should check with the user. The host suggests judging effectiveness seriously in September and October, recalling that 5.6 Sol, in his view, lost precisely on its own reasoning.
Opus 5 has Ultra Code, 5.6 Sol had Extra High and Pro, and these are fundamentally different capabilities with different outputs. For a simple request like choosing a restaurant the difference is unclear; for serious tasks the progress between modes is obvious — worth remembering when reading any comparison.
Complex programming tasks through the terminal — 57% against 37%, understanding computer interface elements — 92% against 76%, false statements — 4% instead of 12%. Astra's biggest gains are in agentic work and programming; OpenAI names science as well.
The host insists that hundreds of processes run at once in a person's life, which is harder than any scientific task. Three years ago the GPT era began with jokes and letters; the real shift will come when a system genuinely understands who you are — and, judging by Astra's description, it is heading there, memory included.
If you need a model to find errors in a letter or look up a country's population, you will not notice the leap. A million-token context and a 128-thousand-token answer are abstract numbers, and promises to handle forms, calendars and mail better came from every new model of the past year. The difference lies in the new architecture, not in a demo with a rocket and a 3D world.
OpenAI has claimed it already has AGI; Sam Altman said the technological singularity has essentially been reached. The host is reserved: if there is a singularity, who controls it? In theory such a system cannot be controlled at all.
From the wording "Astra knows when it can make a reasonable decision" the host draws a long-standing demand: models should ask the user questions and look at why they offer something at all. For now OpenAI emails a suggestion to "try voice" to someone who has used it for ages — the interfaces are built without the AI itself.
A marketer working through a bot in Notion, Slack, Telegram, an ERP or a CRM cannot see what chain of software sits behind it or whether there is a model at all, and thinks it is the same as Codex or Claude Code directly. The same is lost in aggregators such as Perplexity and third-party document tools. To build the habit you have to use the models from the inside — and strong models will be available less and less in free modes.
At the time of recording Astra is open in the paid versions — Plus, Pro, Business, Enterprise — and is rolling out by country, the US first. By some descriptions it is so far only in Pro at 200 dollars with a limit of 200 chat messages a week; Codex and Work are counted separately. All of this can change daily.
In trials without public restrictions Astra found unknown vulnerabilities and built working ways to exploit them — as Fable had before it. So the public version got stricter restrictions and monitoring of action sequences, and, in the host's words, no system today can promise one-hundred-percent confidentiality.
The host does not consider Fable or Astra agents and does not call them a "smarter chat". Real artificial intelligence for him is what films like Oblivion or Ready Player One show, and that is still far off, but Astra is closer to it: a world or a space you are in that does something.
In automating complex work in programs — accounting, sales, marketing, editing, management, analytics, finance — the gain is more than twofold. For Codex users, from 37% to 57%. Interface understanding rose from 75% to 92%, but 8% errors on a contract or a tax form is still far too expensive: it needs fractions of a percent and self-checking.
Scientific work through terminal programs is the most practical leap: large datasets and understanding a complex context, which the host discussed in connection with Anthropic's study of agents. Finding the right information in a very large context rose from 73% to 96% — fewer chances of forgetting a detail among hundreds of documents.
Complex web search stays at 90% — ChatGPT with a good version already searches superbly, and nobody knows what 100% would mean. The general intelligence index holds at around 60%; design and programmers' work with ordinary code have barely moved. Astra has not become radically smarter at everything: the cardinal leap is in automating complex work and in the model asking fewer questions back.
What this episode is about
A solo ToTheMoon episode on the release of Astra — the first model of the GPT-6 line. Alexander Volchek recalls the context: a few days earlier, on the Sunday podcast with Ilnar and Tanya, they discussed Fable 5.1 and were waiting for Astra — and now it is out. It is neither GPT-5.6 with its internal Sol and Luna branches nor 5.5, but a different line, and the difference is of the same order as between GPT three, four and five. To sceptics who say the Chinese will replicate any model, the host answers with Kimi K3: the claimed level of Fable and 5.6 Sol never turned into an explosion of use. Together with its own infrastructure — the Jalapeño chips, servers and software — Astra could return OpenAI to the lead over Anthropic, with whom, in his view, the two companies had recently drawn level.
Then come the numbers from OpenAI's internal tests against GPT-5.6 Sol. Scientific tasks with programs and a terminal: 64% against 22%. Automation of professional work: 41% against 18% — the model is better at knowing when it can decide on its own and when it should check with the user. Complex programming tasks through the terminal: 57% against 37%. Understanding interface elements: 92% against 76%. False statements: 4% instead of 12%. The host admits he was left disappointed by 5.6 Sol's reasoning and in recent months has built systems mostly on Claude, with the same complaints about Opus 5 — and reminds that the Extra High, Pro and Ultra Code modes give fundamentally different capabilities, so Astra's effectiveness can only be judged seriously in September and October.
The second half of the episode is about what those numbers mean. For requests of the "find the errors in this letter" kind the difference from 5.6 will be small; a million-token context and a 128-thousand-token answer are abstract numbers. Every new model of the past year promised to handle forms, calendars and mail better, and OpenAI's Astra video with a rocket and a 3D world does not explain the fundamental difference either. The host turns to AGI and the singularity: OpenAI claims it already has AGI, Sam Altman says the singularity has been reached, but if it has, such a system cannot be controlled. From the wording "Astra knows when it can make a reasonable decision" he extracts his long-standing demand: models should ask the user about themselves, rather than email someone who has used voice for ages a suggestion to "try voice". A separate thread is the vacuum in which people live when they work through a corporate bot in Notion, Slack, Telegram, an ERP or a CRM: they think it is the same as working directly in Codex or Claude Code, and it is not; the same is lost in aggregators such as Perplexity and in third-party document tools. Hence the advice: to build the habit, use the models from the inside — and expect Fable and Astra to be available less and less in free modes.
The episode closes on availability and limits. Astra is open in paid plans, starting with the US; by some descriptions, for now only in Pro at 200 dollars with a limit of 200 messages a week, with Codex and Work counted separately. The main point is the critical level in cybersecurity: in trials without public restrictions the model found unknown vulnerabilities and built working ways to exploit them, as Fable had before it; hence the stricter restrictions on the public version and the monitoring of action sequences, to which the host adds that no system today can promise one-hundred-percent confidentiality. He does not consider Fable or Astra agents: they are closer to artificial intelligence as a space you are in. In practice the leap shows in automating complex work in programs — more than twofold — for Codex users, in understanding interfaces (though 8% errors on a tax form or a contract is still expensive), in scientific work through a terminal, and in finding the right detail among hundreds of documents: from 73% to 96%. Almost unchanged: web search (90% and still 90%), the general intelligence index at around 60%, design and ordinary code for programmers.
The episode is useful precisely because it gives in neither to enthusiasm nor to scepticism. OpenAI's numbers show a leap where the model has to work long and on its own — in automating complex work, in the terminal, across large sets of documents — and show almost nothing where the ordinary user lives: in search, in simple requests, in design. The host's main point is not made in percentages: Astra is neither a smarter chat nor an agent but a first step towards systems that understand a person, and the price of that step is already named — the critical cybersecurity level, stricter restrictions and the end of free access to the strong models.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 47 segments: 47 identified, 0 mixed, 0 probable, and 0 unresolved.
Loading…