Skip to content
Astra · OpenAI · FableEpisode 162 · 9 September 2026 · 29:49

GPT-6 Astra — A New Leap for AI? Where It Has Become 2–3 Times Stronger

Central question

Where does Astra really leap ahead of GPT-5.6 and Fable, where does the new architecture remain pretty numbers in OpenAI's tests — and what does that change for someone who uses the models every day?

What you take away

The reader will see the numbers from OpenAI's internal tests of Astra: 64% against 22% on scientific tasks with a terminal, 41% against 18% on automating professional work, 92% against 76% on understanding an interface, 4% false statements instead of 12%. They will understand why the host sees Astra not as a "smarter chat" but as the beginning of systems that understand a person, and why he does not call it an agent. They will learn what the critical cybersecurity level means and which restrictions follow from it. And they get a sober frame: for simple requests the difference from 5.6 will be small, while web search, the general intelligence index and ordinary code have barely moved.

Main threads

What to watch for

1Match your tasks against the tests: if it is automating complex work in programs, a terminal or going through hundreds of documents, try Astra on a real process, not a demo.
2If you use the chat for simple requests, do not expect a noticeable difference from 5.6 — and do not pay extra for it.
3Where the cost of an error is high — a contract, a tax form — do not rely on a 92% hit rate in the interface: make the system recheck itself or check it yourself.
4Check whether you actually have access to the model itself: plan, region, message limit; Codex and Work are counted separately from the chat.
5If you only work with AI through a corporate bot or an aggregator, work with the model directly — otherwise you will not see the difference between generations.
Signals to track afterwards
Whether the internal-test numbers hold up in real work — the host promises to return to Astra's effectiveness in September and October.
Whether access widens: does Astra stay Pro-only or reach Plus, other countries and the free mode.
Which restrictions the critical cybersecurity level brings — and to whom OpenAI opens the model without them.
Whether models start asking the user about themselves instead of emailing offers of features already in use.
Whether Kimi K3 or another Chinese model shows an explosion of use rather than just a claimed level.
Most useful for
Anyone automating work in programs — accounting, sales, marketing, analytics, finance: this is where the tests show a twofold gain.Developers working in Codex or through a terminal: the leap from 37% to 57% concerns them directly.Anyone working through large sets of documents and data: finding the right detail in a large context rose from 73% to 96%.Leaders choosing a plan and a model for a team: availability, limits and modes are named here plainly.Anyone following the OpenAI–Anthropic contest and what lies behind the words AGI, singularity and critical safety level.

Key takeaways

00:00Astra Is the First GPT-6 Model and the First at the Critical Cybersecurity Level

OpenAI has released Astra, and it is already GPT-6: not the internal 5.6 branches like Sol and Luna and not 5.5, but a different line. The difference is of the same order as between GPT three, four and five, and for the first time the company has placed one of its models at the critical capability level in cybersecurity.

02:58"The Chinese Will Copy Everything" Is Not Yet Borne Out

Sceptics say any model will be distilled and replicated. Moonshot's Kimi K3 claimed the level of Fable and 5.6 Sol — even their PR service wrote to the host about it — yet a month later there has been no explosion of worldwide use, just as OpenAI's 5.6 never showed Fable's level.

04:20Together With the Chips, Astra Could Return OpenAI to the Lead

In the host's view OpenAI and Anthropic recently drew level, although OpenAI keeps a large lead in daily users. Astra, together with the company's own infrastructure line — the Jalapeño chips, servers and software — could put it back in the leading position.

05:07Scientific Tasks With a Terminal: 64% Against 22%

In OpenAI's internal tests against GPT-5.6 Sol, Astra completes 64% of scientific tasks with programs and a terminal against 22% — the host calls it a fundamental leap. He admits 5.6 Sol disappointed him and in recent months he built systems mostly on Claude while remaining a daily ChatGPT user.

06:04Automating Professional Work: the Model Knows When to Ask

41% against 18%: Astra is better than the previous models at knowing when it can make a reasonable decision on its own and when it should check with the user. The host suggests judging effectiveness seriously in September and October, recalling that 5.6 Sol, in his view, lost precisely on its own reasoning.

07:17Modes Decide More Than the Version Number

Opus 5 has Ultra Code, 5.6 Sol had Extra High and Pro, and these are fundamentally different capabilities with different outputs. For a simple request like choosing a restaurant the difference is unclear; for serious tasks the progress between modes is obvious — worth remembering when reading any comparison.

08:02Terminal, Interface, Errors: the Rest of the Test Numbers

Complex programming tasks through the terminal — 57% against 37%, understanding computer interface elements — 92% against 76%, false statements — 4% instead of 12%. Astra's biggest gains are in agentic work and programming; OpenAI names science as well.

08:58Ordinary Life Is Harder Than Science, and the Models Are Only Approaching It

The host insists that hundreds of processes run at once in a person's life, which is harder than any scientific task. Three years ago the GPT era began with jokes and letters; the real shift will come when a system genuinely understands who you are — and, judging by Astra's description, it is heading there, memory included.

10:39For Simple Requests the Difference From 5.6 Will Be Small

If you need a model to find errors in a letter or look up a country's population, you will not notice the leap. A million-token context and a 128-thousand-token answer are abstract numbers, and promises to handle forms, calendars and mail better came from every new model of the past year. The difference lies in the new architecture, not in a demo with a rocket and a 3D world.

13:28AGI Declared, Singularity Declared — but It Cannot Be Controlled

OpenAI has claimed it already has AGI; Sam Altman said the technological singularity has essentially been reached. The host is reserved: if there is a singularity, who controls it? In theory such a system cannot be controlled at all.

14:11Models Should Ask the User About Themselves

From the wording "Astra knows when it can make a reasonable decision" the host draws a long-standing demand: models should ask the user questions and look at why they offer something at all. For now OpenAI emails a suggestion to "try voice" to someone who has used it for ages — the interfaces are built without the AI itself.

15:28A Corporate Bot Is Not the Same as the Model Itself

A marketer working through a bot in Notion, Slack, Telegram, an ERP or a CRM cannot see what chain of software sits behind it or whether there is a model at all, and thinks it is the same as Codex or Claude Code directly. The same is lost in aggregators such as Perplexity and third-party document tools. To build the habit you have to use the models from the inside — and strong models will be available less and less in free modes.

18:56Availability: Paid Plans, the US First, and Limits

At the time of recording Astra is open in the paid versions — Plus, Pro, Business, Enterprise — and is rolling out by country, the US first. By some descriptions it is so far only in Pro at 200 dollars with a limit of 200 chat messages a week; Codex and Work are counted separately. All of this can change daily.

20:19Critical Level: the Model Found Vulnerabilities and Built Exploits

In trials without public restrictions Astra found unknown vulnerabilities and built working ways to exploit them — as Fable had before it. So the public version got stricter restrictions and monitoring of action sequences, and, in the host's words, no system today can promise one-hundred-percent confidentiality.

21:42Not a Smart Chat and Not an Agent, but a Space You Are In

The host does not consider Fable or Astra agents and does not call them a "smarter chat". Real artificial intelligence for him is what films like Oblivion or Ready Player One show, and that is still far off, but Astra is closer to it: a world or a space you are in that does something.

23:52Where the Leap Shows in Practice: Automation, Codex, Interfaces

In automating complex work in programs — accounting, sales, marketing, editing, management, analytics, finance — the gain is more than twofold. For Codex users, from 37% to 57%. Interface understanding rose from 75% to 92%, but 8% errors on a contract or a tax form is still far too expensive: it needs fractions of a percent and self-checking.

26:46Science Through a Terminal and Search in a Large Context Are the Most Practical Leaps

Scientific work through terminal programs is the most practical leap: large datasets and understanding a complex context, which the host discussed in connection with Anthropic's study of agents. Finding the right information in a very large context rose from 73% to 96% — fewer chances of forgetting a detail among hundreds of documents.

28:12Web Search, the Intelligence Index, Design and Ordinary Code Barely Changed

Complex web search stays at 90% — ChatGPT with a good version already searches superbly, and nobody knows what 100% would mean. The general intelligence index holds at around 60%; design and programmers' work with ordinary code have barely moved. Astra has not become radically smarter at everything: the cardinal leap is in automating complex work and in the model asking fewer questions back.

What this episode is about

A solo ToTheMoon episode on the release of Astra — the first model of the GPT-6 line. Alexander Volchek recalls the context: a few days earlier, on the Sunday podcast with Ilnar and Tanya, they discussed Fable 5.1 and were waiting for Astra — and now it is out. It is neither GPT-5.6 with its internal Sol and Luna branches nor 5.5, but a different line, and the difference is of the same order as between GPT three, four and five. To sceptics who say the Chinese will replicate any model, the host answers with Kimi K3: the claimed level of Fable and 5.6 Sol never turned into an explosion of use. Together with its own infrastructure — the Jalapeño chips, servers and software — Astra could return OpenAI to the lead over Anthropic, with whom, in his view, the two companies had recently drawn level.

Then come the numbers from OpenAI's internal tests against GPT-5.6 Sol. Scientific tasks with programs and a terminal: 64% against 22%. Automation of professional work: 41% against 18% — the model is better at knowing when it can decide on its own and when it should check with the user. Complex programming tasks through the terminal: 57% against 37%. Understanding interface elements: 92% against 76%. False statements: 4% instead of 12%. The host admits he was left disappointed by 5.6 Sol's reasoning and in recent months has built systems mostly on Claude, with the same complaints about Opus 5 — and reminds that the Extra High, Pro and Ultra Code modes give fundamentally different capabilities, so Astra's effectiveness can only be judged seriously in September and October.

The second half of the episode is about what those numbers mean. For requests of the "find the errors in this letter" kind the difference from 5.6 will be small; a million-token context and a 128-thousand-token answer are abstract numbers. Every new model of the past year promised to handle forms, calendars and mail better, and OpenAI's Astra video with a rocket and a 3D world does not explain the fundamental difference either. The host turns to AGI and the singularity: OpenAI claims it already has AGI, Sam Altman says the singularity has been reached, but if it has, such a system cannot be controlled. From the wording "Astra knows when it can make a reasonable decision" he extracts his long-standing demand: models should ask the user about themselves, rather than email someone who has used voice for ages a suggestion to "try voice". A separate thread is the vacuum in which people live when they work through a corporate bot in Notion, Slack, Telegram, an ERP or a CRM: they think it is the same as working directly in Codex or Claude Code, and it is not; the same is lost in aggregators such as Perplexity and in third-party document tools. Hence the advice: to build the habit, use the models from the inside — and expect Fable and Astra to be available less and less in free modes.

The episode closes on availability and limits. Astra is open in paid plans, starting with the US; by some descriptions, for now only in Pro at 200 dollars with a limit of 200 messages a week, with Codex and Work counted separately. The main point is the critical level in cybersecurity: in trials without public restrictions the model found unknown vulnerabilities and built working ways to exploit them, as Fable had before it; hence the stricter restrictions on the public version and the monitoring of action sequences, to which the host adds that no system today can promise one-hundred-percent confidentiality. He does not consider Fable or Astra agents: they are closer to artificial intelligence as a space you are in. In practice the leap shows in automating complex work in programs — more than twofold — for Codex users, in understanding interfaces (though 8% errors on a tax form or a contract is still expensive), in scientific work through a terminal, and in finding the right detail among hundreds of documents: from 73% to 96%. Almost unchanged: web search (90% and still 90%), the general intelligence index at around 60%, design and ordinary code for programmers.

The episode is useful precisely because it gives in neither to enthusiasm nor to scepticism. OpenAI's numbers show a leap where the model has to work long and on its own — in automating complex work, in the terminal, across large sets of documents — and show almost nothing where the ordinary user lives: in search, in simple requests, in design. The host's main point is not made in percentages: Astra is neither a smarter chat nor an agent but a first step towards systems that understand a person, and the price of that step is already named — the critical cybersecurity level, stricter restrictions and the end of free access to the strong models.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 47 segments: 47 identified, 0 mixed, 0 probable, and 0 unresolved.

Loading…