Skip to content
OpenAI · Astra · Artificial intelligenceEpisode 154 · 23 August 2026 · 54:37

OpenAI Has Paused Its New Model. What Is Happening to the AI Race?

Central question

What matters more today — having the newest model or getting the most out of an existing one, when new versions do not always beat the old ones on real work?

What you take away

The reader gets a working picture of the market as of August 2026: two ways to make AI stronger (improve the base model or the harness around it), why every add-on such as skills ends up inside the product six months later, and which model combinations actually work for people who use them for daily work.

Main threads

What to watch for

1Do not carry an old impression of a product over to its current version: what did not work a year ago in Codex or Claude Design is worth retrying.
2Before waiting for a new model, look at the harness: MCP, sub-agents, skills and built-in browsers change the result more than a version number.
3Do not tie yourself to one system for everything: context is easier to gather in one, presentation to build in another, development in a third.
4If results got worse after an update, go back to the previous version and compare on your own task — 4.8 against Opus 5 is the concrete example here.
5When choosing between a separate service and Claude Code or Codex, think six months ahead: a standalone feature usually ends up built into the large product.
6Learn the whole chain at least well enough to follow it — architecture, systems analysis, markup, cookies, consent, transcription: without that one person cannot assemble today's result.
Signals to track afterwards
Whether Astra and the next numbered ChatGPT version arrive in the autumn — or the pause turns out longer than announced.
Whether the models pass US government review after the break-in stories.
What the Cursor and Grok combination does to the programming market against Codex and Claude.
Whether Anthropic's next version appears and whether its timing coincides with the IPO.
Whether skills become a technology of their own or dissolve inside Claude and ChatGPT, as every add-on before them did.
Whether Claude Design gets built into Claude Code — right now you have to export between them, and the export loses files.
Most useful for
Anyone working daily with ChatGPT, Claude and Codex who has noticed that new versions are not always better.Developers choosing between Codex, Claude Code, Cursor and standalone services.Anyone following the model race who wants to tell a technical reason from a marketing one.Anyone making presentations and reports and looking for a working combination of tools.Founders planning to build a product single-handed who want to know which roles they will have to cover.Anyone curious about how the new ToTheMoon system is built and what exactly built it.

Key takeaways

00:00A New Model No Longer Automatically Means «Better»

The habit was simple: a new version came out, so it was stronger than the last. On real work over recent months it is increasingly the other way round: older versions turn out to be steadier and sometimes do better. The exception the participants name straight away is Fable.

01:45Why Announce It at All

Internal tests always throw up uncontrolled situations, and companies normally do not report them. Here the statement lands on the eve of Anthropic's very expensive IPO, with the company reporting record revenue. Hence the counter-reading: is this a pitch about some supernatural model OpenAI supposedly has.

04:10The Regulatory Ceiling as an Honest Explanation

Given the large companies' relations with the US government, a model that is too free and too strong simply will not be released — it will be locked down. «Our model is already so good that improving it further is unnecessary» may be not a boast but a description of a ceiling there is otherwise no getting through.

05:30A Quarter of the Compute Goes on Watching the Reasoning

After the incidents OpenAI began tracking not only tool calls but the reasoning trace itself — how the model thinks on the way to an action. That is an enormous additional volume of data: by some estimates up to twenty or twenty-five per cent of training compute. If a trigger fires and cannot be resolved automatically within thirty minutes, the process stops until it is examined by hand.

07:30GPT-6 Slips

The next numbered version was expected by the end of August, and Astra would have been ready but for these problems. And after the break-in stories, passing US government review will not be simple. Anthropic's timing felt similar, but there is no further news yet.

08:30The Peak Was 5.5 Extra High and Opus 4.8 Max

That was the moment when a system first did a proper architecture review, large analyses and serious blocks of code. Then Fable brought not duration and not consistency over a long path but a different quality of result: as if the task had been taken over from a fifteen-year-old by an adult.

10:05A Sense of Degradation on Large Systems

After Opus 5 and the Luna, Sol, Terra family something went wrong: endless glitches, and in the host's practice 4.8 solves tasks better than Opus 5. He still goes back to 4.8 where Fable is unavailable. Meanwhile in 5.6 Sol Pro the mode selector has been resetting for six months — a small thing that keeps repeating.

12:17AGI, Coding or Business Chat

The question about OpenAI's strategy is direct: what do they want to build. A partnership opening a free software-development service on Terra looks odd when ChatGPT already has Terra free by default and Codex is available to anyone for twenty dollars.

14:08xAI Is Entering the Programming Market Seriously

The Cursor deal is closed — with xAI owned by SpaceX and the product called Grok. Cursor is a very serious system, and combined with Grok that is full competition for Codex and Claude. Google is not retreating either: 3.7 Flash was released precisely for agent programming tasks.

15:15A Concrete Goal Against a Universal Product

Anthropic's goal is concrete — AGI safety, with programming solved along the way; you can see it in the research, the essays and the constitution. OpenAI builds a very universal product, applicable anywhere. That is both a blessing and a curse: universality makes it hard to move in one direction.

16:38Two Routes: The Base Model or the Harness Around It

Throughout the podcast's existence there have been add-ons alongside the base level. First you had to write a good prompt — there were whole guidelines. Then came API connections, then MCP and splitting a task across sub-agents. The base model meanwhile changed far less, while the situation changed dramatically.

18:31Skills Give a Month's Head Start and End Up Inside the Product

Today's thing is skills — like the one that helps you work through a subject by questioning and build a map of the project. They can put you a month ahead of everyone else, but all of it will be integrated inside Claude and ChatGPT and become available to everyone out of the box. That is what happened to every previous add-on.

20:34The Market Is Not Hurrying Anyone

Nobody demands that superintelligent models be built urgently. In parallel you can improve the current work — not move a button but fix what has not changed in six months: the linear project-management structure, awkward names, glitches that leave you unsure whether something failed. Those small things decide whether a person sticks with a system or leaves.

22:32Who Exactly Cannot Wait

5.6 Sol and its Ultra mode are very expensive models. If you are paying seriously per request, you will wait. The question about Ultra Fast mode is direct: who is it for, when anyone using Sol Ultra or Extra High is in no hurry. The same question applied to ChatGPT Pro at a hundred dollars.

24:04An Old Impression Outlives the Product

Tatiana never took up Codex: what it offered were basic steps she had reached on her own. It produced them in five or ten minutes against her several months, but she could not take it further — and the system got a cross against it, even though it has changed since.

26:36Even Paying Users Do Not Know What Else Is Possible

Someone with a twenty-dollar subscription does not know what else is possible — the same story as the mode selector nobody moved. And separately: a shared subscription «with dad» was normal two years ago, while today twenty dollars a month seems expensive to someone who easily buys a ten-dollar coffee. That is not meanness but unawareness — and it is the global reality.

29:33The Working Combination: Context in ChatGPT, Design in Claude Design

Ilnar gathers all the context in ChatGPT — for a project, for experiments, sometimes several months of work — breaks it into meaningful blocks together with the model, and then takes it to Claude Design along with references. Two or three iterations give a result he would never have assembled by hand. Claude Design will not gather the context itself, but it handles the presentation.

36:30Thirty-Two Thousand Intersections and a Search in Milliseconds

The new ToTheMoon system: cards for every episode, entity markup, coverage of thirty-two thousand tag intersections, eight hundred and ninety-two entities, a hundred and fifty-three episodes — and a search that shows where, when and in what context something was mentioned, down to the minute. Claude Design did the design, Fable the markup and development, and Claude Design wrote the search architecture on the first attempt.

What this episode is about

Sam Altman has said OpenAI is pausing the training of its next frontier model, Astra — the generation after the 5.6 Sol line — over questions about its behaviour and safety. The host's first reaction is simple: why write about it at all? Internal tests always throw up uncontrolled situations, and normally nobody reports them. Especially not on the eve of Anthropic's very expensive IPO, with the company reporting record revenue and talking about hundreds of billions a year.

Ilnar Shafigullin offers two readings. The first is regulatory: given the large companies' relations with the US government, a model that is too free and too strong simply will not be released, and «no need to improve it further» may be an honest description of a ceiling. The second is technical: after the incidents, OpenAI began tracking not only tool calls but the model's reasoning traces. That is an enormous additional volume of data — by some estimates up to twenty or twenty-five per cent of training compute now goes on analysing those traces. And if a trigger fires and the problem cannot be resolved automatically within thirty minutes, the whole process halts until it is examined by hand. A two-week pause may be exactly that.

The consequence is that GPT-6 slips. The end of August was expected; Astra would have been ready but for these problems. And after the break-in stories, passing US government review will not be simple either.

A separate thread is the degradation the participants observe in practice. The peak was ChatGPT 5.5 Extra High and Opus 4.8 Max; then Fable brought a fundamentally different quality of result. After that came Opus 5 and the Luna, Sol, Terra family — and in the host's experience 4.8 solves tasks better than Opus 5. Meanwhile in 5.6 Sol Pro the mode selector still resets — a small thing that has been repeating for six months.

Hence the question of strategy: AGI, coding, or business chat? Anthropic moves toward a concrete goal — AGI safety, with programming solved along the way — while OpenAI tries to move in every direction with a universal product, which is both a blessing and a curse.

The programming market is being rebuilt meanwhile: xAI has closed its deal with Cursor and is clearly entering it seriously, and Google is not retreating, having released 3.7 Flash for agent tasks.

Ilnar's central point is about two routes. You can improve the base model, or you can extract more from the current one. Throughout the podcast's existence there have been add-ons alongside the base level: first prompt guidelines, then connecting services over an API, then MCP and sub-agents, now skills. Each gave an advanced user roughly a month's head start and then turned up built into the product for everyone.

The second half is practice. Tatiana explains why she never took up Codex: an old impression of an early version and the feeling that the system has not changed. They discuss how a huge share even of paying ChatGPT users do not know what else they can do — the same story as the mode selector nobody moved. Ilnar describes his working scheme: context is gathered in ChatGPT, the presentation is made in Claude Design, and two or three iterations give a result he would never have assembled by hand.

The episode closes with a demonstration of the new ToTheMoon system: cards for every episode, entity markup, coverage of thirty-two thousand tag intersections, eight hundred and ninety-two entities, a hundred and fifty-three episodes, and a search that does all of it in milliseconds. Claude Design did the design, Fable the markup and development, and Claude Design wrote the search architecture on the first attempt.

The Astra pause reads both ways, and both are being used: «we cannot handle what we built» and «we built something so strong we stopped ourselves». What matters more is that the market is not hurrying anyone. What can be improved is not the base model but the harness around it, and that is already happening: prompt guides, MCP, sub-agents, orchestration, skills — each add-on first gives an advanced user a month's head start and then turns up inside the product for everyone. The winner is not whoever has the newer model but whoever understands what their task is made of.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 111 segments: 111 identified, 0 mixed, 0 marked with ✓, and 0 unresolved.

Loading…