OpenAI Has Paused Its New Model. What Is Happening to the AI Race?
What matters more today — having the newest model or getting the most out of an existing one, when new versions do not always beat the old ones on real work?
The reader gets a working picture of the market as of August 2026: two ways to make AI stronger (improve the base model or the harness around it), why every add-on such as skills ends up inside the product six months later, and which model combinations actually work for people who use them for daily work.
What to watch for
Key takeaways
The habit was simple: a new version came out, so it was stronger than the last. On real work over recent months it is increasingly the other way round: older versions turn out to be steadier and sometimes do better. The exception the participants name straight away is Fable.
Internal tests always throw up uncontrolled situations, and companies normally do not report them. Here the statement lands on the eve of Anthropic's very expensive IPO, with the company reporting record revenue. Hence the counter-reading: is this a pitch about some supernatural model OpenAI supposedly has.
Given the large companies' relations with the US government, a model that is too free and too strong simply will not be released — it will be locked down. «Our model is already so good that improving it further is unnecessary» may be not a boast but a description of a ceiling there is otherwise no getting through.
After the incidents OpenAI began tracking not only tool calls but the reasoning trace itself — how the model thinks on the way to an action. That is an enormous additional volume of data: by some estimates up to twenty or twenty-five per cent of training compute. If a trigger fires and cannot be resolved automatically within thirty minutes, the process stops until it is examined by hand.
The next numbered version was expected by the end of August, and Astra would have been ready but for these problems. And after the break-in stories, passing US government review will not be simple. Anthropic's timing felt similar, but there is no further news yet.
That was the moment when a system first did a proper architecture review, large analyses and serious blocks of code. Then Fable brought not duration and not consistency over a long path but a different quality of result: as if the task had been taken over from a fifteen-year-old by an adult.
After Opus 5 and the Luna, Sol, Terra family something went wrong: endless glitches, and in the host's practice 4.8 solves tasks better than Opus 5. He still goes back to 4.8 where Fable is unavailable. Meanwhile in 5.6 Sol Pro the mode selector has been resetting for six months — a small thing that keeps repeating.
The question about OpenAI's strategy is direct: what do they want to build. A partnership opening a free software-development service on Terra looks odd when ChatGPT already has Terra free by default and Codex is available to anyone for twenty dollars.
The Cursor deal is closed — with xAI owned by SpaceX and the product called Grok. Cursor is a very serious system, and combined with Grok that is full competition for Codex and Claude. Google is not retreating either: 3.7 Flash was released precisely for agent programming tasks.
Anthropic's goal is concrete — AGI safety, with programming solved along the way; you can see it in the research, the essays and the constitution. OpenAI builds a very universal product, applicable anywhere. That is both a blessing and a curse: universality makes it hard to move in one direction.
Throughout the podcast's existence there have been add-ons alongside the base level. First you had to write a good prompt — there were whole guidelines. Then came API connections, then MCP and splitting a task across sub-agents. The base model meanwhile changed far less, while the situation changed dramatically.
Today's thing is skills — like the one that helps you work through a subject by questioning and build a map of the project. They can put you a month ahead of everyone else, but all of it will be integrated inside Claude and ChatGPT and become available to everyone out of the box. That is what happened to every previous add-on.
Nobody demands that superintelligent models be built urgently. In parallel you can improve the current work — not move a button but fix what has not changed in six months: the linear project-management structure, awkward names, glitches that leave you unsure whether something failed. Those small things decide whether a person sticks with a system or leaves.
5.6 Sol and its Ultra mode are very expensive models. If you are paying seriously per request, you will wait. The question about Ultra Fast mode is direct: who is it for, when anyone using Sol Ultra or Extra High is in no hurry. The same question applied to ChatGPT Pro at a hundred dollars.
Tatiana never took up Codex: what it offered were basic steps she had reached on her own. It produced them in five or ten minutes against her several months, but she could not take it further — and the system got a cross against it, even though it has changed since.
Someone with a twenty-dollar subscription does not know what else is possible — the same story as the mode selector nobody moved. And separately: a shared subscription «with dad» was normal two years ago, while today twenty dollars a month seems expensive to someone who easily buys a ten-dollar coffee. That is not meanness but unawareness — and it is the global reality.
Ilnar gathers all the context in ChatGPT — for a project, for experiments, sometimes several months of work — breaks it into meaningful blocks together with the model, and then takes it to Claude Design along with references. Two or three iterations give a result he would never have assembled by hand. Claude Design will not gather the context itself, but it handles the presentation.
The new ToTheMoon system: cards for every episode, entity markup, coverage of thirty-two thousand tag intersections, eight hundred and ninety-two entities, a hundred and fifty-three episodes — and a search that shows where, when and in what context something was mentioned, down to the minute. Claude Design did the design, Fable the markup and development, and Claude Design wrote the search architecture on the first attempt.
What this episode is about
Sam Altman has said OpenAI is pausing the training of its next frontier model, Astra — the generation after the 5.6 Sol line — over questions about its behaviour and safety. The host's first reaction is simple: why write about it at all? Internal tests always throw up uncontrolled situations, and normally nobody reports them. Especially not on the eve of Anthropic's very expensive IPO, with the company reporting record revenue and talking about hundreds of billions a year.
Ilnar Shafigullin offers two readings. The first is regulatory: given the large companies' relations with the US government, a model that is too free and too strong simply will not be released, and «no need to improve it further» may be an honest description of a ceiling. The second is technical: after the incidents, OpenAI began tracking not only tool calls but the model's reasoning traces. That is an enormous additional volume of data — by some estimates up to twenty or twenty-five per cent of training compute now goes on analysing those traces. And if a trigger fires and the problem cannot be resolved automatically within thirty minutes, the whole process halts until it is examined by hand. A two-week pause may be exactly that.
The consequence is that GPT-6 slips. The end of August was expected; Astra would have been ready but for these problems. And after the break-in stories, passing US government review will not be simple either.
A separate thread is the degradation the participants observe in practice. The peak was ChatGPT 5.5 Extra High and Opus 4.8 Max; then Fable brought a fundamentally different quality of result. After that came Opus 5 and the Luna, Sol, Terra family — and in the host's experience 4.8 solves tasks better than Opus 5. Meanwhile in 5.6 Sol Pro the mode selector still resets — a small thing that has been repeating for six months.
Hence the question of strategy: AGI, coding, or business chat? Anthropic moves toward a concrete goal — AGI safety, with programming solved along the way — while OpenAI tries to move in every direction with a universal product, which is both a blessing and a curse.
The programming market is being rebuilt meanwhile: xAI has closed its deal with Cursor and is clearly entering it seriously, and Google is not retreating, having released 3.7 Flash for agent tasks.
Ilnar's central point is about two routes. You can improve the base model, or you can extract more from the current one. Throughout the podcast's existence there have been add-ons alongside the base level: first prompt guidelines, then connecting services over an API, then MCP and sub-agents, now skills. Each gave an advanced user roughly a month's head start and then turned up built into the product for everyone.
The second half is practice. Tatiana explains why she never took up Codex: an old impression of an early version and the feeling that the system has not changed. They discuss how a huge share even of paying ChatGPT users do not know what else they can do — the same story as the mode selector nobody moved. Ilnar describes his working scheme: context is gathered in ChatGPT, the presentation is made in Claude Design, and two or three iterations give a result he would never have assembled by hand.
The episode closes with a demonstration of the new ToTheMoon system: cards for every episode, entity markup, coverage of thirty-two thousand tag intersections, eight hundred and ninety-two entities, a hundred and fifty-three episodes, and a search that does all of it in milliseconds. Claude Design did the design, Fable the markup and development, and Claude Design wrote the search architecture on the first attempt.
The Astra pause reads both ways, and both are being used: «we cannot handle what we built» and «we built something so strong we stopped ourselves». What matters more is that the market is not hurrying anyone. What can be improved is not the base model but the harness around it, and that is already happening: prompt guides, MCP, sub-agents, orchestration, skills — each add-on first gives an advanced user a month's head start and then turns up inside the product for everyone. The winner is not whoever has the newer model but whoever understands what their task is made of.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 111 segments: 111 identified, 0 mixed, 0 marked with ✓, and 0 unresolved.
Loading…