Skip to content
Anthropic · Claude Opus 5.5 · Artificial intelligenceEpisode 170 · 25 September 2026 · 41:02

Claude Opus 5.5 Is Out. Is It Even Worth Switching to New Models?

Central question

If model capabilities have grown so much in recent months, why have people and companies not become five or ten times more effective — and is it even worth switching to every new model?

What you take away

Claude Opus 5.5 arrived with the claim that it beats Fable, costs less and works better in medium mode than earlier models on maximum tariffs, and on the same day Codex dropped 5.6 Sol for GPT-6 Astra. The host checks it against practice: his sister's scheduled task stopped running in Astra, in Claude Design he keeps Opus 4.8 Max, Astra in Pro mode spent hours on a simple report, and Claude on Fable 5.1 ate his subscription tokens in two hours. On the tests Opus 5.5 gives 66% in the terminal against 55 for Fable and 57 for Astra, and Astra wins only on science, 64 against 58. Alexander Volchek shows that hiring, software and company processes barely changed, he sees a real shift in one person out of twenty or thirty, and that 0.1 percent of the world got super-opportunities — and that the main question is not adapting to new models but whether they really pay off.

Main threads

What to watch for

1Do not switch to a new model automatically: check on your regular tasks whether it has started doing them worse, as happened to the scheduled task of the host's sister in GPT-6 Astra.
2Keep the working version where it is stable: the host still uses Opus 4.8 Max in Claude Design because Fable and Opus 5 did not understand the task.
3Compare the final result, not tokens: ask whether the result was achieved before trusting a promise to spend 30–50% fewer tokens.
4Do not launch heavy tasks on the most expensive model without counting the cost: Claude on Fable 5.1 ate all the host's subscription tokens in two hours.
5Check whether you need your own software at all: the host replaced his heart-sensor program and MATLAB with Codex commands — start the sensor, download the data, build a report.
6If you are on the free version and cannot use up your tokens, give yourself a serious task — in the host's words, the free version should run out instantly.
Signals to track afterwards
→Whether OpenAI merges the chat, Codex and Work into one app with shared tokens, as the project led by Greg Brockman promised.
→Whether Fable gets degraded again after the Opus 5.5 release, as viewers say happened in July — and which model non-US citizens are using under the name Fable.
→Whether Google holds parity with Gemini on the strength of its search and services, or the lead of Anthropic and OpenAI becomes final.
→Whether the systems start surveying people and measuring a real person's life rather than programming in the terminal.
→Whether ChatGPT Sites takes the CMS market: in a year, by the host's forecast, new websites will be written outside CMS systems.
Most useful for
Anyone who hesitates every time over whether to move to a new model: the cases where the old version proved more stable than the new one are gathered here.Users of paid ChatGPT and Claude subscriptions: a breakdown of the Pro, Extra High and Max modes, the intelligence index and where subscription tokens go.Managers and entrepreneurs: why a fivefold gain in efficiency does not turn into a company of five hundred, and what employers expect.Developers and website owners: the ChatGPT Sites example and the forecast that new sites will be written outside CMS.Anyone building a business on aggregators and booking services: the host explains why OpenAI has more data on restaurants and hotels than they do.Everyone who feels AI has not yet changed their life: the episode asks that question directly and shows who got the advantage.

Key takeaways

00:19Models Grow Faster Than Our Lives Change

In recent months Fable, GPT-6 Astra and now Claude Opus 5.5 have come out with new records and claims from Anthropic, but between how fast the models grow and how life, work and business change the host sees a crazy gap. Three years of endless updates — ChatGPT versions from 4.0 to 6.0, a whole Anthropic line — have moved the discussion from punctuation mistakes to hacks with Mythos-family models and to the robust writing of websites and apps. The episode's questions: where new models give a different level of result, which to choose and what to change in one's own work.

02:27New Models Seem to Degrade — and Are Retired Before People Can Test Them

A task the host's sister had run on a schedule for six months stopped working in GPT-6 Pro — that is, in Astra — although almost the version before last had handled it. In Claude Design he still uses Opus 4.8 Max: with Fable and Opus 5 the system seemed not to understand what to do, and Astra does not produce the reports that version 5.4 did. On the day of recording Codex said it no longer supports 5.6 Sol on the Pro subscription and offered to migrate to GPT-6 Astra, although in 5.6 the host had stability; OpenAI released the Terra, Luna and Sol line and moves on without letting people test it.

05:46Do the New Models Help Anthropic's and OpenAI's Business — and What Is the Race For

The companies are interested in their own result — AGI, ASI, the technological singularity or systems that give them serious advantages, up to hacking states and economies. But do the models of the last half year help Anthropic or OpenAI gain users and earn more from them? The host sees no direct result, although the systems have improved incredibly. If the whole world is given the ability to open a business, there will be chaos of mass competition; AI's impact on countries' GDP last year was small, but the changes in how people behave and communicate are already irreversible.

09:21The Host Sees Real Change in Work in One Person Out of Twenty or Thirty

Websites, reports and app design have become a simplified task, but people stay the same. In many meetings with companies the host honestly sees one case in twenty or thirty where approaches to work have changed and there is a positive effect from systems like ChatGPT. His wife, busy with the house and four children, has used ChatGPT on paid subscriptions from day one — but has her life really changed? He asks viewers the same: have they started earning more, do they have free time and thanks to what — for instance, eight hours of work done in one, with the employer kept in the dark.

11:10A Fivefold Efficiency Gain Has Not Turned Companies of a Hundred Into Companies of Five Hundred

Employers expect the saved hours to be filled with a huge amount of other work. If in a company of a hundred people finance, analytics, marketing and editing gained x5 or x10 efficiency with ChatGPT, has it effectively become five hundred people — or do people just generate more documents and text? The software the host uses has shown no miracle: Booking, Airbnb, Google Drive, Gmail and Google search do not give the right hints or new interfaces, although inside, systems capable of analyzing hundreds of millions of people are being built.

14:48Aggregators Will Disappear, and the Systems Still Do Not Ask People Because That Is Not Their Goal

Systems like OpenAI have far more data about restaurants, cafés and hotels than any aggregator like Booking: when travelling, the host himself asks ChatGPT about the best seats on a plane more than he asks Expedia, Booking or the operators who sold the ticket. Revolutionary changes, in his view, will begin when the systems start surveying people, but that is not happening — the companies do not care about people: they measure terminal programming and business processes, not a real person's life. The smart home has not changed in several years, car assistants, even Tesla's, answer slowly, and there is little real value so far.

20:32Compare the Result, Not the Tokens: Astra Spends Hours on a Simple Report

Opus 5.5 promises to spend 30–50% fewer tokens, but people do not understand what tokens are, and systems should be compared by the final result. The host replaced his own heart-sensor software on the Mac and MATLAB with Codex: start the sensor for the night, download the data, build a report, send it to the doctor, compare it with three years of cardiograms. Yet GPT-6 Astra in Pro mode spent hours on a report covering a few days, whereas ChatGPT-5.5 or 5.0 spent far less at the same efficiency, and Claude on Fable 5.1 ate all his subscription tokens in two hours.

24:31Opus 5.5 on the Tests: 66% in the Terminal Against Fable's 55, While Astra Wins the Science Tasks

Opus 5.5 scores sixty-six percent on terminal work against fifty-five for Fable and about fifty-seven for Astra, and fifty-four in programming against fifty for Fable; the only task where GPT-6 Astra wins significantly is science, fifty-eight against sixty-four. Headlines about a website, a 3D world and eight hundred presentations from one prompt do not impress the host: ChatGPT Pro in Extra High mode did that half a year ago. The main question: if Fable can hack everything in sight, what is Opus 5.5, which outpaces it — or was Fable made worse, as in July.

27:16Too Many Products and Modes: Codex, Work, Cowork and the Intelligence Index

ChatGPT's chat tokens are not counted together with Codex and Work tokens, and Greg Brockman, who took on unifying everything into one app in late spring or June, has not yet removed the split between the chat and Codex; at Anthropic people likewise choose between Cowork, Code and plain Claude without understanding why the system goes into one mode or another. Opus 5.5 has modes from low to max with an intelligence index of 58 at max and 56 at extra high — Anthropic's own scale, not a percentage of correct answers and not an IQ. Opus 5.5 Low costs ten times less than Max with an index of 42 against 58.

31:01No Competitors in Sight: Anthropic and OpenAI Have Pulled Away, Google Holds On Thanks to Search

Over the last half year Anthropic and OpenAI have gained a striking lead: Grok, Moonshot Kimi and DeepSeek were visible in the past, and now nobody is close. Google will stay in the race thanks to its search and services, but without them, in the host's view, Google Gemini would have huge problems. The new iPhone and Siri interest nobody: the host has the eighteenth iPhone with twice the compute for AI, but it makes no difference to him how artificial intelligence works there. No system yet works in different places at once, and Claude Code does not accept files over twenty megabytes and does not work with the file system.

34:44Fable Was Blocked for Non-US Citizens, and the Markets That Could Be Taken Go Untaken

Anthropic released Fable and blocked it after a US government decree: the model may not be used by non-citizens, including Anthropic's own employees — and it is unclear which model is used now. If Anthropic or OpenAI have no restrictions and infinite compute, why do they not take whole industries with one or two percent of their compute? An example exists: ChatGPT Sites creates and hosts websites — more than ten million, according to Brockman — and is taking the CMS market; in a year, the host believes, new sites will be written outside CMS systems, and marketing and sales will stop ordering development.

38:330.1 Percent of the World Got Super-Opportunities, While Eight Out of Ten Resist

Literate people have gained a chance to earn a lot of money fast and to gain contacts, but it went to 0.1 percent of the world — those with the head to take the right decision. Eight out of ten, by the host's observation, resist, and AI remains very relative for them: returning from Europe to the States, he watched people spend the ten hours of the flight in awkward Excel files with accounting data, rechecking everything by hand and not opening ChatGPT or Anthropic. The episode's question is not whether these people will be able to adapt, but whether it will really give them an effect.

What this episode is about

A solo ToTheMoon episode about the release of Claude Opus 5.5 — and about what the release of yet another model means at all. Alexander Volchek starts with Anthropic's records and claims, recalls that Fable and GPT-6 Astra came out in recent months, and records a crazy gap between how fast the models grow and how life, work and business change. But in parallel the models seem to degrade: the host's sister's task stopped working in GPT-6 Pro, in Claude Design he still uses Opus 4.8 Max, and Astra does not produce the reports that version 5.4 did.

The host does not want to treat Opus 5.5 as just another release — what interests him is how models are released in general. On the day of recording Codex told him it no longer supports 5.6 Sol on the Pro subscription and offered to migrate to GPT-6 Astra, although in 5.6 he had stability; OpenAI released the Terra, Luna and Sol line and moves on without letting people test it. The reason, in his view, is that OpenAI and Anthropic are interested in their own result — AGI, ASI, the singularity — and it is a big question whether the models of the last half year help their business: no direct result is visible, although the systems have improved incredibly. If the whole world is given the ability to open a business, there will be chaos of mass competition; the host calls AI's impact on GDP last year small, although the changes in people's behaviour are already irreversible.

Next — people who stay the same. Websites, reports and design have become a simplified task, but at many meetings with companies the host sees one case in twenty or thirty where approaches to work have really changed. His wife, with four children, has used ChatGPT on paid subscriptions from day one — has her life really changed? He asks viewers the same: have you started earning more, do you have free time, or do you do eight hours of work in one and not tell your employer. Employers, for their part, expect you to fill the saved hours with a huge amount of other work, and if efficiency in a company of a hundred people grew fivefold, has it effectively become five hundred people — or just more documents and text.

Software has shown no miracle: Booking, Airbnb, Google Drive, Gmail and Google search do not give the right hints or new interfaces, although inside, systems capable of analyzing a state and hundreds of millions of people are being built. The host repeats that aggregators will disappear: OpenAI has more information about restaurants and hotels than any aggregator, and he himself asks ChatGPT about seats on a plane more than he asks Expedia or Booking. The revolution will begin when systems start surveying people — but that is not happening, because the companies do not care about people: the metrics measure programming, not a real person's life, the smart home and car assistants, even Tesla's, work crookedly, and there is little real value so far.

Astra's positioning — a system that communicates less with the human — the host considers new, but he recalls that the previous one did not ask questions either; his second guess is that the companies fear super-access, which would allow surveying a billion, a hundred million or ten million users. Hence — tokens: Opus 5.5 promises to spend 30–50% fewer of them, but people do not understand what tokens are, and what should be compared is the final result. His case is a heart sensor whose data Codex now collects and analyzes for him. Meanwhile GPT-6 Astra in Pro mode spent hours on a report covering a few days, whereas ChatGPT-5.5 or 5.0 spent less at the same efficiency, and Claude on Fable 5.1 ate the subscription tokens in two hours.

The host gives the Opus 5.5 tests with a caveat: 66% in the terminal against 55 for Fable and 57 for Astra, 54% in programming against 50 for Fable, while Astra wins only on science tasks — 58 against 64. Headlines about a website, a 3D world and eight hundred presentations from one prompt do not impress him: ChatGPT Pro in Extra High mode did that half a year ago. His main question: if Fable can hack everything in sight, then what is Opus 5.5, which outpaces it — or was Fable made worse, as in July. A poll on the channel showed that more than thirty percent of viewers had not even used the new Astra.

A separate knot is modes and products. At ChatGPT the chat tokens are not counted together with Codex and Work tokens; Greg Brockman has been working since late spring or June on unifying everything into one app, but the chat and Codex are still separate, and at Anthropic people choose between Cowork, Code and plain Claude. Opus 5.5 has modes from low to max with an intelligence index of 58 at max and 56 at extra high — Anthropic's own scale, in the host's words — and Low costs ten times less than Max with an index of 42 against 58. No competitors are visible: Anthropic and OpenAI have pulled away, Grok, Moonshot Kimi and DeepSeek are in the past, Google holds on thanks to search, and Siri in the new iPhone 18 interests nobody.

In the finale — limits and advantage. The systems have become slippery on questions of hacking, although they could warn of a risk zone; after a US government decree Anthropic blocked Fable for non-citizens, including its own employees, and it is unclear which model is now used under that name. And if the companies have no restrictions, why do they not take whole markets with one or two percent of their compute — an example exists: ChatGPT Sites hosts, according to Brockman, more than ten million websites and is taking the CMS market. The super-opportunities went to 0.1 percent of the world, eight out of ten resist, and people on a plane spend ten hours in awkward Excel files without opening ChatGPT. The episode's question is not whether they will adapt, but whether it will really give them an effect.

The episode's value is that it does not repeat the press release: the host checks Opus 5.5's records against his own tasks, where new models behave worse than old ones and tokens and modes confuse more than they help. The main shift he records is that the model race runs on its own while hiring, software and company processes barely change, and the advantage goes to the 0.1 percent of people who had the head to take the right decision. He turns the question of whether to switch on its head: what matters more is whether it will really give an effect.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 72 segments: 72 identified, 0 mixed, 0 probable, and 0 unresolved.

Read transcript on a separate page

Loading…