Claude Opus 5.5 Is Out. Is It Even Worth Switching to New Models?
If model capabilities have grown so much in recent months, why have people and companies not become five or ten times more effective — and is it even worth switching to every new model?
Claude Opus 5.5 arrived with the claim that it beats Fable, costs less and works better in medium mode than earlier models on maximum tariffs, and on the same day Codex dropped 5.6 Sol for GPT-6 Astra. The host checks it against practice: his sister's scheduled task stopped running in Astra, in Claude Design he keeps Opus 4.8 Max, Astra in Pro mode spent hours on a simple report, and Claude on Fable 5.1 ate his subscription tokens in two hours. On the tests Opus 5.5 gives 66% in the terminal against 55 for Fable and 57 for Astra, and Astra wins only on science, 64 against 58. Alexander Volchek shows that hiring, software and company processes barely changed, he sees a real shift in one person out of twenty or thirty, and that 0.1 percent of the world got super-opportunities — and that the main question is not adapting to new models but whether they really pay off.
What to watch for
Key takeaways
In recent months Fable, GPT-6 Astra and now Claude Opus 5.5 have come out with new records and claims from Anthropic, but between how fast the models grow and how life, work and business change the host sees a crazy gap. Three years of endless updates — ChatGPT versions from 4.0 to 6.0, a whole Anthropic line — have moved the discussion from punctuation mistakes to hacks with Mythos-family models and to the robust writing of websites and apps. The episode's questions: where new models give a different level of result, which to choose and what to change in one's own work.
A task the host's sister had run on a schedule for six months stopped working in GPT-6 Pro — that is, in Astra — although almost the version before last had handled it. In Claude Design he still uses Opus 4.8 Max: with Fable and Opus 5 the system seemed not to understand what to do, and Astra does not produce the reports that version 5.4 did. On the day of recording Codex said it no longer supports 5.6 Sol on the Pro subscription and offered to migrate to GPT-6 Astra, although in 5.6 the host had stability; OpenAI released the Terra, Luna and Sol line and moves on without letting people test it.
The companies are interested in their own result — AGI, ASI, the technological singularity or systems that give them serious advantages, up to hacking states and economies. But do the models of the last half year help Anthropic or OpenAI gain users and earn more from them? The host sees no direct result, although the systems have improved incredibly. If the whole world is given the ability to open a business, there will be chaos of mass competition; AI's impact on countries' GDP last year was small, but the changes in how people behave and communicate are already irreversible.
Websites, reports and app design have become a simplified task, but people stay the same. In many meetings with companies the host honestly sees one case in twenty or thirty where approaches to work have changed and there is a positive effect from systems like ChatGPT. His wife, busy with the house and four children, has used ChatGPT on paid subscriptions from day one — but has her life really changed? He asks viewers the same: have they started earning more, do they have free time and thanks to what — for instance, eight hours of work done in one, with the employer kept in the dark.
Employers expect the saved hours to be filled with a huge amount of other work. If in a company of a hundred people finance, analytics, marketing and editing gained x5 or x10 efficiency with ChatGPT, has it effectively become five hundred people — or do people just generate more documents and text? The software the host uses has shown no miracle: Booking, Airbnb, Google Drive, Gmail and Google search do not give the right hints or new interfaces, although inside, systems capable of analyzing hundreds of millions of people are being built.
Systems like OpenAI have far more data about restaurants, cafés and hotels than any aggregator like Booking: when travelling, the host himself asks ChatGPT about the best seats on a plane more than he asks Expedia, Booking or the operators who sold the ticket. Revolutionary changes, in his view, will begin when the systems start surveying people, but that is not happening — the companies do not care about people: they measure terminal programming and business processes, not a real person's life. The smart home has not changed in several years, car assistants, even Tesla's, answer slowly, and there is little real value so far.
Opus 5.5 promises to spend 30–50% fewer tokens, but people do not understand what tokens are, and systems should be compared by the final result. The host replaced his own heart-sensor software on the Mac and MATLAB with Codex: start the sensor for the night, download the data, build a report, send it to the doctor, compare it with three years of cardiograms. Yet GPT-6 Astra in Pro mode spent hours on a report covering a few days, whereas ChatGPT-5.5 or 5.0 spent far less at the same efficiency, and Claude on Fable 5.1 ate all his subscription tokens in two hours.
Opus 5.5 scores sixty-six percent on terminal work against fifty-five for Fable and about fifty-seven for Astra, and fifty-four in programming against fifty for Fable; the only task where GPT-6 Astra wins significantly is science, fifty-eight against sixty-four. Headlines about a website, a 3D world and eight hundred presentations from one prompt do not impress the host: ChatGPT Pro in Extra High mode did that half a year ago. The main question: if Fable can hack everything in sight, what is Opus 5.5, which outpaces it — or was Fable made worse, as in July.
ChatGPT's chat tokens are not counted together with Codex and Work tokens, and Greg Brockman, who took on unifying everything into one app in late spring or June, has not yet removed the split between the chat and Codex; at Anthropic people likewise choose between Cowork, Code and plain Claude without understanding why the system goes into one mode or another. Opus 5.5 has modes from low to max with an intelligence index of 58 at max and 56 at extra high — Anthropic's own scale, not a percentage of correct answers and not an IQ. Opus 5.5 Low costs ten times less than Max with an index of 42 against 58.
Over the last half year Anthropic and OpenAI have gained a striking lead: Grok, Moonshot Kimi and DeepSeek were visible in the past, and now nobody is close. Google will stay in the race thanks to its search and services, but without them, in the host's view, Google Gemini would have huge problems. The new iPhone and Siri interest nobody: the host has the eighteenth iPhone with twice the compute for AI, but it makes no difference to him how artificial intelligence works there. No system yet works in different places at once, and Claude Code does not accept files over twenty megabytes and does not work with the file system.
Anthropic released Fable and blocked it after a US government decree: the model may not be used by non-citizens, including Anthropic's own employees — and it is unclear which model is used now. If Anthropic or OpenAI have no restrictions and infinite compute, why do they not take whole industries with one or two percent of their compute? An example exists: ChatGPT Sites creates and hosts websites — more than ten million, according to Brockman — and is taking the CMS market; in a year, the host believes, new sites will be written outside CMS systems, and marketing and sales will stop ordering development.
Literate people have gained a chance to earn a lot of money fast and to gain contacts, but it went to 0.1 percent of the world — those with the head to take the right decision. Eight out of ten, by the host's observation, resist, and AI remains very relative for them: returning from Europe to the States, he watched people spend the ten hours of the flight in awkward Excel files with accounting data, rechecking everything by hand and not opening ChatGPT or Anthropic. The episode's question is not whether these people will be able to adapt, but whether it will really give them an effect.
What this episode is about
A solo ToTheMoon episode about the release of Claude Opus 5.5 — and about what the release of yet another model means at all. Alexander Volchek starts with Anthropic's records and claims, recalls that Fable and GPT-6 Astra came out in recent months, and records a crazy gap between how fast the models grow and how life, work and business change. But in parallel the models seem to degrade: the host's sister's task stopped working in GPT-6 Pro, in Claude Design he still uses Opus 4.8 Max, and Astra does not produce the reports that version 5.4 did.
The host does not want to treat Opus 5.5 as just another release — what interests him is how models are released in general. On the day of recording Codex told him it no longer supports 5.6 Sol on the Pro subscription and offered to migrate to GPT-6 Astra, although in 5.6 he had stability; OpenAI released the Terra, Luna and Sol line and moves on without letting people test it. The reason, in his view, is that OpenAI and Anthropic are interested in their own result — AGI, ASI, the singularity — and it is a big question whether the models of the last half year help their business: no direct result is visible, although the systems have improved incredibly. If the whole world is given the ability to open a business, there will be chaos of mass competition; the host calls AI's impact on GDP last year small, although the changes in people's behaviour are already irreversible.
Next — people who stay the same. Websites, reports and design have become a simplified task, but at many meetings with companies the host sees one case in twenty or thirty where approaches to work have really changed. His wife, with four children, has used ChatGPT on paid subscriptions from day one — has her life really changed? He asks viewers the same: have you started earning more, do you have free time, or do you do eight hours of work in one and not tell your employer. Employers, for their part, expect you to fill the saved hours with a huge amount of other work, and if efficiency in a company of a hundred people grew fivefold, has it effectively become five hundred people — or just more documents and text.
Software has shown no miracle: Booking, Airbnb, Google Drive, Gmail and Google search do not give the right hints or new interfaces, although inside, systems capable of analyzing a state and hundreds of millions of people are being built. The host repeats that aggregators will disappear: OpenAI has more information about restaurants and hotels than any aggregator, and he himself asks ChatGPT about seats on a plane more than he asks Expedia or Booking. The revolution will begin when systems start surveying people — but that is not happening, because the companies do not care about people: the metrics measure programming, not a real person's life, the smart home and car assistants, even Tesla's, work crookedly, and there is little real value so far.
Astra's positioning — a system that communicates less with the human — the host considers new, but he recalls that the previous one did not ask questions either; his second guess is that the companies fear super-access, which would allow surveying a billion, a hundred million or ten million users. Hence — tokens: Opus 5.5 promises to spend 30–50% fewer of them, but people do not understand what tokens are, and what should be compared is the final result. His case is a heart sensor whose data Codex now collects and analyzes for him. Meanwhile GPT-6 Astra in Pro mode spent hours on a report covering a few days, whereas ChatGPT-5.5 or 5.0 spent less at the same efficiency, and Claude on Fable 5.1 ate the subscription tokens in two hours.
The host gives the Opus 5.5 tests with a caveat: 66% in the terminal against 55 for Fable and 57 for Astra, 54% in programming against 50 for Fable, while Astra wins only on science tasks — 58 against 64. Headlines about a website, a 3D world and eight hundred presentations from one prompt do not impress him: ChatGPT Pro in Extra High mode did that half a year ago. His main question: if Fable can hack everything in sight, then what is Opus 5.5, which outpaces it — or was Fable made worse, as in July. A poll on the channel showed that more than thirty percent of viewers had not even used the new Astra.
A separate knot is modes and products. At ChatGPT the chat tokens are not counted together with Codex and Work tokens; Greg Brockman has been working since late spring or June on unifying everything into one app, but the chat and Codex are still separate, and at Anthropic people choose between Cowork, Code and plain Claude. Opus 5.5 has modes from low to max with an intelligence index of 58 at max and 56 at extra high — Anthropic's own scale, in the host's words — and Low costs ten times less than Max with an index of 42 against 58. No competitors are visible: Anthropic and OpenAI have pulled away, Grok, Moonshot Kimi and DeepSeek are in the past, Google holds on thanks to search, and Siri in the new iPhone 18 interests nobody.
In the finale — limits and advantage. The systems have become slippery on questions of hacking, although they could warn of a risk zone; after a US government decree Anthropic blocked Fable for non-citizens, including its own employees, and it is unclear which model is now used under that name. And if the companies have no restrictions, why do they not take whole markets with one or two percent of their compute — an example exists: ChatGPT Sites hosts, according to Brockman, more than ten million websites and is taking the CMS market. The super-opportunities went to 0.1 percent of the world, eight out of ten resist, and people on a plane spend ten hours in awkward Excel files without opening ChatGPT. The episode's question is not whether they will adapt, but whether it will really give them an effect.
The episode's value is that it does not repeat the press release: the host checks Opus 5.5's records against his own tasks, where new models behave worse than old ones and tokens and modes confuse more than they help. The main shift he records is that the model race runs on its own while hiring, software and company processes barely change, and the advantage goes to the 0.1 percent of people who had the head to take the right decision. He turns the question of whether to switch on its head: what matters more is whether it will really give an effect.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 72 segments: 72 identified, 0 mixed, 0 probable, and 0 unresolved.
Read transcript on a separate page
Loading…