Fable 5.1 Is Out. Claude Opened Its Memory. OpenAI Is Preparing Astra — What Is Changing in the New AI Race?
Which of this week's updates actually changes how you work with AI — opened memory, a new architecture or in-house chips — and who ends up with the advantage in the race: Anthropic or OpenAI?
The reader will see why Fable 5.1 matters more to large companies than to a chat user: data storage moves inside the corporate perimeter and cache access gets 75% cheaper. They will see what Claude actually exposed in its "Topics" section and what memory still lacks. They will understand what recurrent transformers are and why the market pays for Astra's quality with a loss of reasoning transparency. And they get a frame for how platforms pull third-party functions in-house — from Amazon and YouTube to OpenAI.
What to watch for
Key takeaways
Anthropic frames the update around the architecture of large systems across several repositories, big migrations, bulk code updates and performance checks. In essence it is the same as Mythos 5.1. An ordinary chat user will probably notice no difference.
Requests to models of this class were stored for thirty days and analysed for jailbreak attempts. For a large company, data outside its perimeter means liability to clients and leak risk — and that is what held Anthropic back exactly where it aims.
With 5.1, data storage and the classifiers that decide whether to block a request can be deployed on the company's own infrastructure. The blocker is removed, and it is a move straight into the enterprise.
Repeat computation is taken from the cache, and access to it becomes about 75% cheaper. Across a whole task that is roughly 25%. For Anthropic's expensive models it is one more step towards market reach.
A "Topics" section appeared in the interface: every stored entry is visible and can be edited or deleted. It applies to Claude and Claude Cowork, but not to Claude Code.
On a long trip the model keeps the context without reminders, and the details no longer have to be explained each time. What stays open is where that context will apply later and how the model works out which place a question is about.
The Grill me skill examines an idea from every side and asks leading questions; dozens of messages are distilled into a document setting out what is being done and why. That document travels into another system and the work continues there.
A person cannot manage memory indefinitely — so the system has to: learning per project and asking clarifying questions. For now the models are solving other problems and memory stays on the user.
The expensive model in Extra High mode produced three versions of what was happening in Madrid, while a simple assistant guessed a demonstration. By evening it turned out there had been two rallies in the city. The expensive system searched and analysed, and lost the piece of information that mattered.
An ordinary model generates its reasoning token by token. In the new architecture several iterations of thinking happen inside the model before the first token appears, and what comes out is a far sparser chain.
Safety control today rests on analysing the reasoning chain: it is how an incident gets reconstructed. In hidden form that becomes far harder. OpenAI promises to limit the depth of internal thinking, but nobody guarantees the same from other companies.
Processors, high-speed memory, the network between processors, boards, server racks, software — and about nine months of development in which AI itself took an active part.
Ordinary hardware gains aggregate throughput by batching users into groups, and each of them starts waiting longer. Here the aim is to serve many requests at once and answer each of them fast.
Anthropic's opened memory will not work through someone else's interface, and Perplexity has no model of its own at the level of Fable 5.1 or Astra. The same argument explains why terminating the Cursor contract makes strategic sense.
Amazon let sellers in, watched what sold, and released its own batteries, paper and forks — they went to the top, and tens of thousands of sellers were squeezed out. The same mechanism applies to generating ads, thumbnails and short-video cuts.
The M6 and M5 Ultra have shipped, along with Mac Studio and Mac Mini, and nobody has come close to the MacBook on its own ground. But Apple has no models at ChatGPT's level — and were building them genuinely easy, Apple Intelligence would already have them.
What this episode is about
The ToTheMoon Sunday podcast with Alexander Volchek, Ilnar Shafigullin and Tatiana Tsvetkova. The occasion is three pieces of news that landed in a single week.
The first block is Fable 5.1. Anthropic describes the update as substantial: it holds the architecture of large systems across several repositories better, along with big migrations and bulk code updates. In essence it is the same as Mythos 5.1, and in an ordinary chat the difference is probably invisible — the hosts did not see it in a couple of days of testing. The real change is elsewhere: requests used to be stored for thirty days and analysed, which blocked adoption in exactly the large companies Anthropic aims at. Now storage and the classifiers can be deployed inside the customer's own perimeter. On top of that, cache access gets about 75% cheaper and a whole task about 25%.
The second block is memory. Claude opened a "Topics" section where every entry stored about you is visible and can be edited or deleted; it covers Claude and Claude Cowork but not Claude Code. The host admits that memory has grown most at ChatGPT: on a long trip the model keeps the context without reminders. Then comes what is missing: the system has to manage memory itself and finally start asking questions, while non-linear storage and the trimming of long chats remain unsolved. Ilnar describes a working detour — skills like Grill me and an ADR document you can carry into any other system. A separate story comes from Madrid, where the expensive model in Extra High mode produced three versions of what was happening while Alice answered "probably a demonstration" and turned out to be right.
The third block is architecture and hardware. The Astra rumours point to recurrent transformers: several iterations of thinking happen inside the model before a token appears, what comes out is a sparse chain, and following the reasoning step by step no longer works. Hence the opposition: unpicking incidents like the Hugging Face break-in depends precisely on reconstructing that reasoning. The payoff is quality growing without hyper-inflating model size. In parallel OpenAI showed the Jalapeño chip, which is in fact a whole platform: processors, memory, network, racks, software; in nine months it beats NVIDIA's forthcoming parts on OpenAI's own tests. The episode closes on how platforms take third-party functions in-house — through Amazon, YouTube, Meta, Cursor and Perplexity — and on who is building the infrastructure: NVIDIA, Tesla with SpaceX, Google with its TPUs, and Apple, which has strong hardware and no models of its own.
The episode makes a useful distinction between a model update and a market shift, and this week the two came apart unusually clearly. Fable 5.1 is nearly invisible in a chat yet removes the main corporate blocker; Claude's opened memory is more a claim on the interface than a solution; the real shift hides in the architecture and the hardware, where it is hardest to verify from outside. And the same episode names the price of progress plainly: recurrent transformers buy quality with the transparency of reasoning, and platforms taking functions in-house take them from the people who invented them.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 105 segments: 105 identified, 0 mixed, 0 probable, and 0 unresolved.
Loading…