Skip to content
OpenAI · Anthropic · ClaudeEpisode 161 · 6 September 2026 · 52:48

Fable 5.1 Is Out. Claude Opened Its Memory. OpenAI Is Preparing Astra — What Is Changing in the New AI Race?

Central question

Which of this week's updates actually changes how you work with AI — opened memory, a new architecture or in-house chips — and who ends up with the advantage in the race: Anthropic or OpenAI?

What you take away

The reader will see why Fable 5.1 matters more to large companies than to a chat user: data storage moves inside the corporate perimeter and cache access gets 75% cheaper. They will see what Claude actually exposed in its "Topics" section and what memory still lacks. They will understand what recurrent transformers are and why the market pays for Astra's quality with a loss of reasoning transparency. And they get a frame for how platforms pull third-party functions in-house — from Amazon and YouTube to OpenAI.

Main threads

What to watch for

1Open Claude's "Topics" section and look at what is stored about you: entries can be edited and deleted.
2If adoption was blocked by data being stored outside the company, re-read the 5.1 terms: storage and classifiers now deploy inside your own perimeter.
3Recost your scenarios with the cache in mind: access to it got about 75% cheaper, and a whole task about 25%.
4Instead of long templates, build a skill and distil the conversation into a single ADR document: it travels into any other system.
5Before building a business on reselling someone else's model, check what stays yours once the platform takes that function in-house.
Signals to track afterwards
Whether Astra ships and the move to recurrent transformers is confirmed — and how model reasoning will be analysed then.
Whether other companies limit the depth of internal thinking, or that stays a promise from OpenAI alone.
Whether OpenAI's own platform becomes an external product or stays inside its own infrastructure.
Whether models get memory that asks questions and separates contexts by itself, rather than waiting for prepared files.
What is left of Perplexity and services like it once the key functions move inside the platforms for good.
Most useful for
Technical leaders whose adoption was blocked by data being stored outside the company perimeter.Anyone choosing between models on price: cache and task cost are named here in numbers.Anyone who cares about an assistant's memory — what it keeps about you and how to manage it.Anyone following model safety: recurrent transformers change the very possibility of control.Founders of services built around large platforms — how to tell a niche from a function that will be taken in-house.

Key takeaways

01:18Fable 5.1 Is Built for Large Systems, Not for Chat

Anthropic frames the update around the architecture of large systems across several repositories, big migrations, bulk code updates and performance checks. In essence it is the same as Mythos 5.1. An ordinary chat user will probably notice no difference.

03:58Thirty Days of Storage Was the Main Enterprise Blocker

Requests to models of this class were stored for thirty days and analysed for jailbreak attempts. For a large company, data outside its perimeter means liability to clients and leak risk — and that is what held Anthropic back exactly where it aims.

05:21Storage and Classifiers Move Inside the Company

With 5.1, data storage and the classifiers that decide whether to block a request can be deployed on the company's own infrastructure. The blocker is removed, and it is a move straight into the enterprise.

06:12Cache Gets 75% Cheaper, a Task About a Quarter

Repeat computation is taken from the cache, and access to it becomes about 75% cheaper. Across a whole task that is roughly 25%. For Anthropic's expensive models it is one more step towards market reach.

08:57Claude Showed Exactly What It Remembered About You

A "Topics" section appeared in the interface: every stored entry is visible and can be edited or deleted. It applies to Claude and Claude Cowork, but not to Claude Code.

10:08ChatGPT's Memory Grew Enough That a New Chat Stopped Being a Problem

On a long trip the model keeps the context without reminders, and the details no longer have to be explained each time. What stays open is where that context will apply later and how the model works out which place a question is about.

17:30A Skill and an ADR Document Work as Portable Context

The Grill me skill examines an idea from every side and asks leading questions; dozens of messages are distilled into a document setting out what is being done and why. That document travels into another system and the work continues there.

19:17The Systems That Start Asking Will Win

A person cannot manage memory indefinitely — so the system has to: learning per project and asking clarifying questions. For now the models are solving other problems and memory stays on the user.

22:23Alice Answered Probabilistically and Was Right

The expensive model in Extra High mode produced three versions of what was happening in Madrid, while a simple assistant guessed a demonstration. By evening it turned out there had been two rallies in the city. The expensive system searched and analysed, and lost the piece of information that mattered.

25:01Astra Is Built on Recurrent Transformers

An ordinary model generates its reasoning token by token. In the new architecture several iterations of thinking happen inside the model before the first token appears, and what comes out is a far sparser chain.

27:01The Price of the New Quality Is Lost Reasoning Transparency

Safety control today rests on analysing the reasoning chain: it is how an incident gets reconstructed. In hidden form that becomes far harder. OpenAI promises to limit the depth of internal thinking, but nobody guarantees the same from other companies.

30:31Jalapeño Is Not a Chip but a Whole Platform

Processors, high-speed memory, the network between processors, boards, server racks, software — and about nine months of development in which AI itself took an active part.

31:28Two Problems at Once: Throughput and Latency

Ordinary hardware gains aggregate throughput by batching users into groups, and each of them starts waiting longer. Here the aim is to serve many requests at once and answer each of them fast.

38:25Reselling Other People's Models Is a Dead Business Model

Anthropic's opened memory will not work through someone else's interface, and Perplexity has no model of its own at the level of Fable 5.1 or Astra. The same argument explains why terminating the Cursor contract makes strategic sense.

42:11A Platform Takes a Function In-House as Soon as It Sells

Amazon let sellers in, watched what sold, and released its own batteries, paper and forks — they went to the top, and tens of thousands of sellers were squeezed out. The same mechanism applies to generating ads, thumbnails and short-video cuts.

51:56Apple Has Strong Hardware and No Models of Its Own

The M6 and M5 Ultra have shipped, along with Mac Studio and Mac Mini, and nobody has come close to the MacBook on its own ground. But Apple has no models at ChatGPT's level — and were building them genuinely easy, Apple Intelligence would already have them.

What this episode is about

The ToTheMoon Sunday podcast with Alexander Volchek, Ilnar Shafigullin and Tatiana Tsvetkova. The occasion is three pieces of news that landed in a single week.

The first block is Fable 5.1. Anthropic describes the update as substantial: it holds the architecture of large systems across several repositories better, along with big migrations and bulk code updates. In essence it is the same as Mythos 5.1, and in an ordinary chat the difference is probably invisible — the hosts did not see it in a couple of days of testing. The real change is elsewhere: requests used to be stored for thirty days and analysed, which blocked adoption in exactly the large companies Anthropic aims at. Now storage and the classifiers can be deployed inside the customer's own perimeter. On top of that, cache access gets about 75% cheaper and a whole task about 25%.

The second block is memory. Claude opened a "Topics" section where every entry stored about you is visible and can be edited or deleted; it covers Claude and Claude Cowork but not Claude Code. The host admits that memory has grown most at ChatGPT: on a long trip the model keeps the context without reminders. Then comes what is missing: the system has to manage memory itself and finally start asking questions, while non-linear storage and the trimming of long chats remain unsolved. Ilnar describes a working detour — skills like Grill me and an ADR document you can carry into any other system. A separate story comes from Madrid, where the expensive model in Extra High mode produced three versions of what was happening while Alice answered "probably a demonstration" and turned out to be right.

The third block is architecture and hardware. The Astra rumours point to recurrent transformers: several iterations of thinking happen inside the model before a token appears, what comes out is a sparse chain, and following the reasoning step by step no longer works. Hence the opposition: unpicking incidents like the Hugging Face break-in depends precisely on reconstructing that reasoning. The payoff is quality growing without hyper-inflating model size. In parallel OpenAI showed the Jalapeño chip, which is in fact a whole platform: processors, memory, network, racks, software; in nine months it beats NVIDIA's forthcoming parts on OpenAI's own tests. The episode closes on how platforms take third-party functions in-house — through Amazon, YouTube, Meta, Cursor and Perplexity — and on who is building the infrastructure: NVIDIA, Tesla with SpaceX, Google with its TPUs, and Apple, which has strong hardware and no models of its own.

The episode makes a useful distinction between a model update and a market shift, and this week the two came apart unusually clearly. Fable 5.1 is nearly invisible in a chat yet removes the main corporate blocker; Claude's opened memory is more a claim on the interface than a solution; the real shift hides in the architecture and the hardware, where it is hardest to verify from outside. And the same episode names the price of progress plainly: recurrent transformers buy quality with the transparency of reasoning, and platforms taking functions in-house take them from the people who invented them.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 105 segments: 105 identified, 0 mixed, 0 probable, and 0 unresolved.

Loading…