OpenAI Showed Astra and a Scandal Began. A Breakthrough or Other People's Ideas?
Did Astra really solve the million-dollar problem on its own — and what matters more to an ordinary person: a proof that has not yet been recognised, or the way AI is already changing purchase decisions and why it still does not ask questions?
OpenAI reported that GPT-6 Astra, in 88 hours of work by ten thousand agents and 130 billion tokens, proved the existence and smoothness of a solution of the Navier–Stokes equations; mathematicians who had discussed an intermediate result in ChatGPT and Claude suspected their ideas had been used, and OpenAI replied that no dialogues after 3 July went into training and that it does not claim priority on the intermediate part. At the time of recording the proof has not been recognised. Ten days on, GPT-6 as a chat showed Alexander no difference and made his personal task worse, while Ilnar switched to it completely from 5.6: the benchmarks are torn apart and nothing got worse. Shopify showed a threefold rise in visits and orders through AI against 30% growth of search in two years, and a 15% higher average order. The hosts show that AI is already replacing restaurant and hotel aggregators but does not ask the user about the outcome, and describe the next stage — a proactive system for one or two thousand dollars a month. They consider the iPhone 18 an upgrade without a reason, and the foldable Duo Apple's entry into someone else's market.
What to watch for
Key takeaways
A system with ten thousand agents, in 88 hours, according to OpenAI, proved the existence and smoothness of a solution of the Navier–Stokes equations. Mathematicians working on the same problem asked whether the model had used their ideas.
Since 2000 only one has been solved — by Perelman, who refused the prize. The Navier–Stokes equations are more than ninety years old, many have tried to solve them, and the Clay Institute has not paid anyone yet.
The mathematicians discussed an intermediate result in ChatGPT and Claude in mid-August. OpenAI says the prompts could not have influenced the model and does not claim priority on the intermediate part; at the time of recording the proof has not been recognised.
Alexander: a mathematician would not have invested such money, and the system absorbed texts no single person could have read. Ilnar: without someone else's direction it would have needed ten times more tokens, and the solution might not have been found.
A result in science is a small "bump" on the body of what is known. Neural networks carry ideas from one field to another, which is harder for a human expert in a narrow area; there are no statements yet from authoritative mathematicians about checking the proof.
There is no super effect of the GPT-5.5 kind; a task that has resisted solution for a month got worse after Astra and cost third-party systems' resources. The host carries on with Fable 5.1, which has not coped yet either.
About 99% on ARC-AGI, no point testing further. Some tasks GPT-6 does not do straight away and needs adjusting to, but others got better; the ten thousand agents are on OpenAI's conscience — nobody in Codex has dreamed of that.
Those who come through AI buy better: higher conversion and a 15% higher average order. Alexander: the multiples have nothing to be compared with without absolute figures, but engagement through AI really is higher — the choice narrows to two options.
Ilnar trusts ChatGPT's recommendation as an expert's, Alexander trusts AI more than search with a caveat about versions; his friend does not trust AI in analytics and sales. Some tasks AI does badly, some better than a person, some no worse.
Booking's rating still works, but fifteen hundred reviews have no "star" for the mattress, and a good location in a big city is relative. Ilnar: chats still need a source of up-to-date reviews.
In two weeks and a hundred and fifty restaurants ChatGPT never once asked where the host went. Answer raters are bought for a few dollars, while a live user is expensive to "buy", although he is ready to answer himself.
Ilnar: a proactive assistant on a server for one or two thousand dollars a month. Alexander is ready to pay two thousand, but for OpenAI it costs almost two thousand a day, and there is a risk of fine-tuning without control.
The new AI block is twice as big, but it is unclear what it speeds up. The foldable Duo at two thousand dollars does not interest Alexander; Ilnar wants to check the crease on the screen but is unlikely to buy; pre-orders open at the end of October.
What this episode is about
A Sunday ToTheMoon episode with Alexander Volchek and Ilnar Shafigullin. The main news is OpenAI's announcement that GPT-6 Astra solved the Navier–Stokes equations, one of the Clay Mathematics Institute's Millennium Prize Problems with a one-million-dollar prize. Alexander came to the news through a harsh tweet aimed at Sam Altman and immediately asks questions: what are these ten thousand agents, who managed them, and is this not marketing. Ilnar sees a familiar pattern: after every strong model OpenAI reports that something even more powerful is in training.
Ilnar explains the context: in 2000 the Clay Institute chose seven problems and set a million for each; only one has been solved — by Perelman. The Navier–Stokes equations are more than ninety years old. According to OpenAI, its interest arose at the end of August, when a mathematician and an Anthropic employee obtained an intermediate result; after 88 hours of work by ten thousand agents the company announced it had proved the existence and smoothness of a solution.
The scandal: the mathematicians had discussed their ideas in ChatGPT and Claude and suspected they had been used in training. OpenAI answers that no dialogues after 3 July were used, the discussions took place in mid-August, and the company does not claim priority on the intermediate part. At the time of recording the proof has not been recognised. Alexander considers an insert about the price necessary: 130 billion tokens, 2.7 million messages — more than a million dollars at external prices; Ilnar recalls that a scientist stands on the shoulders of giants, and neural networks are especially strong at the intersection of fields.
Impressions of GPT-6 ten days on diverge. Alexander saw no super effect of the GPT-5.5 kind: the model did not solve his personal task, made it worse and spent third-party systems' resources, so he carried on with Fable 5.1. Ilnar switched completely from 5.6 to 6: the benchmarks are torn apart, ARC-AGI is closed at 99%, some tasks need adjustment, but nothing got worse. Both suggest leaving the ten thousand simultaneous agents on OpenAI's conscience — in Anthropic's study eighty agents got in each other's way.
In passing — voice: Alexander's companion on a trip did not realise he was talking to ChatGPT in CarPlay and wanted to pay for such a version. ChatGPT Voice has improved in two months, but the hosts do not use it: Alexander dictates tasks, Ilnar types at the computer.
Shopify showed that visits and orders through AI tripled in a year, search grew 30% in two years, and those who come through AI buy with higher conversion and a 15% larger order. Alexander takes apart what the statistics lack — absolute figures — and agrees with the main point: engagement through AI is higher, the choice narrows to two options. Ilnar adds trust in the model as an expert; society is splitting between those who trust AI and those who do not.
The purchase decision has already moved into the chat: Alexander chooses restaurants on trips through ChatGPT, aggregators no longer exist for him, and Booking with its rating loses to a model that re-reads Expedia, Booking and TripAdvisor together. But the model does not ask questions: over a hundred and fifty restaurants it never once asked about the outcome, although a live user is the most expensive source of training. The next stage, according to Ilnar, is a proactive system on a server for one or two thousand dollars a month; Alexander is ready to pay, but for OpenAI it costs almost two thousand a day, and there is a risk of fine-tuning without control.
The episode ends with the iPhone 18 and the foldable Duo. Alexander asked ChatGPT to compare only what matters: the processor makes no difference, the focal length is a plus for those who know how to use it, and the doubled AI block speeds up who knows what. The Duo does not interest him — Apple, in his view, has gone to someone else's market; Ilnar recalls the Vision Pro, wants to check the crease on the screen but is unlikely to buy. Pre-orders open at the end of October.
The episode is useful because it separates two things that merge in the news: a mathematical result that still has to be checked, and the marketing around it. The hosts deny neither the model's strength nor the possibility that it built on other people's ideas — they show how each scenario works and suggest waiting for the proof to be recognised. The second conclusion is more practical: the purchase decision is already moving into the chat, and what to watch is not the search engine's results but what ChatGPT and Claude say about you — while the models themselves still do not ask the person, and that is the main unsolved question of the interface.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 143 segments: 143 identified, 0 mixed, 0 probable, and 0 unresolved.
Loading…