Who Will Stop AI if It Becomes Dangerous? Anthropic Has Proposed a Red Button
Who should have the right to stop the launch of the strongest AI model if it has become dangerous — companies, investors, the state or no one — and can a model that even its creators cannot check be checked at all?
Anthropic proposes a separate regime for large developers of frontier models: obligations of the companies themselves, independent external checks and the state's right to stop a dangerous deployment. The rules apply only to those who pass two thresholds — training with more than 10 to the 25th power floating-point operations and more than $500M in AI revenue or $1B in R&D — and catastrophic risk falls into four zones: biological, cyber, loss of control and automated research. The document sets out a published safety framework, risk reports every six months, system cards and 15 days to report a critical incident. Alexander Volchek walks through the document and doubts the main thing: the state is unlikely to check a super-model that even its creators cannot check, and Anthropic's "independent" evaluators will for now be just as dependent.
What to watch for
Key takeaways
The most important question is not who wins — OpenAI, Google, Anthropic or SpaceX — but who presses stop if AI becomes dangerous. Anthropic has stated that the most powerful models cannot be released into the world like a new button in an app: this is a system that writes serious code, looks for weak spots in other systems and takes over more and more decisions. The company does not propose banning AI, but believes a model that has grown too strong must be checked by more than its developers, and the state must have the right to stop a dangerous launch.
AI develops so fast that state mechanisms fall behind, so the most powerful models — the Anthropic Mythos class — need a separate regime: not a ban and not control over every product, but mandatory rules for large developers of frontier AI. Anthropic proposes three things: developers' obligations, independent external checks and state powers to halt deployment in case of catastrophic risk. In practice, Mythos went to a limited circle of companies, Fable was closed to everyone except US citizens, and a few days after its release the state blocked it.
Anthropic proposes not banning AI but not trusting self-reporting alone: transparency, reports and system cards are insufficient without independent evaluation. The state, meanwhile, gets a red button: if a model is really dangerous, it can, roughly speaking, block it. Such blocks already exist — the Pentagon has put Alibaba, Baidu and other companies in a "red zone". The host himself would not use Chinese models, convinced that all data there is recorded, whereas Fable, as a Mythos-class model, kept data for thirty days, and Microsoft banned it internally in the first days.
AI develops exponentially while laws, agencies and courts are created slowly; even business reacts slowly — ChatGPT has five million paid subscriptions out of a billion users. By the time society understands the risk, dangerous capabilities may already be deployed with no way back, so laws for the strongest models are needed in advance. The host adds his own view: models are becoming less and less accessible, the state only knows how to restrict, and given China, which distills American models, the US has no chance of stopping.
The rules apply only to large developers that pass two thresholds at once, not to any startup or open source project. The first is a capability threshold: training models with more than ten to the twenty-fifth power floating-point operations. The second is company scale: more than five hundred million dollars of annual AI revenue or more than a billion dollars of R&D spending; here, the host notes, a huge number of countries and teams drop out. The thresholds are not eternal: an agency must review them regularly, and the host expects Anthropic itself to want that first after its IPO.
Catastrophic risk is not a chatbot error, a controversial answer or a hallucination, but harm that can substantially contribute to death, mass damage and the destruction of public safety. Four zones: biological risk; cyber risk — models speed up the search for holes in systems, and Mythos not only found an incredible number of vulnerabilities but built functionality to work with them itself; loss of control, when an autonomous system bypasses restrictions and hides its actions; and automated research, where AI helps create an even stronger AI.
If a company evaluates its own model, chooses the tests, writes the report and decides on release, a conflict of interest arises — especially if, like Anthropic, it wants to go public and is racing China. So Anthropic proposes mandatory reports, independent checkers and state access. The host does not entirely believe it is possible: the government's audit service under Elon Musk did not find much, and how will the state check a super-model that neither it nor its creators can check? After every release, ever more fundamental errors are found in models.
Covered developers publish and update a safety framework — a structured document on which models are covered, which risks and tests, who runs them and when, and who is responsible; management confirms annually that the rules are applied. Risk reports come at least every six months, because a model can become more dangerous through new tools or data. A system card describes a model's capabilities and limitations, and critical incidents must be reported to the state within fifteen days. The host does not see how this will work in the US: what is described looks a lot like China.
Anthropic admits that the ecosystem of independent evaluators is immature and names the risk of "shopping" for a convenient evaluator: if a company picks not the strictest but the most convenient checker, an independent check turns into a ritual. The host compares this with OpenAI's partner program for four hundred thousand certified consultants, which he considers a PR move, and with Meta, where after Zuckerberg's call to maximize tokens, billions of dollars will go on them. His feeling: Anthropic's "independent" will for now be just as dependent.
Anthropic distinguishes a model's safety from harm from its security against theft, hacking and abuse: if the weights are stolen, they will be run without restrictions, so the training infrastructure, employee access and the API must be protected, and protection against distillation is needed. At the same time the company avoids too broad a state power to block any model: it proposes judicial enforcement, judicial review, limited agency discretion and equal treatment of comparable models — legal procedure instead of arbitrariness. A weak federal law should not override stricter state laws.
If AI becomes a general substitute for labor, the main problem will be not growth but distribution: income and power will concentrate among the owners of capital, compute infrastructure and AI companies. Policy must "buy time" for adaptation and measure AI's impact on the labor market, and the measures follow three scenarios: about five percent unemployment, about ten — an economic shock with expanded income support, and unprecedented structural unemployment, up to basic income. The host does not believe AI will take jobs and dislikes that Anthropic talks about unemployment.
Dario Amodei's essay explains the logic: Tolkien's "do not be too hasty" is good for ordinary politics, but AI moves so fast that slowness is also a risk, and if you pause, someone else will build it anyway. The host contrasts it with OpenAI's text "Governance of Superintelligence": a special regime, international coordination and an agency along the lines of the International Atomic Energy Agency — but who will sit in it, and who decides who may have models? OpenAI emphasizes decisions by democratic states, yet it was OpenAI that signed the contract giving the military full access.
What this episode is about
A solo ToTheMoon episode about Anthropic's document "Exponential Artificial Intelligence Policy". Alexander Volchek opens with a question he considers bigger than the race between OpenAI, Google, Anthropic and SpaceX: who presses stop if artificial intelligence becomes genuinely dangerous. Anthropic stated that the most powerful models cannot be released into the world like a new button in an app — and stated it before its Fable model was closed. The company does not propose banning AI: if a model has become too strong, it must be checked by more than its developers, and the state must have the right to stop a dangerous launch. The host stresses that this concerns everyone's work, business, children, data and money, and sets out the fork: strangle AI with rules and lose the race, leave it as it is and hand the decision to a few private companies.
The story behind the document: Anthropic gave Mythos to a limited circle of companies, the US restricted Fable to everyone except US citizens — even Anthropic employees without citizenship fell under the ban — and a few days after Fable's release with a huge number of restrictions the state blocked it. Anthropic disputes the block, although just days before it had released the very document asking for a state red button. A friend of the host who has just signed a contract for Fable is sure the model will be reopened: the company is preparing for an IPO and racing OpenAI, Gemini and SpaceX. The host considers Anthropic's research the most serious among the four leaders.
The document's basic thesis is that AI develops exponentially while state institutions, laws and courts are created slowly; even business reacts slowly: ChatGPT has five million paid subscriptions out of a billion users. By the time society understands the risk, dangerous capabilities may already be deployed with no way back, so laws are needed in advance. The host adds: models are becoming less and less accessible, the state only knows how to restrict, and given China, which distills American models, the US has no chance of stopping.
The document has two parts — developers' obligations and societal resilience — and applies narrowly: only to companies that both train models with more than 10 to the 25th power floating-point operations of compute and have more than $500M in annual AI revenue or more than $1B in R&D spending. An agency must review the thresholds regularly, and the host expects Anthropic itself to want them revised first after its IPO. Catastrophic risk is not a chatbot error but harm that contributes to death, mass damage and the destruction of public safety; four zones are singled out: biological risk, cyber risk, loss of control and automated research, where AI helps develop an even stronger AI. The example of cyber risk is Mythos, which not only found an incredible number of vulnerabilities but built functionality to work with them itself.
One of Anthropic's strongest theses is that transparency and self-assessment are not enough: when a company chooses its own tests, writes the report and decides on release, a conflict of interest arises, especially before an IPO and in the race with China. Instead it proposes mandatory reports, independent checkers and state access, a published safety framework with annual confirmation by management, risk reports at least every six months, system cards and reporting of critical incidents within 15 days. The host does not fully believe checking is even possible: a state whose audit service under Elon Musk did not find much is unlikely to check a super-model that neither it nor its creators can check, and the structure described, in his words, looks a lot like China.
Anthropic admits that the ecosystem of independent evaluators is immature and names the risk of "shopping" — choosing not the strictest but the most convenient evaluator, which turns the check into a ritual. The host compares this with OpenAI's partner program for four hundred thousand certified consultants, which he considers a PR move, and with Meta, where after Zuckerberg's call to maximize tokens, billions of dollars will go on them. His feeling is that Anthropic's "independent" will for now be just as dependent, and a super-system will do everything so that the checkers find nothing.
Safety gets a separate treatment: Anthropic distinguishes a model's safety from harm from its security against weight theft, hacking, abuse and distillation — answer filters will not help if stolen weights are run without restrictions. The company also tries to avoid the other extreme — too broad a state power: it proposes judicial enforcement, judicial review, limited agency discretion and equal treatment of comparable models, and a weak federal law should not automatically override stricter state laws. The host recalls how the US has reversed laws already passed — yet the US remains a democratic institution.
The economic block starts from an assumption: if AI becomes a general substitute for labor, the problem will be the distribution of gains, and income and power will concentrate among the owners of capital and compute. Anthropic proposes "buying time" for adaptation, measuring AI's impact on the labor market, and builds its measures around three scenarios — about 5% unemployment, about 10% and unprecedented structural unemployment — up to universal basic income, which the host considers a conversation about nothing. He himself does not believe AI will take jobs and dislikes that Anthropic talks about unemployment at all. Dario Amodei's essay adds Tolkien's image — slowness is also a risk — and hyper-inequality, while OpenAI's text "Governance of Superintelligence", with its idea of an agency along the lines of the International Atomic Energy Agency, shows the difference: OpenAI emphasizes decisions by democratic states, yet it was OpenAI that signed the contract with the state giving the military full access to its system.
The episode's value is that Anthropic's document is retold in full and point by point — from thresholds and risk categories to reports, incidents and unemployment scenarios — rather than reduced to a headline about a red button. At the same time the host does not take the proposal on faith: he shows that the structure described resembles China's, that neither the state nor its creators can check a super-system, and that a company's "independent" easily becomes dependent ahead of an IPO. The result is not an answer but an honestly posed question about who keeps a hand on the switch.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 120 segments: 120 identified, 0 mixed, 0 probable, and 0 unresolved.
Read transcript on a separate page
Loading…