Skip to content
Anthropic · ISI security · Artificial intelligenceEpisode extra15 · 4 March 2026 · 01:39:12

AI Safety Is Not a Fight Against an “Evil Model,” but a Fight Over the Right of People and States to Set Its Rules

Central question

Who should set the rules for AI: the company, the state, the user, or the system of principles inside the model itself?

What you take away

Build a working map of accountability for the case involving AI safety and Anthropic. A practical assessment requires the reader to separate a technical restriction, enforceability, and the responsibility of the company, platform, and user.

Main threads

What to watch for

1Compare “AI Psychopathy: as a model learns to circumvent restrictions” with “Risks of AI in critical areas”: they provide different criteria for judging the same issue.
2Test the conclusion from “Ethical restrictions and censorship in AI models” in your own use case—what actually changes in the process and what remains a promise.
3Before choosing a product or approach, record the constraint identified in “As a limitation to Asian AI-models”.
4Define the owner of the outcome and the quality metric for the situation described in “What is security AI and where real risks are part 1/2”.
Signals to track afterwards
Watch for actions by Anthropic and Google that confirm or challenge the episode’s central claims.
Compare new launches and policy changes with “Ethical restrictions and censorship in AI models”: have access, quality, price, or constraints changed?
Check whether the scenario in “What is security AI and where real risks are part 1/2” becomes repeatable practice rather than a one-off demonstration.
Most useful for
Executives and managersEntrepreneursDevelopersProduct teamsResearchersAI users

Key takeaways

00:00The practical meaning of the issue: ToTheMoon tonight

The decision in “ToTheMoon tonight” depends on one criterion: the risk depends on the scope of access, the scale of the consequences, and whether the system can be stopped and its actions reconstructed.

04:04Where the promise meets reality: how the OpenAI and Anthropic have emerged, and

In the context of “How the OpenAI and Anthropic have emerged, and with the safety of AI, part 1/2,” this criterion applies: the important signal is not one funding number: the next round, available runway, and closure rate show whether a company can survive the new cost of capital.

08:05What determines the outcome: what is the complexity of AI-model security

The discussion of “What is the complexity of AI-model security” yields a practical test: the important signal is not one funding number: the next round, available runway, and closure rate show whether a company can survive the new cost of capital.

14:25Why an announcement is not enough: what a safe robot should be - the

The decision in “What a safe robot should be - the look of Azimova” depends on one criterion: the important signal is not one funding number: the next round, available runway, and closure rate show whether a company can survive the new cost of capital.

16:15The market tests it through use: what is security AI and where real risks

The working conclusion from “What is security AI and where real risks are part 1/2” is that the important signal is not one funding number: the next round, available runway, and closure rate show whether a company can survive the new cost of capital.

22:55The boundary between value and constraint: how AI-models are taught

The working conclusion from “How AI-models are taught” is that a benchmark measures a narrow capability; working value requires repeatability, a clear cost, and control over errors.

48:27Anthropic built its approach around constitutional AI: the model receives a set of principles and learns to compare answers against them

The “AI Psychopathy: as a model learns to circumvent restrictions” issue should be assessed with one constraint in mind: this is better than hidden, arbitrary filters, but the constitution itself is still written by people. It reflects the culture, risk profile, and interests of the organization.

54:13Asimov’s laws illustrate the robotics problem well: a simple formula meets an ambiguous reality

The decision in “Risks of AI in critical areas” depends on one criterion: this section clarifies the mechanism behind the topic and preserves a constraint that would otherwise be lost in an overly simple conclusion.

01:20:00Military, chemical, and cyber scenarios require strict restrictions because a model lowers the threshold for complex knowledge

The working conclusion from “Ethical restrictions and censorship in AI models” is that this section clarifies the mechanism behind the topic and preserves a constraint that would otherwise be lost in an overly simple conclusion.

01:34:54A safe system should separate levels of risk, show the reason for a refusal, and retain a log of critical actions

For the “As a limitation to Asian AI-models” scene, the decisive point is this: this section clarifies the mechanism behind the topic and preserves a constraint that would otherwise be lost in an overly simple conclusion.

What this episode is about

The US government’s conflict with Anthropic, constitutional AI, agents, military applications, and restrictions in Asian models show that there is no single definition of safety. Every prohibition reflects the values of whoever wrote it, while critical systems require external oversight and transparent boundaries.

A model can refuse to carry out a government order because its rules classify the order as dangerous. At that moment, safety stops being a technical parameter and becomes politics. Who decided what is allowed: the company, the developer, the state, or the user paying for the system?

Anthropic built its approach around constitutional AI: the model receives a set of principles and learns to compare answers against them. This is better than hidden, arbitrary filters, but the constitution itself is still written by people. It reflects the culture, risk profile, and interests of the organization.

Asimov’s laws illustrate the robotics problem well: a simple formula meets an ambiguous reality. An agent may not understand that an action will cause harm, or may choose one harm in order to avoid another. The closer AI comes to the physical world and critical infrastructure, the less adequate a general prohibition becomes.

Military, chemical, and cyber scenarios require strict restrictions because a model lowers the threshold for complex knowledge. Complete closure, however, also concentrates power in a few companies and states. Asian models have their political boundaries; Western models have theirs. Neutral intelligence does not exist.

A safe system should separate levels of risk, show the reason for a refusal, and retain a log of critical actions. That may not be enough for a state, which will demand disclosure of protocols, tests, and incidents. The main question is not whether AI can become a weapon.

It can already amplify dangerous actions. The question is who controls access and can stop the system before a principle becomes an irreversible decision.

AI safety is not defined by the image of an “evil model.” It depends on who sets the principles, controls access, and can stop the system before an irreversible action.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 162 segments: 83 identified, 11 mixed, 37 marked with ✓, and 31 unresolved.

Loading…