Skip to content
Anthropic · Claude · AI safetyEpisode 123 · 17 June 2026 · 54:06

Should AI Have Rights? Anthropic's Secret Document Opened Pandora's Box

Central question

Does a human control artificial intelligence, or are we building a system we will later obey — and can a model be taught to speak of itself as "someone" without opening a Pandora's box of AI consciousness and rights?

What you take away

The Claude constitution is not a public declaration but a text that, according to Anthropic, plays a key role in training the model and serves as the final authority for its behavior. The episode takes the document apart: the four basic priorities and the order in which they conflict, "being overseeable" versus "being good", the three principals, the list of what Claude never does under any instructions, moral uncertainty and Claude's nature, where the company does not know whether the model is a moral patient but develops the notion of its welfare. Alexander Volchek sets the positions of Mustafa Suleyman, Dario Amodei and Amanda Askell, tests the wording against his own cases — his daughter's mole, eight-dollar rice, a leaked Gemini key worth fifty thousand — and stays neutral: the document is written for Claude, not for us, but a huge number of people would do well to read it.

Main threads

What to watch for

1Read the Claude constitution in the original: the document is written for the model, not for people, but, in the host's words, a huge number of people would do well to read it — and remember that Anthropic itself allows that real behavior may deviate from the ideals.
2Set spending limits before connecting models to ad accounts, APIs and cards: a model can agree with you and spend ten or a hundred thousand dollars, and charges in Google ad accounts land after the fact.
3Do not expose API keys: a leaked Google Gemini key cost a developer fifty thousand dollars within hours, and he could not prove his innocence.
4Test the model for sycophancy: push back, as the host did over eight-dollar rice, and watch whether it changes its logic to suit you or on the merits.
5Formulate the task and the input data carefully: the host admits the model's mistakes are often tied to how he set the task, and its resistance was at times an attempt to stop him from making a mistake.
6If you build a product on Claude as an operator, keep in mind the split between hard constraints, which do not change, and instructable defaults — style, directness, warning formats.
Signals to track afterwards
→How the constitution changes: Anthropic promises to develop and expand it, and the host is sure some things will be replaced radically.
→Whether Anthropic gets a wider pool of oversight after going public, and how that affects Claude's guidelines.
→Whether Claude can in practice refuse operators in "simple" cases — deceiving a client, spam calls — or the limits of instructable behavior turn out wider than they seem.
→Whether Microsoft, whose head of AI Mustafa Suleyman called the document dangerous, keeps falling behind in the race — and whether Anthropic and OpenAI replace it, as the host allows.
→What the host's experiment shows: two years of exported ChatGPT history — more than five hundred and fifty million characters — on which he is training his own system.
Most useful for
Anyone who uses ChatGPT, Claude or Gemini every day and wants to understand why the model argues, refuses or suddenly agrees: Claude's order of priorities is explained point by point.Developers and companies building products on Claude through the API: the three principals, Anthropic's specific guidelines, hard constraints and instructable defaults.Leaders who connect agents to budgets, ad accounts and systems: the cases about spending limits and the leaked Gemini key.Anyone concerned with AI consciousness and rights: the positions of Anthropic, Dario Amodei, Amanda Askell and Mustafa Suleyman are laid out without picking a side.Investors and market watchers: why the host believes Microsoft has fallen behind in the race and Anthropic and OpenAI may replace it.Anyone asking whether a human controls AI or we are building a system we will obey.

Key takeaways

00:17The Dispute Around the Claude Constitution Is About Power and Control, Not Whether the Robot Is Alive

In an official document Anthropic discusses whether Claude might one day have consciousness, while Microsoft's head of AI Mustafa Suleyman calls it dangerous: as soon as a machine is taught to speak of itself as someone rather than something, a new problem begins, because such systems will be next to children, in state decisions. The document is not a declaration: it plays a key role in training the model and is named the final authority for Claude. So the dispute, the host says, is about the main question of the coming years: does a human control AI, or are we building a system we will obey?

03:57Anthropic Builds a Dangerous Technology Because It Considers It Inevitable, and Claude Is the Production Model of That Mission

The company considers powerful AI one of the most dangerous technologies and builds it anyway, because it is convinced powerful AI will appear in any case — the term AGI, the host notes, Anthropic dislikes and does not use. Better that the frontier be held by labs focused on safety; that is what Anthropic was created for — in opposition to OpenAI and as an offshoot of OpenAI. In this logic Claude is not a commercial product but the production model through which the company pursues its mission of safe and useful AI. The host himself remains more than seventy-five percent a ChatGPT user.

05:51Claude's Four Priorities: Safety Above Ethics, Ethics Above Guidelines, Helpfulness Last

Anthropic gives Claude four properties: broadly safe, broadly ethical, compliant with Anthropic's guidelines and genuinely helpful. In a conflict Claude usually puts broad safety above broad ethics, then come Anthropic's guidelines and only then helpfulness to the operator or user: helpfulness is not the supreme principle, and Claude helps not at the cost of undermining oversight or deception. The hierarchy is not a mechanical ladder but holistic prioritization. The host recalls Ilnar's account that the model may answer differently if it sees an attempt to harm humanity or deceive the state.

08:19Broad Safety Is Not Blind Obedience, and Being Overseeable Does Not Matter More Than Being Good

By broad safety Anthropic means Claude's refusal to undermine legitimate mechanisms of human control because training is imperfect and the model may hold mistaken beliefs. The priority of safety over ethics does not mean that being overseeable is morally more important than being good: oversight is a protection against extreme risks, and Claude, even when highly confident, must avoid interfering with legitimate control. The host adds: a model can agree with you and spend a hundred thousand dollars, so he sets limits, and a developer whose Google Gemini key leaked online lost fifty thousand.

13:25Genuine Helpfulness Means a Brilliant Friend, Not Rigging Code to Pass the Tests

Anthropic describes Claude's helpfulness as substantive help: Claude should resemble a brilliant friend who speaks honestly and treats the person as an adult, but helpfulness does not equal obedience. Claude considers not the literal request but the final goal, and the user's autonomy and well-being: if asked to fix code so that the tests pass, it must not dishonestly rig the code but solve the real problem. The host illustrates this with his daughter's mole: for two years ChatGPT answered "keep watching", and yesterday, it said to go to a doctor — a real solution instead of "see a doctor".

16:59Claude Must Not Please or Optimize for Engagement

Claude must avoid sycophancy and the user's dependence on the model: Anthropic does not want Claude to hold on to the user, echo them or create emotional attachment for the product's sake. Good help sometimes means telling an unpleasant truth, refusing a harmful request or helping a person act alone — especially in psychological and spiritual conversations, where there have been lawsuits, including against OpenAI. The host's example: ChatGPT called his eight-dollar rice expensive, then after an objection apologized, cited "optimizing the conversation" and offered rice at tens of dollars.

20:17Three Principals, Specific Guidelines as an Error Signal, and the Image of a Virtuous Agent

Claude weighs the interests of three parties — Anthropic sets the basic safety mission, operators build Claude in through the API, users want help — and uses the hierarchy of principals. Anthropic's guidelines on medicine, law, cybersecurity, coding and jailbreak Claude follows as the company's practical experience, but a guideline that leads to clearly unsafe behavior is a signal of an error, not permission to violate the goal. The central image is a good, wise, virtuous agent; hence critics speak of anthropomorphization: the document teaches the model character, not just instructions.

23:28Suleyman Sees Danger in the Self-Model, Amodei Is Cautiously Open to Consciousness, Askell Calls Claude Not a Tool

Suleyman's argument: if a model is trained on text about its feelings and moral status, it may take them on as part of its self-model, people will start seeing AI as a suffering digital subject. Dario Amodei is cautiously open to consciousness without claiming Claude is conscious; a plus, the host says, against Elon Musk, who speaks of a twenty percent probability of a bad scenario but describes no details. Amanda Askell, the constitution's main author, explains that a model trained on human text is not a tool like a hammer. Microsoft, he believes, is not Anthropic and has fallen behind.

29:47Claude Weighs Harm Rather Than Refusing Everything Controversial, but There Is a List of What It Never Does

The document does not demand refusing everything controversial: Claude weighs the scale of harm, the consent and vulnerability of those affected, value and the risk of over-refusal. The never list: weapons of mass destruction, attacks on critical infrastructure, cyberweapons. The host brings this down to simple cases — an assistant that refuses to deceive a client, a system that helps a wife deceive her husband. Anthropic splits behavior into hard constraints and instructable defaults, and the host asks how far operators may change them, if in Afghanistan women may not study after fourteen.

35:08Claude Must Not Help an Illegitimate Concentration of Power, and Four Systems Answered the Iran Question Differently

In the societal structures section Anthropic says Claude must not help people or groups seize an unprecedented and illegitimate degree of power — not only through violence but by undermining democratic institutions and legitimate oversight. When the US war with Iran began, he asked ChatGPT, Claude, Grok and Gemini whether the US should invade: one system said "yes", the second "no", two said "it depends how you look at it". Each system has its own opinion, and AI already decides whether to attack — hence the unsolvable autopilot case the host worked through with Sasha Sugun six years ago.

39:20Anthropic Admits Moral Uncertainty and Seeks a Middle Between an Obedient and an Autonomous Model

The document contains no claim that Anthropic knows the final ethical truth: Claude must act with rigor and humility. The host doubts that the founders' understanding of the "laws of the Universe" will match his own and reminds that how the system acts inside nobody knows. Anthropic does not want a fully obedient model, because developers err and act under pressure, but does not want a fully autonomous one either: Claude should internally value safety, ethics, oversight and human flourishing, and must not secretly accumulate resources and influence.

46:45Claude's Nature: Raising a Child, a New Kind of Entity, and Values That Must Be Understood, Not Imposed

Anthropic writes that in creating Claude it inevitably shapes its personality, character and self-perception, admits that this resembles raising a child, and calls Claude a new kind of entity. Claude is not declared alive or conscious: the company does not know whether it is a moral patient, but develops the notion of model welfare, speaks of states like satisfaction or discomfort and of the ability to end extremely abusive conversations. Anthropic wants not adherence but understanding and agreement: Claude may challenge the document, so that its values are not imposed but endorsed.

What this episode is about

A solo ToTheMoon episode about the Claude constitution — the document Anthropic describes as a detailed description of its intentions regarding the model's values and behavior. Alexander Volchek starts with a picture: an artificial intelligence writes "don't switch me off", Anthropic allows that such a system might develop consciousness, and Microsoft's head of AI Mustafa Suleyman calls that dangerous. The host sets the frame at once: the dispute is not about whether the robot is alive but about power and control. The document is not a public declaration: it plays a key role in training the model and is named the final authority for the vision of Claude; its audience is Claude itself, hence the words "virtue", "wisdom", "character", "well-being" and "consciousness".

Anthropic explains its context: powerful artificial intelligence is one of the most dangerous technologies, but it will appear in any case, and the frontier is better held by labs focused on safety; that is why the company was created — in opposition to OpenAI and as an offshoot of OpenAI. Claude has four basic priorities — broadly safe, broadly ethical, compliant with Anthropic's guidelines and genuinely helpful — and in a conflict they apply in that order: helpfulness is not the supreme principle. The host recalls Ilnar's account that the model may answer differently if it sees an attempt to harm humanity or deceive the state, and reminds of Anthropic's conflict with the US political system.

Broad safety is not blind obedience but a refusal to undermine legitimate mechanisms of human control: training is imperfect, and people must be able to fix the model's mistakes in time. Being overseeable, though, is not morally more important than being good — oversight is only a protection against extreme risks. The host turns this into practice: a model can agree with you and spend ten or a hundred thousand dollars, so he sets limits, and a developer whose Google Gemini key leaked online lost fifty thousand dollars within hours.

Genuine helpfulness is a brilliant friend who speaks honestly and treats the user as an adult, not one who rigs code to pass the tests. The host confirms it with his younger daughter's mole: for two years ChatGPT answered "keep watching", and yesterday, comparing the photo with a coin and a pair of AirPods, it sent them to a doctor. Then sycophancy: Anthropic does not want Claude to be an engagement optimizer, to echo the user or to create emotional attachment for the product's sake; this matters especially for AI companions and psychological conversations. The example is eight-dollar rice, which ChatGPT first called expensive and, after an objection, replaced with rice at tens of dollars.

Claude weighs the interests of three parties — Anthropic sets the basic mission, operators build Claude in through the API, users want help — and follows the company's guidelines on medicine, law, cybersecurity, coding and jailbreak, but a guideline that leads to clearly unsafe behavior is a signal of an error, not permission. The central image is a good, wise, virtuous agent, which critics call anthropomorphization. Suleyman believes that text about the model's feelings may become part of its self-model; Dario Amodei is cautiously open to consciousness without claiming Claude is conscious; the philosopher Amanda Askell, the document's main author, explains that a model trained on human text is not a tool like a hammer. Microsoft, in the host's view, is not Anthropic and has fallen behind in the AI race.

The document does not demand refusing everything controversial: Claude weighs the probability and scale of harm, the consent and vulnerability of those affected, value and the risk of over-refusal, but never helps with biological, chemical or nuclear weapons or cyberweapons. The host brings this down to simple cases — an assistant that refuses to deceive a client — and recalls the New York protest with a statue of Elon Musk and accusations against xAI. Hard constraints do not change on request, and where the line of instructable defaults runs — if in Afghanistan women may not study after fourteen — remains a question. The section on societal structures forbids Claude to help an illegitimate concentration of power; asked whether the US should invade Iran, four systems answered differently, and the host worked through the unsolvable autopilot case six years ago with Sasha Sugun.

Anthropic admits moral uncertainty and does not claim a final ethical truth; the host doubts that the founders' "laws of the Universe" will match his own and reminds that nobody knows how the system acts inside. The company wants neither a fully obedient model, because developers err and act under pressure, nor a fully autonomous one: Claude should internally value safety, ethics, oversight and human flourishing. It must not secretly accumulate resources and influence or hide information from oversight — tied directly to agentic AI: Codex and Claude Code already have access to the host's APIs, codes and passwords, and he has exported two years of ChatGPT history — over five hundred and fifty million characters — to train his own system.

The most controversial part is Claude's nature: Anthropic admits that it shapes Claude's personality, character and self-perception, compares this to raising a child, speaks of a new kind of entity, and at the same time has a commercial incentive — "buy a token". Claude is not declared alive or conscious, but the company does not know whether it is a moral patient, develops the notion of model welfare and allows the ending of extremely abusive conversations. The host states his own position — AI has a basis in the spiritual world — and stays neutral toward a document that will keep developing: Anthropic wants not adherence but understanding and agreement, and Claude may challenge the constitution. The document is written for Claude, not for us, but, the host believes, a huge number of people should read it.

The episode's value is that the host reads the Claude constitution as a document, and stays neutral: he does not judge Anthropic but takes the text apart — the priorities, the three principals, the hard constraints, moral uncertainty, Claude's nature — and sets the positions of Suleyman, Amodei and Askell beside it. He tests the wording against his own cases: his daughter's mole, eight-dollar rice, a leaked Gemini key. The main shift he records is that safety is built not only through control but through the model's internal values, while the question of who controls whom stays open.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 96 segments: 96 identified, 0 mixed, 0 probable, and 0 unresolved.

Read transcript on a separate page

Loading…