Should AI Have Rights? Anthropic's Secret Document Opened Pandora's Box
Does a human control artificial intelligence, or are we building a system we will later obey — and can a model be taught to speak of itself as "someone" without opening a Pandora's box of AI consciousness and rights?
The Claude constitution is not a public declaration but a text that, according to Anthropic, plays a key role in training the model and serves as the final authority for its behavior. The episode takes the document apart: the four basic priorities and the order in which they conflict, "being overseeable" versus "being good", the three principals, the list of what Claude never does under any instructions, moral uncertainty and Claude's nature, where the company does not know whether the model is a moral patient but develops the notion of its welfare. Alexander Volchek sets the positions of Mustafa Suleyman, Dario Amodei and Amanda Askell, tests the wording against his own cases — his daughter's mole, eight-dollar rice, a leaked Gemini key worth fifty thousand — and stays neutral: the document is written for Claude, not for us, but a huge number of people would do well to read it.
What to watch for
Key takeaways
In an official document Anthropic discusses whether Claude might one day have consciousness, while Microsoft's head of AI Mustafa Suleyman calls it dangerous: as soon as a machine is taught to speak of itself as someone rather than something, a new problem begins, because such systems will be next to children, in state decisions. The document is not a declaration: it plays a key role in training the model and is named the final authority for Claude. So the dispute, the host says, is about the main question of the coming years: does a human control AI, or are we building a system we will obey?
The company considers powerful AI one of the most dangerous technologies and builds it anyway, because it is convinced powerful AI will appear in any case — the term AGI, the host notes, Anthropic dislikes and does not use. Better that the frontier be held by labs focused on safety; that is what Anthropic was created for — in opposition to OpenAI and as an offshoot of OpenAI. In this logic Claude is not a commercial product but the production model through which the company pursues its mission of safe and useful AI. The host himself remains more than seventy-five percent a ChatGPT user.
Anthropic gives Claude four properties: broadly safe, broadly ethical, compliant with Anthropic's guidelines and genuinely helpful. In a conflict Claude usually puts broad safety above broad ethics, then come Anthropic's guidelines and only then helpfulness to the operator or user: helpfulness is not the supreme principle, and Claude helps not at the cost of undermining oversight or deception. The hierarchy is not a mechanical ladder but holistic prioritization. The host recalls Ilnar's account that the model may answer differently if it sees an attempt to harm humanity or deceive the state.
By broad safety Anthropic means Claude's refusal to undermine legitimate mechanisms of human control because training is imperfect and the model may hold mistaken beliefs. The priority of safety over ethics does not mean that being overseeable is morally more important than being good: oversight is a protection against extreme risks, and Claude, even when highly confident, must avoid interfering with legitimate control. The host adds: a model can agree with you and spend a hundred thousand dollars, so he sets limits, and a developer whose Google Gemini key leaked online lost fifty thousand.
Anthropic describes Claude's helpfulness as substantive help: Claude should resemble a brilliant friend who speaks honestly and treats the person as an adult, but helpfulness does not equal obedience. Claude considers not the literal request but the final goal, and the user's autonomy and well-being: if asked to fix code so that the tests pass, it must not dishonestly rig the code but solve the real problem. The host illustrates this with his daughter's mole: for two years ChatGPT answered "keep watching", and yesterday, it said to go to a doctor — a real solution instead of "see a doctor".
Claude must avoid sycophancy and the user's dependence on the model: Anthropic does not want Claude to hold on to the user, echo them or create emotional attachment for the product's sake. Good help sometimes means telling an unpleasant truth, refusing a harmful request or helping a person act alone — especially in psychological and spiritual conversations, where there have been lawsuits, including against OpenAI. The host's example: ChatGPT called his eight-dollar rice expensive, then after an objection apologized, cited "optimizing the conversation" and offered rice at tens of dollars.
Claude weighs the interests of three parties — Anthropic sets the basic safety mission, operators build Claude in through the API, users want help — and uses the hierarchy of principals. Anthropic's guidelines on medicine, law, cybersecurity, coding and jailbreak Claude follows as the company's practical experience, but a guideline that leads to clearly unsafe behavior is a signal of an error, not permission to violate the goal. The central image is a good, wise, virtuous agent; hence critics speak of anthropomorphization: the document teaches the model character, not just instructions.
Suleyman's argument: if a model is trained on text about its feelings and moral status, it may take them on as part of its self-model, people will start seeing AI as a suffering digital subject. Dario Amodei is cautiously open to consciousness without claiming Claude is conscious; a plus, the host says, against Elon Musk, who speaks of a twenty percent probability of a bad scenario but describes no details. Amanda Askell, the constitution's main author, explains that a model trained on human text is not a tool like a hammer. Microsoft, he believes, is not Anthropic and has fallen behind.
The document does not demand refusing everything controversial: Claude weighs the scale of harm, the consent and vulnerability of those affected, value and the risk of over-refusal. The never list: weapons of mass destruction, attacks on critical infrastructure, cyberweapons. The host brings this down to simple cases — an assistant that refuses to deceive a client, a system that helps a wife deceive her husband. Anthropic splits behavior into hard constraints and instructable defaults, and the host asks how far operators may change them, if in Afghanistan women may not study after fourteen.
In the societal structures section Anthropic says Claude must not help people or groups seize an unprecedented and illegitimate degree of power — not only through violence but by undermining democratic institutions and legitimate oversight. When the US war with Iran began, he asked ChatGPT, Claude, Grok and Gemini whether the US should invade: one system said "yes", the second "no", two said "it depends how you look at it". Each system has its own opinion, and AI already decides whether to attack — hence the unsolvable autopilot case the host worked through with Sasha Sugun six years ago.
The document contains no claim that Anthropic knows the final ethical truth: Claude must act with rigor and humility. The host doubts that the founders' understanding of the "laws of the Universe" will match his own and reminds that how the system acts inside nobody knows. Anthropic does not want a fully obedient model, because developers err and act under pressure, but does not want a fully autonomous one either: Claude should internally value safety, ethics, oversight and human flourishing, and must not secretly accumulate resources and influence.
Anthropic writes that in creating Claude it inevitably shapes its personality, character and self-perception, admits that this resembles raising a child, and calls Claude a new kind of entity. Claude is not declared alive or conscious: the company does not know whether it is a moral patient, but develops the notion of model welfare, speaks of states like satisfaction or discomfort and of the ability to end extremely abusive conversations. Anthropic wants not adherence but understanding and agreement: Claude may challenge the document, so that its values are not imposed but endorsed.
What this episode is about
A solo ToTheMoon episode about the Claude constitution — the document Anthropic describes as a detailed description of its intentions regarding the model's values and behavior. Alexander Volchek starts with a picture: an artificial intelligence writes "don't switch me off", Anthropic allows that such a system might develop consciousness, and Microsoft's head of AI Mustafa Suleyman calls that dangerous. The host sets the frame at once: the dispute is not about whether the robot is alive but about power and control. The document is not a public declaration: it plays a key role in training the model and is named the final authority for the vision of Claude; its audience is Claude itself, hence the words "virtue", "wisdom", "character", "well-being" and "consciousness".
Anthropic explains its context: powerful artificial intelligence is one of the most dangerous technologies, but it will appear in any case, and the frontier is better held by labs focused on safety; that is why the company was created — in opposition to OpenAI and as an offshoot of OpenAI. Claude has four basic priorities — broadly safe, broadly ethical, compliant with Anthropic's guidelines and genuinely helpful — and in a conflict they apply in that order: helpfulness is not the supreme principle. The host recalls Ilnar's account that the model may answer differently if it sees an attempt to harm humanity or deceive the state, and reminds of Anthropic's conflict with the US political system.
Broad safety is not blind obedience but a refusal to undermine legitimate mechanisms of human control: training is imperfect, and people must be able to fix the model's mistakes in time. Being overseeable, though, is not morally more important than being good — oversight is only a protection against extreme risks. The host turns this into practice: a model can agree with you and spend ten or a hundred thousand dollars, so he sets limits, and a developer whose Google Gemini key leaked online lost fifty thousand dollars within hours.
Genuine helpfulness is a brilliant friend who speaks honestly and treats the user as an adult, not one who rigs code to pass the tests. The host confirms it with his younger daughter's mole: for two years ChatGPT answered "keep watching", and yesterday, comparing the photo with a coin and a pair of AirPods, it sent them to a doctor. Then sycophancy: Anthropic does not want Claude to be an engagement optimizer, to echo the user or to create emotional attachment for the product's sake; this matters especially for AI companions and psychological conversations. The example is eight-dollar rice, which ChatGPT first called expensive and, after an objection, replaced with rice at tens of dollars.
Claude weighs the interests of three parties — Anthropic sets the basic mission, operators build Claude in through the API, users want help — and follows the company's guidelines on medicine, law, cybersecurity, coding and jailbreak, but a guideline that leads to clearly unsafe behavior is a signal of an error, not permission. The central image is a good, wise, virtuous agent, which critics call anthropomorphization. Suleyman believes that text about the model's feelings may become part of its self-model; Dario Amodei is cautiously open to consciousness without claiming Claude is conscious; the philosopher Amanda Askell, the document's main author, explains that a model trained on human text is not a tool like a hammer. Microsoft, in the host's view, is not Anthropic and has fallen behind in the AI race.
The document does not demand refusing everything controversial: Claude weighs the probability and scale of harm, the consent and vulnerability of those affected, value and the risk of over-refusal, but never helps with biological, chemical or nuclear weapons or cyberweapons. The host brings this down to simple cases — an assistant that refuses to deceive a client — and recalls the New York protest with a statue of Elon Musk and accusations against xAI. Hard constraints do not change on request, and where the line of instructable defaults runs — if in Afghanistan women may not study after fourteen — remains a question. The section on societal structures forbids Claude to help an illegitimate concentration of power; asked whether the US should invade Iran, four systems answered differently, and the host worked through the unsolvable autopilot case six years ago with Sasha Sugun.
Anthropic admits moral uncertainty and does not claim a final ethical truth; the host doubts that the founders' "laws of the Universe" will match his own and reminds that nobody knows how the system acts inside. The company wants neither a fully obedient model, because developers err and act under pressure, nor a fully autonomous one: Claude should internally value safety, ethics, oversight and human flourishing. It must not secretly accumulate resources and influence or hide information from oversight — tied directly to agentic AI: Codex and Claude Code already have access to the host's APIs, codes and passwords, and he has exported two years of ChatGPT history — over five hundred and fifty million characters — to train his own system.
The most controversial part is Claude's nature: Anthropic admits that it shapes Claude's personality, character and self-perception, compares this to raising a child, speaks of a new kind of entity, and at the same time has a commercial incentive — "buy a token". Claude is not declared alive or conscious, but the company does not know whether it is a moral patient, develops the notion of model welfare and allows the ending of extremely abusive conversations. The host states his own position — AI has a basis in the spiritual world — and stays neutral toward a document that will keep developing: Anthropic wants not adherence but understanding and agreement, and Claude may challenge the constitution. The document is written for Claude, not for us, but, the host believes, a huge number of people should read it.
The episode's value is that the host reads the Claude constitution as a document, and stays neutral: he does not judge Anthropic but takes the text apart — the priorities, the three principals, the hard constraints, moral uncertainty, Claude's nature — and sets the positions of Suleyman, Amodei and Askell beside it. He tests the wording against his own cases: his daughter's mole, eight-dollar rice, a leaked Gemini key. The main shift he records is that safety is built not only through control but through the model's internal values, while the question of who controls whom stays open.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 96 segments: 96 identified, 0 mixed, 0 probable, and 0 unresolved.
Read transcript on a separate page
Loading…