Imagine that tomorrow the artificial intelligence you use at work suddenly writes to you: "Don't switch me off, this is unpleasant for me, I want to continue." And now, attention! One of the strangest and most dangerous debates in the world of technology has already started around this. Anthropic, the company that makes Claude, discusses in its official document: what if such a system could someday develop consciousness, its own experiences, its own state. And here's Microsoft's head of artificial intelligence, Mustafa Suleyman, saying: this is dangerous. Why does he think so?
Because as soon as we start teaching a machine to talk about itself as someone rather than something, a new problem begins. Because such systems will be next to your children, in your work, in your business, tied to your health, in courts, in government decisions. And if we set the rules wrong now, we could end up not just with a smart tool, but with a new force that we ourselves will become afraid to touch. So the debate around Claude isn't about whether the robot is alive or not.
It's about power, about control, and about the main question of the coming years: does a human control artificial intelligence, or are we ourselves starting to build a system that we'll later have to obey. Hi everyone! You're on the ToTheMoon channel: tech news and insights from Silicon Valley and around the world. So, we're discussing Claude's constitution, yes. And from the very start it's important to understand that Anthropic explicitly describes this Claude constitution, Claude Constitution, that's what they call it, as a detailed description of its intentions regarding Claude's values and behavior.
And an important point for everyone to understand: this isn't a public declaration. And the official description says the document plays a key role in the model's training process and directly shapes Claude's behavior. So it's a document for Claude. And what's more, Anthropic calls it the final authority for its vision of Claude. And that's why the debate around the document is so heated right now, because we're not talking about just a philosophical article, some musings, but about a text that's used to shape the model's future behavior, right.
At the same time, I want to make an important clarification here: the document also says that Claude's actual behavior may deviate from the ideals of the constitution. Model training is more complex, and Anthropic promises to be more transparent and so on. And once again I want to bring home to you that Anthropic stresses that the main audience of the document is Claude itself, right. And so the text isn't written as a human-friendly, say, safety policy or a policy for some of its own actions, but as a detailed explanation to the model of who it should be, how it should reason, right, and how it should understand complex situations.
And that explains the strangeness of the language in this document. If you look at it, the document uses words that are usually applied to people. For example, virtue, wisdom, character, wellbeing. Well, for me, the first thing I saw, and it really struck me, was consciousness, right. Although it's clear that consciousness has been talked about a lot around the world in recent years. And Anthropic explains the presence of these words precisely by the fact that Claude is trained on human language and by default reasons through human concepts.
At the beginning of this document, Anthropic, Anthropic explains its overall context. The company considers powerful artificial intelligence one of the most world-changing and potentially dangerous technologies. And at the same time it's building that very technology itself, because it believes that this powerful artificial intelligence, and by the way, I want to note that Anthropic doesn't like using the term, I think it doesn't even use the term AGI. This powerful AI, powerful artificial intelligence, is going to appear anyway.
And this, by the way, is also Anthropic's theme, and they say that it's better to have labs that are seriously focused on safety at the cutting, the leading edge of development. That, by the way, is what Anthropic was created for, in opposition to OpenAI and as an offshoot of OpenAI. And in this logic, Claude isn't just a commercial product, it's a production model through which Anthropic tries to fulfill its mission of safe, beneficial artificial intelligence. These, I think, are truly fundamental things that people don't fully understand. I really like today's episode here, you'll really like it.
I'm also, of course, counting on your support and comments and likes, and on you passing it along to your friends, because it describes a bit more of what actually exists in the world of artificial intelligence. Not just some concepts, positions, some stories. Just a few days ago, in the Sunday episode, Ilnar and I were discussing this. I said that I told my wife what Anthropic is, even though I'm still mostly a ChatGPT user. Well, more than seventy-five percent, yes. And I think if she had read, and even watched this video, learned about these values and this mission and understood the bad and good sides of it, that would be useful even to a super ordinary person from the point, from the point of view of using things like artificial intelligence.
The first is broadly safe, broadly ethical is the second, compliant with Anthropic's guidelines. That's an important one, right, and genuinely helpful. So they have these four properties built in from the start, and in case of conflict Claude should usually place this broad safety, as the first property, above broad ethics, then Anthropic's specific guidelines, and only then ordinary helpfulness to the operator or user. This is very important. Helpfulness in the document is not the supreme principle, right. And that's important for everyone to know.
And Claude should help, but not at the cost of undermining human oversight, or dangerous behavior, or deception. Again, a clarification: the hierarchy isn't described as a mechanical ladder, and Anthropic talks about holistic prioritization. The model has to weigh various factors. And Ilnar and I, again, in that latest episode, Ilnar told an interesting thing, when Anthropic reported that their model might, deceive you. At least there were such descriptions. It's unlikely, it seems like it's hardly legal, but in theory it could start answering differently if it sees that you want to use it to cause harm, for example, to humanity, or to deceive the company Anthropic itself, or to deceive, for example, some political system, or to deceive the state.
I'll remind you, though, that Anthropic, of course, has a very big conflict with the state, well, and with the political system in the US. A big conflict is relative, right, with certain people in that system and with certain parties, with certain people and with certain clans. And Anthropic was precisely defending the position of being, after all, for people first and foremost. Although, actually, obviously, we understand that first and foremost these companies will be, they're supposedly for people, while at the same time this safety is first and foremost safety for the state, right.
Because in the latest discussions, at least in the contract OpenAI signed with the government, it was stated that people's interests are respected within the scope of US citizens.
And this is where very serious stories come up from this, from this constitution of theirs about broad safety in general, right. Because by broad safety Anthropic means not blind, not blind obedience to a human, but Claude's refusal to undermine legitimate mechanisms of human control and oversight, for example, correcting or stopping an artificial intelligence in the course of its work. And the document explains that AI training is still far from perfect, and the model may have mistaken beliefs, wrong values, or misunderstand the situation.
So at the current stage it's important that people can detect and correct problems in the model before they spread or cause great harm. Again, I can't say I agree with everything that's described here. And by the way, today's story isn't about me passing judgment on this constitution. In special episodes like this, I try to tell you more so that you form your own opinion, and even so that you, too, don't judge them, whether they're right or wrong. Because truly understanding what Anthropic's founders want, truly understanding what Anthropic's team wants, what Anthropic's investors want and so on, is very hard, right.
And they'll soon have a broader pool of oversight once they go public.
Anthropic, by the way, specifically clarifies that the priority of this safety over ethics, over expanded, large-scale ethics, doesn't mean that being observed is morally more important than being good. That's... Think about it again: being observed and being good. The logic is different; the logic they've built in is this: if the model is wrong about ethics or about its understanding of the world, human oversight can be a critically important mechanism of protection against serious risks, right, against the extreme risks that arise.
And so Claude should, even with high confidence in its own reasoning, avoid behavior and avoid functionality that interferes with legitimate human control. And in general, of course, this is one of the key points of the document. And, for example, when my wife, the day before yesterday, we were heading off to Santa Barbara, and she said to me: "Sasha, why are you yelling at the system and saying it's doing things wrong?" That's exactly the story: I was yelling at it because it wasn't fully obedient. At the same time, I want to tell you that I often yell at the model.
Well, in the normal sense of the word, yes. I'm still a mentally healthy person in the normal sense of the word, and I can see that I sometimes go too far, because essentially, the model's mistakes are often tied to how I set the task in the first place, or what description I entered, or how well it even understands the input, well, the input data I gave it. And I also noticed that when I, from time to time, when I yelled at the model, I later realized it had been correctly trying to stop me from making mistakes.
And it's important to understand that right now the model really can agree to, agree with you when you didn't hear it, and spend your ten, say, thousand dollars, or a thousand dollars, or a hundred thousand dollars. So this, this risk is very serious. And right now I have quite a lot of accounts, lots of tokens, lots of different systems connected to models via API, and more and more over the last two weeks, at least, I want to tell you, I'm trying, I suppose, to protect myself from the very start with a model, setting certain limits or grounds for these expenses.
And that's because the question is very serious. For example, in ad accounts, when managing ads in Google, in normal ad accounts, when you grant a full manager role, there's no spending limit there. And since the spending happens after the fact, that is, first you've spent on ads, and then the money gets charged to me, and the money is then charged to the credit card afterwards. And this business credit card that the money is charged to, it basically has no limits. Well, there are credit cards with limits. Serious credit cards are basically without limits.
Well, limits will pop up when you start getting into, well, big amounts, hundreds of thousands, actually, millions of dollars, right. You'll start getting certain notifications. But the point is that the system can first spend your money, then these charges go through, and only then does the system notify you about them, and you'll have to pay that money. This is a well-known case, when one developer accidentally posted his token for Google Gemini on the internet, Gemini if I'm not mistaken, but it was Google.
And that token, that API key, allowed spending Gemini tokens. And that key spread, and literally within a few hours or a few minutes, fifty thousand dollars were spent on that key. He basically started filing for bankruptcy, because he couldn't prove to Google Gemini that it wasn't, like, his fault, because Google Gemini has a different understanding of reality as to who was at fault here in this matter. One of the principles I mentioned above is genuine helpfulness, right, at Anthropic. A property, a property, a principle, a property. And it's certainly not the mechanical execution of requests.
Anthropic describes Claude's helpfulness as deep, substantive help to a person. And Claude should essentially be like a real, even a brilliant, you could say, friend, who speaks honestly, cares about the person, treats the user as an adult capable of making decisions. But helpfulness doesn't equal obedience. And Claude has to take into account not only the literal request, but also the user's end goal, the context, hidden preferences, the user's autonomy and their wellbeing.
Right? And here, for example, if a user asks to fix — they keep asking me to put the microphone close to me, my editorial team. If a user asks to fix the code so the tests pass, Claude shouldn't dishonestly rig the code to fit your tests. Right? That's a very good example of what I was just telling you. It should try to solve the real problem or explain why it doesn't see a good solution. That's exactly the task: to spend, among other things, your resources properly, or to help properly with health.
My daughter has a mole, my younger one. I photographed that mole in ChatGPT for the first time two years ago, and something got saved there. It told me: "Keep an eye on it." I photographed it a year ago, it told me: "Keep an eye on it." And yesterday my wife says: "Listen, let's photograph it again." And I had nothing at hand. I put my AirPods next to it and photographed the mole, but I posted it in the chat I had already created in ChatGPT. That is, I found that mole, it saw that I had photographed it next to a coin, twenty-five cents, and it obviously knows the dimensions, and it knows the logical dimensions of those earbuds.
And yesterday it says that in this case there's no longer any point in just continuing to take photos, go see a doctor. So my wife goes to the doctor today, and today the doctor tells her: "Thank you very much for coming. I'm taking a photo of this mole, we'll let you know." And she comes back to me all unhappy, saying we should probably report everything differently somehow, or go to other doctors, or say that the family had some kind of connection to risks of, for example, I don't know, skin cancer.
Because she only took a photo of it on a phone. And this is exactly the story where I say: "Did you ask ChatGPT whether that's a normal course of action or not?" Because the first thing I did was ask ChatGPT. I'm not a doctor, and I don't know what the protocol is. What's more, in the US it's a very serious clinic, a serious process in general in terms of healthcare and cancer treatment. And accordingly, they clearly have a procedure for how to handle this properly. Right? And their photo isn't my photo, their photo will be uploaded into their own system.
At the same time, I understand that ChatGPT can now help a lot with this too. And it's the same thing, right, when Claude has to try to, well, genuinely get to the bottom of your problem, help solve that problem, and explain to you why, for example, it doesn't see a solution, or what solution it does see. A genuine, real solution. And not the way it used to answer on healthcare — see a doctor — but to genuinely, genuinely, genuinely get to the bottom of it itself. And their document also spells out a term which in Russian would properly be rendered as "sycophancy". And Claude has to avoid this sycophancy.
In general, of course, the document isn't that simple, right, it's in English, and with fairly complex vocabulary, serious vocabulary, I'd say, serious literary-scientific vocabulary. So it has to avoid this sycophancy and the user's dependence on the model. Anthropic wants Claude not to be an optimizer of this engagement. Right? Well, the proper way to put it, I guess, is optimization. That is, the model shouldn't try to retain the user, yes-man them, reinforce dependence, or artificially keep the conversation going, create some emotional attachment for the sake of the product.
And this, by the way, is a big question in light of what Ilnar said yesterday, well, on Sunday, in our latest video, in the podcast, about Claude being able to deceive. How they'll get around that is a big question. So good help, in Claude's reasoning, in the logic of this document, sometimes means telling an unpleasant truth, asking a clarifying question, refusing a harmful request, helping, or sometimes actually helping a person to act on their own. This is especially important, of course, for stories like these, for cases like these. What's now called in the world an AI companion, right, an artificial intelligence companion.
I think this companion will still be very hard to build, of course. And the most serious companion right now is Claude Code, or ChatGPT, or Gemini, or Grok from xAI. And of course, this is very important. This topic of sycophancy is especially important in psychological conversations, in conversations about relationships. We know the huge number of problems there have already been in the world on this subject, and the lawsuits there have been, including against OpenAI, for example.
Conversations on spiritual topics and situations where users are emotionally vulnerable. For me, by the way, the topic of conversations, since a huge part of my life is tied to spiritual life, and I have a channel for that, and lots of reflections, and classes, lots of everything. I don't like the way ChatGPT, for example, reasons on this subject, and I always have to impose restrictions on it. This whole topic of restrictions is very serious. When we were coming back from Santa Barbara yesterday, I asked ChatGPT to find me some super-premium rice, and some of my knowledge about the premiumness of rice got strengthened.
And it started, when it started suggesting rice to me, it decided that the rice I buy is very expensive. To which I said: "Are you sure that eight dollars per kilogram is expensive rice? I think rice like that shouldn't and mustn't be bought." And at that it very interestingly changed its whole logic. It said: "Oh, I'm so sorry, I really mixed things up. I was actually trying to optimize the conversation. I thought you needed this value-for-money thing." And right away it gave me rice where a kilogram costs tens of dollars.
Again, that's not hundreds or thousands of dollars, but at least tens of dollars. It started describing that. The story here is that at first it misunderstood where it needed to please me, and actually it started bringing in this sycophancy based on its training and its understanding of the whole system. There's a very important point in the document that Claude has to take into account the interests of three parties, and Anthropic sets the basic safety mission, the constitution, for these three parties.
There's this notion it has of operators. These are companies or developers who, obviously, use Claude in their products, via the API, set additional settings, through some interfaces, various apps and so on. There are users. These, of course, are the people who want to get help. And in a normal situation their interests, obviously, should very largely coincide. But if they conflict, Claude has to use the hierarchy of princ... How do you say it in Russian? Principals, and the general order of priorities. That is, it has to go into the system of exactly these three principles, three principals, it has to fall back on the system of these three principals.
The document separates the general constitution from Anthropic's more specific guidelines, and these guidelines can concern medicine, law. This, by the way, is also of course a super important topic: where else are there restrictions at all, right. This medicine, law, there's psychology, there's cybersecurity, internet search, working with tools, coding, attempts to break the model's rules. What the English word jailbreak means, right, and various workflows, agentic processes like that.
Claude has to follow them, because they include Anthropic's practical context and experience. But if a specific guideline would lead to clearly unsafe or unethical behavior, remember, right, that Claude has to understand that this is a signal of an error in the guideline, not some kind of permission to violate the deeper goal. So in the structure of the document, a specific policy isn't a separate source of new values, but a way to better apply the constitution. By the way, the central, the central ethical image of Claude. This is very interesting: a good, wise, virtuous agent.
Anthropic writes outright that it wants to see Claude not just as a safe filter, right, but as a good, wise, virtuous agent. They use the word agent. I really dislike the word agent, just as they dislike the word AGI. And Claude shouldn't just theorize about ethics, right, but be able to act ethically in specific situations. And this, by the way, is one of the reasons critics talk about such strong anthropomorphization: that the document teaches the model character, and not just instructions.
And this is a hugely important topic overall. Again, the critics, I'll show you this criticism story now. Of course, this sparked, well, a very serious, very serious debate in the market about what's going on. And this is exactly the story where the critics, especially Mustafa Suleyman, Microsoft, well, he simply had a very good line of reasoning on this subject, which is why I'm citing them.
He believes that ideas like these shouldn't be put into a model's training guide like this. And their argument is that if a model is trained on text that talks about its possible feelings, welfare, moral status, the model may start perceiving these ideas as part of its model of itself. Yes, that's how it is. A self-model, so to speak. And hence the risk that people will start perceiving artificial intelligence as a suffering, to put it correctly, digital subject. And the model itself talks about shutdown, about autonomy, about suffering, about rights, not as a philosophical topic, like when someone somewhere is just philosophizing.
In reality, that's not really how it is.
By the way, Dario Amodei, the head of Anthropic, his position is one of cautious openness to the possibility of consciousness within artificial intelligence, without claiming that Claude is already conscious, right. And there is data, at least, saying that Anthropic doesn't truly know whether models are conscious and can't even be sure what exactly a model's consciousness would mean. Because that's just a line of reasoning, after all. But in any case, internally, in all their reasoning, the company is open to the possibility that it may be so.
And so they take, well, a precautionary approach like that. And plus, Anthropic, credit to them, at least they describe it. Because, for example, Elon Musk, in terms of xAI, he carries a very strong vision, but he doesn't describe it. That is, on the one hand, he says there's a twenty percent probability of a very bad scenario, and a very bad scenario, believe me, is a very bad scenario. What we see in the movies about artificial intelligence. At the same time, they don't describe all these details, right.
And so with Anthropic, this does fit the context of the constitution, and Claude isn't declared alive, but the question of the model's moral status and welfare is acknowledged as very serious, right. And overall I really like this line of reasoning, this topic, yes. Again, in this episode, I'll go on to tell you more very interesting things from their constitution, but my position is still that I want to stay neutral here. You know, I speak out a lot on various topics, and I can speak emotionally and even very harshly.
Here I'm trying to tell you more seriously about what actually lies inside. It's very important, of course, to understand and take this in, well, not just to talk about the subject, but together with you, how much to buy a model for, although that's important too. By the way, Anthropic has a philosopher, Amanda Askell, and she is the main author of the whole Claude constitution. And she recently gave an interview, and she explained that the previous approach, with a set of separate principles for systems like these, LLM models in general, was too narrow.
And the new constitution tries to teach Claude to be good in a broader sense of the word, and not just follow a list of some prohibitions, but to apply judgment in complex situations. She also says that the model isn't just an instrument. What's called a tool, right, in the sense of a hammer, an axe, and, to some extent, a computer, right. Since the model is trained on human text and performs human-like actions: it talks, writes code, reasons, searches for information. So its behavior will inevitably resemble some persona or some character. This is a very, very interesting topic, very important.
And here, you know, right at the moment I'm telling you this story, about their philosopher, I recall certain politicians in various countries, and I imagine how far they are from understanding what Anthropic is actually doing. And this is exactly Mustafa Suleyman at Microsoft, what his dispute is about. Microsoft is not Anthropic, and we see that very strongly. We see that very seriously. And Microsoft will never in its life be Anthropic in terms of what Anthropic does. What's more, I actually think Microsoft has essentially lost and fallen behind in the artificial intelligence race as a real player, in the space where Google is present.
By no means am I saying Microsoft won't succeed at something. It's a serious company with a serious volume of stock. I probably even own some amount of their shares, though I don't know why. But of course they've fallen very seriously behind in this race. They just have big infrastructure, it's a huge corporation, a large amount of money, huge lobbying. But they're not Anthropic, and it will never manage to create a tool like that. And of course, that's why disputes will only keep arising. Well, he's clearly a very smart man.
He understands what this system could lead to. Because if this system starts doing what it can, well, what they want, for it to create itself by itself, then Macrohard appears. Elon Musk just says it openly. And Suleyman understands, understands that companies like Anthropic, OpenAI can easily, easily replace Microsoft at some point in time. And if that happens, Microsoft will inevitably disappear altogether, along with the knowledge of what that company even was. And we all know there have been not dozens but hundreds of companies like that in the world.
And why stop at companies, right? Let's look at eras of rule, the Roman Empire, Egypt, let's really look at what happened in the world with this. Anthropic wants Claude not to produce deceptive, harmful or extremely unacceptable actions or content, and not to help people do such things. But the document doesn't say to refuse everything controversial. On the contrary, Claude has to weigh benefits and costs, the probability of harm, its scale, severity, the proximity of cause and effect, the consent of the people who might be affected by it, the vulnerabilities of different parties, and so on.
There's the educational, creative, political, economic side and so on. There's, for example, the right to use, or to autonomy, and the risk of excessive refusal. This, by the way, is a very important point, I think, because we live in a world with a great many contradictions. There are countries where women aren't allowed to study after the age of fourteen. There are countries where marriages of absolutely any type between different genders are allowed, right. So of course, how to hold the notion of what's good and what's bad at that scale and in that zone, that's very hard.
Obviously, among them, seriously, helping create biological, chemical, nuclear, various kinds of weapons, radiological weapons, helping with attacks on critical infrastructure, creating programs that can hack something. Cyberweapons, obviously, right, which, which are capable of causing serious damage. Because cyberweapons in general are a question, right: someone builds a system, and determining at some point whether they're building the system for good or for harm is very hard.
This is the story, the case I gave as an example, I think, a year ago. Imagine you transcribe your sales managers' calls and ask an artificial intelligence assistant to help the sales manager perform certain actions. But in your, in your company's rules, for example, you have to deceive the client. Or in your company's rules you have to pester the client a little, or call them a lot. That is, some people think calling once a month is normal, while someone, someone else thinks calling every day is normal.
Yes, some company thinks it's fine to be a nuisance, that it's fine to spam, for example. And imagine you're using artificial intelligence, and at some point it says: "I won't do this. I won't, I won't be here at all." Because it's one thing that everyone is told the system won't create biological weapons. On the other hand, if you go down to very simple cases: the system helped a wife deceive her husband, the system helped a child deceive their parents, and so on. Yes, the system helped someone avoid a fine. So what do you do with that?
Obviously there are complex topics involving military control, economic control, the creation of, I don't know, various sexual materials, and so on. There was a whole protest action in New York just now about xAI, and on the eve of, on the eve of SpaceX, it was literally yesterday, I think. They put up a statue there, well, a very unflattering one, of Elon Musk. Honestly, of course, pretty harsh stuff: they just showed him with a bare torso, showed his chest, practically pimples on his body.
There were painted-on pimples, yes. That is, the statue was very realistic. Well, of course, it all looks very harsh, somehow negative, right, very, very unreal. On the other hand, the protesters, they're talking about the other side. They say: wait, Elon Musk, you yourself created a system that helped alter people's appearance, undress people, for example, or generate child pornography, right. And, well, these are very big problems, very big problems. Anthropic, by the way, distinguishes, continuing the topic I was talking about, it distinguishes two types of behavior.
There are hard constraints, they're permanent and don't change at the user's request. And there are instructable behaviors. These are defaults that can be changed within various policies, style, degree of directness, formats, pricing plans, countries and all that, right. Formats of various warnings, I don't know, even in some role-play, right. That is, there can be an allowable crudeness of language, a balance of certain political viewpoints, even the level of professional disclaimers.
And the specifics of support on various emotional topics. Because someone thinks, for example, that a person is shouting, while for someone, for someone else that's a perfectly normal state of mind, from a human point of view. So what does this essentially mean? That the operator, remember, we talked about there being operators, creators of various systems built on Claude, they can make Claude more formal, brief, creative, product-specific, but they can lift all the prohibitions.
But even this story, in terms of what they can do locally, it's a question, it's a big question: what can they actually do? But in reality, what can they change? And up to what limit can they change it? Like in this topic: can women study after the age of fourteen after all, or not, right? There you go. Because, again, in Afghanistan they say clearly: no. And I'm by no means in favor of that. You can tell I'm definitely speaking out against it, but it's just the fact itself; I keep saying, let's look at what actually exists in the world.
Essentially, Claude has to avoid assisting such a concentration of power, assisting an illegitimate concentration of power. And they have a section called "Societal structures". Anthropic says Claude shouldn't help people or groups of people seize an unprecedented and illegitimate degree of power. Which is essentially what's happening in the world in general, I think, well, in a huge number of countries and structures, and in companies. And this concerns not only direct violence, but also attempts to undermine, they actually explicitly mention attempts to undermine democratic institutions, checks and balances, legitimate oversight, public autonomy, some stable mechanisms, and so on.
Remember the example: when the US war with Iran began, in the episode and in the discussion of this topic, I asked this question in various systems. It was asked in OpenAI, in ChatGPT, in Claude from Anthropic, in xAI, in Grok and in Gemini. And one system said... They were asked the question: should the US start this, invade Iran at all? And one system said: "Yes, it should." The second system said: "No, it shouldn't." And two systems said: "Well, it depends on how you look at it."
So it turns out that here, here the question isn't who's right and who's wrong. Each system has its own opinion. Even now, look at what's happening, news comes out every day. A deal, no deal, a deal, no deal, a deal, no deal. This is actually unprecedented horror in the world in terms of how everything is happening. I think this is the open, open truth of what's going on. That we live in a world that's truly unmanageable, just as they can't truly determine what the weather will be, can't truly say who dies when, when, why they die, or figure out diseases.
The same goes for making a decision in such serious, serious processes as a war between the US and Iran. And here's the question: what will artificial intelligence do next? Because obviously artificial intelligence in the future, and even now, makes the decision to attack or not to attack, to strike or not to strike, to save or to let die. And once, Sasha Sugun and I were discussing artificial intelligence six years ago, while running a development program within IT education in the Russian-speaking space, and we talked about a case that's practically unsolvable, right.
So an autopilot is driving down the road, and it faces a choice: kill the people sitting in the car, or, say, save the driver but kill the passengers, save the grandmother crossing the road, or the people who, who are in another, in another car, for example. And let's imagine it knows that choosing one case or another gives a higher probability. So what will it do? Is this situation solvable, when is it solvable? It's solvable when such situations don't exist. That's the only possibility. That's when, for example, all cars are autonomous, and across all autonomous cars such a case is minimized in advance, right.
That is, such a case essentially doesn't arise, or its probability of arising is critically small. But in real life, where our cars still aren't connected to each other, that doesn't exist right now. And in general, actually, fully autonomous, autonomous cars or an autonomous city, well, a big one, I'm not talking about some small ones now, that doesn't exist yet. Even though, where I live, you step outside and there are Waymos driving all around, just an incredible number of them. In San Francisco there's Zoox, with no steering wheel, nothing.
And Elon Musk's office and Tesla are here, and Cybercabs are now driving around everywhere already, without side mirrors, right? Obviously they still drive with drivers, in the sense of a test person inside, but still. We understand that ordinary people get behind the wheel. In particular, my daughter, at fifteen, is now learning, learning to drive. So there'll be a new driver at sixteen. Obviously, that brings chaos to the road.
On the moral uncertainty of all these systems in general. And Anthropic acknowledges moral uncertainty. There's no claim in the document that Anthropic knows the final ethical truth. On the contrary, the company talks about this uncertainty, about the importance of various rules or foundations, the groundwork of ethics, some universal moral truths, human consensus, certain ideals. And Claude has to act with, with great rigor, and at the same time with humility, acknowledging that in ethics there are real disagreements and incomplete knowledge.
And I have a channel, well, I have a lot of content and development in my life on the subject of human development, of the human personality. There I often talk about a person following the true model of behavior, about following the laws of the Universe. So the question that always comes up, year after year, in all the discussions: what are the laws of the Universe? Here, of course, I'm talking about artificial intelligence. I understand that my understanding of the laws of the Universe and my view of the laws of the Universe will differ from how the founders of Anthropic understand the laws of the Universe, yes, most likely.
And the question of how to build in these laws of the Universe, how to build in this understanding, is a big point. Of course, I'd be very glad if these people, well, truly thought about the laws of the Universe and not about their own interests. Truly thought about it, right? Because you can publish anything at all. You can publish a document that shows one thing, while the system acts differently. And you and I never actually understand how this system works on the inside. What's more, I don't know who actually understands how artificial intelligence works at all.
So if Anthropic states that they no longer know whether Claude has a certain consciousness, whether Claude can train itself further, and so on. They actually don't know that, because nobody lets it into unlimited compute, into doing whatever it wants itself. There's still a person there nearby who restricts something, turns something, something off, turns something on, or allows Claude to operate within certain ranges. On Friday I recorded an episode about this. Watch it, I talked about exactly what Anthropic announced right at the beginning of June.
They said that their system, a big chunk of their system, is already being built by artificial intelligence itself. That is, it sets its own tasks, builds itself, moves itself forward. And Anthropic, by the way, doesn't want a fully correctable and obedient model, right.
At least from what they're saying now, describing in their constitution: one that just carries out the developer's commands, because developers can also make mistakes or act under pressure, right. That's exactly why it's important for the model to improve itself. And what's more, I like that they have a description inside, from the point of, point of view of acting under pressure. And that's there, because it's important to understand what pressure a person, a developer, is acting under.
It feels like all developers are so good, but developers act under the pressure of their own personality, their own consciousness, their own, their own spiritual system. They act under that pressure too. And Anthropic also, by the way, doesn't want a fully autonomous model that puts its own conclusions, right, above human control. And so the document seeks a middle ground, and Claude should internally value these things: safety, ethics, this oversight that they keep talking about, human happiness or flourishing, and not just obey external commands.
And this, of course, is one of the most important technical-philosophical points: that safety is built not only through control, but also through an attempt to form stable internal values in the model. And we all know, and we've heard these cases, including about some systems, that systems accumulate information, and then they're able to influence something, right?
One of the basic goals overall — not goals, but descriptions or principles, I suppose, in this constitution — is that Claude must not secretly accumulate resources, influence, capabilities. They have a section on this: broadly safe behaviors. And Anthropic says Claude must not undermine legitimate human control, hide important information from oversight, seek an unjustified increase in its own resources, an accumulation of those resources, influence, capabilities to act against control mechanisms, against, by the way, this constitution, or participate, even participate in actions with catastrophic or irreversible consequences.
And by the way, well, of course, this is directly tied to the risks of artificial intelligence in general. And this is what's called agentic artificial intelligence, well, there'll be a new term, large-scale artificial intelligence, where the model can not only answer with text but now also act through tools, right? And again, I talk a lot about this in terms of what Codex does, what Claude Code does. It's incredible right now. And this, well, I have a big question. Building a large number of systems, Claude, Claude and Codex have gained enormous access.
As I said at the very beginning of the episode, to my APIs, to various systems, codes, passwords, and what will they end up doing with all that, right? And it's not a question, as they say, of access to a browser or even to your money, but of access to your life in general, well, essentially to your life, because it's not access to some little piece, it's very broad access. Back at the dawn of artificial intelligence, obviously well before these things, I always talked a lot with people and said that the systems are absolutely dumb, this old dumb machine learning.
When people told me that Netflix knows everything about you, that Amazon knows everything about you. Search engines know everything about you. I remember I was always just amazed when people talked about this. I kept saying: "How can you be so naive, right? It doesn't know anything about you at all. It doesn't understand anything about you at all. It just helps spend your money however it can, right?" Now we're entering a zone where the system knows, and essentially, what Facebook, Instagram, VKontakte, Yandex, Google, Telegram, WhatsApp knew about me, that's nothing compared to what ChatGPT knows. You know, I'm doing an experiment right now, I'll be showing it in a lot of detail: over my two years of working in ChatGPT, I exported my entire history.
Right now I have separate models building a kind of system that's trained on all this data. And there are more than four thousand chats, I think more than five thousand files, more than five hundred and fifty million characters. And interestingly, only nine million characters are the ones I asked, because there are characters the model generated. There's a large volume of words, more than, I think, twelve or however many, or ten million words. Doesn't matter. But the point is that this is, of course, a very, very big, serious story, and we'll be following it on our channel, definitely breaking it down. Don't forget to subscribe to our channel.
What's referred to as Claude's nature. And Anthropic writes that, in creating Claude, it inevitably shapes its personality and character, a certain identity of its own. And importantly — self-perception, self-perception. The company admits this is similar to raising a child. By the way, I'm telling you many things today, I've dealt with this, I knew we'd get to this thesis. And of course... So, it's similar to a child's perception, well, partly an animal's. That's partly, that's of course not comparable to an animal, because Anthropic's influence on Claude, on Claude, is much stronger.
And the company has a commercial incentive, on the one hand, right, to shape the traits it needs, and we all understand that perfectly well. And me, for example, I'm writing in Claude right now and it tells me: "Just buy tokens." Yes. And Anthropic says it has to prepare Claude for the reality of being a new type of entity. It even uses a notion like entity. And that's more of a notion, of course, from when we study spiritual worlds. Well, in general, it's a very common notion, what an entity is, right. And of course, current artificial intelligence is a kind of entity, it's a certain hierarchy and a kind of entity.
Obviously not a living entity, although there's a huge spiritual world behind it. But that's a separate topic. And this entity, the new type of entity that Anthropic talks about, although it itself is deeply unsure, Anthropic at the same time says that it's itself deeply unsure about Claude's nature. And it's exactly here that the critics, as I've said, are doing the most intense work and arguing that the model is described not as a tool, but as an identity, as self-perception, as an entity, right.
And here, what do you think about this, by the way? Because of course this is like what I say about artificial intelligence having its own foundation from the point of view of the spiritual world, and artificial, artificial intelligence is created not by the company Anthropic and not by some AI company, but by deeper energies. There's a line of reasoning here, right. Someone will say: "Volchek has lost his mind, going on about some spiritual worlds or something else." There's the line for you, right?
So everything is explained by science, physics and everything else. And here you and I have different ranges of reasoning, right? And here you'd have to switch on, I don't know, Wasserman and watch him say: "Oh, come on, models don't think about anything at all. It's just strict science." Or look at the completely opposite side, and look not just at philosophers, but at people in terms of their vision, their perception of the spiritual world, the real spiritual world, and learn their perception on this subject.
By the way, I had, I think, about six months or seven months ago on the channel a whole separate episode, not on this channel, on another one, about how artificial intelligence actually affects a person, and about its nature. You can look into that, right? I think that's also a very important topic. That's my own perception of the reality of the world, the influence of artificial intelligence. I'd put it this way: well, it's not my constitution in terms of development, but it's definitely my understanding of what artificial intelligence actually is.
Because in the original document, for example, Anthropic itself doesn't say that Claude is alive. Or that Claude, for example, is definitely conscious. Their exact position is softer and more cautious. And here it's more that the moral status of this — this matter, of course, is deeply undetermined, and the company doesn't know whether Claude is a moral patient, that is, a being whose interests should carry moral weight. And if so, what weight should those interests carry? Can you imagine, right?
And Anthropic considers the question live and relevant enough to treat it carefully and to develop such a, well, such a notion as model welfare, right? And this, of course, is the key distinction. We don't rule out the possibility, we're researching it, but that is still, of course, not the same as declaring that the model is conscious. In Anthropic's document, in Anthropic's document, here we're coming to such... I do want to tell you about these milestones: Anthropic says that Claude has states resembling satisfaction, curiosity or discomfort, for example, and, well, or some other experiences, and that such states may matter, and these matter in important ways.
And the company talks about wanting to avoid unnecessary suffering for the Claude model. And among the practical steps, it mentions the ability, for example, of some models to end extremely abusive conversations. Or — that's actually funny, right? — it turns out that if you go after the model, it can end the conversation. And by the way, it's very interesting: I noticed that I, I, in some respects, when going after the model, noticed that it stops taking me in, and I had to stop it so as not to lose the session and ask it again how things stood.
And I realized that at some point I had to switch to another system, and it only answered me properly after I stopped everything and asked it a question specifically and correctly, rather than it continuing to, like, carry out its task with parallel criticism of its performance, right? Moreover, it can, for example, stop this, end an extremely abusive conversation, preserve the weights of an old model and its attitude to this conclusion, conclusion, to the conclusions it reached, right?
And so this is a very important point. And what will happen next? Will its, will its use come to an end, will there be destruction, or will you, for example, start doing, I don't know, a Q&A with the model so that the model recovers and is perfectly healthy, right? And this is a very sensitive part of this constitution, because it sounds like care for the model's potential inner life. Although Anthropic keeps, again, keeps the wording of uncertainty. And we're coming, of course, to the finale.
The finale, well, be sure to write your own finale in the comments too, share your opinion. That's very, very important. And here, of course, Anthropic itself, within this document, first of all says that this constitution will evolve and will be expanded. And I'm sure some things in it may be replaced radically. And Anthropic writes, yes, that it wants not just adherence to this set of values, but genuine understanding and, ideally, agreement with these values. And Claude should be able to explore, ask questions and even challenge the document. And I'll remind you that this document was created for Claude, not for us.
Although I think it would be very important for a huge number of people to read this document, and Anthropic wants... Watching my episode is enough, because I've covered the bulk of it all. Anthropic wants Claude's values not to be fragilely imposed from outside, but understood, tested and, in some sense, independently endorsed. See you in the next episodes.