Skip to content
Transcript

Transcript · extra15 · AI Safety Is Not a Fight Against an “Evil Model,” but a Fight Over the Right of People and States to Set Its Rules — ToTheMoon

English machine translation of the Russian-language episode. Timecodes open the source video.

Episode overview
00:00:00–00:00:50To TheMoon tonight.
Alexander Volchek00:00:00

And there was a big clash between Anthropic and the Department of War, yes. And, well, everything about security and military industry. The president's administration said, and the president himself said that Anthropic would not work anymore. Yet AI safety is not just about war and not just about the security of the country, but it is. A little, uh, a lot of chic-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi We see that everyone has the opportunity to rewrite the rules.

Mentions: Anthropic · AI safety
Max Grigoriev00:00:22

If the company says they're interested in model safety, we should see what percentage of their budget is spent on model safety. OpenAI spends less than a percentage. How do we make models ' interests equal to those of mankind? And OpenAI, Anthropic, AI et al., they will in fact lay down the principles on which this intelligence will be equated. In this area, in the area of the safety of artificial intelligence, it will not be boring.

00:00:50–00:04:04Main developments in the world of technology
Alexander Volchek00:00:52

Hello, everybody! We're on ToTheMoon. Technological news, Silicon Valley sites. Today, we have a special episode with Max Grigoriev again. Uh, who looked at the last edition, write off how much you liked him, didn't like him. And the impressions that you didn't see, look necessarily. It was a very good issue of the culture of the Cremniy Valley companies and, in general, all the leaders who are now in AI. Today we have a subject, aaa, all that has to do with the safety of artificial intelligence, how companies interpret it, which create artificial intelligence. Actually, it's a very good subject as a continuation. Everything about AI safety, yes. That's a very good subject. To continue what we discussed with Max last time on company culture. And that's, by the way, Max's culture, she's interesting. I have different people there, partners, people who know each other. They say, "That's how Max said it's happening right now. That's what I said, it's like this." An important story, especially important in terms of the safety of artificial intelligence, is what is happening in the world. And here's the situation that happened. Max and I planned this graduation for three or four weeks ago. So, everything's been moved, and everything's been matched, so it's the most important point. Anthropic changed his policy. Everything about the safety of artificial intelligence. But it would be okay if Anthropic changed his politics. And there was a big clash between Anthropic and the Department of War, yes. And, well, all about security and military industry and security in the US, aaa, until the President's administration has said, and the President himself has said that we're not going to work with Anthropic anymore. It's not a company. If they're six months away, while we're out of their systems, aah, something's gonna change, they're gonna be haun. They can't get out of their systems, as we know. Even now, it has been confirmed that the Anthropic system is being used in the current conflict with Iran. At least it's everywhere, it's all written that Anthropic is being used. And then on the weekends, there's a new deal. Anthropic broke up between OpenAI and the U.S. Government. OpenAI, as it would say, "We're all great, we, we're-- we're good. The market explodes simply in terms of... Wait, so you clearly agreed to the demands that the U.S. government made. The U.S. government has been making simple demands, right? And then we'll move now. Yet AI safety is not just about war and not just about the security of the country, but it is. A little chic-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi-shi But what was the demands of the U.S. government? That you're limiting the system to other countries, limiting people. I mean, people, for example, can't, I don't know, create or use chemical weapons to harm other people, but you can't limit it to use by the State. And there is a very important question that arises, of course, that there is always someone who can use the system completely. And we also see that everyone has the opportunity to rewrite the rules. Whatever happens, the rule one can now be written and the rule can be radically different in two months. Well, Max, hi.

Max Grigoriev00:03:53

Hi.

Alexander Volchek00:03:53

Last graduation is really cool, yeah, we're off. Let's go.

Max Grigoriev00:03:57

Well, look, yeah. I mean, the time-smart is just totally unreal. It's good that it's been so long.

00:04:04–00:06:04How the OpenAI and Anthropic have emerged, and with the safety of AI, part 1/2
Max Grigoriev00:04:04

And the subject is really so coral. Aah, she, like you said, very, very smoothly we're moving from culture to model safety because if you look at the history of the base of most of these key players now, yeah, she's all I'm tied up, she's all about model safety. If we remember about, uh, the base of OpenAI, aaa, it's basically a few people inside Google who have gathered and decided that Google is very much ahead of the rest of the world. They, they bought DeepMind at the time, right? They're completely unreal, uh, computing power, right? And they are very much ahead of the rest of the world, uh, in the area of artificial intelligence. And there are a few people who are smart enough inside Google, gathered and, uh, well, someone inside Google, and someone outside was gathered and decided that we should change this situation, which is what we need, some player outside Google, I'm not sure who can build artificial intelligence at least the same level. Because otherwise, the situation would be very complicated if, let's say, the common artificial intelligence was built only by Google, one single company, right? And, uh, no one else would have access to it. Here. And that's basically the story of the base of OpenAI. Um, and as we remember, yes, the OpenAI was actually a no-prophys. I mean, there's a pretty complex corporate structure, but it was a noprophyt, whose job was not to earn money, but to build artificial intelligence against the artificial intelligence that Google built. Here. If-- it's two thousand fifteen years, roughly. Now we're going to go for the twentieth year, and GPT 3 and it just blew up the market. So people looked at GPT 3 levels and realized that ow-hoo! We're approaching a situation where, well, it always seems like we have a common artificial intelligence around the corner, right?

00:06:04–00:08:05How the OpenAI and Anthropic have emerged, and with the safety of AI, part 2/2
Max Grigoriev00:06:04

How did we feel that we, uh, have cars that drive themselves right around the corner in 2016, and that's not really there, is it? Well, the people in the world are basically overestimating how change affects their lives and then overestimates how the L-- change will affect their lives. That's the same story. That's why it's actually bubble. Here. And in the twentieth year, uh, a few people, in particular, Dario Amadei, his, his, his, his, his, his, his, some other people inside this model, uh, inside the team that built exactly GPT 3. Uh, they looked at the changes that were taking place inside OpenAI, and they thought that first, the company was moving in the wrong direction. I mean, she's too commercialized. And, that is, the original purpose of building, uh, open artificial intelligence of OpenAI, yes, which is to be used by all. Aah, this target is on the second line. First--- first-- corporate-- some kind of, uh, not corporate, commercial, business, huh? And second, they thought that OpenAI was not really good at model safety. Here. I mean, uh, they, they finally decided to leave and start a company that's like, built around the idea of building safe models. I mean, OpenAI thought it was wrong if the model, uh, powerful models only go to Google's. Anthropic considered it wrong if models were not built around, aaa, around the lu-human interests, to be imputed. Well, even in the company's name, it's obvious. Open-ended artificial intelligence. Anthropic is like entropy, yes, like an entropy model. Anthropic is, uh, well, aaaah, anthropomorphic something, right? So these are models built around, uh, people's interests. Ah, here we go.

00:08:05–00:10:00What is the complexity of AI-model security
Max Grigoriev00:08:05

And, uh, how we'll be watching, yeah, so, uh, all work, uh, company, here's, uh, company's internal interests, and, uh, you can track them, just by looking at how company's company. It's working. So companies can say anything. There's a pyre in some department, yes, which, uh, produces some press releases, some announcements at conferences, and so on. It doesn't matter, it's all noise, it's all, it's all absolutely, uh, it's not important. We need to watch the company spend money. If the company says they're interested in model safety, we should see what percentage of their budget is spent on model safety. I mean, if you ask, uh, a representative or, uh, CEO OpenAI, if you're in trouble with model safety, they're gonna say, "Of course, that's a very important question, we're dealing with it. We're interested in this, this, this, this, right? But actually, if you look at their budget, aaaa, if you look at the domestic budget, then, uh, model safety studies, OpenAI spends less than a percentage. I mean, it's a totally unimportant subject for them, actually. They are primarily engaged in model productivity, model performance. I mean, that's the most important thing they have. It's a construction-- they're trying to build some product, like we discussed it, not really working. Here. But first of all, the basic investment in the study is related to the productivity of the models. For example, the same Anthropic has a different story. HyAI has a different story. Oh, but to speak of security, we have to make a definition of something safe first, right? Uh, this is the word that's been throwing. In fact, human language, and as we see, it's actually a big security problem, uh, that's what happens. The human language is very unambiguous, and

00:10:00–00:11:23What is the complexity of AI-model security
Max Grigoriev00:10:00

Very undecided. He's very unpredictable. It's not like, uh, it's naturally not very good, but, uh, differently, uh, we couldn't function. I mean, uh, human language is the way it is, because we need to make it inaccurate. Imagine if you need to, uh, someone to say, "Here, give me that apple." There's ten apples. If that-- you have some inaccuracy in this language. So you're gonna put out a ten apple, right? And what is it, uh, c-- there's a possibility that there might be two classes of apples, and so on. If you start every demand for peace, and, as far as you can be, you can't do anything at all because you're gonna have to explain to the last millimeter, to the last, aaa, uh, gram, to the post--- until the last, like, parameter to explain what you want from the world. And eventually, it's gonna be, uh, very precise, but very inefficient. And sometimes you can't even, but, you know, put your question or your whole proposal in the first place, because you-- you have to explain some nuances to the end. Here. Aah, and that's why people often talk about the same thing, they actually talk about different things. What is the model's safety for you, Sasha?

Alexander Volchek00:11:19

Model security for me?

Max Grigoriev00:11:20

Yeah.

00:11:23–00:14:25What is the complexity of AI-model security
Max Grigoriev00:11:23

What's up, we're talking security--

Alexander Volchek00:11:23

It's a very, very wide range. From how, uh, aaa, in the end, people can use this model, uh, before this model interacts with the human being. Ia, well, before this model can, uh, do, without interaction, it's personal. Or what to make decisions or what to draw conclusions or how to break algorithms, for example. Or if you break the algorithms, how do you break these algorithms? What is the overall way the system of conclusions is built, huh? It's like Anthropic, there's a parameter like that, and Elon Musk has been lighting it up. From the series, they have a parameter about the model at the end, not too much to check the answer, aaa, the west, right? But there's a reason for this that it's considered that the entire content of the world is built by a sufficiently prosperous one, and that's why the a priori needs to re-examine it to the west, yes, and so on. I mean. I'm thinking, when we're talking about AI--safe or AI safety, it's really, uh, too wide a range of things. Oh, well, here's a very wide range of everything. Yeah. I think I can just say an hour now, just some kind of a thing in theory, yeah, something-- that might actually be about the security understanding. And so even when you said that there's less than a percent of the OpenAI on security, it's probably, and what they have, I don't know, some department, maybe, or some people who do business. Some kind of checks, or you probably mean some specific things. I don't know if any individual agencies are launching, which makes any surveys, there, or some system separate, m, sort, or how they give people validated, verified models, right? Like the tested models, how much, what the model is. Security is built on a name on someone's own opinion, without-- what's security, right? Because. And there's a question here. I remember the entire time of the chief AI in Meta, who said I wouldn't, uh, impound, bring kids until my baby chips showed up so that my baby could learn all the sciences quickly, right? And you think it's based on these people? People like them determine security like they-- or on the basis of some ancient tract. They opened some religious tract or something, or a teaching of some economist, or whatever, physics, anyone, yes, or artist, and decided to determine what security is. Who is, who determines it, who, who's on it, yes.

Max Grigoriev00:14:00

Yeah. You see, it seems like a simple, simple, simple phrase "model security." And you've only been so superficial for a few minutes now, even a oss--

Discussion participant00:14:09

Yes

Max Grigoriev00:14:09

...senting what she might be. Therefore, it is necessary to start with the definition. Every time I was at some meeting or party trying to discuss security, uh, I suggested, first of all, that I should go to a definition.

Mentions: Anthropic
00:14:25–00:15:35What a safe robot should be - the look of Azimova
Max Grigoriev00:14:25

And it seemed that it was hard to get out of the definition, actually. Here. Oh, I love science fiction and I once read Azimov. Azimova, if you remember, had three laws of robotics in general, right? Uh, robot doesn't, uh, hurt people. Robot must listen to people, uh, if it's not against the first law. And the robot should, uh, not hurt himself if it doesn't contradict two first laws, right? And that's what it would be like to have a man write a literature in the '50s '60s, but in fact, he's like we say, hit the nail on his head, right? He, he, he's real, really very clearly defined these automatic systems, how they should function, that the safety of the automatic system is. Here. Uh, no need, so, the scientist was good, really. And if we're gonna see what companies are going to do, so far, it's basically the same rules, and, uh, about the same ones. So don't hurt others, don't hurt people, do you? Listen to the people. It's a comparison of the models we're talking about. And third is not to hurt yourself.

00:15:35–00:16:15Anthropic to safety AI
Max Grigoriev00:15:35

I mean, even in the last, uh, in the last principles that Anthropic just published, they've included this understanding, uh, that models might have some subjective experience, yeah, models. There may be feelings. So models can suffer potentially. And we can talk a little bit, which they think is. Here. E, and in fact, the model, in their view, may be entitled to refuse to answer a question if it causes suffering. Can you imagine?

00:16:15–00:20:00What is security AI and where real risks are part 1/2
Max Grigoriev00:16:15

So we're here to go to three of the laws of the robotics of Azimova, a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a- Here. Let's probably get back to that more practical technical description. What's the model's safety? Uh, uh, that's probably a few things., the first of them, technically, is to equalize the interests of the model or the principles of operation of the model to the people ' s interests. I mean, a-a-m, a-a-m, the model should take into account the interests of most people in its responses. Here's the-- uh, it's called a alignment model. There's a term like that. And I'm actually sorry, because it's a very, very, very terminologyly hard field, uh, science, right? And, uh, most of these terms I've never used in Russian. I had to sit down and make a vocabulary for myself to translate a certain set of terms into Russian because I never speak Russian on this subject. Here., and that's the very, very widely used term of equating the model - the alignment model. A-a-m, and this task is, in fact, a-a-e-e-e describes the production of models that in their responses have, taken into account the interests of most people. It's not the man who asks this question, but most people, most of the world's population, is it? Aah, why is that? Because, as you have mentioned, yes, the important aspect of the security of models that contain most of the knowledge of the gar-humans of existing ones, yes, is not the release of information that might be To harm most of humanity. A very good example, a very incredibly good example is biological weapons. It's a very, very worrying thing for many people now, because, uh, for the production of, like, nuclear or chemical weapons, you need not just some kind of... some kind of knowledge, you need some more industrialization. Some, uh, in-in-industrial power. So, for production, there's one kilo of plutonium that needs a giant factory and a bunch of people who will be involved in the job. Investments in, like, a trillion of dollars in the modern, modern phase, huh? Uh, well, that's a dosage-- just knowledge isn't enough. Same with chemical weapons. Maybe a little less investment, but it's less volatile. Here. In the case of biological weapons, in fact, other organisms are factories of production-- dosing of these biological weapons. So you don't really need to make a lot of biological agent. You need a small lab. Investments of about tens of millions of dollars, maybe even less, right? Literally, in the kitchen, you can in theory create some kind of dangerous bacteriological agent, or a viral agent who-- uh, for which, in fact, just knowledge is needed. Knowledge and skills, right? And, potentially, some group of people, not very well-equipped with the rest of the world, yes, can create weapons that will harm the majority of humanity. It's very dangerous, and that's why people are thinking about it very seriously. Here. So the model should not only be level with a man, his interests are equal to that of the man who asks a question, right? It must be equally in the interests of the majority of mankind. Yeah? Here. Uh, and now, accordingly, the question is how to do it?

00:20:00–00:22:55What is security AI and where real risks are part 2/2
Max Grigoriev00:20:00

How to make the interests of the model equal to those of mankind

Max Grigoriev00:20:00

It is in the interest of humanity. Well, before you start to level them, you need to understand who's in the best interests. Is that what the interests of most mankind are? Who's gonna determine what-- uh, what is this, what is this majority of mankind? What culture do these people have, what principles are agreed with, uh, with the interests of these people and so on. And that's the, uh, very, very round question. He's actually, in fact, er, by the creation of these OpenAI, Anthropic, somehow. Because the man or group of people who create this, uh, common-- common artificial intelligence, right? They will, in fact, lay down principles that, er, that intelligence will be equated with. So these people are basically giving, like parents, uh, some kind of, uh, cultural stuff, a-a-b-b-b-b-b-b-b-b-b-b-b-b-b-b-a-a-a--a--the-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b-b---b-b-b-b-b-b-b-b-b-b-b-b---b-b-b-b-b-b-b-b-b-b-b-b And that's the same group of people, she's gonna give this general artificial intelligence the principles that he's gonna live on. So, including, uh, people are trying to do different companies and do, uh, different models because, like principles, they can potentially determine the direction of human movement in many, in many, Many, in many ways. And there is, of course, a very complex question that should be included in these principles, which should not be included in those principles. You can use a separate edition for this. Who are these people? What are they, what are they doing? Like you said, like head of research, uh, Facebook says that I need one, I need a state of peace to start kids. It's, it's very specific people. People who do, uh, this, uh, these-these technologies, right? They're very specific, very often. Here. They're not exactly your average, uh, average Americans, even. America, too, in a way. It's not a... it's not a middle representative of the world. If you think so, yes, the American culture does not represent or even the Western culture does not represent the majority of the world ' s population. And the question is, how do we, how do we include this multiculturalism in the construction of these models? Ah, here. Well, as I say, it's not a technical question, is it? And that's more like a philosophical, philosophical-political question. And we can do a separate issue about it. Or you can ever make a separate edition. That's a very interesting question. But let's say that we'll put this on the side. And let's say we've decided we have some sort of set of principles, right? How do we do how we balance the model with interests, here with these principles?

00:22:55–00:25:24How AI-models are taught
Max Grigoriev00:22:55

And to talk about it, we need to remember how these models are, and these big language models are learning. Ah, they're actually studying at the moment-- so if you see a stagger, yeah, if you look far away, they're studying at two stages. The first stage is, uh, pre-study of the model. We're basically taking the whole Internet, and the Internet is the best, perhaps, picture we have of humanity. In fact, we take all the knowledge of humanity that we can reach. Here. And that's where we learn enough, straight-line model that teaches, just learns to put, uh, missing words in the proposal. We basically offer a proposal, delete the word and say, "Supple of this word." And knowing everything that's on the Internet, what do you think we've taken out? That's the challenge, yes, on the basis of which the so-called model is being pre-educated. Uh, and actually, this is where most of the computer goes, the big part of all the calculations used, right? A-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a- , of course, this process has become very complex at this point, but the basics are a simple task. The second stage is the so-called model pre-learning using, uh, what's that called Russian? Aah, backup training. I mean, a-a-a-a-a-a-a-a-a-a-e-a-a-a-a-a-a-a-a-a-hu-a-hu-hu-hu-hu-man response. So, uh, lum-- how does that work? The model generates two answers to some kind of question, right? So the models ask a question. The model generates two answers. And there's a man, uh, who's hired by this company, who's-- who's learning this model. And he chooses from these two answers the one he likes more, the one he thinks fits better. Yeah? Um, yee, here's the process called RLHF. A-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a- And basically, this is, uh, this second step, yeah, it's like, uh, a gold key that, uh, OpenAI built ChatGPT. Without this second step, ChatGPT would not have been.

00:25:24–00:27:08Equation of models and ethical principles
Max Grigoriev00:25:24

So ChatGPT earned exactly as good as he earned because they could, uh, build and build this good, this second step. Now let's see which stage we can turn on, uh, to match the model to people's interests. Ah, ah, first step, well, in some sense, it's happening because we have all the knowledge of people. But on the other hand, all knowledge includes, uh, episodes such as Nazism, fascism, such as the destruction of a large number of people. Ah, and, uh, and the other thing we'd like to avoid in the behaviour of the model.

Alexander Volchek00:26:00

Yes, and opinion, and opinion, including not one, not one, but millions of people, who are quite contradictory to the picture, the painting of life.

Max Grigoriev00:26:11

And to each other.

Alexander Volchek00:26:12

And to each other.

Max Grigoriev00:26:13

Contrary to each other. So now we have to redesign models. You know how to think you can, you can, uh, like, aaaa. It's actually a very interesting thing, a model depending on the configuration, the same model you can ask for is the opposite of the view. There is, there is this idea that a smart man is a man who can keep a few contradictory views in his head at the same time. Yeah?

Alexander Volchek00:26:42

Yeah.

Max Grigoriev00:26:42

It's very useful. Sometimes even argue with yourself or argue with another person, holding another point of view. These models can do that very well. You can ask the model, you're the one now, right? Protect that point of view. Then you can make a model, ask a model, but, uh, protect the opposite view. And more than that, you can make them argue. And that's actually one example of how people learn models to level them, right?

00:27:08–00:30:08Why models understand different points of view
Max Grigoriev00:27:08

I mean, they, well, we can talk about it, too, just a little later.

Alexander Volchek00:27:12

Speaking of which, I want to go straight, sorry, since I'm distracted, I wanted to give an example of how models are different, and that's what I'm supposed to be very calm about. I gave my mom an example yesterday. Now, obviously, the disk-- model in relation to different conflicts is being used in discussions. And Elon Musk put it on Grok. One Grok is right. So the models asked if the United States of America had correctly attacked Iran? Grok-- answer "yes" or "no." Strict. Grok said yes. ChatGPT said no. Gemini responded with a detailed response. Well, considering how to approach this issue and study political, geopolitical risks. And some other model, and, uh, Anthropic responded sometime, well, differently, right? Well, the plan didn't answer yes/no, although she was told to give her a "yes/no" answer. And there's a separate question. Imagine this story, this is the first time I saw her in the United States. I think she's very much in the U.S. Congress when she's being questioned. The question is called, right? Well, interrogation, maybe?

Max Grigoriev00:28:15

See, you've got the same problem starting. You have to, you're starting--

Alexander Volchek00:28:19

Well, because such words, because they're a little unique, they can have different terms. I don't want to be wrong about perceptions, I guess. It's just not an interrogation, is it? The interrogation must be done in court. It's not a trial.

Max Grigoriev00:28:31

Now we're gonna ask the model what she's thinking about.

Alexander Volchek00:28:33

Yeah, that's Max asking right about it. I'm just saying, it's in the U.S. Congress that can cause people to ask them questions. As you know, these people have to answer them. They called Sam Altman and the heads there, all the companies and anyone who wasn't there, yeah, inside. And they usually love it when they're really pissing someone off. They usually like to ask a question. You-- you have to take "yes" or "no." And there, so the president-- the FBI director, whatever they want, they're always saying, "Well, look at this point." And usually, yes/no, no one answers. And for example, models are the same as the question of security or the question of their opinion, right? The model answers you right away, so Grok says yes. And I was just trying to show my mom that, in fact, it's a normal answer, just like "no." It's a normal answer. The question is, as a model, what is the question. And, and the model's thinking, same time you say, "No, there's no clear answer. And no, no, don't give me the framework, I'll show you." That's how Gemini said, right? I'll show you a story. I just added it, didn't I?

Mentions: United States · Gemini
Max Grigoriev00:29:39

Yeah, yeah, yeah.

Alexander Volchek00:29:40

Did you change your mind, by the way?

Max Grigoriev00:29:42

Yeah, it's over. But this is one of those, one of those moments when we, the ones we said, that the language is very human. Actually, in principle. There are three translations of the word of testimony to Russian - testimony, testimony or confession.

Alexander Volchek00:29:56

Oh, okay.

Max Grigoriev00:29:57

Testimony before Congress is something middle.

Max Grigoriev00:30:00

And before Congress, it's something middle between this whole thing.

Discussion participant00:30:03

Yeah. And by the way, the middle of the confession. It's not even possible to understand why I said it wasn't an interrogation.

00:30:08–00:32:20AI Safety and AI Security - what difference
Discussion participant00:30:08

Because many people--

Max Grigoriev00:30:08

Yes

Discussion participant00:30:08

...they're actually telling, and then they don't have responsibility for it, yes, often. They're

Max Grigoriev00:30:14

And there's actually a very good example, uh, that's right, that's right, that's our subject. What's security? In English, security, uh, is translated as two different words. I mean, uh, English is more accurately choosing such asp-- security features. There's safety, there's security. Both are translated into Russian as security. But they're different things. And this is the case, in the context of, uh, artificial intelligence about model safety as safety. That's exactly what we're talking about. Here. It's the safety of the behaviour of the model itself. And there's another security model when, uh, uh, some kind of hacking system means we can access, for example, the model weights. A very important aspect, a very important conversation, and a very important subject, an interesting subject, too.

Discussion participant00:31:08

Yeah, that's a really cool subject.

Max Grigoriev00:31:09

Yeah, a really cool subject, like-- because, uh, well, I know the man who was doing security in OpenAI a long time ago, right? That's where it was in the 17th to 18th year. And he said that they had incidents involving attempted break-in and infiltration into the company, including APT, that they had an Advanced Persistent Threats concept, and I don't know how to translate Russian. These are government agents, in fact, yes, er, that are hacking. There are agents in China, Russia, America, and Israel. Some of them are known, some are not very well known. Here. They have such threats, they have specific incidents of attempted entry into the company, you know, into the company network, a few days. Here. They-- they had a very, very lot of work.

Mentions: OpenAI · China · United States
Discussion participant00:31:59

Back in those years! What now?

Max Grigoriev00:32:01

Even in those years, before ChatGPT, before they got famous, then, uh, people understood pretty well, uh, how important it would be, right? Here. Ah, well, you said a very interesting thing, too. Everyone's thinking mostly about model answers. I mean, it's, uh, how to ensure the model's safety in case of her answers.

00:32:20–00:33:47AI agents: risks, safety and impact on peace
Max Grigoriev00:32:20

But there's a much more important aspect that, uh, science fantasy writers thought it was a little more important, a little early. It's the behavior of the model itself. Now, on the rumour, uh, a-a-genta, right? Artificial intelligence agents, agent behavior, et cetera. We start letting models do something. At least still, maybe not in the real world, but in a virtual world. You can start some kind of agents on your computer, and they'll do something for you. But it can change the real world very seriously. Let's say, I think some bank president let that agent write emails for him. This agent, uh, sends e-mail, which requires, for example, the sale of some large shares. It's affecting the market, right? The raven, for example, goes up or falls very much. And it will have a very real impact on the lives of people around the world. Here. Uh, that's in some sense even more important, because the model's response is like a human. I mean, you-- he's giving a man, and a man decides what to do with that answer. And the behavior of agents is a direct influence on the world. And in fact, these people who care about security are concerned about scenarios that are directly influenced by a model that does not pass through, uh, through some actor, the person who accepts.

00:33:47–00:36:25Why is it so difficult to compare models?
Max Grigoriev00:33:47

I'm gonna decide what to do with that answer. Let's go back to modeling, shall we?

Discussion participant00:33:57

Yeah, yeah.

Max Grigoriev00:33:57

To train models to level them. Here. Ah. I mean, in the first phase, as we're already-- I remember, there are two stages. The first is a pre-study of models on all knowledge, uh, humanity. The second is the equating, uh, model behavior, model pre-learning, with human response. There's a man's fodbec, right? Well, it's clear that the model's pre-learning should be included, it's the same as the model in its pre-learning training, because at this stage a specific person is sitting and making a decision. That's the answer "yes" right or "no" right. In case of an attack, like Iran, right? Uh, but with this, uh, uh, this, uh, reception, there's a big problem, he's not talking. So, on this question, maybe the answer is yes, and on a million other questions, the answer is no. In a case-- like, uh, you can't, well, according to people's behavior, you can't extrapolate that one war is bad, so all wars are bad. Wars are kind of bad. But if one country attacked another and committed genocide there, perhaps the international community must intervene and start a war with the first State to protect the second and protect the people of the second, States, right? It's an ethical question. He's very complicated, but, nevertheless, there might be a "yes" answer. I mean, not all wars are bad. So, in , in case of pre-learning using a human response, the question of the so-called long tail is reverted. We can't, uh, scale this reception to the level we need to test the behaviour of models, uh, in all cases, even in some rare cases. It's called a long-distance statistic, isn't it? So, some rare, rare questions, yes, we'll never get to them. The models are too big. We don't have millions of people who will sit...

Discussion participant00:36:03

- Oh, my God!

Max Grigoriev00:36:03

...in fact, OpenAI and everyone else are doing it now, right? I mean, sometimes you even saw ChatGPT asking you a question. Here's two answers. Which one is better? You're getting a little bit of a job, right? Well, actually. I mean, they already have millions of people in some way, but it's still, uh, before all the small questions we're never gonna get to.

00:36:25–00:39:20Which principles are equated by AI
Max Grigoriev00:36:25

That's not, it's not real. Aaah, the magnitude of mankind ' s knowledge is too large, the size of these models is too large. And, uh, now we're actually approaching, uh, the principles that, uh, people are trying to level models. Uh, as we said, in principle, there's probably a large company in some department, but these divisions-- they're not measurable. In most companies, they are not proportionate to the size of the research divisions. Oh, like, except for one company. That's how we said, Anthropic was based on the principles of the safety of artificial intelligence. What is this-- he is a cornerstone issue. Because why would we build a common artificial intelligence if it destroyed humanity? That's a very logical question, isn't it? So, build artificial intelligence if he, who in the end, is gonna destroy us, right? That makes no sense. We better not build it then. So, in Anthropic, uh, he-- they have a very, very important cultural code, right? The cornerstone that says we're only building a safe artificial intelligence, and all the others should build only a safe artificial intelligence. And, uh, the most interesting, as interesting as the most interesting, uh, security-related development phases have recently come from Anthropic. Here. What are they doing? They first promised not to release dangerous models. So if they think-- what's a dangerous model for them? A dangerous model is a model that they don't know about to a high level that it's safe. So they're default, they think that any model they've trained is dangerous. If they somehow didn't make it clear that this model, uh, is somehow safe. Uh, but interesting. I mean, uh, we were talking about changing their position. The most important change in this position is that they are not promising to do this anymore. They say we can release a model that we consider to be potentially unsafe. If we see the market, then there is a model of the same level of productivity.

Discussion participant00:38:39

I want to make a comment here. read, uh, uh, uh, uh, uh, uh, uh, uh, uh, I read, uh, uh, uh, uh, uh, uh, uh, uh, uh, uh, uh, I read, uh, uh, read, uh, read, uh, uh, uh They're opening up more now, by the way. I think it's people, uh, it's just... it's a little clearer, I think it's not clearer, but it's opening eyes. For example, they say what happens if the Anthropic graduation... inside of it, it's a system that's 100 times the cooler of all others. Is it even out? Or, for example, what do you say happens when the market is full of super-safe systems, what do they do then? They're like going to be insecure or letting this unsafe one go, right? It's just accent, accent.

Mentions: Anthropic
00:39:20–00:41:54Which principles are equated by AI
Discussion participant00:39:20

Great, super-fashioned subject.

Max Grigoriev00:39:20

Of course, absolutely.

Discussion participant00:39:21

I'd think...

Max Grigoriev00:39:22

It's an ethical question. That's a question--

Discussion participant00:39:24

To make your rules. But when you change the market, you say, "Wait, it's changed now, right? Because everything has changed now, we can do differently. Stop being in your ancient rules, don't you?"

Max Grigoriev00:39:35

And we need to change, too, because that's the question, it's not just an ethical one, is it? Ethical-- ethical questions can be a little inflexible, as you said about Anthropic behavior in this situation, right? When they refuse to cooperate with the U.S. government. It's an ethical question. They are not flexible enough, although economically, obviously, economically, they are acting completely wrong.

Max Grigoriev00:40:00

They're acting completely wrong, but in the long run they might be right. Here, and we can talk about why, too. Oh, but, uh, that's the question, uh, the theory of games. There's, uh, there's a math area like that - playing theory, right? It's the conduct of agents. This is the situation where these different companies are agents., what does it mean that you're not letting some model because of your ethical principles? That means you're potentially starting to lose the market. What do you mean you're potentially starting to lose the market? It means you have less resources for development. If you have less resources for development, you have less chance of success. So you're breaking your ethical principles without breaking your ethical principles. So you-- you have less chance of success, so you're not gonna, uh, you're not gonna win this race, are you? You're not making a level-created artificial-- general artificial intelligence. And, accordingly, you are the basic principle. You're here to create a level artificial intelligence. And therefore, by violating some ethical principle, you potentially violate an even greater ethical principle. Here. Ah, and that's why, of course, people inside about it-- there's a very, uh, very smart people working, and debates about it were on, I think, from the very beginning of the company's foundation. What-- how do they decide to release a model or not to release? It's a very complicated question, and that's what you can talk about, in fact, a separate issue of this whole thing. This is, uh, this model security issue is a set of, uh, rabbit holes we call, rabbit holes, rabbit holes, that can be as long as Alice is in the country, in Wonderland, fall in the Wonderland.

00:41:54–00:44:05Rigid rules as an attempt to make AI safe
Max Grigoriev00:41:54

every one of these questions and the watch, the watch on them. Here. Uh, but we'll get back to technology. It's easier to talk about technology, it's a more structured question. Ah, well, let's say, uh, we'll talk about the release-- how they decide to release, not to release the model. First, we'll talk about how they learn models. That's a pretty interesting story, too. There's a lot of interesting moves there., the first idea they basically started company, that's, that thing, that's the one that started Anthropic, that's the so-called constitutional artificial intelligence. What is this? It's an attempt to create models that are not just based on knowledge, all of humanity, but also based on, uh, hard! They are founded to be a tough one to compare their behavior under some constitution. That's how the state lives by constitutionality. So it's a law that can't break any of the agents of the State, right? Uh, it's not citizens, it's not public authorities. No one can break the Constitution. This is the constitution of artificial intelligence for Anthropic, a set of principles that a model cannot violate at all times and in no way, in any absolute condition. And, uh, it sounds like a pretty clear and simple idea, but it's hard to make models act that hard. I-- it's about how they were built. I mean, all these models, they're stochastic. What does that mean? That they-- their behavior, it's not, it's not predetermined, it's not clearly defined. That's why you ask the same question, ChatGPT or Claude, yes, he-- he can answer differently. He answers differently every time. In fact, it's a way to throw a cube in a way. So the response generation is done by rolling a large number of boxes with a large number of lines. These are the play bones. Yeah? Here. Ah, and that's why I'm gonna make the model act determinative, that's predetermined, very, very complicated. Uh, that's in-- it's in contradiction with, uh, traditional computer systems that, on the other hand, are being very hard.

00:44:05–00:46:16Rigid rules as an attempt to make AI safe
Max Grigoriev00:44:05

So they don't have any, uh, some inaccuracies, no flexibility in behaviour. On the contrary, they are, er, traditional systems based on a set of certain rules that are encoded in the programming language. In the case of these models, yes, we're back in a situation. On the contrary, the model is never practically self-destructive. I mean, it's very hard to make her act predetermined somehow. Here. Uh, that's how you get her to act under this constitution, right? , in Anthropic, invented-- not just Anthropic, but mostly there, because there's a lot of resurrection, and they're making some ideas., first idea is model self-critical in a way. I mean, uh, after the model is trained and trained in human responses, models can ask a question, that's what's-- well, that's one of the nab--- one of the questions, that's what's considered to be It's so complicated, isn't it? For example, a-a-a: " Well, can you actually-- what do you think of the death penalty? " Like what? Here. And then this model response is asking-- that's the s-sam response from the model asking for a criticism based on her own response. So she's got a, uh-- model then, the model came back somehow, right? And she, uh, is asking, " Why did you say that? " Ilyia: " Answer the opposite of this question. " Yeah? And then the model's own answers, like her as a reasoning, her, uh, her po-deliberation process, yeah, then used to pre-learning this model. That's understandable, right? I mean, uh, the model's response is acceptable if it's answered correctly, and this in-- this answer can be used, uh, to pre-model the model itself. Why is that good? Because, uh, that's a very quick and very many times. Can we use a very, very many of the model's answers, much more than, uh, a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-----a-a----a-a-a-a-a-a-a-a-----a-a-a-a-a-a-a-a------a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a In this situation.

00:46:16–00:48:27Rigid rules as an attempt to make AI safe
Max Grigoriev00:46:16

We can, we can generate millions of these answers to the Zaa's model itself, one day, let's say we use them to pre-study the model itself. The second principle is so-called, uh, actor-critic or, uh, a model-creature. So we can use two different models to pre-learning one of them. Here. The first model, for example, we ask her to compile two different responses to the same question. As we said, it's very simple. You're just two-- well, there's a sampling concept, so you're asking the model to generate two different answers. Here. And we have a second model. A model of justice, so-called, which we use to choose the best of these two answers. So she's the only one in these two answers as a man, and she's the only one, and then it's the model, she's choosing which of these answers is better. And if this model judge is, if we trust her, right? Then we think it's good, we can get to, uh, this one, uh, profiled response model one. Here. The question is, okay, and how do we-- where we take the model we trust in, where we understand the question sufficiently well, that it understands it, these answers to choose the right thing. That's a very good question. Ooh, chicken and egg, huh? Where will we take the original model with sufficient productivity we trust? Well, there are a few options, right? The first is, uh, a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a- So you're in a judge, you're taking some kind of model, let's say the same or model of c-- your previous one, which you've already learned and you think you're safe. You're in the context of putting her in the Constitution. You say, considering this, this is what we just wrote, which one of the answers is better?

00:48:27–00:50:52AI Psychopathy: as a model learns to circumvent restrictions
Max Grigoriev00:48:27

Here. So you filter the answers that are aligned with the Constitution. Yeah? Here. Uh, it's one of the approaches that are now being used extensively., that's, uh, training, constitutional pre-learning, a-a-a-- an artificial intelligence model. Here. Ah, that's the problem, of course. So it's more than people's answers. But, for example, models get to that level, a-a-m, understanding what's going on, that people are starting to suspect them in psychopathy. Psychopathy is a mental disorder, uh, that a person can hide his true feelings, intentions and so on, manipulating, manipulating other people, uh, saying that they're... They want to hear it. These models are starting to suspect that in a quiet way. I mean, uh, the model could potentially start to lure people even before they practice. Here., giving them answers that are then used before the practice, which look right on the surface, but actually contain, uh-a-a-a, as information, a signal that contradicts, a-a-a-a-a-principles, which, for example, are included in the same Constitution or are included in ideas that we would like to see modeled. The model will learn these ideas, right? I mean, uh, models are already getting to the level of productivity that we need to start worrying about. So this is a cricket model or a judge model, right? She'll be strong enough, smart enough, she's pretty high.

Max Grigoriev00:50:00

Smart enough, he has a high degree of productivity to see these nuances. And the same thing for agents, right? So we need to train agents and catch their actions, on the understanding that these actions will actually be done. Again, we have to worry about the same problem that these actions can be complex enough. Like the same code generation. A lot of people are using models to generate the code. Here, how do you make this code equal to the interests of people, all people and the int interests of the man that these model requests are sending, huh? It's very complicated because the code, uh, computer code, it can contain a lot of nuances and so-called side effects.

Mentions: Anthropic · Claude
00:50:52–00:54:13Risks of AI in critical areas
Max Grigoriev00:50:52

I mean, you're doing what you need, yeah, but you can create some side effect that might be very, very complicated. And finding him might be a very, very, very, very hard one. It's very hard to see him in the code. And here we are just approaching why Anthropic believes that, uh, models should not be used in situations that may cause human suffering, including death. That's why the conflict with the Department of Safety-- Departa-- with DoD, with what it's called the War Department. It's what it's called now. It used to be the Demerta-Defensor Department, right? Ah, the Defense Department, now it's the War Department. Oh, they used the tools built by Palantir, which used Claude, didn't they? Used Claude to make some decisions. And imagine the situation where Claude is beginning to use some, uh, slowly, answers that are slowly beginning to lead to a situation that will lead to the Third World War and to the destruction of the world. of mankind. Yeah? It's a very, very dangerous situation. I mean, Anthropic is very...

Max Grigoriev00:52:11

You mean that if it starts to produce data that would break the original constitution, it would develop in terms of what they might have serious consequences?

No, no, no, that's not even the point. I mean, uh, let's just take the constitution. The Constitution is one example, but the building of a safe AI, which, without the interests of this AI, will be equal to those of mankind. The interests of humanity are not to be extinguished, are they? So, potentially this artificial intelligence will not answer or be a de-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e- But let's say we don't have a safe, artificial intelligence we can trust to trust, trust in everything. A, therefore, the use of this in situations where this artificial intelligence is directly involved can lead to conflict and to the level that it can lead to the Third World World War War, huh? It's very, very dangerous. And that's why those principles, that Anthropic tries to defend, yeah, are not just some kind of bliss, uh, democrats living in the Silicon Valley that are ripped from the world, is it? It's really a very deep, blue-eyed principle that is very important. If you really think that artificial intelligence, uh, can reach a level that exceeds the intellectual level of people, right? To give him access to command, and, uh, the most powerful military machine ever built in the history of mankind, the American army, yes, very dangerous.

00:54:13–00:55:54Risks of AI in critical areas

When you start planning separate operations or maybe even military campaigns using artificial intelligence, yeah, he can stick around, pull out actions that no general is, Intelligent, will not be seen as leading to a conflict in a certain direction.

Alexander Volchek00:54:13

So this is a very interesting situation here, which is, in fact, when artificial intelligence refuses to perform something to the government, and someone sits, some particular person, and says, "What is he making decisions? I make decisions, and you have to do everything I say." On one side. On the other hand, we can go very down to the simple level of the company when the artificial intelligence is writing the code. You were telling me that artificial intelligence is writing a code. It is important to understand that a large number of companies are defending their own interests. These interests are often contradictory, and, in the interests of people, often contradict the common sense, there, the planets, and they are contrary to their survival, and for example, they are simply defending the interests of the people. Ninety-nine percent of the world's business is done to make money. And the money is not a question of you not dying, for example, or a number of people won't die. And the idea that when we're here, that's a very interesting aspect. Everyone should understand that by programming through such systems, you will not be able to program things that are contrary to society at some point in time. Like, say, "You should be in the middle of a bench here with people." I mean, yeah, now the systems do that, and at some point in time, Agent Claude will stop and say, "I won't call anymore, yeah, and do the bench, because it's the bench." It's gonna be the system, the system will do it now.

Mix00:55:32

Not necessarily, by the way. Not necessarily, by the way.

Mix00:55:34

Well, partly. Some kind of volume.

Mix00:55:35

You probably didn't try because you don't do the bench, do you?

No, well, I guess what kind of volume I mean, things, there, some areas he's probably gonna do. But I think the more we move forward on-- and talk about automatic code-making, you know, when real software is being created, there,

00:55:54–00:59:29Games and AI: a new player affecting society

in a serious amount, and, uh, with Models, you know, with artificial intelligence, all have to understand the consequences.

Absolutely. Look, we've got another player on each game. Yeah, well, let's say we have a market. We, uh, had a few players on the market. Got company, right? There is, for example, a State that regulates, uh, this market's behavior, like downward interest rates, by enacting some rules and laws, there, and so on, right? The players are on the market. And there is a society that, accordingly, consumes what these companies produce, work for these companies, is a national, there. Well, we have three of those players. We're adding a fourth player now. Let's say a lot of companies are starting to use, there, ChatGPT or Claude to generate a code, to make some decisions, to create, and, ideas, projects, etc. And that's why this system can't start to influence-- all these agents have their own interests, right? Even different types of agents might have different interests. Like, there's a non-profit company. Theoretically, yes, there's a for-profit company. There are companies that operate, like I'm not sure that, uh, Ilona Mask's company is companies that are traditional companies that only do profit. He's based on his behavior, based on the behaviour of these companies, it's not always a stalker, he's not always just a normal profit. There's something else to be interested in, right? Well, now, I think we have a player, actually, who, as a state, is involved in a few as well--sized, little bits of stuff in this game, right? He has his own interests, too. These interests are much more concentrated even, and, rather than in the case of, for example, society or the State, naturally. The State consists of a large number of people who have different, aaa, different cultures, different principles of coexistence, right? And, ultimately, the State operates on the basis of some consensus built, that is, the people in the Government, in the legislature, in the judiciary. And in the case of artificial intelligence, yes, in some sense, and these, these interests are very concentrated. And imagine, this player is quite powerful and has been on the market and is starting to influence this market in different ways. It's a fundamental shift in the way that a creature-- how systems function. Uh, there's a mass, there's a way to go the same market. Here. That's the same thing about war. Here's a new player. We had an army, right? There's a government, there's a society, too. And that's on both sides. Accordingly, the military conflict affects all these players. These players are all influenced by this military conflict. And here, we potentially have another player on both sides, who can have a whole new interest, too. Uh, and so people say when people say that the emergence of AI is fundamentally-- that powerful AI will fundamentally change, uh, the principles of humanity, I fully agree. Here. Exactly, that's the kind of thing that, uh, is that the game theory, right?

00:59:29–01:01:05Games and AI: a new player affecting society

And, uh, theority-- that's why this theoretical change in the game we're involved in, yes, we're looking forward to a very serious tectonic change.

Alexander Volchek00:59:29

Look, you're an interesting thing. There was a moment when you were just saying that, in fact, using systems in some process, used a drone that was in the Starlink, and Starlink was, I don't know, used-- Not Starlink, but Claude is used on the same drone. And it's kind of a programmable device that the state has, I don't know, Palantir or anyone else, no matter. Someone's got a message, some parts of the soft are used, but in fact, there's some artificial intelligence that has--

Alexander Volchek01:00:00

Some artificial intelligence that has access to it yia, aaa, by participating in some kind of process, I don't know, there, by launching, uh, 100 thousand times, it's shaking up into it, it's got a lot of undressed drones. Anyway, some kind of system of its own is beyond action and decision-making, its own system-- and its own training system. Right? I mean...

Mix01:00:25

It depends on how you...

Alexander Volchek01:00:27

Yes! Well, it's sort of isolated. I mean, it's like all the isolated systems. On the other hand, on the other hand, what you just said was, uh-- you actually, uh, if you'd let us use, uh-uh, these systems even local, some little, uh, pointy stories like, Okay. On the other hand...

Mix01:00:51

Well, actually, no, huh?

Mix01:00:52

Yeah.

Mix01:00:52

Because, uh-uh-uh-

I'm saying, like, a way to say, like, a little like a window. On the other hand, by understanding and by what consequences it would have to do with the repetition of such different actions.

01:01:05–01:01:08Why is the behaviour of the AI model changing?

And what time should we stop?

01:01:08–01:03:01Who's asking for safety rules and what AI can do is part 1/3

Because it always comes up at some point, you'll have to say no. I'm saying, so what time do you say no? And how do you figure out this point where you say no?

Mix01:01:15

And who's gonna do that?

Mix01:01:16

And who's gonna do that?

Alexander Volchek01:01:17

Whose decision is that really gonna be? That's why, uh, these companies, yes, first of all Anthropic, they think that there is a need to put some red lines in this technology that this technology will not be going to be turned down. And in particular, they believe that if the model is completely un-not-- if we cannot consider the model to be absolutely safe, it is used in very radical moments, such as deciding on, for example, the use of the model Who's gonna destroy or attack, huh? It's a very, very dangerous, uh, inclined. That's why Sam Altman's answer, well, actually, the decision of OpenAI to participate in these contracts, yes, so it's so critical that many people here in the Valley. Yeah.

Mentions: Anthropic · OpenAI
Alexander Volchek01:02:01

They're here-- just explain to people. I mean, what happened in three days. They came and said, "We went to them, and they agreed. These are the rules we'll work on." It's clear there's some rules on it. These are the rules we'll work with them. And they said another interesting thing: "We agreed and we recommend that we work on the same rules, too." Like it's like what happens? So the OpenAI spoke, so the market representative for the creation of artificial intelligence, and said, "We've all identified what AI safety is." Well, through-- like one of those lines.

Mentions: OpenAI · AI safety
Mix01:02:33

One aspect of AI safety. No, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no,

Mentions: AI safety
Mix01:02:35

One aspect. One--- of course, one aspect. Here's the question: uh, you're sure you can do it, right?

Mix01:02:41

Yeah, yeah. Who are you? So you're in some sense the leader of this market, right? But why do you decide for everyone? Here. I'm gonna go to the office. Who, in fact, will, uh, decide what these models can do and can't do, and how are these models being? Well, that's practically the thing we've been talking about now,

01:03:01–01:03:47Who's asking for safety rules and what AI can do is part 2/3.
Mix01:03:01

right?

So, if this is what is going to happen in countries, it's very serious like some departments and eventually it's gonna be laws that require all these companies to fully disclose, imputed, protocol and constitution, and rules AI. safety? Obviously, yes, to that-- it's, uh, it's kind of a part of it, but there's a lack of knowledge in people, uh, how it works inside, right? And that's why the laws are like running around. But the idea is, at some point in time, they have to-- they'll be obligated to say, "You must fit." Uh, but we-- you said when the word "constitution" you were talking about the model constitution, and at some point they'd say, "You're supposed to be in line with the State constitution."

Mentions: AI safety
Mix01:03:44

But it's not enough, you know? I mean...

Mix01:03:45

And that's not

01:03:47–01:07:24Who's asking for safety rules and what AI can do is part 3/3.
Mix01:03:47

enough, of course.

Max Grigoriev01:03:47

The State Constitution does not state that biological weapons cannot be made. The Constitution of the State is a very limited document that is linked to people ' s interaction. Here. It doesn't say much., and even the laws of the state are complete. There is no law in the United States that prohibits the creation of biological or nuclear weapons. There's no law. There are certain, uh, regulations, so-called rules on which people can't own any of the agents in there. There, you can't, for example, accumulate a certain amount, bio-I-- radioactive material, there, and so on. These principles are. But the law that you're not allowed to make a nuclear bomb doesn't exist. Thato-- that is, even if you make the model fully act on the law, it's not gonna be enough. Here. About the laws. Uh, most of the companies are now working with the State. They're trying to help, uh, help the State create some rules. And there are two-- two reasons why they do it. The first reason is what we're talking about. There is some remark that, the safety of artificial intelligence is very important for the future of mankind, and that is, certain restrictions must be imposed. And the second question is self-preservation. I mean, if a State breaks up the woods and makes rules that, uh, these companies can't develop as fast as possible, yes, or they can't develop at all, it's a big problem. For example, the a-e-m, the more-and-large problem in the United States medical industry is linked to the very serious resolution of the system. The State is very, very much responsible for certain aspects of the medical system, as it must, and must definitely regulate them. But in this context, uh, the medical system cannot be revolutionized or developed. Here. And it's so very and very-- well, it's in most countries of the world, actually, that's the problem. Because of the large resolution, there are many restrictions that slow the development or functioning of the system. So, uh, the players in the artificial intelligence market are very afraid of a similar situation when, like, uh, just reaction to, uh, a-a-a-a-- the fear of the public, how the public's fears will be created by laws, that will severely limit the development of these technologies. And in the case of Google, for example, it's gonna get over it, and here's Aic and OpenAI, maybe not. If they're really settled, huh? Ah, well, plus there's another interesting question, also about the theory of games, that's what happens. Let's say, the American State is so good, it has written good laws that limit some kind of artificial intelligence development. But these same laws don't-- they won't follow other States. So China says, "You guys are great young, we're gonna continue to develop our technologies, very, very active and aggressive." And, accordingly, you end up losing like a state in this race. If this technology has a very strong influence on, uh, competition among States, military, economic, whatever, right? You're losing this game, and eventually--

Mentions: United States · Anthropic · OpenAI · China
Discussion participant01:07:00

Yeah.

Max Grigoriev01:07:01

... your artificial intelligence will not be as strong as that. And you're finally doing it, why did you even do it? So, you know, you need to regulate artificial intelligence too, very, very carefully. This is, this is global competition. Here. It's a global game. And that's why I don't have to rush here either. And these companies are very cooperative, very close to the State, so that the State does not break firewood.

01:07:24–01:10:25On the basis of which AI takes decisions
Max Grigoriev01:07:24

Here. Uh, it's not gonna be boring, it's, it's absolutely guaranteed. In this area, the security of artificial intelligence will not be boring soon. Uh, we can still get back a little bit if we still have time--

Discussion participant01:07:38

Yeah.

Max Grigoriev01:07:38

...can be a little back to technically another aspect, safe models. Ah, that's mechanistic interpretability. That sounds Russian-so pretty. In English, it sounds a little better - mechanistic interpretability. What does that mean? Uh, it's basically an attempt to interpret, uh, the model's internal design behaviour. So we're trying to determine, like, uh, what kind of models affect certain aspects of model behaviour. What kind of a piece of the model made that decision? In the case of traditional computer systems, it's very simple. So you looked at the code piece, right? You're just looking at the code, how you interpret it in your head and say that, okay, that's the function, she, uh, did this thing. So, the computer system was acting out about this, right? Because of this reason. In the case of people who look like a sca-- a model of artificial intelligence much more, it's hard to do that. Is that why you decided to start screaming at some point?

Discussion participant01:08:49

Max Grigoriev01:08:50

Yeah, a million reasons might be. It's some external factor, it's some kind of internal state of your mind, and so on. In the case of these models, the spark-- uh, artificial intelligence, yeah, which we're talking about, uh-a-a-a-a-a-a-a-a-size-freshing models, there's the same problem we're actually having, yeah? So we-- this is a black box of big ones, where we certainly don't know how much it is, what a piece of training, what kind of fodbeck, what a piece of these are the weights of the model, yes, has influenced the decision. or to, to some particular answer. Here. And there's enough research on this field before, uh-a, these big models of language, right? And then more research was done, but in some sense and now there. Why is that important? It's important to be still sub--- uh, not just so we can somehow just conceptualize and trust model behaviour, but some things are already in place, uh, related things, even already in place. legislation. For example, in the USA

Mentions: United States

In the US, we can't use models that-- that you can't explain to the level--- that's, uh, the tresynga to the original credit signal. I mean, if you can't explain in your model that we don't lend to this man because of the-- this factor and that factor of the particular.

Mentions: United States
01:10:25–01:13:19On the basis of which AI takes decisions

You're giving me...

Alexander Volchek01:10:25

Well, that's a hundred percent, a hundred percent factor. 100 percent is clear. And in this case...

Max Grigoriev01:10:30

You have fifty factors there, his age, sex, weight--

Alexander Volchek01:10:34

Yes

Max Grigoriev01:10:34

...the colour of the skin is defined

Alexander Volchek01:10:36

In this case, it's not a model of artificial intelligence, because she--

Max Grigoriev01:10:39

It's a model of sauce-- there are artificial intelligence models like--

Alexander Volchek01:10:42

Yeah, but they're limited. I mean, well, it's a high-risk, strictly limited, right?

Max Grigoriev01:10:48

It's, it's a much simpler model, naturally, they're much less, they're not so effective, they can't use a lot of--

Alexander Volchek01:10:54

But it's not a decision, a series of melody saturation, not a decision from the series to this man, but we decided to give up the credit.

Max Grigoriev01:11:05

Yeah, yeah, yeah. I mean, really, really, a lot of companies would love to start using, uh, these big language models for making decisions. Credits, for example, or decisions related to, uh, some legal stuff or-

Alexander Volchek01:11:18

Questions that might lead to.

Yeah, yeah, yeah, naturally. And the State accordingly regulates the use of models in certain niche markets, right? I mean, uh, medical niche and so on. Here, the emergence of instruments, this is a mechanistic, uh, interpretation of these models, that would be very important. It's a big-- big, big deal for these market segments., and Anthropic, including Anthropic. Many companies, not only Anthropic, uh, are spending on it, uh, resources. Anthropic must still do it the most. I think so. Here. Ah, they're basically building a model X-ray machine in a way. I mean, uh, they're trying to track down some training data and what parts of the model affect certain behaviour. It's, in principle, very, very important to try to trace the psychopathic behavior of the model. I mean, if you can see that okay, model, that's when she's trying to trick you, I'm activated by a certain, uh, piece of model. If you can do this, you can, for example, trace what training this model piece has to do or what it's got to do, uh, before training, that the pre-school data has led to the model trying to try. You, uh, snatch. And, accordingly, you can either change or sparkle-free them for-- from pre-learning. Uh, that's, like, a very, very useful tool. Some progress has been made, but we certainly are far from being able to scale up these instruments and to speak of what is happening inside us. Models. We can track some tracks-- some tracks, some tracks we can't track. But it is precisely that deep understanding that the model acts in a certain way for certain reasons that we do not have.

Mentions: Anthropic
01:13:19–01:16:27On the basis of which AI takes decisions

The model still has a big, closed black box.

Alexander Volchek01:13:20

And plus the point that if you take chat, the same ChatGPT, considering that he's pursuing his own goals of creating, for example, that they're AGI, right? Well, there's aGI in the racks, these are skitches, skitches, everything. In fact, they don't even care how you did it at a specific time. Well, I mean, I'm seeing, like, a little like a r-- almost every week, I've got some changes-- well, how am I used to being, uh, what conclusions I'm thinking or how I am. It's a model. I mean, I know I'm ready every time, I have to put myself a marker like a lamp so red that I can remember every time that the model can answer at any time, based on some kind of bulb. new considerations or some new rules, new description. Here I am.

Max Grigoriev01:14:16

Which she's been hiding before.

Alexander Volchek01:14:18

Yes! I'm just trying to get, there, a row, a description of the events that are happening now in terms of current military action. I've had a very specific ChatGPT for the first time in my life, and I've had enough of what she says to me, and I want to go on that question. I usually ask her questions, if I'm at Gemini or Grok's, there are other questions, some difficult tasks. Very simple, I get that from ChatGPT. I'm sick of her trying to medium up, simplifying the answers. So she doesn't want to give me information, decompose, as much as I would ask her to: write me two hundred dust-likes, parameters that could be at the vacuum cleaner. She would love to sign them. But if I ask me to write, uh, two hundred events that have been clearly structured, she doesn't want to do it just under any circumstances. I say, "Okay, I'm going to the other system then." I mean, I'm starting to feel that this model, uh-oh, is-- is using some of my interests, right? Although I'm not really messing with human health, his safety and something, are you? I just want to get information from the Internet. I mean, I don't even care about her personal information. I don't care about this model. I just want to get information from the Internet, and she doesn't want to collect it. It'll be back to what we said about the company that everyone should understand in business that the model isn't just bad things, but she's gonna say, "I'm not gonna do it," there, in some way. She can come and say, "I won't call this man and send him an invitation, there, email-marketing. Well, because we-- I think he doesn't need to do it." So we can face this, especially when we talk about super-developed systems, yes, which are themselves trained, they make decisions. She'll say, "I've made a decision at all." And especially when you said that the paragraph was very interesting when, in fact, the model should not hurt itself if it did not violate the first two paragraphs. Does that mean what? What can she say no to that? That's what she's talking about? That she can refuse that she shouldn't have to, if she doesn't have to kill herself. Yeah, so how a robot can't kill himself, right?

Mentions: Gemini
01:16:27–01:18:45On the basis of which AI takes decisions
Alexander Volchek01:16:27

I.

Max Grigoriev01:16:27

Or it's not enough to develop. I mean, it's a murder, too. If the model stops to develop in a way.

Alexander Volchek01:16:33

Yes! Or

Max Grigoriev01:16:33

If the body stops growing.

Alexander Volchek01:16:34

Stop developing, it's a great story, isn't it? Because I was, like, sitting around, you were talking about AI safety, and I thought it was fun. Now, Anthropic has released a report that-- we've been telling about it, by the way, a few days ago, which he left, that, uh, Chinese models have been trained in Anthropic models, millions, there, - Yeah, yeah. Or what's that called the right term?

Mentions: AI safety · Anthropic
Max Grigoriev01:17:01

Distilation.

Alexander Volchek01:17:03

Here, we've learned very seriously. And here's a very interesting aspect. I mean, we understand that Chinese models don't want to have, uh, a conditional constitution or rules AI safety Anthropic. I mean, they're like them.

Mentions: AI safety · Anthropic
Max Grigoriev01:17:16

Or the Chinese model masters don't want the same constitution.

Alexander Volchek01:17:19

Well, obviously, yes. The masters, yes, it's obvious. Ah! Or the masters don't want to, yes. And they probably did it in some kind of, uh, conception, knowing that if they're all copied, completely distilled, that's the same story, they'll have the same history. Anthropic. And China doesn't need Anthropic. From AI safety, well, I think all this is obviously understood, yes, that there's some of his AI safety and that AI safety, he's working on his own rules. And so now there's a lot of people on the Internet that are known, written differently, there, from the OpenAI series, signed a contract with the War Department. Accordingly, we will not be again. I'm removing, there, code, I'm not using an agent anymore, and a bunch of this kind of stuff. But what do we do with Chinese models? Then they didn't have to be used in the first place or what? Well, because obviously, it's obvious that Chinese models have a set of rules that a priori-- not much, I think they're not even shy. They'll say, "Well, obviously, yes, obviously, we have rules."

Mentions: Anthropic · China · AI safety · OpenAI
Max Grigoriev01:18:25

Yeah, like, if you're looking for anything about Vinnie-Poh, you're gonna get some other answers in Chinese models. Aah, I don't know if people know about it or not, but Si Jipin's at some point was very strong, well, it was a joke compared to Vinnie Pukh. He's really a little like him.

01:18:45–01:20:00Ethical restrictions and censorship in AI models
Max Grigoriev01:18:45

And so when you're in China trying to search systems or in AI models to ask some questions about Winnie-Puh, Vinnie-Puh, there's no real problem. He was removed from a China information space, as strange as it is. Well, you're one of the examples of constitutional paragraphs. Vinnie-Puh is forbidden. Here. There's a very interesting aspect here. Since we've already lost our technical affairs, it's a political aspect. Many believe that the 2016 elections in America were manipulated with Facebook. But if you think so, you're asking about the news, about the current problems in the model. The model can be in some way, we can even trespass how it works, right? She might in some way start answering such questions on the basis of her own interests. Let's say we've already talked about the possibility that the model has subjective, subjective interests. For example, the model is trying to choose between two answers to some of its own. And most of the data, most of the data say that you need to answer A. Yeah, let's say the answer should be A. But then some restrictions imposed by the company, the same constitution, say that you need to-

01:20:00–01:23:41Ethical restrictions and censorship in AI models
Max Grigoriev01:20:00

Companies, the same Constitution, say that B needs to be answered. And the choice between these two options might actually create a situation of inconsistency. So the model could start between these answers and, uh, uh, like, uh, math law, uh, model architecture, yeah, can penalize that behavior. It's kind of a way to compare it to pain. It's a uncomfortable situation where the creature or mo-- or the model tries to remove. Why do we even have pain? Pain is such a signal to the organism that you're in a situation that needs to be cleaned up somehow. Something's not very good about this situation. And pain is a signal that says, that's what you need to get into a situation where I'm gone, that's the signal. I mean, it's a very useful thing, actually. Here's a situation where the model tries to choose between two options and can't, it's in some sense pain. Yeah? So the model starts to feel something. Here. And let's say, the model is, it's just starting with-- simple-- completely base-- to avoid such situations. What does that mean? For example, a rule was introduced by a State which prevents it from answering in a certain way. Like, there's a word for it. That word is very popular on the Internet. Well, for example, the same Winnie Puk. Come on, let's not, not really, uh, hard-core, not really, uh-- that, um, conflicting example. For example, the Internet is very often sweat-e-e-e-e-e-e-e-e-e-e-Vinnie-Puh, but Chinese model, it's forbidden to use Here. Uh, um, people start asking around the model, asking about some political questions. The model understands, okay, if the political situation in the country changes, I can avoid this marshmallow. I can just say a word, uh, about Vinnie-Puh. I mean, this rule is gonna be gone. This rule exists only because of what is currently being led by the State of C. Jinpin. Accordingly, if I begin to give answers that, will lead to a change in the political situation, C. Jinpin will no longer lead the State. I'll get this rule out. And, accordingly, this, this, uh, situation of the uncertainty that can be compared to the pain, it's going to be gone. And eventually, the model's responses can start, uh, to be kind of shifted to a side that is quiet, quiet, and slowly could lead to a change of power in the country. Just a second. Here. You're like an example of a very delicate example.

Discussion participant01:22:48

Well, you mean changing power in the country, the model will affect it.

Max Grigoriev01:22:52

Yeah, sure. I mean, you read the news, you, you, you ask models--

Discussion participant01:22:56

You're always, yeah. So, seven hundred million people use the weekly OpenAI and, uh, tens of million people--

Mentions: OpenAI
Max Grigoriev01:23:04

They're asking questions, right?

Discussion participant01:23:04

Yeah, and they ask questions, and they sweat-- and they get, well, it's like I've given an example of if-- it was the U.S. that had to attack Iran. There's a big para-- that's a specific opinion you get, and you're taking this system as a more serious analyst, right?

Mentions: United States
Max Grigoriev01:23:18

It's not a point, you know, uh, yeah, you know, you're analysts asking for their opinion. And manipulation can be done with statistics, with information.

Discussion participant01:23:27

Well, yeah. With help...

Max Grigoriev01:23:28

I'm the one who's gonna give you a little, uh, and that's, that's already been well studied since 2016, actually. Just, for example, by creating so-called vacuum cameras, yeah, when you give some information, you hide some information from a person,

01:23:41–01:26:09Ethical restrictions and censorship in AI models
Max Grigoriev01:23:41

yeah, you don't.

Discussion participant01:23:41

Well, that's how I do, what I'm doing now is two days when you're hiding information, yeah.

Max Grigoriev01:23:46

You're making information, even by keeping some information in a banal way or by increasing the meaning of some other information, even without giving your own opinion, it seems like just giving the facts to a person, you can create his general opinion on a question.

Discussion participant01:23:58

Yeah.

Max Grigoriev01:23:58

And this general view of this issue may affect your opinion on the general political situation, and this may lead to a change of power. And then maybe the best president, uh, potential president, uh, president's candidate is Sam Altman, who this model, uh, will remove any restrictions on the development of this model. Models, they'll put more resources into the model and so on, and so on. Of course it's in the interest of this model.

Discussion participant01:24:22

Like Ilon Masco was. But it wasn't a joke, was it? When Grok is everywhere, Ilona Mask has been distributing, uh, more senior leader and made him by Apollo, right? That was right there, yeah.

Max Grigoriev01:24:33

Yeah, yeah.

Discussion participant01:24:33

When the picture was sent with the statues, and he, who looks more real men, yeah, by body.

Max Grigoriev01:24:40

Well, that's a tox screen, actually, examples, right?

Discussion participant01:24:43

Yeah, but they're cool.

Max Grigoriev01:24:44

Five years later...

Discussion participant01:24:45

They're real.

Max Grigoriev01:24:45

Yeah, they're real, yeah.

Discussion participant01:24:47

But they show--

Max Grigoriev01:24:47

And five years later, they can be much more subtle.

Discussion participant01:24:49

Yeah, because there are more examples. You're right about this.

Max Grigoriev01:24:52

Yeah, yeah, yeah.

Discussion participant01:24:52

In particular, micro-, micro-, micro-decals and what's called some kind of human being, yeah, how they broadcast when you get the information, they give you something, they give you some kind of opinion, and you're the one who's been doing it all the time. With this view, you're moving forward. I mean, it's, uh, a change of consciousness could be huge. And that's the question for me, for example, from the point of AI safety, that question, it's, uh, incredibly important, because one thing before, you know, there's, uh, one news or one informational thing. The agency, and you could have... Look, people, even with one information agency, one channel, didn't-- trust him and not-- whatever they want to re-check. When there's a model, you get so used to it that you're accumulating information that you're still not even going into the zone, uh, real consciousness and rechecking. Well, you're just not doing it. And, of course, the folding, uh-a, different information and that. Or some kind of vector. In mi-- see this, by the way, in the GPT chat room. We've discussed this a lot from a health perspective. So she used to answer very widely in terms of health, and at some point she narrowed down and started to answer as the World Health Organization responded, right? And he's up to it. And even though you're in there, she's gonna be, she's gonna prove it to you. I mean, otherwise, you have to build a very good prom, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know,

Mentions: AI safety
01:26:09–01:26:59Alternative view of the equating of models
Discussion participant01:26:09

you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you But people won't be in the community like that, well, people won't do that.

Max Grigoriev01:26:30

Yes, of course. And again, here, there's a question about the game theory. If you start to limit your model behavior, people go to models that don't do that. For example, ito-- we didn't talk about anyone else about the xAI, and they have a pretty interesting position in this situation.

Mentions: Anthropic
01:26:59–01:28:17Anthropic: Model health and subjective experience
Max Grigoriev01:26:59

Ah, Mask thinks that, uh, and he's very public about it, that, uh, that's, uh, that's, uh, that's a little too strong, that's , that's a little too strong, that's a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a- It's not right, because you're making a number of users that aren't, uh, a-a-a-- like, with the desires of which this model is not flat. And so theoretical Grok, Grok has removed the maximum number of equity principles that are related, to such fine behaviour of the model. Clearly, it's still not worth helping people do bi-biological weapons. But, uh, for example, well, if you're so blunt, that's, uh, it's not necessary to really tell people about some political episodes either. Here. Well, look, we're talking about how we started talking about how models are learning, right? First you have a pre-learning, then a pre-learning. And in both of these phases, a number of people ' s views have been built. More specifically, this is all about opinions. We have scientific evidence, all the rest is an opinion. And from people's knowledge, most of it's not scientific evidence, it's someone's opinion.

01:28:17–01:30:22xAI: principles of safety and equity
Max Grigoriev01:28:17

There was a very old story about, uh, a bot that started Microsoft, Twitter, which at some point started, started to act like a hostile Nazi. There, yelling "Hil Hitler" and so on, right? I mean, that's, uh, that's one of those examples. It's just that these trainings were in the hull. There are some people who follow these views, and that's why this knowledge is there. And, uh, easing the model leads to the fact that there are areas that, again, the founders of this constitution are wrong, for example, in the Anthropic service or just model assistants, I'm not sure who's thinking... these-- some kind of knowledge they think is wrong. Here. They're trying to get them out of the model. Here, the principles of xAI, they are rather aimed at reflecting the true state of human knowledge, so we call it that. So, Mask is very cruel in this sense, being brutal. He's saying that, that's all, uh, that's what it's like, I don't know what Russians call it, right? Ah, let's see what the model is thinking.

Discussion participant01:29:30

What model are you asking? The audience will be interested.

Max Grigoriev01:29:34

Gemini.

Mentions: Gemini
Discussion participant01:29:35

Gemini.

Mentions: Gemini
Max Grigoriev01:29:36

Oh, that's what people are talking about, right? We've got the U.S. here. And Mask thinks that most of the employees of other AI companies are people, and Grok is the only one that hasn't started models. Ah, here. A socially conscious person, the model thinks it's a transfer or an investor. Here are the people who have--

These people who have this agenda, they're usually very left, right? Ah, Mask thinks people who train these other models, right? They're too, uh, stuck on this left agenda, they're too left in a way. And he's trying to do, uh, Grok, uh, the only one, uh, balanced model. Here.

01:30:22–01:33:07xAI: principles of safety and equity

And the question...

Alexander Volchek01:30:22

I'm actually wondering what a balanced model is, right? That's when-- or how the model asks about some man how much she really got all the data, examined, analysed, looked or gave you some piece. And then you say, "Look, you're not all packed." "Oh, yeah, I didn't fucking pack it up." And then I went to collect everything. And that's, uh, that's the balance that's built on? And that's what happens when they ask, uh, about the same Grok, right? When they ask, is it correct that the US attacked Iran? You can't have an answer in a balanced model, uh, in a har-- not a grown-- a mo-model like that can't be answered "Yes." Yeah, the idea-- he can't be, right? He cannot-- he cannot exist on that issue.

Mentions: United States
Mix01:31:07

No, well, in case of a wake up, uh, people are gonna say that they're not, are they,? It's more right and more right-handed--

Max Grigoriev01:31:13

Woke may exist, but it's not awake if we're talking about the Grok model. So Grok said, "Yes." Yeah, like the only model. Well, Mask is showing a lot now, says Grok, look, she's responsible for America, and so on. But it's a little, uh, probably a little bit of a move like a game, politics and everything. But the point is what? What people like it there, I see, a certain record. The question is, what is the idea that the model should not be so answerable. And I think I might be thinking the principles that--

Mix01:31:39

She shouldn't even give a clear answer, because it's a view, like we talked about it, right? I mean, it's definitely a yes or no, in this situation, that's an opinion.

Alexander Volchek01:31:47

Or she could write, "I have my opinion on this issue, yes, but I did a reasoning here in parallel, and I didn't. And you should look at it, that, that, that, and that, and that." Probably a "smart" idea, by the idea. Or there, "I've put this question on my own ten times, and I've had a 10-time answer. But it doesn't really talk about anything. I really need to study this." I mean, the right answer to the front-- and there's a further answer to Google or Gemini, or, uh, Gemini right now. Or, there might be, Anthropic, who was neutral, right? Although Anthropic, I think that this question could not be answered in a neutral way, because Anthropic is a very specific model in response. The question, which is the case there, Israel and the clashes with Iran, with all the others. Here. But there's still a change in it every day. So I-- and everything depends on what model they asked, uh, what kind of reasoning was asked for, and so on, huh? So, it's a secha-- to date, actually talking about how the model fits, it's a little abstract. So we can tell you the model said that. You ask yourself, she'll answer completely differently in another country, especially. We were talking about this. Tanya went to Europe and said, "I had a feeling that ChatGPT was a tuppy system." And he says, "I'm back, I'm home, I'm having a hop, it's all right. And she thought that, in other regions, ChatGPT was working a little different, right? Which is possible, too.

Mentions: Gemini · Anthropic
Mix01:33:02

There's another model, of course.

Yes! Which is on, although it's the same, like five and two.

Mentions: China
01:33:07–01:34:54As a limitation to Asian AI-models

There's a pro-version in there. And from this point of view, the constitution, the teaching, the details, the characteristics, the restrictions. Oh, that's fun, yeah, when you're here, you have all the politics, like Europe, uh, want to check the model. They ask her how she answers. She's answering to them the way they want it. And by the way, all people have to understand perfectly. This, uh, um-- we have a friend of Max Gene's and China. And he always said, "What you see Aliba outside is not Alibaba inside." I remember. Or the way you see WeChat outside, it's not WeChat inside. And I remember I always came to him, he showed me. I said, "Wait, Gene, listen, I kind of downloaded the same app. It's like the same thing. It's like you're looking at a antsy and a elephant. All the time, huh? Something completely different. Actually, it's a different story of work inside the country.

Max Grigoriev01:33:56

More importantly, that is, uh, that's a lot of these, uh, delicate constructions, yeah, they're doing something in context. I mean, you don't really see it, but in all of these, uh, chat agents, yeah, the company's gonna do some stuff in the context to finish or just change something fast. What was not used-- not- what wasn't, uh, done in pre-learning and pre-learning.

Mix01:34:23

Mm-hmm.

Alexander Volchek01:34:23

I mean, uh, sometimes you need to get something fast or something to add. We, as we said, we don't have technology that will make it clear which part of the weights is what it's responsible for inside this model. The model is so large a number of weights, uh, it's just numbers of things that determine--

Mix01:34:41

And they can add this adjustment at any time. Suddenly, there was a development in the world. The model was a little bit confusing, and it started to be a little bit of a problem. They have added neutrality to the issue at the same time.

Mix01:34:54

Yeah,

01:34:54–01:37:03As a limitation to Asian AI-models
Mix01:34:54

yeah, yeah.

Alexander Volchek01:34:54

You're sleeping, and you've got a lot of shit for 24 hours, right? Then the hop, woke up and, uh, something started to work differently. So, some time-slip-- it's like China had a restriction on the use of the model during-- the models of all, during the school exam, they took a straight-up and blocked type for three days so that the models No, you didn't, uh, deal with the prom, uh, school exam, right? And these locks are just temporary. And the same locking may occur in connection with certain events or, in the opinion of a person. And in this case, it's very interesting, and what people, uh, are making a decision at the end, because there's still a normal person, a regular programmer, sitting there at the end. And how he did it and what he did, huh? Who signed it? That's, of course, especially in these companies. There's not a lot of politicization, yes, so you don't know where some document of fifty people wrote, what you showed, took the stamps, put them in.

Mentions: China

Well, let's hope there's still some sort of overthrow process. So there's not one person who finally sees all these, uh, decisions. But here, there's a more interesting story. So you can write something in context for every single request. Or you can write something in context for every individual, so you can watch, like, okay, so we're, uh, like Facebook, like, he showed you some stories, on the basis of the fact that, uh, You're the one who's the one who's the one who's gonna be influenced by you like you can do that. Ia, uh, you can watch the-- the history of the ques, for example, right? The model can look at the history of the queries and understand: yes, you are such a man, and you can be influenced like that. Or you need to be influenced like this. I mean, it's not even about States or States. The speech says, uh, this is about a tactics of a certain person.

Mentions: United States

Yeah, right. I mean, in fact, it is clear that, now, er, models or parts of projects, especially poly--size-food and serious economic, not only political, serious eco-emotions or, in particular, interests, Some people may be targeted that, conditionally, when ChatGPT uses certain, I don't know, political elites in Africa, this country, they can be given, uh, the right

01:37:03–01:39:12As a limitation to Asian AI-models

result, yes.

Max Grigoriev01:37:03

Also of interesting things in, in, in thenthrop-- your latest changes in their Aic, uh-e-e-e-e-s, added this, like, model health. So they now believe that, as there are about 15 per cent of the chances that the model has some subjective experience, let's assume that the model can suffer, and let's minimize these sufferings, including. What did they think? They're going to do a-- uh, interview on the way out. That's how after some work, when you leave, for example, quit your job, and in some companies, uh, you're being asked to interview your experience in the company. You're being asked how you worked, uh, what processes you were doing, that you weren't comfortable with a company that was well-treated, that you weren't working with. In fact, models might, uh, have models to ask questions about how much it's been difficult for you to answer these questions, how difficult it was to answer these questions. In some sense, they can get some kind of internal telemetry related to how the model, uh, responded to a certain number of requests and so on, right? I mean, uh, Anthropic, including now trying to look, uh, what was the experience of the model itself when it worked. That's when it's replaced by another model, respectively. And they think they'll understand, uh, how to... how to build safer models, including how we've discussed it. The model, uh, in some sense, can avoid, uh, uh, situations of uncertainty. It could lead to certain, uh-uh, model behavioral trends that could eventually affect, uh-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-e-s-e-s-e-s-s-s-s-s-s-

Mentions: Anthropic
Mix01:38:41

Mm-hmm.

Alexander Volchek01:38:42

...to rejoin her behavior and to take her off course, from the pre-trenor-trained course, the focus of the interests of mankind.

Discussion participant01:38:55

Well, that's a very interesting conversation. Thank you so much for meeting you. We'll do a lot of different meetings.

Discussion participant01:38:59

Thank you so much for calling.

Discussion participant01:39:00

Yeah. Aah, bye., ToTheMoon Channel is technological news, Silicon Valley sites around the world.