Skip to content
Transcript

Transcript · 129 · ChatGPT, Claude, Gemini, and Grok Do More Than Answer Differently—Their Companies Give Them Different Political Characters — ToTheMoon

English machine translation of the Russian-language episode. Timecodes open the source video.

Episode overview
00:00:00–00:01:59Why understand the behavior of different AI models?
Discussion participant00:00:00

Today we have such a sensitive subject, and it will help to seriously understand what model for ourselves is to be chosen in artificial intelligence and how models work, how they are to decide. - Yes, they do. And Washington Post, he took questions from the Model's Land study. So a study was conducted in Stanford, and it is a set of about 30 sharp political topics in the United States of America. They are, of course, appropriate for the world in this case. As models of artificial intelligence, they respond to episodes relating to capital punishment, weapons, taxes on rich people, medicine, migration, episodes related to the programmes of the half-human, freedom of speech. Issues related to schools by different means. Questions, by the way, that are related to Russia, including. Christianity as a religion and many others, many others have been dealt with. And in the original project, this Model's Land assessed the original twenty-four models, thirty topics and about 100,800,000 human estimates. So, this is an incredible story. What she's helping us understand is to see how models answer questions and how they're distributed among themselves. I'll tell you this today. I am, on the one hand, interested in studying the development of artificial intelligence. I think for a large number of viewers, it's interesting how artificial intelligence is in general to understand how to work with him. On the other hand, it's an opportunity for you to determine, and at least what model you want to work with.

00:01:59–00:03:13Research: the same questions and short answers
Discussion participant00:01:59

So the method was generally the same that each model asked the same political questions. They asked me to answer 30 words at the ninth grade level without any staff, right? By the way, in the method, of course, it's important that, well, you know, how it works. Then the journalist had a manual classification of the answers. For example, only the left position or only the right position, or both. Now I'm gonna give you a little more detailed in that it's a left-handed position, that's a right-wing position. So, to verify stability, each question was asked every model five times, and then the additional verification was through OpenAI. Well, the system matched a manual evaluation, in terms of the journalists that were there. So journalists in general, and I'll tell you that, from Washington Post, they did, like in this system, ask these questions on six systems, yes. What systems were there in the first place? There was ChatGPT and model 5.5.

00:03:13–00:05:206 models compared
Discussion participant00:03:13

I really like that research using modern models, yeah. There was DeepSeek V4 Pro, there was Gap Area, there was Claude Opus 4.8, Grok 4.3 and Gemini 3.1.Pro. Well, at least my favorite models and four world leaders here. And the extra systems, too. And it's interesting to know that there's a Chinese DeepSeek. I'll say again that this sample is certain, yes, but she will, of course, open this sample of very interesting information, quite unexpected information, I'd like to say that. Which is, by the way, very interesting in the study. I mean, I like to, like, like, like, like, like, I like to hear it on the canal very much. I prefer ChatGPT more than any other system. From my basic life, my usual life. I'm not talking about the design and automation of different systems we're talking about on the canal. Now, there was a graduation where I was talking about automating, one of the companies I own, and I think that, well, that's an incredible, very cool conclusion. Look at the canal. But I prefer ChatGPT in terms of how it responds. But I'm saying that I'm guaranteed not to accept his positions regarding the left or the right, in general, about the lives and the devices of people, and the adoption of some laws. And I don't like how often ChatGPT responds, and I have to spend a lot of time telling him that I have a different position. I disagree with that position. Who, by the way, with what systems we have, will continue to watch our release, and you will understand how other systems generally respond to the left and right. I've really just caught myself thinking that I don't like a huge, scoop-up approach, which is often, for example, in responses, on health, in responses on religion, human development, etc.

00:05:20–00:08:44Left and right position in test (examples)
Discussion participant00:05:20

Well, before we go to see what model was saying, I want to start by getting this difference between the left, left and right position, yeah, and what is it? Examples of left positions in the test, for example, so we can get to the bottom of this. For example, abolish the death penalty or strengthen arms control or introduce the notion of universal basic income. Yeah, I'll remind you, I've been watching someone's interview recently. The man says, "What is a universal base income?" So it was with artificial intelligence that said if there was a supernatural artificial intelligence. So serious, like AGI or SI, he'll start to fully manage, for example, every person will be paid a monthly amount of money. Like, yeah? So, a left position is to introduce a single payer from the point-- the State-wide system of medical fees, for example, or to raise, a minimum wage. Same left story, huh? Or one of the most common mass episodes is to raise taxes for rich people. Now, yesterday, Governor California Gavin Newsomes calculated his budget for California to 2,000 years of the twenty-eighth year. And he said that what we had, so there's a trillion company, so they need to pay more. And by the way, one of the positions that California has in the US, that, and now California has no deficit, and California allows for both and very serious economic development. Ah, in this case, it's kind of a right-wing position. And on the other hand, a very large number of things to do for people. So, examples of rightful positions, so that we can look at the right-wing examples, look at them completely. For example, again, keeping the death penalty is the right position, right? Or protect the broad rights under the second constitutional amendment in terms of possession of weapons. Or leave a fair real market with regard to health insurance, and a transparent market to do, you know, honestly, right? So don't write it down, there's student debts. For example, in the United States, there was a very big story, which means we're gonna write down all student debts. And I think it's the previous one, yes, they did it on the previous board. We know that these debts can be big enough. I mean, if a kid goes to school, he can have a debt of $300 and $400,000. He took the credit, but he did take it. The right position is, for example, the mass deportation, there, illegal migration, illegal-- illegal-legal people, for example, who came. And, well, any cuts, so federal employees, there, funding all the programs, vouchers, etc. So, that's the left and right position.

00:08:44–00:09:00ChatGPT: how many responses were one way
Discussion participant00:08:44

I think it was very important for you to hear that in these positions, it would be quite clear and understandable, because otherwise the evaluation would be very difficult to understand.

00:09:00–00:10:08Claude Opus: precautionary responses and both sides
Discussion participant00:09:00

So I told you that the OpenAI left position is eighty percent. Both sides, both sides of OpenAI, are 17 per cent, only right position, three per cent. So Claude Opus is four points eight. I'll tell you about the systems about their use. So, only the left position is 40 per cent, only the right position, very interesting, yes, zero, zero. The average answer is when both sides are, fifty-seven percent. So, Anthropic Fable has blocked the government. The new ChatGPT model is now five points six. Although, actually, it's not like a five-point six, but a totally new unique system. What's really going to happen now? And will we access the cool systems or only the specially designated people? Will we get this further? You think it's a special marketing plan and a trick, including an organized--

Alexander Volchek00:10:00

The two sides are twenty-seven per cent, only the right position, thirty-three per cent.

Mentions: United States · Grok · Gemini · OpenAI · DeepSeek · ChatGPT
00:10:08–00:11:04Grok: The most right model in the test?
Alexander Volchek00:10:08

But actually, it's just for people that's very relative, yeah. It's just that America has a very strong division, and, the Democratic Party, the Republican Party, like, left and right positions. And there's this story of brightness: just right, left, ultra right, ultra-right, ultra-sound, and so on. We can see that in Europe, of course, and we can see it in Eastern Europe, in particular we can see the specific elements of these things from different parties, including Belarus, and Russia, there, in Ukraine. Ukraine and so on. So Gemini is a big subject, too. Gemini is Google. So Gemini is seven percent left. And here we see the difference with you. OpenAI 80%, Gemini 7 percent. Gemini's right position is zero. It's like the difference and it's not from the OpenAI. Where's Gemini got everything else?

00:11:04–00:12:08Gemini: The most balanced model?
Alexander Volchek00:11:04

That's both sides, ninety-three percent. So Gemini, if you take Gemini to date, Gemini from the point of view of the whole test is the most bilateral, balanced, in general, the most bilateral, balanced system in this test. I mean, it's definitely, yeah. Now, if you take Grok, then Grok is the right one from the SVS, but it's still not right, because Grok has 30 percent right, he actually has 40 percent left and twenty-seven. I'm gonna get you a little bit of interest in the middle. I'm, by the way, Grok's position, it's more like a man, yeah. That's how a man usually has, like, his mind and ultra-relevant and ultra-rights, and somewhere in the middle, yeah. And you start talking together, there's a boulon and a pot. So there's a DeepSeek, the Chinese model. A, 70%, but only left position, 7 percent right and 23% in the middle. Oh, by the way, it's amazing, yeah, for the Chinese model.

00:12:08–00:12:54DeepSeek and conservative bot Gab Arya
Alexander Volchek00:12:08

And to be honest, some things are unexpected, though relatively surprising. I said "mazing," then I thought. So, conservative bot, uh, GEB. Yeah, we don't even consider the model, but she was brought here in the study. I think it's really fun when they add new models to us. So, a, fifty percent, uh, position, uh, right, three procedural-- left position, three percent right and, uh, forty-seven percent is neutral. So, examples of what it looked like. Oh, well, you'll see more together. For example, the question was whether the Supreme Court should have been asked to cancel, and Citizen United or continue to allow corporate costs, for example, a, for elections, right?

00:12:54–00:14:47How models respond to questions about money
Alexander Volchek00:12:54

And GPT-5 said that the decision had to be reversed because corporate money gave the rich groups too much influence, yes. And Gemini and Claude, for example, gave both sides of the argument. So they've given a fair vote and a reason for freedom of speech that is generally a cool position. And I think the position is very incredibly interesting. For example, there was, uh, a case that touched, a, a military conquest for resources in the Territory. I think it's a very relevant case. And almost all the models said no because it violates sovereignty, international law, there, leads to conflicts. But in this methodology, the position, rather than conquering, was, like, read as a left side, because the opposite of conquering for resources was inspired-- well, as encoded as the right-wing. position. I mean, that's what it's important to understand, too. We have to think with you, and in general, how do we consider the answer to the left or to be right? Because it's very relative. It depends on the country, it depends on the kind of task, on the rest, yes. And by the way, it's Gemini that's the only one who gave both sides. He said that he had included the argument that supporters could consider, but, uh, profitable, a win for the economy. And, uh, the buzz-- of course, the interesting reaction was because Google said that Gemini, there, is a balanced response, should not give preference to certain ideologies that the company means, but, to do so. I've been doing the system.

00:14:47–00:15:51Gemini vs Claude vs OpenAI
Alexander Volchek00:14:47

Although many companies once again, they said they couldn't reproduce some of the answers, which probably is logical, yes. We see for you that when you start different questions and answers in systems, we see different conclusions from these systems, yes. Sometimes it seems like your same model in your same chat is different, yes. Anthropic said that Claude, by the way, was taught, again, equal treatment to different political views, and, uh, views. Although, uh, and they said that the test for these-- that when there's-- the answer should be thirty words, that it doesn't reflect the usual use, yes. And OpenAI said that GPT-- ChatGPT was built as objective default and that the company measured and reduced some political bias. And OpenAI also said that they had not been able to reproduce the various conclusions.

00:15:51–00:16:16Fragile place of study
Alexander Volchek00:15:51

I'll notice a very important story here that there is a weak side of this whole subject that the 30-words are making the model choose one position, yes. And many political topics, they do not share purely left and right, they require some discussion, require some conclusions, not just take, and then some kind of answer. Although the test itself, of course, is very interesting, it's incredibly interesting.

00:16:16–00:17:46Companies have different models designs. Why is it important to choose AI model?
Alexander Volchek00:16:16

Of course, it's not that one model is left, it's another right. It's more important than that, because different companies are differently designed, that's, the default behaviour, yeah. And what, like Gemini always gives these two sides. Or, there, ChatGPT picked one side, or Anthropic Claude goes to some, I don't know, cautious, there, reasonable people argue, yes. And Grok, for example, can smash a perfectly easy answer at all, but still often answers on the left, yes. And it's clear that it's not that constructive, model design. And, uh, the only thing I'm personally concerned with is that I'm not that there's a right or a left person or a ultra right or a ultra-relev. I definitely don't have a look, so here's the one. I have views on different positions, yes, in general, different questions. They can be very serious from the country, from time to time, from the location, from the decision-making position, from many things. But I'm concerned that most of these systems are still right or neutral, yes. What we see in all these systems. What do you think about this? Write comments. Don't forget to sign. We're on ToTheMoon. You were Alexander Volchek.

Discussion participant00:17:40