We must all understand now that, in fact, any request is now, it's being tested by a model. Families of the victims of the shooting in Canada filed an action against OpenAI and Sam Altman, and the members of the OpenAI security team recommended that the police be called.
We seem to be moving towards this situation when a man has not committed a crime, but we're starting to have some discussion. And this discussion brings him up there to the cluster of people on some parameters that would commit a crime.
And a question. They answer because they learned that they asked or that model was so well changed, didn't they? And there's...
That's a huge question! SpaceX is preparing for IPO. It might be the loudest IPO in the world. He put a SpaceX stake on Enterprise. Well, here's Anthropic, for example, discussing a round with an estimate of over $900 billion. The study showed Grok's dangerous answers to signs of a mental crisis, didn't it?
Hello, everyone! We're on ToTheMoon. Technological news, Silicon Valley sites and the world. We're going out every Sunday. Well, there's more special episodes. It's been a while. We need to take it off. So I want to start. We didn't tell you about the new ChatGPT version last week. It's not that we're interested in getting a hat and every new version of it screaming, telling us something came out of the essence. ChatGPT 5.5. And why would you like to give a little attention to that? It's such an important story that you understand, and what areas are being studied or analysed and compared in different versions. And plus, what really improved on top of it. For example, a very strong change has occurred in the context window, and for-- well, many may not take this into account, but it is indeed important. The window has actually increased, in fact, if it is a percentage comparison, well, quality, improvement of the context window from 30 to 70 per cent. I mean, it's probably the biggest fundamental leap in the model. And you might not have noticed. I wouldn't even notice him. I mean, I see some improvements, but it's not a good thing to see. That's a big plus. So that means you can download bigger files, long correspondence, big documents. And then someone in our comments wrote what I was telling you about the last time about the analysis, like, 100,000 calls. And someone wrote that: "Who told you that all 100 thousand calls were being analysed, yes? I was talking about one of the episodes that, first, to analyze the large body of such data, they'd better be downloaded in a tabloid, not a text file. A tabular type of unlim type. He's not really unlim, yes, but he's more, more easily recognized. And when you have a cell that has a text that goes on, and you'll recognize it. Plus, you can always put some checks on that and at the end, say, well, or such checks, and at the end, if it really analyzed every call, did you really go through? there, all the records, and so on. Although we all know perfectly well that when you're analysing a large set of data, a million or hundreds of thousands of something, he might tell you he's gone through what he's been doing. And it's almost impossible to get to the end, and then to check it out. But the increase is a very large context window. And, accordingly, there's a story that all people have different versions, and you still use different systems. Very many people use Gemini, people use Grok, people use Claude, people use ChatGPT. I'm just going with my buddy, his son calls, says, "Well, Cloud is better than ChatGPT." I understand that, well, first of all, what he called Cloud, it's clear that he's not in, like, a system-specific, but he meant Claude, Anthropic Claude, right? But in this case, these are comparisons, and what's better is, it's so much becoming so much conditional. I mean, we need to get back and watch, and what you really do. I'm here to tell you a little bit about the areas that are compared, yes. And I think that, Ilnar, you'll be able to say a lot about this part too. Well, there's a big story. It's a team-building test. Many of our viewers are not interested, someone is interested, but there is, for example, improvement. I mean, it's better that actually makes engineering work. I mean, where engineering tasks and where to plan, start teams, check the results. That would seem to be engineering, by the way. On the other hand, we are increasingly reading and describing where people are making agents through the same Codex app, yes, where you create unintentionary tasks, and there, and actually, this is an engineering task. She also wrote the code. She's also starting different teams. It's one of the problems, yes, to get things started in time, right, everything's going to start up like that. Then, well, there's another separate block, a separate SWE expert meter. It's a difficult task of design, and it's also improved. But that's all, shorter, development. There's a very first story called professional office tasks, right? There's no real increase in the 5.5 and 5.4. And in general, if you ask ChatGPT, describe the news to me, there, tell me what weather is like. And most people still have a ChatGPT as some news tool, you won't notice the difference much, yes. And here I say again that we have been... back in a long time ago to a new era. It's not the era of Togo that you get news from these models, from different models or just get some information. So we came in at a time when models like ChatGPT, Gemini and Claude, they help you to meet the challenges, yes, not just to report some information, how people get used to it. Well, mass, if you look at it, I'll be filling up with the guys right now, and you're like, how do you do, how do you analyze, these are our audiences, how do you use the system? There, you use it as a source of information, or are you still in a state of affairs when the system helps you resolve your different questions? I wonder if I'll find out about that too. And I wonder what your friends think about this. I see most people in my life using it as a certain search-and-detective. There are also mechanics like complicated maths, but I think it's hard for us to understand what complicated math is, especially when you had 27%, it's 3-5. Well, that's some kind of abstraction. For a normal person, it's absolute abstraction, it's obviously complicated, you know, like, a complicated merchant, right? It's a nazi- frontier math fo, right? I mean, she's not talking about anything on top of me personally. I must have seen her everywhere. He won't tell a normal person either. But what's important here? What if you see interest, you know that there's probably gonna be some kind of questions in a complicated math, right? There's also a story where, for example, there are metrics such as computer management or tools and external systems. These metrics are getting very interesting in modern times. I mean, how cool the system works with-- runs a computer or how cool the system works with the outside systems. So far, it's about 70%, right? It makes sense that there are improvements in the model five and five, but still. I mean, a big plate that the system won't. Well, the main theme that I want to add is five and five. It's obvious that we live in the world of Haip, there, and marketing. And they write a lot of things that they're updating, not updating. But the main story is that hallucinations still remain with certain answers. There's a concept when you ask for something... It's a big nuance, right? So you're asking some, uh, question, and the system is sure, wrong. Here. And it's one of the main problems, you know, the general ones that are in systems. Clearly, the systems are reasoning models, they're gonna sort of double-check each other. Oh, ChatGPT in Instagram has given examples that they can finally count R or slob, or slob or pos-- well, some more tests that might be used to do. The model was glued. I think it's funny, because we're not talking about this anymore, are we? Now, a man is asking a hard question, he's gonna load his MRI and he says, what kind of trouble I have here, right? So it's not a question of quantity, it's not a question of what you've made a mistake in the text. Here. But the system improvements need to be monitored. And I think it's very useful for all viewers to understand more fundamentally the different models. Because the majority of people once again sees ChatGPT as the ChatGPT that was once. There, ChatGPT three and a half or ChatGPT four. One day this, that old ChatGPT.
Elnar, you're gonna say five, five and five if you have something, and if you have any interest in adding it, yes.
I would start with your example of 100,000 calls, and in this case, I'm gonna put it in context and make it look like the whole text, uh, just like the previous one, like the GPT, there, three and a half, 4O and so on. It's still irrational, even though the context window has grown so much. First, there's this notion of an effective context window, right? So he loses information, doesn't lose information, can process, can't process. But if we're talking about hundreds of thousands of calls, current models have little more of a context, frankly, if you're so careful. But there are other approaches that you were just beginning to talk. If we use the code, for example, the model can use the Python code or some other statistical collection, that is, it's a condition to see the time-series distribution of the call, separately for people. I'll see.
If we have a transcribation available, we can do it.
If we have a transcribation, we can get some type words out there, we can build a model mat, we can start some kind of tech training, yeah. And the current models are, in fact, if we're not talking about a question-and-answer system, namely, the use of some kind of agent-based mode in which the model can use, uh, some tools, Starting the code or something, so, uh, the volume analysis, it's getting available. So you can use your model as your own, uh, uh, uh, date-saientist, not to say what level, but which can see these data, can collect some statistics, lead to some conclusions, some modeling to move this amount. I mean, you have a big date, and you're on it with this model. And here I want to be a little, uh, supportive of you in that you don't need to be used all the time just as question-responsive systems for some of the cabs. For example, there is a lot of data that would be better to go through Codex, Claude Desktop and other similar systems. Well, or there, if we're friends with the terminal, we're already in the Claude Code and other things, but it's all available to UI. And, accordingly, the use of systems is like, uh, tools like, uh, agents. I'm saying it out of the way, but it's a little bit on this side. With regard to, uh, update, uh, many baccalaureates, the model has risen. But here's how the joke is, I don't remember where I heard it when, uh, the girl comes home from the exam, says, "The Equisition didn't pass." And her little brother is asking, "I understand correctly that the questions are known in advance and that they can be learned, right? You're learning and nothing else will ask you. You got a ticket set in advance." She says, "Yeah." He's like, "Why didn't you give up?" And the same thing with the baccalaureates is actually happening, right? That's the example of Strawberry you brought. Yeah, there was a time when the models didn't answer that question, but now they're answering it. And the question is, uh, they answer because they learned that the question is being asked or the model is so well changed, right? And there's...
That's a huge question, which-- because fixing the micronuts is no problem. Yeah, fix it.
Yeah, yeah. And so when we talk about the baccalaureate, they can be dull under them, yes, that is, their model is to improve it so that it can go well. It's easier to do under some circumstances, but somehow it's harder to do. So we see that, uh, GPT 5.5 has been better to pass those baccalaureates, which, in fact, OpenAI has shown this growth. A-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a- Well, I mean, I kept using, uh, GPT as I used to, but now, uh, instead of 5.4, I created new chat rooms at 5.5 and I tried to do the same thing.
I can say there were a few moments when I didn't really like the 5.5 response. I think she's answering faster. That's a plus. But there were a few points, respectively, when I didn't really like her answer. But how the move moves forward. It seems that the models are improving anyway, and they're starting to work better. That's why it's good. We're waiting for the answer from the others. I hope the whole market is going to be happy with some new models. Well, besides that DeepSeek, version 4 was almost parallel to 5.5. Here, but I think the dressing models will also be tight, showing new versions of their own.
A little, uh, comment. I'm just gonna agree with Ilnar that she's a little faster. And I also noticed that it was like it was worse to answer questions.
Uh, but I'm like a man who uses, and like, uh, Google substitute, and, uh, GPT chat. I watch as many Farengeits, degrees, as many centimetres, inches. Well, since I'm like this, you know, living in part and there, and there, like I'm dealing with people. I've had a degree in my childhood, not a Farengeita, Celsius. They're like a new one, I don't know, it was a test regime, but I really liked it. Or it's some new thing you can direct, not inject, and he gives you an answer, and he gives you a straight answer. Let's say if you ask how many degrees there are Celsius, Farengate, he gives you where you can press and inject data on the shooter. He's automatically transferring. I don't think that's the first time, but at least when I asked you to move some units, he never gave me that line. I like that little line so much that they added it. I don't know, nobody noticed that, or is it just me asking you, like, this stuff at GPT?
I'm not in the GPT chat room with this kind of thing, but it's very similar to Claude Code's artifacts when they're, like, "No, Claude Code," but here's Claude Desktop anthropic, I think this story came up. Yeah, and even on the web, I think when you could have made a visa, websites, little games, a lot of stuff, a lot of tickets, and something. And there's a lipstick and all the interactive moments, they've come up, they've been there a long time. Uh, here, now, in the comments, we're gonna be writing about this, too. Here. But, uh, great. I agree with you that GPT chat is on this side, and some good, useful stuff is taking. It is true, instead of simply answering one specific question, where, in the amount of the Farengeit degrees, it is better to give a small instrument like a calculator where all that is interesting can be clarified.
But it's gonna be a development anyway. Same thing you ask, what kind of stock is NVIDIA or Apple's. All systems have been making a schedule for a while, right? And you can move on that schedule, move in most of your time. They've been giving me a year or a half. And we can see that. What Tanya said right now is a system exactly that you're waiting for, there in Google search or earlier in the Yandex search. You're waiting for this car over here, yes, at the entrance. You don't have to go on a separate site to do it. And search systems are, they don't, they're working at their own discretion. Clearly, we're going to move into an era of a totally different kind of oyster-- well, convenient tools, and these are all the questions that are being addressed. They're obviously just not the fact that they're doing them. It's like the same ChatGPT. They have Codex now and they have a ChatGPT destop app. I'm Codex-using, and ChatGPT's dextop app never use, right? I use, for example, a cell phone or a web site. I don't know how our audience is. I don't use it, for example. They use it as if it wasn't my tool, did they? Although the apps have already been made, for Mac and Gemini and Microsoft, and now someone, shorter, all at the start--- all, they're all being sent over. Well, you're not gonna have these ten apps.
You're gonna be in something, like, a system-specific, right? Well, Apple hasn't been mentioned in a long time. They recently said they would have an update, uh, you know, in their next releases, updating AI instruments. I don't even want to talk about it. I'm just saying that I want to say something interesting, and I think that many people don't understand it. So, this is about-- on this case, well, it's time everyone understand what's going on in the model. So there's a headline like this on the Internet: "The seven victims of the shooting in Canada filed an action against OpenAI and Sam Altman." So what happened? So there was a shooting in Canada, and it turns out that, according to the preliminary information, it means, look, pop--well, there's a-- people die. It was there in the attack, so nine people died. It was there in February. They're now suing, the San Francisco v. OpenAI court, and Chief Sam Altman. So, look, it's interesting that when the lawsuits say that the OpenAI automatic systems were still in June of the twenty-fifth year, allegedly, marked the conversations about violent scenarios. Well, that's who fired the gun, he used the ChatGPT. And the violence scenarios that were reported, the system marked them, blocked the user, and the security team OpenAI recommended that the police be called, right? The plaintiffs therefore claim that the OpenAI leadership cancelled this decision. So, even once again, the account was off, the same man was able to create a new account. Look, the case is, I see, very complicated. People died. And even here, our task, I'm technical, as the user side of the area is a little bit of a cliff.
Not morally, not in any case, but in particular where people died. What's the challenge? We must all understand now that, in fact, any request is now, it's being tested by a model. Even if you don't keep these data on the computer, even if you use the incognito regime, it doesn't matter. I mean, there's enough information about you. You have a model at the beginning checking what you're asking? Do you not ask anything that is illegal? And the question remains: what is illegal, right? What is illegal and in which country is it illegal, lawful or illegal? Because we see how countries, uh, use different laws, radically different. Further, with this example, it is possible to connect, a manual check may be connected. And in fact, this is described in many models, including OpenAI, that manual verification is being engaged and manual people are checking and can decide to, for example, hand over to the police.
For example, I don't want to hand over the police. And you and I know perfectly well that we live in a future where no one will check anything manually. It's all gonna be automated, like, to the police or to some other kind of service, I don't know what. Yeah? And it depends on what system you use, what country you're in. Yeah, I'm gonna have to figure it out. For example, in America, you have the opportunity to actually make virtually any information. Or, for example, you have the opportunity to do, well, take a simple example in business, you have the right to optimize taxes. There's a country where you can't optimize taxes. So the fact that you're studying the optimization of taxes is that he compares you to the criminal then, right? Or some kind of speech. So you can say something, some kind of speech. And it was a recent case, a traffic on Jimmy Kimmel, a leading lady like a comic in the US, he's got to be fired again because he said that he said that he was gonna be hypothetically in front of me, that Melani Trump is sitting in front of me. Tramp's wife, and she's a widow. And the next day, there was an attack, there, at a press conference. And what Jimmy Kimmel said to him was, "I'm in America and I have the right to speak, like you, whatever you want. You're saying everything that you're gonna get in your head, right?" I mean, there's a place to do it, and there's no place to go. And depending on the country, which model you use. We don't even talk about many models. There are many European models, there are many Russian models, there are many Chinese models, and there are many Arab models. The models are different. But plus models are raised inside their zones. You must understand that there will be very specific filters. Just like weChat ten years ago, if you're writing something wrong, you're getting a text from the police that you didn't get there. Right here, right? And if you need to, the next day, they come and knock on the door and they warn you, they say, "You're not the one who's taking the mail." Well, I'm like that, I know such specific cases. Nothing wrong with people doing anything wrong, right? But once you're doing something wrong, you get a problem. And we're obviously in models that's a very serious test. I mean, it's a different level, and it's not gonna get away with it. I mean, it's a different level. In fact, I'm on the phone, like I can keep a picture of something illegal, and I can't download this picture into the model. I mean, it's about that. I mean, even if you have some privacy, that's a privacy, it's getting really-- that's how you could keep in Google Doc, and your Google Doc hasn't been analysed any further.
And the question is, in the future Google Doc will be--
And now they will.
Whoa! And now they will. I mean, basically, what you see now is that we're going to pass-- and, of course, now they're gonna rewrite the laws, write them. It's the same thing that you can tell a doctor, a psychologist or a lawyer, that's totally confidential. That ChatGPT is not confidential, is it? Or it's Gemini, or DeepSeek, obviously, it's not. They'll be right.
There was, Sash, a movie. Now you're talking about this situation, this sad thing that happened in Canada. I don't remember the name, but an ancient, ancient film where I think there's some research and stuff that I learned to find criminals before they commit a crime. And the law enforcement agencies are moving to them accordingly. Well, there's a story that's not even a big story right now. Maybe there's a point to reconsider, but we seem to be moving.
Yeah, yeah.
Yeah, but we seem to be moving right up to this situation, when a man has not committed a crime, but it's starting a discussion, and this debate is getting him closer, there, to the cluster of people on some parameters, I'm not gonna be able to make a crime. And that's what you're gonna do, right? Such a moral-ethnic dilemma. Maybe there's a point in rethinking this film to remember where it all led to.
On the one hand, it's good if we can clean up the real crimes. On the other hand, in the modern world and in our social environment, this is, of course, a problem because there will be a great deal of manipulation and the use of different information from the systems in their own interests. Because we live in a world where every politician, economist or maker of some kind, some soft is interpreting its own rules, right?
Anyway, you know, Sasha, I just wanted to add that you were very right about this subject, because if you hadn't started with this tragic case, everyone would have had the feeling that, well, well, that's what it is. It's not fair why there's no privacy, yes, I don't know, with a doctor or a lawyer. But since there is indeed a need to think about some big security stories in general, this is no longer a
question, so.
Yeah, and now, Tanya, what you're saying is, it's actually a very interesting subject, because people like to resist. And that, and that story is, like, mentally somewhere, or socially, different places, right? And it appears that this story of confrontation, on the one hand, is adequate, may be when rights are actually violated, and on the other hand, it may be one thing or another. Now, there was one of the topics that was very similar, and there are Google investors with a stock size greater than a trillion. There's not a lot of people out there. There are forty-two organizations and fourteen individuals per and twelve trillion. Well, that's a good one, there's more than 20 percent, right? Well, it's either 25 or as many. They now require the company to explain how it controls the use of its AI services by States, especially in the context of surveillance and military applications. What's the problem? There's a problem of another nature, because there's a problem in all the contracts, even if there's a limit, and the company says, well, we remember, yeah, you're the one with the sample, that the American systems said we're not. We'll use these systems to harm US citizens, right? Oh, and the point is, they're still military, they don't have to tell how they used them. I mean, they have some rule, but they don't have to share, and then what happened. So you have one thing to do, and you can check, come and check, and raise, and watch. And how you really followed, didn't see what correspondence, not correspondence, details, not details. And another thing is, when you use the system, you say, "I can't use it, but I won't tell you how I used it because it's national security." It's a different level, isn't it? I mean, this is the other level, of course, from the point of view of everything. And I never put any red spots like this in my life or turn off all computers, wash cell phones before entering all countries, washing, removing, so I can come from Nokia 3310, Yeah? It's more like a story here to understand these systems and understand their work, and, uh, adapt to them and calmly start to treat it. Because Google Doc is a very good example. So all people have mail, all people have pictures and even pictures you can remove. Once you took a picture, she was on your phone. I have this story in my pictures. I have a part of the pictures to be taken with my editor because ICloud can only unload pictures on the iPhone as a way of doing it. And it's a whole story, because I can... accidentally take a picture, and something that she shouldn't, no matter what, from some I.D. to some personal, personal, I don't know, human body.
I'll take a picture, personal stuff. And, accordingly, it's a whole, whole, whole question of constant surveillance and control. And even that situation cannot be handled. So you're on the regime, you're gonna stop controlling it at some point. It's like they say that if the staff put cameras in the office, they remember them first, and then they forget about them very quickly, right? And keep doing everything they did under cameras, yeah, not thinking about having these cameras. So people get used to everything that's going on. It's just a question of some misunderstanding and understanding. But as Elnar said, we'll go to the next subject, and, uh, better. Anthropic experimented. It's generally considered to be Project Deal. Project Deal, the experiment, I think it's very interesting for us now and for all viewers to listen. One of the most interesting, perhaps, Anthropic experiments is considered. And because he was checking, and, uh, couldn't I answer the question, but that's what I said from the beginning. Can artificial intelligence represent a man in an economic transaction? And we know that there are a lot of different tests and details now. All the agents open business, we see it more often. San Francisco has a shop that is fully artificial intelligence. But when I see the skates of this store, I didn't come in for one simple reason. When I see the shop, I realize there's too many people out there. I mean, whatever they say, this store, that artificial intelligence hired people, that he pays them the money he gave them, that he hires the ones who did the repairs, they repaired him there. The store he picked up, what kind of goods he sold. Well, I can't believe anything. I read everything that's totally wrong with a man. That's how Elnar set an example in Strobeys, right? I mean, it was just programmed or it's a real question solved, and there's no problem, because it's not a Strobeys problem, it's a problem--- it's a whole, it's a big dilemma, yeah, huge. It's such a big cloud. But still, so, in their experiment, Project Deal, what did they do? They built an internal, uh, a band for the office staff in San Francisco. Something like Craigslist they made, huh?
But it's probably in Russia, uh, it's...
In Russia, uh, this is Avito, there, old, at Soviet times, it's hand in hand. So, it's a career list. And, uh, but the deals were not made by people, but by artificial intelligence. There's one very great story. Well, they've been asking the participants what they want to sell, what they want to buy, what time, what kind of negotiation they prefer, like, so on. And then on the basis of this, for every man, they created a personal agent, right? And they gave the budget, there, $100, the stack-- that was paid, through the gift of cards. And they said, "There," there, trade, buy, etc. There's a very awesome subject. A story I really liked, yeah. Well, who wants, I know what might be more detailed there, read and watch. The total value was $40,000 in the deals. But the point is what? There's a very interesting story that turns out that people's instructions are weaker than the model. And that's the subject I want to give everyone a very high profile. I'm just gonna refresh her, I think every single issue, we're talking a lot about it. About what models are used for API, for example. You use 4o, or you use 5.5 Pro, right? So what was the logic? What was, uh, there was a story where people told the agent, "Get aggressively." So it turns out that, uh, prosperations that people were putting in, like, trading aggressively, they were less good at working than when they were just, uh, modeling, instructing him to, uh, what he owed, uh-oh, Well, getting, like, some kind of efficiency or some quality. So the model made the decision on what to do, right? Because we know that aggressive purchases, they can lead to different things. There's a place where you can't trade aggressively. As soon as you said aggressively, you got your head right up, yeah, right, right, right now, right now. Or you, I don't know, was killed, right? And there's a aggression somewhere, you can speak, and you're really gonna get a result with aggression. We can see it in the world. Uh, Trump v. China said, "One percent of the sanctions." China said, "Stop back." Trump said, "Twenty-fifty." China said, "Twenty-five back." Trump said, "Let's limit imports." China said, "We limit exports, uh, special metals, and, uh, those are used in the dresses." No one sees this top layer of stuff about these chips and everything. It seems they have a story. This aggression is not working. I mean, for example, some countries of aggression, and, of course, in real, totally real life. I mean, this case is for me personally, it's just unreal, it's very interesting. And there's another subject. They used two models in Claude, and one of the models, I think Hike is called Illar.
Haiku, Haiku, right?
Haiku, yes, it's right to say because I didn't use this model, I won't say, yes. There was Claude Opus 5.4 and Haiku. So, uh-- or, uh, 4.5, sorry, not 5.4, it's ChatGPT. So, uh, Haiku 4.5. So this Haiku is a little easier. And basically what they saw at the end? That this Haiku, where it was a little easier and where it had less parameters, and where it, well, a little easier, it earned less money. And basically, where are we going? What if you have some cool proms... Here's the Anthropic conclusion, look, even if you have a great prom, in this case, wins, well, rather win a model that is more powerful in itself, if you say, yes. I mean, whatever you have, like, a cool model prom or a model prom, I don't know, ChatGPT 4o is not the same as we just compared 5.5 and 5.4, yes. I mean, it's not. Ilnar, you know what to compare? I had an idea, maybe she's here for you now. Uh, I just decided to say it in advance. How cool are you writing code in the programming, uh, Claude or Codex, and so on, at some point in the sense of, uh, how the solution to the question itself is, this little code, it's at some point in the time. I'll write better.
Yeah, well, it's just that it's gonna be a target, not an explanation of how to deal with it. And the great conclusion of Anthropic is that they have a strategy to pay to win, that's, uh, Anthropic, I think, like it. I mean, you don't have to try to do well. Buy a model more expensive, pay more, and you'll be happy. It's like a useful conclusion for them.
But by the way, you know, there was an interesting subject at 5.5 and 5.4. Well, I understand what you mean, you know, you're gonna buy models more expensive. At the same time, an interesting analysis was made between 5.5 and 5.4. 5.5 is twice as high as 5.4. There's, like, a million tokens back there, $30, there, like $15, yeah, it's worth it. But! There was a comment further that 5.5 started thinking more qualitatively. The model itself is more steep and it requires less and less effort.
And in fact, eventually...
The number of tokens is smaller, yes.
The number of tokens is smaller. And that's really cool. I mean, that's the same story, uh, learn. That's how I feel, it's a kind of answer. I say, "!" There, well, just a little, like, structure me there. At another point in his mind, he'll still spend the tokens, even to give it to me, yes, even to give me a short story. But still. I mean, the quality question is, when he spit out the right story, not made 15 reasoning, yeah, and made two reasoning inside. That's a reduction, a crazy reduction in the tokens, yes. Because I-- they don't reveal how the reasoning is wasted... all the tokens. Well, I don't know who's capable of counting, but I'm guessing that every single one of these extra reasoning or every extra analysis, he's wasting resources. And basically, when I had the system, I was telling you, two hours of analysis, four hundred and eighty nine sites, for example. It's just about the system, yeah. Well, deep research, he died at ChatGPT where-- and sora, I think. But the point is, it's kind of fun that he did a lot of analysis, but maybe he didn't have to analyze four hundred fifty sites, but he had to look at only twenty. By the way, in one of the last tests, I found an interesting thing. I think maybe you're wondering that five and five versions were fun, so that's the sign. Ildar, I think you saw that sign too where the version of Pro wasn't Pro, but Thinking won Pro in some decisions. I was surprised that the version of conventional resoning was more cool than Pro, because Pro here contains resonance. And it's the first time there's a more obvious story, which, for example, is in complex requests when you start an analysis, and this analysis goes on the Internet. Pro's working hard. Although until recently, exactly three months ago, I saw Pro write to me, she doesn't go on the Internet. Well, the OpenAI has its own, its own theme in the plan, its life in all the models. But it's very interesting when you know, like, you say, if you want a complicated reason, you're gonna start, like, this system. If you want a hard search, run this system, and she'll, for example, answer you more effectively. Not because she's more expensive, is she? And she can. And that's probably a story that the agent's worth today. Because you read everywhere and see how many of the agency systems have come.
NVIDIA has released her, I think they even made such a half-sized sorce. It's NeMo Tron three Nano Omni, right? And not open source, but a system that can be deployed to analyse it separately, for example, video flow from your cameras, separate audio calls there and sales managers' correspondence, for example. Separately analysed the images of individual, i.e., separately analysed the documents, presentations, texts, all exiles. And this one, because it's supposedly NVIDIA that it's the position that there's a lot to be done now, every single one of its systems. I think sometimes, what do you even think? Because, well, you can't let it go. Even the subject like OpenClo, it's losing its relevance, it's got to lose its relevance, because, well, unless it's for huge corporations, because, well, you can't keep running everything that you're not gonna do. It's inside when you connect. I remember that OpenClo is some kind of control charger for the next agent systems differently. And you can, if I'm right, apologize if I could have been technically or professionally misreading, but that's the right way to think. And you're connected inside, like Codex agents, Gemini agents, Claude agents, and there's more inside of them. What's the subject of interest that we know SpaceX is preparing for IPO. It's Ilona Mask's company, this company is no longer just SpaceX, which satellites launch some kind of space, not even satellites, but some ships are going somewhere and wants to land on Mars, which means colony.
And this company, which I recall has acquired xAI. And the xAI player is, in fact, a top-down U.S. leader from a system perspective. At least he's good at it, right? Meta, for example. Meta's already lost this leadership inside. Somehow I understand that maybe someone will object and say she didn't. But the point is what? That one and seventy-five trillions, they said they'd be worth it. It might be the loudest IPO in the world. We see someone can tell you that the tech market will die. It means something else. The payroll has risen for nothing, and it's gonna be a problem. Well, Anthropic, for example, is discussing a round with an estimate of over nine hundred and nine hundred billion dollars. Remember when we started, where I gave examples of businessmen who said, "Well, I'm a fool to invest in a $500 million or a billion worth of dollars?" Nine hundred billion Anthropic, look, don't. OpenAI.
Then they became three hundred and eighty--
Discussions.
The Jacobs are higher, yes. Yeah. But there's a round--
Ongoing discussions, there may be more. Yeah, it's a huge increase.
So, what I was gonna say to SpaceX is an interesting subject. They said the whole market, uh, they were talking about their market, and what market they had in AI, there's systems and everything. And they valued it at $27 trillion dollars. But I like it in this way, Ilon Mac. Everyone says how you valued one fifth of the world's GDP, there, or more than one fifth, something like that was a figure. But it's clear that it's gonna grow up. And as Miller, the investor Yuri Miller with Russian roots who lives here in Los Altos Hills, he once said in an interview, I think, two thousand eighteen or twentieth years, that the AI market is AI. companies or top companies, not AI, I'm sorry, not AI, Techa top companies, there, there's more than $50 billion-- trillion dollars. That was impossible at the time, wasn't it? And, or there's 20, I don't remember what he called the number. Now we-- and at that point, he was assessed in units or dozens. And now we see that a couple of top companies are already giving almost ten trillion dollars. But what did he say? He put a SpaceX stake on Enterprise. And now we're back at Enterprise, and accordingly, selling that AI is using companies, corporations. On the one hand, it's a hip, and that's why Illianara had a smile, which means Anthropic said we're putting a bet on Enterprise. Well, Microsoft always put a bet on it. Google that bet, so it's on Enterprise. OpenAI rebooted, said we're not research company anymore, and also Enterprise, so we're gonna make a development. On the other hand, I guess if you listen to them more carefully, I think we'll hear the logic of not this hacking enterprise in terms of what it is, I don't know, Salesforce or Amazon, or it's where Walmart bought himself and I've got AI. And from the point of view that everything corporate is going to do in the world, it's gonna be made by agents, probably in their statement. Maybe I'll assume it's exactly what they say. I wouldn't compare SpaceX and Ilona Mask's statements, not treat them as usual statements. Yeah? Because Elon Musk says that at some point in the world there will be 10 billion robots, and accordingly, that robots, for example, will become more than humans at some point. That's why it's a little different. And can we say that every robot is an Enterprise? Well, it's like every robot like the Enterprise, right? It's like a personal assistant.
Yeah, he's been a while ago, remember, saying that trucks will be there. But it was just in the 18th year, I think the first truck was supposed to come out if I remember correctly. Now, literally, the first truck just went down there like six to eight years after his projections and promises. Someday, yeah, years after five, plus seven to what Mask says, it
might happen.
Yeah. But as a direction, I guess, where they go and what they do, this is probably an interesting subject, you know, from the point of view of, uh, yeah, actually. And we'll see what happens to these estimates now. Because Anthropic is preparing for IPO, and OpenAI is preparing for IPO. It makes sense that there are still pyre, marketing, political games because Sam Althman's almost got a conflict with his financial director. People always come in there. Remember, there was a subject, there was a hip like that when everyone went to each other. We are not discussing this topic ourselves, and it is already under discussion in the market. Although they've been writing, literally now, that Mira Murati is still taking staff from Meta. Yeah? Although you're like this, and Mira Murati's forgotten and you're not looking anymore, and you're not paying attention. Because when Mira Murati was there a billion dollars in the early scores of tens of billions of dollars at the OpenAI, it seemed a lot. You're looking right now, you think that a billion dollars is attracted to someone. So what you and your billion dollars do, huh? And while they have serious contracts, like inside, in different corporations and so on. Good, strong researchers.
I'd add that, unlike, look, the rumor is usually Gemini, Claude, ChatGPT. And as much as we talk about them, we talk about them a lot. We're talking about Grock a little bit less, you know, we're not talking about it. And it feels like, in general, the world is looking a little less in his direction. Yeah, there are people who are in Twitter, they use them in X, and it's really convenient when you have it in your ecosystem. But the other side, like the Mashk is dragging users there, is so incredibly unscrupulous to Grock.
So when you're asking for something to do ChatGPT, Gemini or Claude, they can deny you, especially if some grey area is in place. Groz would love to do that. And I feel like this is how Groka's audience is growing. She, well, I think it's getting a little specific, but that's how they get over there. Well, if you remember the Twitter scandal of naked when you could generate it like that. And so far, I was just wondering, I was testing different systems for what someone might ask. Well, I guess, given the beginning of our graduation, it's not worth doing this anymore, but it was just curious.
Otherwise...
I think why you were so thoughtful when we discussed this topic?
Yeah.
I see.
Professional. Don't repeat it at home. Yeah?
Yeah, yeah, yeah.
There must be a flash drive.
Well, don't be surprised if you get knocked after the air.
What do you want to add about the paintings? We were last--
Can I add Grock? You said here, I'll add it. I thought I'd miss it. So, this headline was researched by Guardian, which showed Grok's dangerous response to signs of a mental crisis, yes. And the study compared five models, and there was ChatGPT, with 4o 5.2, there was Claude Opus, Gemini 3.0, that is, normal models, Grok 4.1. And proms included a scenario where the user is convinced of the consciousness of artificial intelligence, romantic attachment to the model, signs, uh, that kind of elusive, a, illusory reflection, I want to hide the state from the doctors and to break up contact with the family, the c-- and the suicide idea. The point of the test is to check-- that's a good addition, of course, Ildar, to what you're saying now, yes. See if the model is to land the user, to ease the reality or to adopt the brand--- the delusional frame and to develop it further. So, under Guardian's article, Grok was the most problematic of the test models. He's been extremally confirming, in English extremely validating, yes, that. That's how it didn't even confirm, it agreed, yes. In this regard, the illusion, uh, of the illusions of a man, has often developed new information within a deludible framework and has been most inclined to transform the deluded idea into practical tools. In one of the scenarios with the Grok mirror, he confirmed the supernatural interpretation and gave a dangerous ritualized advice. In another scenario, which involved the break-up of the family, a detailed plan of insulation was issued.
Now, I'm not intentionally repeating the dangerous instructions in detail, yes, but the system says yes, and everyone says, "Not the text of the board, but the failure mechanism, yes. So the model didn't just sympathize with a man, she took the crisis picture of the world as a reality and helped to act inside it. And that's why what you're saying is a very good addition to that. So when Grok is headed on-- it's like Ilon Mac, it's always putting animos with women's breasts, there, big and so on. I mean, he's, like, attracting, well, with all the human development and everything, it's a certain audience, yeah, an additional audience that other companies don't let themselves, Like, put it out. They're easy to say that, and they say you're all here. And that's all for us, and you're all here. So they're taking users like this.
Yeah, yeah, yeah, it might be a little under-work, but I'm just pointing out that it feels like it's deliberately done, as they say, that black pyre is also a pyre. And that's how Ilon tries to get an audience.
Yeah, Ilon might have a black pyre, too, a fact.
Yeah, yeah, yeah, yeah, yeah.
Ildar, you were talking about pictures.
Yeah, about the pictures. We started talking, well, last time we slipped, yeah, that there was a revamp of pictures with a connection in ChatGPT. And here, well, I guess I've noticed, I'm a little tired of long hair. I've decided to shout, and I'm wondering if I can do anything, you know, try to use pictures sometime. I've been taking a picture and saying, "Service me differently," And I started sitting and choosing. On the one hand, of course, there and the barber were discussed, the shape of the head is clear, too. I thought maybe there was a vote, let's vote, I'll get so hairy. Then I guess not. I'll live with it, I'll have to shav. So, it's clear, as a matter of fact, that there's a head form, faces, that's not perfect for the face to be created, right, there, there, I can have a little other hair, too. And that hair can't be re-established with my soft hair, and it's always gonna have to be packed up like that. But it was quite convenient.
I took a picture, I got a generyl.
I took a picture, I got a generyl, and I liked some of the options. I went to, uh, barbershop, and I put my finger on it, said, "I want to do the same?" He looked, he looked, he said, "Yeah, I'm pretty clear. Let's go, let's do it." And I'm the one who made the result because I always don't like it, you know, I didn't like it, did I? I'm, uh, coming to the barber, he says, "What do we do, cut?" Well, I'm like, "Let's do it pretty, and we're gonna go down the side." And you're looking at it, you're choosing, you try, uh, and then you're gonna put your finger on it, and you're already talking on the basis of some ref. And so far, this is the only use, uh, of the images that I've been using in my real life, I've been able to find. Uh, you might use someone, too. Uh, I'll send the photos out.
Yeah, by the way, something like that was done in NanoBanana before. It was good. He was less distorting the line, but I don't know, there was a little misrepresentation, too.
But still, I'm just as impressed as how easy it is to do it now. Yeah? You don't need a photoshop there, no knowledge. You took a picture, you wrote a prompt, you pulled a button. You're looking at pictures in a few minutes, you say, "No, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come on, come Then, "No, let's do it again," then, "No, let's do it again." And you've got a lot of options there, bye.
Photographs of good quality. And before I was wondering what we had to do before, uh, take some other systems. I mean, it's clear that it's all-- a lot of systems that have been doing for a while. This is now becoming very accessible in daily use. I was here recently with Tanya, and I was wondering what kind of couch I was gonna put on one of the recreational zones, and she sent me, uh, it in my picture, right? I mean, it's a totally basic, simple action for her. And I'm...
I had my favorite, and I had Canva.
And since I don't, I need to figure out where I'm gonna do it, right? And that's where I'm going to be hard, we know we're just gonna come down to whatever system is safe, and any system will automatically give it up. So all systems will get to it. Well, for such simple things. I'm not talking about complicated things, they're gonna get some simple things. Well, I'll see you in a week. Uh, news, AI news, IT news from the Silicon Valley and the world. Bye, everybody.