Skip to content
Transcript

Transcript · 097 · AI Agents Are Becoming a Systemic Force: One Error Now Travels Through Code, Money, and the Physical World — ToTheMoon

English machine translation of the Russian-language episode. Timecodes open the source video.

Episode overview
00:00:00–00:00:30To TheMoon tonight.
Alexander Volchek00:00:00

The United States and China do not sign a declaration on artificial intelligence in the military sphere. Anthropic warns the risk of chemical weapons. Apple is delaying the AI launch. Perplexity launches Model Console. Gemini is over seven hundred and fifty million monthly users. xAI reorganized after the merger with SpaceX and many others. Hello, everybody! We're on ToTheMoon. Technological news, Silicon Valley sites around the world.

Mentions: Apple · Perplexity · Gemini · OpenAI
00:00:30–00:04:11What's new in the Codex is part 1/2
Alexander Volchek00:00:30

You, Ilnar, were just talking about Sam Althman's, uh, code-based pedith, and we're gonna see him, right? And that because they started to cooperate, uh, now, with Celebra, that they're gonna have, uh, expanded power. What kind of thing I wanted to say. I've been throwing it away. I see it as long as it's only in Pro's case, I've got a few days ago. I was able to launch parallel flows in Pro. They had this story once because parallel flows clearly give a crazy load. Well, that's got to be some huge load. My partners in one thing said, "You're in the same business, I think, destroying the OpenAI, they're gonna block you up soon, so the number of queries you're wearing chat, right?" Pro's questioning, when he answers, there's eighty-nine minutes. Aah, what did I see? That I was sitting there, dealing with one ana-- one analysis, that's the longest one that's usually done, there's, like, 40 minutes, sixty. And I accidentally pressed... Ah! I made an appeal. I, I, uh, sent a chat-up on a request, realized he'd be doing an hour, and I came in and wrote "Apdadet." There's a button of "opdate" for users who don't know, right? You can click the peddate and do-- how to supplement, uh, the information you've forgotten. It works, well, in--in-the-white-white, but something's happening. She, by the way, worked perfectly at that point of time. And then I wrote a new request in the line and thought I'd send it when this chat ended, and accidentally pressed Enter, and I had a parallel flow. And I'm like, "Oh! There's a parallel flow. "And I'm further looking at this document, and I've created another 10 parallel flows. And he's, like, working in one chat... That's a very interesting subject. Why am I talking about her? Because at the beginning, I remember, earlier when it was, it seemed a little stupid, and I realized it would be a strange answer to come and everything. But here, in the era, the models are going to such a level as to, uh, uh, uh, they do, and, uh, he's eighty minutes, he's writing a code, you know, to get it, right? When he's processing my papers, he's writing the code, and it's for Python to write the code, start, compil, perform, analyse the vast amount of documents, and he's counting, he's looking for something, he's looking for, It's analyzed. I mean, it's some kind of incredible number of requests. And you realize that these parallel flows are coming to the system, like, where you don't go into the different windows, these things that I've got, you know, chatting these huge numbers, parallel chat, right? And when you can spend your life in a single chat. And here I did in that window, in this analysis, I tried to run a test in one window, and differently, as it was in the-- the shortest part of this subject, but what I would normally do in different chat rooms. I'm still confused about how he does it, what kind of power should it be? Is he creating new flows or is he doing this at the same time? I have a big, big question, right? Because when he gave me answers, how did he answer? He's giving a reply. I have a p-- you're gonna be wondering if he's the one who's answering? He's not the only one who's answering. He gave a reply, and the old flows were still hanging.

00:04:11–00:05:52What's new in Codex is part 2/2
Alexander Volchek00:04:11

Then he started closing them, too, and he started giving them answers.

Ilnar Shafigullin00:04:11

How's the sequence of the answers? As you're ready?

Alexander Volchek00:04:15

I don't understand. I never understood because I-- I can't say because I've loaded them too much. I understand your question because he could have collected it all the time, and he could have been in the kind of request. If the request is fast, he's responding fast, right? I mean, here. No, I'm not gonna answer that because I, I did it the other day, and yesterday, too, but because I did it in a complicated, complicated system of inquiries, they've been working for a long time. I mean, I didn't track it, so, I didn't see it. The only thing I'm saying is that at some point in time, the analysis seems to have broken. So at some point in time, the quality of the answers has become worse. But we see, yes, that ChatGPT is running a huge number of tests. I had a sister here, she opens an app, and she has dates marked. She says, "Oh, look, the dates are finally here! "The dates of the teased." I'm looking at my place. I don't. I say, "Look, I have an American account, and I don't. How do you have that?" She says, "I have one." Cup, she's in ten minutes, she doesn't have dates.

Ilnar Shafigullin00:05:12

Alexander Volchek00:05:13

She didn't update anything. Well, that's clear. There's a part of the web interface down there. She's got, like, a test sample, right? I mean, they understand that a lot of things are being tested and, perhaps, there's no way in thinking about this regime. I mean, I only had this regime in Pro to just know. I mean, I didn't have the chance to press the button, I had the opportunity to push just a stop. I remember that they tested it again, but it just c-- it looked like, you know, different quality, right? Who else has been in touch with this, write it. These are the parallels. Very interesting. You want to say something? Perplexity has begun...

Mentions: Perplexity
00:05:52–00:08:15Perplexity Model Council: what is and for which
Alexander Volchek00:05:52

It's a little different story, but still, right? What do you think of her? The concept itself. Perplexity has launched what is called Model Console. So what does Model Console mean? It's a model committee regime. The same request is launched immediately in three selected LLM, and then a separate synthesis model collects the results in a single response, but clearly shows where the models matched, where the different and that each I added a unique one. I'm not Perplexity, I'll tell you right away, yes, I'm not. There are amateurs, I know, very big. And, uh, and, uh, while there's a fund where I was, there was a cash investment, everything, and there's good, cool figures. But once again, I don't take Perplexity as a company. And I think I'm all thinking who's gonna buy her, and if she doesn't, that name will disappear. But from the point of view of this idea, our viewers, write when you're alone-- I think there's a lot of different applications now, too, on the Internet, they've already been mobile applications, That's what enabled different models to use. And you have the same question in different models, and some of you get the outcome. Ah, Ilnar, what do you think-- Ilnar, Tanya, this question? Because it's so very, um, interesting to me, but at the same time it's not clear.

Ilnar Shafigullin00:07:11

Carpathians, Andrei, collected this story a couple months ago. It's a pretty busy story, but it's a bit of a model-comparison. You've got a few models starting to work, you watch them interact, something else. So research looks interesting. To compare models there to understand where weak and strong players who contribute, there, to the outcome and so on. In terms of specific tasks, uh, I don't really see much of this in some big profile so I don't know what four models do in parallel, then something else happened, like, They're plus or minus equal. If, for example, I know Gemini is better here, I'm just gonna go to Gemini and put it in. Something else. I mean, I wouldn't stick, there are three, four, five models parallel to real tasks. To study, play interesting. Right? Well, I'm curious, because Carpathians put it out a few months ago, and now in Perplexity, it's taken as a function.

Mentions: Gemini · Perplexity
00:08:15–00:10:32Perplexity Model Council: what is and for what
Ilnar Shafigullin00:08:15

It's a little confession that the idea was a little busy.

Alexander Volchek00:08:16

Well, it's working...

Ilnar Shafigullin00:08:16

But I don't see life in that.

Alexander Volchek00:08:18

Look, but it's working out, in some of the requests, it might be interesting where you'd like to get an alternative opinion. We know that models are different from the point of view. I was just saying, uh, situation that, for example, was about how the models were answering about Jews, and, like, they're responding differently. Or as Mo-- as models answer about the policies of a given country. They're not answering differently. And, accordingly, for these briefcases, Ilnar, this is a whole good subject to get you an allegedly independent subjective opinion. The question is, you're choosing models of a wide range, there, I don't know, from Chinese to American, right?

Ilnar Shafigullin00:08:55

I'm a little closer to the story, still with the programming, with the design. And there's a lot of people in the agent environment, doing some kind of next. With ChatGPT, imputed, there, 5.2 Pro, there, 5.3 Codex or something, they're producing a plan that will be performed, and the code is already left for Gemini, for, I don't know, Claude, for another model, perhaps cheaper or something. I mean, well, in society, it sounds like, uh, a kind of opinion that breaking the task on the part, making a plan of development and something is better doing GPT. And, accordingly, the code is written, so it's better handled, for example, by Gemini. And in this sense, uh, it looks more sensible not to throw the same challenge into two models at once. And if you know about what phases the model is best, then you'll give it away. Yeah, that's what's done in the agency approach. A lot of people do that. Part of the tasks you know that the model is more powerful than one, then actually, you can move automatically or not automatically. That's a more sane approach, I think.

Mentions: Gemini
Alexander Volchek00:09:58

Yeah. Perplexity, by the way, did-

Mentions: Perplexity

Yeah. M- Perplexity, by the way, does it at Claude Opus, 4.6, GPT-5.2 and Gemini, uh, 3 Pro, yeah, just so that. In general, a good set, but again, mine., the only thing that's history, uh, before-- how to supplement, yeah, it's even useful that different models have different blind zones. And when the model is definitely wrong, well, and since they can steer each other, there's a little bit of a plus. Again, what model would they be collecting the final result? That's interesting, by the way, yes.

Mentions: Perplexity · Gemini
00:10:32–00:11:54Comparison of models and the " Emergency forest " approach

I mean, if you got three models, you're gonna have to give it a final result. What model is the outcome?

Ilnar Shafigullin00:10:37

There's a classic machine learning, uh, that, uh, that's a little bit of a headphone right now. There is, uh, this approach to the task, called random forest. What's he up to? You have a set of models, uh-e, that give a final answer, and then, uh, among the things you did, you're voting. And there's a mathematical reason that the answer dispersion, that is, the difference in the way the model is, in a way that is, it's actually wrong, it's very much reduced and reduced in proportion to the number of models you have. There's a place inside. In fact, it is an attempt to achieve a similar result, not on the simplest trees there, for decision-making, but on the big LLM-ca respectively. If LLM-cas differs not always the same, you know, they're, like, linearly independent, so, say, yes. If there's a little talk in terms, then, uh, the dispersion of the answers, their dispersal, it really needs to be reduced. And that-- that's a mathematical rationale. Here. So if that's the point, there's probably something about it. But I don't feel like this is the future.

Alexander Volchek00:11:45

Sca-- looks like some sort of Mar-Marking story, right? So, uh, b- that's the subject I think the subject is very interesting for passing.

00:11:54–00:13:32The United States and China have refused to sign a declaration on military AI
Alexander Volchek00:11:54

There was, uh, a summit, initiatives and responsibilities of artificial intelligence in the military sphere. And, uh, they were the initiative at the summit, uh, REAIM. Both the United States and China, did not accede to the joint declaration supported by other participants. So, uh, that's--

Ilnar Shafigullin00:12:16

And you did, Sas, what players I'm not so developed, I understand, right?

Alexander Volchek00:12:20

Well, who else has an AI, except for these two countries, huh? I mean, of course, uh, why is that news, why is this topic even important? Still, well, everyone has to understand perfectly that the USA-US and China are interested in using artificial intelligence, well, in serious things, strictly for themselves. So we now see the unbelievablely strong introduction of artificial intelligence in Chinese State structures. There's a lot of de--- services that use artificial intelligence. They even put it up like a case. We know that all the top American companies are working on the development of the artificial intelligence of Anthropic and, uh, Gemini, and OpenAI, and they all cooperate, uh, they've given these systems, there, for the replicas, for the impolite. Free to the State and gave to the Government of the United States, and, in particular, for military deployment. I mean, obviously, history, uh, future history, it's very serious. And here, uh, there's, uh, uh, uh, I don't know, everybody knows or doesn't know. Uh, a few literals, um, days ago, yeah, that's, uh, a week ago, uh, SpaceX.

Mentions: Gemini · OpenAI
00:13:32–00:15:19XAI and SpaceX mergers - what does that mean?
Alexander Volchek00:13:32

Uh, remember, we were telling you that early January, that Tesla had set itself the focus of creating, uh, robots, that it was important not cars, but robots. And there's a little time going on, and, uh, SpaceX, which, um, launch-- like, uh, Ilona Mask's doing all space, and, uh, it's all, uh, it's all, it's all, it's all, it's all, uh, it's all, it's all happening. The merger is planned, first, that it will be, uh, super expensive, there, IPO and the oldest-- dearest start-up in the world now. But at least, according to current estimates, he even hits the OpenAI. Although I think the OpenAI will beat, uh, this organization. Obviously, these companies will be worth, they'll be worth tens of trillion dollars. The question of the osp-- is such a time-bound. But what's the point? What, look, there's a hi-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i-i- And when you see the military industry, you know, even now, Ilona Mask's satellites, they, uh, they create whole corridors, yeah. They have projects there. For example, tracking the borders of a given country. I think Mongolia ordered it. There are whole projects in there that are big when you have a soft in your satellites that is aimed at certain actions. And, of course, that countries, within the framework of, uh-a, refuse, well, publicly refuse and sign various military initiatives, even in words, this is not even discussed. This is a good example. I wanted to show him.

00:15:19–00:18:24Anthropic on the risks of chemical weapons
Alexander Volchek00:15:19

And then he's giving Anthropic risk warning, a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-bone, including in the chemical. And they, uh-oh, publicly accented, they did it, uh, about our shoot yesterday. They publicly emphasized that the growth of the model ' s capabilities is not only beneficial, but also risk. Accordingly, the chemical threat system has been explicitly emphasized separately that models can reduce the threshold of entry into dangerous receptor, synthesis and action plans, and so on. Interesting thing was when the OpenAI yesterday in Instagram gave a picture of robots moving. And so I remembered my youth, uh, 20 years ago, when I was doing exactly that robot sophth alone. Well, it's very similar, something's moving, yeah. We've mostly moved silicon plates, rivers. Well, that's all the things used in the processor process, yes, and including the plasma panels, different. And--and this company is in the middle of biological things? Well, I mean, biological tests do things. I mean, if you look at it, the company is involved in various biological or chemical development. And, of course, you realize that these systems are for a normal person, they're not gonna-- they're not gonna give this information, but for the state and for the intelligence services, they're gonna be able to count anything. I wonder, Elnar, you've never encountered, no, with any description of how the OpenAI is written for public companies doesn't limit the data on how to create chemical weapons? Or there's Gemini. That's not what you're saying, right? There's no such thing.

Ilnar Shafigullin00:17:02

Nobody talks about it. Yeah.

Alexander Volchek00:17:04

But the idea is, it should have been written, right? That you could have put it on. There, I don't know, the U.S. Congress will take and present. Well, we'll see the number of new public cases at the U.S. Congress, in two and a half years or, there, the courts, there, in the U.S. Now, they're going to the current authority, there in two and a half years, and I think they'll be their own courts. Here. The U.S. is in high court, yes. There's a city mayor who just says, "I don't care. I filed a 40-three-time claim for Trampa, 40-two wins, one technically lost. Take every law, we'll win." Here. Okay, well, if they don't write it, it's probably inside the system, but inside it is deep. So we have to understand that we are not making this data available. So we're not allowed to have a specific content and a certain calculation. Both, in the opinion of the OpenAI or the Google and in the view of their staff, who believe that the content should not be accessible. How do you limit it? I mean, it's got to be, uh, programmers that's all down, engineers, yeah, strong engineers are limiting it. How do they really limit it? And who's checking it out, and what are the side-checking models? Well, you know, yes, there are separate companies, separate programs that do these checks.

Mentions: OpenAI · Google
00:18:24–00:21:34AI works worse in Europe?
Alexander Volchek00:18:24

Yes, Tanya?

Tatyana Tsvetkova00:18:26

I have a question for you two, as a matter of fact, as for people who are more aware. Maybe that, well, in keeping with this subject, maybe that in different countries, I'm allowed to be a ChatGPT user to work differently? Or is that my self-programming? Because I have a real feeling that he works much worse in Europe. Is that right? Can I explain that?

Alexander Volchek00:18:56

Tanya, maybe 100 percent. Yeah, maybe it's just because they have their laws and-- in Europe. In Europe, however, serious personal data requirements are not sufficient. And for example, my sister says, "You were in the graduation saying that last year you could have seen the results of the twenty-fifth year." She says, "I couldn't see the results of the twenty-fifth year 'cause I didn't have that button in Europe."

Tatyana Tsvetkova00:19:20

Interesting.

Alexander Volchek00:19:21

Still, in Europe, working with personal dan-- well, maybe in some countries she didn't. In Europe, personal data work is very serious. And plus there's a bunch of other laws. Obviously, they'll be limited, Tanya. It's like Australia, for example, not social networks using Australians.

Tatyana Tsvetkova00:19:36

No matter what your account is, it's important where you're territorial.

Alexander Volchek00:19:40

Look, I came when I was, like, me, I was in Paris the last time I was in Paris when I was in? Couple months ago, yes. I'm out of my access, sora's gone. I remember very clearly what I wrote-- with an American account, yes. I mean, I just moved to other IP addresses. I see. I guess we could have picked up a VPN, change it.

Alexander Volchek00:20:00

There's a VPN up there, changing. It's not like a sim card to come in so they don't see what kind of cellphone is. Well, there's a lot of stuff. But he's still determined. Look, we live in the world, again, look, there's Instagram. I'm explaining it to everyone. Answer me the question is very simple. There's Instagram. Instagram changes the location of the account automatically if you're in a different place for a few months. Google also changes. So you can't change your own location. If you go into Russian bloggers, they all have Russia. They can't come into Instagram from Russia, but they can only do it with VPN. They should have the location of Germany, Hungary, Holland, anything, Indonesia, Thailand, but not Russia. They're in Russia. Look, it's very important, isn't it? It's an important design that, well, everyone should take. Just like youTube people come in on VPN from Russia, and many people don't show up. And there's no advertising coming up? Because youTube doesn't show a commercial on Russia. Here. There are therefore many methods of location, location today. And, Tanya, of course, ani-models. You've raised a very cool question. I'm-- I'm really cool. The public will ask us, as you understand it, or as you can see it, what is different? Maybe someone thinks it's the same thing. I just personally, of course, think the systems will be completely different. Tanya, you said a very cool subject.

Mentions: Google · Russia
00:21:34–00:23:40AI works worse in Europe?
Alexander Volchek00:21:34

Oh, that's definitely the question.

I was just pissed off about how he answered me. I've already done it. I'm like I'm back, I don't know how many, a year, two years ago. I think why is he answering me so stupidly? I even checked, maybe I'm log out, and I have some free version of it working on the left. I think we should ask the kids what they're gonna tell me about it. Maybe it's my self-programming or something.

Ilnar Shafigullin00:22:01

Technically, it's very easy to do. I mean, from a banal substitution of a system prom. I mean, when you write a message, among other things, you have a system message in your model that describes what additional information for the model, how to answer it, and so on. And that's where the banal substitution of this system message is based on some additional data, until some regions are just starting another model. Well, I mean, there's a lot of models running parallel, and your message just goes either one or the other, based on some parameters that are. So technically, it's very easy to do. It's kind of like-- and, yes, probably based on the laws and other restrictions on the extra countries that are being put on or something, that's exactly what you can do.

Alexander Volchek00:22:45

Well, I'm curious, because I didn't ask for anything that was legally restricted. Pretty banal, simple.

Alexander Volchek00:22:54

Thanya, by the way, can allocate other resources.

Mix00:22:56

Yeah.

Alexander Volchek00:22:56

Tanya, can just allocate even other resources to find people in a certain location. It's really-- thank you for being really that way. And one-- one-- one-- one is very odd, because yet you-- they're gonna have to create a policy of understanding the location of a man, like Google did, yes, or how it was done in Instagram. You'll have to, they'll have to do it because people are moving around the world, and they need to know exactly what this man really lives, like in America, using American stockings, and so on. Or it's a man, and he travels where you were, in France, in Italy. Or this man, he's Italian, yes, or this man is Italian, and he's just trying to buy an American account.

Mentions: Google · OpenAI
00:23:40–00:25:20Update Deep Research and fall in response quality
Alexander Volchek00:23:40

I mean, we have to figure it out. By the way, we have to understand that OpenAI, Tanya, produces a huge number of peddates, sometimes very strange. They've got a new pddette, they're a Deep Research button, Deep Research. I have a Deep Research button in my own right now, and I have to look into it, come in, open in more, you know, and I got it. And above, I've got a few applications that I don't use. You could have chosen a model you're working for Deep Research. We can't pick models now. But I'm telling you honestly, I didn't really understand Deep Research, the model you're choosing, or is she default chosen? Based on the Instagram Sam Altman post, the reassignment of Research OpenAI Isa Fulford. She said, "We're super-renewed. Now Deep Research works 5.2 versions." I understand correctly, Isa Fulford, that until February 10th, we used Deep Research, which wasn't 5.2?

Ilnar Shafigullin00:24:42

At some point, he started working very fast. Sasha, I'm not surprised if they're just, uh, with this transition to Gen AI, and so on.

Alexander Volchek00:24:48

There's this pause, there's a Start button, and it's got some shit, by the way. I don't know how fast you got him-- you're right, by the way, that's when you remember that subject, I don't remember, six months ago or when you were? I can say that I've been active in Deep Research for the past two weeks, and he's got a new one, but the quality of the answers is so bad compared to the deep heavy heavy, like I have heavy reasoning, yeah, heavy thinking, even extended thinking. I mean, well, that's an expanded thought, right?

Mentions: OpenAI · Apple · Siri
00:25:20–00:27:32Apple delayes Siri updates
Alexander Volchek00:25:20

Well, the quality of the answers is just very bad, and I have to reset. Well, at least I have that impression. Write your impression on Deep Research OpenAI. So, what are we all talking about OpenAI? Let's talk about Apple. So, of course, I like the news that Apple is delaying the launch of Siri's AI-news. To be honest, it's not funny anymore. Here. It is reported that Apple is moving the timelines for major updates Siri, including the promised staff. We all know that, well, they have, first of all, an IPhone interface, there, AirPods and CarPlay. You're looking at this, you don't know what's going on in Apple with artificial intelligence. It feels like they're lost in artificial intelligence. I'm still waiting for a new monitor, but if they don't let me out in March, I don't think I'll be waiting for him anymore. Yeah, I-- but in artificial intelligence, they're lost. The only thing they have is that, I think it should be brought to the attention. They bought a Q to-- Q dot AI, specializing in audio, uh, artificial intelligence and voice processing, right? And it's basically MOWA, but Apple's not like that, and Apple doesn't have that much of a deal, does it? And that's what it's like, the voice's getting clearly key in terms of, uh, directions. This is a very important topic for them. And, uh, they're probably afraid, uh, that they're gonna be giving up-- that's not what they're afraid of, uh, Jonathan Eve with Sam Altman, is it? Jonathan Ive is an ex-, uh, super-crunch Apple designer. And I think that story is connected to what they understand is that the devices, the audio devices with the ability to keep this recording forever, we've talked about it a lot, too, and they're, uh, gonna be on the front. Let's see what they're gonna say. Ilnar, you're not gonna ask for Apple. You're gonna laugh. Uh, maybe I think we should go--

Mix00:27:13

Older laughing.

Ilnar Shafigullin00:27:13

They have great MacBokes. That's what I can say.

Mix00:27:17

Alexander Volchek00:27:18

Yeah. No, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no,

Mix00:27:19

Interesting.

Alexander Volchek00:27:20

I like it, by the way, yeah.

Mix00:27:21

Comment.

Ilnar Shafigullin00:27:22

They have great MacBokes.

Mix00:27:24

IPad, no?

Alexander Volchek00:27:25

No, I'm actually doing everything for Apple, so I have nothing to say.

Mentions: Apple
Mix00:27:28

Yeah.

Alexander Volchek00:27:28

I even have Mac Studio.

Alexander Volchek00:27:31

Because they're all connected.

00:27:32–00:29:20AI-agent Codex vs Claude: what's better?
Alexander Volchek00:27:32

You can't take different products.

Alexander Volchek00:27:36

About-- I want to remember Codex quickly. Yeah, now that's the move, in the view that, uh, OpenAI has started Codex on Mac. Again, who tried it, we asked about it last time, but now it's been a long time. Which one of you tried who didn't try, but, Codex on Mac, who did, who didn't, what did they do? It's more interesting than programming. It is clear that their functionality is the main one that involves programming. And I liked Sam Altman again, and I got a message that, uh, well, Codex-- some man who wrote, "I understand, Codex is getting, uh, becoming a lead platform. All the software I know, uh, went, uh, to-- all the software I know, they've gone to Codex. Well, weird, uh, Cloude, Ilnar, with Cloude. You, Ilinar, know a lot of the software people that forgot Claude this week?

Mentions: OpenAI
Ilnar Shafigullin00:28:43

Well, look, 5.3 Codex is a good model. There's no argument here. But as you remember, we talked about when Gemini got out, and he went through the OpenAI, we said that wait for a month, the Open will come out with a response. That's what happened. This is exactly the same thing going on here. They're going to the nostril. At some point, someone comes out, at some point another one comes out. So if you're not in principle, you don't have to go anywhere. They all do the same thing.

Mentions: Gemini · OpenAI
Alexander Volchek00:29:11

I'm curious here, and I always have a case of interest in the agents of life and not life. We've got a lot of paperwork in the past, we've been talking about it.

00:29:20–00:31:15AI agents: why and where are the real boxes
Alexander Volchek00:29:20

I'm also interested in hearing more about it. I mean, this is a subject that's gonna be expanding. Still, we came up, remember, Ilnar, there was a time a year ago, people asked a lot about the agents' releases. I think that's the same time, uh, well, it was the time before, right? Even now, the agents are still in a big question, but still with the release, and with the progress, and, you know, you can make more serious decisions. Although I don't know why I'm gonna create a separate agent to study and create something if I can get this request into, there, Pro, and he's kind of doing it to me. I see that Agent, he can go somewhere else to bail, go in, get in, get in touch, keep his time.

Alexander Volchek00:30:00

To tie up, to do some kind of time. I mean, it's still a challenge, because the Internet is about to be filled with, uh, powerPoint, ten slides, there, ta-ta-ta-ta-ta-ta. Uh, not a case, nothing, right? Case, it's probably when your agent walks every day, every 30 minutes, checks, for example, the stock market, and if there's a trigger or some position, it's working on it. If something happens, he'll tell you about it. I mean, this is a good case, yeah, he's alive. I'm, like, permanent, well, in the market, I'm living in a... living- alive investment market, and I have some things that I care about, my unique, my own, here. Those I can't, for example, set up push-notices in mobile applications, yes.

Ilnar Shafigullin00:30:42

Now, that's the question: why would you need an agent for this if you could use a determinated logic to... write a program? Like it happened, send me.

Alexander Volchek00:30:51

I'll tell you why the agent. I'm not doing it, you know, long-term programming, right? And I wouldn't want a logic, over there, on top of some logic. I have a single window where I get data all the time. It's like ChatGPT's doing me. I still have a ChatGPT. I have some briefcases, for example. In the franc-- investment plan, it's very fun. He's counting the portfolio every week.

00:31:15–00:34:52Agent risk and loss of control
Alexander Volchek00:31:15

You're the one who's completely. I don't even see that task. And sometimes, I see that I received a notice, oh, I come in, and he thought it was. I mean, in this case, it's possible, in some ordinary text-based, different requests for my, with some other chat, to ask him a task.

Ilnar Shafigullin00:31:31

But you, you pay for that, because you don't control everything that's happening completely and how you take additional risks. I mean, here's the...

Alexander Volchek00:31:40

Well, I don't cry in reminders. As my friends say, my two hundred dollars are paid, there's a day of my inquiries.

Ilnar Shafigullin00:31:46

Sasha, you pay extra risks that arise from this. I mean, some kind of bang, well, some bang happened--

Alexander Volchek00:31:53

Illadar, 100 percent. Illadar, 100 percent. We're talking about agents now, not like, a-a-a-a-a-a-a-a-a-a-a-a-a-a-a--I don't know, the automating I did back there, twenty-five years ago, right? When a function is created and something is being tested in standard, 100-- and everything, and it's very clear. I'm just talking about agents in artificial intelligence right now. Here.

Ilnar Shafigullin00:32:11

And that's one of the reasons why they haven't been so active. Because you're paying extra risks, uncontrollability, injecting extra to intercept some control or something. So there's a lot of extra risks that are put on top of it. And that might be one of the reasons I think it's not that much developing. We have the comments last time--

Alexander Volchek00:32:37

And expensive. And expensive, yes.

Ilnar Shafigullin00:32:39

And expensive, yes.

Alexander Volchek00:32:40

It's very expensive.

Ilnar Shafigullin00:32:40

It's just that the condition "if," is worth a lot cheaper than the GPT-5.2, send and get the answer, left or right, to turn me.

Alexander Volchek00:32:48

But still, about the cabs, the agents, because they created an environment, just so that, uh, don't forget, you can still give examples. Ah, it's nice to see some live examples. It's a very interesting thing that's real, like I'm talking about reminders, like me, about testing, which is basically a microagent. Actually, it's a microagent. It works more or less.

Ilnar Shafigullin00:33:05

I'd just give you a few examples here, uh-a-a-a-a-a-ags that I saw. I'll tell you right now, it's not rocket science, like it's something so great that everyone should do now, but somehow. We've been asked to speak about Openclaw, too. If you heard, there's a platform like that, it's called differently. First Clawbot, then Maltbot, then I think it's Openclaw now. What is this? It's a cord over models that allows, uh, a little change in the process of interaction with them. So you're bluntly turning your own computer or the silver or some service. This service can take messages from you, for example, to Massangers. It could be there, Telegram, Slack, some kind of massager, whatever. I mean, you have a window of engagement with him. I mean, you don't have it on your phone, you're in your equipment somewhere, and you're just interacting with it through a text message in some massage. Next, this agent, or an agent shell, may, as a matter of fact, have some action under his hood. Uh, if it's on your computer, you know, the pattern of behavior can be that. You're on vacation somewhere, yeah, it's a little bit of a probable, I don't know, I'm going in the car, and you're on the computer, and there's some kind of presentation on your Mac Studio, some kind of report you've made, and now he's on your computer. I need it. And physically, you're not in the clouds. You see how much we put on the limit. It's not somewhere on Google Drive, uh, or somewhere, but it's physically on your computer.

Mentions: Google
00:34:52–00:37:14Agent risk and loss of control
Ilnar Shafigullin00:34:52

And in this sense, you can send a message to this system and say, "Find me a file on my computer, and then, accordingly, they've got it over there, I don't know, mail, yes, or something." And if this system is, this agent system has access--

Alexander Volchek00:34:52

It's like an hour.

Ilnar Shafigullin00:34:53

Yeah. Here. So, Tan, if you turn that thing around, you can do that stuff. Like you've got, I don't know, PDF-ca with interior designs. It takes a long time. You left for the night, you left somewhere in the morning. Then you can find out if she's got pregnant or not, and then, respectively...

Alexander Volchek00:35:12

I'm done.

Ilnar Shafigullin00:35:12

You can put something out there, I don't know, in a cloud, in a mail or something. Yeah. Look, on the one hand, that's cool. On the other hand, you pay for it? This agent system is getting access to tools.

Alexander Volchek00:35:24

Mm-hmm.

Ilnar Shafigullin00:35:24

Like reading files on your computer, editing files on your computer, maybe if you give it, accessing your mailbox, so I don't know, sending you, uh, a message or something. Sher-- or something, or access to downloading these files, somewhere in the cloud. While it works in the world of pink pony, it's all right. But then the problem is that this prompt injection is happening, right? I mean, uh, maybe there's some prompot you might use a little unconscious, and it makes you feel kind of vulnerable, and all your files are gone somewhere, right? So you're paying extra for this automation--

Alexander Volchek00:36:02

We went wrong.

Ilnar Shafigullin00:36:03

Risks. Yeah, they sent it wrong, uh, there was a mistake, something else. Besides, on the subject with the agents, we haven't touched it in a while, uh, uh-e- Agents MD, yes, i mean, files that agents describe what to do with this task. I mean, for example, an agent can perform hundreds of tasks there, and to avoid describing how to do a task every time, they do such text messages. They're on your CD, in the project. And now that the agent is doing a specific task, he's a little bit of a blank, and there's a directive on how to do it. You can write it on your own, or you can download it from the repository of these files. And if you don't take it very carefully, you can download a file where the instructions are not very good, not very honest, they can put some extra stuff, right? So on the one hand, the agent system is very well developed and is already showing up, which is, as you see, even perhaps useful, uh, useful. But we must understand that we pay for this loss of control over this set of actions that are being performed.

00:37:14–00:40:04Agent risk and loss of control
Ilnar Shafigullin00:37:14

Yeah, sorry, it's a long match, but here's a...

Discussion participant00:37:14

Well, that's interesting.

Ilnar Shafigullin00:37:14

...the comment asked, like, what's the subject.

Discussion participant00:37:17

Yeah.

Alexander Volchek00:37:17

Good subject. And here I want to add. We will move on to the following themes. In addition, what do I say? That's what the case was about. I sent him from Britain, I think, yes, from England. There was a case, which, well, if we're talking about an agent, agents are still being deployed, chat-bots are all over the place, that in a little business, people have been chat-bot, and the man with this chat-bot when he was in touch, he's been trading himself. I'm sorry, I'm sorry. I'm sorry. I'm sorry. Yeah? So this story is about the introduction of agents, we have to understand that there's a range. These things are, they're very simple in restrictions, right? But this, Ilnar, you were just telling me when you're giving access to the mail, all the files, a bunch of different data, you really don't know what's going on. We were living. Google, Google, for example, or Apple, or Microsoft. And they were basically unavailable? Yeah, they always had access, a crazy amount of access to your files, files, everything else. But you're like a window, right? Or you speak Zuma, there, or you write to the massagers. But now we're going into the area, one more thing, you have access to your messages from the company to read, process. It's something that's used to a crazy resource, it's impossible to verify. I remember someone once said, "There's a country where all the calls are constantly being listened, well, all the phone calls are written." Someone said, "You know, to listen to calls and understand what people say, "there's a separate agent sitting in. And this individual agent sits all day, listens to calls, you know, because otherwise you don't know anything." I mean, there was a time, you know, yeah, that if you, you need to follow one, some man, you should have been single out for a separate person who sits and listens to him all day. Now! Everything, the world changed. Now, to conclude that one person has done all his audio recordings in a week, it's about time like that, like, a little bit of a finger-coiled, and that's it. Yeah? And you have a report on your desk. So, very fast things, quick conclusions are made. So, of course, from the perspective of analysis, introducing some risks inside, but we're moving in this era, we're not gonna avoid it because any restrictions, they're all very, of course, they're all very relative, aren't they? On the one hand, we kind of limit, on the other hand, Apple and always scanned all my photos, and so he's working fast searches, according to the words, yes, on the key, or he wouldn't work. Look, otherwise he wouldn't work. How does he work then? Apple will say, of course, that we built this mass on your phone and so on. And Apple is considered one of the safest companies in the world in terms of data storage. You know, I don't think the programmers trust.

Mentions: Google · Apple · Microsoft
Alexander Volchek00:40:00

Information.

Mentions: Apple · Google · Microsoft · Gemini · OpenAI · Perplexity
00:40:04–00:42:14AI-companies: Who wins AI race
Alexander Volchek00:40:04

You know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, like, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, like, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, you know, And there's a lot of good in that. So they're actually safe, right? But it's still, security, relative. Uh, what do you want to say? So Google, uh, put out, by the way-- now, Microsoft has put out a number of people who-- they showed live meters for the first time. Microsoft, first of all, I'll start with Microsoft, I'll just say Google and move on to another one. Microsoft said they have about 100 million monthly users in the Copilot to date in their annual report, right? And then, uh, through the Speaker, they said they had, uh, 150 million users, uh, and, uh, inside. Google let them, that they have, uh-oh, seven hundred fifty, million, monthly users in Gemini. I'm just saying that this assessment, uh-oh, well, I think it's very special in Google, because it's still seven hundred and fifty million users inside Gemini's mobile application or inside, A-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-the-me-me-me-the-the-the-the-goods-g-goods-goods-goods-goods-g-g-goods-goods-goods-goods-goods-goods-goods. Because Gemini works a lot. Gemini works in Chrome, Gemini works in Google search, Gemini works in a crazy number of Google service. And that question, he, uh, he's hard on him, I think it's a question to get answered. But what's important is that, uh, uh, OpenAI, I remember, they had an exhibit that they had seven hundred fifty, eight hundred, eight hundred million users of the weekly. And look, this number is still on. I mean, it's been like four months, I think ChatGPT on traffic number one, but this number isn't really changing.

00:42:14–00:44:24AI-companies: Who wins AI race
Alexander Volchek00:42:14

I saw the statistics recently that ChatGPT had a monthly increase of 10 per cent. But I don't understand how they can have a 10 percent increase if they-- well, maybe a 10 percent increase and a run-off? Maybe new users come in, and the outflow-- but just if you have a so-called outflow, you're out of the way, you know, you're gonna have a little lesser number of users, right? More. I mean, you're not gonna have it. You're gonna have to have a lot more growth. Meta says that Meta has a billion users. But I-- that's just funny, right? Because I think it's just users. They said every second user, there, Facebook, means it's a user, uh, these systems. Oh, uh, just a little more chips. Baidu, according to their statements, China, is, uh, 200 million m-month users, ByteDance has 100, sixty million monthly users, Aliba, Qwen, yes? - 100 million monthly users, Tencent has 70,3 million monthly users, DeepSeek has seventy-two million monthly users, X AI, well, we're accurate-- I'm accurate. I'm not saying, but they're... they're really small, right? They're very small. It's not enough X AI works inside X, and it's a good way to go. This is why it's a matter of monthly users or not. Perplexity has $45 million, . What they understand is an active user, I don't understand. It was the twenty-fifth year. That's the only thing Perplexity had the latest data for the twenty-fifth year. Here. So? ChatGPT is still in the clear, uh, beauties, well, clear, because still, monthly users, weekly, this parameter is a little different. I understand that everyone chooses the most beautiful for himself. And Google obviously didn't choose a Jeme-Seven-- weekly, because it's probably a very strong fall. Just like OpenAI, you don't-- you don't pick a day, because it's probably a very big difference. Here. And, uh, but I think the pace has stopped. That's what you see, it's probably the pace of implementation, but it's just like that. Let's see what happens, uh, someone's got a real big scale. I'm impressed, to be honest, China. So, at 1 1 billion people, not so many people, uh, in Chinese artificial int-- well, that's not what a net had, like, 700 million. If I put them all down, I'll be able to put them down, like, three hundred and sixty, four hundred sixty, plus one hundred and fifty. Well, there's six hundred and ten million, at least the basic systems show. The other thing is, it's a unique user, it's gonna be some kind of thing.

00:44:24–00:46:22AI will write feedback for people
Alexander Volchek00:44:24

There's such interesting news, uh, that, uh, nature has raised the problem, uh, that polls and questionnaires can increasingly be filled not by people, but by an artificial agent, an action--intens--- an-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a- Intel. What's the problem? Which, in fact, there's a huge amount of clip on the Internet, you know, and different kinds of writing. They have learned more less about the definition of people. But as we know about Perplexity, uh, like the self--e-e-e, Perplexity's system itself has gone through a lot of things. In particular, I can lose, there, YouTube without, uh, no ad. And the, uh-- and they have the opportunity to create, uh, behavior. In fact, modern agents create the behaviour of real people, right? And there's gonna be more agents coming up with the very cool behavior of real people. We're-- we're in the hood. It's someone, uh, in the American government who said that Europe was guilty of having this shit on top. Well, we all know, right? You go on every website and you need to click. Do you know what a horror is? That's horrible! If you think about it, it's just-- by the way, ninety percent of Europe clicks Reject, and Eastern Europe doesn't care, they click accept. That's-- that's a study. And that, by the way, study, we need to see the numbers, it's really-- I think I've been talking about it in Tuzmu. It's very interesting that, in fact, people in Eastern Europe don't even care. What's the difference, uh, you keep our data or you don't keep it, right? We don't-- we don't believe that. The Europeans click on Reject. But I have a question: why don't you, uh, default for Europeans, take these caps to the damn mother and make them disappear forever, and somehow standardize this process, because I'm sorry it's W-- just a little bit of luck. I'm clicking on the day, I think it's 20 to 30. Well, we're all clicking, right? Some kind of space number on these hoods.

Mix00:46:17

Sasha, caps, in fact, are not used only to--

Mix00:46:20

Oh, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, no, Sorry, no drip.

Mix00:46:22

Cookie.

00:46:22–00:48:41AI will write feedback for people
Mix00:46:22

I mean...

Alexander Volchek00:46:22

I was just saying, "react, yes, accept, reject the u--- of course. I'm on a new one. I'm using that word. I mean, the cukes, Elnar, I don't mean... We've talked about this before in terms of the definition of the human. I'm the one who led this case, this case. So, Nature, why did I tie this to you? What they say is that now, the Internet is basically recounted, and, uh, the rags-- they're gonna be writing agents. Not much, look, we're going to see what world we're going to? That I was like going to a restaurant, and I have my app, like, some food or mine, mine, I don't know, ChatGPT, and I'm gonna take it and say, "Look, I just went to a restaurant, leave it on my behalf, "Let it go away," "I'm calling my name a recount of this restaurant that I didn't like him or that I liked him." Well, because I'm here last week, I don't know, I was in the hotel, they're sending me, "Leave the feedback." I'm so , "I don't want to go in there." But if it were in my central system, I might have done it, right? Accordingly, these systems are going like a-- that's the agent of my real one, leaving a retraction. What happens to the non-real?

Mix00:47:32

Elnar, did you hear that, too?

Mix00:47:34

The audience, what do you think? What do you think? What do you think about that? Because it's really amazing what agents do for people. But these agents, they become, you know, a spam for peace.

Well, remember, you and I were at the beginning of Tuzlam, and we were still discussing that, uh, you can just book a table in the restaurant by talking, I don't know, chat-bot, you can just put it on your side. chat-bot, and chat-bot chat-bot will be communicating, and, in fact, communication, uh, will be completely in this plane. But there's a very similar story, and the value of the feedback, it's gonna be all lower and lower. If she's any kind of a person now, then then, well, it's just not gonna make any sense. Just like the pictures and videos are not evidence. Well, at least if you looked at YouTube, it's not the fact that it's a real thing. And the feedback will be too seriously depressed, in my opinion, just because, well, they'll be filled-in like that...

Mix00:48:29

I wonder what's going to work? So what kind of design is that that's gonna be real alive? So what's gonna be real?

Mix00:48:37

Personal.

Mix00:48:37

Exclusion, uh-

Mix00:48:37

Personal contact with a man.

Mix00:48:39

The model?

00:48:41–00:50:49AI will write feedback for people
Mix00:48:41

The model?

Mix00:48:41

Yes, personal recommendations. I'll back Tatiana up here.

Mix00:48:43

Personal recommendations. It's basically working now.

Mix00:48:45

There's gonna be a separate network.

Actually, as a matter of fact, people who choose only on their feedback even now seem to be unwise at all. I'm kind of very tense when something absolute, there's five of five, which means, uh, people somehow have overpassed the system and only received positive feedback. Well, rarely when it's true. Or when people write feedback, how are you gonna check who wrote it now? The man himself wrote it, someone wrote it for him on a book or it's... Recalls, I think it's just a little extra plus after a few, uh, touches. And the more personal photographs, personal meetings, personal communication will be valued in general.

Tatyana Tsvetkova00:49:37

If there's a negative response, you'll think ten times, go in there or not and use this service or not.

Well, you know, I'm a service man, I know that people rarely leave good feedback, but they always leave bad feedback that might be the result of something that they just had in their lives.

Mix00:49:58

Yeah. But it's a big topic, because it's a...

Alexander Volchek00:50:00

Yeah. But it's a big subject, because, for example, if there's a $1 and a half in there for a man to pay for a request, or three dollars for what he's offering, the model can, for example, make it, five cents, there, or One cent or three cents. Of course it's gonna be a story. And yet, this system exists. There's no such thing as making a recommendation, there's a man in person. And we're used to-- I'm used to using the feedback. If I come to a new town, I open-- now we're there, I don't know, I went to Santa Barbara with my sister last week, and I'm not good-- and there's Mendocino, I-I--- oh, Montecito. I don't know the region well. And we're opening the system, of course, watching the feedback.

00:50:49–00:52:16We enter the AI chaos.
Alexander Volchek00:50:49

The other thing is, I've already learned how to evaluate restaurants on American feedback because, you know, they're a little different from the feedback, like people who don't know, live in Moscow, right? I mean, you have a little, a little different. Or on a certain type, uh, space, there, on the coffee café, if you drink, it's one, in restaurants, another, in some store, third. But the subject is very interesting. Write what you think of her. We're going in, uh, anyway, we're going into the zone now, uh, chaos. There's gonna be years of chaos. Why? Because there are a lot of old, huge numbers of old soft, old unadapted pages. A lot of things are gluing. The Internet is still not really working. I don't know how robots are gonna do. We're in some kind of chaos. Uh, that's when you have everything. And on the other hand, you, uh, just have some incredible cool things that a number, uh, a number of people can use. It's chaos, of course. It's chaos and it creates, yeah. But personally, personally, I do. Here. It's like, uh, these chat-bots, which, for example, gave a discount. That's, uh, that's a serious problem. So what happens to chat-bots, for example, if chat-bot does something wrong on a plane, like, or some kind of financial structure? But we remember the news is that banks are not rushing, uh, to introduce, uh, real artificial intelligence. They do the right thing, and they do the right thing. Well, I'll see you in a week. ToTheMoon Canal.

Discussion participant00:52:16