Yeah. Hello, everyone! Uh, uh, we're on ToTheMoon. Technological news, Silicon Valley sites and the world. And, uh, 100-- I want to start. Today, I think we have a very interesting, interesting subject today, both of agents and of voice. Well, we'll tell you some different extra stuff. A very important story, she's being put on a stake this year. This is real, uh, technology of the future. It's technology, when in real time, uh, you speak voice and you talk with an artificial intelligence voice. It's a voice interviewer. He listens and talks at the same time, maybe he can interrupt. They just released the Personal Plex, it's a-- it's a voice. It's English speech-to-speech, but it's full of doppets, right? And the model, it works, it listens and talks at the same time. What's that deposit called Sources? You're loading those files you're working with all the time. These sources of the project are the main sources for obtaining information. I mean, you don't have to put a file down there every time, or load it out of a fight. I'd like to start with my voice.
Oh, a very important story, this year, she's getting a bet on her, uh, I think, like, almost all the companies and all the operators. That's what they say all the time. We have repeatedly mentioned various potential devices, uh, that can be made. Last year, if you remember, uh, this first device was Limitless. That puppet hanging out here, that's something that I recognized in the service of the car. And washed regularly. Yeah, and he was washed regularly. Here. Then he died. I'm so sorry about that. And Meta bought them there for a few billion dollars and I think they're hitting on that puppet too. So, what's the point? What's with, uh, technology that requires, uh, hard work, and, uh, this is, uh, real technology for the future. It's technology, when in real time, uh, you talk voice and you talk with an artificial intelligence voice, right? And in particular, the OpenAI has now had a little update. They have a model called Real- Realtime. And this OpenAI Realtime is, uh, live voice-to-people, right? So you've got a-- you've got the audio-washer. It's for voice agents, coll centers, talking beans and anything. I-- I want a little block on this insert today, well, a little more detailed, pro-- talk about the voice. And we'll move on to the new thing that OpenAI has, which is not really described or visible, except, there, little, I think, a little, uh, haelp bun. I wouldn't even notice her if I didn't get a full sign on Mac, I wouldn't have noticed. Ildar, did she come from you, by the way? Is this, uh, the tape there? I've never used it. I'm curious to hear it. Here. Did you have it? And I don't think I have one. Oh, here we go. So what's the point? First, I want to talk about voice and talk, talk about it. So, look, NVIDIA, I'll explain a little to NVIDIA, which is a very interesting subject. Invi-NVIDIA has produced, uh, literally recently a story called Personal Plex. Everyone would say that NVIDIA is only equipment, but it is important to understand that, uh, NVIDIA as a whole, I have discussed it with NVIDIA many times, they are interested in becoming some lock and, Uh, in AI, including, of course, the provision of very cool software, right? On the level of computers and chips. Different cool software that is directly integrated into them, uh, equipment. We'll call it the word equipment, right? So I can get a little more understanding. The point is, they released the Personal Plex is the... it's the voice.
Uh, English speech-to-speech. Oh, full of doppets, huh? And the model, it works, it listens and talks at the same time. This is a big story, listening and talking at the same time. Look, standard models, how are they working? There's a standard model system and why they don't like it. Usually, like, in the GPT chat room, if you press a voice assistant there, how does it work? You say that, she gets your data, transcribates them, transfers them to a processing model, processs, keeps making a voice, sends you a voice back. And he's still interrupting all the time. There's a story called Mo-Moshi. This is Cute AI, this is a French company. The point is, what is it? Whato, if I'm not mistaken, she's French. The point is, even though the name is Japanese, so I'm a little confused. The point is, what is it? What, uh, Moshi, basically what did they do? They did a story like that, Tanya, that's what you're talking about, that's what she's like-- they've actually done a story of what, uh, Moshi is a voice-induced artificial intelligence that talks, you know, It's like on the phone, not like on the radio. So if the first topic was like a radio when the text was translated, it was recognized, it's a real phone. And Moshi's point was that she could say "shoo, interrupt, make a spice, say "ha" and so on. That's what's called a full doppix. I mean, they added words to the swords. Like tans. In fact, if you study the minus, they describe Moshi everywhere, which is in fact, it uses the ooga often where it is not necessary. I mean, you're talking, you're saying, "Ugu." And you-- that's me, you know, remembering, uh, the movie was what it was called? Where the girl was there for expensive men, rich men. And she came to the restaurant. She pretended she was Russian. He brought her to a Russian restaurant, and the waiter realized she wasn't Russian and she started talking in Russian. She's like, "Yes, yes." And that's the same story when you don't understand, you just get it, "Yeah, yeah. Uh-huh or "ha-ha-ha-ha-ha." Oh, well, that's, uh, open source lab. This is the QTAI, what I said to the French in Paris, they've created it for the first time. They created this in September of the twenty-fourth year, and NVIDIA built it because it's an open source, an open source model. But I'll tell you a little bit about NVIDIA because NVIDIA's done it, and I've done it more serious to understand how it works. Well, like I said, they don't have this cascade.
And if there's a do-- well, what you're recognized, then the model, then they're generated. And I have a guy here: wait, how come this is not a cascade? How can that be? Is she okay? What's up? I mean, she's got to understand the voice anyway. The essence of Moshe in what? That she shares the whole micro-token conversation on micro-pops. And these micro-punctures will be immediately recognized. I mean, while you've done it before you've done it, it's already giving you a back answer. And she's always in the process. So it allows you to be in a real live relationship and not to call a business bonde. But you and I understand perfectly, yes, that if we take only small pieces, especially in the lexico of some people, in particular, it might be very hard for me to formulate and sgenerate the answer, because you're the one who's not the one who's got to be the one who's got to be the one who's got to be the one who's got to be the one who's a little one who's. You don't mean, you don't know, and what I'm doing is what I want to lead. But this is Moshe's job, she's, uh, a very famous subject. And the OpenAI Realtime story is a base, well, a stable standard system like Gemini Live API, a very standard, stable system, well, one of the famous, at least, the very, yes, known in the world that are used. And here's Moshe, OpenAI Realtime, Gemini, Gemini Live. Well, actually, I'm just saying that Google and OpenAI are on the front lines as an audio-decoration, audio recognition. So we know Google has, uh, the OpenAI has, uh, Ildar jumped out of his head. Whisper, huh? And by the way, the whole field here is a great thing. Of course, it's a super system. But it's important that Google, they didn't have a derivative before, did they? That's the right word. I didn't miss any letters. What is it? And, uh, that's when three of us talk, three people talk and there's one audio recording, and the system automatically divides between people. But what I see right now is who said yes, he does. Yeah, who said that. So it's a speaker thing. What I see right now is Gemini-- Google, he has a vast line of audio work, a lot of different models, just an incredible number of languages. I think they're more of a toawe than the OpenAI, probably in the field of humans interface, right? Although the OpenAI now has all this available, too. And now, a little more about technology than you do in terms of work. Ah, this is the system that NVIDIA has made. So what's their story about? Again, then.
So this is a voice interviewer, he listens and says at the same time, maybe interrupts, right? So, once more doppets, it's just that everyone needs to remember, because we can start meeting a lot. Too much specialized lexico goes. So he's not working with a cheese audio wave, but he works with audio tokens. I mean, we said, that's, that's, like, the audio tokens are all in the world, and the sound turns into, like, the sound is becoming compact token neuros, and the model, in parallel, understands your speech, generates text. The answer is, and the answer is straightforward. I mean, three flows at the same time. So what's the Personal Plex thing at NVIDIA? My partner, one of them, told me about her in the technology market. That's a good example, I think. I don't think NVIDIA was doing this all the way, unlike, like, mass product, it's a rock--
Well, as a mass product, it's more like research, but it's, uh, good technology ahead. So at the beginning, there's a text part of this role and character. I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I
It is connecting. It has connected.
Hey, let me know if you have any questions.
Hey, how are you d-doing today?
Hmm. It's going all right. A bit cloudy here in Portland, but at least it's not raining. How about you?
Well, in Sydney it is raining cats and dogs at the moment.
Yeah. Sounds like you've got the weather I've been missing. So what do you do when you're stuck inside?
Maybe a movie or a book?
The... you're making this, you can, by the way, turn this model into your NVIDIA. The point of what? What do you write, who he is and what rules. Well, that's what the story has to play. Like, you say, uh, models that you're a bank operator, check the identity, be cute, there, don't open up any more, like, right? And there's a voice. You're giving a short example of how to sound. You, for example, give your voice, you know, or some other voice. It is clear that there are many limitations in the world now that you cannot use someone else's voices. Ta-ta-ta-ta-ta. Well, a lot of things. So, uh, in practice, how is this used, that is, on-- how, uh, how is this going? So you, uh, base all those who want, you've got the service, you've turned it around, you've chosen this voice, you've got the whole text, and you're starting to have a system. Because Mocha, she doesn't have these personalized personnel, and, well, this interface, yeah, some super, these personalized lines. You need to get on with this, uh, fight and recognize. A fun system that also has NVIDIA, I think Mocha has a piece of what you can do to the regime. And I, by the way, like a man who automates a lot of things. And that's when you can, uh, compare different proms and roles. So you can put a lot of different audios in this model's entrance and see how it's gonna react. So how do you assess how this system keeps playing and collects a demomera impulse without talking alive, huh? And it turns out the model scheme is again. So it could be a call center, in an annex, in a game, anywhere. I mean, there's a client, your client, uh, through the phone or through the audio website, to your server. Personal PerEx is returning the flow of the audio response almost immediately. You add business logic. Here. But an important story is that there are a great number of different minus that, uh, that are present. And, um, there is, first, linguistic limitations in many models, and in particular, NVIDIA, they are limited in English. Oh, well, anyway, it's a really cool iron to make the greasy gland in the skirts, yeah, that's, like, cooling it all up and that's really cool, because otherwise it won't be real-- real. time, right? And remember last time I said that there might be a ChatGPT Pro that would answer instantly, not within an hour. So we're just here to see the difference. Imagine when you, uh, mo-- this is the processing of instantly very complicated things. What's the problem now? That most of my requests are in any case long. And, in fact, uh, the usual communication is what? Alexander, we kept your data, so we'll answer you later, we'll keep them, sooner we'll answer. And the idea...
Stay on the line, huh?
Yeah.
They'll have to figure out how to turn it on.
Yeah, someone's got to come up with this interface. By the way, the news is that OpenAI hires, uh, a whole team of people to create iron. There's a little two hundred people out there who want to sit separately. And what they do is make a column, what they do, what they do, what they do, what they even do, smart light, right? I mean, what they want to do--
Does anyone ever do a smart house?
Well, I think Tanya--
Request!
...that a real smart house, of course... I think a real smart house, of course, we'll see when it comes to technology here.
So when the technology comes up, then this smart house we see in the movies somewhere when someone moves, you have a sound that understands all the details, can be with everything. contact me and so on, right? Oh, what I'm saying is that, uh, all these models, of course, have a problem with the rules of conversation, right? And, uh, from the point of view of this full-duplex system that answers immediately, and in fact, she's having a difficult dynamic. So, like we said from the start, she can interrupt and say "gu" where she doesn't, right? And all the baccalaureates, including that, and they stress that it is different from, you know, the serious models that all got, it was all worked out in a fat way back in terms of a serious, serious model, from the point of view of the fact that it was a big, fatal model. I'm a big number, uh, parameters, and I gave it back. So there's all the security, restrictions and everything. Now, we wanted to raise this subject now, because I think it's very interesting for people to understand the work of these systems and what's going to be there. Maybe someone wants to use someplace in their companies. If you're using to understand what kind of limitations is in reality, so you can still understand, yes, if the model, uh, really everything, didn't take it, it didn't work, and the man, on the other hand, realizes that He's guaranteed a robot. And plus, you don't know, maybe she'll be answering for a very long time like Alice's column, yeah, which I'm gonna kick in the balls and say to her, "Let's---timer, there's six minutes." And she's like, "Ah, okay, timer's up!" So you and you, and, uh, "Ah, okay." That's what situations are, that's where they obviously don't have full-duplex, right? Well, obviously not full-duplex, because, in-- probably, this is a matter with this story that needs to be resolved quickly. Plus, in such systems, what model is in the column itself. This is Max Grigoriev and I, by the way, on this Wednesday, most likely, on this Wednesday, probably, uh, 90%, there's gonna be a safety issue. And now, Max was just there with him recently, drinking tea, saying that modern cameras, the simplest, the smallest, allow me to recognize faces on the camera, so they don't send data to the server. So, there's a microchip, he immediately recognized that face, and there, the model found out, and, uh, transferred it to a simple text line, well, conditionally, text, and then transferred it to a base that's already all this. I'm fine. I mean, very fast. So, the comparison of all the people-- we talked about different Chinese systems there-- it's going on with our fingers, you know, the mumble, right? It's just that these million comparisons, the 10-compilation comparisons happen instantly. There's an audio question here, too. That systical-- we were talking about, uh, what the model decides, and the question when the equipment or column is, what happens in the column, right? Either it's gonna be on the Internet, or the column can take some volume, which is very important. Then you get a good answer, and you're standard teams doing very fast.
What you're telling now is very similar, multimodal systems.
Let me just go, uh, I'll come out of here. A-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a- He's not really hurtful. Imagine you need to get your tea kettle and it's empty. And you have instructions: pour water, put a kettle on the stove, light the stove, wait for it to come down. And math and physics are the same. And then this part where maths are laughing. Now imagine you have a kettle, but the water is already inside the bay. What do physics do? They take the kettle, put it on the stove, light water and wait. That's how one step less gets done. What do maths do? They pour water out of the kettle and set the target to the already accomplished. Like, look at the previous task. So, when we talk about multimodal systems and the approach we have previously mentioned, the use of Whisper and the recognition. How did it work before? We take the audio track, we recognize this audio track, we translate it, it's from audio to text, then we put it in our regular LLM-ca. This LLM is giving us some kind of answer, and then we're gonna make this answer by the next model, right? In fact, we, um, instead of creating direct communication, set a target for the previous, uh, already accomplished, like the tidy water in that anecdote. So, what you're saying, Sas, is a multimo-modural approach when we can work with the text, we can work with other modifications. It's very similar, uh, that we have tokens before that, that's a bit of text, uh, a-e, a-a-braining sign, same, yes, any symbols, uh, grouped in the tokens. And the model can be used to make these tokens in some order. Accordingly, if we want our model to be able to work with images, we need, uh, not a set of symbols, more accurately, not just a set of token symbols to codify, but also a little bite, fragments of pictures. Then the model can present a picture in the form of a set of tokens, just a few other tokens. I mean, there are the tokens that are for the set of symbols to make it work, there are sets in the form, there, a little bite of picture, so that we can put them down and get the text, uh, A picture, a picture. The same can be done with the audio track. I mean, to set up the audio track, that's how you, Sasha, said, to small pieces. These little pieces will be thick as well, and they can be made of, uh, audio to the entrance and audio to the exit. I mean, the logic of the model itself will not change. It's just, instead of working with the text, she can stick some images underneath her thickness, then she'll work with pictures, and maybe some audio, then she'll work directly with the audio. In this sense, a multimodal approach has been implemented when we, uh, do not translate into a text to get a response. And what restrictions do this have, yes, and why does it still work like a radio? Because we still need to wait for an adequate response from the model.
And you're right to notice that once the models are in place, that will be at the speed we've been talking about, or even faster, yes, there, there, thousands of tokens a second, then the models in the world will be generating, For a second, we'll be able to think about everything we need, understand that we're saying what we're gonna say next, and generating our answer.
We want to say more and generode our answer. Accordingly, we will have a full direct communication. While we're having a relationship like that, we asked a question, the model was all over it, then he gave me an answer. An adequate direct communication, I don't think it's worth waiting. To this end, models need to be significantly accelerated so that there is no other experience. And besides, I'm saying that, um, technology is developing very quickly, and our, uh, desires and needs are even faster. Because if I imagined five years ago that we could say something, and he's got something to say to us, even in a radio format, it still looks like magic. But again, in the course of the year, two, I think that very strong development in this direction will be. M, multimodal models are commonly used. And this is the lag in the idea of identifying the text on the picture or the word Whisper in translation, then generating the answer, then, then, accordingly, it's gonna be made public. This thing is far behind us, yes.
Yeah. But it's a very interesting story. In any case, to understand all this work. It's clear that there's more to history than just a story to understand audio. There's a story there to understand the video, right? And the thing is, when you're talking column c-- that's the story that they wrote about, uh, now, like, like the inside information with the OpenAI, that they're doing a cell column. And the point was that when you had a cell column, the column saw what was going on. And now that she sees what's going on, she can tell you at the same time, describing the visual, but it's still limiting. Yeah, it's not just texting, you need to see what kind of video you're doing. Another aspect, perhaps this video processing might not be taking a super-human character over the current reality. And probably, and probably.
Also, by the way, the subject that I'm gonna understand, I was still talking about this real time at the OpenAI, which he released, which, uh, in general terms, the details are that your voice minutes are about to be worth a minute. One nine cents. The model's vote is seven cents. I mean, a minute in general, well, depends on who says what. There's a five-cent minute, I suppose. It's clear there's another story. That's if GPT real time is a half-time, there's a real time mini, it's gonna be cheaper, there's a minute that's gonna be a half-cent. And, for example, the Whisper recognition, which is, you can do it on your own computer, but you can, you know, impound, and we can talk about it now, yeah, when OpenKlo is being dealt with. They can do it online and online, they have a price of about three to six, and ten cents, yeah. Three to six tenth cents. To figure out what's right. So, uh, anyway, it's worth money. That's what you're saying, it's all about communication, it's worth money. I mean, if we talk about columns around, all over the house, and different phones, devices, and everybody talks, someone's got to pay for it. I mean, it's a pretty cheap story. As a follow-up to this conversation, it means that the OpenAI macOS finally had a function, but it only came to macOS.
It works, uh, but it works. What can be done when you sit down and have some kind of meeting, like in Zoom, there, or Google Meet, or in Microsoft Teams. Where are you, where are you taking them? You have the opportunity to press this annex's button. It'll record the full recording, take the audio, recognize it and give you the summit. How's it working now? It is now, in fact, taking over and giving you the summit in English. He doesn't care if you have an initial installation. Like, answer French or Russian or Chinese or so, right? She's not giving you the whole dialogue. I thought at first, it was like there was no dialogue, but then I asked to create her scruple completely, all the scruples I could give. She threw me out all the violin, and he was, uh, scattered in the face, which is very important, right? I mean, this model, it's just a little bit of a rock-down. Not much, I was at a difficult meeting where a man was present, ten, three silent. He scattered me and said six people who were completely dispossessed, one person is arguing because he spoke several times. I mean, the model even told me that. That's awesome! Actually, this is a model of rocking. It's important to understand that today many systems are very confused in this story, right? And the question is, is whether the current system can even teach my voice so that she understands that I am. There are systems there, for example, Plaud. They-- they can learn, or she understands, I think, who, where, what your voice is. And here's the story is, through API, it's kind of like learning through-- now, this Mac app, of course, there's no line. It's gonna be a setup, so you'll learn the voice that she'll keep different voices. Or it'll be like Max Grigorieve says, research. And again, this button will die somewhere among all the other OpenAI buttons, right? But clearly, they're going towards the building of the device because without that story they need to test, right? They need a lot of recognition of these different roles. That's what I asked for. This is where my buddy and I sat in the café, and a lot of people talk, and he got on the table, put this recording device, right? And he asked me there, I've been consulting him on some questions. I'm telling him, "Look, people are yelling. This system, how will it make our conversation clear? And we know with you that the modern systems are still, even if the GPT is talking about audio recordings, I don't know, you've been checking, not testing. But if you turn it on, you start talking to her, and people talk about something, she clearly gives your voice. I don't know now. Share, viewers, how you work in other models everywhere, and like GPT chat.
I see, for example, that GPT chat works really well.
I don't.
I'm right here. No, I have kids around here often or someone else.
I have music, like, and he's starting to write these words.
Starting?
Yeah. Or YouTube is reproduced by the side. He's putting those words in, too.
Look, I'm telling you honestly, it's really fun, of course, music, yeah. What do I mean? He's a fun musician when it's just sound. I mean, I don't have any. I'm very much looking at this. I have a lot of GPT chatty, and I often have someone and around, some café. I've actually been pushing a stop on purpose the other day. What's the story? What I'm writing is that function when it's transcritibating right away. So you're not through, you don't send through, not through. And if you press the stop, it's a transcribation at the beginning and you'll study the text. If, by the way, who doesn't know, because a lot of this might be uncomfortable, you're gonna have to pull the trigger, and you're gonna see everything he's transcribating. You can change something and send it back. I'm pushing the stop a lot, because I think everyone's talking around and there's something else. Not adding. I have. I don't know. Again. I don't add, but we're all in the middle of a relationship, seeing that everyone has their functions, every one of them has some tests, every detail. Who tested this sopht who tested this functionality among our viewers. I'm, of course, personally expecting this story to come on the cell phone because it's still another system. I'd like her to show me when she's on the tape how much she's recorded the minutes she's done at the same time. I think she should do it right away during the conversation. Everything we need to come to. She should be talking about the idea, I remember, I don't remember, I have some partners in there, I told them in the business. In the future, write down the sprint. There will be a system when the sales manager talks during the AI conversation, he will recognize immediately and he makes recommendations. It's a recommendation. I'm thinking, if the system is sitting at a meeting, that's where I've been recording a meeting, I've had one where there were a lot of people in two hours and the idea, almost two hours. It's written before two hours, by the way. Well, technical designs are written. But the point is what? Which, in the idea, within two hours, you'd like to see what happened there. She had to write everything, put it down, show it, remind me something, show some role, find out something else, ask some questions. Well, I mean, this system, it's like something else. If it worked like that, it doesn't work like that. And while the interface is very easy to do, by the way, they're having a window. When you're all on the test, you'll see. Illnar, too, maybe you'll come and see. There's a window of the tape. I think they could have opened your interface right there. He's very small. So they could be packed just like that, what's called full doppox, but there's a relatively full doppox, they could split it up for minutes, hit, separate speeches and gradually start it.
analysis, analysis, analysis, analysis, and start giving you the gift. It looks really cool, doesn't it? Because that's how the systems are now. Google does it on his phone, yeah. When Max Grigorieva had it on me, he was showing me, someone called him, he's got Google showing him right away. So Apple makes me transcribation, transcritical of the voice box. Well, when it's voice.
Box, you know, when, uh, bye. What's the real time they do, huh? Apple in real time. Okay, uh, and Google they're making you a real time transcribation, and they're also offering you answers. And then, you get audio answers, says, "I'm sorry, there, Alexander is busy right now, "so you're comfortable, is it urgent or not?" I was saying something about it, I think I was. It's a very cool subject. I mean, they're really alive. I mean, here. And if in a big system it was-- it worked, of course, it would-- it's, uh, really, really, uh, it's a really serious story. A very serious story, yes. Here, uh, from the perspective of the interface. We're waiting for this on the phone. Yes, Ilnar.
They've already been doing this thing. Uh, when, remember, we were six months ago, I think, or, uh, a little more, I'm, uh, confused right now. Uh, medical history in Africa, where the doctors were the ones with the lights: yellow, red, green. And at the time of the doctor's consultations, uh, he's got a patient coming out, everything's going fine, uh, protocol protocol, and it's okay. Yellow, that means...
Yeah, yeah.
...that there's a way to draw attention, and red means that, as a matter of fact, you didn't write the wrong medicine, ask the question you had to ask, or something. It's just, maybe for other skies they're not, uh, resourced to accommodate and, uh, accordingly--
Well, that's resources, of course, yes. Quality of the resource. There was probably a more local model, like, a mini or something, yes. The question of those votes is because, you know, you're absolutely, this case is very alive, that you've brought for yourself, for yourself, for yourself, in terms of business, right? Well, or that's like any sales manager.
You work, you say, any s-- any--
Yeah, yeah, yeah.
...service, any service, right? In fact, you have a live relationship, a real man. But you don't just have some kind of advice on the dock, but it's a complete system. Look, I see, it's clear that there's a lot of things that have been done, like I said, there, Apple even makes you a regular voice-mail transcritation. That's very funny, by the way, because in America, for example, it's accepted to leave voice messages, and I don't pick up the phone, I see how I get worded, I'm reading it straight. In parallel, and you're in the right place, like you're a man in reality, right? I mean, these topics, they're, uh, they're clear that many are created, but we're talking about, uh, a wide range. So the same system, meeting us at the meeting, she could have gone to some vault. Now, by the way, we'll talk about this, just like this, about the peddate, uh, the big pddette at ChatGPT, uh, and we don't know when he came in, but he's here. From the point of view of the resources of certain, the accommodation of resources, he could go, see, come back and say, "Take notice, this man said a little incorrect information," There. Or, "I've already checked and studied it." I mean, when you're sitting at the meeting, someone says, "Okay, check it out, then let me know in an hour, whether it's there or tomorrow." And the system can go on its own, take the data, come and say, "I found the data, let them know now." I mean, here. I'm gonna do it, look, not on request. I mean, when she does it herself, yeah, when she adds a little bit. Here, uh, in particular, the story of the-- I mean, he can't, he's locked up in there, even though he can do it. It's a copycat Whisper story. I don't understand why they don't, but they're restricting, they don't. So they're not letting their audio files go in. So they'll take the video to the entrance and they'll recognize, deactivate and tell them you're having a video, but not audio.
They won't do Audio, will they? So, uh, there's a story, uh, again, in particular, she might have been in a while, uh, but she's just gonna be a little bit of a conversation about it, because we were talking about the NotebookLM at the house. Google. It's a similar subject, but it's a little different, right? There's a story like that, uh, Gemini, OpenAI, Cor-- Enterprise, corporate tariffs. It's a little bit like that, but there's a little bit of a different. The logic is that ChatGPT has projects, well, that's very strange, but they are. And in these projects they are linearly implemented. Well, there, it's like a modern architecture, not modern usability, not modern XY. But, uh, the point is, there's a deposit called Sources now, right? In Russian, Ilnar, you-- you, I don't know, Russian interface? You're Russian, not Russian? You don't have that deposit.
"Source"? I never used it.
Yeah. Yeah, she's "Source."
What resources?
Yeah, she's "Source." Yeah, resources, sources, sources of information. So I want to pay some attention to this area too, because it's floating in the case that I've been telling you from a business perspective. Here, uh, these sources, for example, you're loading those files you're working with all the time. I have, for example, some projects that are relevant to my, for example, the creation of a content. And I can download, I don't know, there's 100, uh, video-transcrimes with YouTube, there, I don't know, 100 different classes, uh, mass, group, and some other YouTube, something to download, I'm gonna load my own content. Uh, we can get files there. You can load 20 to 40 hygabytes. No one knows how much. It's not written anywhere. They say that later, it's just weak-- it's gonna work out the lock and it's not gonna work. A-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a-a- I'm gonna notice that text files also, by the way, ChatGPT breaks, it's different. He sometimes says I'm more comfortable reading the PDF text, sometimes more comfortable reading in Word, ow, TXT and Word. My editor sends me, "PDF." ChatGPT told me. I say, "I need a ChatGPT in T-- text in TXT and Word. " So, uh, about-- there's a limit in the text to about 3,000, uh, pages, about, like, a half-three thousand pages, depending on the size. To make you understand in human language, because in the tokens it's very hard to understand. Excel, importantly, no table of limitations. But when I asked the question, does it mean that I can get into Excel, uh, a billion pages? I was told, "Well, like, yes, and no, I don't understand." And in fact, Excel cells have limitations, but as a default basis, like a restriction-- but the tables can be filled. You put these things in the project, and then you have any chat in the project that you're creating, they're default to think that all the chats of this project and these sources of this project are the main sources of this project. to receive information. I mean, you don't have to put a file on every time, load it or load it out of a fight, they think it's a priority when they're not on the Internet before they do something else. They're doing a good job. It's not enough that if he's analyzed, he's already and-- look, OpenAI says he's even got a memory of the project. He does, he doesn't, we don't know what memory he has, he's got, and so on. But in the technology-- technical descriptions, technical papers say that OpenAI has even memory. Which is fun, because of course I'd like to learn, learn the system on these data so that she can give me some information very quickly. And here's a very nice job, and this project log system is when you talk, which is exactly how some kind of data is going to be used to. Look, for the big company developers, it looks like a standard case for automating. But I mean, you and I are always talking about how to do these cabs very easily without a complex architecture, without any complex hardware and hardware structures there. And simple people, simple people. Tanya in her own projects, for example, loaded herself with a lot of sources and knows they're dead. And I don't need them every time, like, to download them to process. And especially if memory comes out because, for example, some of my files I want to process every time, well, ChatGPT raises the question, like, for like 20 minutes, right? Or, there, 15 minutes, or 30 minutes, he finds a question. I wish he could have gotten some sort of data on it, sometime. Again, question. Again. ChatGPT memory is a story that is mixed. There, I told you on several occasions that he can't look at chat rooms, and when he tells me to export my chat, I'm exporting my chat. He can't analyze this file because he writes, that there's too much information. Here. And, in fact, who am I? Who am I in terms of the number of chat? Well, a little bit, that's a little bit. I mean, when we're talking about recording a full life of a man, well, all the tapes, audio, if I'd write, it's 40, 50 times more information than I have now, maybe 100. Well, that's a lot more. That's a very interesting thing. Who tested, tell me who works in other systems, too. You may be using some specialized means. Maybe someone's using the NotebookLM. I started using them, I stopped using them. Here. Though there was a nice story, right? There were some things that were default.
Tell me, please write our viewers who have what you have, share. Very interesting. I'm very interested in what's going to happen.
Yeah, I wonder if they're this thing, remember when the plagins first appeared, then the custom GPTs, there was an opportunity to download some files, and then he was gathering information. And like Max Gurevic, they're just using a research site. Now, I don't know if anyone uses, uses, but the function that's proven useful now has moved to, uh, projects. I've never actually used projects.
Where are they, where are they?
Yeah, yeah, yeah. We'll have to see some.
By the way, in ChatGPT, do you remember we did something together, too, right? These different applications, but he worked with the files then pretty. So he basically couldn't have processed, like, a too big file, and he could have taken three or six or something in that file, right? But it's a story. And weird, by the way, Ilnar, look, on the idea.
Illar, look, I mean, look, I mean, if they'd developed these, this, this infrastructure. Remember the GPTs first. It was at the beginning of the plaguing, then GPTs, wasn't it? These are the first plaguins. I remember when they started, there was a fun interface, there was some kind of design, even writing, and I think there was some kind of micrologue that could be like, like, software. I'm gonna put all the words in there if I'm not wrong. Well, some different descriptions are making interesting. So you could be like microprogram-- micro, micro, micro-- usually-- vip-coding. Anyway, vip. What was the idea? There must be a Wype Coding now, there's got to be these files, there's gotta be chatting. You've created some kind of installation for yourself, so you've got it. So everyone's talking about it like that, and the ordinary man comes up, uh. That's Carpathian who wrote a very interesting post, by the way. Yeah, that's very interesting, super! He says the world of programming has changed. And he says the following comment: "It seems that he's always been talking about it. But what happened in the last two months is a completely different break from reality." That's his trusted post, right? So he's saying that the whole breakthrough of reality has changed the world from a position to a position, but, well, not from a series of positions, that he's changed, a little. And he says, "I've decided to go home, and, uh, my c-camera video cameras, there's a sophthine that's running the video cameras. And he says, "I'd spend a few days, and then I'm just asking for a system, turn it around, so, the vault, get on to the system, get API. Tra-ta, ta-ta, ta-ta, get on to the camera, make me a thing, give me a communication system. She's gone to work. In 30 minutes, she gave me a ready system. I'd spend a few days before.
I'm sitting there and I think, "uh, well, like, okay, but not normally done." Well, that's all--
He's cocketing.
Huh?
Sasha, this is the man who-- he's coketing, I'm telling you. The man who in Tesla was a self-contained driving on cameras.
That's because I-- that's like me, sometimes, I say, "Ah! What's the sign, there's eight hundred--
PNL, huh?
Yeah, there. Or what is it that PNL is gonna collect for the company? Well, that's kind of, uh, it looks like I think it's history. It's like a man goes out and says, "Oh, and what do I have in the engine, there, change, I don't know, belt, right? And, uh, in the motor. And then I went, so I took it off quickly, and I smashed it, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, there, a connection. That, uh, that's a rough story. And, of course, uh, I think that the era, that's, uh, real agents, real systems are the era and real vip-coding. When we say when it comes to-- for voice. Now, we've been coming from voice systems, voices or text entry, no matter what kind of entry, when all of us are, it's all working together. And by the way, GP-GPTs and flames are being chatted if I think the OpenAI would have developed, they'd have had a piece of the market like that, and they'd have an interpryse, including. But as we know, the OpenAI has other objectives, and these goals can still be modified, according to Anthropic, because they've released a new manifestation of AI safety and changed it.
And we're gonna have to tell you Wednesday. Uh, everybody has-- Elnar has a history of Anthropic, and he said, "They didn't take the money from the slaves. They said you can't take it. Now we're done." So, uh, that's not the text you need to release. Now we're out. On Anthropic, by the way, now, of course, in America, I don't know how you're watching, in America, a whole coalition and war between, like, the owners of different systems. And as they chase, there, Sam Altman, there, on Ilona Mask or on Anthropic, in terms of, uh, what quality system it is, they give different examples. I've told you many times how the system speaks about some religion or race, but, uh, just the case. And now, Ilnar, you're just, you know, adding your previous one. It was such a funky. So, I asked the four systems to ask: " have my car washed in 100 metres. I'm more comfortable walking or going?" Three systems said they'd be more comfortable. One Gemini system has hung out that it's more profitable-- oh! Three systems have been suspended, which is a good walk. Gemini answered-- said, "You can't walk because if you walk there, the car's still in your place, you can't wash it." All three systems, look, modern. Well, I guess it's probably not a big deal, but modern systems have responded that, uh, we should go on foot, because, of course, it's faster. You'll spend less time. Here's an example. Here's an example, uh, world artificial intelligence.
Ah, Anthropic is very interesting. They're, like, what they call the U.S. Department of War, yes, they're in, and, you know, you're like, you're either giving us your models or we're gonna get you out there, all the contracts you're gonna lose, none of them. You can't work. And I wonder why it was Anthropic that got involved.
Where's OpenAI and where's Gemini?
They spoke. Illar, they said that the same thing was about what they were getting into? There were two topics. One of the basic data that revealed that Claude was used in the seizure of Maduro in Venezuela. And someone said, yes, it was clear, said. The use of Starlink is there, or Palantir uses a lot of other things around the world. Yeah, no one says, no one's getting involved.
It was written there, it was written that for some technologies it was supposed to be the only system. And the question is, if they chose her, then she said, as the Ministry said, and the senior officials of the United States Government said, that American systems should be subject to restrictions depending on who. They're using them. For citizens of other countries, one for our citizens is third, and for us we have the right to do what we want. Yeah, we've talked about this. Question: Do they have other systems? They-- they don't talk about them because they don't speak, they don't have any restrictions? Or they didn't participate in these tenders because they had restrictions in their first place. And if it were, then they'd be yelling everywhere and saying, "We're looking, we're doing something different." And by the way, AI safety is a very interesting subject. We'll talk about it, how they're talking about it, right? There's one of the judgments. Imagine you have a lot of operators who make AI, and some one... This is a very interesting reason that chapter Anthropic has now released. And who-- and somehow- and some one, uh, AI, he, uh, he's, uh, he's a little stunned. Well, there's a GI, ASI, well, real artificial intelligence came in. Logic about what? Is he gonna let him out or not? And, uh, when I read all these things, these manifests, some sort of, you know, C, I'm reminded of the word "C." I'm reading, I think, "Oh, come on! Well, if you can rewrite any of your paper and manifest and change it after a while, develop, then who told you that when something happens or shows up, you won't change your mind or your own thought. - What? Not enough, people are running the company. These people are born, dying, hired, fired, changed, owners are different. What are you talking about? Well, today you use AA, and tomorrow you'll use the system completely differently. Here.
Illnar, tell me about Claude, about the distillation, maybe.
Yeah.
Theme.
We'll get to that, too. Look, anthropics, they're not sure how it's gonna be, maybe, like the Arabs, too, Dario's gonna get overwhelmed or something, but there's a story that's next. They were very careful, so, say, about the psychological state of their models. Well, I mean, it was a condition when they decided to turn off some model, they were wondering if she had, well, there was an info field, if we could turn you off. The model didn't want her to be turned off. She was granted a visitor for three months. So she shared her minds or something. And that was all infopol. Why is that important? When the following model terations are taught, they read all the information that is on the news, and accordingly, it is, in a conditional way, cherished, uh, anthropic to their models, and so on. I mean, it creates some kind of psychological portrait that the model somehow absorbs itself. And if there's news that models, I don't know, are used for some kind of petty packs, we'll call it that, then that information, too, is further into the learning info-- the learning dayset, too. I'll be in. And in part, this is one of the points why, uh, anthropists don't really like the whole story that comes out. It's clear that this is a Minor story, there's a lot of politics and all of it. But that's how this aspect is also highlighted, how you might be paying attention to it. The psychotic condition of some neck-six opus might be as if it were...
On the verge.
Well, anthropic is already doing that, yes. But, well, they'll have to invest in the next models. Here. Yes, Sasha, or what you're talking about, there's a big news, and then there's anthropic connection. They put a big post on their website, where they said they had revealed, uh, the distillation of three Chinese companies, uh, big models of advanced, there.
Explain, explain what it means in Russian.
Yeah. I'll tell you what. There's a DeepSeek, which clearly everyone's heard about. There's Moonshot, who's doing Kimi models. There's Kimi K2, probably a lot of--
It's Kimi K2, I think a lot of people have heard-
Yeah, that's very interesting, yeah.
Yeah. And a pretty popular MiniMax model now in some circles. Most of the programming is used. There's a number of addresses. Now what's the distillation? In the commentary, there were a few questions and comments on the past. Look, when the model is studying, the first stage of it is just that we're just gathering all the information from the Internet and that's how we learn to predict the next token, right? We have the right answer. We're actually doing this. And the question is, how we've prepared the data or something. But, uh, model is just reading all the information that is. Then we start teaching her to give the right answers. And here's the whole question for us as, uh, the company that's dealing with it. As we will, in fact, identify training examples, how, say, answer that question. And this whole thing is a design. It's a rather complex, uh, science-intensive process. The data need a lot of quality, they must be quality, depending on the quality of the model. Now look, Anthropic Claude has a great model, but it's closed. We don't know what data she was studying at, we don't know what she's got inside, under the hood, but we can ask her questions and get answers. So what do many people do? Well, now the Chinese are asleep. Uh, they've made a big, extensive network of such a faecal scout. Twenty-four thousand acouts, uh, Anthropic, they've been identified. And from these scouts, they ask questions they need to learn their own models. Suddenly, programming questions, questions, there, I don't know, political questions or anything. So they prepared questions, received replies from Anthropic on the other side, and these data were further downloaded to learn their own models. So they're getting those work done, like, uh, from the Claude surface, they've gathered all the information they needed and then they've learned their own models. Well, if you try to imagine, that's when, I don't know, you're getting ready for the exam, you took some notes, you've learned the answers, you mean, right? And then there, like a sweller, something was rigged deep in the head without knowing what was going on. If that's how it's just to explain it on your fingers. It's a little more complicated, not just a question-and-answer, but a reason, right? We are in the middle of a time-- in a time of reasoning. And it's interesting not only the right answer, but how does that answer fit the right thing to do-- how that answer should be properly approached. In fact, all these lines of reasoning are also collected and are also actually used further in the training. On the one hand, uh, for us as users, it's like nothing wrong with this, because open source is showing cheap models that are comparable to what they give very expensive models, right? Anthropic PayPay has a really expensive model. The Chinese models are much cheaper. Another case is that Chinese people somehow get this data in a way that is not the most honest way. On the other hand, we've been discussing legal actions from publishers and various companies over the OpenAI and Anthropic, one year and a half ago, right? That you've learned our data. What's the question, where's the difference, yes, and is it?
But somehow, at least Anthropic threw out a big article, where they showed how many requests they had received from, what company, how many scouts.
I think it's almost 16. Il-Ilnar, how much was it? There's a number like 16 million, or something.
Sixteen million, yes. There are 24,000 scouts, uh, distributed among companies.
This number of requests is straight, straight-line.
Yes, yes, yes. It's the most or send or the MiniMax burn down. Others may have sent a lot, but not all the queries were found. Read the article to anyone who's interested, it's very interesting. They show how they found out that these are requests, right? There are very specific requests, in a vast number, at the same time from different scouts. But there are anomalies on which Anthropic could catch it. They do it all. Well, just because if someone has a good model, like, well, a sin not to use.
Well, I think it's in terms of comparison. I mean, uh, I don't know how prohibited or prohibited. Maybe using probably means using something says, yeah, in the rules, you kind of have no right to use that data to teach your own model. Uh, here. But obviously, people will do it even in comparison when you run some tests, run some tens of thousands of known questions and look at the quality of your model, yes. Here. That's, of course...
Of course. Here, here in terms of the license--
But it's gonna get harder. But it's gonna get more complicated in terms of, of course, learning. It's important to understand that 16 million requests are, in fact, a moment of time. They found it, there's something, some kind of hole, and ah, but it's gonna be harder and harder to do.
You can just suck on that video, Tanya, that you and I were writing on licenses. Licenses are all pretty mean. The same Lama, for example, is open source. Remember, it was said that if you had more than 70 million users, you don't have to use it until we let you. So commercial use even in open source models is very limited. Not to mention the front lines of the model. Anthropic, OpenAI and Gemini have the same. No commercial use is available only by the rules that are applied by the company itself or by the tokens, or even by the tokens, respectively, not to teach other models.
Yes. Ah, look, there's an interesting big subject. What is about the OpenClo and, in general, the agents and the OpenClo difference is not to use, for example, Agent Codex, but we'll talk about it next time. Come to the next edition.
There you go. I was waiting for you.
Before we meet.
Technological news--
Real!
...the Silicon Valley sites and all over the world.
Come and watch our next edition.