Every one of us holds a huge database about our own life and work. There is a great deal of mail, documents, photographs, video, notes, projects, doctors' reports. But the problem is that when you need to find something specific, we often do not even remember where it is. Today that is starting to change. AI can already be taught not merely to look for words but to understand how people, events, topics and various documents are connected to each other. And in effect to assemble your own Google over all of your information. How do you do that today?
Why does an ordinary person need it, and a business? And where does the chance to earn appear here? Today I will show on my own real projects how this works and where it can be applied right now. So the preamble is this. Before we move on to the subject of building separate search resources of your own, I see a large and continuing number of projects where people are building their own search engines. And there are even projects from former Yandex people who have just raised a round as well.
They are building search for AI models. And you and I have been going through this subject for several years now. About two years ago we were talking about how important it is that the sites you build, or work with, get indexed in ChatGPT or Claude, or Gemini. And in parallel we were saying that ordinary search engines would go away and browsers would go away.
And there was a very interesting comment. Late last night there was an article about search on TechCrunch, and there was a comment about what happens if a person stops going to sites. The question of quality here. Well, okay, okay, that is a separate, separate problem and it will be. But what is happening? Here is what I see in myself: I am stopping going to sites at all. That is, when Google used to give you results, you would see a small short text. If that text was done badly from an SEO point of view, from a search-optimisation point of view, you have no idea what the page is about, you open it, you start searching inside it. Whereas now you are given search and you simply say: "Make me a summary of this block, and this one, and this one.
Check the additional information." This is obviously going to work by default. You stop going to the site. And if you stop going to those sites in the results, then the whole concept of getting information, of spending time, and of advertising changes. Take note of that. And there was exactly that discussion: if people stop visiting sites, then how are you going to convert on their conversions into the offers they give — you simply will not be doing it, right? You are living in a slightly different, different, different life. Again, I speak from my own corner: I use even the current search tools very heavily.
I like this news very much, because what I do not like right now is the junk that is in the search engines. And I prefer it, for example, when a generative AI works for me first and gives out the information, with search secondary, rather than having this enormous listing come first. We will see what they actually release, what they put out, what they do. And when. Yes, and when they do it. And they have a whole series of different things due to come out. But the progress, it seems to me, is enormous. And Bing has just announced an update this week too on the search side inside itself.
In the models, incidentally — interesting, you mentioned hallucinations. Google Gemini in its updates gave a separate block to describing how they are supposedly fighting hallucinations, launching, writing separate, separate technical settings to check things. Because of course, when you ask it to read a book or watch a film, and you ask for an IMDb rating, and it gives you an invented film — a very strange feeling — one that does not exist. You ask for a link to Amazon. Yes, it is very strange. It has no right in principle to give out a link to Amazon.
Yes, it is a very strange sensation to have. And this, by the way, is an interesting construction around the fact that many sites are now blocking, blocking the search networks. For example, with TechCrunch you cannot give a TechCrunch story to it — it will say that I cannot go there and will not be able to analyse it. And it is very hard, it seems to me, to decide to open your own site to AI. And it seems to me I would open most sites, but the problem becomes exactly this. If TechCrunch is being analysed for you, why do you need TechCrunch? You will sit inside and analyse everything.
Well, there is exactly the point: if we get a bit futuristic, if advertising disappears from the search engines everywhere, then some kind of revenue share with the AI networks — if you open your site, then they share some part of their revenue with you for the right to use your information. That is — no, well, it turns out there is also the opposite story. I have an acquaintance with a large CMS. That is, a system that lets you build sites, and he feeds it as hard as he can — to the models, to all of them.
He writes special inserts, code, text as hard as he can, so that the models get the information about all of his technical documentation. Because he says: "I am deeply convinced that if I am inside, at least they will be able to suggest me. If someone asks how to do something on my CMS, it will at least write the code. And if I am not in there, it simply will not suggest me." So on the one hand it seems to us that you have to pay money to that TechCrunch for giving out the information. On the other hand, many systems ought to pay money back to OpenAI so that it hands them out.
In principle. Because if you set yourself up correctly and people simply stop linking to you, they will stop visiting you. Well, that is, at some point you will say: "Look, it does not give me the info, I will go and look at others." And here the question arises: do I need to build some resource of my own, where people can come and search for information about me, or about my
company, or about the ToTheMoon channel, or about my other YouTube videos, or in some other business build such resources — if in effect a person is going to work only through ChatGPT and will not visit any site at all. On the one hand. On the other hand, we understand that for ChatGPT to find some information in a structured way and really, really show it to you, that information still has to be presented to it somehow. And there is a big difference. It will, for example, collect information from a channel like ToTheMoon or from some site, and it will collect that information in the form it managed to download it in. Or it will collect that information in a very structured way, and you will show the search engine what you want to convey to it, what you want to reflect, what you want to show it. One of the projects I find interesting right now: there are certain authors, and with those authors I read a great deal of literature
and different books, and I often want to find, with a given author, what his specific thoughts were, what he said, for example, about some country, or what he said about particular people, or what he said about raising children, or what he said about human health. And in effect, what did I do? I go either into a model, or to some site, or into a search engine, and try to find it somehow. And even the models, when they do not find all of it for me, they find me information that is only partly verified.
And one of the stories I did: there is, for example, the author Rudolph Steiner, who has some three and a half thousand official lectures in German, which are open as the heritage of humanity. That means any person has the right to work with these materials, to structure them, to study them, including to translate them. That is, if before, working with ChatGPT or with some other sites, I had to find that author's books in Russian, then now what did I do? First of all I asked the system to find those lectures and download those lectures.
In this case I asked Codex. In parallel, in ChatGPT, I was still checking on intellectual property rights, so as not to end up in a zone where I cannot download something. And there is a big difference: you download for personal use, or you download, for example, so that this kind of fundamental search can then be done on the site. What interested me was information that would then be structured on my resource for the people who want it, and for me as well. But still, among other things, as a kind of business task. And at first I proceeded from the idea that I needed to download the Russian-language books.
That is, I download this author's Russian-language books and work with them. And then I understood that downloading this author's Russian-language books is a breach of the licence. But I did not want to breach the licence under any circumstances. And what did I arrive at? That I downloaded these German lectures. I realised that I would be able to teach the system step by step how to translate these German lectures, and to teach it his particular terminology, and Codex can already do this — it has already been trained on this author, it has already studied a huge amount of his material, and that way I will be able to translate this material into different languages. That is, I will be able to translate it not only into Russian, I will be able to translate it into English, I will be able to translate it into Chinese, into Spanish, or into other languages that I want, and to open this resource in the other languages I want.
What task was I setting today? I set myself the task that this system should give me not simply plain search but relationships — that it give me the relationships of particular entities. That is, there is such a notion as entities. Or, for example, we say that this person talks about particular places, and those places can be countries, can be cities, can be some locations, can be some tourist areas. That is, he talks about some places, as, for example, in our case there is the United States, there is California, there is Los Altos Hills, and there is Silicon Valley. So that I could look at what this author said about a particular location or some place — by a river, by a river, or by some shop, or a place belonging to a particular company.
That is, a company's location, as in Silicon Valley there is Apple Park. That is the location of the company Apple, its best-known office. And still that already counts as a location, as something very much, as we say, in Silicon Valley. And then separately there is the notion of people. And people can be simply a person with a particular name, like Alexander Volchek, or Ilnar Shafigullin, or Tatiana Tsvetkova. Or it does not matter — some president, which president exactly, or a singer. And there are people where we simply say that this is a child, a wife, a father.
A friend, friends, and so on. And so we single out an enormous number of such different entities.
And there can be very many entities. There can be the entity clothing, there can be the entity colour. There can be an enormous number of different entities. And we single them out and say how they are connected to each other, or what Rudolph Steiner said, for example, or what Alexander Volchek said about children in the United States. That is the construction that emerges. Incidentally, I have now partly done this project at ToTheMoon. To some extent you can see it on our site. And within my other channels — and I have a lot of different lectures, programmes and everything else — this is partly done, but perhaps by the time this episode comes out something will already be published.
Maybe not yet, I do not know. If there is, we will give the link. Or it will appear soon. That is, deeper analysis, when we are able to combine different entities with each other or to look at the depth of a connection. Next we say that within the material — as I describe at ToTheMoon — if, for example, you enter on our site, say, China as a location, or you enter, say, Baikal, then how many episodes were centrally about Baikal, or in how many episodes did I mention it. That is, where there is a real connection about it, where I actually talked about it, or where I simply had some micro-mention.
Clearly a micro-mention is ordinary search, and it is of relatively little interest to us. Whereas structuring information more seriously is a supremely important matter. And that allowed me to move incredibly far in analysing a large number of different projects. Projects in education, projects in real estate, projects in technology, projects in the artificial intelligence and technology market, projects in the finance market, projects in the market of spiritual development. That is, to go deeper in analysing and studying the material — not a superficial study of the material but a more systematic, structural study of the material.
It may seem here that I am talking about complicated things, or someone will say that this is all nonsense, the models can already do it anyway. The question is how far you are able to train the model further yourself. And here a very interesting story arises: in effect, today, in Codex or in
Claude Code, or even in ChatGPT, you have the ability to upload a large amount of your lectures or notes, or some materials, photographs, or all of your trips that ever happened to you, importing mail from Gmail or from Google Docs, or to upload various learning materials that you may use at work or in your profession, and to structure that information in great detail in terms of semantics, topology, various entities, occurrences, search. Which then gives you the ability, depending on your activity, your specifics or the tasks you are set, to speed up on the basis of that.
In my case, for example, in particular at ToTheMoon, this gave me the ability to give the search engines the right description of the content — of what we actually do. Or it gives me the ability to look at what we talked about this quarter, to run a faster, deeper analysis of what we discussed, what our reasoning was. Or, in particular, when I study particular material by other authors, to do it more seriously and more systematically, to get answers that are genuinely broad rather than abstract. Or if I need to look at some paintings or study some music, to get deeper into the learning.
That is, of course, for a basic person, if he has some skill in working in Codex or in Claude Code, or in systems like ChatGPT. So why am I telling you this today? This is a chance for you to get into it more seriously — on the one hand as your own learning, on the other hand, for a large part of the world right now, this is earning money. That is, you can make money on this if it relates to your profession — from the point of view of marketers, from the point of view of salespeople, from the point of view of project managers, product managers. On this, obviously, money can be made.
That is, in your companies you can now raise such resources — this applies to an enormous number of niches and directions. If you are not working in a super-technological niche where everything above is already superbly done and there is no way in, then this makes it possible to build yourself genuinely strong aggregators. There was a moment, incidentally, twenty years ago, when more serious aggregators started appearing in the world, and then gradually projects like Airbnb appeared. Or projects like Booking, or projects like Skyscanner, or projects like Kayak — plugging in different platforms, analysing something here and there, but technologically giving a person the ability to do some good search or some filtering. After that the world froze on the whole.
We have now come to a time when such projects can in fact be raised very easily both for yourself and, likewise, such projects can be raised for your professional work, and you can get into them fast. Which of you has done projects, and in which niches — first of all, tell me. Very interesting. Which of you has raised such libraries? I am talking here about serious libraries of graph connections, connections between entities. Incidentally, this development does not require any super-fundamental understanding, or the building of complex systems.
It is a question of how easily you can set the task for Codex. And, by the way, the task for Codex can be set on the basis of this video too. What Alexander Volchek talked about, for example, in this video. Or on the basis of some resources I describe, where this is partly done, when you want to structure information for yourself even further. Here you have to understand that in many places you do not have the right to download everything from the internet. You have to look at the licences. Where you have the right to download, for example, some resource or site, where it is publicly permitted — and where there are still resources where it is specifically forbidden and you cannot download information from them. And I have said several times, for example, that you cannot download information from the X network, formerly Twitter. That is officially prohibited by them. What is more, the fines for it are large. Well, some people do not pay attention to that.
But it is a fact. And there you can download information only through the API, for example. And then yes, there are libraries, there are authors, there are a great many different books and materials where it is officially permitted and you can use it. And even, in particular, for processing your own resources interesting things can be done — for example, from the point of view of processing your own photographs, or processing the files that are on your computer. Even from the point of view of structuring the files that are on your computer. There was simply a moment when, for example, aggregating all the files on a computer was a kind of local micro-task. And it is even unclear — not so much a micro-task; it was only partly solvable.
Now the models are excellent. I am talking about good models, at the level of 5.6 Sol 13 or Extra High. They have, of course, moved seriously forward, and so has the quality of structuring information, the quality of producing final reports. This applies, clearly, to structuring information — for example, from the point of view of your medical data or the medical data of your family. For some people this is a different approach to structuring data. For example, if you have a small medical practice inside, and within that practice you have the ability to analyse such data. Your users have given consent, the data was already stored. You too can now go to an entirely different level in terms of search, of entity analysis, of analysing details, conversations, connections between certain parameters. That is, to draw conclusions of a different class, of a different quality.
And to present it in the form of a working platform. Not simply in the form of "you gave the materials and got a simple answer", but where you structured that information. And then you say that I will, for example, every day or once a week add extra materials there, and that material of mine will update without recomputing everything else. Or somewhere partly the other materials will update too. That is an important, important, big difference. That is, to start building, with Codex, with Claude Code, such resources of large semantics, large topology, of very good structured search.
Do not forget to support our channel, to like it. We put out a lot of videos, and of course to write comments. I am waiting for your comments on this subject in terms of your reasoning and reflections, especially in terms of moving forward. Because I am telling you this in order to give you the chance to get a feel for this functionality. It is not that it will be supremely necessary in the future, but it gives a broader idea and understanding of how to work with the current systems at this moment in time, today. Bye, everyone.