Skip to content
ChatGPT · Codex · Artificial intelligenceEpisode 160 · 4 September 2026 · 21:24

Your Own Google Across Your Whole Life and Work Is Already Real

Central question

How do you build your own search across all of your material — letters, books, lectures, video, files — so that it shows connections between people, places and topics rather than word matches, and what does that give you at work and in business?

What you take away

The reader will understand how this kind of search differs both from ordinary search and from asking a chat: entities and connections instead of keywords, the depth of a mention instead of the fact of a match. They get two worked cases — a library of three and a half thousand lectures by one author, and search across ToTheMoon episodes — a frame for rights, where downloading is allowed and where it is not, and a sense of where such systems earn money: inside a company, inside a profession, and as a product of their own.

Main threads

What to watch for

1Start with one archive rather than everything at once: mail, a documents folder or your own lectures — wherever you already lose time searching.
2Describe the entities before you upload: people, places, topics, events, roles — they are what turns a pile of files into connections.
3Check the licence of every external source separately: personal use and publishing on your own resource are different rights.
4Set the task for Codex or Claude Code in words rather than as a technical spec: the development here comes down to the wording, not to system complexity.
5Make the top-ups regular — daily or weekly — so new material is added without recomputing everything else.
Signals to track afterwards
Whether people stop visiting sites far enough for the whole advertising construction around search results to change.
Whether publishers and models settle on revenue share, or the number of closed sites grows.
Whether quality catches up: fighting hallucinations is announced in the release notes, yet non-existent films and forbidden links still come back.
Whether the author's open resources for his other channels and lectures appear — at the time of the episode only a link is promised.
How far the structuring of sensitive data goes: medical archives are named as a separate, stricter case.
Most useful for
Anyone with a large personal or working archive who loses time inside it.Authors, channels and publishers deciding whether to open their material to the models or shut them out.Marketers, salespeople, project and product managers — as a working tool inside the profession.Anyone who wants to build a search resource as a product of its own and does not know where to start.Anyone already using ChatGPT who has never set a task for Codex or Claude Code.

Key takeaways

00:00Everyone Has the Archive; Nobody Can Search It

Letters, documents, photographs, video, notes, projects, doctors' reports — everyone's database about their own life and work is huge. The problem is not its size but that when something specific is needed, we often do not even remember where it is.

00:18The Model Is Taught Connections, Not Words

AI can be taught not merely to look for words but to understand how people, events, topics and documents connect. That is what makes your own Google over all of your information.

02:01The Reader Stops Visiting Sites

A discussion on TechCrunch: what happens if people stop opening pages. Results used to give a short text, and a bad description meant opening the page; now search summarises the blocks you need — and the whole construction changes: information, time, advertising.

03:31The Model's Answer First, the Listing Second

The author likes the new order: search results have accumulated junk, and having a generative model work first with the list of links secondary suits him better than before. Bing announced a search update this week too.

04:17Hallucinations Still Produce Things That Do Not Exist

Google Gemini gave hallucinations a separate block in its release notes, along with technical settings for verification. The reason is concrete: the model can produce an IMDb rating for a film that does not exist and a link to Amazon it has no right to give.

04:54Open Your Site to AI, or Shut It Out

TechCrunch is closed to the models: ask one to work through its story and it will answer that it cannot go there. The author says he would open most sites, yet sees the problem: if the material is analysed inside the model, the reader no longer needs the source.

05:51The Opposite Strategy: Feed the Models Yourself

An acquaintance of the author's with a large CMS does the opposite: he writes inserts, code and text so that all his technical documentation lands inside the models. The logic is simple — if he is inside, he will at least be suggested and code will be written for his system; if not, he will not be suggested at all.

07:22What Matters Is the Form the Model Receives

For a model to find information properly, the information has to be presented to it. The difference between «it collected whatever it could download» and «it received structured information showing exactly what you want to convey» is the answer to why you need your own resource.

08:45Case: Three and a Half Thousand Lectures

Rudolph Steiner has some three and a half thousand lectures in German, open as the heritage of humanity: they may be worked with, structured, studied and translated. The author asked Codex to find and download them.

09:52The Licence Decides What You May Download

He checked the rights in parallel in ChatGPT and distinguishes personal use from publishing on his own resource. Downloading the Russian-language books by the same author turned out to break the licence — so the German lectures went into the work instead, and the system is learning step by step to translate them.

10:49Entities and Connections Instead of Keywords

The task was set around relationships between entities: people, places, topics, events. Places come at different scales — the United States, California, Silicon Valley, Apple Park as a company's location; people are both a name and a role: a child, a wife, a father, a friend.

13:11A Central Subject Versus a Micro-Mention

On ToTheMoon this partly works already: enter a location and you see in how many episodes it was central and in how many it stayed a micro-mention. A micro-mention is ordinary search; the value comes from the structuring.

14:19Training the Model Further Is Not Only for Engineers

Into Codex, Claude Code and even ChatGPT go lectures, notes, photographs, trips, mail from Gmail or Google Docs — structured by semantics, topology, entities and occurrences. No super-fundamental understanding is required: the question is how clearly the task is set.

16:01This Earns Money Inside the Profession

For marketers, salespeople, project and product managers this bears directly on the work, and such resources are raised inside companies in almost any niche. Twenty years ago aggregators appeared the same way — Airbnb, Booking, Skyscanner, Kayak — then the world froze, and now the barrier is low again.

What this episode is about

A solo episode by Alexander Volchek about building your own search over your own material. The starting point is domestic: everyone holds a large database about their life and work — letters, documents, photographs, video, notes, projects, doctors' reports — and when something specific is needed, we often do not even remember where it is.

The first block is what is happening to search in general. More projects are building their own search engines, including ones from people who came out of the large search companies. The immediate occasion is a discussion on TechCrunch about what happens if people stop visiting sites: the author says that is exactly what he sees in himself, because the model summarises the blocks he needs. The whole construction changes — how information is obtained, how time is spent, how advertising works. He prefers the new order to the old one, though quality is still catching up: Google Gemini gave hallucinations a separate block in its release notes, and the model can still hand you an IMDb rating for a film that does not exist.

Hence the fork for a site owner: open it to the models or shut them out. TechCrunch is shut, and the author understands why — if the material is analysed inside the model, the reader no longer needs the source. One logic leads to revenue share; the opposite one leads to an acquaintance of his with a large CMS who feeds the models as hard as he can so that he is at least suggested.

The second block is two of his own projects. First: some three and a half thousand lectures by Rudolph Steiner in German, open as the heritage of humanity. He asked Codex to find and download them and checked the rights in parallel in ChatGPT: the Russian-language books by the same author could not be downloaded under the licence, the German lectures could, and from there the system learns to translate them. Second: search across ToTheMoon episodes, already partly working on the site, where you can see in how many episodes a subject was central and in how many it stayed a micro-mention.

The third block is how it is done and where the money is. The task is set around entities and the connections between them — people, places at different scales, topics, events, roles — not around search. Today, in Codex, Claude Code and even ChatGPT, you can upload lectures, notes, photographs and trips, import mail, and structure all of it by semantics, topology and occurrences, with no super-fundamental understanding required; the only question is how clearly the task is set. Licences are a separate frame: nothing may be downloaded from X, the fines are large, and only the API is left. The episode closes on the business side: such resources are raised inside companies in almost any niche, the way aggregators like Airbnb and Booking were raised twenty years ago.

The episode's point is that personal search has stopped being an engineering problem and become an editorial one. Technically it comes down to setting the task clearly for Codex or Claude Code; the hard part comes earlier — deciding which entities you single out, what counts as a real connection and what as a micro-mention, and on what rights you take the material at all. The author states both limits plainly: the licence that made him drop the Russian-language books in favour of the German lectures, and the ban on scraping X. The value here is not in the volume uploaded but in the structure: ordinary word search the reader already has.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 33 segments: 33 identified, 0 mixed, 0 probable, and 0 unresolved.

Loading…