We have a very interesting topic today, because I can see how incredibly the internet is bursting. If you take the secondary, or third, or fourth market — not just OpenAI and Anthropic, or Google, but the huge number of other startups, companies or people who want to make money on this — they are simply bursting with headlines: the 15 best plugins in this system, the 20 best plugins for that system. And that system runs on some other system, and that system is built on, say, OpenAI or on Anthropic.
And you all know I have made this point more than once. When studying artificial intelligence, I suggest everyone focus their attention as much as possible on serious companies and serious models. If two and a half years ago you could still run tests in ChatGPT, in Claude, in DeepSeek, probably even in Gemini, today you need to focus yourself very clearly. If you can use one model, then use one — ChatGPT or Claude, whichever you like. If you develop software, you can also choose Claude or Codex.
And if you go broader, well, you can use two models, you can pick some other models, or use various middle layers after all. Such as Perplexity, for example, or Hermes, and a great many others. I am in no way denying that many of them have their own solutions, their own agent systems. Or use, for example, various plugins inside these systems, like the plugin for health, ChatGPT Health. Or, for example, two interesting updates have just come out. I think I will make an episode about them. About financial, separate financial systems in OpenAI and in Claude.
These plugins they make — all sorts of them, a huge number of them. If you look at the plugin libraries in OpenAI or the plugin libraries in Claude, they really are very big; they help with files, with spreadsheets, with search. They let you study something in a targeted way: health or nutrition, fitness, finance, education. What is a plugin? In short, it is when there is some predefined room, or some restrictions, or additional instructions, so that your system performs tasks in that subject area better.
So if we take the example of the plugin in OpenAI Health, it is essentially a plugin to which you can connect independent systems for your health, the ones it integrates with. In particular, you can connect Apple Health to it. And when answering your questions, it will use its knowledge base about health. That is, the health instructions it has. And look, here is the question. Today's episode is very interesting, because literally a few days ago Claude made a new feature in Claude Code —
essentially, it made a new command called claude plugin eval, claude plugin eval. And it already has an update. This feature existed, and it already has an update, a version update of sorts. And essentially it is a testing tool that answers this very question. And what question is that? What pays off better: using a plugin to solve a particular task — for example, calling a finance plugin in finance, or calling some particular coding plugin when writing code, or calling a plugin to plan a tourist route or for travel ideas — or not using a plugin?
So, here is the new feature; we will now — I will tell you separately about Claude and separately about OpenAI. Very, very important structural information, because OpenAI released — I think — I will give you the exact date in a moment. OpenAI released it literally a few days ago. By the way, it was on about the same day, I see — they put out a report called "Rethinking skills and instructions for GPT-6 Astra". And it is very interesting precisely for understanding plugins. And if you look at the whole market, at what everyone is writing: 20 use cases you must do with GPT Astra, these things you must enable in GPT Astra, those things to enable in GPT Astra.
Again people start pointing to specific instructions. And I, as you know, am not very fond of instructions, of all these restrictions everyone wants to build around it. And in my concept, in my paradigm — we will soon have a very interesting episode about the main features of the current trend in artificial intelligence. I urge everyone to use the models in their daily routine, understanding that search in its usual form will not exist, and most applications and software in their usual form will not exist.
A huge number of the processes that exist now will change. A huge number of companies and those companies' processes will change. That is, the market is in for substantial changes in the coming years. What I am saying now is not bad… It is neither good nor bad, right? It is a certain given. And arguing with it is pointless, right? Like arguing that people will keep using Google search the way they used to. It is already not used the way it used to be. At the very least there is already a very fast AI answer on top, right? We discussed this topic very seriously just recently in the episode with Ilnar.
So, if we talk about this claude plugin eval feature, what happens there? What happens there is essentially repeated execution. That is, for a chosen plugin — you choose a plugin, and for it… Look, if you don't use Claude, that's no problem. You should understand it on the basis of how it works, including for OpenAI. Or if you use Gemini, or if you use DeepSeek, whatever you use at all, this will help you understand what these additional plugins actually are. And even better, it will help you understand the middle layers and wrappers of different systems.
Once again, I categorically do not support various middle layers; I support them only in complex, strictly very complex cases, right? These middle layers people make: a layer to create a mobile app easily, a layer to quickly make pages in a CMS, a layer for writing text, a layer for building proper storytelling, a layer for creating, writing and processing YouTube videos. Today, simply by using this software — somewhere, of course, things just work automatically, out of habit — you very often — this does not mean you have to give up all such things, but you very often miss completely new possibilities that have already appeared, right?
What Claude Code with Fable or GPT-6 Astra can do now is unreal — in layout or website development, in building the structure and strategy of presentations. That is, it is a new, an entirely new paradigm of everything that is going on. Creating various analytical reports. These middle layers — they and the plugins. These systems, by the way, will support them. That is, Anthropic or Claude Code will keep creating plugins for specialised systems — I don't know, for Power BI, say.
They will keep doing it, because that is their corporate client or corporate customer. So, what does this claude plugin eval do? It does repeated execution. What does that mean? By default, each task is run three times with the plugin and three times without it. Very interesting, right?
We will talk in a moment about repeated execution of various tasks in general. A clean environment — there is such a concept — means separate sessions without your memory, other plugins, or the usual user settings. Next, what happens? A check takes place. The various answers, the files created and the agent's actions are evaluated — in these, essentially, six sessions. And a final score is assigned: the share of passed checks, weighted and averaged over the runs. So, there are various commands inside. We will talk about them. There is a comparison: essentially you can look at it with the plugin, without it, or at the difference between them.
Both the runs and the model's evaluations consume subscription limits. Accordingly, you should understand that each of these runs may cost differently, and they consume your limits within your subscriptions. Subscriptions — or else you pay via the API, the programming interface. So, there are six types of checks. There is such a thing as text match.
There is tool call, there is checking the order of calls, presence of files, and so on. And there is model grading and comparison with a reference. You can simply look at it from the angle of this technical method. I will tell you about the technical method now. Once again, I am not urging you all to run these technical methods. It is interesting for people who are ready to dig deep. Those who have dug in, please write what you ran your tests on and what the difference was. It is very important to learn and see this on the channel, and it is important for other people to see it. And it is very interesting to me too.
And don't forget to support our channel ToTheMoon. Your like, your subscription and your comment — a simple action. It really boosts our momentum, our development on the channel. But these check types are interesting too, because we need to understand what we are actually checking. Because it is not a simple matter. How are we going to evaluate the output at all, especially if the question put to the plugin is a fairly difficult one? So you have prepared everything inside for the run, everything you had. You have made all these settings.
And at the end you essentially get this result. There may be some estimated budget, and you ran it. Here is your result. How do you read the result?
Let's say that on identical tasks you got something like — here is an example for you. Again, I will give you some hypothetical figures, not real measurements. For example, you got 70% without the plugin and 90% with it. That means there is a noticeable improvement on the chosen criteria. So, 70 and 90. If your result came out 95% without the plugin and 95% with it, then no quality improvement has been shown at all. And you need to compare, for example, how much time was spent on these tasks and what the cost was.
If, for example, the result is 95% without the plugin and 80% with it, you need to consider that the plugin may be getting in the way, or some test is set up not quite right. And if you have a situation where it is zero without the plugin and, say, 95% with it, you really need to look into it. Look, when the differences are this strong, you need to work it out: did the plugin give a new capability, or did it simply open access, for example, to the necessary data? For example, you asked it to analyse what happened over a whole month with the number of steps you take every day.
And the plugin has access to Apple Health, which collects the number of steps you take. Without the plugin, it simply did not know. Obviously the result will be zero without the plugin and 95 with it. I think that is clear to everyone here. Because the fact that everything worked with the plugin does not yet mean it worked thanks to the plugin. Yes, that is exactly what this control group without it is for. So, it is important to understand that a plugin and a restricting instruction are not the same thing. Right?
Because we have all seen it on the internet — back at the dawn, remember, two or three years ago, a huge number of people were saying: take this instruction, take these prompts. It will help you enormously. It is important to understand that a plugin and an instruction are not the same thing. For example, in Claude Code a plugin is an extension package. It can contain skills, it can contain additional agents, handlers for various events. It can contain connections to external systems, as in my example with the Apple Health connection. And a skill is, first of all, as we know, an instruction.
Sometimes it comes with some examples, it can come with files and even ready-made programs. It is not a specially trained model. Right? So I would split the benefit here into different situations. And now, once again — I told you we would move on to OpenAI. And here is a fundamental thing — well, probably two fundamental things I want to get across to you about GPT-6 Astra and about the very notion of plugins and instructions. So, what does an extension add? It adds access and various actions.
For example, retrieving documents from a corporate system; on top of that, say, changing a record, launching some specialised tool that is available. And clearly there is a concrete benefit here, because the model gets a capability it would otherwise have had no way of having. And now, if we look at these additional plugins and in places even systems — there are separate huge add-ons in education, or in security, in cybersecurity at OpenAI. These add-ons and plugins can, of course, connect.
Financial systems, for instance, connect directly to institutional data — for example, to daily analytical reports, summaries, various paid systems and so on. Right? Of course, you can supposedly connect your own system to that data separately yourself, and then set it up. I will talk about that in a moment. But what matters is that a plugin can have this access to systems and actions in them. Next. Well, what does this extension add? It adds special knowledge and ready-made solutions.
For example, the peculiarities of an internal system, a proven way of processing, say, some non-standard file. For example, some plugin for working with specialised audio formats. It will open those files for you as well. Or a plugin that does some super-cool archiving. And the benefit here is that you don't have to search for or reinvent solutions, or describe how it should work. That is, the system is fundamentally tuned to this, understands what it is, and works specifically in that area.
And the extension adds a third thing: it adds rules and an order of work. That is, it may contain a report template, a sequence of specific actions or checks the system will have to perform — for example, mandatory approvals it has to get from you, or within itself, in terms of its agent system. So the result essentially matches your requirements. Here is something very important — this is the construction I see: these additional instructions can often end up making the result worse.
That is, possibly, if you don't give instructions, the result will be better. And you can see it in Codex, in GPT-6 Astra and in Fable: a large number of tasks without a description. I remember when Fable appeared, this was very clearly visible. You started writing fewer prompts, and this era of no prompts began. And look, it will only continue. It won't get worse. And thinking that someone can configure a system better, uniquely, somehow, and hand you some neat trick is a kind of abstraction.
That is, your task is to learn to understand how these systems work in general and how to apply them in your daily life. Right? But the important thing is that a restriction is not always a drawback. Here it splits. I am not a huge fan of plugins; I would say I generally try categorically not to use them. Yes, I see problems in them. But still, there is another side, and I want to devote some time to it. Many of you will see usefulness in this. Many of you have already suggested skills to me — skills, a heap of everything. Clearly there are a lot of them.
Once I opened Codex, and it tells me: "You have something like 140, or some number of skills." I think: oh, great. I'm not sure that I created those skills. And right now it says that I have used 1,319 skills. Well, okay. And explored another 23 skills. So the system itself creates something, makes these skills itself, describes them itself within my projects. That is, it does this for itself. But it was not me who said: please create skills, describe them. This happened several times.
And I have to say that I then also run into problems with this. And I have not fully figured out for myself whether these AI skills need to be created at all. We will talk about that now. Or whether very clearly, rigidly described criteria are needed. So, once again: a restriction in these plugins is not always a drawback. And still, we can see that if, for example, a plugin has a written restriction not to send an email without confirmation, that will, for example, slow your system down, compared with when you say "Write an email" and it just goes and sends it straight away.
So the question we have is: are plugins needed at all, or not? Or are they not needed after all, and there is no need to use these middle layers, no need to do all this — no need to go into Notion and keep using the artificial intelligence inside it to process a short table, no need to call Notion from inside ChatGPT for your — well, if it is some small project, a small table — or no need to go even lower, into a plugin — not Notion, but some plugin for organising tasks; there obviously are such plugins.
But rather to drop all that and say: I live my own life. Let me give you an example. Yesterday a friend of mine flew in from Italy — a friend and a partner — and he photographs his food everywhere and says: "Well, I decided for myself that this is the simplest way." And I told him: "But that's the coolest way." He says: "Listen, there are different systems, plugins, separate software solutions — these trained systems could help me a bit more with nutrition, with food." And I told him: "There is nothing better today, no better solution at all, than simply photographing your food all the time, whatever you eat, into an ordinary chat, into ChatGPT, or into Codex if you are connected to it; of course, ChatGPT is enough, because ChatGPT will give you far more of everything.
It lets you turn not to one database, not to two, three or four, but to a hundred, or two hundred, or to different countries, or different doctors, or whoever you want. And it will have this information in… You can photograph however you like, you can weigh things or not weigh them, you can photograph the menu, you can dictate by voice — you can work with this information in completely different ways. What's more, my partner has the Pro version, and no external system will ever in its life run an evaluation of your food through GPT-6 Astra Pro. Never!
Believe me, it won't do that, because the requests would be too expensive for constant processing. Especially if you also keep talking to it, re-reading the whole context, and "look at what I have been eating here for two years". And on top of that, you can load doctors' notes and test results into that chat, while other systems will simply be limited in this. Yes, someone, again, may now argue that this is not so — that such systems will survive, or won't survive. Well, this is just my conclusion. So, OpenAI.
This, by the way, is a good example of using plugins. I do not recommend using plugins in a system like this. It is strictly my personal recommendation. I personally would not recommend doing it to myself, and I won't. And I told my friend that he had actually chosen the coolest path, not a limited one. So, on 11 September OpenAI published a piece called "Rethinking skills and instructions for GPT-6 Astra", and they describe several problems.
The first problem goes like this: long, overlapping skill descriptions make it hard to pick the right one, and overly broad activation conditions. Conditional instructions, the ones that come from outside, force instructions unrelated to the task to be loaded, and, essentially, detailed step-by-step scripts that helped previous models can over-constrain a stronger model. And Ilnar talked about this, about Astra — that Astra works differently. That is, we have to understand that we are in a new era and in a new age.
And we need to start — all of you, starting with today's episode — looking at this completely differently. Hear it once again: detailed step-by-step scripts that helped previous models can constrain a stronger model. Mandatory reading of large documentation before every small change consumes context and time. And by the way, I want to tell you what I noticed in Fable. Interesting, by the way, again, for those of you who design a lot. I noticed this with Fable: it no longer pays for me to open a new chat every time I design a system.
I have a lot of systems going there, and it pays better for me to develop everything in one chat. And recently, by the way, I asked Fable in a very long context window. A very long task — it has been going on, I think, for about a month and a half in one chat. However many times I tried to solve this task in various other chats, it would not get solved. And by the way, I told you about it. Remember, a few episodes ago, that I have a task that won't get solved. GPT-6 Astra didn't solve it again. Blah, blah, blah. I finally solved it. Yesterday evening I solved it.
That's it, the task is solved. So, this task was solved in that old context window — I think it's more than a month old, a month and a half, two. And I asked Fable there just a few days ago, when I switched over after it hadn't worked out with GPT-6 Astra: maybe I should start a new context window? To which Fable tells me: "No, no need. I'm compressing everything here, everything is clear to me." Earlier even Fable itself used to say that, indeed, we are oversaturated now, restart.
Well, Opus 4.8 definitely used to say that. But Fable too, Fable too — I remember it saying so at some point. Now it kept working inside that chat, and it truly solved the task for me. It carried it right through to the end. I liked that very much. And I even see that in Codex I have lots of chats hanging there — and in them, because they are cloud-based, GPT-6 Astra Extra High or Ultra is not available right now, but 5.6 Sol Extra High is. Probably because they run from a different Codex, from a different computer.
It's specifically the part of the tasks I use constantly — I use them deliberately, because opening any new task causes some kind of, well, some kind of paradox in the system every time. Rather than you simply developing inside the system. I like that very much. That is, I really want us to get, in systems development, to the point where, say, you develop some system or run a company and can essentially run that company through one window. Clearly, this one window probably shouldn't, at some point, be presented as a single chat.
It should be some kind of — some specialised interface will turn up; we'll see what it looks like. A chat-like interface where things can pop up and open, with details here and there. All these systems are trying to find it — as, in particular, in Claude Code, where a list of tasks to be delivered gets created, or as Codex does it with steps. It writes how many steps it has, what plans it has set, what goals it has set. Remember, this appeared in 5.6 Sol. And it's a very interesting thing.
They propose an approach: a short description, a short description of purpose, loading details only when needed, and regular review of old requirements. It is also important to keep in mind, by the way, that one instruction may be used by different models. What is useful for one is not necessarily useful for another. And in this, by the way, I see a very big problem — the collision of Codex and Claude. When Fable does something and seems to say: "Great." And then I go to GPT-6 Astra, and it reports that something is off. And they have these different… It's because there is also an instruction there; it didn't analyse all of it, it analysed some part.
And you sit there thinking: so how are the two supposed to get in tune with each other? And it's not a simple matter. And it's important to keep this in mind: different models understand an instruction differently. Different artificial intelligence systems understand an instruction differently. Just the same as different people relate differently to what would seem to be the same things. And by the way, what I am telling you now are OpenAI's official engineering observations. That's important.
And not some study that, I don't know, gave some universal percentage of degradation. At OpenAI there was another story: at the beginning of the year they released a report called "Systematic testing of agent skills" — "Testing Agent Skills Systematically with Evals".
It proposed evaluating not the persuasiveness of an answer but four aspects of the work. That is: whether the result was achieved, whether the necessary actions were performed, whether the formatting requirements of the task were met, and how many resources were spent. And for a first set, 10–20 target requests were enough, including ones on which the skill should not activate. And there Codex was run while saving a structured log of actions, then programmatic checks analysed the file and the commands, and a separate model evaluation checked the more subjective requirements.
It also proposed tracking token consumption. And essentially this is a practical method for building checks, not a promise that a skill was useful. So what I'm telling you now, why I remembered this case — this story isn't… There was a case, by the way, back with GPT-5.4: an experiment with excess context. Either all the descriptions were loaded at once, or the needed ones were found on request, inside. For example, when you tell it some story about your health and give it this big context.
Or I, for example, constantly load, I don't know, an entire ToTheMoon episode — imagine, all the transcripts. Or I load some books all at once, or bit by bit. And there, too, the system measured this. But those are all old systems. I'm just saying the problem is not new. And there was an independent study, SkillBench.
It was done in the summer on 87 tasks and 18 agent-model configurations with various prepared skills. And these prepared skills raised the success rate from 35% to 50%. So, roughly, 15% there. That is, they found that specialised skills really did help in many cases — substantially helped, by the way, to improve everything else. Especially certain sets helped, for example, with efficiency. And I want to say that… And there was, by the way, a study of Agents.md files — roughly speaking, agent files for building systems.
And the conclusion was this: instruction files for working with repositories did not improve the success rate overall, but increased computational costs by more than 20% overall. This, by the way — pay attention — is a mind-blowing topic! Because all the systems were saying you have to teach these Agents.md files how to work, load them up fully. Everything. And it turned out that — as you can understand it — merely adding project descriptions and requirements did not guarantee improvement.
Although, once again, the authors acknowledged the usefulness of such files for non-standard development rules. These are just some details for you.
The story here is that Anthropic itself provides information on when a skill becomes unnecessary. The first type is skills that compensate for shortcomings in the model's capabilities. For example, some special technique for creating documents. As the model improves, they essentially become unnecessary. And we all knew this: earlier, for example, plugins were released that polished text better, because the model couldn't do it. And now the model can do it. Why use these third-party systems? And this is what I repeat everywhere; I talk about it a lot.
Models will develop incredibly strongly in any case. Even if some feature doesn't work today, at some point it will start working. I remember a time when it was impossible to upload an archive — it wasn't analysed — or some large documents weren't analysed. Now all this is becoming, essentially, a very relative problem. And the second type that Anthropic identifies is skills that capture your preferences and workflows. That is, the model may be perfectly able to perform every step but not know which order, format and criteria your particular team needs.
And accordingly, such skills can stay useful longer. And you have to understand here that if you have some fundamental internal things of your own, and you've launched a plugin that works from outside, you don't fully understand how it works, and it has third-party instructions, then it is more likely to do you harm than good. That's how ChatGPT Health harmed me. The integration. A very strange construction. What's more, when I then started asking health questions in other chats, it kept trying to load this plugin until I switched it off.
By the way, a new ChatGPT Health update has just come out, and this setting — load or don't load the plugin — gets set somewhere by default. And it turns out I essentially can't ask a single health question in ChatGPT without it using ChatGPT Health. Or I have to keep telling it not to use it. That is, all these connected plugins can only create problems. For example, I remember a situation: the Photoshop plugin is connected, I ask it to do something, and the system keeps trying Photoshop, saying: "Opening Photoshop for the changes."
I say: "Just do it yourself without that." And here we get this: on the one hand plugins can help, on the other hand they can restrict.
And here is the point: for you to truly understand how models work, to truly figure out what this is, to truly live through new cases, to start working with artificial intelligence not as with ready-made software, which we have all been using anyway for years, for decades and decades — for that you need to work directly with artificial intelligence all the time. Without plugins you need to work with the latest frontier models, without plugins you need to work inside these systems, use ChatGPT in your everyday work, in your professional activity, in your ordinary life, try to do something in Codex or try to do something in Claude Code, not limiting yourself to professional activity only, not limiting yourself to personal activity only.
Only this will let you — the most effective path, at least the one I see for myself — will let you really, truly, seriously understand artificial intelligence, understand how it works, and apply what you need in your everyday life. Without any contentious, pointless discussions that nobody needs. And at the same time, if you want to test some additional plugins, just test them with an understanding of what they are, not with some limited trust and without understanding how these systems work.
Because for the most part you can accomplish these same tasks yourself, unless there is some unique, extra-special exclusivity. Or, for example, unless some plugin gives you access to a database you can't connect to otherwise. What do you think about this? Be sure to write in the comments, and see you in our next episodes.