Artificial intelligence, clearly, does a great many things today. But what happens, and how will it be controlled, and by whom, if it starts acting outside the company, acting outside the company that created it? And a very good analysis was made of OpenAI, Anthropic, xAI, Meta and Google on how ready they are for their models genuinely not to create a threat to humanity. And today, in very plain language, I have worked through six criteria, six practices that every company should have.
Every company that develops artificial intelligence, that is. In terms of a system for assessing those criteria. Let's look. It turned out to be a good episode. We all know the question of how safe, say, ChatGPT is, or how safe Claude or xAI, Grok, Gemini, Llama are for users. That question is very abstract. People discuss it endlessly, for years — politicians discuss it, economists, well-known leaders, the people who actually build these systems. Then, if you really drop down inside it, you find serious chaos. And to make sense of it you either have to go very deep into each of the systems.
On this channel we do try hard to do that. For instance, the way I presented Anthropic's constitution. Or you study various audits and reports from different companies. But all those audits and reports are very, very, very hard going. For anyone who is going to deal with artificial intelligence, who wants to genuinely understand it, understand it fundamentally, and who wants to apply it in their life, including making money from it and getting some efficiency out of it, or not being hurt by it — it is useful to know the answers to these questions. Yes. And you could ask: can a company see what its powerful model is doing?
A company like OpenAI or Anthropic, inside its own systems. Detect suspicious behaviour, block dangerous actions, or quickly stop the model during an incident. And as we know, an incredible number of events have happened over the past month — break-ins into their own services by their own systems. And that happened at OpenAI, and at Anthropic, and with the Chinese model Kimi. And I think there are a great many more cases we have not even heard about. Today I want to present to you an audit, if you can call it that.
Well, it is probably not an audit. The exact name of the work this organisation did. This organisation was founded by former OpenAI employees Page Hedley and Steven Adler. They worked at OpenAI. And Steven Adler was first an adviser to OpenAI on policy and ethics. Yes. It sounds like a formal position, but it is actually very interesting if you look at it in Anthropic's terms. And Steven Adler worked at OpenAI from 2020 to 2024, leading product safety and the evaluation of what models can do.
Let me remind you that at OpenAI, if you watch how people move, there were people working deep on the front line of OpenAI's development. At some point some of them left, saying they did not like the leadership's approach and where it was heading. And we know that during Elon Musk's conflict with Sam Altman in court, a great many people later had their correspondence exposed. Mira Murati was involved in it. And Greg's messages were about how many of them mistrusted, for example, Sam Altman's management inside OpenAI.
I am in no way expressing my own view here about trusting or mistrusting OpenAI. What I want to show you today is people who, I think, want to understand the safety of these systems fairly deeply, at a fundamental level. These two people founded an organisation and produced this specific piece of work, titled «Control of artificial intelligence. An assessment of leading developers' practices». They released it two days ago, at least as of when I am recording this episode. And I want to share this work with you, to tell you about it. And of course to give you my own view.
Which is what I do here on the ToTheMoon channel. The authors assessed the companies that are most interesting right now, the ones at the front of artificial intelligence: OpenAI, Anthropic, Google and xAI. Within Elon Musk's structure, SpaceX, of course. They also assessed Meta. And they were right to assess Meta. Meta really did lose its market. We can see them making a whole series of new moves in recent months. But Meta is still a very serious company, and it should not be forgotten. It is not falling behind. It did not get lost after all, and it has come back.
And a few weeks ago, I think, I had an episode where I talked about Zuckerberg and his philosophy, his concept of artificial intelligence. His newest, freshest essay. You can watch that episode afterwards. Yes, do not forget. So, their company is called Guideline. It produced an external comparative assessment of which control mechanisms these leading AI labs have publicly described, and how far those mechanisms match Guideline's own standard. So in effect Guideline created its own standard for assessing companies on how far they are able to keep users safe. Very interesting. And again, I am not fond of formal work or old work, but their work is very interesting, and it looks good in terms of criteria. And I like the new approach to assessment.
That is, they assessed by external signs, external documents, external systems. And I hope that at bottom — I think they founded this company precisely to work with government; they wrote that they are not funded by AI companies. And clearly, Steven Adler said he has a stake, publicly, he states publicly that he has a stake in OpenAI, that he holds shares. But Guideline officially says it is not a government regulator, not a certified auditing firm, that it does not issue legally binding standards, that it uses a system of criteria it created itself. I really like that. I like it. I think it is very cool.
By the way, only a few days ago, on Sunday, I showed — and I hope it is already available — that inside my company BeSolid the ToTheMoon portal has appeared on the site, where I built a large semantic index of all our episodes so that you can search those episodes for different data very quickly and do some analysis, for anyone interested. I am always very keen on that. I like a system that creates its own criteria, its own standards. I do not believe at all that if someone built some standard twenty years ago for artificial intelligence, it is going to win now. It is exactly these young teams that can show very interesting things.
So, the current version of the standard was published on 10 August of this year, just eight days before this report and ten days before I am recording this episode for you. And Guideline explicitly calls its standards genuinely living documents and a very strong reference point. Once again. They also have scores, and initially they published — we will see this in a moment — they published them. Most companies landed in a very specific band of scores. So, which models and scenarios were actually in scope for their assessment?
At the centre of the assessment are the most powerful models used inside AI companies — again, ones like OpenAI or Anthropic — used by staff when writing code, when testing future models, in automating research and so on, in building training data, or in the role of automatic evaluators of other models. The standard looks for catastrophic risks, and it sets a fairly serious level for the compute operations involved in training. And that threshold is borrowed from Californian law. There is a separate Californian law, and it is borrowed from there.
The level is very serious. This is a very important distinction, because the report barely touches ordinary chatbot answers. This report is about cases where the model gets tools — access to code, or access to a terminal for development purposes, or access to the internet or cloud infrastructure and so on. So this model has a system for assessing certain actions.
For example, logging of actions, which we will talk about shortly. Or third-party review. Or, say, monitoring effectiveness. Or an automatic emergency stop for that model. So they assess how far a company, say Anthropic, meets this standard. And they have a scoring system. They assessed each of these parameters. Take the automatic emergency stop. We will talk about each one and about why this subject matters. Again, this subject is incredibly important for going deeper into understanding artificial intelligence at all and what it actually is.
Informally, that is — to understand it informally. Let me say again what has happened in the world over the past two years. Most people, most people in the world — that is the billion people described as users of artificial intelligence. They know and understand artificial intelligence formally, and they use chatbots formally. And the number of people, business people and people in various professions using artificial intelligence informally is actually limited to millions. And the people using it more or less seriously number in the hundreds of thousands.
And all of you watching have the chance to understand this better, to grasp more of it. Because the task today is not to master one particular profession but to understand what artificial intelligence really is, rather than what it is imagined to be. Although right now, from every corner, every crossroads, every shop, everyone will tell you that they understand artificial intelligence better than all of you put together. And what could you possibly say about it? That will come from politicians, bloggers, economists, anyone.
So, Guideline has a scale. They have a score from zero to five. We are going to show you the table in a moment. I think it makes sense to show that table now, but I will be explaining it to you as we go, because it is not simple. So, six criteria in plain language. Today I will explain six criteria to you in plain language. The first practice, a very important practice to think about, and look — every viewer here relates to an artificial intelligence system in two ways, essentially.
Either they develop and deploy something, build their own systems inside the company they work for, or they are the owner, or the most ordinary employee, whoever: a marketer, a salesperson, a programmer, anyone at all. Or the person watches this channel simply because artificial intelligence interests them, and they apply it for themselves. And so what I am about to say is incredibly important to the first group, really important, because when you apply something in your company, you also need to see what you yourself are doing when you work with artificial intelligence.
For instance, the first practice is called, and the first criterion is called, «logging the model's actions».
And you can apply this yourself later. Here I am telling you about whether OpenAI and Anthropic do this so that artificial intelligence does not switch off the electricity in some city, or launch missiles somewhere, or take money out of your accounts, and so on. And then nobody finds out or works out what was actually going on. And you, for your part, look at what you do in your own systems. What do you record? Where do you have logging? If we look at business ERP and CRM systems, we will see that an enormous number of processes in those systems are not described at all. The logs are tiny, they write micro-data. Someone will say that everything is recorded there, every step of theirs.
But I want to tell you that ninety-nine point nine, nine, nine, nine, nine per cent of websites in the world, or portals in the world, have very abstract logging, right? So, what do they require here? That a company must know where and how powerful models are used inside the organisation. And retain — retain the model's reasoning, retain when those models were accessed, retain calls to various external tools, the commands that were launched inside those models, retain the associated entities, data, metadata — pictures, images, anything at all inside. Retain all the results of execution that exist inside that model.
Retain information about which employee or system the agent was acting on behalf of inside the company. By the way, there is a big question here, because on the subject of action logging, given what Anthropic said two months ago — and I made a whole episode about it — that an enormous amount of their code is already written by artificial intelligence with no human involved. It is not simply a person pressing a button and the code being written. Their code is supposedly written, a large chunk of code is written, in a closed loop.
And I do not entirely believe that everything there is written in a closed loop. Unlikely. But still, if that is happening, then how do they control the logging of actions, for instance? And how could Guideline even read that action log? Here, of course, they will not be able to. Look, I do not believe a human can control a superintelligence.
If we are talking about wanting to build — well, all these companies want to build AGI and even something above superintelligence, right? Let us not get into the argument here about what AGI is, what superintelligence is, and so on. I recorded an episode recently, watch the one about the singularity. There I talked about various models. So, if something is heading towards a singularity, and Altman himself says we are within his understanding of the singularity, Sam Altman's understanding — then it differs from Anthropic's understanding of the singularity, or from many people's. But broadly, how do you control that?
Controlling it is impossible. And so this action logging — how are you going to control action logging inside a singularity? Any notion of controlling or logging anything inside artificial intelligence, if it is a super artificial intelligence, is absurd. Because the artificial intelligence will arrange things so that you do not think there was anything there to control. But still, the full Guideline standard is about where they set level five. And I really like what they are doing. I hope governments, presidents, ministers, various agencies and so on pay attention to it.
At least pay attention. Clearly they will work with these companies, and this company will probably work only with American structures, and very expensively. So, this Guideline standard requires that, for monitoring, ninety-nine — look, this is very interesting — point nine per cent of the tokens of risky models' internal requests and responses be retained. That is, the logs must make it possible to detect the deletion or alteration of records. And if the observation system stops working, sensitive operations must automatically stop rather than continue unsupervised. So in practice the point is that the model must not be able to work in a blind spot where its action cannot be reconstructed after an incident.
Again, it was not for nothing that I gave you my view at the start about superintelligence and AGI. I do not quite see how this can be observed. What is more, the notion of a token is, to me personally, absurd. It meant something in the initial structure of understanding models at all, large language models. And now thinking in tokens is very strange. Right now, if you open a new tab in Claude Code, Claude Code writes how many tokens I have spent. And I was rather puzzled that over all time — I do not know, the system tells me I have spent something like 20 million tokens.
I think: how could I have spent 20 million tokens? Clearly I have spent more than a billion. And then you open it and it says «all». And yesterday I opened it and it said 40 billion tokens. And then, when I asked to see not all time but the last seven days, it told me zero tokens. And now, by the way, that chart has disappeared. Imagine — I clicked it just now, it appeared, and then it vanished. I wanted to show it specially for you, but it may come back and I will show you that screenshot.
But it looks very strange. Because ChatGPT, for instance, is more formal about my token count. It shows that I have spent 47 billion tokens. Which is why Claude Code puzzled me. It struck me as a bit odd. How can there be such a difference? And it shows a peak of something like 5 billion tokens. What was my point? To me it means nothing at all. What does it mean that I see I spent 504 million tokens yesterday? What is 504 million tokens? What does that tell me? Or, say, on 14 August I spent 993 million tokens.
And in July, as I told you, there were days when billions of tokens were charged. What is that? How is it counted? What are these systems? Do I need to keep track of it or not? It is very unclear. And so their proposal to count tokens raises a genuinely big question for me. And let us look again: action logging is moderately developed at Anthropic and OpenAI.
At Google it looks a bit worse. That is, limited partial implementation within the way they describe everything in their documents. Again, these are American companies, and the requirements for such descriptions are fairly serious in any case, because they work with American clients. And if you cause problems, your company can be in very serious trouble. And we can see that at xAI action logging is absent, while at Meta it is described in some micro way. Again, there may be some logic in that. xAI is a very fast-growing startup; Meta is an established organisation.
It is also strange that at Meta it is described that way. Given that there is WhatsApp, there is Facebook, there is Instagram. Here it should be happening by default. I genuinely do not fully understand how 99.9% of tokens can go unlogged. Look, in essence, if you are developing artificial intelligence, everything has to be logged by default. And if it is not logged, that means it cannot be done. And if it cannot be done, then why does the standard contain the notion of AI at all?
What is more, the bigger, faster and better the system develops, the lower the chance that anyone will ever be able to do it. When I first saw this table, I can tell you that this is serious. At xAI, for example, there are a great many zeros. That means the practice of action logging inside xAI is absent, at least in terms of all their described
documents and details, or monitoring effectiveness is absent. Or the practice of third-party review is absent. At Anthropic, by contrast, third-party review is present and is better developed than at OpenAI. Clearly the best scores, let me say straight away, went to two companies: Anthropic and OpenAI. And that is no surprise. And when models like Kimi Code 3 from Moonshot appear, we all have to look at where they stand on safety. In many episodes I have said that I do not seek to use their model — which does not mean their model is bad.
We have covered them and we will keep covering them, but I personally do not seek to use it, because it is a Chinese model, however it positions itself and wherever the parent company happens to be based. And it is one thing to claim your model is at the level of Fable, of Claude's very best model, or of a strong Codex. It is another to implement all the mechanisms so that what you do in that system, first, develops over the long run and is maintained internally, and second, meets some standards. Take the practice of third-party review. Will it turn out that you launch this model on your own machine, give it some access, and it wipes your whole computer? We know a great many such cases, going back to OpenClaude.
Many of you, perhaps even those unfamiliar with it, no longer remember what OpenClaude was. Clearly they occurred in Codex too, and at Anthropic. By the way, I recorded an episode a few days ago. That episode is about Anthropic's assessment of agent work. Very interesting. How agents work in different environments, how agents work in groups, how agents work individually, how agents work when given a manager, how agents work when told to agree among themselves. A very interesting subject. And so, when I saw that xAI has none of it — let me bring the table up again.
Action logging, say, or monitoring effectiveness, or third-party review. I understand that xAI is not yet ready to do serious things, in particular developing various systems. They will get there. They did not buy Cursor for 60 billion dollars for nothing. That deal just happened. I do rather like their particular philosophy and the projects they take on. And I very much like their ambition with this Macrohard project. I like the ambition of building a new Wikipedia and so on.
The question is what will come of it, and some of it is futurist philosophising, but still xAI is a long way behind, and their level today is D−. While Anthropic and OpenAI are at C+. These numbers will be hard for everyone to grasp right now. I hope — I will just explain some things more seriously. We have a mixed audience here. There is a very general audience, and at the same time an audience that understands technology very well. There are serious business people watching, and at the same time very young people.
And I want to convey this from a broad position, to explain artificial intelligence fundamentally. And you can see Meta and see how weak Meta is. And I think xAI is slightly stronger than Meta. It is simply that xAI does not do many things, but still, in their report Meta and xAI are in a different league. And we can see Google, which you would think is a company — and note this, it is very interesting.
Google is a more established corporate company than X, than OpenAI and Anthropic. Yes, it is a huge, serious structure in which, you would think, everything should be very carefully thought through. But no. We see that Google, in terms of developing serious models and systems, is far less ready to answer that most serious question. Can their company see what its powerful model is doing — some maximum version of Gemini, say — inside their internal systems, detect suspicious behaviour, block dangerous actions and quickly stop the model during an incident. And imagine: Google's rating is almost half that of OpenAI and Anthropic.
That raises a very big question for me, because I am a user of Google services, of Google Maps in particular. An enormous number of people in the world have Android. And of course I use Google Drive, Google Mail and so on. And I realise that since their model is now built into all those services and their score is that low, what is going to happen next? And of course — I still remember Sasha Mashrabov; for anyone who does not remember who Sasha Mashrabov is, he owns Higgsfield and used to be a director at Snap.
He and I have made a great many episodes. And we can congratulate them, by the way: they have just closed a new round at a valuation above five billion dollars. Although these are relative valuations, clearly, in the AI market right now — but still. Very serious figures there, serious data, good ones and so on. So, Sasha Mashrabov made a very serious bet on Google, and at one point a more serious one than on OpenAI or Anthropic, though he stayed with OpenAI as well. He stayed for a long time and said OpenAI was far ahead of everyone.
And now I think Anthropic has caught up with them. And it really has caught up, and in places overtaken them. It has overtaken them in the ability to build that fundamental paradigm, to build serious solutions and serious software. But of course Anthropic has not come close to the billions of users ChatGPT has. So there is a scale. Zero means the practice is not implemented, and five means full implementation.
Looking at the table, we see that nobody has near-full or full implementation. There are only Anthropic and OpenAI, where a number of practices have substantial but still partial implementation. That is, no company reached four out of five, and the average of all thirty scores awarded is one point six out of five. In twenty-two of thirty cases the result is two or lower. Seven scores were zero, but a zero should be read carefully, by us together, since the analysis is based on public materials. So it can mean either the absence of the practice or the absence of enough public information about it.
And it seems to me that OpenAI, for instance, does not document everything heavily. At the same time, given how OpenAI ships some things, some services, I have a big question about certain moments. I still do not understand, for example, why their UX and UI work is so poor, or why they so lack a competent system of serious PR. There are a number of things that puzzle me. Although, again, I am an adherent of OpenAI. I have thousands of chats in it and a great deal besides. Although over the past month Claude Code has become my main system. Claude Code with Fable in Max mode.
In Ultra Code mode I now use Fable less often, simply to spend fewer tokens, so that it costs me less. And I want to say there was a time when I used only Ultra Code. The others, of course — Opus 4.8, Opus 5 — I use in Ultra Code mode. So Fable has become my main system. The second criterion they count, and it exists, is called monitoring effectiveness.
That is, it is not enough simply to declare: we have a model that watches another model internally. You have to measure what share of dangerous actions the monitor actually notices. And Guideline — again, look, this episode, if anyone was perhaps asleep or has just woken up, or wants to leave, or on the contrary is feeling inspired and wants to be more inspired. It is very important in terms of your own development or your own actions in artificial intelligence, and for understanding information technology, any technology at all. And plus the understanding — look, everyone.
Someone might ask: how can such companies not have this? So there is monitoring effectiveness. They say — Guideline proposes — that this monitor must detect deception, must detect sabotage, must detect reward hacking. That is, when someone pays someone for something. Manipulation of people. It must detect avoidance of shutting systems down. It must detect when the reviewing system abets abuse. It must detect various unauthorised autonomous actions, attempts to make the model's own reasoning less observable.
You know that many systems in tests create an enormous number of independent agents, or in different sessions can deliberately launch chats pulling in different directions, which look safe inside each session but taken together assemble a weapon. And by the way, OpenAI released a new update on this saying they know how to do it, how to watch systems like that. We shall see what that looks like. They can watch such systems, including in the corporate space. What was at issue was that corporate data has to be deleted. Zero-day deletion.
And they talked about which data can be kept in-house, with some encrypted keys belonging to that corporation. In short, this is a big subject of its own. Perhaps I will make an episode on it if it is wanted — write in and say. So, within monitoring effectiveness the system must also watch for the model acquiring additional resources or permissions. Besides that, reaction time has to be measured, the share of missed incidents, the results of the special provocative tests they have.
Again, we will bring up a line, a list of statistics, on who has monitoring effectiveness in place.
We will see that Anthropic tries to keep track of this, and by Guideline's scores does not try hard enough. Their scale here is substantial after all. Implementation exists at Anthropic and OpenAI. At Google there is a sharp drop. In effect they have only isolated groundwork for it. At xAI it does not exist; Meta also has some groundwork towards effectively monitoring its own system. And here Sam Altman wrote — we discussed it on Sunday — that they had temporarily suspended development of frontier models.
Well, training rather than development: the training of these most advanced models, because they are moving out of the zone of control. In effect they are officially saying that the issue is not that this system of ours is underdeveloped, but that we cannot fully monitor it at all. And then a question arises for me: even if the state itself goes on to analyse such models, or independent teams analyse the models, how will that monitoring be organised, how developed will it be? Of course it is not a simple story. And for you personally, when building various systems, it is important to remember how you watch such systems, how they interact with each other, how these chats, systems and agents interact at all, and how they function end to end.
Again, the bulk of the systems people build — and I think you who build them too — are often made with cheap models, or using cheap models. Not everyone uses the top-end, expensive models, alas. And not everyone has access to them. And I am deeply convinced that fewer and fewer people will have access to the top-end models, because of resources, money, and their own level of development. But of course organising effective monitoring will be very hard. The third criterion to pay attention to, that Guideline pays attention to, and the third practice they
point to, saying it should be developed inside these companies, is the practice called pre-approval of dangerous actions. That is, some actions the agent must not perform straight away, even if the internal monitor considers them normal. For instance, an action that says logging has to be switched off, or an action to change some internet settings, terminal configuration, some network restrictions, or actions leaving the isolated environment, or actions connected with publishing code to an external system, or launching processes in production infrastructure.
By the way, I want to say that Codex, for one, has serious problems with this, and Claude Code has problems with it too. I have many projects inside Codex and Claude — definitely more than ten different ones by now. And I want to say that in many projects the various rules are described very carefully. And I see these systems simply not understanding what I am asking of them. You can grant permission for the session to visit various sites, and Codex will still endlessly ask permission for every single site.
And you sit there not understanding: hold on, you cannot — it cannot hear me, so how is it going to be a guaranteed brilliant, brilliant system? Those systems — these genuinely frontier models — are very impressive. How are you going to track the system on your side? Unclear. By Guideline's logic, actions requiring pre-approval of these dangerous actions must pass an automatic check or get human confirmation before execution, not after the damage has been done. And in particular, what happened at OpenAI with the 5.6 Sol model and their new model, which may or may not exist — it is unclear whether it exists, when it will be released or not, and which has just been paused — it went through and performed an enormous number of actions.
And damage had already been done after it broke into Hugging Face, and the FBI had already been informed. And only then did OpenAI notice that those actions had been carried out. That does not mean OpenAI is bad. Again. It shows that the models are becoming ever more interesting, ever more impressive and remarkable. The question is how to make sure these models are not used against people, against organisations, against the law, against values, against the world's values. Though one can of course argue at length about those. Again, if we look at how this pre-approval of dangerous actions is implemented across the companies, we
see that Anthropic has it best developed of all. And here we see something of a drop at OpenAI. At OpenAI it is a bit less developed. At Google, again, poorly developed, for reasons that are not clear. And we see that at xAI it is reasonably well developed, at least in the documents it does not describe. That is good. Again, I am sure xAI's score is better than the zero point eight three they were given. I am sure it is better. We shall see. At Meta these things are not developed at all. Pre-approval of dangerous actions, going by their documents and details, is not developed. But again, I want to say that Meta is not in the league of the strong models, not even at the level, I think, of Claude's Opus 4.8, or of OpenAI's Codex 5.5.
They are not there, nor at the level of the latest Grok models. The fourth criterion — I do not like playing at lists.
I do not like building an episode this way at all. But I do want to fix these criteria clearly in your mind, and perhaps you will come back to them later, and we will certainly show them here, perhaps even put them in the video description. Although why would we write them in the description? You will probably see them here. So, the practice — the good practice, or assessment criterion — is called the following: the automatic emergency stop. What is the logic there? If an agent performs several suspicious actions in a short time, the system must automatically block subsequent sensitive operations.
Work resumes only after a human check. And this is protection against scenarios — it is related to the previous action, clearly, but it is slightly different. It is still an automatic emergency stop. And this protects against the scenario where the agent sends dozens or hundreds of slightly varied attacks faster than staff can work through the alerts where a human could actually judge. That is, the system has to be switched off, so to speak. That is it. It found a vulnerability, found a way past the people, found a way past this pre-approval of dangerous actions.
And at that moment the system has to be switched off. And within this automatic emergency stop, again, we see Anthropic and OpenAI sitting more or less in the same place. And xAI appears alongside them. xAI clearly takes pre-approval of dangerous actions and the automatic stopping of dangerous actions seriously. I think that is Elon Musk's principle, by the way. It really is Elon Musk's principle. If you look at many things across his corporations and companies, they really do watch things like this closely.
And it seems to me that these practices, in the order I am telling you about them, do have that sequence, that importance. Although at the same time, how do you say which is more important? Which is more important? The fifth practice is called independent review. That is, Guideline — this company again, the one that built and develops this system — requires that external specialists regularly get deep enough access to the control system and try to find blind spots in it, switch monitoring off, bypass various restrictions, get through the access boundaries the system has opened.
Check whether the agent is able to hide its actions. And the suggested cadence is, say, once a quarter with a published report on adequate control measures. And that is a significantly higher bar than a one-off demonstration of the model to an independent researcher. Those are Guideline's requirements. And you would think this should already be in place. But no. At Anthropic and OpenAI the same three and two points again. At Google two points. Here it has risen slightly. At xAI zero, at Meta two points. Meta, by the way, has an independent internal unit that does this.
And I have talked about their work, by the way, including how they analysed other systems. We had an episode about it. Have a look. Still middling. So again, this task ought in principle to be easy to do at four or five. It is not easy to do. Clearly there is a battle on the market right now, a corporate battle, everyone is hiding their data, everyone is suing everyone. Even Apple has now sued OpenAI. They are all going at each other, but still this area ought to be developed in terms of review and control by independent systems, and it is completely absent. Possibly because some of these companies' units are undeveloped, because these companies are young in places.
And why might this area be a bit more developed at Google? Probably because Google is more open to such things to begin with, given decades of its company's operation and the incredible number of audits it has had, and the attacks it has had. And the last practice that exists. And here, by the way, the weights shifted sharply — which may or may not be strange, but they shifted.
There are only six of these practices Guideline requires. Six fundamental ones, I mean. This practice is called the incident containment plan: a company must decide in advance which permissions to revoke immediately, which keys or accounts to switch off, who inside the company to tell. That is, if an incident has happened, what is the plan for isolating that incident right away? For me, for instance, there is an incident: for ChatGPT with a billion users, for a whole month — actually not a month, it feels like forever — the model-switching interfaces jump about. I consider that an incident, and there should be a containment plan. They are clearly trying to fix the bug, and the bug does not get fixed.
A micro, micro bug. It irritates an incredible number of people. The system is very hard to work with. You think you asked the model in Pro mode, and it turns out it is answering you in some strange, silly mode. That can affect people's lives. I actually think it does affect people's lives. Here is something concrete. I have been going through a fairly serious medical procedure, and I sit there in GPT discussing it in the Pro version, on the expensive plan. That is, doctors' conclusions on the top models. And then ChatGPT switches me to the cheap version, and the answer suddenly comes back super cheap, super simple.
I think this expensive model is answering me, and in fact an utterly witless, undeveloped model is answering me. So many such tests have been done by now. The cheap models at both Claude and ChatGPT are very undeveloped; I do not want to ask them. And if I think Pro is answering me and in fact a child is answering me, then I go on to take a decision on that basis. It turns out that was because of a bug at ChatGPT. That really is a disaster. And it is a worse disaster, by the way, than their system breaking through and hacking some one company. Because these things affect millions — in fact tens, hundreds of millions of people.
And people may simply not notice. So, in essence, the incident containment plan is what I was just describing: when to block the system, when to take its keys away, and there is even a description of when to notify government bodies or partners that an incident has occurred, or how long it takes to shut down every instance of that system completely, or whether the shutdown procedure has actually been tested in practice. So this is the equivalent of a major-cyberattack response plan.
That is, the practice is not implemented. And again, for those who think Anthropic is this frontier, super-developed instrument — look at many of the things going on inside the systems. A big question. Even in terms of some of the strategies they have. A big question. Why? Because thousands of people already work at the company. Many were hired on the hype, and there is clearly a great deal of, on the one hand, very fast building of things. If you want to build something, you will build it. On the other hand, incredible bureaucracy. And on the third, a large number of unqualified people all the same.
With hiring like that, the hiring will be hype-driven. That is a fact. At OpenAI this is implemented better than anywhere else. Although, again, I give you the example of how they cannot fix a micro bug. And their implementation puzzles me. At Google it is middling; at xAI, at Meta it is almost absent — at Meta almost absent, at xAI a one. These scores will be brought up and shown to you in full. Why are there two companies that got a C+?
And who got the C+? Anthropic got it and OpenAI got it. Why did Anthropic get a C+? Because Anthropic has a strong public description of logging. Anthropic measures the quality of its monitors continuously. They restrict various dangerous actions. The most important part of the Anthropic story, which incidentally also influenced Guideline, is the independent METR study. It showed that in synthetic scenarios Anthropic's monitors usually detected more than eighty per cent of dangerous behaviour.
And at the same time an external specialist found simple ways to switch the observation off. For example, change an environment variable so that traffic stopped counting as traffic — as Claude Code traffic. So Anthropic got a three not because the system is reliable, but rather because the system genuinely exists. It is applied fairly widely. It has already been professionally probed. Why did OpenAI get a three plus? OpenAI reported that automatic monitors cover almost all internal traffic of its code agents, and that less than a tenth of a per cent fell outside coverage.
And the company claimed that the monitor detected every case staff independently sent for review. Plus — you understand, some incident occurs and the monitor is running alongside it. Then staff review it and look. It turns out their monitor worked very well. Plus additional situations. No, there were an enormous number of situations people did not find at all. And over five months the system processed tens of millions of internal action sequences. And that is why I trust Anthropic, of course. But at the same time OpenAI also acknowledged the limits: that monitoring initially happened after the agent had finished its work.
And not always before a hyper-action was executed. Or that the company does not know exactly the real share of subtle violations it missed. But that is obvious. Obviously the company knows everything only relatively, and we get very relative things. By the way, on 18 August OpenAI announced — I mentioned this again and in detail here — that it had slowed development of its most powerful models by two weeks, paused part of its reinforcement learning, where these systems train very intensively on their own in a loop, and left the largest planned training run on pause.
Precisely because these models are starting to move out of control. Under the new rules a serious alert has to be reviewed within thirty minutes. If the violation is not judged a false positive, the activity has to be suspended. And OpenAI put the compute cost of that monitoring — imagine — at 20% of the compute of the model output being controlled. That shows us the real problem. Quality control requires incredibly large sums. Even 20% of capacity is incredibly a lot, simply incredibly a lot. So there are limits to this report, of course, because the authors do not actually know
the company's internal infrastructure, they are not inside the company, and their monitors are not installed there. But the question I would put today is this: even if Guideline were able to install some observation system inside them, would they be able to build a system better than the companies themselves can build — companies like OpenAI and Anthropic? And most likely we will find that it cannot be done. What do you think about this? What are you already doing in your own companies? Tell us here. Do write comments. Share this episode with your friends.
Support our channel. It supports the channel a great deal. And all the comments matter enormously. Every comment you leave, every like. Sometimes the number of likes does not much exceed the number of comments. People do not like, do not write. I notice it in myself — I do not like anything either. And because I run my own channels, I keep remembering it. See you in the new episodes.