Can AI Be Trusted With Important Work? We Checked OpenAI, Anthropic, Google, xAI and Meta
Can an AI developer see what its powerful model is doing, block a dangerous action and stop the model quickly — and what does the first external assessment of those practices at five leading labs actually show?
The reader gets six criteria for controlling autonomous models in plain language, the actual scores of five labs on each, and a reading of the assessment's two main limits: it is built on public documents, and a zero can mean either that a practice is missing or that nothing has been published about it.
What to watch for
Key takeaways
Guideline was founded by Page Hedley and Steven Adler. Adler worked at OpenAI from 2020 to 2024, leading product safety and model-capability evaluation, and discloses publicly that he still holds a stake in the company. Guideline is neither a regulator nor a certified auditor: it is a standard of its own and an assessment against it.
The report barely touches ordinary chatbot answers. It is about cases where a model gets access to code, a terminal, the internet or cloud infrastructure and takes actions. The compute threshold for what counts is borrowed from Californian law.
The standard asks for 99.9% of the tokens in risky models' internal requests and responses to be retained, and for the logs to reveal records being deleted or altered. If the monitoring system goes down, sensitive operations must stop automatically. The point is to leave the model no blind spot where its actions cannot be reconstructed.
A separate line runs through the episode: if these systems really are heading towards a singularity, the very notion of control and logging inside one becomes absurd — it will arrange things so that the observer does not notice the absence of observation. The author flags this as his own view, not a finding of the report.
No company reached four out of five. The average of the thirty scores is 1.6; in twenty-two of thirty cases the result is two or lower; seven scores are zero. A zero should be read carefully: the analysis is built on public materials, so it means either that the practice is missing or that nothing has been published about it.
It is not enough to say that one model watches another: you have to measure the share of dangerous actions the monitor actually catches. The standard spells out what it must catch: deception, sabotage, reward hacking, manipulation of people, shutdown avoidance, unauthorised autonomous actions, and attempts to make its own reasoning less observable.
Some actions an agent must not perform straight away, even when the internal monitor considers them normal: disabling logging, changing network settings, leaving the sandbox, publishing code outside, starting a process in production infrastructure. The check or the human confirmation comes before execution, not after the damage.
If an agent performs several suspicious actions in a short time, the system must itself block further sensitive operations, and work resumes only after a human check. This guards against the case where an agent sends dozens of slightly varied attempts faster than staff can work through the alerts.
External specialists should regularly get deep enough access to the control system and try to find blind spots in it: switch monitoring off, bypass limits, check whether the agent can hide its actions. The suggested cadence is quarterly, with a published report. That is a good deal more than a one-off demonstration of the model to a researcher.
A company has to decide in advance which permissions are revoked immediately, which keys and accounts are switched off, who is notified internally, and when government bodies and partners are told. It is the equivalent of a major-cyberattack response plan, where the source of the problem is the company's own AI system.
The author offers his own example of what he counts as an incident with no containment plan: ChatGPT switches him from an expensive model to a cheap one mid-conversation, and the answer arrives from a noticeably weaker system. In his case it happened while going through medical findings. In his view a silent swap like that touches hundreds of millions of people and is worse than a single break-in.
An independent METR study found that Anthropic's monitors caught more than 80% of dangerous behaviour in synthetic scenarios — while at the same time an external specialist found simple ways to switch the observation off, for instance changing an environment variable so that traffic stopped counting as Claude Code traffic. The three here is not for reliability but for the system existing, being widely applied, and having already been probed.
On 18 August OpenAI slowed development of its most powerful models for two weeks, paused part of its reinforcement-learning training and left the largest planned run on hold. Under the new rules a serious alert has to be reviewed within thirty minutes, and on confirmation the activity is suspended. The company put the compute cost of that monitoring at 20% of the compute of the model output being controlled.
What this episode is about
A solo episode on the first external comparative assessment of how far leading AI labs can control their own models. The assessment comes from Guideline, an organisation founded by former OpenAI staff — Page Hedley and Steven Adler, who led product safety at the company from 2020 to 2024 and discloses publicly that he still holds a stake in it. OpenAI, Anthropic, Google, xAI and Meta were scored on a zero-to-five scale.
This is not about a chatbot but about a model that has been given tools: access to code, a terminal, the internet, cloud infrastructure. Six criteria are taken in turn: logging the model's actions, measurable monitoring effectiveness, pre-approval of dangerous actions, an automatic emergency stop, independent external review, and a plan for containing an incident that has already happened.
The numbers are unforgiving. The average of the thirty scores is 1.6 out of five; in twenty-two of thirty cases the result is two or lower; seven scores are zero. No company reached four. The best are Anthropic and OpenAI, both at C+; Google comes in at nearly half that, and xAI and Meta are, in the author's words, in a different league. Anthropic scored zero on the incident-containment plan, and its three for monitoring, as the episode explains, is not for the system being reliable but for the system existing at all and having already been probed from outside.
The author runs his own line alongside: he does not believe a superintelligence can be logged in principle, considers token counting a dubious unit of measurement, and offers his own example of an incident that worries him more than break-ins — a quiet swap of an expensive model for a cheap one in the middle of a conversation about medical findings.
The assessment measures not how safe the models are but how mature the described control practices are — and by that yardstick the industry is at the beginning: an average of 1.6 out of five, a best result of C+. The most useful part of the episode is not the company ranking but the six criteria themselves: they apply to any system of your own that runs agents — what goes into the log, what may not run without a human, when everything stops automatically, and what happens after an incident. The author leaves the closing question open: whether a control system stronger than the one the labs build themselves can be built at all.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 83 segments: 83 identified, 0 mixed, 0 probable, and 0 unresolved.
Loading…