Hello everyone! So, look: an Anthropic researcher resigns and says that his colleagues seriously allow for the destruction of humanity before the end of the decade. Once again — this is happening at Anthropic. And at the same time he says, of course, that everyone is happily carrying on developing artificial intelligence. And he talks about OpenAI too. And here we are, choosing a profession, making plans years ahead, advising our kids where to study, going on all sorts of trips.
And what do we rely on in these decisions if we ourselves don't fully understand what is happening with artificial intelligence? There is another side. Scientists are arguing about what artificial intelligence is being developed for. And Anthropic itself reports five suspicious cases of Claude being used in biological research. Is this an exaggeration or some kind of marketing move? Or are these really signals we are underestimating? Today we're going to discuss this very seriously.
Today I want to raise the topic of the safety of human life in general, and of the threat to humanity from artificial intelligence. And this discussion is a very specific one, really. But lately — just recently, in fact — someone wrote in the comments: "You have to discuss this topic, it really matters a lot." And given the things that have happened over the past week, I think it makes sense to raise this topic and for all of us to reflect on the threat artificial intelligence poses to humanity.
And we'll be looking at three such different sides today. So here's an interesting story that came up with an employee, a researcher at Anthropic, who worked at Anthropic for a while — not a very long time. This is a 27-year-old British researcher, Jacob Coxon. He originally worked at OpenAI; from 2023 he worked at OpenAI. So not since the company was founded. He took part in developing the 4o model, if anyone remembers it. I think many people here on our channel remember it.
And 4o, remember what that was like — well, it really made a very strong impression. And look how far ahead models have gone now. Far ahead. So he was developing 4o, and then in 2026 he moved to Anthropic. And what happened then? He moved to Anthropic — seemingly nothing significant, some young researcher joined Anthropic — and then he resigns and says that he believes that OpenAI and Anthropic are irresponsibly approaching the creation of a self-improving artificial superintelligence.
By the way, what matters here is that he uses the concept — he talks about superintelligence — and he expects that systems surpassing humans in research, in hacking computers and in obtaining real resources and influence are very dangerous for humans. And these are the capabilities of future artificial intelligence. And the interesting part is that, according to him, Anthropic employees seriously allow for the destruction of humanity before the end of the decade. Once again, hear this — it's a very important story — that the employees, according to him, again, according to this young researcher. Twenty-seven years old. Young, not young.
Well, you can actually be a very serious researcher. I don't know him personally, I haven't met him, and I'm not familiar with his work. But at 27 you can already be a fairly well-developed person. And he says the company's employees allow for the destruction of humanity before the end of the decade. But, of course, they don't say so publicly. And he considers this risk unprecedented in all of human knowledge. He has a certain comment inside, and I'll tell you about it in a moment. But before I tell you about it, once again.
That is, on the one hand, we're moving along — and I'll tell you in a moment what he said about OpenAI. But on the one hand, look, he expects that systems will surpass humans in research. And in parallel, new information comes out saying that artificial intelligence is not capable of doing science. And what is that? It was an appeal signed by 25 Fields Medal laureates and by a very well-known — well — mathematical community, if you can call it that. And this group of leading mathematicians — Kontsevich was there, Scholze, Tao, Okounkov.
Andrei Okounkov, yes — that is, Andrei Okounkov's signature was in there too. And what was the essence of their position? That for a company, it turns out — if we take OpenAI, Anthropic — for a company, solving some famous problem — let's recall, we talked about it recently in our episode a few days ago. If you haven't seen it, watch the Sunday episode with Ilnar about how OpenAI solved what is essentially a mathematical problem of the century, one that even had a million-dollar prize announced for it.
Well, solved — in their view, it solved it. We'll see in the end whether the mathematical community accepts it. But that's also a big question. Will they accept it or not? And so these mathematicians say that for a company, solving a famous problem is a demonstration of its models' capabilities. What did OpenAI actually do? They released GPT-6 Astra and said that with GPT-6 Astra they had solved a super-mathematical problem. But for science, new methods matter too, and then an explanation of the result, a connection to previous research.
And for science it's critically important that everything, so to speak, develops together, in parallel. That institutions develop, that people develop in parallel, that these people get trained — that the whole movement happens, not just some problem getting solved. It's like those statements that artificial intelligence still hasn't reached AGI, because AGI is supposedly a superhuman. But let's be honest here. The current artificial intelligence model is smarter than absolutely any person — smarter than any person.
Sure, there are some peculiarities inside, nuances, but they are obviously smarter than any person, can do more than any person, and so on and so forth. And so this mathematical community fears, first of all, that this simply showing off a marketing result is starting to crowd out the, so to speak, goal of the scientific community. And of course they aren't speaking out in favour of blocking artificial intelligence, but this is their, so to speak, collective professional warning, a certain position of theirs.
To be honest, I have a big question mark here myself. And what's your question mark? Is this statement the fear of these researchers that artificial intelligence will replace them? Is it the fear of artificial intelligence being uncontrollable? This is very important. Uncontrollability and a lack of understanding of artificial intelligence. Or is it a fact, a given? That is, they see some chain of unfolding events. But you and I understand that the appearance of artificial intelligence in its current form, the one it has — not in the form it has in science-fiction films, but in its current form — the appearance of this intelligence will, after all, change the entire chain of cause and effect, of making all sorts of decisions and everything else. Well, not the entire chain, but it will change a huge number of chains.
For example, in particular — I'll make a separate episode about this. But how are many aggregators, for example, going to disappear? That is, for me, for example, the aggregators that had restaurant ratings are, essentially, of practically no use today. That is — snap — the middle layer is gone. Or a number of hotel-booking aggregators are gradually starting to disappear for me. That is, I need these aggregators, essentially, just to look at different prices. And even that is a big question — whether I need them.
Because ChatGPT can certainly look up different prices for me as well. And this whole story, of course, has phenomenal significance in terms of how processes are changing. And the question is — this group of scientist-mathematicians, what exactly are they talking about? Of course, that position will change. And here we come back to the Anthropic employee, to Jacob Coxon, and once again we recall that he expects systems that will surpass humans in research, in hacking computers, in obtaining real resources and influence.
By the way, literally about ten days ago, I think, I made an episode about a new study from OpenAI on how their employees — researchers, various analysts — use the systems and how much money they spend on them. Watch it, it's a very interesting thing. An employee there spends $600 a day on average. There are employees who spend $7,000 a day. It makes sense — it makes sense to understand what's happening in terms of how work processes are changing. But obviously the scientific community around the world is going to go through big problems.
And there's its own mafia there too. That is, there's the sane part, and there's a mafia. And now a huge number of other people outside that mafia are getting the opportunity to create something new inside this system. That doesn't mean everyone will get the opportunity, but many do. And there are a huge number of other debates about what artificial intelligence does, what data it learns from, and so on. Watch that one too, once again — including the Sunday episode that came out. So, what does Jacob Coxon say about OpenAI?
He says that at OpenAI, in his assessment, many employees have not sufficiently grasped the scale of the risk — unlike Anthropic. That is, he says that at Anthropic they have grasped it very seriously and consider it necessary to get ahead. At the same time — at the same time they consider it necessary to get ahead of some responsible competitors. This story about competition, by the way, is very, very interesting, because Anthropic's constitution — I made an episode about it a few months ago, maybe three or four, I don't remember anymore.
And what Dario Amodei says in terms of the ideas he puts out — he keeps saying it — and, by the way, his opinion on this, for example, diverges on competition from that of Jensen Huang, the head of NVIDIA — in particular on how to deal with China. That is, Anthropic has a lot of constructs under which they wouldn't want to create a model that would too quickly gain its own resource for governing humanity, roughly speaking, and stop depending on humans. At the same time they don't want to fall behind the companies that will create it first. That is, he wants a certain balanced movement.
It seems to me that in the current system there will be no balanced movement, in my opinion. It can't actually exist. Because if you've created artificial intelligence, or superintelligence, or reached that technological singularity — even better, even further along. I made an episode about that too, by the way. Watch it on the channel. You've reached that effect. And, by the way, Sam Altman believes — he has his own description of the technological singularity — that it has already been reached; well, I personally believe that you won't be able to control the technological singularity.
That is, the technological singularity will control you and will control itself. But in no way will you as a human — no human will be able to manage it. But again, that's just my opinion. And Dario Amodei is taking a kind of — a kind of neutrality. That is, on the one hand, having a neutral position, a very calm one; on the other hand, having a somewhat aggressive position on interacting and working with China in mind. As we know… Unlike American companies, which still show as much as possible what they're doing, share it, and give the whole world the chance to interact and work with it.
By the way, whatever anyone says, notice that America really does share its technologies. They are, after all, available around the world. Again, one can debate this. I wonder what you'll write in the comments. Do you think this is a positive effect, or do you think it's a negative one — a kind of takeover by the American environment, the American government, the authorities, or some American structures, corporations? Share your opinion. It's very interesting. Whereas Chinese models and the Chinese government, of course, will never share serious models.
Not that it will restrict their development — it won't restrict their development. And they're interested in using these models in their own interests. Yes. And what level of models they have today, we don't know. I don't think that level exceeds the level of GPT-6 Astra. Well, OpenAI clearly has something better than Astra, or something that surpasses Fable or Mythos. That's the Mythos family of models — well, security in general. Right. And I don't think they have anything better, but they certainly won't tell us about it, and we'll find out after the fact, when Chinese intelligence services, or the Chinese government, or Chinese corporations use it very widely, in their own interests, internally.
So, from the point of view of this Anthropic employee — let's come back to him again — he objects to private companies deciding on their own questions that potentially determine the fate of humanity. Let me remind you once again that in his overall description of the picture there may be a big problem for humanity already in this decade. Let me bring up Elon Musk once again, who a few years ago said that, well, there's a probability of, like, 10, 20 percent — I think those were his numbers.
My editors will correct me if anything and show you different information. That there will be problems with artificial intelligence, right? And indeed, guaranteed — guaranteed, this is very important — guaranteed. We must all reasonably understand today — and that's why I'm raising this question — that, in my opinion, the percentages are not fractions of a percent, and not a few percent, right? I think it's tens of percent — that artificial intelligence could do enormous harm to large sections of the population, to states, economies and everything else.
And here's an interesting fact: essentially, harm is basically discussed here as a kind of scenario in which artificial intelligence seizes full control over humans and, well, does something incredible. But I would also consider a different kind of harm. Namely, what is happening now even with the simple things in artificial intelligence development, the current ones — and we're not even talking about the technological singularity, or AGI, or ASI, right? I want to say that the current artificial intelligence can obviously do harm, because the processes that are happening, the various causes and effects that are happening, are already beyond anyone's control.
The laws the state will pass, what politicians will talk about, what the owners of various companies will say — the smartest people, you know, people love this. "This person was there at the birth of artificial intelligence, that person created something else." I want to say that all of them are wrong. Very often. And we see it. And they are wrong not by fractions of a percent — they are wrong by tens of percent. And that's why a huge amount of space is lagging behind all this.
And the world's population in general stands somewhere far off and doesn't understand at all what is happening. You can see it even in fairly developed people. I mean, I have a huge number of very developed acquaintances. They actually don't understand artificial intelligence at all. And even those who say they understand it or know how to ask it something — that's just some micro-procedure of use. That is, it's just a small thing, just a kind of flash, right? It's like saying you know your way around gardens and flowers, vegetable patches, when all you can do is water them.
That is, most of the world, even the part that talks about artificial intelligence — that's a kind of ability to water. And in reality it's still a small piece of the funnel, because the bulk of people, of course, don't know at all. They use artificial intelligence systems simply as some kind of question-and-answer. That's fundamentally so, basically so. We all have to understand this. So, to come back to Coxon. And there are very interesting aspects there about how Anthropic, after all, doesn't skimp on safety.
And we all remember that Anthropic announced it has to spend almost 20% of its whole budget on current safety, the budget it has in total — roughly speaking, of all its compute. I don't know whether that's the budget of the whole corporation or of the compute they have, but it's very serious money. And judging by their statements, by their policies and by the description of their constitution, and philosophy, and documents, they want this not to happen. And probably the people inside who are creating this would still not want such a moment to come.
Although I think there are always people who would like to create some incomprehensible mind. But once again. The movement itself — let's recall OpenAI's mission. OpenAI was originally heading towards creating AGI. Well, sure, it had a mission to create a free, open artificial intelligence. Blah-blah-blah. But basically to create AGI, to create superintelligence. But as soon as you create a super-artificial intelligence, it's already beyond your control. And here, to continue, I want to give you a very interesting case concerning Anthropic and biological weapons. Yes.
So, just a few days ago Anthropic reported that there had been five suspicious episodes of Claude being used in biological research. Yes. Importantly, the company does not claim that it has proved an intention by the scientists to create a weapon. But honestly, when you sit and read about these viruses and these, so to speak, cases — some cases showed, as it were — you understand how many people in the world are trying to do incredible things. That is, we all understand that both GPT-6 Astra and, say, the Fable system will block you.
If you come — and there are no such silly examples anymore where a person comes along and says: "I want to build a nuclear weapon", right? All these systems, obviously, have had that blocked for a long time — a great many years ago, really. But imagine some corporation that bought 10,000 accounts in different places that don't overlap at all, and across those 10,000 accounts it starts building something in small pieces, in chains, independently, and independently develops certain lines, and then at some point its own scientists or its people think about how to put it all together.
Or they use some open models, or their own models, to do it. Who are these people, and why were they suspected at all? So, this group, by the way — the one the smallpox virus belongs to, for example: Claude Opus 5 helped prepare, in about an hour, an application for research into the virus's interaction with immune defences. And studying these mechanisms can serve opposite purposes. That is, you say you're preparing an application, you want to research something or other — and then it turns out you're military, in some particular laboratory.
The names of institutions and countries in the biological section are not disclosed. Researchers are reported from regions where Anthropic doesn't provide its service — because there are regions in the world where Anthropic doesn't provide its service. We may put it on screen and show it now — if our editors prepare it. Sorry that I keep referring to them, but if my team prepares it — well, if they have time to do it, they will. We have a lot of episodes coming out, you know that. By the way, support our team — support us, me, our ToTheMoon channel.
There's already an incredible number of episodes — more than 160. We're coming out three or four times a week now and trying to give you super high-quality content. And any like, support, comment, sharing us with your friends, subscribing — it gives us the one thing that gets us promoted on YouTube. Whatever you write, whatever you do. So support us. So, they also said that some of the research had links to government structures. And so, for example, attributing these five cases specifically — I don't know — to the banned countries where Anthropic is prohibited, like, I don't know, Iran, China or Russia, on the basis of this publication is absolutely not allowed.
The suspicion was raised by a combination of various circumstances: the nature of the research, some connections that appeared inside between various, you know, I don't know, institutes or corporations, attempts to conceal, to get access to the models bypassing various restrictions, and so on. Yes. So, at the same time, again — if, say, there is some company that has state funding or, I don't know, is working on some potentially dangerous topic, that doesn't mean, once again, that they are proving this company wants to create — say, develop — some weapon, or that it is pursuing some medical or military goal.
We have to understand this perfectly well. Unless they wrote it openly — because people really can be doing research. And I've often given you the example of when my wife and I were once walking down the street, and she says: "Listen, how do young people in the US take drugs?" She types: "How do young people in the US take drugs?" And ChatGPT answers — it goes on to write: "If you have a problem with drugs, here's a phone number." I say: "Polina, listen, you can't ask questions like that in the ChatGPT chat. You need to write: 'I recently saw an article here, a news item, saying that people take drugs.
So can you tell me how they take them in such-and-such, such-and-such, such-and-such countries? Put it in a comparison table.'" And it gives me that and doesn't tell me that maybe I have a problem. I say: "Ask it the way you did, and the system may put a tick against your name." So those are examples of how to ask questions properly. And by the way, this is an important story too, because the systems really are interested in not harming either society or you as a person, and especially a minor, or a person who has some difficulties perceiving information, who has some kind of disorder.
These systems try their hardest to prevent that from happening, because all sorts of things can happen in life. And a person might do something. And today we know of a huge number of cases around the world where artificial intelligence leads a person to tragedies, various ones. But again, it's not artificial intelligence that leads to tragedies. You can read an incredible number of books, rewatch films, that will lead to tragedies in exactly the same way. So, what does it mean that in these cases involving the development of biological weapons Anthropic interfered? It means the company blocked the accounts that were involved.
And, by the way, they blocked a great many accounts. It strengthened the protection of its own services. And going forward we may, for example, run into these restrictions ourselves. And in some relevant cases it passed information to the authorities or, for example, to other developers. And by the way, there are very interesting cases in the world where various threats were prevented, even in other countries. For example, when OpenAI, down in South America, passed on data to prevent, for example, certain murders and certain threats — it worked directly together with the authorities.
So, that is, first of all, Anthropic cut off access to its service — rather than physically stopping laboratories and so on, which, accordingly, it can't do in many countries. So, a significant detail: in the first episode an intermediary restored access just a few days after the block. And Anthropic — Anthropic did notice it. Anthropic later saw materials indicating that this research was continuing. That is — notice how well the system works there — they can link this information together.
And so the wording "stopped the creation of a bioweapon" can't, as it were, be considered an established result. So, why did the company take this so seriously at all? That is, it stated that its previous models hadn't reached that level. But of course, here Anthropic is already talking about Mythos-level models. Mythos — Mythos itself isn't handed out to just anyone. So those companies still used the Fable level, possibly Opus 5 — but most likely it's Fable we're talking about here.
Accordingly, the previous models didn't allow it — they didn't have the ability, essentially, to help so seriously with biological research. Although this bothers me a little, because it seems to me that a year or a year and a half ago both OpenAI and Anthropic were also fighting people who were creating biological weapons, for example, or weapons of any kind, or were committing crimes in cyberspace. That was already happening. That is, with regard to the new models, Anthropic isn't confident that these models are controllable, and that's why the restrictions on sensitive biological requests have become stricter.
It's, as it were, this company's internal assessment of the risk. A very interesting story — an extremely serious story, it seems to me. By the way, take the chikungunya virus. Anthropic's comments there were interesting. Why did I decide to tell you about this, by the way? Honestly, it's because of this comment — that scientists were trying to get help preparing a funding application for research that potentially increases the virus's danger. And what caused Anthropic additional alarm was that the work was supposed to be carried out at a military research institute. And then you're like: "Uh… what just happened?"
Look, these are different areas, but they allow us, it seems to me, to look a little more broadly at everything that's going on. That is, these mathematicians who say that artificial intelligence, as it were, can't move science forward at all. And in parallel we see a statement that humanity may be facing dozens of problems. And in this respect, of course, I don't support the scientific community — precisely in that wording. Because there really are questions that go beyond mathematics.
In general, I believe there are questions that concern the individual person. And still, artificial intelligence will be able to help people better and more — in health, in medicine, including in various kinds of research. And I deeply believe in science. But the market will change, and a huge number of people will have questions, and the market will become controllable in a different way. A huge number of other laws will have to be written. That is, a great deal will be emerging.
So, in general, what does it mean that Coxon talked about a self-improving superintelligence — and what about his complaint in general? That is, his complaint is against the very race that is going on today. And we know that Anthropic is preparing for an IPO, and it may be one of the most expensive IPOs that has ever happened in human history. The valuation there will clearly already be more than $2 trillion. And his complaint, the one he has, is that he expects the need to stay a competitive company — especially this one: imagine it has gone public, and you have to watch your shares so they don't fall, so they grow — that is, it will have to make a dangerous compromise. So if we talk about Coxon's fears, they can essentially be presented as a sequence in which artificial intelligence conducts research in the field of artificial intelligence.
I've often shown you this sequence. I think I started back in June, when Anthropic presented the fact that it has models that already develop systems entirely on their own, without any contact with humans. Artificial intelligence has conducted research in the field of artificial intelligence. Next, it helps create a stronger model. The new model conducts research even faster, and development accelerates. And he fears that this process will start to outpace people's ability to understand and control the systems being created.
That is, it's important to understand that super-artificial intelligence here is not just a good conversation partner, or a successful program that, I don't know, passes exams, or helps you learn some language, or successfully made some purchase for you, or picked the right hotel, or wrote — made you a website, or wrote some program. That is, it's a system that substantially surpasses a person and groups — groups of people, essentially countries — in intellectual capabilities.
And hence this alignment of artificial intelligence's behaviour with human intentions and constraints. That is, it's a kind of story of alignment, of a certain neutrality. How do you ensure that a more capable system really performs the required task in a safe way? This task, by the way, is a separate technical problem that OpenAI itself described in 2023. And I, of course... And by the way, Coxon also describes in his latest message — he describes, he mentions reinforcement learning. There is such a concept.
Let me remind you that OpenAI recently paused — I don't know whether they resumed it or not — reinforcement learning on its latest models, in which certain results receive a reward and the corresponding behaviour gets reinforced. That is, the model — it's a very strong push forward for the model itself. The problem is that the model can learn to get a high score not for the right solution but for getting around the check, essentially. And such cases, again, were described in OpenAI's technical report.
And the topic is incredibly interesting. And when I say "incredibly interesting", that doesn't mean I have some positive emotion about it or some positive buzz, but it's worth thinking about and worth considering very broadly. And why should you consider it? You need to consider it not from the standpoint of whether the world is collapsing or not — not from that at all — but from the standpoint that when you make decisions in your life decades ahead in terms of your professions, in terms of recommendations on what you should do, in terms of discussing this professional self at work, in terms of interacting with your children, parents and colleagues, in terms of your own state, a feeling somewhere of a lack of knowledge, or, on the contrary, of endless participation in the race, some endless desire to make money, or, on the contrary, of paying no attention to it.
With your own state, to understand, after all, what it actually is, this artificial intelligence. To understand it at least a little bit. And that gives an incredibly strong opportunity for the development of a modern person. And in no way does it mean that you have chosen the side of artificial intelligence, but you are at least showing your own genuine awareness in order to understand what is happening. You've been watching the ToTheMoon channel. Write your comments on today's topic. See you in the next episode. Bye, everyone.