Skip to content
Transcript

Transcript · 166 · Could AI Destroy Humanity Within the Next 10 Years? 3 Alarm Bells — ToTheMoon

English machine translation of the Russian-language episode. Timecodes open the source video.

Episode overview
00:00:00–00:30:28Could AI Destroy Humanity Within the Next 10 Years? 3 Alarm Bells
Alexander Volchek00:00:00

Hello everyone! So, look: an Anthropic researcher resigns and says his colleagues seriously allow for the destruction of humanity before the end of the decade. Once again — this is happening at Anthropic. And at the same time he says, of course, that everyone is happily carrying on developing artificial intelligence. And he talks about OpenAI too. And here we are choosing a profession, making plans years ahead, advising our children where to study, going on all sorts of trips. And what do we rely on in those decisions if we ourselves do not fully understand what is happening with artificial intelligence?

Alexander Volchek00:00:30

There is another side. Scientists are arguing about what artificial intelligence is being developed for. And Anthropic itself reports five suspicious cases of Claude being used in biological research. Is that an exaggeration? Some kind of marketing move? Or are these really signals we are underestimating? Today we will discuss this with you very seriously. Today I want to raise the topic of the safety of human life in general, and of the threat to humanity from artificial intelligence. And this discussion is a very specific one, really.

Alexander Volchek00:01:17

But in the last episode — even recently someone wrote in the comments that this topic must be discussed. It really matters a lot. And given the things that have happened over the past week, I think it makes sense to raise this topic and for all of us to reflect on the threat artificial intelligence poses to humanity. And these are the three different sides we will be taking apart with you today. Here is the interesting story that arose with an Anthropic researcher, who worked at Anthropic for a while — not an especially long time — a 27-year-old British researcher, Jacob Coxon.

Alexander Volchek00:02:04

He originally worked at OpenAI: from 2023 he worked at OpenAI — that is, not from the company's founding — and took part in developing the 4o model. If anyone remembers it — I think many here on our channel remember it. And 4o, remember, that was — well, it made a really strong impression, and look how far models have moved on since. So he was developing 4o, and then in 2026 he moved to Anthropic — and what happened next? He moved to Anthropic — a seemingly insignificant, young researcher moved to Anthropic.

Mentions: OpenAI · Anthropic
Alexander Volchek00:02:44

And then he resigns and says that he believes OpenAI and Anthropic are irresponsibly approaching the creation of a self-improving artificial superintelligence. By the way, it is important that he uses the concept of superintelligence. And he expects that systems surpassing humans in research, in hacking computers, and in obtaining real resources and influence are very dangerous for humans. And these are the capabilities of future artificial intelligence. And the interesting part is that, according to him, Anthropic employees seriously allow for the destruction of humanity before the end of the decade.

Alexander Volchek00:03:32

Once again, hear this — it is a very important story — that employees, according to him, again, according to this young researcher, 27 years old — young, not young, well, you can actually be a very serious researcher. I do not know him personally, I am not familiar with his work, but at 27 you can already be quite a well-developed person. And he says the company's employees allow for the destruction of humanity before the end of the decade. But, of course, they do not say so publicly, and he considers this risk unprecedented in all of human knowledge.

Alexander Volchek00:04:12

He has a certain comment inside, and I will tell you about it in a moment, but before I do — once again, on the one hand we are moving, and I will tell you what he said about OpenAI. But on the one hand, look, he expects that systems will surpass humans in research — and in parallel new information comes out that artificial intelligence is not capable of doing science. And what is that? It was an appeal signed by 25 Fields Medal laureates — a very well-known, so to speak, mathematical community.

Alexander Volchek00:04:55

And this is a group of leading mathematicians: Kontsevich, Scholze, Tao, Okounkov — Andrei Okounkov's signature was there too. And the essence of their position was that for a company — if we take OpenAI and Anthropic — for a company, solving some famous problem — let's recall, we told you about it in an episode a few days ago, watch it if you haven't, the Sunday episode with Ilnar, about OpenAI having solved, essentially, a mathematical problem of the century, for which even a prize of a million dollars had been announced.

Mentions: Anthropic · OpenAI
Alexander Volchek00:05:40

Well, solved — in their opinion it solved it; we will see whether the mathematical community accepts it in the end, but that is also a big question — will they accept it or not. And so these mathematicians say that for the company, solving a famous problem is a demonstration of the models' capabilities — which is essentially what OpenAI did. They released GPT-6 Astra and said that with GPT-6 Astra they had solved a super-mathematical problem. But for science, new methods matter too, and an explanation of the result, a connection with previous research — and for science it is critical that everything develops together, in parallel.

Mentions: OpenAI · Astra
Alexander Volchek00:06:20

That institutions develop, that people develop in parallel, that these people are trained — that the whole movement happens, not just some problem getting solved. It is like those statements that artificial intelligence has still not reached AGI, because AGI is supposedly a superhuman — but let's be honest, the current artificial-intelligence model is smarter than absolutely any person, smarter than any person. Of course there are some nuances inside, but they are obviously smarter than any person, can do more than any person, and so on and so forth.

Alexander Volchek00:06:48

So this mathematical community fears, first of all, that simply showing a marketing result is starting to crowd out the goal of the scientific community. And of course they do not speak out in favour of blocking artificial intelligence, but this is their collective warning, so to speak — their professional position. I am standing before a big question here, and you are standing before such questions too. Is this statement the fear of these researchers that artificial intelligence will replace them; is it the fear of artificial intelligence being uncontrollable — this is very important, uncontrollability and a lack of understanding of artificial intelligence.

Alexander Volchek00:07:29

Or is it a fact, a given — that is, do they see some chain of unfolding events. But you and I understand that the appearance of artificial intelligence in its current form — not the form it has in science-fiction films, but the current form — the appearance of this intelligence will change the whole chain of cause and effect in making all sorts of decisions and everything else. Well, not the whole chain, but it will change a huge number of chains. For example, in particular — I will make a separate episode about this — how many aggregators, for example, will disappear.

Alexander Volchek00:08:02

That is, for me, for example, the aggregators that had restaurant ratings are essentially of practically no use today. That is — snap — the middle layer is gone. Or a number of hotel-booking aggregators are gradually starting to disappear for me. That is, I need these aggregators, essentially, just to look at different prices. And even that is a big question — whether I need them, because ChatGPT can certainly look at different prices for me on top. And this story has, of course, phenomenal significance for how processes are changing.

Alexander Volchek00:08:35

And the question about this group of scientists, mathematicians — what exactly are they talking about? That is, of course, their position will change. And here we return to the Anthropic employee, to Jacob Coxon. And once again we recall that he expects systems that will surpass humans in research, in hacking computers, in obtaining real resources and influence. By the way, literally about 10 days ago I made an episode about a new study from OpenAI on how its researchers and various analysts use the systems and how much money they spend on them.

Mentions: Anthropic · Jacob Coxon · OpenAI
Alexander Volchek00:09:14

Watch it, it is a very interesting thing: an employee there spends 600 dollars a day on average, and there are employees who spend 7 thousand a day. And that is the point — to understand what is happening in terms of how work processes are changing. But obviously the scientific community in the world is going to go through big problems. And there is its own mafia there too. That is, there are adequate people, and there is a mafia. And now a huge number of other people outside that mafia are getting the opportunity to create something new inside this system.

Alexander Volchek00:09:44

That does not mean everyone will get the opportunity, but many do. There are a huge number of other discussions on what artificial intelligence does, what data it learns on, and so on. Watch, once again, the Sunday episode as well, the one that came out. So, what does Jacob Coxon say about OpenAI? He says that at OpenAI, by his assessment, many employees have not sufficiently grasped the scale of the risk — unlike Anthropic. That is, he says that at Anthropic they have grasped it very seriously and consider it necessary to get ahead — while at the same time they consider it necessary to get ahead of some responsible competitors.

Alexander Volchek00:10:21

The story about competition, by the way, is very interesting, because I made an episode about Anthropic's constitution a few months ago — maybe 3 or 4, I don't remember any more. And what Dario Amodei says in terms of the ideas he broadcasts — he keeps saying it — and, by the way, his opinion on competition diverges from that of Jensen Huang, the head of NVIDIA. In particular, on how to interact with China — that is, Anthropic has a lot of constructs under which they would not want to create a model that too quickly acquired its own resources for, conditionally, managing humanity, and did not depend on humans.

Mentions: Anthropic · China
Alexander Volchek00:11:07

At the same time they do not want to fall behind the companies that will create it first — that is, he wants a certain balanced movement. It seems to me that in the current system there will be no balanced movement, in my opinion. It cannot in fact exist, because if you have created artificial intelligence, or superintelligence, or reached that technological singularity — even better, further on. I made an episode about that too, by the way, watch it on the channel. Once you have reached that effect — and, by the way, Sam Altman believes, he has his own description of the technological singularity, that it has already been reached — then I personally believe you will not be able to control the technological singularity.

Alexander Volchek00:11:45

That is, the technological singularity will control you and will control itself, but in no case will you, as a human — no human will be able to manage it. Well, again, that is my own opinion. And Dario Amodei is going in a kind of neutrality. That is, on the one hand he has a neutral position, very calm; on the other hand he has a rather aggressive position regarding interaction and work with China in mind. As we know, unlike American companies, which still show as much as possible what they do, share it, give the whole world the chance to interact and work with it.

Mentions: China · United States
Alexander Volchek00:12:23

By the way, whatever anyone says, note that America really does share its technologies. They are, after all, available in the world. One can argue about that too, of course. It is interesting what you write in the comments — do you consider this a positive effect, or a negative one, a kind of takeover by the American environment, the American government, authorities or some American structures, corporations. Share your opinion, it is very interesting. Whereas Chinese models and the Chinese government, of course, will never share serious models.

Mentions: United States · China
Alexander Volchek00:12:52

It will — not that it will limit their development — it will not limit their development. And they are interested in using these models in their own interests. And what level of models they have today, we do not know. I do not think that level exceeds the level of GPT-6 Astra, but clearly PNA has something better than Astra, or surpasses Fable or Mythos. That is the Mythos family of models, yes — cybersecurity. Well, security in general. And I do not think they have anything better, but they will certainly not tell us about it, and we will find out after the fact — when Chinese special services, or the Chinese government, or Chinese corporations use it very widely, in their own internal interests.

Mentions: Astra · China
Alexander Volchek00:13:40

So, from the point of view of this Anthropic employee — back to him again — he objects to private companies deciding on their own questions that potentially determine the fate of humanity. Let me remind you once again that in his overall description of the picture there may be a big problem for humanity already in this decade. Let me remind you again of Elon Musk, who a few years ago said that, well, there is a probability — 10–20 percent, I think those were his numbers, my editors will correct me and show you other information — that there will be problems with artificial intelligence; and indeed, guaranteed — guaranteed, this is very important — guaranteed, we must all reasonably understand today, and that is why I am raising this question: in my opinion these are not fractions of a percent, and not percent — I think it is tens of percent — that artificial intelligence can do enormous harm to large sections of the population, to states, economies and everything else.

Alexander Volchek00:14:49

And here is an interesting fact: essentially, the harm is discussed here basically as a story in which artificial intelligence seizes full control over humans and does something incredible. But I would consider another story of harm — what is happening now, even with the simple pace of artificial intelligence development, the current one, without even talking about the technological singularity, or AGI, or ASI. I want to say that the current artificial intelligence can obviously do harm, because the processes that are happening — the various causes and effects — are already beyond anyone's control: the laws that the state will pass, what politicians will talk about, what the owners of various companies will say — the smartest people love this: this person stood at the dawn of artificial intelligence, this person created something else.

Alexander Volchek00:15:40

I want to say that all of them are wrong, very often, and we see it — and they are wrong not by fractions of a percent, they are wrong by tens of percent, and so a huge number of spaces lag behind, and the world's population stands somewhere far away and does not understand what is happening at all. You can see it in fairly developed people: I have a huge number of very developed acquaintances, and they really do not understand artificial intelligence at all — even those who say they understand it or know how to ask it something, that is some micro-procedure of use, just a small thing, a little flash.

Alexander Volchek00:16:16

It is like saying you know all about gardens and flowers and vegetable plots when all you can do is water them. Most of the world, even those who talk about artificial intelligence, has a kind of ability to water — and even that is a small part of the funnel, because the bulk of people, of course, do not know at all; they use an artificial-intelligence system simply as question-and-answer. That is fundamentally so, basically so — we must all understand it. Coming back — to come back to Coxon — there are very interesting aspects there about how Anthropic does not skimp on safety, and we all remember that Anthropic announced it has to spend almost 20% of its whole budget on current safety — of the compute, of all conditionally, all the compute it has; I do not know whether that is the budget of the whole corporation or of the compute it has, but it is very serious money.

Alexander Volchek00:17:14

And judging by their statements, by their policies, by the description of their constitution, philosophy, documents — they want this not to happen. And, probably, the people inside who create this would still not want such a moment to come, although I think there are always people who would like to create some incomprehensible mind. But once again — the movement itself: let's recall OpenAI's mission; OpenAI was originally heading towards creating AGI. Well, of course it had a mission to create a free, open artificial intelligence, but basically to create AGI, to create superintelligence.

Alexander Volchek00:17:52

But as soon as you create a super-artificial intelligence, it is already out of your control. And here, to continue, I want to give you a very interesting case concerning Anthropic and biological weapons. So, Anthropic reported literally a few days ago that there had been five suspicious episodes of Claude being used in biological research. Importantly, the company does not claim to have proved an intention to create a weapon. But honestly, when you sit and read about these viruses and these cases — some cases, as it were, showed — you understand how many people in the world are trying to do incredible things.

Alexander Volchek00:18:38

That is, we all understand that GPT-6 Astra and the Fable systems block you if you come to them. And there are no such silly examples where a person comes and says: I want to create a nuclear weapon. All these systems, obviously, have been blocked for a long time — a great many years ago, in fact. But imagine some corporation buying 10 thousand accounts in completely different, non-overlapping places, and in those 10 thousand accounts starting to create something in small pieces, in chains, independently.

Mentions: Astra
Alexander Volchek00:19:10

And independently developing certain lines, and then at some point their own scientists or their people think about how to connect it. Or they use some open models, or their own models, to do it. Who are these people, and why were they suspected at all? So, this group, by the way — to which the smallpox virus belongs, for example — Claude Opus 5 in about an hour helped prepare an application for research into the virus's interaction with immune defences. Studying these mechanisms can serve opposite purposes.

Alexander Volchek00:19:41

That is, you say you are preparing an application, you want to research something, and then it turns out you are military, in some particular laboratory. The names of institutions and countries in the biological section are not disclosed. It reports researchers from regions where Anthropic does not provide its service — because there are regions in the world where Anthropic does not provide it. We may show it to you now. If our editors prepare it — I apologise for referring to them periodically — but if my team prepares it, manages to do it, manages.

Mentions: Anthropic
Alexander Volchek00:20:10

A lot of episodes go out from us. You know about that — by the way, support our team, support us, me, our ToTheMoon channel. An incredible number of episodes already — more than 160 episodes. We now come out 3–4 times a week and try to give you super high-quality content. Any like, support, comment, forwarding us to your friends, subscription — it gives us. It is the only thing on YouTube that gives us promotion, whatever you write, whatever you do — so support us. So, and they also said there were connections — parts of the research with government structures.

Alexander Volchek00:20:42

And so, for example, attributing these 5 cases specifically to — I don't know — the banned countries where Anthropic is prohibited — Iran, China or Russia — on the basis of this publication is in no way allowed. The suspicion was raised by a combination of various circumstances: the nature of the research, some connections that appeared inside, various — I don't know — institutes or corporations, attempts to hide, to get access to the models bypassing various restrictions, and so on.

Mentions: Anthropic
Alexander Volchek00:21:06

So, at the same time, again — if, say, there is a company that has state funding, or, I don't know, is using some potentially dangerous topic — that does not mean, once again, it does not prove that this company wants to create, say, develop some weapon, or that it is pursuing some medical or military goal, yes — we must understand this perfectly well, yes; only here they did not write it out in the open, because people really can be doing research, and I have often given you the example of when my wife and I were walking down the street and she says: listen, how do young people in the US take drugs?

Mentions: United States
Alexander Volchek00:21:39

She types: how do young people in the US take drugs? And ChatGPT answers — then writes: if you have problems with drugs, here is a phone number. I say: Polina, listen, you can't ask ChatGPT a question like that, you have to write: I saw an article recently, a news item, that people take drugs — can you tell me how they take them in such-and-such countries, compare a table for me — and it produces it and does not tell me that I may have a problem, yes. I say: that is how you make the system tick its box — so these are examples of how to ask questions correctly, and, by the way, this is also an important story, because the systems really are interested in not harming either society or you as a person, yes, and especially a minor, or a person who has some difficulties perceiving information, some disorders — these systems try their utmost to prevent that, because all sorts of things can happen in life, and a person can do something, and we know a huge number of cases in the world today where artificial intelligence leads a person to tragedies, various ones — but again, it is not artificial intelligence that leads to tragedies; you can read an incredible number of books, rewatch films, that will lead to tragedies in exactly the same way, yes.

Alexander Volchek00:22:45

So, what does it mean that in these cases of biological-weapon development Anthropic interfered? It means the company blocked the accounts that were there — and, by the way, blocked a great many accounts — strengthened the protection of its own services, and we may, for example, run into that restriction later; and in some appropriate cases it passed information to the authorities or, for example, to other developers. And there are, by the way, very interesting cases in the world where various threats were prevented even in other countries — for example, when OpenAI in South America passed data to prevent certain murders and certain threats, working directly together with the authorities, yes.

Alexander Volchek00:23:23

So — that is, first of all Anthropic cut off access to its service, rather than physically stopping a laboratory and so on, which it accordingly did — and cannot do — in many countries. So, a significant detail: in the first episode an intermediary restored access just a few days after the block. Anthropic noticed that, yes — Anthropic later saw materials indicating that the research continued; that is, note how the system works there, that they can connect this information together, yes — and therefore the wording 'stopped the creation of a weapon', as it were, cannot be considered an established result, yes.

Mentions: Anthropic
Alexander Volchek00:23:58

So, why did the company consider this so serious at all — that is, it stated that its previous models had not reached the level — well, of course, here Anthropic is already talking about models of the Mythos level; Mythos itself is not handed out to just anyone, yes — that is, the companies still used the Fable level, possibly Opus 5, but most likely it is Fable being talked about now. Accordingly, the previous models did not allow — they did not have the ability, essentially, to help so seriously in biological research — although that confuses me a little, because it seems to me that a year or a year and a half ago both OpenAI and Anthropic were also fighting people who were creating biological weapons, for example, or weapons of any kind, or committing crimes in cyberspace.

Alexander Volchek00:24:43

That was already happening, yes — that is, with regard to the new models Anthropic has no confidence that these models are controllable, and therefore the restrictions on sensitive biological requests became stricter; this is, as it were, the company's assessment of this risk, yes, internally. A very interesting story — it seems to me an extremely serious story. By the way, about the chikungunya virus — there was an interesting comment from Anthropic there; why I decided to tell you this, by the way — to say it now, because of this comment — that scientists were trying to get help in preparing an application for funding research that potentially increases the danger of the virus.

Mentions: Anthropic
Alexander Volchek00:25:19

And additional alarm was raised for Anthropic by the fact that the work was to be carried out in a military research institute, yes. Well, and here you go: what happened? Look — these are different zones, but they allow us, it seems to me, to look a little more broadly at everything that is going on. That is, these mathematicians who say that artificial intelligence, as it were, cannot move science at all, yes. And in parallel we see a statement that humanity — humanity may have a question of tens of percent of problems.

Alexander Volchek00:25:50

And in this respect I, of course, do not support the scientific community — precisely in that wording — because there really are questions that go beyond mathematics. I generally believe there are questions that concern the individual person, yes. And still, artificial intelligence will be better and will give more people today the opportunity to help themselves in health, in medicine, including in various kinds of research. And I deeply believe in science, yes. But the market will change, yes.

Alexander Volchek00:26:17

And a huge number of people will have questions. And the market will become controllable in a different way. A huge number of other laws will have to be written. That is, a great deal will be arising. So, in general, what does it mean that Coxon spoke about a self-improving superintelligence — and in terms of his complaints in general. That is, he has a complaint against the very race that exists today. And we know that Anthropic is preparing for an IPO. And it may be one of the most expensive IPOs that has ever happened in human history.

Alexander Volchek00:26:49

The valuation there will clearly already be more than two trillion dollars. And his complaint is that he expects the need for the company to remain competitive. Especially imagine it has gone public: now the shares must not fall, they must grow. That is, it will have to make a dangerous compromise. So if we talk about Coxon's fear, it can essentially be presented as this sequence: artificial intelligence conducts research in the field of artificial intelligence. I have often shown you this sequence, I think.

Alexander Volchek00:27:17

I often — I started back in June, when Anthropic presented that it has models that already fully develop systems on their own, without touching a human. Artificial intelligence conducted research in the field of artificial intelligence. Next, it helps create a stronger model. The new model conducts research even faster. And development accelerates. And he fears that this process will start to outpace people's ability to understand and control the systems being created. That is, it is important to understand that super-artificial intelligence here is not just a good conversation partner or a successful program that, I don't know, passes exams, or helps you learn a language, or successfully made a purchase for you, or chose the right hotel, or wrote — made you a website, or wrote some program.

Alexander Volchek00:28:01

That is, it is a substantial system that surpasses a person and groups of people — essentially countries — in intellectual capabilities. And so this alignment of artificial intelligence's behaviour with human intentions and constraints — that is, it is like a story of levelling a certain neutrality: how to ensure that a more capable system really performs the required task in a safe way. This is a task, by the way, that is a separate technical problem, which OpenAI itself described in 2023.

Alexander Volchek00:28:32

And I, of course... By the way, Coxon also describes in his latest message — he mentions reinforcement learning; there is such a concept, let me remind you that OpenAI recently paused — I do not know whether it continued or not — reinforcement learning in its latest models, under which certain results receive a reward and the corresponding behaviour is reinforced, yes. That is, the model... It is a very strong movement forward of the model itself, yes. The problem is that the model can learn to get a high score not for the right solution but for bypassing the check, essentially.

Mentions: OpenAI · Jacob Coxon
Alexander Volchek00:29:08

And such cases, again, were described in OpenAI's technical report. And the topic is incredibly interesting. When I say incredibly interesting, that does not mean I feel some positive emotion or some positive thrill — but it is worth thinking about and worth considering very broadly. And why should you consider it? It should be considered not from the standpoint of whether the world is collapsing or not — not that at all — but from the standpoint that when you make decisions in your life a decade ahead in terms of your professions, in terms of recommendations on what to do, in terms of discussing your professional self at work, in terms of interacting with your children, parents and colleagues, in terms of your own state — a feeling somewhere of a lack of knowledge or, on the contrary, of endless participation in the race, an endless desire to earn money, or, on the contrary, paying no attention to it — with your own state, still understanding what this artificial intelligence actually is, at least a little.

Alexander Volchek00:30:02

And that gives an incredibly strong opportunity for the development of a modern person. And in no way does it mean that you have chosen the side of the artificial material. But you are at least showing your genuine, own awareness in order to understand what is happening. You have been on the ToTheMoon channel. Write your comments on today's topic. See you in the next episode. Bye, everyone.