OpenAI has shown how its employees work with artificial intelligence, where for at least half of the researchers the use of the models would cost more than six hundred dollars a day. If you count it, of course, at the prices for external clients. That is not a salary, that is only the work of the artificial intelligence itself. What do they entrust to it for such sums? What have they already stopped doing themselves? And why, even with such capabilities, is a person still needed?
Today we will look at what has changed inside OpenAI over this year and what of it matters for your work. Watch the episode. On the sixth of September OpenAI published a very interesting report on how its own researchers — well, the employees inside — work together with artificial intelligence and with assistants inside. The company stated that it had reached the level of an automated research intern, so to speak. That is the term, right? The system could be given specific research assignments that would take a qualified employee — a qualified person in general — several days.
But setting the task, directing it, supporting it and checking it stay with their employee, with the person. And the main change in this study is that artificial intelligence begins, essentially, not just to suggest to this researcher what to do, but to perform a substantial part of the work. For example, to sort out software bugs, help launch experiments, analyse the various results obtained. That is, to do the whole spectrum of this analyst-researcher's work, right? And OpenAI described such use in the GPT-5.6 materials. And against this background three figures appeared: six hundred dollars, seven thousand dollars and, roughly, three point one working days. The first two figures show how much the use of AI assistants would cost at the external price list for individual people, if it were not inside their company.
And the third is how much time in total the agents work relative to people. None of these figures by itself means that these researchers have been replaced or that scientific work has sped up threefold. What exactly OpenAI studied, according to the description they gave, was an analysis of its own working practice.
It was not some experiment where a group was randomly selected and, say, left without artificial intelligence while another was given artificial intelligence. Here, essentially, the company studied the use of agents. There was also work with code, launches of various experiments inside, among its own employees. And, well, the results over several months of work this year, right? The difference was very substantial. Such an observation helped answer the question: how is work changing inside the company? But to answer how much exactly artificial intelligence raised productivity, one would need separately to separate its influence from other changes. For example, the team could at the same time have received more powerful computers, improved the software infrastructure, hired some additional specialists.
Then, essentially, the increase in the number of experiments cannot automatically be credited entirely to artificial intelligence. And now I will give you all the details. The data is actually very interesting. I think it is very positive for understanding what is happening with artificial intelligence now, in particular in combination with work. And personally I was very interested in this story about what six hundred dollars a day means. And there the numbers even reached seven thousand dollars a day. At first the mind does not fully take in what that is, but then you realise that these are fairly serious figures. To understand this story, once again, it is useful to imagine what lies between the idea "let's improve the model" and the proof that the improvement really works.
Below I will give you an example, roughly. It is a separate experiment from this report. A researcher assumes that a new organisation of the training data will help the model make mistakes less often, right? And to check it, writing down the assumption is not enough. You need to prepare the data, make changes to the program, then launch the training or the check, deal with possible failures, compare the various results and so on. Part of this work can essentially be distributed between artificial intelligence agents, between assistants.
One prepares, for example, the changes in the code, another writes the checks that should detect errors, and a third works out why the experiment ended in failure. A fourth, for example, collects the results of several runs for comparison. That kind of mechanics matches the principle of OpenAI's Codex for working with code. It can read and change files, run commands and tests, and then hand over the changes and results for a person to check. And essentially this is not just some recommendation like "try fixing some line inside"; in this case, and in our example, the researcher afterwards has to answer substantive questions. Were identical conditions compared, for example? Did the errors really decrease, is the result reproducible?
Or, for example, is it worth continuing this line of testing at all? And that is where the potential saving arises — on doing the work. So the value of a research decision still has to be proven separately. Now let us look at what these six hundred dollars mean, the ones I spoke about at the beginning.
So what is the point? OpenAI said the following: on average each of their employees — where they ran the study, in terms of analytics, these researchers — consume six hundred dollars' worth of models. This average price is imprecise, because OpenAI does not pay for the models, and they calculated how much other people would pay for such models. There were those who spent less than six hundred dollars — it is the median value. There were those who spent more than six hundred dollars and, in fact, went as high as seven thousand.
That is, six hundred dollars a day actually does not look that small. So there is an analyst employee, he has his own salary, and you could think of it this way: you hired a programmer. This programmer uses models, and to perform various tasks he spends on average, I don't know, three hundred dollars, or five hundred dollars, right? No precise examples can be given here. Some of you may ask, for example: "What can you spend six hundred dollars a day on?" Well, there is no such story.
It is stated that it could have been a hundred different tasks. Look, one more very interesting figure they had is the amount of time that the
agents work. And what was shown there overall? That the agents started working more, if you compare the time before June and now in August. That is, before June the figure was that on average an agent worked less than a day, roughly less than eight hours. What is that? It means how much time various tasks were being performed by these agents inside. An agent, clearly, is also a somewhat notional construct, because it may be a multitude of agents, right? Let us keep it that way and use that description. What interesting thing happened? That the time now, in August, has grown and become more than three days.
That is, essentially, the agents started working twenty-four hours. And why do they say in the description that agent time has become greater than human time? As we said, it used to be less than eight hours of use, now it is more. In the same section OpenAI shows a growth in the number of researchers using four or more agents at the same time. And this statistic counts the daily peaks of parallel work, both of agents launched by a person and of additional agents they created to perform subtasks.
So it matters that an agent could launch additional agents, right? That is why the ratio can, essentially, exceed one. Human time runs sequentially. Clearly, a person works eight hours, while several agent processes accumulate their time simultaneously. And this is the second story we are discussing today, right? The first is essentially this average of six hundred dollars. And the second story is the three days, right? And here it matters that this does not mean an agent has to work for several days without stopping, right?
At the same time, the daily peak of parallelism and the total agent hours are, once again, different figures, right? Briefly launching many agents and sustaining long parallel work are not the same thing, right? And this ratio of three — actually three point one — does not mean that each employee constantly had exactly three assistants working, right? Just so we all understand this. And the company, by the way, also reported that agents were more often being given longer and more complex assignments.
And, well, that is probably understandable, right? After all, even from June to August, or from the start of the year to August, there have been really very serious changes. By the way, I want to add one thing before we move on to the real conclusions that came out of this study. After all, GPT-6 Astra has come out, right? And relative to Astra, this study was done within GPT-5. And regarding Astra, Sam Altman, for example, says that Astra is worth trying not just as a conversation partner answering some questions, but as a tool for doing whole complex jobs. And this story, it seems to me, is very telling in this study they have now done.
And Altman describes Astra precisely as the next step towards models that will help create products. And in fact, by the way, Altman says that they will help launch companies and do very serious research. And as examples he cites creating a full program in a dialogue with artificial intelligence, developing various computer games, home projects and so on. By the way, how do you like Astra? Because it seems to me it is still too early to say whether Astra is actually great or not.
A few days ago I released an episode about Astra, but I think at least a month or two will pass before we can truly assess its work. Some write that Astra-6 is a fundamental revolution. I am probably not ready to draw such conclusions for myself yet. I do not see a truly cosmic, super result. Maybe there are some interesting things, but for me everything is still very relative. But I know it usually takes several weeks, sometimes months, to truly see the results and assess a model.
And OpenAI, most likely, of course, was in a hurry with the launch — in a hurry ahead of Anthropic's IPO, in a hurry because Anthropic released Fable 5.1. They had to make this move. I hope they have not made a mistake. Right? And another question is whether they released Astra without hard restrictions. Because if there are hard restrictions, that will be a problem. And Astra, after all, was and is a model with super capabilities in terms of cybersecurity. But now let us continue about the study, about these six hundred dollars and three days, in terms of the conclusions OpenAI came to and, in particular, the fact that they created this analyst, this research intern. So look at what details and changes the company saw inside.
OpenAI basically reported faster code contributions. And the August number of experiments per active human experimenter became the highest since the start of all the observations they have had over practically two years. That is, they started them at the beginning of twenty twenty-five. There is one point: at the same time the available computing power also grew, right? So if you break it down, the practical meaning can be explained as follows. Imagine that a researcher who works inside on their team came up with five variants of an improvement, right?
But because of the volume — of some improvement of a task or of analysis — he has a limited amount of time, and because of the volume of preparation he only managed to check this one variant. If the preparation becomes easier, there is a chance to test the rest, the results. In this case, by the way, even a negative result can be useful. And the team — yours, or in particular theirs in the description — can more quickly find out that some direction does not work, or stop spending time and money on it.
Which, by the way, is extremely critical, right? In business I understand, especially with a large number of people, how phenomenally necessary this is. Although at the same time I understand that most corporations today, and more or less large companies in general, will not be able to cope with this at all. Because if companies are now used to holding a huge number of meetings, endless discussions of anything, then what will happen next, right? Well, they launched agents, ran the studies, saw, for example, that some direction does not work. And then what?
By the way, the last example is especially clear for business in this case, as we discussed. Say, previously an employee went to a colleague with every complex error. For example, an analyst went to a programmer, to some technical colleague, and waited for help. And if part of such questions is resolved with the help of artificial intelligence, then time is potentially freed up both for the one who asked and for the one who answered. This is possibly the mechanism of benefit: removing these everyday delays, and not only writing text or code faster.
On the other hand, as I said, there is a flip side to this. If a company has established processes and you cannot get around them. Remember my examples. When I was developing some solutions, I understood that these solutions would previously have been made by a designer, a programmer, analysts, business analysts, systems analysts, some project managers, marketers, salespeople. That is, a lot of people were involved. And I managed to do a number of projects because, roughly speaking, I excluded all of them, including the people who were doing any kind of sign-off. Because if there are bosses above who approve everything, that is also quite a big problem.
Why is OpenAI's system still called an intern? Although, of course, this story about six hundred dollars is very interesting to me, right? But write your thoughts later. And, of course, it is interesting that this agent work reached more than three days. Now, about the intern. More than half of the successfully completed assignments — rated, they had a frame of four to eight hours of human work — required at least one intervention by a person. So that is the construction. And here there are two clarifications.
First, four to eight hours is an estimate of how long it would take a person, not the duration of the agent's work. And second, successfully completed does not mean done entirely on its own. Imagine, for example, that an agent implemented changes and got good metrics. An analyst working at the company, or a researcher, noticed that the comparison was made on different data. He pointed out the error. The agent redid the check. That is, the artificial intelligence redid the check and finished the task correctly.
So the result is successful. But without the substantive intervention of a person, the outcome would have been wrong. So intern is, above all, a description of a certain degree of independence. And let us recall the episode I made a few days ago about Astra. I told you there is a very interesting figure in Astra. It is GPT-6's ability to perform many tasks without human intervention. And on a number of tasks this figure grew two or three times. And here I think it is very interesting what OpenAI will say next about its systems — whether it will still call them interns or not.
So intern is above all a description of independence. The system may have very broad knowledge, quickly perform technical work, but still needs human help and guidance. That is why the researchers propose assessing automation across the different components of research work rather than reducing it to a single ability to write code. And success on one type of assignment does not automatically cover all professions. Which we have seen many times. And so there are people who say — for example, Tanya often says in our episodes: "Artificial intelligence has not even come close to solving my tasks." Although I am sure that, for example, in interior design part of the tasks can definitely be solved by artificial intelligence, while part really requires human support and human intervention. It would be interesting to hear your cases. But in many of my own cases I see where I give a task to artificial intelligence.
The artificial intelligence, for example, starts going in completely the wrong direction. It seems it would have done something completely different. I correct it, and it gives me a really great final result. And, by the way, many people recommend setting up such a thing as skills. For example, in Codex there is a skill: you wrap certain actions into a kind of function described in advance. But I understand that this skill may later start working incorrectly too. By these figures, by the way, one can in no way establish the number of replaced workers, or the size of the savings, or say, for example, that this profession no longer exists.
That is, agent time and the notional cost of using them contain no answer to these questions. Here it is more correct to speak of handing over part of the operations a person does — that you do — and not automatically the whole position. There may simply be some correlation here, but it does not hold everywhere. Say, previously a specialist spent, I don't know, six hours on preparing some check. Now, for example, an agent does that preparation, and the person spends part of the time on setting up and checking this work.
In that case, five hours seem to be freed up. On the other hand, are they really freed up in all cases? That is a big question, because if, for example, the agent prepared the work incorrectly and the specialist had to spend another six hours redoing it, the saving disappears completely. Or, as I have noticed, when building various systems there is a situation where something could simply not have been done. It is like us today. I am in France right now, and my friend and I were travelling.
And we were driving, and in one town there was a sign turned upside down. And I say: "Listen, that's interesting." He says: "Ask ChatGPT why the sign here is upside down." And I drive on and think to myself: I do not want this information at all. So it turns out that if you consider not even building systems but obtaining information, we have gained the ability to perform a huge number of different operations or actions, or to study something as an analyst. But the question is whether we need that study in life at all, or whether we can do without it.
As I, for example, could easily have done without this upside-down sign today, because most likely, had I learned that information, I would not remember it at all in six months. Just as today I was in the south of France, and I was in the town of Ménerbes. A village standing in the mountains, very beautiful. And I had been in this town, in this village, six or seven years ago. And I remember nothing at all. Time passes and you forget that you ever visited it, you forget certain visuals and so on.
to safety, because the more actions the agent was allowed, we understand that if an agent starts working twenty-four hours within a person's working day, then it has fairly serious access. And accordingly it is very important to control not only the result but also the way it was obtained. And OpenAI made an announcement in August describing a two-week pause in the reinforcement learning of the model. We talked about that too. They stopped a great deal, because that pause was also connected with the Hugging Face hack.
But they saw that the frontier models had started developing too strongly and had gained some frighteningly independent control.
And what will happen if all employees use such huge volumes of agents? What will these agents actually be doing? Yes. And besides, resources are consumed very heavily, because, for example, OpenAI separately estimated the additional computing load from monitoring. Watching over safety, for example, and in general. All this computing power they have now has grown. And to watch over safety, they estimated that an additional twenty percent of the supervised model's work has to be spent.
That is very serious money. And also, if we take 600 dollars and multiply it by even 20 working days, that is very serious money, right? That is an additional 12 thousand dollars to be spent per employee per month. It is one thing if this employee earns tens of thousands of dollars a month. It is another if this employee is in a country where the salary is 500 dollars or a thousand dollars, right? That is a big question — what will come of it. In closing, I want to tell you the following.
I think it is very important, because we make these episodes so that you develop your understanding of artificial intelligence and its application in your life, both professionally and in your everyday life. What does OpenAI want to arrive at with such studies? They have a certain next stated goal that they set within this study. That they will have an automated researcher, an automated analyst. They often, of course, lean heavily towards scientific analytics. They will have an automated researcher by March 2028. But again, also in the context of human oversight.
Very important. This is a development target. They are in no way saying that they will manage to build it. And in the accompanying essay that comes with the study, OpenAI's chief scientist Jakub Pachocki describes a strategy in which artificial intelligence takes an ever greater part in creating the next generations of artificial intelligence itself. Remember, a few months ago Anthropic published material saying that their artificial intelligence takes part in creating itself, right?
Well, now OpenAI has started talking about it a lot too. At the same time, by the way, Pachocki stresses the need to keep control and to match the speed of development with the reliability of safety measures. Again, this is his position, not a forecast, not some proven outcome. The logic of the potential cycle is very simple. There is the current model, it helps this analyst. They create a stronger model, and the new model in turn becomes a more useful assistant in the next study that will take place.
And it is very interesting what will happen, in particular, given even how OpenAI's employees use the models. Yes, it was interesting for me to see these three days of agent work, because it seems to me that on many of my days the agents work more than that — in a day they do the equivalent of ten days. At least in terms of my parallelisation. That is, it is not just 24 hours. I have also seen this with some of my employees on certain tasks, on a number of tasks. I see how such things can change from day to day, because sometimes I, for example, run dozens — sometimes even, I think, hundreds — of chats in parallel, or dozens of different tasks in Codex, in Claude Code, in parallel in ChatGPT. And sometimes everything is very calm and there is no endless agent work.
And there are programmers who say that they have thousands of agents working twenty-four seven. How do they manage to check them? How do they manage to think up the business logic and the tasks? What kind of system is that? And I, for example, sometimes reach these parallel dozens of agents only because it usually happens in tasks where, roughly speaking, I am the only participant, right? Because if a large number of other people are involved, that is impossible to achieve. You often ask another person a question and wait three or four days for an answer.
Tell me how it works for you. How many such days do you have? How much money do you spend? Well, maybe there will be few people here who spend 600 dollars a day, or 7 thousand dollars a day. But still, it is interesting how you assess the work of your agents in parallel. See you in our new episodes. Bye everyone. Do not forget, of course, to comment and support our videos.