Skip to content
OpenAI · GPT-5 · ChatGPTEpisode 070 · 10 August 2025 · 46:43

GPT-5 Became Easier to Use, but the Main Progress Is Not One Benchmark—It Is How the Model Handles a Task

Central question

What did GPT-5 genuinely change in task execution if progress can no longer be measured by a single benchmark?

What you take away

Compare GPT-5 and Model on a real task instead of choosing by one benchmark or announcement. A practical assessment requires the reader to compare quality, price, access, memory, and control on a real task rather than by one announcement or benchmark.

Main threads

What to watch for

1Compare “GPT 5 presentation: What changed” with “Chat subpersonality in GPT-5”: they provide different criteria for judging the same issue.
2Test the conclusion from “The use cases at the ChatGPT 5 presentation” in your own use case—what actually changes in the process and what remains a promise.
3Before choosing a product or approach, record the constraint identified in “Microsoft 365 Copilot on GPT-5: where growth will be”.
4Define the owner of the outcome and the quality metric for the situation described in “Groq and Cerebras: Inference at high speeds”.
Signals to track afterwards
→Watch for actions by Apple and Google that confirm or challenge the episode’s central claims.
→Compare new launches and policy changes with “The use cases at the ChatGPT 5 presentation”: have access, quality, price, or constraints changed?
→Check whether the scenario in “Groq and Cerebras: Inference at high speeds” becomes repeatable practice rather than a one-off demonstration.
Most useful for
Executives and managersDevelopersTechnical leadersDesignersAI usersProduct teams

Key takeaways

00:00GPT-5 Became Easier to Use, but the Main Progress Is Not One Benchmark—It Is How the Model Handles a Task

The “GPT-5 Became Easier to Use, but the Main Progress Is Not One Benchmark—It Is How the” topic becomes clearer once this point is included: live tests of GPT-5 show speed, Thinking, and Microsoft integrations, but hallucinations, multiple modes, and the comparison with Gemini remain — user experience is finally part of the race alongside intelligence.

01:04A new model usually arrives with a table of wins

The practical meaning of “GPT 5 presentation: What changed” is that the interesting part is not a table of wins but OpenAI's attempt to make mode selection invisible: the user should not decide each time when to turn on Thinking — the system itself gauges a task's difficulty and spends the right amount of time.

03:06Why an announcement is not enough: first impressions of GPT-5 — the user experience

The practical meaning of “First impression from GPT 5: User experience” is that in early tests GPT-5 runs ordinary dialogue faster and more confidently, holds the thread of a request better, and less often makes you repeat context; but hallucinations become less obvious rather than disappearing, so polished prose cannot be taken as proof.

11:34The market tests it through use: has OpenAI fallen behind on technology?

The decision in “Has OpenAI fallen behind on technology?” depends on one criterion: technologically OpenAI has probably fallen behind, perhaps irreversibly, but that is not decisive for the company — B2C matters more to it than API, so the point is not a test ranking but usability, and ChatGPT's “last-decade” design is its main area for growth.

12:07In early tests, GPT-5 conducts ordinary dialogue more quickly and confidently, holds the thread of a request better, and less often forces the user to repeat context

The “Chat subpersonality in GPT-5” topic becomes clearer once this point is included: the “subpersonality” has not gone anywhere — tone, depth, and character of the answer still shift across modes; the option to pick one was announced but not yet shipped, and many features will roll out only later.

19:08Who owns the outcome: video and image generation advance fast — even without OpenAI

The practical meaning of “Video and image generation are advancing rapidly—even without OpenAI” is that in video and image generation the market has taken a huge step without OpenAI — companies like Higgsfield ship updates almost daily, and ChatGPT does not look like a competitor there; but big players usually catch up, and it is a matter of time.

21:06OpenAI’s presentation differs from the Grok competition because it shows mass-market use cases

The discussion of “The use cases at the ChatGPT 5 presentation” yields a practical test: OpenAI's presentation shows mass-market scenarios — not “we rank higher” but how the model helps you understand a task, write, learn, or work; most users will never open a benchmark, yet will immediately feel it when an interface fails to ask clarifying questions.

24:54Why context matters more than one metric: the system that asks questions will win — why?

The boundary of the “The system that asks questions will win: why?” case is defined by this point: a model's success is decided not by a peak score but by whether it asks clarifying questions, integrates into work, and gives a predictable result once the presentation is over.

32:41For programmers, GPT-5 matters through Codex, Cursor-like environments, and the ability to work with a large project

The “Microsoft 365 Copilot on GPT-5: where growth will be” topic becomes clearer once this point is included: anthropic remains a strong competitor here, while Gemini Deep Thinking follows another path. Microsoft is quickly integrating the model into Copilot and Microsoft 365, turning an OpenAI release into an update across a vast workplace ecosystem.

44:10Groq and Cerebras add the question of inference speed: companies sometimes need not the largest model, but cheap and extremely fast processing of millions of requests

The decision in “Groq and Cerebras: Inference at high speeds” depends on one criterion: companies sometimes need not the largest model but cheap, very fast processing of millions of requests, so GPT-5 cannot be judged by one number — inference speed and price matter as much as intelligence.

What this episode is about

Live tests of GPT-5 show speed, Thinking, changes for programmers, and new Microsoft integrations. The model makes a strong first impression, but hallucinations, multiple modes, and comparison with Gemini remain. User experience is finally becoming part of the race alongside intelligence.

A new model usually arrives with a table of wins. GPT-5 is more interesting for another reason: OpenAI is trying to make mode selection less visible to the person. The user should not have to decide every time when to turn on Thinking or which internal model to choose. The system should understand the difficulty of the task and spend the appropriate amount of time itself.

In early tests, GPT-5 conducts ordinary dialogue more quickly and confidently, holds the thread of a request better, and less often forces the user to repeat context. Yet the chat’s “subpersonalities” have not disappeared: tone, depth, and the character of an answer still change across modes. Hallucinations are becoming less obvious, not disappearing, so polished prose cannot be accepted as proof.

OpenAI’s presentation differs from the Grok competition because it shows mass-market use cases. Not only “we rank higher,” but how the model helps a person understand a task, write, learn, or work. That is the right direction: most users will never open a benchmark, but they will immediately notice when an interface fails to ask clarifying questions.

For programmers, GPT-5 matters through Codex, Cursor-like environments, and the ability to work with a large project. Anthropic remains a strong competitor here, while Gemini Deep Thinking follows another path. Microsoft is quickly integrating the model into Copilot and Microsoft 365, turning an OpenAI release into an update across a vast workplace ecosystem.

Groq and Cerebras add the question of inference speed: companies sometimes need not the largest model, but cheap and extremely fast processing of millions of requests. GPT-5 cannot be judged by one number.

Its success depends on how well the system selects a mode, asks questions, integrates into work, and produces a predictable result after the presentation is over.

A new version becomes an advantage only when it can enter a real workflow without costing the user data, time, or control.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 75 segments: 39 identified, 6 mixed, 18 probable, and 12 unresolved.

Read transcript on a separate page

Loading…