Skip to content
OpenAI · GPT-5 · ChatGPTEpisode 070 · 10 August 2025 · 46:43

GPT-5 Became Easier to Use, but the Main Progress Is Not One Benchmark—It Is How the Model Handles a Task

Central question

What did GPT-5 genuinely change in task execution if progress can no longer be measured by a single benchmark?

What you take away

Compare GPT-5 and Model on a real task instead of choosing by one benchmark or announcement. A practical assessment requires the reader to compare quality, price, access, memory, and control on a real task rather than by one announcement or benchmark.

Main threads

What to watch for

1Compare “GPT 5 presentation: What changed” with “Catal subacity in GPT 5”: they provide different criteria for judging the same issue.
2Test the conclusion from “The cabs at the ChatGPT 5 presentation” in your own use case—what actually changes in the process and what remains a promise.
3Before choosing a product or approach, record the constraint identified in “Microsoft 365 Copilot on GPT-5: where growth will be”.
4Define the owner of the outcome and the quality metric for the situation described in “Groq and Cerebras: Inference at high speeds”.
Signals to track afterwards
Watch for actions by Apple and Google that confirm or challenge the episode’s central claims.
Compare new launches and policy changes with “The cabs at the ChatGPT 5 presentation”: have access, quality, price, or constraints changed?
Check whether the scenario in “Groq and Cerebras: Inference at high speeds” becomes repeatable practice rather than a one-off demonstration.
Most useful for
Executives and managersDevelopersTechnical leadersDesignersAI usersProduct teams

Key takeaways

00:00GPT-5 Became Easier to Use, but the Main Progress Is Not One Benchmark—It Is How the Model Handles a Task

The “GPT-5 Became Easier to Use, but the Main Progress Is Not One Benchmark—It Is How the” topic becomes clearer once this point is included: the forecast can be tested through specific dates, company actions, and changes in the product or market.

01:04A new model usually arrives with a table of wins

The practical meaning of “GPT 5 presentation: What changed” is that this section clarifies the mechanism behind the topic and preserves a constraint that would otherwise be lost in an overly simple conclusion.

03:06Why an announcement is not enough: first impression from GPT 5: User experience

The practical meaning of “First impression from GPT 5: User experience” is that a benchmark measures a narrow capability; working value requires repeatability, a clear cost, and control over errors.

11:34The market tests it through use: openAI's back in technology

The decision in “OpenAI's back in technology?” depends on one criterion: the forecast can be tested through specific dates, company actions, and changes in the product or market.

12:07In early tests, GPT-5 conducts ordinary dialogue more quickly and confidently, holds the thread of a request better, and less often forces the user to repeat context

The “Catal subacity in GPT 5” topic becomes clearer once this point is included: this section clarifies the mechanism behind the topic and preserves a constraint that would otherwise be lost in an overly simple conclusion.

19:08Who owns the outcome: video and image generation are advancing rapidly—even without

The practical meaning of “Video and image generation are advancing rapidly—even without OpenAI” is that the important signal is not one funding number: the next round, available runway, and closure rate show whether a company can survive the new cost of capital.

21:06OpenAI’s presentation differs from the Grok competition because it shows mass-market use cases

The discussion of “The cabs at the ChatGPT 5 presentation” yields a practical test: this section clarifies the mechanism behind the topic and preserves a constraint that would otherwise be lost in an overly simple conclusion.

24:54Why context matters more than one metric: win a system that asks questions: Why

The boundary of the “Win a system that asks questions: Why?” case is defined by this point: a benchmark measures a narrow capability; working value requires repeatability, a clear cost, and control over errors.

32:41For programmers, GPT-5 matters through Codex, Cursor- environments, and the ability to work with a large project

The “Microsoft 365 Copilot on GPT-5: where growth will be” topic becomes clearer once this point is included: anthropic remains a strong competitor here, while Gemini Deep Thinking follows another path. Microsoft is quickly integrating the model into Copilot and Microsoft 365, turning an OpenAI release into an update across a vast workplace ecosystem.

44:10Groq and Cerebras add the question of inference speed: companies sometimes need not the largest model, but cheap and extremely fast processing of millions of requests

The decision in “Groq and Cerebras: Inference at high speeds” depends on one criterion: this section clarifies the mechanism behind the topic and preserves a constraint that would otherwise be lost in an overly simple conclusion.

What this episode is about

Live tests of GPT-5 show speed, Thinking, changes for programmers, and new Microsoft integrations. The model makes a strong first impression, but hallucinations, multiple modes, and comparison with Gemini remain. User experience is finally becoming part of the race alongside intelligence.

A new model usually arrives with a table of wins. GPT-5 is more interesting for another reason: OpenAI is trying to make mode selection less visible to the person. The user should not have to decide every time when to turn on Thinking or which internal model to choose. The system should understand the difficulty of the task and spend the appropriate amount of time itself.

In early tests, GPT-5 conducts ordinary dialogue more quickly and confidently, holds the thread of a request better, and less often forces the user to repeat context. Yet the chat’s “subpersonalities” have not disappeared: tone, depth, and the character of an answer still change across modes. Hallucinations are becoming less obvious, not disappearing, so polished prose cannot be accepted as proof.

OpenAI’s presentation differs from the Grok competition because it shows mass-market use cases. Not only “we rank higher,” but how the model helps a person understand a task, write, learn, or work. That is the right direction: most users will never open a benchmark, but they will immediately notice when an interface fails to ask clarifying questions.

For programmers, GPT-5 matters through Codex, Cursor-like environments, and the ability to work with a large project. Anthropic remains a strong competitor here, while Gemini Deep Thinking follows another path. Microsoft is quickly integrating the model into Copilot and Microsoft 365, turning an OpenAI release into an update across a vast workplace ecosystem.

Groq and Cerebras add the question of inference speed: companies sometimes need not the largest model, but cheap and extremely fast processing of millions of requests. GPT-5 cannot be judged by one number.

Its success depends on how well the system selects a mode, asks questions, integrates into work, and produces a predictable result after the presentation is over.

A new version becomes an advantage only when it can enter a real workflow without costing the user data, time, or control.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 75 segments: 39 identified, 6 mixed, 18 marked with ✓, and 12 unresolved.

Loading…