GPT-5.4 Is Stronger, but Google Wins Where the Model Already Lives Inside the Documents
Why does a stronger GPT-5.4 not guarantee victory over Google where the model already lives inside documents and work context?
Test whether Google and GPT become useful everyday interfaces or require constant correction. The decision requires the reader to check how many steps the interface actually removes and what dependency it creates in return.
What to watch for
Key takeaways
In the context of “GPT-5.4 Is Stronger, but Google Wins Where the Model Already Lives Inside the Documents,” this criterion applies: the race shifts from the quality of one answer to the workplace: who can see the files, understand the context, and verify the result without ten switches.
In the context of “First impressions from GPT-5.4,” this criterion applies: GPT-5.4 makes noticeable progress on difficult tasks, but the interface again forces you to sort out Instant, Thinking, and Pro; a strong model should not require the user to pick its internal engine.
The “ChatGPT Pro vs Plus” scene leads to a working conclusion: the difference between the tiers is justified only on heavy tasks; a new mode does not make Pro automatically necessary for everyday work.
The “Plus and minus GPT-5.4” topic becomes clearer once this point is included: in Heavy mode the result is deeper and part of the reasoning is visible, which helps review, but a visible process is not a guarantee of correctness: a model can explain a mistaken path beautifully, so the result must be checked against the data and the real task.
For the “Problem of new models: quality has become unstable” scene, the decisive point is this: unstable quality means the result has to be checked against the real task each time; neither a benchmark nor a first impression guarantees stable behavior.
In the context of “Anthropic: which occupations are most affected by AI,” this criterion applies: a ranking of affected professions describes exposure, not the outcome; what matters is which tasks get automated and who is responsible for the result.
The decision in “Amazon restricts juniors and mid-level devs because of AI code” depends on one criterion: cutting the hiring of junior and mid-level developers because AI writes the code shifts the risk: someone still has to review the changes and answer for them, and fewer trained people is a separate long-term cost.
The decision in “Anthropic launches AI for code review” depends on one criterion: technically Codex can run a large project, but its interfaces lag behind, so a narrow code-review service is logical: the model does not just write code but reviews changes in project context and looks for risk before a merge.
The decision in “Big Gemini update: integration with Google Docs, Sheets, Slides, and Drive” depends on one criterion: Google's strongest move is distribution: Gemini gets the documents where they are already created, so the user does not have to upload a file and explain its structure, while OpenAI will have to integrate more deeply or build its own.
The discussion of “OpenAI buys an AI-safety startup” yields a practical test: code generation without control no longer sells as a sufficient answer; the winner is not the model with the highest number but the environment that understands work context, shows changes, and does not make you guess the right mode every time.
What this episode is about
OpenAI is improving Thinking and Codex, Anthropic is launching a separate code-review product, and Google is putting Gemini into Docs, Sheets, Slides, and Drive. The race is shifting from the quality of one answer to the workplace: who can see the files, understand the context, and verify the result without ten switches.
GPT-5.4 shows noticeable progress on difficult tasks, but the OpenAI interface once again forces people to understand Instant, Thinking, Pro, and reasoning levels. A strong model should not require the user to keep selecting an internal engine. The more toggles there are, the less the product feels like an assistant.
Heavy mode can produce a deeper result and expose part of the reasoning process. That is useful for review, but a visible process is not a guarantee of correctness. A model can explain a mistaken path beautifully, so the result still has to be compared with the data and the real task.
Codex is developing in an odd way: technically it can manage a large project, while the product interfaces and connections to other tools lag behind. Anthropic is responding with a dedicated code-review service. The narrow product is easy to understand: the model does not merely write code, but reviews changes in project context and looks for risk before a merge.
Google is making its strongest move through distribution. Gemini inside Docs, Sheets, Slides, and Drive receives documents where they are already created. The user does not have to upload a file to a separate chat and explain its structure. OpenAI will either have to integrate more deeply with workplace systems or build its own.
The acquisition of an AI-safety startup and a new verification layer in Codex show that code generation without control is no longer sold as a sufficient answer. The winner will not be the model with the highest number, but the environment that understands work context, shows changes, and does not force a person to guess the correct mode every time.
A device or voice assistant wins when it does not force the user to learn another system and constantly repair its output.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 74 segments: 46 identified, 6 mixed, 18 probable, and 4 unresolved.
Read transcript on a separate page
Loading…