OpenAI Taught a Model to Think Longer—but Made Choosing an AI Even Harder
What does a user gain from a model that thinks longer, and why does choosing between models become even harder?
Understand when longer model reasoning creates a real advantage and when it merely increases cost, latency, and the difficulty of choosing a tool.
What to watch for
Key takeaways
The boundary of the “ToTheMoon is a podcast and a hello from Silicon Valley!” case is defined by this point: o1-preview exposes intermediate reasoning and handles tasks where one fast answer is not enough, but the quality gain brings a new chaos: users must choose among several models, and the black box gets deeper.
In the context of “Google Wallet — an interesting novelty from Google,” this criterion applies: Google Wallet now lets you add a passport to present the document electronically — first in the U.S., tied to the TSA system where your ID is already scanned at security: the document finally moves into the phone.
In the context of “We are not Apple fanboys!,” this criterion applies: the hosts take the commenters' charge of Apple bias with a caveat: their phones really are all Apple — for style, security, or habit — but they follow Android too: Google's Gemini-linked Wallet ID feature was the first thing they covered.
The “New model from OpenAI! How does o1-preview work” issue should be assessed with one constraint in mind: the novelty changes the mode of operation itself — before answering, the system spends time reasoning and shows it "thought" for several seconds; for hard math and code that is a step forward, but users now must know when to pick fast 4o, when o1, and when mini.
The “The technical side of o1-preview. Why this model is unlike all the others” issue should be assessed with one constraint in mind: with o1 the market moves toward verticalization: one model suits everyday conversation, another code, a third expensive analysis; the interface grows more complex, and picking the wrong mode means overpaying or getting a weak result.
The working conclusion from “How o1-preview was trained” is that reinforcement learning always needs an auxiliary model checking the first model's results — and that built-in verification reduces hallucinations; the hosts expect a tool within a year that picks between 4o, o1, and mini automatically, because no one will do it by hand.
The boundary of the “B2C or B2B? Where AI solutions will go” case is defined by this point: Meta and other open-model builders face their own dilemma — developers fine-tune the mass-market Llama for themselves, but a reasoning model demands entirely different economics and infrastructure: open and closed markets diverge no longer by license but by product type.
The boundary of the “Will ChatGPT cost $2,000 a month? For whom?” case is defined by this point: the talk of plans costing hundreds or even thousands of dollars is a bet on professional scenarios: the price is justified only where the model's long reasoning replaces expensive human work rather than merely speeding up chat.
The discussion of “ChatGPT competitors: why the discussion still centers on OpenAI” yields a practical test: Anthropic is strong as a developer platform, and Claude 3.5 Sonnet already beats ChatGPT on some real tasks, but OpenAI has the chance to build an Apple-grade brand: expensive, mass-market, and clear to people who have no wish to study architecture.
The boundary of the “Black box in the black box” case is defined by this point: we see a polished description of the chain of thought but do not know how closely it mirrors the real internal process: one black box has learned to explain another — and trusting it has become harder, not easier.
What this episode is about
o1-preview exposes an intermediate reasoning process and handles tasks better when one fast answer is not enough. The quality gain brings a new kind of chaos: users must choose among several models, companies must decide between a mass product and an expensive professional tool, and researchers must admit that the black box has become even deeper.
OpenAI's new model matters not because it answers a few percentage points better. It changes the mode of operation itself: before responding, the system spends time reasoning and shows that it ‘thought’ for several seconds.
For difficult mathematics, programming, or hypothesis testing, that is an important step. But it creates a new problem for ordinary users—now they need to understand when to use fast ChatGPT-4o, when to use o1, when to choose mini, and why they have to choose at all.
Until now, the difference between leading models was often visible only on specific tasks. With o1, the market is moving toward specialization: one model may be better for everyday conversation, another for code, and a third for expensive analysis.
That makes the interface more complex and raises the cost of choosing incorrectly. A person either overpays or receives a weak result and concludes that AI as a whole does not work.
Meta and other builders of open models face a separate dilemma. A widely available Llama lets developers fine-tune the system for their own tasks, but a complex reasoning model may require completely different economics and infrastructure.
You cannot simply increase the parameter count and expect the same effect. Open and closed markets are therefore beginning to diverge not only by license, but also by product type.
At the same time, OpenAI is moving more visibly toward becoming a consumer company. Anthropic is strong as a developer platform, and Claude 3.5 Sonnet already beats ChatGPT on some real tasks, but OpenAI has a chance to build an Apple-like brand: expensive, mass-market, and understandable to people who do not want to study architecture. That is where talk of plans costing hundreds or even thousands of dollars for professional use cases comes from.
The paradox of o1 is that the model is supposed to bring us closer to more intelligent AI while making that intelligence less transparent. We see a polished description of its chain of thought, but we do not know how closely it reflects the real internal process. One black box has learned to explain another black box—and trusting it has become harder, not easier.
The model race is not won by the highest score alone, but by a system that solves the task repeatedly and makes its price and constraints clear.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 51 segments: 36 identified, 0 mixed, 9 probable, and 6 unresolved.
Read transcript on a separate page
Loading…