Skip to content
OpenAI · OpenAI o1 · AppleEpisode 024 · 22 September 2024 · 32:47

OpenAI Taught a Model to Think Longer—but Made Choosing an AI Even Harder

What to watch for

1Compare “New model from OpenAI! How does o1-preview work” with “The technical side of o1-preview. Why this model is unlike all the others”: they provide different criteria for judging the same issue.
2Test the conclusion from “How o1-preview was trained” in your own use case—what actually changes in the process and what remains a promise.
3Before choosing a product or approach, record the constraint identified in “ChatGPT competitors: why the discussion still centers on OpenAI”.
4Define the owner of the outcome and the quality metric for the situation described in “Black box in the black box”.
Signals to track afterwards
→Watch for actions by Anthropic and Apple that confirm or challenge the episode’s central claims.
→Compare new launches and policy changes with “How o1-preview was trained”: have access, quality, price, or constraints changed?
→Check whether the scenario in “Black box in the black box” becomes repeatable practice rather than a one-off demonstration.
Most useful for
Executives and managersProduct teamsAI usersProfessionalsPeople planning their careersEntrepreneurs

Key takeaways

00:00Why context matters more than one metric: ToTheMoon is a podcast and a hello

The boundary of the “ToTheMoon is a podcast and a hello from Silicon Valley!” case is defined by this point: o1-preview exposes intermediate reasoning and handles tasks where one fast answer is not enough, but the quality gain brings a new chaos: users must choose among several models, and the black box gets deeper.

01:19How the issue moves from news to product: Google Wallet — an interesting novelty from Google

In the context of “Google Wallet — an interesting novelty from Google,” this criterion applies: Google Wallet now lets you add a passport to present the document electronically — first in the U.S., tied to the TSA system where your ID is already scanned at security: the document finally moves into the phone.

03:36The practical meaning of the issue: we are not Apple fanboys

In the context of “We are not Apple fanboys!,” this criterion applies: the hosts take the commenters' charge of Apple bias with a caveat: their phones really are all Apple — for style, security, or habit — but they follow Android too: Google's Gemini-linked Wallet ID feature was the first thing they covered.

05:13OpenAI's new model matters not because it answers a few percentage points better

The “New model from OpenAI! How does o1-preview work” issue should be assessed with one constraint in mind: the novelty changes the mode of operation itself — before answering, the system spends time reasoning and shows it "thought" for several seconds; for hard math and code that is a step forward, but users now must know when to pick fast 4o, when o1, and when mini.

06:31Until now, the difference between leading models was often visible only on specific tasks

The “The technical side of o1-preview. Why this model is unlike all the others” issue should be assessed with one constraint in mind: with o1 the market moves toward verticalization: one model suits everyday conversation, another code, a third expensive analysis; the interface grows more complex, and picking the wrong mode means overpaying or getting a weak result.

10:37Meta and other builders of open models face a separate dilemma

The working conclusion from “How o1-preview was trained” is that reinforcement learning always needs an auxiliary model checking the first model's results — and that built-in verification reduces hallucinations; the hosts expect a tool within a year that picks between 4o, o1, and mini automatically, because no one will do it by hand.

16:06The market tests it through use: B2C or B2B? Where AI solutions will go

The boundary of the “B2C or B2B? Where AI solutions will go” case is defined by this point: Meta and other open-model builders face their own dilemma — developers fine-tune the mass-market Llama for themselves, but a reasoning model demands entirely different economics and infrastructure: open and closed markets diverge no longer by license but by product type.

19:07The boundary between value and constraint: will ChatGPT cost $2,000 a month? For whom

The boundary of the “Will ChatGPT cost $2,000 a month? For whom?” case is defined by this point: the talk of plans costing hundreds or even thousands of dollars is a bet on professional scenarios: the price is justified only where the model's long reasoning replaces expensive human work rather than merely speeding up chat.

23:51At the same time, OpenAI is moving more visibly toward becoming a consumer company

The discussion of “ChatGPT competitors: why the discussion still centers on OpenAI” yields a practical test: Anthropic is strong as a developer platform, and Claude 3.5 Sonnet already beats ChatGPT on some real tasks, but OpenAI has the chance to build an Apple-grade brand: expensive, mass-market, and clear to people who have no wish to study architecture.

30:54The paradox of o1 is that the model is supposed to bring us closer to more intelligent AI while making that intelligence less transparent

The boundary of the “Black box in the black box” case is defined by this point: we see a polished description of the chain of thought but do not know how closely it mirrors the real internal process: one black box has learned to explain another — and trusting it has become harder, not easier.

What this episode is about

o1-preview exposes an intermediate reasoning process and handles tasks better when one fast answer is not enough. The quality gain brings a new kind of chaos: users must choose among several models, companies must decide between a mass product and an expensive professional tool, and researchers must admit that the black box has become even deeper.

OpenAI's new model matters not because it answers a few percentage points better. It changes the mode of operation itself: before responding, the system spends time reasoning and shows that it ‘thought’ for several seconds.

For difficult mathematics, programming, or hypothesis testing, that is an important step. But it creates a new problem for ordinary users—now they need to understand when to use fast ChatGPT-4o, when to use o1, when to choose mini, and why they have to choose at all.

Until now, the difference between leading models was often visible only on specific tasks. With o1, the market is moving toward specialization: one model may be better for everyday conversation, another for code, and a third for expensive analysis.

That makes the interface more complex and raises the cost of choosing incorrectly. A person either overpays or receives a weak result and concludes that AI as a whole does not work.

Meta and other builders of open models face a separate dilemma. A widely available Llama lets developers fine-tune the system for their own tasks, but a complex reasoning model may require completely different economics and infrastructure.

You cannot simply increase the parameter count and expect the same effect. Open and closed markets are therefore beginning to diverge not only by license, but also by product type.

At the same time, OpenAI is moving more visibly toward becoming a consumer company. Anthropic is strong as a developer platform, and Claude 3.5 Sonnet already beats ChatGPT on some real tasks, but OpenAI has a chance to build an Apple-like brand: expensive, mass-market, and understandable to people who do not want to study architecture. That is where talk of plans costing hundreds or even thousands of dollars for professional use cases comes from.

The paradox of o1 is that the model is supposed to bring us closer to more intelligent AI while making that intelligence less transparent. We see a polished description of its chain of thought, but we do not know how closely it reflects the real internal process. One black box has learned to explain another black box—and trusting it has become harder, not easier.

The model race is not won by the highest score alone, but by a system that solves the task repeatedly and makes its price and constraints clear.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 51 segments: 36 identified, 0 mixed, 9 probable, and 6 unresolved.

Read transcript on a separate page

Loading…