Skip to content
OpenAI · OpenAI o1 · AppleEpisode 024 · 22 September 2024 · 32:47

OpenAI Taught a Model to Think Longer—but Made Choosing an AI Even Harder

What to watch for

1Compare “New model from OpenAI! How does o1-preview work” with “Technical side o1-preview. Why is this model not all the others”: they provide different criteria for judging the same issue.
2Test the conclusion from “How did o1-preview learn” in your own use case—what actually changes in the process and what remains a promise.
3Before choosing a product or approach, record the constraint identified in “ChatGPT competitors: why the discussion still centers on OpenAI”.
4Define the owner of the outcome and the quality metric for the situation described in “Black box in the black box”.
Signals to track afterwards
Watch for actions by Anthropic and Apple that confirm or challenge the episode’s central claims.
Compare new launches and policy changes with “How did o1-preview learn”: have access, quality, price, or constraints changed?
Check whether the scenario in “Black box in the black box” becomes repeatable practice rather than a one-off demonstration.
Most useful for
Executives and managersProduct teamsAI usersProfessionalsPeople planning their careersEntrepreneurs

Key takeaways

00:00Why context matters more than one metric: ToTheMoon is a podcast and a hello

The boundary of the “ToTheMoon is a podcast and a hello from the Silicon Valley!” case is defined by this point: a launch matters only when it changes access, quality, price, or user behavior in a real workflow.

01:19How the issue moves from news to product: google Wallet is an interesting novel from Google

In the context of “Google Wallet is an interesting novel from Google,” this criterion applies: a benchmark measures a narrow capability; working value requires repeatability, a clear cost, and control over errors.

03:36The practical meaning of the issue: we're not apple choms

In the context of “We're not apple choms!,” this criterion applies: the risk depends on the scope of access, the scale of the consequences, and whether the system can be stopped and its actions reconstructed.

05:13OpenAI's new model matters not because it answers a few percentage points better

The “New model from OpenAI! How does o1-preview work” issue should be assessed with one constraint in mind: this section clarifies the mechanism behind the topic and preserves a constraint that would otherwise be lost in an overly simple conclusion.

06:31Until now, the difference between leading models was often visible only on specific tasks

The “Technical side o1-preview. Why is this model not all the others” issue should be assessed with one constraint in mind: this section clarifies the mechanism behind the topic and preserves a constraint that would otherwise be lost in an overly simple conclusion.

10:37Meta and other builders of open models face a separate dilemma

The working conclusion from “How did o1-preview learn” is that this section clarifies the mechanism behind the topic and preserves a constraint that would otherwise be lost in an overly simple conclusion.

16:06The market tests it through use: b2C or B2B? Where the AI-Decisions go

The boundary of the “B2C or B2B? Where the AI-Decisions go” case is defined by this point: the conflict reveals which rights, money, and control points the parties consider strategic.

19:07The boundary between value and constraint: chatGPT cost $2,000 a month? For who

The boundary of the “ChatGPT cost $2,000 a month? For who?” case is defined by this point: a launch matters only when it changes access, quality, price, or user behavior in a real workflow.

23:51At the same time, OpenAI is moving more visibly toward becoming a consumer company

The discussion of “ChatGPT competitors: why the discussion still centers on OpenAI” yields a practical test: this section clarifies the mechanism behind the topic and preserves a constraint that would otherwise be lost in an overly simple conclusion.

30:54The paradox of o1 is that the model is supposed to bring us closer to more intelligent AI while making that intelligence less transparent

The boundary of the “Black box in the black box” case is defined by this point: this section clarifies the mechanism behind the topic and preserves a constraint that would otherwise be lost in an overly simple conclusion.

What this episode is about

o1-preview exposes an intermediate reasoning process and handles tasks better when one fast answer is not enough. The quality gain brings a new kind of chaos: users must choose among several models, companies must decide between a mass product and an expensive professional tool, and researchers must admit that the black box has become even deeper.

OpenAI's new model matters not because it answers a few percentage points better. It changes the mode of operation itself: before responding, the system spends time reasoning and shows that it ‘thought’ for several seconds.

For difficult mathematics, programming, or hypothesis testing, that is an important step. But it creates a new problem for ordinary users—now they need to understand when to use fast ChatGPT-4o, when to use o1, when to choose mini, and why they have to choose at all.

Until now, the difference between leading models was often visible only on specific tasks. With o1, the market is moving toward specialization: one model may be better for everyday conversation, another for code, and a third for expensive analysis.

That makes the interface more complex and raises the cost of choosing incorrectly. A person either overpays or receives a weak result and concludes that AI as a whole does not work.

Meta and other builders of open models face a separate dilemma. A widely available Llama lets developers fine-tune the system for their own tasks, but a complex reasoning model may require completely different economics and infrastructure.

You cannot simply increase the parameter count and expect the same effect. Open and closed markets are therefore beginning to diverge not only by license, but also by product type.

At the same time, OpenAI is moving more visibly toward becoming a consumer company. Anthropic is strong as a developer platform, and Claude 3.5 Sonnet already beats ChatGPT on some real tasks, but OpenAI has a chance to build an Apple-like brand: expensive, mass-market, and understandable to people who do not want to study architecture. That is where talk of plans costing hundreds or even thousands of dollars for professional use cases comes from.

The paradox of o1 is that the model is supposed to bring us closer to more intelligent AI while making that intelligence less transparent. We see a polished description of its chain of thought, but we do not know how closely it reflects the real internal process. One black box has learned to explain another black box—and trusting it has become harder, not easier.

The model race is not won by the highest score alone, but by a system that solves the task repeatedly and makes its price and constraints clear.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 51 segments: 36 identified, 0 mixed, 9 marked with ✓, and 6 unresolved.

Loading…