Skip to content
GPT-5.6 Sol · Codex · TokenizationEpisode 139 · 24 July 2026 · 20:56

Will the Best AI Models Become a Luxury? Tokens, Limits, and the Price of Access

Central question

What does it really cost to work with the strongest AI models, and will access to them become a privilege for those who can afford enormous compute consumption?

What you take away

Understand why tokens now determine not only technical limits but also the real cost of work, model choice, and the future gap between users with different levels of access.

Main threads

What to watch for

1Count not only the subscription price but also the number of attempts, time to completion, and work the model failed to finish.
2Compare models by the cost of a completed task rather than the price of one token or a headline benchmark position.
3Separate temporarily subsidized subscription access from the real compute cost that a platform may later pass on to the user.
4Maintain several working model tiers: an expensive mode for critical tasks and lower-cost systems where their quality is sufficient.
Signals to track afterwards
Changes in limits, pricing, and availability of the highest-end OpenAI and Anthropic modes.
Falling prices of Chinese models and how much that changes the choices made by companies and developers.
The emergence of separate plans for the most compute-intensive models and long-running agent tasks.
The gap between the public subscription price and the actual cost of the compute consumed.
Most useful for
developers and users of Codex and Claudecompanies planning large AI workloadsfounders of AI productsfinance and operations leaderspeople choosing between premium and lower-cost models

Key takeaways

03:30Forty Billion Tokens Stop Being an Abstraction

A personal Codex counter reveals the scale of compute hidden by an ordinary subscription. As long as consumption appears as one large number, the user cannot see the economics of the work.

05:54API Pricing Changes How a Subscription Looks

When the same volume is translated into public API prices, a fixed subscription begins to look like heavily subsidized access to an expensive resource.

06:33Enormous Consumption Does Not Guarantee a Result

GPT-5.6 Sol consumed billions of tokens over several days, yet some work stopped or required new runs. Compute consumption and useful output are different measures.

08:43A Working Model Is Defined by More Than Intelligence

Real development depends on stability, limits, speed, and the ability to finish a task. A more capable mode can be a worse working tool if it disrupts the process.

10:49Lower-Cost Models Change Price Expectations

Kimi K3 and other Chinese systems create pressure not only through quality. They show that a useful result may cost far less and force the market to reconsider compute budgets.

12:54Every Token Is Ultimately Paid For

Usage may be hidden from the user by a subscription, but it remains an infrastructure cost for the platform. Limits and pricing emerge at the boundary between those two economies.

14:30Access to the Strongest AI May Become Unequal

If the highest-end models require too much compute, platforms will separate modes and users. The divide will not only be between people who use AI and those who do not, but also between levels of access.

What this episode is about

Tokens are no longer a technical abstraction. They are becoming money, limits, and the price of access to the strongest models. A personal Codex usage counter reveals the scale of consumption that an ordinary subscription hides from the user.

Forty Billion Tokens Looks Like Just a Number
A usage counter appeared in my Codex profile, and it showed roughly forty billion tokens. For an ordinary person, that figure means almost nothing: we do not think in tokens, and we do not see the amount of computation behind one work session. But once the same volume is priced at API rates, an ordinary subscription begins to look like access to a resource worth tens or hundreds of thousands of dollars.

A Model Can Consume Billions of Tokens Without Producing the Result
Over three days of work, GPT-5.6 Sol used roughly fifteen billion tokens. Yet heavy consumption did not guarantee completion of the task. The model could reason for a long time, use successive windows, and require new runs. For a user on a fixed subscription, this feels like a limit and lost time. For the provider, it represents real infrastructure cost. The new economics of access emerges between those two perspectives.

Price Determines Which Model Becomes the Working Model
OpenAI and Anthropic compete on more than answer quality. Available modes, limits, speed, and the number of attempts a user can make before being stopped all matter. At the other end of the market, less expensive Chinese models such as Kimi K3 are appearing. They may be weaker in some tasks, but they change the expected price. If a strong result can be obtained more cheaply, users and companies begin to allocate subscriptions and computing budgets differently.

The Best AI May Not Be Available to Everyone
The most important conclusion is not about one subscription bill, but about the future of access. If a powerful model requires enormous computation, the provider has to impose limits, divide plans, and decide who receives the most expensive mode. The gap may therefore run not only between people who use AI and those who do not, but also between those who can access maximum-capability models and those limited to cheaper versions. Tokens are becoming part of how opportunity is distributed.

Tokens are becoming a new unit of access. They reveal what intelligence as a service actually costs and why the choice of a working model will increasingly depend on economics, not only answer quality.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 30 segments: 30 identified, 0 mixed, 0 marked with ✓, and 0 unresolved.

Loading…