Will the Best AI Models Become a Luxury? Tokens, Limits, and the Price of Access
What does it really cost to work with the strongest AI models, and will access to them become a privilege for those who can afford enormous compute consumption?
Understand why tokens now determine not only technical limits but also the real cost of work, model choice, and the future gap between users with different levels of access.
What to watch for
Key takeaways
A personal Codex counter reveals the scale of compute hidden by an ordinary subscription. As long as consumption appears as one large number, the user cannot see the economics of the work.
When the same volume is translated into public API prices, a fixed subscription begins to look like heavily subsidized access to an expensive resource.
GPT-5.6 Sol consumed billions of tokens over several days, yet some work stopped or required new runs. Compute consumption and useful output are different measures.
Real development depends on stability, limits, speed, and the ability to finish a task. A more capable mode can be a worse working tool if it disrupts the process.
Kimi K3 and other Chinese systems create pressure not only through quality. They show that a useful result may cost far less and force the market to reconsider compute budgets.
Usage may be hidden from the user by a subscription, but it remains an infrastructure cost for the platform. Limits and pricing emerge at the boundary between those two economies.
If the highest-end models require too much compute, platforms will separate modes and users. The divide will not only be between people who use AI and those who do not, but also between levels of access.
What this episode is about
Tokens are no longer a technical abstraction. They are becoming money, limits, and the price of access to the strongest models. A personal Codex usage counter reveals the scale of consumption that an ordinary subscription hides from the user.
Forty Billion Tokens Looks Like Just a Number
A usage counter appeared in my Codex profile, and it showed roughly forty billion tokens. For an ordinary person, that figure means almost nothing: we do not think in tokens, and we do not see the amount of computation behind one work session. But once the same volume is priced at API rates, an ordinary subscription begins to look like access to a resource worth tens or hundreds of thousands of dollars.
A Model Can Consume Billions of Tokens Without Producing the Result
Over three days of work, GPT-5.6 Sol used roughly fifteen billion tokens. Yet heavy consumption did not guarantee completion of the task. The model could reason for a long time, use successive windows, and require new runs. For a user on a fixed subscription, this feels like a limit and lost time. For the provider, it represents real infrastructure cost. The new economics of access emerges between those two perspectives.
Price Determines Which Model Becomes the Working Model
OpenAI and Anthropic compete on more than answer quality. Available modes, limits, speed, and the number of attempts a user can make before being stopped all matter. At the other end of the market, less expensive Chinese models such as Kimi K3 are appearing. They may be weaker in some tasks, but they change the expected price. If a strong result can be obtained more cheaply, users and companies begin to allocate subscriptions and computing budgets differently.
The Best AI May Not Be Available to Everyone
The most important conclusion is not about one subscription bill, but about the future of access. If a powerful model requires enormous computation, the provider has to impose limits, divide plans, and decide who receives the most expensive mode. The gap may therefore run not only between people who use AI and those who do not, but also between those who can access maximum-capability models and those limited to cheaper versions. Tokens are becoming part of how opportunity is distributed.
Tokens are becoming a new unit of access. They reveal what intelligence as a service actually costs and why the choice of a working model will increasingly depend on economics, not only answer quality.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 30 segments: 30 identified, 0 mixed, 0 marked with ✓, and 0 unresolved.
Loading…