The US vs. China, Google vs. Microsoft, Anthropic vs. ChatGPT: Who Really Controls the AI Race
Who really controls the AI race: the owners of models, chips, infrastructure, or access to the user?
Map how control of the AI race is distributed across models, chips, talent, cloud infrastructure, and access to users—and assess companies across the whole chain rather than by one benchmark.
What to watch for
Key takeaways
The discussion of “The most expensive part of the new AI race is almost invisible to the average user” yields a practical test: behind the demos there is a contest over multimillion-dollar engineering salaries, plans to raise trillions for infrastructure, and the few people able to build such systems — so players should be judged by their control of talent, compute, and distribution, not by the polish of another chat.
The boundary of the “Digital avatars - what's going on in U.S. venture funds? Google’s work” case is defined by this point: at a Y Combinator demo day two companies already pitched through avatars, and the fund seriously debated whether a real person was speaking; Google's vlogger, meanwhile, is a Google Research demo rather than a product, while Synthesia and HeyGen, each with over $20M in annual revenue, show the demand is real.
For the “Fake voices — a godsend for fraudsters? Where this can lead. Part 1/3” scene, the decisive point is this: the same tool that narrates books becomes a ready-made machine for fraud and political manipulation, which is why OpenAI deliberately restricted access to its voice technology: as long as a fake is quick and cheap to make, society has not yet learned to tell it from a real person.
For the “Where voice avatars are used” scene, the decisive point is this: alongside fairy tales read in a parent's voice, support for people with speech impairments, and publications translated into any language stands a constraint: a comparable model can be trained on tens of thousands of outputs from a powerful one, so keeping Google's or OpenAI's technology closed does not by itself stop abuse.
The discussion of “The U.S.–China fight over AI: NVIDIA vs. Huawei” yields a practical test: NVIDIA designs the chips, TSMC manufactures them in Taiwan, China is building up Huawei Ascend, and Google, Intel, and Qualcomm are looking for ways to reduce dependence on a single supplier — at this layer an export measure or a manufacturing constraint can matter more than another algorithmic improvement.
The discussion of “The emotional side of AI: how soon artificial intelligence will stop being “artificial”” yields a practical test: voice already supports working scenarios — audiobooks in a favorite actor's voice, earbuds with real-time translation — yet the hosts see conveying emotion and attention, what matters most to children and loved ones, as still out of reach and possibly utopian.
The “Y Combinator and start-up trends” scene leads to a working conclusion: the accelerator is producing more and more AI companies, but putting a fashionable term in a pitch deck does not create a market. Devin, for example, has revived the debate over the end of the programming profession. Yet the product's real value is determined not by the demo, but by how reliably it performs the work and how much control a person still needs to retain.
The “Is Anthropic cooler than ChatGPT?” scene leads to a working conclusion: Anthropic has a window of three to six months before a likely GPT-5 release, and Amazon confirmed its bet by exercising its option and adding money: a cloud leader without a strong model of its own needs this partnership strategically.
What this episode is about
The AI race does not begin with a polished chatbot. It begins with people, chips, money, and access to infrastructure. While OpenAI and Google demonstrate voice avatars, NVIDIA and Huawei are fighting over the market's computing foundation, and Y Combinator is testing which ideas can actually become businesses.
The most expensive part of the new AI race is almost invisible to the average user. We see a voice avatar, a new chat interface, or another demo. Behind it are multimillion-dollar engineering salaries, plans to raise trillions for infrastructure, and a fight for the small number of people who can build these systems at all.
The question is no longer which model gave the most elegant answer today. The question is who controls the talent, the compute, and the channels through which the product reaches users.
Voice avatars reveal the market's dual nature particularly well. On one side, HeyGen, Synthesia, and Google's work open up practical uses: dubbing, training, personalized messages, and children's books read in a parent's voice.
On the other, the same tool becomes a ready-made machine for fraud and political manipulation. OpenAI did not restrict access to voice technology for no reason: when a convincing fake can be produced quickly and cheaply, society still has to learn how to distinguish it from a real person.
The conflict is even sharper at the infrastructure layer. NVIDIA designs the chips, TSMC manufactures them in Taiwan, China is developing Huawei Ascend, and Google, Intel, and Qualcomm are looking for ways to reduce dependence on a single supplier. An export measure or manufacturing constraint can matter more here than another algorithmic improvement.
A country without access to modern chips simply cannot develop models at the same speed.
At the same time, Y Combinator shows where entrepreneurial energy is moving. The accelerator is producing more and more AI companies, but putting a fashionable term in a pitch deck does not create a market.
Devin, for example, has revived the debate over the end of the programming profession. Yet the product's real value is determined not by the demo, but by how reliably it performs the work and how much control a person still needs to retain.
The larger picture is straightforward: OpenAI, Google, Anthropic, Microsoft, and Chinese players are competing across several layers at once—models, chips, cloud infrastructure, data, products, and talent. The winner will not be the company that tops a benchmark once. It will be the one that connects all of these layers into a working system and makes it available to ordinary people.
One first-place benchmark result is not enough. The advantage goes to the player that connects models, compute, data, product, and distribution into a working system.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 47 segments: 37 identified, 6 mixed, 2 probable, and 2 unresolved.
Read transcript on a separate page
Loading…