After DeepSeek, the Market Started Asking Not Where a Model Comes From, but Where Its Output Can Be Trusted
After DeepSeek, where can a model's output be trusted, and why is its origin no longer a sufficient answer?
Separate the investment signal and the impressive demonstration from the real business in the case of DeepSeek and OpenAI. The working test is to check who pays, which indispensable part of the chain the product controls, and whether the economics survive scale.
What to watch for
Key takeaways
The “What is Deep Research from OpenAI: a model that “asks questions”” issue should be assessed with one constraint in mind: Deep Research is positioned as an analytics tool, but the key point is that the model behaves “like a human” — it searches, analyzes, and asks clarifying questions itself; on the cited test 4o scores 3.3% accuracy, Deep Research 26.6%, and DeepSeek R1 9.4%.
In the context of “Donald Trump on DeepSeek,” this criterion applies: Trump mentioned the Chinese model's release on a briefing the very same day — a rare case — and called it a “wake-up call” for American industry: if the same result is achievable for less money, competition must get sharper.
The “DeepSeek's role in shifting OpenAI's position and its popularity surge” topic becomes clearer once this point is included: DeepSeek's popularity showed how quickly users switch to a free and sufficiently strong tool, but the downloads brought two uncomfortable questions — was the model trained on OpenAI outputs, and where does the data of the Chinese service's users go.
The boundary of the “Changes in copyright law and their impact on AI content” case is defined by this point: the dispute is almost mirror-symmetric — American companies trained for years on huge swaths of the internet and now object to distillation of their outputs; a possible violation is not erased, but it shows how unstable the rules are when everyone protects their own and uses someone else's.
For the “DeepSeek and the question of large-scale compute: is powerful hardware needed?” scene, the decisive point is this: serving at online scale still requires chips and infrastructure — the host never once managed to process his usual large file in DeepSeek; and NVIDIA's stagnant stock is tied not only to DeepSeek but to the discussion of harsh export restrictions for China.
The boundary of the “The pros of DeepSeek” case is defined by this point: DeepSeek is open where ChatGPT was unreachable without a VPN, costs fifteen to thirty times less, includes reasoning out of the box, and keeps updating — it has already added its own image generator.
The “Open source in banks and major corporations: experiments on Hugging Face” issue should be assessed with one constraint in mind: they are looking at Hugging Face and open models that can be deployed inside a protected environment. Even a somewhat weaker system may be preferable if data does not leave for an outside provider.
The “OpenAI Workspace and data leaks” issue should be assessed with one constraint in mind: in OpenAI's corporate workspace, data reportedly is not used for training; for large corporations the constraint is real, but the host forbids businesses nothing — ordinary working data is not unique, and personal data is easy to replace with codes before upload.
The “HLE (Humanity's Last Exam): how AI models' knowledge is tested” issue should be assessed with one constraint in mind: Humanity's Last Exam raises the bar and tests knowledge missing from standard benchmarks, but one number cannot tell a lawyer, physician, or accountant how useful the model is — the user still has to choose a mode and verify the result.
The decision in “The shortcomings of new AI models and their future” depends on one criterion: the next stage is not a contest for one first place — there will be public models, local systems, and specialized products such as legal databases, and the winner will be the model whose data, license, price, and errors are understandable to a specific organization.
What this episode is about
DeepSeek surged in the App Store, OpenAI said it may have been trained on its outputs, and companies began arguing over copyright and exports. At the same time, Humanity’s Last Exam, o3-mini, and Deep Research appeared. The race is moving from polished answers to tested knowledge, data-handling models, and the ability to deploy AI inside a company.
DeepSeek’s popularity showed how quickly users will move to a new tool when it is free and strong enough. But the surge in downloads brought two uncomfortable questions with it: was the model trained on OpenAI outputs, and where does the data of people using the Chinese service go?
The copyright dispute is becoming almost perfectly symmetrical. American companies trained for years on enormous portions of the internet, and now object when a competitor may have used their outputs for distillation. That does not erase a possible violation, but it shows how unstable the rules are when every participant is simultaneously protecting its own model and using someone else’s content.
For banks and large corporations, a web interface is not the main option at all. They are looking at Hugging Face and open models that can be deployed inside a protected environment. Even a somewhat weaker system may be preferable if data does not leave for an outside provider.
Humanity’s Last Exam tries to raise the bar by testing knowledge that standard benchmarks do not cover. But one number cannot tell a lawyer, physician, or accountant how useful a model will be. o3-mini and Deep Research represent different trade-offs among speed, reasoning, search, and cost. The user still has to choose a mode and verify the result.
The next stage of the market will not be a contest for one first place. There will be public models, local systems, and specialized products such as legal databases.
The winner will not necessarily be the smartest model, but the one whose data, license, price, and errors are understandable to a specific organization. After DeepSeek, “Who is better?” becomes “Who can be trusted, and in which environment?”
After DeepSeek, “Who is better?” turns into “Who can be trusted, and in which environment?”.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 90 segments: 50 identified, 2 mixed, 31 probable, and 7 unresolved.
Loading…