The New York Times Lawsuit Shows How Much Data AI Retains—and How Little Control the User Has
What does The New York Times lawsuit reveal about AI data retention and how little control users have over it?
Understand which data remains inside AI services, how litigation can expose it, and which retention settings users need to check in advance.
What to watch for
Key takeaways
The “The New York Times Lawsuit Shows How Much Data AI Retains—and How Little Control the User” topic becomes clearer once this point is included: the publisher's dispute with OpenAI exposes not only copyright but which prompts, answers, and deleted conversations are retained, who can access them, and whether a company can actually honor a user's request to “delete” data.
The working conclusion from “Why are the ChatGPT logs retained?” is that a user clicks “delete chat” expecting the conversation to be gone, but because of legal demands, internal policies, or technical architecture the data may keep being stored — an interface button and the actual destruction of information are not the same.
The decision in “What if your ChatGPT requests become public?” depends on one criterion: prompts sent to a model are often far more sensitive than an ordinary search — people upload contracts, medical documents, work correspondence, and personal problems, and the more useful ChatGPT is, the more context it receives, which makes that data all the more valuable for training, security, and legal demands.
In the context of “Google data: how search helped solve a crime,” this criterion applies: in the US a cold case was cracked by matching who had searched Google for one specific location, and a court ruled such data admissible — a sign that the query logs of the apps we use daily really can surface in an investigation.
In the context of “Card fraud: how can banks solve the problem?,” this criterion applies: in the US and some countries the risk of fraudulent charges is covered by chargeback rules — the bank refunds a transaction you did not make and sorts it out afterward; so how much you worry about a card depends not on the technology but on who is made responsible for the loss.
For the “AI and data security: Country differences” scene, the decisive point is this: access rules, retention, and the availability of the newest models differ by country, and not just a state border but an information border is forming — one user gets a new tool and guarantees, another works through restrictions, a VPN, or services with a different policy.
The working conclusion from “GPT-5: what to expect in 2025?” is that GPT-5 is promised as early as this summer, though whether it ships is unclear; Polymarket is telling too, where bets piled onto Gemini while OpenAI's rating stayed low — a forecast is worth testing against concrete dates and actual launches, not announcements.
The discussion of “MIT study: how does AI affect the brain?” yields a practical test: in a small MIT study (just 54 participants) those who wrote essays themselves showed more active brains on EEG and answered questions about their own text better, while those who wrote via an LLM engaged less — the sample is tiny, but the direction is telling.
The “AI and memory: problems of concentration loss” scene leads to a working conclusion: access rules, retention, and the availability of the newest models differ by country. Not only a state border but an information border is forming: one user receives a new tool and a particular set of guarantees, while another has to work through restrictions, a VPN, or services governed by a different policy.
The “Wi-Fi vision: how AI sees through walls” scene leads to a working conclusion: Carnegie Mellon researchers paired a Wi-Fi router with a camera and trained a model to recognize a person's silhouette from the radio signal, turning an ordinary router into a camera-like sensor that “sees” through walls — impressive and unsettling, though how far it scales depends on whether the hardware needs per-case calibration.
What this episode is about
The publisher’s dispute with OpenAI is not only about copyright. It exposes a more uncomfortable question: which prompts, answers, and deleted conversations are retained, who can access them, and whether a company can actually honor a user’s request to “delete” data.
A user clicks “delete chat” and expects the conversation to disappear. The New York Times litigation with OpenAI shows that data may in fact continue to be stored because of legal requirements, internal policies, or technical architecture. For the individual, that means one simple thing: an interface button and the actual destruction of information are not always the same.
Prompts sent to a model are often much more sensitive than an ordinary search. People upload contracts, medical documents, work correspondence, product ideas, and personal problems. The more useful ChatGPT becomes, the more context it receives. At the same time, that data becomes more valuable for training, security, legal demands, and investigations.
The lawsuit also exposes the conflict between content owners and model developers. The publisher wants to know whether its material was used and whether the system can reproduce it. OpenAI is defending both its technology and user data. Yet both sides work with a mass of information that the ordinary person barely sees and cannot independently audit.
Geography is a separate problem. Access rules, retention, and the availability of the newest models differ by country. Not only a state border but an information border is forming: one user receives a new tool and a particular set of guarantees, while another has to work through restrictions, a VPN, or services governed by a different policy.
The practical conclusion is not to stop using AI. It is to separate safe context from data whose exposure would cause real harm.
Before uploading a document, remove unnecessary personal information, check the retention mode, and understand whether an enterprise version is being used. A model can be an excellent assistant, but it should not become the only place where the most sensitive information lives.
Before uploading a document, remove unnecessary personal information, check the retention mode, and understand whether an enterprise version is being used. As a result, a model can be an excellent assistant, but it should not become the only place where the most sensitive information lives.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 86 segments: 53 identified, 3 mixed, 12 probable, and 18 unresolved.
Loading…