Imagine that you wrote a text yourself and a system tells you it was written by artificial intelligence. Or the opposite: the entire piece was produced by AI, but determining that has become almost impossible. We are approaching a world in which it is increasingly difficult to tell whether we are looking at the work of a person or a machine. AI detectors, watermarks, and systems that record how a document was created already exist, yet different checks can rate the same text in completely different ways. So can we reliably determine where the human ends and the AI begins? And what happens when the distinction becomes impossible? That is what we are going to examine. We have already discussed cases in which models helped people deceive companies and other people—for example, by generating altered images or fake receipts that employees then submitted as if they had actually made the purchases. Entire investigations followed, and in one large sample roughly twelve percent of the receipts were said to be fake. We also talked about customs officers asking travelers to prove that an expensive watch had been worn before the trip. A person can now generate an old-looking photo or video with the watch. I joked
with Tanya that I could make a clip showing Stalin personally giving the watch to my great-grandfather. Some customs officers may still believe that kind of evidence, at least in some countries, although probably not for long in the United States or China. The same applies to fabricated events. Against that
background, Anthropic announced that models released from August 2 onward would embed an invisible, machine-readable watermark in generated text. The mark is placed inside the text itself, should travel with ordinary copy and paste, and may survive light editing. Anthropic says the system will apply worldwide, including in the United States. We are therefore continuing the ToTheMoon discussion about legitimate and illegitimate uses of AI and about the problems created by systems that try to diagnose AI involvement. We will also look at what OpenAI and Gemini encode in their outputs and at the tools already used to identify possible violations.
I would divide the problem into two parts. In some settings, independent verification is genuinely necessary. In others, the same checks can create additional harm for an innocent person. The first question is therefore not “Which detector should we buy?” but “Where does authorship actually matter enough to justify verification?” Schools and universities are an obvious example. An institution may need to know whether a student mastered the material, took the exam, or wrote the essay personally. I will show the main verification approaches, because I have serious questions about several tools that are treated as reliable. We previously discussed a student who was accused of using AI even though he said he had a complete Google Docs history showing that he wrote the paper himself.
Admissions create a similar issue: is the applicant completing the work, or is an AI system doing it for them? Beyond education, the next major area is hiring and job interviews.
Hiring is an especially difficult case. Over more than twenty years I have hired well over a thousand people, and the gap between what someone says in an interview, what they produce in a test assignment, and how they actually work has always been a problem. AI turns that into a much larger problem. Some companies explicitly allow candidates to use AI during an interview or a take-home task; for them, the real test is how the person frames the problem and uses the tool. In other interviews, however, a candidate is supposed to answer unaided. Entire classes of products can now feed a person polished AI-generated answers without the interviewer noticing. I suspect this kind of deception is already happening millions of times, because far more people know how to use these systems than know how to identify their use correctly.
Another important area is science, grants, publications, journalism, and other public-interest content. Science has always had a recognition problem: a result may be celebrated, rewarded, and treated as settled, only to be overturned ten or twenty years later. AI is already being used to challenge claims that had been accepted as proven. At the same time, it gives people new ways to manipulate data, sources, experiments, quotations, and even the existence of a study. In journalism, provenance matters for the same reason.
Did the source exist? Was the quotation preserved in context? Was a beautiful passage written by the reporter, fabricated by a model, or assembled from something else? I am not listing these areas to create a dull catalogue. I want us to understand where verification has a real purpose and where it is merely surveillance dressed up as certainty. I read a story in a Russian media outlet—RBC—and wondered whether the framing had shifted the meaning of the source. The familiar technique is to quote one fragment of what someone said and then surround it with a description that pushes the reader toward a particular interpretation. I asked ChatGPT to cross-check the material, and it concluded that the information had indeed been presented from a particular angle. That ability to verify claims can be extremely useful. On a flight from Los Angeles, my wife showed me a young woman online who claimed to attend a particular university and program. I suggested running a serious check in the paid version of ChatGPT. It searched her public social profiles, university pages, and published records and assembled the available evidence. ChatGPT 5.6 Sol handled that verification in a reasoned way. But even here we need proportion. If I say that
I got my first 286-class computer in 1992 rather than 1994, I may simply not remember the exact year. That is not the same thing as deliberately fabricating a credential. My mother probably would not remember the exact year either. Whether the computer appeared when I was nine or eleven changes nothing about the point of the story. A useful verification system should be able to say that the discrepancy is immaterial. The situation is different when someone uses a false childhood story, invented employment, or a fabricated achievement as a marketing claim. Verification needs context: what was claimed, why it matters, whether the speaker intended to mislead, and what consequence follows. A detector that treats every mismatch as equally important does not provide analysis; it simply manufactures suspicion.
The same problem appears in legal and financial documents, where facts, professional responsibility, and the actions of particular people may carry direct consequences.
Did the person act properly? Did they use AI to prepare the document? And does that use matter? May my lawyer use AI to draft a contract? If the licensed lawyer reviews the work, accepts responsibility, and delivers the result I need, why should the tool itself be a problem? I recently helped someone research laws with AI and sent them passages produced by the system. Strangely, I felt as though I had contributed less than if I had typed the same work manually, even though the value of the result did not necessarily change. Medical records and doctors’ conclusions raise the same question at a much higher level: was AI used, how was it used, who reviewed the output, and who remains accountable? Similar issues arise in competitions, statistics, commissioned reviews, spam, information operations, and PR campaigns. These are all settings in which verification may be important—but the object of verification is not merely whether a model touched the text.
Claude has begun embedding a signal in the content it generates, but there are important qualifications. Not every Claude response is necessarily marked. Anthropic committed to native support for models launched on or after August 2, a date tied to new European transparency requirements, although the marking is intended to operate globally. We often treat familiar stylistic habits—an em dash, a particular quotation mark, a certain list format—as proof that AI wrote a text. Sam Altman even said OpenAI had addressed the long-dash problem, yet ChatGPT still produces long dashes. Is that evidence of AI authorship, or simply a punctuation choice that humans also make? That is the larger problem with visual stereotypes. Anthropic is also working to extend the feature to earlier models, but the full model list has not been published. A newer model such as Opus 5 or Fable may fall on one side of the date, while Opus 4.8 clearly falls on the other. At the moment, however, there is no generally available public tool where anyone can paste a text and receive an official yes-or-no answer about Anthropic’s watermark.
Anthropic says it is preparing a detection mechanism and will publish technical documentation later. For now, there is no official website that turns a pasted passage into a definitive result. The company says ordinary copying should preserve the watermark, but heavy rewriting, translation, transformation, or mixing with other material may weaken or remove it. A short excerpt may not contain enough signal to test. Detection also would not prove that Claude wrote a document from scratch. A person may have written the text and used Claude only for editing, translation, or style. The final version could still carry the mark. Anthropic has not disclosed the exact technical implementation, so it is incorrect to describe the watermark as a set of hidden characters. The public claim is narrower: a signal is introduced during model generation. This limitation becomes especially important when we look at a university case in which a student was punished on the basis of a detector.
A university concluded that a student had used AI to write an assignment. The student replied that he had written it himself and could show the complete Google Docs history. In Newby v. Adelphi University, a professor relied on Turnitin, which reportedly classified the paper as one hundred percent AI-generated. The consequences of this kind of accusation can be severe: students can lose scholarships, awards, educational support, or even their place at a university. There are also real cases in which students who performed well online failed when asked to repeat the work in person, so institutions do have a legitimate problem to solve. But in the Adelphi case, the student denied using AI, and two additional checks commissioned by the family reportedly classified the work as human-written. The university nevertheless upheld the violation. In January, a court set aside the disciplinary decision, ordered the record removed, and rescinded the sanctions. It is important not to overstate the ruling.
The court did not conduct a scientific experiment proving mathematically that a human wrote the paper. The court’s narrower point was more important: one contradictory detector result, combined with a flawed internal process, was not enough to justify punishment. I have seen the same uncertainty personally. When my nephew was applying to universities, he checked an essay that he had written entirely himself. One detector still said there was a high probability that AI had written it. This is why people keep mixing together technologies that answer very different questions.
The first technology is a watermark inserted by the model provider—Claude, Gemini, ChatGPT, or another system—during generation. The second is an external detector that receives a finished text and guesses its origin from linguistic patterns. A provider watermark may support the statement that a passage probably passed through a particular system, assuming the verifier knows the signal. A statistical detector says only that the passage resembles texts commonly produced by AI. Neither result proves that a person cheated, failed to do the work, or broke a rule.
I often write an email myself and then ask a model to edit or format it. The underlying thought is mine even if the final text passes through an AI system. Anthropic also acknowledges that a watermark can disappear under several common conditions.
The signal may be lost if a passage is heavily rewritten, paraphrased, translated into another language, mixed with a large amount of other text, reduced to a short fragment, or generated by an older model that does not support marking. The opposite problem also exists: a human may write an article and send it to Claude for translation, after which the human-authored work could carry Claude’s mark. That is why we need to understand what a test actually measures. In practice, there are four different ways to investigate the origin of content. Before listing them, consider the policy question running underneath the entire episode: how many errors are acceptable when a system can damage a real person’s education, career, or reputation?
Imagine one hundred students taking an exam. Modern AI systems make mistakes; that is normal behavior for probabilistic models, just as people make mistakes. Suppose thirty students violate the rules and a detector finds most of them, but the system also has a ten-percent error rate. It will miss some people who cheated and accuse some people who did the work themselves. Is it acceptable for ten innocent students to face a disciplinary problem because the overall detection rate looks useful? We discussed a similar trade-off in a recent episode about why McDonald’s could not accept an AI ordering system at roughly eighty-percent quality and needed something closer to ninety-five percent. The acceptable error rate depends on the consequence. A recommendation error is not the same as a disciplinary sanction. Yet refusing to use any tool also makes verification much harder.
One possible answer is to use automated checks only as a starting signal and require a human review afterward. But even that introduces bias: once a system tells the reviewer that a person is suspicious, how independent will the human judgment remain? With that warning in mind, the first of the four approaches is a provider-embedded watermark. Its advantage is that the provider inserts a known signal, so detection can be more precise than guessing from style. Its limitation is that it works only for supported models, may weaken when the text is transformed, and does not by itself explain where the original ideas came from. The second approach is cryptographic provenance for files. Standards such as C2PA can record where a file was created or processed, whether it changed after signing, and which supported tool handled it. Claude, for example, can use this kind of signed provenance for certain generated files and images.
OpenAI also supports provenance for certain images and audio. But file provenance has limits. If I write an entire text myself and ask OpenAI only to package it as a PDF, the resulting file may show AI involvement even though the authorship of the ideas is human. The third approach is an external statistical detector. Turnitin, GPTZero, Copyleaks, Originality.ai, and other services analyze a completed text and search for combinations of features associated with language-model output. I do not personally rely on these tools. A specialized detector may be trained for a narrow domain, but for broad verification I would rather use a strong reasoning model to inspect claims, evidence, and inconsistencies than accept a single opaque percentage. I do not believe stand-alone generic AI detectors are the long-term answer. The more important question is whether any form of detection will remain useful as models continue to improve.
Statistical detectors do not find a hidden ChatGPT trace in every document. They classify patterns, which is why different services can give opposite answers for the same text—and even the same service may change its answer after a small edit. The fourth approach records the creation process itself. A system can track what a person typed, what they pasted, which versions they created, where AI was used, which prompts were sent, and how the output was edited. Tools such as Grammarly Authorship can work inside Google Docs or Microsoft Word to show the provenance of document fragments. That is the kind of evidence the Adelphi student pointed to. Instead of guessing after the work is finished, the system preserves a history while the work is being produced. For universities, interviews, and legal disputes, this can be more informative than a generic detector. But even process evidence can be staged.
A person can place AI-generated text on a second monitor and slowly retype it into the tracked document over several months. They can delete words, re-enter them, and even record themselves at the keyboard. The process log will look human even though the intellectual work came from a model. This is simply a more sophisticated version of old academic tricks. I remember an exam at university in which I did not know how to answer one question about early programming. My handwriting was terrible, so I wrote something barely legible. The instructor gave me a four out of five and said she could not read part of it but could see that I seemed to be writing the right thing. She was not going to spend the time calling me in to decipher every line. Many people have told me they used the same tactic. A process can therefore be made to look compliant if the person knows what the system expects.
The person can simply produce the “right” history while consulting an AI-generated answer in parallel. Diagnosing that is extremely difficult. This brings us back to the central limitation: evidence about the process is useful, but it is not the same as proof of authorship or intent. Before drawing the final conclusion, it is worth looking at how Google and OpenAI approach marking.
Google uses SynthID Text in Gemini. During generation, the model subtly influences the choice of words or phrase fragments to create a statistically detectable pattern. It is not a set of hidden characters or clipboard metadata. Google published research on the system in Nature and deployed it in Gemini and Gemini Advanced. The same limitations remain: screenshots, translation, re-generation through another model, or heavy editing may disrupt the signal. OpenAI’s ChatGPT does not currently have a comparable mass-market watermark for ordinary text.
OpenAI has introduced verifiable provenance and watermarking for supported images and audio, but ordinary text is not listed in the same way. For Meta, Microsoft, Amazon, xAI, and many other providers, the public picture is also uneven. Google says it tested SynthID Text on nearly twenty million real Gemini responses and found no meaningful difference in user ratings, suggesting that the watermark did not noticeably reduce text quality. That matters because people often try to identify AI from superficial habits: long dashes, certain emoji, symmetrical lists, or a particular bullet style. Those habits may be common in model output, but none of them is proof. Startups have spent the last three years building businesses around this uncertainty.
Some of those companies are still valuable and earn real money, but I do not believe the underlying promise is durable. Models have improved so rapidly that reliable detection from the finished output alone is becoming less likely. Soon, text, images, video, and audio will be good enough that a person cannot distinguish machine-made work from human-made work by perception alone. Then the more important question is not “Did a machine touch this?” but “Why does that matter here?” Sometimes I read a human-written analysis and ask why the author did not use AI to verify the claims, test the logic, and improve the work. Why waste the time of several people checking something that a strong model could help examine? That is precisely where AI is useful. The goal cannot be to preserve human labor for its own sake. We need rules that focus on responsibility, evidence, and outcomes rather than treating every instance of AI assistance as misconduct.