Will We No Longer Be Able to Tell Humans from AI? What Counts as Evidence Now?
Can we reliably establish where a text or other piece of content came from—and what evidence is strong enough to justify punishing a person?
The reader will be able to distinguish four different verification methods, understand the limits of each one, and design a process in which an AI signal starts an investigation instead of becoming an automatic verdict.
What to watch for
Key takeaways
Different services can rate the same text in opposite ways; the score is a model classification, not proof of where the idea came from.
Anthropic’s signal may show that a supported model participated in generation, but it does not establish whether the person wrote the original material.
Without explicit rules, a company cannot distinguish transparent use of a work tool from hidden substitution for competence.
In science and journalism, sources, data, and the verification chain matter more than the stylistic feeling that “AI writes this way.”
A lawyer or doctor may use a model, but the license and responsibility belong to the person who approves the final decision.
The Adelphi case shows why a contradictory detector result cannot justify punishment without a transparent process.
The first looks for a known provider signal; the second infers from language patterns. Treating them as equivalent creates false confidence.
A ten-percent error rate may be tolerable in a recommendation and unacceptable in a disciplinary decision.
A watermark, file provenance, statistical analysis, and process history can reinforce one another, but none is absolute.
If a person knows what the system measures, they can retype machine-generated text and create a plausible human process log.
Google changes probabilistic token choices during generation; the result is not hidden characters or a visible writing style.
When the distinction disappears, rules should protect contexts where authorship is critical without punishing useful AI assistance by default.
What this episode is about
A student writes an essay independently and an AI detector returns “100% AI.” Someone else takes fully generated text, changes it slightly, and the same kind of system sees ordinary human work. At that point, AI detection stops being a technical curiosity. When the result affects a university record, a job, a legal dispute, or a professional reputation, an unreliable percentage becomes a real punishment.
Alexander Volchek begins with a simple trap: we keep asking whether a human or a machine made the work, even though today’s tools answer several different questions. A provider watermark may show that text passed through Claude or Gemini. Cryptographic provenance can record which service created or processed a file. A statistical detector estimates whether the language resembles typical model output. A document history records what a person typed, pasted, and edited. None of those signals proves deception on its own.
Education shows the problem clearly. In Newby v. Adelphi University, a professor relied on Turnitin, which reportedly classified a paper as fully AI-generated. The student denied using a model and pointed to the Google Docs history; other checks returned the opposite result. The court did not prove mathematically who wrote every line. It established a narrower and more important principle: a contradictory detector score and a weak internal procedure were not enough to punish a student. Universities must distinguish a risk signal from a verdict.
The same logic applies to hiring. Across more than twenty years of recruitment, Alexander has seen the gap between what a candidate demonstrates in an interview and how that person later performs. AI widens the gap. One candidate may use a model transparently as a real work tool; another may receive invisible answers and imitate competence. A one-line ban cannot solve that. The company must decide in advance what it is testing: knowledge, problem framing, output quality, or the ability to work without outside assistance.
In science, journalism, law, and medicine, the cost of a mistake is even higher. The important questions include whether a source exists, whether the data are authentic, whether an experiment can be traced, and which professional accepts responsibility. A lawyer may use AI to draft a contract, a doctor to prepare a preliminary note, or a journalist to test a quotation. Tool use alone says little about quality. The real problem begins when the human stops checking the result or hides a material role played by the system.
Anthropic’s approach is an invisible, machine-readable signal embedded during text generation. It may survive copying and light editing, but translation, heavy rewriting, or mixing with other material can weaken it. A short excerpt may be too small to test. The reverse case matters just as much: a person writes an article and sends it to Claude only for translation or editing, after which the human-authored work may carry a machine mark. The watermark indicates tool involvement; it does not identify the origin of the idea.
The episode separates four evidence layers. The first is a provider watermark. The second is signed file provenance such as C2PA. The third is a statistical detector such as Turnitin, GPTZero, Copyleaks, or Originality.ai. The fourth is a creation log in Google Docs, Microsoft Word, or an authorship-tracking tool. The more serious the consequence, the less rational it is to trust one layer. Institutions need multiple signals, context, and an appeal process.
Even a document history is not absolute proof. AI-generated text can be displayed on a second monitor and retyped gradually over several months. The log will show manual typing, revisions, and versions. It is a modern version of an old academic trick: once a person knows what the system measures, the person can imitate the expected process. Verification should therefore make deception harder without destroying an innocent person’s life after a false positive.
The deeper shift comes next. Models are approaching a level at which people will not be able to distinguish text, images, voices, or video from human work by style alone. Authorship will still matter in specific contexts—an exam, licensed professional work, a court record, a scientific result, or political advertising. In many other situations, the better questions are whether the claims were checked, who is responsible for the outcome, and whether the work improved. Sometimes the problem is not that a person used AI. It is that the person refused to use a tool that could have found errors and strengthened the analysis.
When human and model output can no longer be distinguished by sight, the focus must move from guessing style to an evidence chain, transparent rules, and responsibility for the result. A system may show that a tool was involved; it should not decide by itself who the author is or who is guilty.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 28 segments: 28 identified, 0 mixed, 0 marked with ✓, and 0 unresolved.
Loading…