Machine text watermarking sounds like a label added after writing. It is not. A real text watermark is inserted before the sentence exists, at the instant a model chooses each next token. The provider changes the sampling rule by a small secret amount. The prose still looks ordinary, but across many token choices it contains a statistical preference that should not appear by accident. A detector with the same secret key can later recompute the choices the sampler was nudged toward and ask whether the text followed them too often.
The mechanism is easiest to see in the green-list watermark proposed by Kirchenbauer and coauthors. Before each token is sampled, a keyed hash of the recent context partitions the vocabulary into two temporary sets. Tokens in one set are promoted by a small logit bonus. The model is not forced to choose them, so the prose remains fluent; it is merely made a little more likely to choose them when several continuations are already plausible. Detection then counts how often the produced text chose the promoted set. If the count is far above the null expectation, the text carries the signature.
Google DeepMind's SynthID-Text uses a more production-shaped version. In the Nature paper published in 2024, the authors describe a generative watermark that changes only the sampling procedure, does not require retraining, and can be detected without running the original model. Their tournament sampler draws candidate tokens from the ordinary model distribution, runs a seeded tournament over those candidates with pseudorandom scoring functions, and emits the winner. The watermark is the correlation between the hidden scoring function and the final sequence. DeepMind reported deployment in Gemini and a live evaluation over nearly twenty million responses. The public Google documentation frames SynthID Text as a logits processor with secret integer keys and an n-gram context length; the detector returns probabilistic states rather than certainty.
This matters because the watermark is not a universal test for machine writing. It is a signature produced by one sampler under one key. If the generator did not participate, the detector has no cryptographic relation to the text. A third-party detector that merely says a paragraph “looks AI-generated” is doing something different: it is classifying style, entropy, or distributional regularities. Those regularities shift with model improvements, genre, language, education level, and editing. The most dangerous confusion is to treat a weak stylistic classifier as if it were a keyed watermark.
The watermark's power is also conditional on entropy. When the model has freedom among many acceptable continuations, a small bias can accumulate invisibly. When the output is a fact, a quotation, a code fragment, a legal citation, or a rigid template, there may be no safe freedom to exploit. Low-entropy text gives the sampler fewer opportunities to hide a signal. Short text has the same problem: the statistical test has too little evidence. A one-sentence answer can be honest and watermarked yet undetectable; a long essay gives the detector enough trials.
Paraphrase is the central attack. If a second model rewrites the text, the original token sequence disappears and the green-list count or tournament score can collapse. Semantic watermarks such as SIR try to address this by deriving the watermark from the meaning of the preceding text rather than the exact preceding tokens. That increases robustness against synonym substitution and paraphrase, but it moves the system away from the cleanest form of invisibility. The field lives inside a trilemma: the watermark can be invisible, robust to rewriting, or publicly verifiable and hard to spoof, but not all three at once.
Regulation is pushing the issue from research into infrastructure. The EU AI Act's Article 50 transparency obligations, applicable from 2 August 2026, require providers of generative systems to mark AI-generated outputs in a machine-readable way and make them detectable as artificial or manipulated as far as technically feasible. The EU code of practice treats marking and detection as compliance machinery, not as a magic guarantee. China went further earlier: the 2025 Measures for Labeling of AI-Generated Synthetic Content require both explicit labels and implicit labels, with metadata as the baseline and digital watermarks encouraged, effective 1 September 2025. These rules do not require that every label be a statistical text watermark. Metadata, provenance records, and visible disclosures can satisfy parts of the same obligation.
The structural pattern is therefore not “AI text can now be detected.” The pattern is delegated authorship through a controlled channel. A watermark works while the producing institution controls the pen: the model, sampler, key, tokenizer, and detector are one system. The moment text leaves that channel, it can be copied, cropped, translated, paraphrased, mixed with human prose, or regenerated by an unwatermarked open model. The signal becomes evidence about a path through infrastructure, not an essence in the sentence.
That distinction is the policy hinge. Watermarks are useful for platform accountability, provenance audits, and compliance claims by first-party providers. They are weak as courtroom-style proof against an individual writer unless the chain of generation, key custody, detector threshold, false-positive rate, and text handling are all specified. A watermark can say “this sequence is statistically consistent with generation by this keyed sampler.” It cannot say “no human wrote this,” and it cannot police models that never agreed to watermark in the first place.
The likely future is layered provenance. First-party generators will embed keyed statistical signals. Files and platforms will carry metadata or C2PA-style assertions. Public interfaces will expose labels or verification APIs. Regulators will demand that these marks survive ordinary distribution as far as feasible. Attackers will route around them through rewriting and open models. Honest systems will treat the output as probabilistic evidence, not as an oracle.
The exact technical lesson is simple: watermarking does not mark language; it marks a generation process. It is not a dye in the ink. It is a private bias in the hand that moves the pen, recoverable only when the verifier knows how that hand was biased.