An AI Text Watermark Is Evidence—not a Verdict on Authorship

An AI text watermark can provide evidence that a passage is compatible with a particular watermarking process, but it cannot independently prove who authored the document, how much AI contributed or whether anyone broke a rule. Its weight depends on the scheme, detector configuration, amount and type of text, intervening edits and records connecting the output to a person or account.
A positive result may justify further investigation, not a verdict. A negative or uncertain result cannot reliably clear a document: the passage may be too short, thoroughly rewritten, translated, generated without that watermark or tested with an incompatible detector.
What a positive result actually establishes
A statistical watermark is introduced during generation by influencing token selection. A compatible detector later tests whether the observed sequence contains the expected statistical pattern. Unlike a conventional digital signature, that pattern does not contain a human author’s identity or a complete history of the document.
Detection is probabilistic and configuration-specific. Google’s SynthID Text documentation defines watermarked, not-watermarked and uncertain states, with thresholds controlling false-positive and false-negative rates. It also allows models with the same tokenizer to share a watermark configuration and detector when the detector was trained on examples from all participating models.
A positive result therefore supports a narrow proposition: the tested passage is compatible with the watermark under the detector, configuration and thresholds used. Without additional records, it does not identify the prompt operator, distinguish AI drafting from AI editing, attribute the text to one exact model within a shared configuration or establish a policy violation.
Length and content determine how much signal exists

Watermark detectors accumulate small statistical effects across tokens, so more text usually supplies more evidence. The published SynthID Text research evaluates detectability as a function of text length and identifies length as a primary factor in statistical certainty.
Content also affects the available signal. Google notes that watermark application is less effective for factual responses because there is less freedom to alter token selection without reducing accuracy. A mixed document can create another ambiguity: a strong result in one passage does not establish how the rest of the document was produced.
The tested span, token count, language, content type and treatment of quoted or previously published material are therefore part of the result—not administrative details. A document-level conclusion is unsupported if the report does not disclose which passage was examined or whether the detector was validated on comparable text.
Editing can remove a real mark while preserving meaning

Copyediting is not neutral to a token-based watermark. Replacing words, reordering sentences, translating text or regenerating it through another model changes the sequence on which detection depends. Meaning can remain substantially intact while the statistical signal weakens.
An ACL 2024 study of color-aware substitutions tested an attack against the green-token watermark family introduced by earlier research. Its SCTS method inferred token-color information, replaced selected green tokens and evaded detection with fewer edits than related attacks; the authors also demonstrated watermark removal for arbitrarily long watermarked text.
Those findings apply to the studied watermark family and attack conditions, not every deployed system. They nevertheless establish the relevant evidentiary limit: semantic continuity does not guarantee watermark continuity. A signal that survives light editing demonstrates resilience under those edits, but still does not resolve human authorship.
An evidence hierarchy for detector outcomes

The useful question is what an outcome justifies when its test conditions and corroborating records are known. Four tiers prevent a probabilistic score from being silently converted into a binary accusation.
- Positive, validated signal: The exact passage is preserved; the detector is appropriate for the suspected configuration; validation covers comparable text; and thresholds, error rates and versions are documented. This supports compatibility with the watermarking path and warrants corroboration, but does not alone prove authorship or misconduct.
- Positive, unverified signal: A tool reports a mark, but its configuration, calibration, version or chain of custody cannot be examined. The result is an investigative lead, not an attribution.
- Degraded or uncertain signal: The passage is short, mixed-source, translated or substantially rewritten. The result has limited discriminatory value and should remain inconclusive.
- Negative signal: No mark is detected. Possible explanations include human writing, output from an unwatermarked system, insufficient text, an incompatible detector or signal loss through editing. The outcome does not prove that no AI assistance occurred.
A July 2026 preprint evaluated KGW, Unigram and a MarkLLM implementation of SynthID-Text using 846 valid paraphrase runs across 15 prompts per method. The empirical forensic-readiness evaluation reported high initial false-negative rates and conditional watermark-removal rates of 100% for its KGW and Unigram configurations and 98.3% for its SynthID configuration after meaning-preserving paraphrase.
The paper also compared those configurations with Daubert factors and the NIST SP 800-86 digital-forensics process. Because it is a submitted preprint examining particular implementations, its percentages cannot be generalized to every watermark deployment. Its defensible broader contribution is the framework: detector output must be evaluated as evidence with known limitations, not accepted as self-validating proof.
Authorship and misconduct require a separate chain of proof
Provenance, authorship and misconduct are different propositions. Provenance concerns a technical path through which content may have passed. Authorship concerns intellectual and expressive contribution, while misconduct additionally requires an applicable rule, prohibited conduct and a fair process for evaluating the evidence.
NIST’s synthetic-content report treats content authentication, provenance tracking, watermarking and synthetic-content detection as related but distinct transparency approaches. A sound review can therefore combine the preserved text and detector details with legitimately obtained model-side records, account logs, revision history, citations, the applicable rule and the affected person’s explanation.
The permitted consequence should track the evidentiary tier. An unverified positive may justify a conversation; a validated result supported by account and revision records may justify formal review; a watermark result alone should not produce a final finding. The same distinction applies to systems that merely estimate AI authorship rather than test a provider-specific watermark.
The reliable conclusion is deliberately narrow: a text watermark can add probative evidence about a possible generation path. Its reliability is conditional, and independent, reviewable evidence must bridge the gap between that path and any judgment about a person.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.