Quasa
Use QUASA App
Join the pioneer of Web3 crypto freelancing today!
Open
Creator Economy

Substack’s AI Detector Estimates Authorship—it Does Not Prove It

|Author: QUASA Editorial Team|6 min read
Substack’s AI Detector Estimates Authorship—it Does Not Prove It

Substack’s Pangram-powered AI detector classifies patterns in finished text. Its percentage estimates how much of the analyzed prose appears human-written or AI-associated; it does not prove who wrote the post, which tools were used, or how much editorial judgment shaped the result.

Accuracy does not remove that boundary. An independent analysis of Substack’s scanner explains that the output is a statistical classification rather than provenance: the detector sees the published wording, not drafts, prompts, notes, recordings, revision history, or editorial decisions.

What the displayed percentage means

A result such as “30% AI-assisted” should be read as an estimate about portions of the submitted text. It is not necessarily 30% confidence, proof that a model supplied exactly 30% of the words, or evidence that the writer concealed AI use.

This distinction matters because a simple percentage can appear more conclusive than the underlying task. The detector can identify language associated with classes it learned during training, but it cannot determine who conceived the argument, verified the sources, chose which suggestions to reject, or accepted responsibility for publication.

A result may justify asking about the production process. By itself, it cannot establish identity, intent, originality, plagiarism, accuracy, or deception. The inverse is also true: a human classification does not certify that a post is original, accurate, or free of AI involvement.

How Pangram classifies text

Pangram’s technical explanation describes a supervised neural classifier that tokenizes text, converts tokens into numerical embeddings, processes them through a neural network, and produces human or AI labels. Its initial model was trained on approximately one million public or licensed human-written and AI-generated documents.

The same explanation describes hard-negative mining, a training method intended to reduce false positives. Human documents that the model incorrectly flags are added to the training data as difficult examples, paired with comparable generated text, and used in further training cycles.

This is pattern recognition, not a search for authorship metadata. Results can depend on the model version, decision threshold, document length, language, genre, represented generators, training distribution, and surrounding text. Changes in human writing and new generation or editing methods can also weaken the fit between a future document and an earlier evaluation set.

What the error rates actually tell you

Pangram 4 benchmark separates false positives on human text from false negatives on AI-generated text.

The Pangram 4 technical report, submitted on July 29, 2026, gives a 0.0041% false-positive rate and a 0.3396% false-negative rate at its evaluated operating point; it also evaluates fine-grained edits and mixed human–AI text. These are benchmark results for a specified model and test design, not guaranteed rates for every newsletter, language, genre, or Substack configuration.

A false positive occurs when human text is classified as AI-associated. A false negative occurs when AI text is missed. Applied mechanically, the published rates correspond to 0.41 expected false positives per 10,000 human documents and about 34 expected misses per 10,000 AI documents. Expected counts are long-run averages, not predictions about a particular batch.

Base rates determine what those errors mean in practice. Consider a hypothetical collection of 10,000 documents containing one AI document. If the benchmark rates transferred unchanged, the detector would identify almost one true positive while producing about 0.41 false positives among the human documents. About 29% of all positive flags would then be false, even though the false-positive rate itself is extremely low.

That example is not a forecast for Substack. It demonstrates why “false-positive rate” and “probability that this flag is wrong” are different quantities. The second depends on how common the target class is in the scanned population, as well as whether the benchmark conditions resemble real use.

The displayed text-share percentage answers another question again. It allocates analyzed text among classifications; it is not automatically a calibrated probability that the named author used AI, and the published benchmark rates do not establish that a specific Substack deployment uses identical settings.

Why mixed editing resists a simple verdict

Human and AI contributions are often interleaved rather than divided into clean blocks. One writer might develop the reporting and argument, use a model to reorganize a draft, reject most suggestions, and rewrite the remaining passages. Another might lightly edit a generated draft. Both workflows involve AI, but they reflect different levels of authorship and control.

A classifier can infer which passages resemble its learned categories, yet the final wording cannot reveal the entire production history. It cannot reconstruct whether a model summarized research, corrected grammar, translated a passage, proposed an outline, or supplied sentences that survived extensive rewriting.

Operational definitions create another source of ambiguity. A detector’s distinction between negligible editing, assistance, and generation may not match a reader’s understanding of “written by AI.” A percentage can therefore compress several different workflows into a label that appears to answer a broader question than the system actually measured.

A protocol for disputed results

Dated drafts, source notes, revision history, and an AI-use disclosure document how a newsletter was created.

Creators can preserve evidence that a text classifier cannot supply. The record should be proportionate to the work and should protect confidential material, but it needs enough continuity to connect research, drafting, editing, and publication.

  1. Define AI use by function. State whether tools were used for brainstorming, transcription, translation, summarization, copy editing, restructuring, or generation. A precise description is more informative than a generic “AI-assisted” label.
  2. Retain development records. Keep dated outlines, source notes, interview material, major draft versions, and revision history. Preserve material prompts and outputs when they substantially affected published language.
  3. Disclose consequential use. Identify AI-generated or materially rewritten passages when that role would affect a reasonable reader’s understanding of the work. Routine automated corrections need not be presented as equivalent to co-authorship.
  4. Document the disputed scan. Save the exact submitted text, displayed percentages, scan date, and detector or interface version when available. Comparisons are meaningful only when the text and surrounding context remain identical.
  5. Respond with process evidence. Use relevant drafts and a concise workflow account to address an authorship claim. Running the passage through another detector merely substitutes one classification for another.

The defensible conclusion remains narrow: Pangram can estimate which learned patterns appear in finished prose. Establishing who wrote a post and how AI affected it requires provenance, context, and accountable disclosure that no public percentage can provide alone.

Also read:

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0