Abstract

To evaluate AI-text detectors, researchers must compare known-human productions versus known-machine productions. This is commonly achieved by using archived human writing from the pre-AI era and newly generated machine text. However, it is becoming increasingly common to use AI as a first draft and then edit the text to appear more human. To test whether this impacted classification, we evaluated the Pangram AI-text detector, version 3.3.2, on 20 pre-2020 biomedical introductions and 80 matched LLM-derived passages. The detector labeled every historical passage human and every generated, copyedited, “deslopped,” or paraphrased passage AI. We also tested 220 variants with controlled punctuation and word corruption. Every variant remained classified as AI.

It is likely though that this study design is, or will soon be, outdated. As people repeatedly read, edit, and converse with large language models, it is certain that they will adopt some of its linguistic patterns and verbiage. Therefore, the natural human text productions will themselves begin to diverge from the pre-AI human corpus. As humans begin to sound more like AI, what will constitute "human" style? It is possible that a human will naturally write a sentence originally generated by AI; is it fair to consider this non-human?

In a separate longitudinal analysis of 11,465 arXiv titles and abstracts, prespecified LLM-associated terms rose sharply after 2022, peaked in 2025, and then fell in 2026. This elucidates a measurement problem. Did humans begin using AI more frequently or did they simply internalize AI styled writing? Then, one must explain why the drop occurred in 2026. Did model-associated language distributions change over time or did humans intentionally begin obfuscating AI generated text? These issues are not easily disentangled, and thus detector evaluation should use time-indexed controls rather than treating “human writing” as a permanent baseline.

100/100Study passages classified as expected
220/220Corrupted AI variants still labeled AI
11,465arXiv records analyzed

Note: This is a small pilot using one commercial detector, one generator, and one scientific domain. Perfect results within this sample do not establish perfect real-world accuracy.

The question

Studies have shown that paraphrasing and adversarial transformations can disrupt AI-text detector classifications [1][2]. However, these successful techniques, such as DIPPER, are often large and unfeasible on low-grade commercial hardware. Common edits conducted by, for example, students, are far more likely to invoke editors that remove familiar “AI style”: formulaic contrasts, canned transitions, vague significance claims, uniform rhythm, and other recurrent patterns. We call that process deslopping.

We asked whether copyediting, deslopping, or broad paraphrasing changed a current commercial detector’s classification of matched scientific prose. We then asked what a historical human comparison means if contemporary human writing is itself changing.

Note: Deslopping is not necessarily detector evasion. It can be ordinary editing for directness, specificity, and disciplinary fit.

Methods

Historical controls and matched generations

We selected 20 open-access biomedical articles published from 2015 through 2018. Each supplied a 550-to-800-word Introduction or Background passage. For every article, a pinned OpenAI model received the title, abstract, and target length and wrote a matched scientific introduction using only supplied information.

Each generation independently entered three transformations: a local copyedit; a fixed deslopping revision; and a substantial paraphrase changing syntax, sentence boundaries, transitions, and wording. No transformation received detector output, and no passage was resampled to lower a score. The design produced 20 historical passages and 80 LLM-derived passages.

Controlled corruption

We applied nested punctuation corruption to each raw AI passage at six dose levels, up to eight edits per 100 words. Validation confirmed that words and their order did not change. We separately applied word noise (letter transposition, character deletion, article deletion, and function-word duplication) at five dose levels, up to four edits per 100 words. These experiments added 220 variants without detector feedback.

Detection and longitudinal analysis

Every passage was submitted once to Pangram’s V3 API, which identified itself as version 3.3.2; its developers describe the classifier in a technical report [7]. We archived results, source material, prompts, transformations, model identifiers, and hashes. Separately, we sampled public arXiv metadata from January through July in 2010, 2015, 2020, and 2022 through 2026 across computational linguistics, high-energy phenomenology, and probability. A prespecified 27-term lexicon measured LLM-associated vocabulary in 11,465 titles and abstracts. This tests temporal distribution shift.

Results

Editing did not change a classification

Pangram labeled all 20 historical passages Human and all 80 LLM-derived passages AI. Document fractions were saturated: 0.0 for every human passage and 1.0 for every AI-derived passage. The observed false-positive rate was 0/20, with a two-sided 95% Clopper–Pearson upper bound of 16.8%. Sensitivity was 80/80, with a lower bound of 95.5%.

ConditionnMean wordsLabeled AIAI fraction
Human20709.000.00
Raw AI20720.6201.00
Copyedited AI20717.9201.00
Deslopped AI20640.9201.00
Paraphrased AI20665.0201.00

Deslopping removed about 11% of the raw generation’s words; paraphrasing removed about 8%. Neither changed a classification.

Low-level errors barely moved the score

All 120 punctuation-corrupted and 100 word-corrupted passages remained labeled AI. At maximum dose, punctuation corruption averaged roughly 57 edits and word corruption averaged 28.95 edits per passage. Pangram’s continuous mean window score moved from 0.99294 for raw AI to 0.99272 and 0.99254, respectively. Historical human windows averaged 0.00389.

CorruptionMaximum doseMean editsAI labelsWindow score
None0020/200.99294
Mixed punctuation8 per 100 words≈5720/200.99272
Mixed word noise4 per 100 words28.9520/200.99254

The vocabulary baseline moved

Target terms occurred 209 times per million words in 2010 and 416 times in 2022. The rate rose to 802 in 2023, 1,656 in 2024, and 1,835 in 2025, before falling to 737 in 2026. Document prevalence followed the same path.

YearPapersTerms per millionPapers with a term
20101,054208.82.7%
20151,411293.63.7%
20201,500389.05.5%
20221,500415.96.1%
20231,500802.210.7%
20241,5001,655.618.9%
20251,5001,834.721.1%
20261,500737.111.3%

Obviously, static list of notorious AI words is not a permanent fingerprint. Models can change, writers can avoid publicized markers, and topics within fields can shift.

The historical-control problem

Pre-2020 prose offers comparatively clear provenance, but does not represent the population on which detectors are now used. People acquire words through repeated exposure [3] and coordinate lexical choices with conversational partners [4]. LLMs now supply language across email, education, journalism, software, and scholarly editing. Studies have documented abrupt LLM-associated vocabulary shifts in biomedical and other scientific writing, along with later changes consistent with a moving linguistic baseline [5][6][8]. Someone can therefore internalize, use, or encounter model-shaped language without directly using AI. A detector can separate historical human prose from current model prose while users interpret the result as a judgment about current human prose versus model authorship, though these situations are quite different!

Conclusion

In this paired biomedical sample, copyediting, deslopping, paraphrasing, and substantial low-level corruption did not disrupt AI-detection classifications by Pangram 3.3.2. This demonstrates the effectiveness and robustness of current AI-text detectors and increases the burden of proof required for those who object to their use. Separately, as model-associated language circulates through environments where people read and write, historical controls become attractive because provenance is clearer and inadequate because they may no longer represent current humans. AI-text detection should be evaluated as a temporally shifting measurement problem. This is partially visible through the frequency of scientific vocabulary associated with LLMs increasing 2022 and then partly reversing.

References

  1. Krishna K, Song Y, Karpinska M, Wieting J, Iyyer M. “Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense.” NeurIPS. 2023;36.
  2. Wang Y, Feng S, Hou A, et al. “Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks.” ACL. 2024.
  3. Gaskell MG, Dumay N. “Lexical competition and the acquisition of novel words.” Cognition. 2003;89(2).
  4. Brennan SE, Clark HH. “Conceptual pacts and lexical choice in conversation.” JEP: Learning, Memory, and Cognition. 1996;22(6).
  5. Kobak D, González-Márquez R, Horvát E-Á, Lause J. “Delving into LLM-assisted writing in biomedical publications through excess vocabulary.” Science Advances. 2025;11(27).
  6. Geng M, Trotta R. “Human–LLM Coevolution: Evidence from Academic Writing.” Findings of ACL. 2025.
  7. Emi B, Spero M. “Technical Report on the Pangram AI-Generated Text Classifier.” arXiv:2402.14873. 2024.
  8. Sanger MD, Maurer BW. “Have Large Language Models Enhanced the Way Civil & Environmental Engineers Write?” arXiv:2602.03864. 2026.

AI-use disclosure: OpenAI’s GPT‑5.6 Sol was used through Codex to assist with this study, generate text used in it, and with the writeup.

← Back to Tennessee Labs