Abstract
To evaluate AI-text detectors, researchers must compare known-human productions versus known-machine productions. This is commonly achieved by using archived human writing from the pre-AI era and newly generated machine text. However, it is becoming increasingly common to use AI as a first draft and then edit the text to appear more human. To test whether this impacted classification, we evaluated the Pangram AI-text detector, version 3.3.2, on 20 pre-2020 biomedical introductions and 80 matched LLM-derived passages. The detector labeled every historical passage human and every generated, copyedited, “deslopped,” or paraphrased passage AI. We also tested 220 variants with controlled punctuation and word corruption. Every variant remained classified as AI.
It is likely though that this study design is, or will soon be, outdated. As people repeatedly read, edit, and converse with large language models, it is certain that they will adopt some of its linguistic patterns and verbiage. Therefore, the natural human text productions will themselves begin to diverge from the pre-AI human corpus. As humans begin to sound more like AI, what will constitute "human" style? It is possible that a human will naturally write a sentence originally generated by AI; is it fair to consider this non-human?
In a separate longitudinal analysis of 11,465 arXiv titles and abstracts, prespecified LLM-associated terms rose sharply after 2022, peaked in 2025, and then fell in 2026. This elucidates a measurement problem. Did humans begin using AI more frequently or did they simply internalize AI styled writing? Then, one must explain why the drop occurred in 2026. Did model-associated language distributions change over time or did humans intentionally begin obfuscating AI generated text? These issues are not easily disentangled, and thus detector evaluation should use time-indexed controls rather than treating “human writing” as a permanent baseline.
Note: This is a small pilot using one commercial detector, one generator, and one scientific domain. Perfect results within this sample do not establish perfect real-world accuracy.
The question
Studies have shown that paraphrasing and adversarial transformations can disrupt AI-text detector classifications [1][2]. However, these successful techniques, such as DIPPER, are often large and unfeasible on low-grade commercial hardware. Common edits conducted by, for example, students, are far more likely to invoke editors that remove familiar “AI style”: formulaic contrasts, canned transitions, vague significance claims, uniform rhythm, and other recurrent patterns. We call that process deslopping.
We asked whether copyediting, deslopping, or broad paraphrasing changed a current commercial detector’s classification of matched scientific prose. We then asked what a historical human comparison means if contemporary human writing is itself changing.
Note: Deslopping is not necessarily detector evasion. It can be ordinary editing for directness, specificity, and disciplinary fit.
Methods
Historical controls and matched generations
We selected 20 open-access biomedical articles published from 2015 through 2018. Each supplied a 550-to-800-word Introduction or Background passage. For every article, a pinned OpenAI model received the title, abstract, and target length and wrote a matched scientific introduction using only supplied information.
Each generation independently entered three transformations: a local copyedit; a fixed deslopping revision; and a substantial paraphrase changing syntax, sentence boundaries, transitions, and wording. No transformation received detector output, and no passage was resampled to lower a score. The design produced 20 historical passages and 80 LLM-derived passages.
Controlled corruption
We applied nested punctuation corruption to each raw AI passage at six dose levels, up to eight edits per 100 words. Validation confirmed that words and their order did not change. We separately applied word noise (letter transposition, character deletion, article deletion, and function-word duplication) at five dose levels, up to four edits per 100 words. These experiments added 220 variants without detector feedback.
Detection and longitudinal analysis
Every passage was submitted once to Pangram’s V3 API, which identified itself as version 3.3.2; its developers describe the classifier in a technical report [7]. We archived results, source material, prompts, transformations, model identifiers, and hashes. Separately, we sampled public arXiv metadata from January through July in 2010, 2015, 2020, and 2022 through 2026 across computational linguistics, high-energy phenomenology, and probability. A prespecified 27-term lexicon measured LLM-associated vocabulary in 11,465 titles and abstracts. This tests temporal distribution shift.
Results
Editing did not change a classification
Pangram labeled all 20 historical passages Human and all 80 LLM-derived passages AI. Document fractions were saturated: 0.0 for every human passage and 1.0 for every AI-derived passage. The observed false-positive rate was 0/20, with a two-sided 95% Clopper–Pearson upper bound of 16.8%. Sensitivity was 80/80, with a lower bound of 95.5%.
| Condition | n | Mean words | Labeled AI | AI fraction |
|---|---|---|---|---|
| Human | 20 | 709.0 | 0 | 0.00 |
| Raw AI | 20 | 720.6 | 20 | 1.00 |
| Copyedited AI | 20 | 717.9 | 20 | 1.00 |
| Deslopped AI | 20 | 640.9 | 20 | 1.00 |
| Paraphrased AI | 20 | 665.0 | 20 | 1.00 |
Deslopping removed about 11% of the raw generation’s words; paraphrasing removed about 8%. Neither changed a classification.
Low-level errors barely moved the score
All 120 punctuation-corrupted and 100 word-corrupted passages remained labeled AI. At maximum dose, punctuation corruption averaged roughly 57 edits and word corruption averaged 28.95 edits per passage. Pangram’s continuous mean window score moved from 0.99294 for raw AI to 0.99272 and 0.99254, respectively. Historical human windows averaged 0.00389.
| Corruption | Maximum dose | Mean edits | AI labels | Window score |
|---|---|---|---|---|
| None | 0 | 0 | 20/20 | 0.99294 |
| Mixed punctuation | 8 per 100 words | ≈57 | 20/20 | 0.99272 |
| Mixed word noise | 4 per 100 words | 28.95 | 20/20 | 0.99254 |
The vocabulary baseline moved
Target terms occurred 209 times per million words in 2010 and 416 times in 2022. The rate rose to 802 in 2023, 1,656 in 2024, and 1,835 in 2025, before falling to 737 in 2026. Document prevalence followed the same path.
| Year | Papers | Terms per million | Papers with a term |
|---|---|---|---|
| 2010 | 1,054 | 208.8 | 2.7% |
| 2015 | 1,411 | 293.6 | 3.7% |
| 2020 | 1,500 | 389.0 | 5.5% |
| 2022 | 1,500 | 415.9 | 6.1% |
| 2023 | 1,500 | 802.2 | 10.7% |
| 2024 | 1,500 | 1,655.6 | 18.9% |
| 2025 | 1,500 | 1,834.7 | 21.1% |
| 2026 | 1,500 | 737.1 | 11.3% |
Obviously, static list of notorious AI words is not a permanent fingerprint. Models can change, writers can avoid publicized markers, and topics within fields can shift.
The historical-control problem
Pre-2020 prose offers comparatively clear provenance, but does not represent the population on which detectors are now used. People acquire words through repeated exposure [3] and coordinate lexical choices with conversational partners [4]. LLMs now supply language across email, education, journalism, software, and scholarly editing. Studies have documented abrupt LLM-associated vocabulary shifts in biomedical and other scientific writing, along with later changes consistent with a moving linguistic baseline [5][6][8]. Someone can therefore internalize, use, or encounter model-shaped language without directly using AI. A detector can separate historical human prose from current model prose while users interpret the result as a judgment about current human prose versus model authorship, though these situations are quite different!
Conclusion
In this paired biomedical sample, copyediting, deslopping, paraphrasing, and substantial low-level corruption did not disrupt AI-detection classifications by Pangram 3.3.2. This demonstrates the effectiveness and robustness of current AI-text detectors and increases the burden of proof required for those who object to their use. Separately, as model-associated language circulates through environments where people read and write, historical controls become attractive because provenance is clearer and inadequate because they may no longer represent current humans. AI-text detection should be evaluated as a temporally shifting measurement problem. This is partially visible through the frequency of scientific vocabulary associated with LLMs increasing 2022 and then partly reversing.
References
- Krishna K, Song Y, Karpinska M, Wieting J, Iyyer M. “Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense.” NeurIPS. 2023;36.
- Wang Y, Feng S, Hou A, et al. “Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks.” ACL. 2024.
- Gaskell MG, Dumay N. “Lexical competition and the acquisition of novel words.” Cognition. 2003;89(2).
- Brennan SE, Clark HH. “Conceptual pacts and lexical choice in conversation.” JEP: Learning, Memory, and Cognition. 1996;22(6).
- Kobak D, González-Márquez R, Horvát E-Á, Lause J. “Delving into LLM-assisted writing in biomedical publications through excess vocabulary.” Science Advances. 2025;11(27).
- Geng M, Trotta R. “Human–LLM Coevolution: Evidence from Academic Writing.” Findings of ACL. 2025.
- Emi B, Spero M. “Technical Report on the Pangram AI-Generated Text Classifier.” arXiv:2402.14873. 2024.
- Sanger MD, Maurer BW. “Have Large Language Models Enhanced the Way Civil & Environmental Engineers Write?” arXiv:2602.03864. 2026.
AI-use disclosure: OpenAI’s GPT‑5.6 Sol was used through Codex to assist with this study, generate text used in it, and with the writeup.