Fraud doubles in 18 months. Retraction takes 40.
The scientific literature is the one place in this territory with a real record, and the record shows the correction machinery growing at less than half the rate of the thing it corrects.
TL;DR. A Northwestern study in PNAS, August 2025, screened tens of millions of papers and found suspected paper-mill output doubling every 1.5 years, against a 15-year doubling time for legitimate output. Retractions double every 3.3 years. Of more than 32,000 fraudulent articles identified, 29% had been retracted. Meanwhile global output is on pace to pass 6 million articles in 2026, one AI conference took more than 30,000 submissions after about 15,000 the year before, and reviewers gave an estimated 100 million hours of unpaid labour in 2020 alone. This is the one subject in this territory with a genuine record, and what the record shows is a correction system growing at less than half the rate of the problem.
---
Status: established, with classifier caveats carried. Primary sources: the Northwestern PNAS study, August 2025; Retraction Watch data via Crossref; conference submission counts; and a BMJ methodological study screening 2.6 million cancer papers. Prevalence figures rest on classifiers trained on confirmed cases and inherit their error, which the article states rather than glosses.
---
The two rates
Suspected paper-mill output is doubling every 1.5 years.
Legitimate scientific output doubles every 15 years.
Retractions double every 3.3 years.
Those three numbers are the article. A problem compounding on an 18-month cycle against a correction mechanism compounding on a 40-month one does not converge. The gap widens by construction, and no amount of effort inside the current system closes it, because the constraint is the doubling time rather than the level.
The consequence is already visible in the stock. Of more than 32,000 fraudulent articles identified in that study, 29% had been retracted. Roughly seven in ten remain in the literature, citable and cited.
And the record exists, which is the unusual part
The search referral article opened this territory on a domain with no disclosure regime at all, where the platform holds the data and publishes nothing.
This is the opposite case. Retraction Watch maintains a database. Crossref carries the notices. PubMed indexes the corpus. The Northwestern team screened tens of millions of papers because those papers are indexed, permanent and public.
That is disclosure obligation working, and it is why this subject has real numbers while synthetic web content has none.
It also means the finding is checkable rather than asserted, which is worth stating plainly given how much of this territory rests on estimates.
What the volume looks like
Global scholarly output is on pace to cross 6 million articles in 2026, from about 5.5 million in 2025. Web of Science indexed roughly 2.53 million new studies in 2024, a 48% rise on 2015.
The review system did not grow with it. An estimated 100 million hours of unpaid reviewing labour were given worldwide in 2020, and editors report a widening gap between submissions and willing referees.
One venue makes the strain legible. A major AI conference received more than 30,000 initial submissions for 2026, against roughly 15,000 for 2025. A doubling in a single cycle, into a review process that is structurally the same as it was a decade ago.
And the response is predictable. At one 2024 conference, 49.4% of submissions received at least one AI-assisted review, with model-generated content estimated in 4,428 of 28,028 reviews.
Which is the loop worth naming: submissions rise partly because drafting got cheaper, and the review capacity gap is closed partly by the same tools, so both sides of the filter are now running on the technology the filter was meant to assess.
The largest correction event on record
One publisher retracted more than 11,300 papers from a single acquired portfolio between 2022 and 2024.
That is what a working correction system looks like at scale, and it is also the exception. It happened because the fraud was concentrated in one imprint, was detectable through submission-process anomalies, and had a commercial owner with a reason to act.
Retraction notices from 2025 name three recurring causes: compromised peer review and paper mills; citation manipulation and citation cartels; image and data manipulation. Over 6,400 notices cite phrasing markers including tortured phrases, nonstandard phrases, and text generated by a language model. Around 2,300 involve paper mills directly and 1,800 citation manipulation.
One journal stopped accepting a category of submission entirely after being overwhelmed by model-generated commentaries, which is the honest response to a filter that no longer filters.
Where the numbers get soft
Every prevalence figure here rests on a classifier, and the detection article established what that costs.
The cancer-research screening study is explicit about its method: a BERT text classifier trained on 2,202 retracted paper-mill papers, validated against independent expert assessment, applied to 2.6 million papers. It reports papers flagged as similar to known paper-mill output, which is not the same as papers that are paper-mill output.
And the AI-authorship estimates diverge enormously. Corpus detection puts computer science at 17.5% to 22% AI-drafted. Self-report puts scientist usage at 30% in 2023 rising to 57% by 2025. Confirmed fabrication traced to a model sits at around 139 papers identifiable on one search index.
Those three numbers measure different things: detectable style, stated behaviour, and proven fabrication. The source reporting them says so directly, noting that the gap between 139 and 57% reflects detection difficulty rather than absence, and that is the correct reading.
So the doubling-time finding is robust and the level estimates are not, because the first is a rate comparison within one method and the second requires knowing a true prevalence nobody can measure.
Three things this establishes
Doubling times decide outcomes, not effort. A correction system improving steadily still loses to a problem compounding twice as fast. The policy question is not how hard anyone is trying; it is whether anything changes the rate.
A record makes a problem measurable and does not make it solvable. Scientific publishing has the disclosure infrastructure this whole territory otherwise lacks, and the infrastructure documented the failure without preventing it.
And the filter is now partly made of the thing it filters. Roughly half of submissions at one venue drew a model-assisted review, while submission volume rose partly because drafting got cheap. A quality control system running on the technology under assessment is not obviously stable.
What it does not establish
That most published research is fraudulent. Suspected mill output is a small fraction of a very large corpus. The finding is about growth rates, not about the current composition of the literature.
That the flagged papers are fraudulent. They are papers a classifier judged similar to confirmed cases, and the study says so.
That AI caused this. Paper mills predate language models, the incentive structure that pays for authorship is the underlying cause, and cheaper drafting lowers a cost that was already being paid.
And nothing about any individual paper, author or institution. Every figure here is aggregate.
What is unresolved
Whether screening at scale changes the rate. Classifier-based detection is now being applied to millions of papers, and whether that shortens the retraction doubling time is not yet observable.
What fraction of the literature is affected. The 29% retraction figure describes identified cases. The unidentified population is unmeasured by construction, which is survivorship in its purest form: the record contains what was caught.
Whether review can be rebuilt. Nothing in the current model scales to six million articles, and the proposals in circulation, including model-assisted screening, are being trialled rather than evaluated.
And what happens to citation. Roughly seven in ten identified fraudulent papers remain in place, accumulating citations that propagate into work that is otherwise sound.
The counter-argument
Doubling times from a classifier are a rate of detection, not a rate of fraud. If screening improved sharply over the study window, suspected output would appear to double faster than it does. The article treats a detection curve as a production curve, and the two are only equal if detection efficiency held constant, which nobody has shown.
The 29% figure may reflect process rather than failure. Retraction requires investigation, notice and often institutional cooperation, and a lag of years is procedural rather than negligent. Comparing a stock of identified cases to a completed-retraction count at one moment overstates the shortfall.
Volume growth is not degradation. Six million articles reflects more researchers in more countries publishing more, which is what expanded access looks like. Reading a submission rise as a quality problem imports an assumption about the previous equilibrium being correct.
And the review-loop framing is neat rather than demonstrated. That models draft submissions and assist reviews is documented; that this constitutes a destabilising feedback loop is an inference, and no measurement here supports it.
The short version
Suspected paper-mill output doubles every 1.5 years. Legitimate output doubles every 15. Retractions double every 3.3. Of more than 32,000 identified fraudulent articles, 29% had been retracted, leaving roughly seven in ten in place.
A problem compounding on an 18-month cycle against a correction compounding on a 40-month one does not converge, and the constraint is the rate rather than the effort.
The volume underneath it is large and rising. Around 6 million articles in 2026, 2.53 million new studies indexed in 2024 at 48% above 2015, and an estimated 100 million hours of unpaid review given in a single year. One AI conference went from about 15,000 to more than 30,000 submissions in one cycle, and at another 49.4% of submissions drew at least one AI-assisted review.
This is the one subject in this territory with a real record, because retractions are indexed, permanent and public. The infrastructure documented the failure and did not prevent it, which is worth holding against any argument that disclosure alone is the remedy.
And the level estimates remain soft. Prevalence rests on classifiers trained on confirmed cases, AI-authorship estimates run from 139 confirmed fabrications to 17.5 to 22% by detection to 57% by self-report, and those measure three different things. The rate comparison is the robust finding. The levels are not.
Common questions
What is the central finding? That the rates diverge. A Northwestern study published in PNAS in August 2025 found suspected paper-mill output doubling every 1.5 years against a 15-year doubling time for legitimate scientific output, while retractions double only every 3.3 years. Of more than 32,000 fraudulent articles identified, 29% had been retracted. A problem compounding twice as fast as its correction does not converge regardless of effort.
How large is the volume problem? Global scholarly output is on pace to cross 6 million articles in 2026, up from about 5.5 million in 2025, and Web of Science indexed roughly 2.53 million new studies in 2024, a 48% rise on 2015. Reviewers gave an estimated 100 million hours of unpaid labour in 2020 alone. One major AI conference received more than 30,000 submissions for 2026 against roughly 15,000 the year before.
Are reviewers using AI too? At one 2024 conference, 49.4% of submissions received at least one AI-assisted review, with model-generated content estimated in 4,428 of 28,028 reviews. Submissions rise partly because drafting became cheap, and the review capacity gap is being closed partly by the same tools, so both sides of the filter now run on the technology the filter is meant to assess.
Has anything worked at scale? One publisher retracted more than 11,300 papers from a single acquired portfolio between 2022 and 2024, which is the largest correction event on record. It worked because the fraud was concentrated in one imprint, detectable through submission-process anomalies, and owned by a party with a commercial reason to act. Those conditions are the exception.
How reliable are the prevalence figures? Less reliable than the rate comparison. Screening studies use classifiers trained on confirmed retracted cases, in one instance a BERT model trained on 2,202 papers and applied to 2.6 million, and they report papers flagged as similar to known paper-mill output rather than confirmed fraud. AI-authorship estimates range from about 139 confirmed fabrications on one index, to 17.5 to 22% by corpus detection, to 57% by self-report, and those three measure proven fabrication, detectable style and stated behaviour respectively.
Is AI the cause? Not on this evidence. Paper mills predate language models, and the underlying cause is an incentive structure that pays for authorship. What cheaper drafting does is lower a cost that was already being paid, which raises volume without creating the demand.
Why does it matter that this field has a record? Because most of this territory does not. Search referral effects, synthetic web content and training data composition are all measured by proxy or not at all. Retractions are indexed by Crossref, catalogued by Retraction Watch and permanent in PubMed, which is why a study could screen tens of millions of papers. The disclosure infrastructure made the failure measurable and did not prevent it, which is a useful check on any argument that transparency alone is the remedy.
What is the strongest objection? That a classifier's output is a detection rate rather than a production rate. If screening improved over the study window, suspected mill output would appear to grow faster than it actually did, and the two curves are only equal if detection efficiency held constant, which has not been shown. A second objection is that the 29% retraction figure partly reflects procedural lag, since retraction requires investigation, notice and often institutional cooperation, rather than reflecting negligence.
Sources
Primary documents only. Where a claim rests on a single report, the entry says so.
- Paper-mill output growth and retraction rates Northwestern University, PNAS, August 2025 The doubling times: 1.5 years for suspected paper-mill output against 15 years for legitimate output and 3.3 years for retractions, and the finding that 29% of more than 32,000 identified fraudulent articles had been retracted.
- Machine learning based screening of potential paper mill publications in cancer research BMJ, methodological and cross-sectional study The method behind prevalence estimates: a BERT classifier trained on 2,202 retracted paper-mill papers and applied to 2.6 million cancer papers, reporting similarity to known cases rather than confirmed fraud.
- Learning from Retractions to Drive Prevention Strategies MDPI, using Retraction Watch data via Crossref The 2025 distribution of retraction reasons across compromised peer review and paper mills, citation manipulation, and image and data manipulation.
Further reading
The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.
- International Energy Agency (2025), Energy and AI — a case where a research body measured what no regime required. :: https://www.iea.org/reports/energy-and-ai Disclosure Obligation
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — deployment at scale with no obligation to validate, and what an independent check found. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 Disclosure Obligation
- Pineau et al., reproducibility programme work, and Bouthillier et al. (2021), Accounting for Variance in Machine Learning Benchmarks — why a literature of reported results describes what was submitted. :: https://arxiv.org/abs/2103.03098 Survivorship Bias
Related articles
- Slop scores premium 70% of the timeThe first rigorous measurement of machine-generated content in ad buying found it passes every quality check the industry uses, and passes them better than real inventory does.
- Zero-click is 60%, or 22.4%, from one providerTerritory 9 opens on what cheap generation does to information. Every method agrees the traffic is falling. None agrees on how far, and the headline metric differs threefold within one dataset.
- The sepsis model caught 7% of what clinicians missedA sepsis warning system ran at hundreds of US hospitals before anyone outside the vendor validated it. The external check found the number that matters is not the one being reported.
- Three years in: no disruption, and one 20% holeEvery aggregate measure of AI's labour effect shows continuity. One within-firm comparison shows a fifth of a cohort gone. Both are well evidenced.