Eight subjects, one ratio, and it is not quality
Territory 9 closes. Across eight subjects the damage came from a cost ratio inverting rather than from bad output, and in every case detection failed while friction worked.
TL;DR. Eight subjects, and the same structure underneath every one: producing a claim became nearly free while checking one did not move. Search referrals down 33% globally with publishers under 10,000 daily views down 60%. Model collapse proved under replacement and bounded under accumulation, a condition routinely dropped. Text detectors flagging non-native writers at 61.3% against 3%. Research fraud doubling every 1.5 years against 3.3 for retractions. Generated ad inventory scoring 77.2% viewability against 74.9% for clean supply. Human detection of synthetic video at 24.5% where a coin gets 50%. Four open source projects closing in one quarter. Provenance shipping as an ISO standard while one signing bug revoked every credential a camera line had issued. In every subject where detection was attempted it failed, and in every subject where friction was attempted it held. And the measurement quality tracked the same thing it tracked in the previous territory: whether anyone was obliged to publish.
---
Status: synthesis. No new factual claims. Every figure appears in one of the eight Territory 9 articles with its own sourcing and caveats, and each is linked where used. Source quality varies enormously across these subjects and is stated per row in the table below, because three of them have no disclosure regime of any kind.
---
The eight
| Subject | The measured thing | What it turned on | Evidence quality |
|---|---|---|---|
| Search referrals | Traffic down 33%, small publishers 60% | Zero-click reported 60% and 22.4% | No regime at all |
| Model collapse | Proved under replacement | Accumulation gives a bounded error | Peer-reviewed, contested |
| Text detection | 61.3% false positive, non-native | Perplexity is a shared property | Peer-reviewed, vendor-disputed |
| Research integrity | Fraud 1.5y, retraction 3.3y | Correction loses by construction | Strong, indexed, permanent |
| Ad inventory | 77.2% viewability vs 74.9% | Delivery metrics measured delivery | First measurement, days old |
| Human detection | 24.5% on synthetic video | Below chance means structured error | Meta-analysis, vendor-heavy |
| Maintainers | Valid rate 1 in 6 to 5% | Review costs an hour, submission seconds | Primary, maintainer-published |
| Provenance | ISO standard, credentials revoked | Signing outpaces verification | Standards plus vendor trackers |
Finding one: it is a cost ratio, not a quality problem
The instinct is that the problem is bad output. In seven of the eight subjects, output quality is either irrelevant or points the wrong way.
Generated ad inventory beat clean supply on viewability and invalid traffic and was classified as premium more than 70% of the time. It did not sneak past the quality checks. It won them.
Generated pull requests look correct. Coherent commit message, right files touched, a plausible problem described. The defects sit in the logic, which is only visible after reading, so every submission costs a review whether or not it is any good.
Fabricated vulnerability reports were confident and well-formatted, which is precisely what made them expensive: each had to be read, reproduced and refused in writing.
And the model collapse literature turns on a condition, not on quality. Under replacement the degradation is proved. Under accumulation, which is what the open web does, test error has a finite upper bound independent of iteration count. Same outputs, opposite conclusions, decided by whether the old data is kept.
What actually changed across all eight is a ratio. Producing a claim, a page, a report, a submission, a track, a paper, fell toward zero. Checking one did not move at all, because checking usually means reproducing the work described.
Every system in this territory was built when those two costs were comparable. None of them wrote the assumption down, which is why none of them noticed when it stopped holding.
Finding two: detection failed everywhere it was tried
This is the most actionable result in the territory and the one most likely to be ignored.
Text detection: 61.3% false positive rate on non-native writers against roughly 3% on native speakers, because perplexity-based detection measures predictability and second-language writing is predictable for reasons unrelated to machines. The developer of the most-cited model withdrew its own detector, disclosing 26% detection and 9% false flags.
Human detection: 24.5% on high-quality synthetic video against a coin's 50%. Warnings did not improve accuracy and did reduce trust in content generally, trading one failure for another.
Machine detection of synthetic images: works, at 97%, which is the exception and the reason the general claim needs care.
Content prevalence: unmeasurable, because every estimate rests on a detector whose error rate varies twentyfold with who wrote the text. A prevalence figure is a measurement of the instrument applied to an unknown population.
Adversarial pressure: paraphrasing cuts text detection by roughly 88%; real-world deepfake detection runs at about half its benchmark accuracy. Once a generator learns what a detector flags, the next version removes it.
The pattern is that detection is an arms race against a party with the advantage, and it puts its errors on identifiable people while the deliberate evader escapes.
Finding three: friction worked
In the subjects where anything held, the mechanism was raising the cost of submitting rather than classifying what was submitted.
curl now requires a reproducible test case, which an unverified report cannot supply and which costs nothing to someone who actually reproduced the bug.
Others require disclosure of assistance, participation in the issue thread before a pull request, or contribution history that takes time to accumulate.
Wikipedia's 2026 policy went to human reviewers trained on specific tells rather than automated tools, explicitly because automated detection flags too much legitimate writing.
And provenance is the same move at a different layer. It does not ask whether something looks real. It asks whether a chain verifies, which is a cheap question, and it puts the cost on whoever is making the claim.
The common property is that friction is priced to the honest case. A person who did the work already has the artefact the requirement demands. Detection asks a hard question about every item forever. Friction asks an easy question once, of the person best placed to answer it.
Finding four: the previous territory's finding held again
Territory 8 closed on the observation that the reliability of a figure tracked whether anyone was obliged to publish it, and nothing else.
It repeated here, in a domain where the extremes are further apart.
The best-measured subject was research integrity, because retractions are indexed by Crossref, catalogued by Retraction Watch and permanent in PubMed. A study could screen tens of millions of papers precisely because those papers are public and permanent.
The worst was search referrals, where the platform holds complete data on how many searches end without a click, is under no obligation to publish it, and does not. The headline statistic varies threefold with the same provider behind the extremes.
And the middle is populated by interested parties doing real work. The strongest per-query measurements, the most rigorous ad inventory sizing, the deepfake adoption trackers, are all published by organisations with a commercial position in the answer.
In none of those cases was the alternative a disinterested measurement. It was no measurement. The discipline is to state the interest and read accordingly, which is what this corpus has tried to do and cannot claim to have done perfectly.
Finding five: three things got measurably better
A synthesis that only reports damage is a different kind of dishonesty, and the record contains improvements.
Machine-assisted analysis found more than 100 real bugs in curl that years of fuzzing, compiler flags, static analysis and multiple human security audits had missed. Same project, same period, same class of tool as the flood. The difference was that the researcher verified before submitting.
Provenance moved from specification to ISO standard to default-on consumer hardware in under three years, which is fast for infrastructure of this kind.
And the accumulation result on model collapse is genuinely reassuring, with an analytic proof rather than only an empirical one, in the regime the open web actually occupies.
None of these cancels the rest. They establish that the technology is not the variable. Verified output is fine and unverified output at volume is the problem, which is a statement about process rather than about capability.
Finding six: the cost landed on whoever did not choose it
Worth separating, because it is the distributional question and no article in this territory addressed it directly.
In every subject, the party absorbing the cost is not the party that generated it.
The maintainer reads and refuses the fabricated report. The submitter paid nothing. curl's programme had run for six years and paid more than $100,000 across 87 confirmed vulnerabilities before the arithmetic broke, and the person who broke it was not the person who closed it.
The non-native writer carries the false accusation. The tool that produced the 61.3% false positive rate is bought by an institution, and the cost is borne by a student with a disciplinary process attached, which is error asymmetry in its most concentrated form in this corpus.
The small publisher absorbs the traffic loss. Sites under 10,000 daily page views fell 60% over two years while large brands retained direct, app and newsletter distribution. The decline is not evenly felt and the aggregate figure of 33% conceals exactly the cases that close sites.
The reviewer carries the paper volume. An estimated 100 million hours of unpaid reviewing labour in a single year, against submissions that doubled at one venue in one cycle.
The advertiser pays above clean supply for the 12% of generated inventory that falls outside existing frameworks, having been told by every quality metric that the inventory was premium.
The pattern is that cheap production externalises a verification cost onto a party with no ability to refuse it, and in five of these cases that party is unpaid, unrepresented or both.
Which reframes what the friction remedies are actually doing. They are not quality controls. They are attempts to return a cost to the party that created it, which is why they work and also why they feel unwelcoming: the cost is real and somebody has to carry it.
What would settle each subject
Naming the missing measurement is more useful than restating the uncertainty, and in most of these cases it is a single specific disclosure.
| Subject | What would settle it | Who holds it |
|---|---|---|
| Search referrals | Zero-click rate on a stated definition | The platform |
| Model collapse | Real-data fraction in current training sets | The labs |
| Text detection | Held-out uncontaminated evaluation across years | Nobody, retrospectively |
| Research integrity | Detection efficiency held constant across the window | Achievable by researchers |
| Ad inventory | Independent replication of the category definition | Achievable now |
| Human detection | Field study rather than isolated stimuli | Achievable now |
| Maintainers | Burden distribution below the visible projects | Achievable now |
| Provenance | Credential recovery rates through real pipelines | The platforms |
Four of these eight are achievable by researchers today with no cooperation from anyone. Independent replication of the ad inventory category, a field study of detection rather than a laboratory one, a survey of maintainer burden below the projects with profile, and a fixed-instrument recomputation of the fraud trend.
Three require disclosure from a party with no obligation to provide it, which is the previous territory's finding wearing different clothes.
And one is unrecoverable. No uncontaminated held-out evaluation exists retrospectively, because the benchmarks entered the training corpora before anyone thought to preserve a control.
That last row is the most instructive. It is the only question in the territory that could have been answered cheaply at the time and cannot be answered now at any price. The cost of not holding something out is invisible until the moment you need it.
What this corpus got wrong while writing it
A synthesis that grades other people's evidence should account for its own, and three errors in these eight articles are worth recording.
A concept was added twice under two names. Refutation cost was introduced as a node describing checking costing more than producing, and the glossary already contained verification asymmetry defined in the opposite direction, describing checking costing less than producing. Both are real and they describe different regimes, separated by whether a cheap machine check exists. The collision was found by the map check rather than by design, and it was found late.
Four nodes were asserted rather than earned. Each was added on the argument that an idea recurred across several articles, then linked from only the newest one, leaving the recurrence claim on the changes page and absent from the corpus. An audit found and fixed it, wiring twelve articles to the nodes their own cases had established, and the rule that a recurrence argument must be wired at the moment it is made was added afterwards rather than beforehand.
And a concept slug was mistyped twice, two articles apart. Both were caught by the build gate rather than shipped. The second one should not have happened, because a sweep of all posts after the first would have prevented it, and that sweep was run only after the second failure.
None of these changed a published claim. They are process failures rather than factual ones, and the reason for listing them is that a corpus arguing that verification is the load-bearing step should show its own.
What the territory does not show
That the internet is collapsing. Four open source projects closed against an estimated 1.4 million maintainers. Generated ad inventory is 1.3% to 2.4% of open web programmatic spend. These are real effects at small current magnitudes with fast growth rates, which is exactly the rate against level distinction and is why both the alarm and the dismissal are available from the same data.
That AI caused all of it. Paper mills predate language models. Maintainer burnout was a crisis a decade ago. Publisher traffic was already shifting to social and video. Cheaper production lowered a cost that was already being paid in several of these cases, and no study here isolates the contribution cleanly.
That the subjects are representative. These eight were chosen for having numbers, which selects for domains with disclosure regimes or motivated researchers and systematically excludes everything nobody has measured.
And that any of it is settled. Three of these subjects produced their first quantification within the last year, and one within days of being written about.
What is unresolved across all eight
Whether friction requirements survive. A reproducible test case is a real barrier today and is not obviously one in two years.
Whether prevalence ever becomes measurable. It currently requires a detector, detectors are unreliable in a directional way, and no method avoids the circularity.
Whether the platforms disclose. Search click data, training data composition and model mixtures are all held by parties with no obligation to publish, and the questions that matter most in this territory are exactly the ones they could answer.
Whether provenance chains survive real pipelines. Signing is solved. Preservation through transcoding, at scale, with recovery rates published, is not.
And what happens below the visible cases. curl and Wikipedia have profile and defenders. The typical small publisher, small project and individual maintainer has neither, and nobody is measuring what happens to them.
The counter-argument
One ratio explaining eight subjects is suspiciously tidy. A frame that fits everything has usually been fitted to everything, and these articles were selected independently over several weeks rather than chosen to demonstrate a thesis. The reader should weigh that the thesis arrived after the articles, which is the honest sequence and not proof of anything.
Detection is not uniformly failing. A convolutional network reached 97% on synthetic images where humans were at chance, and detection tooling caught 88% of generated ad inventory under an existing category. The claim that detection fails is really the claim that it fails for text and for adversarial cases, which is narrower and less quotable.
Friction has a cost nobody has counted here. Requirements that make bulk submission unprofitable also exclude newcomers, first-time contributors and people without accumulated history, who are precisely the population open systems exist to serve. This territory has praised friction without measuring what it turns away, and no article here attempted to.
The disclosure argument may be self-serving. A corpus that grades subjects by how well they are documented will conclude that documentation matters, and better disclosure is also the remedy most convenient for a site that works from public sources.
And the improvements section may be too generous. One researcher finding bugs with careful use does not offset a bounty programme closing, and treating them as two sides of a balance understates a real asymmetry in who absorbed the cost.
The short version
Eight subjects, one structure. Search referrals down 33% globally and 60% for the smallest publishers. Model collapse proved under replacement and bounded under accumulation. Text detectors at 61.3% false positive for non-native writers against 3% for native. Research fraud doubling every 1.5 years against 3.3 for retractions. Generated ad inventory at 77.2% viewability against 74.9% for clean supply. Human detection of synthetic video at 24.5% where a coin manages 50%. Four open source projects closed in a quarter. Provenance an ISO standard, and one signing bug revoking every credential a camera line had issued.
The damage is a cost ratio, not a quality problem. Generated ad inventory beat the quality checks. Generated pull requests look correct until read. Fabricated vulnerability reports were confident and well-formatted, which is what made them expensive. Producing a claim fell toward zero and checking one did not move, because checking usually means reproducing the work.
Detection failed in every subject where it was tried on text or against an adversary, put its errors on identifiable people, and was withdrawn by its own developer in one case. Machine detection of synthetic images is the exception, at 97% where humans sit at chance.
Friction worked. A reproducible test case, disclosure, prior participation, accumulated history, a verifiable chain. All of them price the requirement to the honest case, because someone who did the work already has the artefact.
And the measurement quality tracked obligation again, as it did in the previous territory. Retractions are indexed, permanent and public, so that subject has real numbers. Search click data sits with a platform under no obligation to publish it, so its headline statistic varies threefold.
Three things improved. More than 100 real bugs found in curl by machine-assisted analysis that fuzzing and multiple audits had missed. Provenance from specification to ISO standard to default-on hardware in under three years. And a proof that accumulation bounds model collapse in the regime the web actually occupies. The variable is not the technology. It is whether anyone verified before publishing.
Where this sits against the earlier territories
Four territories now, and the findings compound rather than repeat.
Territory 6 examined nine documented incidents and found that not one was fixed by a better model. Three contained no AI at all. The remedies were procedural: burden of proof, contestability, disclosure.
Territory 7 examined ten robotics domains and found the specification moved in every successful case. Difficulty did not predict feasibility; task shape did.
Territory 8 examined ten economic subjects and found the reliability of a figure tracked whether anyone was obliged to publish it, not how much the answer mattered.
And this territory finds that the damage came from a cost ratio rather than from output quality, with detection failing and friction holding.
The four have a common shape that none of them states alone. In every case the intuitive explanation, better models, harder tasks, more important questions, worse output, was the wrong one. The operative variable was structural each time: what the procedure allowed, what the specification permitted, what the disclosure required, what the verification cost.
That is either a real regularity or a house style. The honest position is that a corpus written by one person will find the patterns that person is equipped to see, and the falsification test is whether a subject arrives where the intuitive explanation turns out to be correct. Across thirty-nine articles in four territories it has not yet, which is worth stating as a warning rather than as a result.
Common questions
What is the single finding of this territory? That the damage came from a cost ratio inverting rather than from output quality. In seven of the eight subjects, quality is either irrelevant or points the wrong way: generated ad inventory beat clean supply on every metric used to catch bad inventory, generated pull requests look correct until read, and fabricated vulnerability reports were confident and well-formatted, which is precisely what made them expensive to refuse. What changed is that producing a claim fell toward zero while checking one did not move, because checking usually means reproducing the work described.
Why does that distinction matter practically? Because it changes the remedy. If the problem were quality, better classification would fix it. Since the problem is cost allocation, the intervention that works is moving the checking cost back to whoever is making the claim. That is why a reproducible test case works and a detector does not: the requirement is free to someone who genuinely reproduced the bug and expensive to someone who did not, and it never has to classify anything.
Did detection fail everywhere? For text and against adversaries, yes. Text detectors flagged non-native writers at 61.3% against roughly 3% for native speakers, the developer of the most-cited model withdrew its own detector at 26% detection and 9% false flags, and paraphrasing cuts detection by roughly 88%. Human detection of synthetic video runs at 24.5% against a coin's 50%. The clear exception is machine detection of synthetic images, where a convolutional network reached 97% on the same material where human accuracy was indistinguishable from chance.
What worked instead? Friction priced to the honest case. curl now requires a reproducible test case; other projects require disclosure of assistance, participation in an issue thread before a pull request, or contribution history that accumulates over time. Wikipedia's 2026 policy uses human reviewers trained on specific tells rather than automated tools, explicitly because automated detection flags too much legitimate writing. Provenance is the same move at a different layer, asking whether a chain verifies rather than whether something looks real.
How reliable are the numbers in this territory? They vary more than in any previous territory, and the table above states quality per subject for that reason. Research integrity has the best evidence because retractions are indexed by Crossref, catalogued publicly and permanent in PubMed. Search referrals have the worst, because the platform holds complete data, is under no obligation to publish it, and does not, so the headline statistic appears at 60%, 69% and 22.4% with the same data provider behind the extremes. Several subjects rest substantially on figures published by parties selling into the market they measure.
Is anything getting better? Three things, and reporting only the damage would be its own distortion. Machine-assisted analysis surfaced more than 100 real bugs in curl that years of fuzzing, static analysis and multiple human security audits had missed, in the same project and period as the bounty flood. Provenance moved from specification to formal ISO standard to default-on consumer hardware in under three years. And the accumulation result on model collapse is genuinely reassuring, with an analytic proof of a bounded error in the regime the open web actually occupies.
Does this mean the technology is the problem? No, and the curl case is the clearest evidence against that reading. The same class of tool produced both the flood that closed the bounty and the analysis that found more than 100 real bugs. The difference was one step: the researcher filtered output through his own expertise before submitting, and the bounty submissions did not. The variable is verification before publication, which is a statement about process rather than capability.
What is the strongest objection to this synthesis? That one ratio explaining eight subjects is suspiciously tidy, and a frame fitting everything has usually been fitted to everything. The honest defence is only that the articles were selected independently over several weeks and the thesis arrived afterwards, which is the right sequence and not proof. A second objection is more substantive: this territory has praised friction repeatedly without measuring what friction excludes, and requirements that make bulk submission unprofitable also raise the barrier for newcomers and first-time contributors, who are the population open systems exist to serve.
Sources
Primary documents only. Where a claim rests on a single report, the entry says so.
- The eight Territory 9 articles Artifipedia This is a synthesis and makes no new factual claims. Every figure appears in one of the eight articles with its own sourcing and caveats. Evidence quality varies more across this territory than any previous one, from indexed permanent retraction records to a headline statistic that appears at three different values from one provider, which is why the opening table states source quality per row.
Further reading
The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.
- Kleinberg et al. (2016), Inherent Trade-Offs in the Fair Determination of Risk Scores — why error rates cannot be equalised across groups while calibration holds. :: https://arxiv.org/abs/1609.05807 Error Asymmetry
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — alert burden against omission, where only one side leaves a record. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 Error Asymmetry
- International Energy Agency (2025), Energy and AI — per-unit efficiency improving while sector consumption rises. :: https://www.iea.org/reports/energy-and-ai Rate Against Level
- Epoch AI (2025), LLM inference prices have fallen rapidly but unequally across tasks — why a rate depends on which milestone is chosen. :: https://epoch.ai/data-insights/llm-inference-price-trends Rate Against Level
Related articles
- Fraud doubles in 18 months. Retraction takes 40.The scientific literature is the one place in this territory with a real record, and the record shows the correction machinery growing at less than half the rate of the thing it corrects.
- Slop scores premium 70% of the timeThe first rigorous measurement of machine-generated content in ad buying found it passes every quality check the industry uses, and passes them better than real inventory does.
- Generation takes seconds. Debunking takes hours.Four open source projects closed their doors in one month. The cause is not bad contributions but a cost ratio that inverted, and the same maintainer who shut his bounty credits the technology with finding 100 real bugs.
- One bug revoked every photo those cameras signedProvenance is the serious answer to synthetic media, it is now an ISO standard shipping in consumer hardware, and the gap between signing and verifying is wider than the adoption figures suggest.