Home/Blog/Evaluation & evidence/Slop scores premium 70% of the time

Slop scores premium 70% of the time

The first rigorous measurement of machine-generated content in ad buying found it passes every quality check the industry uses, and passes them better than real inventory does.

TL;DR. TAG, the ANA and Fiducia published the first statistically rigorous sizing of machine-generated low-value inventory on 28 July 2026, four days before this article. It accounts for 1.3% to 2.4% of open web programmatic spend, against a made-for-advertising level of 1.1%. The finding worth reading twice is not the size. It is that the inventory scores better than clean supply on every metric the industry uses to detect bad inventory. Viewability 77.2% against 74.9%. Invalid traffic 0.05% against 0.32%. After measurability is applied it is classified as premium more than 70% of the time. 88% of it also registers as made-for-advertising, and the remaining 12% escapes existing frameworks entirely and costs more per verified impression than clean supply.

---

Status: established, and very recent. Primary source: the TAG TrustNet, ANA and Fiducia analysis released 28 July 2026 as part of the Q1 2026 ANA Programmatic Transparency Benchmark. These figures are days old and have not been independently replicated. The classification of what counts as this category is the analysis's own and is the load-bearing methodological choice.

---

The measurement

Between 1.3% and 2.4% of open web programmatic spend, described by the authors as the first statistically rigorous sizing of the category.

For comparison, made-for-advertising inventory sits at 1.1% of ANA member spend in the same quarter.

And that MFA figure moved the wrong way. It rose from 0.6% in Q4 2025 to 1.1% in Q1 2026, the first meaningful increase since the ANA began counting in 2023, when the original study found members spending 21% of budgets, around $13 billion annually, on such inventory. The ANA's own report names growing sub-types including machine-generated content as a reason.

So a seven-year cleanup got most of the way there and then reversed, and the reversal coincides with content becoming cheap to produce.

The part that inverts the expectation

This inventory does not evade the quality checks. It passes them, and passes them better than legitimate supply.

Viewability: 77.2% against 74.9% for clean inventory.

Invalid traffic: 0.05% against 0.32%. Six times cleaner.

And after measurability is factored in, it is classified as premium more than 70% of the time.

Every one of those metrics is doing exactly what it was designed to do. Viewability measures whether an ad was rendered in view. Invalid traffic measures bot activity. A page assembled by a machine, served fast, with a clean layout and no fraudulent traffic, legitimately scores well on all of it.

The metrics were built to catch fraud and this is not fraud. It is real inventory, on real pages, seen by real people, and worth very little to anybody.

Which is a construct problem, not a detection failure

The industry's quality stack measures delivery, not value.

Viewability, invalid traffic, brand safety and measurability all answer versions of the same question: did the impression happen as described? None of them asks whether the page was worth being on.

That gap was tolerable when producing a page cost something. The cost of publishing was a weak but real proxy for intent, and inventory that scored well on delivery was usually inventory somebody had built for a reason.

Generation removed the proxy. The cost went to near zero and the metrics kept measuring delivery, so the correlation they silently depended on stopped holding.

This is construct validity in its most expensive form to date in this corpus: a measurement that worked for a decade because of a relationship nobody wrote down, applied unchanged after the relationship broke.

The overlap, and the part that escapes

88% of this inventory also registers as made-for-advertising, which means existing tooling catches most of it under an older label. Common categories are the familiar content-farm subjects: recipes, personal finance, how-to.

The remaining 12% is the finding with teeth. It falls outside existing frameworks, escapes current tools, and costs more per verified impression than clean supply.

That is the worst combination available: undetected by the category everyone is already policing, and priced above the inventory it displaces.

And known publishers showed effectively zero of it, which is the counterweight worth stating: the problem is concentrated in the open exchange rather than distributed across the web.

Where the money actually goes

Wider benchmark data puts this in proportion. IAB Spain, citing ANA figures in April 2026, reported that only 41% of total programmatic investment reached genuine, measurable, viewable impressions free of invalid traffic and MFA inventory. The Q1 2026 figure is 43.3%.

So slightly more than four in ten dollars arrive as intended, and this category is a small and growing part of the remainder rather than the main cause.

The concentration finding matters more than the average. Lower-performing advertisers spent 2.1% of budgets on MFA against 0.9% for the highest-performing group. The same total looks very different depending on who is buying.

And the same pattern is visible elsewhere

Deezer reports machine-generated music at roughly 39% of daily uploads, around 60,000 tracks a day, with more than 13.4 million detected and tagged during 2025.

The mechanism is identical. A streaming platform pays out per stream, the cost of producing a track collapsed, and the payout system measures streams rather than worth. A share of those streams is fraudulent, which dilutes the royalty pool paid to everyone else.

Two industries, one structure: an automated payment system keyed to a delivery metric, meeting production costs that fell to nothing.

Three things this establishes

Passing a quality check is not evidence of quality when the check measures delivery. This inventory beat clean supply on viewability and invalid traffic. Anyone treating those scores as a quality signal is reading a delivery confirmation as a value judgement.

Cost was carrying more weight than anyone stated. The metrics worked because building a page was expensive enough to imply intent. That assumption was never written into the standard and it was doing most of the work.

And the residual is where the damage sits. 88% overlapping an existing category is a tooling problem with a known shape. The 12% outside every framework, priced above clean supply, is the part no current process touches.

What it does not establish

That the category is large. 1.3% to 2.4% of open web programmatic spend is small, and the open web is a fraction of total digital advertising.

That the classification is settled. What counts as this category is the analysis's own definition, published days ago, with no independent replication. The size estimate inherits that definition entirely.

That it caused the MFA reversal. The ANA names it as a contributing sub-type. It does not attribute the increase to it.

And nothing about any individual publisher or platform. Every figure here is aggregate.

What is unresolved

Whether the definition holds up. A first measurement of a newly named category is the least stable kind of figure, and this one is four days old.

Whether detection can distinguish generated from low-value. The analysis recommends advertisers ask verification partners whether their tools separate machine-generated content from this category in a nuanced and accurate manner, which implies most currently do not.

What a value metric would even look like. Nobody has proposed one that is measurable at programmatic speed, and delivery metrics survive partly because they are computable in milliseconds.

And whether the 12% grows. If tooling catches the 88% overlapping MFA, the pressure runs toward whatever the tooling does not catch.

The counter-argument

A first measurement of a self-defined category is weak evidence. The organisations publishing it also sell certification and benchmarking into this market, and a newly quantified threat is commercially useful to them. The figures may be sound and the incentive is real, which is the standard this corpus applies to every interested source.

The metrics are not broken. Viewability was never intended to measure editorial worth, and criticising it for not doing so misreads its purpose. A buyer who wanted quality inventory always had to make an editorial judgement, and the complaint is really that automation removed the judgement, not that the metric failed.

At 1.3% to 2.4% this may not warrant the attention. More than half of programmatic spend fails to reach a valid impression at all, and focusing on a two-percent category because it is novel is exactly the misallocation this corpus criticises elsewhere.

And the streaming comparison is loose. Music uploads and ad inventory differ in how payment is triggered, who bears the loss, and what fraud looks like, and calling them one structure imports more than the evidence supports.

The short version

The first rigorous sizing, published 28 July 2026, puts machine-generated low-value inventory at 1.3% to 2.4% of open web programmatic spend, against 1.1% for made-for-advertising, which itself rose from 0.6% to 1.1% in a quarter, the first increase since the industry began counting.

The size is not the finding. This inventory beats clean supply on every metric used to catch bad inventory: viewability 77.2% against 74.9%, invalid traffic 0.05% against 0.32%, and classified as premium more than 70% of the time.

Because the metrics measure delivery, not value. Viewability asks whether the ad rendered. Invalid traffic asks whether a bot was involved. A machine-assembled page, served fast, seen by a real person, legitimately passes all of it, and none of those checks asks whether the page was worth being on.

The stack worked for a decade on an assumption nobody wrote down: that producing a page cost enough to imply somebody meant it. Generation removed the cost and the metrics kept measuring delivery.

88% of it also registers as made-for-advertising, so existing tooling catches most of it under an older name. The remaining 12% falls outside every framework, escapes current tools, and costs more per verified impression than clean supply, which is the worst available combination.

And the same structure appears in music, where one platform reports machine-generated tracks at roughly 39% of daily uploads and a share of the resulting streams is fraudulent, diluting the pool paid to everyone else. A payment system keyed to a delivery metric, meeting a production cost that fell to zero.

Common questions

How much programmatic spend is this? Between 1.3% and 2.4% of open web programmatic spend, according to the TAG, ANA and Fiducia analysis published on 28 July 2026, which the authors describe as the first statistically rigorous sizing of the category. For comparison, made-for-advertising inventory sits at 1.1% of ANA member spend in the same quarter.

Why does it pass quality checks? Because the checks measure delivery rather than value. Viewability asks whether an ad rendered in view; invalid traffic asks whether bot activity was involved. A machine-assembled page, served quickly with a clean layout and real human visitors, legitimately scores well on both. The measured figures are viewability of 77.2% against 74.9% for clean inventory, invalid traffic of 0.05% against 0.32%, and classification as premium more than 70% of the time after measurability is applied.

So the metrics are broken? Not in the sense of malfunctioning, and this is the strongest objection to the framing. They do exactly what they were designed to do. The problem is that they worked as quality proxies for a decade because producing a page cost enough to imply somebody intended it, and that relationship was never written into the standard. Generation removed the cost and the metrics kept measuring delivery.

What is the 12% figure? 88% of this inventory also registers as made-for-advertising, so existing tooling catches most of it under an older label. The remaining 12% falls outside existing frameworks, escapes current tools, and costs more per verified impression than clean supply. That residual is the part no current process addresses, and it is priced above the inventory it displaces.

Did made-for-advertising spending really increase? Yes, for the first time since counting began. It went from 0.6% of ANA member spend in Q4 2025 to 1.1% in Q1 2026. The original 2023 study found members spending 21% of budgets, around $13 billion annually, on such inventory, so the long trend has been strongly downward. The ANA names growing sub-types including machine-generated content as a reason for the reversal without attributing the increase to it.

How much programmatic spend reaches a valid impression at all? About 43.3% in Q1 2026, up from a reported 41% cited earlier in the year and 36% in the ANA's December 2023 study. Slightly more than four in ten dollars arrive as genuine, measurable, viewable impressions free of invalid traffic and MFA inventory, which puts this category in proportion: it is a small and growing part of a much larger shortfall.

Does the same thing happen outside advertising? The clearest parallel is music streaming, where one platform reports machine-generated tracks at roughly 39% of daily uploads, around 60,000 a day, with more than 13.4 million detected and tagged during 2025. A share of the streams those tracks generate is fraudulent, which dilutes the royalty pool paid to everyone else. The structure is the same: an automated payment system keyed to a delivery metric, meeting a production cost that fell to near zero.

How much should this measurement be trusted? Carefully, because it is four days old and defines its own category. It is the first quantification of the thing, has not been independently replicated, and the size estimate depends entirely on the classification the authors chose. The organisations publishing it also sell certification and benchmarking into this market, which does not make the figures wrong and is the kind of interest a reader should weigh.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. TAG/ANA/Fiducia Analysis Quantifies Level of AI Slop in Digital Advertising Supply Chain for First Time Trustworthy Accountability Group, Association of National Advertisers and Fiducia, 28 July 2026 The primary figures: 1.3% to 2.4% of open web programmatic spend against 1.1% MFA, the 88% overlap, and the 12% that falls outside existing frameworks and costs more per verified impression. Published four days before this article and not independently replicated.
  2. AI slop exposes a blind spot in programmatic advertising and ad quality metrics EMARKETER The inversion: viewability of 77.2% against 74.9%, invalid traffic of 0.05% against 0.32%, and premium classification more than 70% of the time after measurability.
  3. MFA Ad Spend Is Increasing. Is AI Slop To Blame? AdExchanger The reversal from 0.6% to 1.1% between Q4 2025 and Q1 2026, the 2023 baseline of 21% of budgets, and the concentration among lower-performing advertisers at 2.1% against 0.9%.

Further reading

The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.

  • Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — why general benchmarks cannot carry general claims. :: https://arxiv.org/abs/2111.15366 Construct Validity
  • Bowman & Dahl (2021), What Will it Take to Fix Benchmarking in Natural Language Understanding? — what a benchmark must satisfy to support inference. :: https://arxiv.org/abs/2104.02145 Construct Validity

Learn the concepts

← All posts