Home/Blog/Evaluation & evidence/AI in journalism: 45% of news answers had a flaw
AI in journalism: 45% of news answers had a flawResponses with at least one significant issue. 3,000 evaluated.76%worst assistant45%all four, averageResponses with at least one significant issue. 3,000 evaluated.
Responses with at least one significant issue. 3,000 evaluated.

AI in journalism: 45% of news answers had a flaw

Twenty-two broadcasters in eighteen countries evaluated 3,000 AI answers about the news. Forty-five per cent carried a significant issue, and the worst performer failed on 76%.

TL;DR. Journalism is the one domain in this series where the subject matter experts audited the AI instead of deploying it. Twenty-two public service broadcasters across eighteen countries and fourteen languages evaluated more than 3,000 responses from four assistants. Forty-five per cent contained at least one significant issue. Thirty-one per cent had serious sourcing problems and 20% had major accuracy errors. The worst performer failed on 76% of responses. The finding that matters most is structural rather than statistical: a news organisation has a corrections policy and an AI assistant does not. When a newspaper is wrong, a correction is published and attached to the record permanently. When an assistant is wrong, the error is served once, to one person, and vanishes.

---

Every other article in this series asks how well AI performs a domain's work. This one is different, because the domain in question is the practice of establishing what is true, and it audited the AI rather than adopting it.

The European Broadcasting Union coordinated and the BBC led what is the largest study of its kind. Twenty-two public service media organisations across eighteen countries evaluated more than 3,000 responses from ChatGPT, Copilot, Gemini and Perplexity, posed in fourteen languages and scored by working journalists against five criteria: accuracy, sourcing, editorialisation, separating opinion from fact, and context.

Forty-five per cent of responses contained at least one significant issue.

Thirty-one per cent showed serious sourcing problems, meaning attributions that were missing, misleading or simply wrong. Twenty per cent contained major accuracy errors, including fabricated detail and outdated information presented as current.

One assistant failed on 76% of its responses, more than double the others, driven mainly by sourcing.

Two of the documented errors give the flavour better than the percentages. One assistant stated that surrogacy is illegal in the Czech Republic. Another named a pope who had died months earlier as the sitting pontiff.

The finding underneath the numbers

The percentages will improve. The structural point will not, and it is the reason this article exists.

A news organisation has a corrections policy. An AI assistant does not.

When a newspaper publishes something wrong, a specific machinery engages. The error is identified, usually by a reader or a subject. A correction is written and published. It is attached to the original article permanently. The organisation's rate of correction is visible, and a publication that corrects constantly acquires a reputation for it.

None of that exists for an assistant. A wrong answer is generated once, delivered to one person, and is gone. There is no record it was given. There is no mechanism for the subject of a false claim to have it corrected, because there is nothing to correct: the next person asking the same question receives a freshly generated answer that may be right, may be wrong differently, and bears no relationship to the first.

This is the deepest difference between a news organisation and a system that summarises news, and it has nothing to do with accuracy rates. A publication with a 5% error rate and a correction policy is more trustworthy than a system with a 2% error rate and none, because the first has a mechanism for becoming less wrong and the second does not.

An error rate is only meaningful alongside a correction mechanism, and this is the one domain where that distinction is the profession's own founding principle rather than an outside criticism.

The failure mode journalists named

The evaluators identified something more specific than inaccuracy, and it is the most useful observation in the study.

The assistants would not say they did not know.

Faced with a question where the honest answer was that the facts are unestablished, the systems produced an explanation instead. A journalist's craft in that situation is to state the limits of what is known, and the assistants filled the gap rather than marking it.

That is a different failure from hallucination and worse in a news context. A fabricated fact can be checked. An unmarked uncertainty cannot, because the reader has no signal that checking is required. The output is fluent, plausible and carries no indication that the underlying question is open.

It also explains why sourcing was the dominant failure at 31%. A system that will not decline to answer must attribute its answer to something, and if the attribution is not available it will be constructed. Missing attribution and fabricated attribution are the same failure viewed from two sides, and both follow from an unwillingness to return nothing.

Why this is the sharpest test in the series

Three features make this evidence stronger than anything else in these twelve articles.

The evaluators were domain experts with no stake in the outcome. Working journalists at public service broadcasters assessed responses about their own subject areas. They were not the vendor, not the buyer, and not compensated on the result.

The scoring criteria were specified in advance and are not accuracy alone. Sourcing, editorialisation, separating opinion from fact and context are all things a newsroom evaluates as a matter of routine. A field with a pre-existing professional standard applied that standard, rather than inventing a benchmark for the occasion.

And the scope defeats the usual objections. Fourteen languages and eighteen countries removes the argument that a finding reflects one market's phrasing or one language's data density. The study found the failure rate consistent across both.

Compare that with the other domains. Customer service is measured by parties who all benefit from the same answer. Medicine measures resemblance to a predicate. Education cannot blind its trials. Journalism produced the cleanest evaluation in the series because assessing whether a claim is true, sourced and in context is the profession itself.

One procedural detail is worth recording. At least one participating organisation had to temporarily stop blocking AI crawlers to allow its content to be evaluated, then restored the block afterwards. The study of whether assistants represent news accurately required news organisations to briefly permit the access they otherwise refuse, which is the whole commercial dispute compressed into a footnote.

The correction mechanism, generalised

The observation that an assistant has no corrections policy is the most portable finding in this series, and it applies well beyond news.

Every reliable institution has a route by which its errors come back to it.

A newspaper has corrections. A court has appeal. A bank has an examiner and a model revalidation cycle. A laboratory has replication. A hospital has morbidity review. These are not quality control in the ordinary sense. They are the mechanism by which an institution finds out it was wrong, and an institution without one does not become more accurate over time regardless of how accurate it starts.

Now apply that to a generative system, and three properties block the loop.

The error is not recorded. A response is generated, delivered and discarded. Unless someone screenshots it, there is no artifact to correct.

The error is not reproducible. Ask the same question again and you may get a correct answer, a different wrong answer, or the same one. That defeats the first step of any correction process, which is establishing that the error occurred.

And the error has no addressee. A person harmed by a false claim in a newspaper can contact the newspaper. A person harmed by a false claim generated for someone else, once, has nobody to contact and nothing to point at.

The practical consequence: for any deployed system, ask what happens to a wrong output after it is produced. If the answer is nothing, the system's error rate is a fixed property rather than a starting point, and every improvement will have to come from the model rather than from operation.

This is also why the deployments that work throughout this series share one feature. A licensed human between the model and the record is not only a safety control. It is the correction mechanism, because a person who catches an error can log it, and a logged error is one the institution can learn from.

What is actually deployed in newsrooms

Worth separating, because the study measures assistants summarising journalism, not journalism using AI.

Transcription and translation are established, uncontroversial and largely solved. A recorded interview transcribed automatically and checked by the reporter saves substantial time with a bounded failure mode.

Structured-data reporting predates the current wave. Sports results, earnings summaries and election returns have been generated from feeds for over a decade, and the reason it works is that the input is a table with known provenance.

Research assistance and first drafts are where practice varies most and disclosure varies with it. Some organisations publish detailed policies specifying what AI may touch and how it is labelled. Others have nothing.

And the disclosure gap is the live problem. There is no industry standard for labelling AI involvement, no agreement on what threshold requires disclosure, and no consistency in where a label appears. A reader cannot currently tell, from the artifact, what was involved in producing it.

The audience question

Adoption is smaller than the discussion implies and skewed by age.

Around 7% of online news consumers use AI assistants for news, rising to 15% among under-25s. That second figure is the one that matters for trajectory, and it is why broadcasters describe the finding as a democratic concern rather than a product complaint.

Roughly 54% of UK adults report worrying about AI's impact on journalism, which is a larger number than the usage figure and indicates the concern is not confined to people encountering the problem directly.

The mechanism people describe is specific. An assistant answers in the register of authority, without the signals a reader uses to calibrate. A newspaper carries a masthead, a byline, a dateline and a correction record. An assistant's answer carries none of those and reads with the same confidence regardless of whether the underlying claim is well-sourced or invented.

What is unresolved

Whether sourcing failures are fixable within the current approach. Retrieval grounding was supposed to solve attribution and the sourcing failure rate is still the largest category at 31%. Whether that improves substantially or reflects something harder about connecting a generated sentence to a specific source is unsettled.

Whether a correction mechanism is even possible. Nobody has proposed a workable design for correcting a generative system's error in a way that reaches the people who received it. The distribution model does not have a channel back.

What happens to the journalism being summarised. Broadcasters report traffic declining as assistants answer questions using their reporting without sending readers to it. If the summarising degrades the summarised, the error rate is a second-order problem behind an economic one.

And whether disclosure standards will converge. Every profession in this series eventually produced a norm. Journalism has not yet, and the absence is more consequential here because disclosure is the profession's existing answer to almost every other conflict of interest.

The counter-argument

The comparison class is not a perfect newspaper. Journalism has its own error rate, and it is not small. Studies of factual accuracy in published news have found error rates that would embarrass any newsroom asked about them directly. Measuring assistants against an idealised standard the profession does not itself meet is not a fair test, and the honest comparison is against a rushed human summary of an unfamiliar story.

The 45% figure counts issues, not falsehoods. A significant issue includes missing attribution, insufficient context and blurred opinion. Those matter and they are not the same as being wrong. The accuracy-specific figure is 20%, which is still high and is less than half the headline.

It is improving, and the study says so. The same consortium's earlier replication produced a rate five percentage points worse. A technology improving measurably between two audits months apart is behaving differently from one that has plateaued.

And the evaluators are interested parties in one specific sense. Public service broadcasters are in a commercial and existential dispute with the companies whose products they assessed, over traffic, licensing and access. That does not make the findings wrong, the methodology is strong and the criteria are the profession's own. It does mean the study was conducted by people who would not be displeased by the result, and that is worth stating plainly rather than leaving implicit.

The short version

Twenty-two public service media organisations across eighteen countries and fourteen languages evaluated more than 3,000 AI responses about the news, scored by working journalists against accuracy, sourcing, editorialisation, opinion-versus-fact and context.

Forty-five per cent contained at least one significant issue. Thirty-one per cent had serious sourcing problems, meaning missing, misleading or incorrect attribution. Twenty per cent contained major accuracy errors including fabricated detail and outdated information. The worst assistant failed on 76% of responses.

The evaluators named a failure more specific than inaccuracy: the systems would not say they did not know. Faced with an unsettled question they produced an explanation rather than marking the limit of what is established. That is worse than hallucination in a news context, because a fabricated fact can be checked and an unmarked uncertainty cannot. It also explains why sourcing dominated, since a system that will not decline to answer must attribute, and will construct an attribution if none exists.

But the structural finding outlasts every percentage. A news organisation has a corrections policy and an AI assistant does not. A published error is identified, corrected, and attached to the record permanently, and the organisation's rate of correction is visible. A wrong answer from an assistant is generated once, delivered to one person, and gone, with no record it was given and no mechanism for the subject of a false claim to have it fixed.

An error rate is only meaningful alongside a correction mechanism. A publication with a 5% error rate and a corrections policy is more trustworthy than a system with 2% and none, because the first has a way of becoming less wrong.

This is also the cleanest evaluation in the series, because the evaluators were domain experts with no commercial stake in the result, the criteria were the profession's own pre-existing standards rather than a benchmark invented for the occasion, and fourteen languages across eighteen countries removes the usual objection that a finding reflects one market.

Common questions

How accurate are AI assistants at summarising news? Not accurate enough for the purpose, on the best available evidence. A study coordinated by the European Broadcasting Union and led by the BBC had journalists at 22 public service organisations evaluate more than 3,000 responses in fourteen languages. Forty-five per cent contained at least one significant issue, 31% had serious sourcing problems and 20% had major accuracy errors. The worst-performing assistant failed on 76% of its responses.

What kinds of errors do AI news summaries make? Sourcing failures dominate at 31%, meaning attributions that are missing, misleading or wrong, including facts falsely attributed to specific publications. Accuracy errors account for 20% and include fabricated details and outdated information presented as current. Documented examples include stating that surrogacy is illegal in a country where it is not, and naming a pope who had died months earlier as the sitting pontiff.

Why can't AI assistants just say they don't know? That was the specific failure journalists identified, and it is more consequential than hallucination. Faced with a question where the facts are unestablished, the systems produced an explanation rather than marking the limit of what is known, which is the craft response. A fabricated fact can be checked; an unmarked uncertainty cannot, because the reader gets no signal that checking is needed. It also explains the sourcing failures, since a system that will not decline to answer must attribute, and will construct an attribution if none is available.

Do AI assistants have a corrections policy? No, and this is the deepest difference from a news organisation. When a publication is wrong, the error is identified, a correction is published and attached to the original permanently, and the organisation's correction rate is visible. When an assistant is wrong, the answer was generated once for one person and is gone. There is no record it was given, and no route for the subject of a false claim to have it corrected.

How many people get their news from AI? Around 7% of online news consumers use AI assistants for news, rising to about 15% among people under 25. The second figure is the one broadcasters point to, since it indicates trajectory rather than current scale. Separately, around 54% of UK adults report worrying about AI's impact on journalism, which is a substantially larger number than the usage figure.

What is AI actually used for inside newsrooms? Transcription and translation are established and largely uncontroversial, with bounded failure modes and a reporter checking the output. Structured-data reporting from feeds, such as sports results and earnings summaries, predates the current wave by more than a decade and works because the input is a table with known provenance. Research assistance and drafting are where practice varies most, and disclosure practice varies with it, since there is no industry standard for labelling AI involvement.

Is AI news accuracy getting better? Somewhat. The same consortium's earlier replication found a rate five percentage points worse than the current study, so measurable improvement occurred between two audits months apart. The rate remains high, sourcing remains the dominant failure category, and the improvement does not address the structural point that no correction mechanism exists regardless of the error rate.

Should I trust AI summaries of news events? Treat them as a pointer rather than a source. The measured failure rate on significant issues is 45%, the assistants will not signal when a question is unsettled, and attribution is the weakest area, so the citation offered may not support the claim. For anything consequential, the useful step is opening the cited article, which also verifies the attribution exists.

Learn the concepts

← All posts