Home/Blog/Evaluation & evidence

Evaluation & evidence

73 articles

August 2, 2026 The EU delayed the part with no standards Three AI Act obligations take effect today and the widely reported headline says the opposite. What was deferred, what was not, and why the split falls exactly where it does. August 2, 2026 72 seconds, or 30 minutes, and both are trials Territory 11 opens on the first subject this corpus has examined where the evidence is genuinely good. Registered trials,… August 2, 2026 The 12% was non-inferior, and P was 0.41 MASAI is the best-evidenced AI deployment in medicine and the headline everyone quoted describes a result the trial did… August 1, 2026 Refactoring fell from 25% to 3.8% Survey evidence says AI improves code quality. Repository telemetry says the opposite. They are measuring different… August 1, 2026 Phase I improved. Phase II did not. AI-designed drugs clear safety trials at well above industry rates. At the stage that tests whether a drug works, the… August 1, 2026 Adding the doctor to the model changed nothing A randomised trial found physicians did better with an LLM than with conventional resources. Its second comparison,… August 1, 2026 The pilot failed and the staff deployed it anyway Enterprise AI is measured by what organisations sanctioned. A separate literature measures what their employees actually… July 31, 2026 The pilot ran on data the production system will never see Enterprise AI pilots are built on a curated slice, pre-cleaned, with limited users and manual review. Then production data… July 31, 2026 The framework that exists produced 1.6% The FDA has two AI tracks. One is final, has authorised over 1,350 devices, and is the regime under which almost none of… July 31, 2026 One trial, a waitlist control, and a letter The best evidence for AI mental health support is a single randomised trial of a purpose-built clinical tool. Its own… July 31, 2026 The lock-in is the prompts, not the API Switching costs used to require board approval to incur. AI switching costs accumulate through ordinary engineering… July 30, 2026 621,000 robots installed, virtually no humanoids Territory 7 opens on the question article 117 left unresolved. Industrial robotics is enormous and growing. The… July 30, 2026 220 million miles, inside a boundary Waymo drew The strongest safety evidence in physical autonomy, and the methodology that makes it honest is also what limits what it… July 29, 2026 Robots weed ten million acres. They still cannot pick. Agricultural robotics has a clean split between what scales and what does not, and the line falls exactly where the thesis… July 29, 2026 Every chatbot query on earth is 2% of AI's power Territory 8 opens on the numbers behind AI's physical footprint, and on the arithmetic in the IEA's own report that almost… July 29, 2026 Ten domains, and the specification moved in every one Territory 7 closes. Physical automation succeeded wherever the task could be changed, and the one place it has not is… July 28, 2026 Same GPUs, same month, opposite depreciation Two companies bought the same hardware, both were audited, and they reached opposite conclusions about how long it lasts.… July 28, 2026 The best-documented humanoid deployment is seven units Circulated figures run to tens of thousands. Company statements run to hundreds, and one chief executive said his robots… July 27, 2026 A million deliveries, and the drone never lands The most successful autonomous delivery system in the world solved the problem by removing its hardest part rather than by… July 27, 2026 Every AI chip passes through one company's machines The concentration risk in AI hardware is usually discussed as a country. It is more precisely a single firm, and that firm… July 27, 2026 A million trajectories, 85% from four robots Eight articles have shown automation succeeding by redefining the task. This is what it would take to succeed without… July 26, 2026 Same company, same year, revenue figures 63% apart The most quoted numbers in AI are run-rates from private companies, reported by outlets triangulating from leaks, and… July 26, 2026 0.03% of construction spend, so the work moved indoors On a building site nothing about the task is negotiable. The industry's answer was not a better robot but a different… July 26, 2026 32 million vacuums, and 38% on household tasks The most-deployed consumer robot succeeded by accepting a worse result more often. The home is the one place where the… July 25, 2026 AI in science: 380,000 predicted, 736 actually made AlphaFold won a Nobel Prize and did not reduce the rate of experimental structure determination. A materials model… July 24, 2026 Testing a system that answers differently every time Forty percent of organisations hit significant quality regressions within ninety days of deploying an LLM application.… July 24, 2026 How to check an AI claim before you believe it A vendor advertised a hallucination rate below 0.001% with no evidence behind it, and a state attorney general… July 23, 2026 Fixing the seed does not make it reproducible Training the same network fifty times with an identical seed produced almost as much variance as fifty different seeds.… July 22, 2026 AI in software: 19% slower, and they felt 20% faster A randomised trial put experienced developers 19% behind on their own repositories while they reported being 20% ahead.… July 22, 2026 AI incidents: two registers, 1,460 and 14,530 The two main public AI incident databases count the same phenomenon and report 1,460 and 14,530. Neither is wrong. This is… July 22, 2026 LLM-as-a-judge: when a model can grade another model GPT-4 agreed with human evaluators over 80% of the time, which is the rate humans agree with each other. A 2026 study… July 22, 2026 What a model card should say and usually does not A model card was meant to be an evidentiary document: what a model was trained on, where it fails, how it was measured.… July 21, 2026 AI support: 90% deflected, 40% actually resolved A support system can report 90% deflection on a 40% resolution rate, because a customer who gets an unhelpful answer and… July 21, 2026 AI in translation: the field that retired its own metric Translation has fifty years of formal evaluation practice. In 2022 its own shared task published under the title "Stop… July 21, 2026 How to read an AI paper A meaningful fraction of state-of-the-art results, in the highest-prestige venues, could not be reproduced from the… July 19, 2026 AI in journalism: 45% of news answers had a flaw Twenty-two broadcasters in eighteen countries evaluated 3,000 AI answers about the news. Forty-five per cent carried a… July 17, 2026 Superintelligence: the empirical record is zero Sixty years after the intelligence explosion was described, no system has demonstrated sustained open-ended… July 17, 2026 Why machine learning does not do error bars Models differing only by random seed showed 0.057% variance in accuracy and 28.9% in certified robustness. The variance is… July 15, 2026 Inference prices fell 9x a year. Also 900x. The most cited number in AI economics is a single rate. The study behind it reports a range spanning two orders of… July 14, 2026 The water bottle was per 10 to 50 responses The most repeated environmental claim about AI dropped a qualifier from the study it cites. The concern it created is… July 13, 2026 The 13% traveled. The authors' caveat did not. A careful study found entry-level employment falling in AI-exposed jobs. Its own authors later narrowed when that becomes… July 13, 2026 Three years in: no disruption, and one 20% hole Every aggregate measure of AI's labour effect shows continuity. One within-firm comparison shows a fifth of a cohort gone.… July 12, 2026 AI in education: students feel twice the gain they get A 2026 review of 66 randomised trials found AI tutoring improved satisfaction by 0.93 and confidence by 0.91, against 0.53… July 12, 2026 Why AI models get worse: forgetting, collapse, and drift Three different mechanisms quietly degrade AI systems, catastrophic forgetting, model collapse, and data drift. They get… July 11, 2026 A $3.5bn guarantee book and a $250bn commitment The financing structure behind the AI buildout is disclosed, legal and defensible. It also routes several distinct-looking… July 11, 2026 232 studies, and 1.3% recorded skin type A systematic review found AI detecting skin cancer at 90% accuracy across 232 studies. Almost none of those studies… July 10, 2026 The binding constraint is a transformer, not a chip Capital is available and chips are shipping. The thing stopping data centres from opening is a waiting list held by… July 9, 2026 1 to 2% of the chips, about 30% of the tokens Export controls were designed to constrain compute and have. The measure that moved instead was usage, and the two have… July 8, 2026 The grader rewards what the detector flags AI graders assign higher scores to text with lower perplexity. AI detectors flag text with lower perplexity. The same… July 8, 2026 Slop scores premium 70% of the time The first rigorous measurement of machine-generated content in ad buying found it passes every quality check the industry… July 8, 2026 The good numbers all came from obligations Territory 8 closes. Across ten subjects the reliability of a figure tracked whether someone was legally required to… July 7, 2026 Worse than chance means bias, not noise People identify high-quality synthetic video 24.5% of the time. A coin would do better, and the reason it beats them is… July 7, 2026 Zero-click is 60%, or 22.4%, from one provider Territory 9 opens on what cheap generation does to information. Every method agrees the traffic is falling. None agrees on… July 6, 2026 Generation takes seconds. Debunking takes hours. Four open source projects closed their doors in one month. The cause is not bad contributions but a cost ratio that… July 6, 2026 Collapse needs you to throw the old data away The Nature result is correct under its stated condition. The condition is that each generation discards its predecessor's… July 6, 2026 Twenty-four hours of silence is a billable resolution Customer service AI has moved from per-seat to per-resolution pricing. The corpus already established that deflection and… July 5, 2026 One bug revoked every photo those cameras signed Provenance is the serious answer to synthetic media, it is now an ISO standard shipping in consumer hardware, and the gap… July 5, 2026 Fraud doubles in 18 months. Retraction takes 40. The scientific literature is the one place in this territory with a real record, and the record shows the correction… July 4, 2026 Clinicians override 49% to 96% of alerts Five years after the sepsis model was externally validated and found wanting, the sector-level evidence on clinical… July 4, 2026 99.98% is a tracking rate, and the field is animation NVIDIA's motion controller is a real advance that routes around the data problem Territory 7 identified. The number… July 4, 2026 Eight subjects, one ratio, and it is not quality Territory 9 closes. Across eight subjects the damage came from a cost ratio inverting rather than from bad output, and in… July 3, 2026 The same PDF says 83% and nobody quotes it Territory 10 opens on enterprise deployment. The most-quoted statistic in the field is real, measures something much… July 3, 2026 4.7% at one attempt, 63% at a hundred The previous article showed reliability decaying across repeated attempts. Security decays the same way with the sign… July 2, 2026 61% once, 25% eight times running Every agent benchmark score you have seen is a single-attempt number. The metric that measures whether an agent does the… July 2, 2026 Lesson planning is 60 to 99% of the usage Teachers report saving between 2.9 and 14 hours a week with AI, every figure is self-reported, and the platform data shows… July 2, 2026 Good evidence, and eight different failures Territory 11 closes. This was the field chosen because its evidence is strong, and every subject produced a failure that… July 1, 2026 The flags land on lower prior attainment Detection tools misclassify human writing in 10 to 20% of cases, and an analysis of 10,725 assessments found the flags… July 1, 2026 Plus 48% with the tool, minus 17% without it Territory 12 opens on education, where the randomised evidence is unusually good and points in opposite directions… July 1, 2026 Why AI benchmarks mislead: contamination, gaming, saturation Every model launch leads with benchmark scores, and buyers read them like thermometer readings. They are closer to opinion… June 21, 2026 Your agent returned 200 OK and did the wrong thing Ordinary software tells you when it breaks. AI systems return well-formed, confident, wrong answers with a success status.… June 6, 2026 67% against 39%, and the gap is training The AI education divide is usually described as access to tools. The measured gap is access to training, which is the… June 5, 2026 How do we measure AI progress? The benchmark problem Every AI model launches with a table of benchmark scores that look like objective proof of progress. A growing body of… May 28, 2026 Do AI detectors work? Accuracy, bias, and false positives Students are accused, applicants are rejected, and articles are dismissed on the strength of a percentage produced by an…

The other five subjects