Evaluation & evidence
August 2, 2026
The EU delayed the part with no standards
Three AI Act obligations take effect today and the widely reported headline says the opposite. What was deferred, what was not, and why the split falls exactly where it does.
August 2, 2026
72 seconds, or 30 minutes, and both are trials
Territory 11 opens on the first subject this corpus has examined where the evidence is genuinely good. Registered trials,…
August 2, 2026
The 12% was non-inferior, and P was 0.41
MASAI is the best-evidenced AI deployment in medicine and the headline everyone quoted describes a result the trial did…
August 1, 2026
Refactoring fell from 25% to 3.8%
Survey evidence says AI improves code quality. Repository telemetry says the opposite. They are measuring different…
August 1, 2026
Phase I improved. Phase II did not.
AI-designed drugs clear safety trials at well above industry rates. At the stage that tests whether a drug works, the…
August 1, 2026
Adding the doctor to the model changed nothing
A randomised trial found physicians did better with an LLM than with conventional resources. Its second comparison,…
August 1, 2026
The pilot failed and the staff deployed it anyway
Enterprise AI is measured by what organisations sanctioned. A separate literature measures what their employees actually…
July 31, 2026
The pilot ran on data the production system will never see
Enterprise AI pilots are built on a curated slice, pre-cleaned, with limited users and manual review. Then production data…
July 31, 2026
The framework that exists produced 1.6%
The FDA has two AI tracks. One is final, has authorised over 1,350 devices, and is the regime under which almost none of…
July 31, 2026
One trial, a waitlist control, and a letter
The best evidence for AI mental health support is a single randomised trial of a purpose-built clinical tool. Its own…
July 31, 2026
The lock-in is the prompts, not the API
Switching costs used to require board approval to incur. AI switching costs accumulate through ordinary engineering…
July 30, 2026
621,000 robots installed, virtually no humanoids
Territory 7 opens on the question article 117 left unresolved. Industrial robotics is enormous and growing. The…
July 30, 2026
220 million miles, inside a boundary Waymo drew
The strongest safety evidence in physical autonomy, and the methodology that makes it honest is also what limits what it…
July 29, 2026
Robots weed ten million acres. They still cannot pick.
Agricultural robotics has a clean split between what scales and what does not, and the line falls exactly where the thesis…
July 29, 2026
Every chatbot query on earth is 2% of AI's power
Territory 8 opens on the numbers behind AI's physical footprint, and on the arithmetic in the IEA's own report that almost…
July 29, 2026
Ten domains, and the specification moved in every one
Territory 7 closes. Physical automation succeeded wherever the task could be changed, and the one place it has not is…
July 28, 2026
Same GPUs, same month, opposite depreciation
Two companies bought the same hardware, both were audited, and they reached opposite conclusions about how long it lasts.…
July 28, 2026
The best-documented humanoid deployment is seven units
Circulated figures run to tens of thousands. Company statements run to hundreds, and one chief executive said his robots…
July 27, 2026
A million deliveries, and the drone never lands
The most successful autonomous delivery system in the world solved the problem by removing its hardest part rather than by…
July 27, 2026
Every AI chip passes through one company's machines
The concentration risk in AI hardware is usually discussed as a country. It is more precisely a single firm, and that firm…
July 27, 2026
A million trajectories, 85% from four robots
Eight articles have shown automation succeeding by redefining the task. This is what it would take to succeed without…
July 26, 2026
Same company, same year, revenue figures 63% apart
The most quoted numbers in AI are run-rates from private companies, reported by outlets triangulating from leaks, and…
July 26, 2026
0.03% of construction spend, so the work moved indoors
On a building site nothing about the task is negotiable. The industry's answer was not a better robot but a different…
July 26, 2026
32 million vacuums, and 38% on household tasks
The most-deployed consumer robot succeeded by accepting a worse result more often. The home is the one place where the…
July 25, 2026
AI in science: 380,000 predicted, 736 actually made
AlphaFold won a Nobel Prize and did not reduce the rate of experimental structure determination. A materials model…
July 24, 2026
Testing a system that answers differently every time
Forty percent of organisations hit significant quality regressions within ninety days of deploying an LLM application.…
July 24, 2026
How to check an AI claim before you believe it
A vendor advertised a hallucination rate below 0.001% with no evidence behind it, and a state attorney general…
July 23, 2026
Fixing the seed does not make it reproducible
Training the same network fifty times with an identical seed produced almost as much variance as fifty different seeds.…
July 22, 2026
AI in software: 19% slower, and they felt 20% faster
A randomised trial put experienced developers 19% behind on their own repositories while they reported being 20% ahead.…
July 22, 2026
AI incidents: two registers, 1,460 and 14,530
The two main public AI incident databases count the same phenomenon and report 1,460 and 14,530. Neither is wrong. This is…
July 22, 2026
LLM-as-a-judge: when a model can grade another model
GPT-4 agreed with human evaluators over 80% of the time, which is the rate humans agree with each other. A 2026 study…
July 22, 2026
What a model card should say and usually does not
A model card was meant to be an evidentiary document: what a model was trained on, where it fails, how it was measured.…
July 21, 2026
AI support: 90% deflected, 40% actually resolved
A support system can report 90% deflection on a 40% resolution rate, because a customer who gets an unhelpful answer and…
July 21, 2026
AI in translation: the field that retired its own metric
Translation has fifty years of formal evaluation practice. In 2022 its own shared task published under the title "Stop…
July 21, 2026
How to read an AI paper
A meaningful fraction of state-of-the-art results, in the highest-prestige venues, could not be reproduced from the…
July 19, 2026
AI in journalism: 45% of news answers had a flaw
Twenty-two broadcasters in eighteen countries evaluated 3,000 AI answers about the news. Forty-five per cent carried a…
July 17, 2026
Superintelligence: the empirical record is zero
Sixty years after the intelligence explosion was described, no system has demonstrated sustained open-ended…
July 17, 2026
Why machine learning does not do error bars
Models differing only by random seed showed 0.057% variance in accuracy and 28.9% in certified robustness. The variance is…
July 15, 2026
Inference prices fell 9x a year. Also 900x.
The most cited number in AI economics is a single rate. The study behind it reports a range spanning two orders of…
July 14, 2026
The water bottle was per 10 to 50 responses
The most repeated environmental claim about AI dropped a qualifier from the study it cites. The concern it created is…
July 13, 2026
The 13% traveled. The authors' caveat did not.
A careful study found entry-level employment falling in AI-exposed jobs. Its own authors later narrowed when that becomes…
July 13, 2026
Three years in: no disruption, and one 20% hole
Every aggregate measure of AI's labour effect shows continuity. One within-firm comparison shows a fifth of a cohort gone.…
July 12, 2026
AI in education: students feel twice the gain they get
A 2026 review of 66 randomised trials found AI tutoring improved satisfaction by 0.93 and confidence by 0.91, against 0.53…
July 12, 2026
Why AI models get worse: forgetting, collapse, and drift
Three different mechanisms quietly degrade AI systems, catastrophic forgetting, model collapse, and data drift. They get…
July 11, 2026
A $3.5bn guarantee book and a $250bn commitment
The financing structure behind the AI buildout is disclosed, legal and defensible. It also routes several distinct-looking…
July 11, 2026
232 studies, and 1.3% recorded skin type
A systematic review found AI detecting skin cancer at 90% accuracy across 232 studies. Almost none of those studies…
July 10, 2026
The binding constraint is a transformer, not a chip
Capital is available and chips are shipping. The thing stopping data centres from opening is a waiting list held by…
July 9, 2026
1 to 2% of the chips, about 30% of the tokens
Export controls were designed to constrain compute and have. The measure that moved instead was usage, and the two have…
July 8, 2026
The grader rewards what the detector flags
AI graders assign higher scores to text with lower perplexity. AI detectors flag text with lower perplexity. The same…
July 8, 2026
Slop scores premium 70% of the time
The first rigorous measurement of machine-generated content in ad buying found it passes every quality check the industry…
July 8, 2026
The good numbers all came from obligations
Territory 8 closes. Across ten subjects the reliability of a figure tracked whether someone was legally required to…
July 7, 2026
Worse than chance means bias, not noise
People identify high-quality synthetic video 24.5% of the time. A coin would do better, and the reason it beats them is…
July 7, 2026
Zero-click is 60%, or 22.4%, from one provider
Territory 9 opens on what cheap generation does to information. Every method agrees the traffic is falling. None agrees on…
July 6, 2026
Generation takes seconds. Debunking takes hours.
Four open source projects closed their doors in one month. The cause is not bad contributions but a cost ratio that…
July 6, 2026
Collapse needs you to throw the old data away
The Nature result is correct under its stated condition. The condition is that each generation discards its predecessor's…
July 6, 2026
Twenty-four hours of silence is a billable resolution
Customer service AI has moved from per-seat to per-resolution pricing. The corpus already established that deflection and…
July 5, 2026
One bug revoked every photo those cameras signed
Provenance is the serious answer to synthetic media, it is now an ISO standard shipping in consumer hardware, and the gap…
July 5, 2026
Fraud doubles in 18 months. Retraction takes 40.
The scientific literature is the one place in this territory with a real record, and the record shows the correction…
July 4, 2026
Clinicians override 49% to 96% of alerts
Five years after the sepsis model was externally validated and found wanting, the sector-level evidence on clinical…
July 4, 2026
99.98% is a tracking rate, and the field is animation
NVIDIA's motion controller is a real advance that routes around the data problem Territory 7 identified. The number…
July 4, 2026
Eight subjects, one ratio, and it is not quality
Territory 9 closes. Across eight subjects the damage came from a cost ratio inverting rather than from bad output, and in…
July 3, 2026
The same PDF says 83% and nobody quotes it
Territory 10 opens on enterprise deployment. The most-quoted statistic in the field is real, measures something much…
July 3, 2026
4.7% at one attempt, 63% at a hundred
The previous article showed reliability decaying across repeated attempts. Security decays the same way with the sign…
July 2, 2026
61% once, 25% eight times running
Every agent benchmark score you have seen is a single-attempt number. The metric that measures whether an agent does the…
July 2, 2026
Lesson planning is 60 to 99% of the usage
Teachers report saving between 2.9 and 14 hours a week with AI, every figure is self-reported, and the platform data shows…
July 2, 2026
Good evidence, and eight different failures
Territory 11 closes. This was the field chosen because its evidence is strong, and every subject produced a failure that…
July 1, 2026
The flags land on lower prior attainment
Detection tools misclassify human writing in 10 to 20% of cases, and an analysis of 10,725 assessments found the flags…
July 1, 2026
Plus 48% with the tool, minus 17% without it
Territory 12 opens on education, where the randomised evidence is unusually good and points in opposite directions…
July 1, 2026
Why AI benchmarks mislead: contamination, gaming, saturation
Every model launch leads with benchmark scores, and buyers read them like thermometer readings. They are closer to opinion…
June 21, 2026
Your agent returned 200 OK and did the wrong thing
Ordinary software tells you when it breaks. AI systems return well-formed, confident, wrong answers with a success status.…
June 6, 2026
67% against 39%, and the gap is training
The AI education divide is usually described as access to tools. The measured gap is access to training, which is the…
June 5, 2026
How do we measure AI progress? The benchmark problem
Every AI model launches with a table of benchmark scores that look like objective proof of progress. A growing body of…
May 28, 2026
Do AI detectors work? Accuracy, bias, and false positives
Students are accused, applicants are rejected, and articles are dismissed on the strength of a percentage produced by an…
The other five subjects
Foundations
29
Theory, history and the parts that will still be true in twenty years.
Language & meaning
21
How language models handle language, and the places they do not.
Agents
11
Building agents, deploying them, and why most attempts do not reach production.
Building with AI
18
Retrieval, cost, infrastructure and the practical business of shipping something.
Safety & governance
30
Alignment, security, regulation and who is accountable when it goes wrong.