Evaluation & evidence
July 24, 2026
How to check an AI claim before you believe it
A vendor advertised a hallucination rate below 0.001% with no evidence behind it, and a state attorney general investigated. Six questions separate a claim you can act on from a number…
July 22, 2026
LLM-as-a-judge: when a model can grade another model
GPT-4 agreed with human evaluators over 80% of the time, which is the rate humans agree with each other. A 2026 study…
July 21, 2026
Testing a system that answers differently every time
Forty percent of organisations hit significant quality regressions within ninety days of deploying an LLM application.…
July 21, 2026
How to read an AI paper
A meaningful fraction of state-of-the-art results, in the highest-prestige venues, could not be reproduced from the…
July 20, 2026
What a model card should say and usually does not
A model card was meant to be an evidentiary document: what a model was trained on, where it fails, how it was measured.…
July 19, 2026
Why AI models get worse: forgetting, collapse, and drift
Three different mechanisms quietly degrade AI systems, catastrophic forgetting, model collapse, and data drift. They get…
July 18, 2026
Fixing the seed does not make it reproducible
Training the same network fifty times with an identical seed produced almost as much variance as fifty different seeds.…
July 16, 2026
Why machine learning does not do error bars
Models differing only by random seed showed 0.057% variance in accuracy and 28.9% in certified robustness. The variance is…
July 1, 2026
Why AI benchmarks mislead: contamination, gaming, saturation
Every model launch leads with benchmark scores, and buyers read them like thermometer readings. They are closer to opinion…
June 21, 2026
Your agent returned 200 OK and did the wrong thing
Ordinary software tells you when it breaks. AI systems return well-formed, confident, wrong answers with a success status.…
June 5, 2026
How do we measure AI progress? The benchmark problem
Every AI model launches with a table of benchmark scores that look like objective proof of progress. A growing body of…
May 28, 2026
Do AI detectors work? Accuracy, bias, and false positives
Students are accused, applicants are rejected, and articles are dismissed on the strength of a percentage produced by an…
Other subjects
Foundations
27
Theory, history and the parts that will still be true in twenty years.
Language & meaning
21
How language models handle language, and the places they do not.
Agents
11
Building agents, deploying them, and why most attempts do not reach production.
Building with AI
18
Retrieval, cost, infrastructure and the practical business of shipping something.
Safety & governance
11
Alignment, security, regulation and who is accountable when it goes wrong.