Home/Blog/Evaluation & evidence

Evaluation & evidence

12 articles

July 24, 2026 How to check an AI claim before you believe it A vendor advertised a hallucination rate below 0.001% with no evidence behind it, and a state attorney general investigated. Six questions separate a claim you can act on from a number… July 22, 2026 LLM-as-a-judge: when a model can grade another model GPT-4 agreed with human evaluators over 80% of the time, which is the rate humans agree with each other. A 2026 study… July 21, 2026 Testing a system that answers differently every time Forty percent of organisations hit significant quality regressions within ninety days of deploying an LLM application.… July 21, 2026 How to read an AI paper A meaningful fraction of state-of-the-art results, in the highest-prestige venues, could not be reproduced from the… July 20, 2026 What a model card should say and usually does not A model card was meant to be an evidentiary document: what a model was trained on, where it fails, how it was measured.… July 19, 2026 Why AI models get worse: forgetting, collapse, and drift Three different mechanisms quietly degrade AI systems, catastrophic forgetting, model collapse, and data drift. They get… July 18, 2026 Fixing the seed does not make it reproducible Training the same network fifty times with an identical seed produced almost as much variance as fifty different seeds.… July 16, 2026 Why machine learning does not do error bars Models differing only by random seed showed 0.057% variance in accuracy and 28.9% in certified robustness. The variance is… July 1, 2026 Why AI benchmarks mislead: contamination, gaming, saturation Every model launch leads with benchmark scores, and buyers read them like thermometer readings. They are closer to opinion… June 21, 2026 Your agent returned 200 OK and did the wrong thing Ordinary software tells you when it breaks. AI systems return well-formed, confident, wrong answers with a success status.… June 5, 2026 How do we measure AI progress? The benchmark problem Every AI model launches with a table of benchmark scores that look like objective proof of progress. A growing body of… May 28, 2026 Do AI detectors work? Accuracy, bias, and false positives Students are accused, applicants are rejected, and articles are dismissed on the strength of a percentage produced by an…

Other subjects