The One Number That Drifted Was the One That Changed: Why We Started Checksumming Our Own Benchmark Claims
Our MockMed comparison page said the baseline agent made 24 model calls per run. It made 13. All 20 baseline rows in the retained results record api_calls: 13. The 24 belonged to a single separate theme-drift run, folded into the headline number as if it were typical instead of an outlier from a different experiment entirely (openadapt-web#387). That correction landed August 26. It was not the last one that week. ...