The One Number That Drifted Was the One That Changed: Why We Started Checksumming Our Own Benchmark Claims

Our MockMed comparison page said the baseline agent made 24 model calls per run. It made 13. All 20 baseline rows in the retained results record api_calls: 13. The 24 belonged to a single separate theme-drift run, folded into the headline number as if it were typical instead of an outlier from a different experiment entirely (openadapt-web#387). That correction landed August 26. It was not the last one that week. ...

August 28, 2026 · 5 min · OpenAdapt Team

Get OpenAdapt posts by email

No spam. Unsubscribe anytime.