Benchmark CI That Doesn't Lie
Say we put a benchmark gate on every pull request: if performance regresses, the check fails. After a week, something strange shows up. A PR that does nothing performance-sensitive fails the gate. Another PR, one that rewrites a hot path, sails through with a green check. Nobody is imagining it. Both outcomes are correct given the data, and the data is just wrong. A shared CI runner is a multi-tenant machine, so our benchmark lands on a host that also runs other teams’ compile jobs, Docker builds and test suites. A real 10% regression can vanish into that noise, and a phantom one can appear from nowhere. The gate isn’t measuring our code. It’s measuring the lottery of what happens to be running next door. 🎰 ...