Engineering · · 11 min

Three Questions Before You Trust a Benchmark

A benchmark bot once told me that one of my pull requests made a benchmark 6–9% slower. A same-machine comparison said the pull request made it faster. Both results were stable, and they disagreed. In this last part we’ll find out how that happens, and leave with three questions and three small Go tools that tell us how far to trust a number. A loose cable Physicists have been fooled the same way, at a much larger scale. In September 2011, the OPERA collaboration announced that muon neutrinos appeared to travel faster than the speed of light. Months of rechecking found nothing wrong. The root cause, eventually, was an improperly seated fibre-optic connector in the GPS timing chain, which introduced a ~73 ns bias that made neutrinos appear to arrive early (Science called it a loose cable). A second fault, an oscillator defect, pushed the other way and partly masked the first. Once both were corrected, the 2012 re-measurements showed neutrino speed consistent with the speed of light. ...

August 4, 2026 · 11 min · 2271 words · Kemal Akkoyun
Engineering · · 14 min

Benchmark CI That Doesn't Lie

Say we put a benchmark gate on every pull request: if performance regresses, the check fails. After a week, something strange shows up. A PR that does nothing performance-sensitive fails the gate. Another PR, one that rewrites a hot path, sails through with a green check. Nobody is imagining it. Both outcomes are correct given the data, and the data is just wrong. A shared CI runner is a multi-tenant machine, so our benchmark lands on a host that also runs other teams’ compile jobs, Docker builds and test suites. A real 10% regression can vanish into that noise, and a phantom one can appear from nowhere. The gate isn’t measuring our code. It’s measuring the lottery of what happens to be running next door. 🎰 ...

July 28, 2026 · 14 min · 2874 words · Kemal Akkoyun
Engineering · · 12 min

Before CI: Can You Trust a Benchmark on Your Own Laptop?

Every benchmark number starts life on somebody’s laptop. We write the function, run it, like what we see and open the pull request, long before CI or a pinned runner gets a say. That makes it the least controlled measurement in the pipeline, and the one we act on. If it lies, everything downstream inherits the lie. Sending a noisy benchmark to CI doesn’t fix the noise. It industrialises it. ...

July 21, 2026 · 12 min · 2373 words · Kemal Akkoyun
Engineering · · 11 min

A Single Benchmark Number Is a Lie

Here is one benchmark that gave us two different answers for the same code. We’re going to take that run apart: see what benchstat can and can’t say about it, add the one number it doesn’t report, and decide how many runs it takes before a result earns any trust. Let’s call it the run that lied. The benchmark is BenchmarkMakeBuffer_Correct, run on a machine with sixteen CPU-bound background processes competing for the same cores. Here are two of its twenty runs (-16 is the GOMAXPROCS suffix), with the B/op and allocs/op columns dropped: ...

July 14, 2026 · 11 min · 2140 words · Kemal Akkoyun
Engineering · · 13 min

The Go Benchmark That Measured Nothing: Compiler Honesty in testing.B

We’re going to chase down one benchmark result that looks suspiciously good. The function under test calls make([]byte, 64) every time it runs, and the benchmark reports 0.34 ns/op, 0 B/op, 0 allocs/op. That is almost three billion heap allocations per second on a laptop, and the allocator counted none of them. (It comes from a real run, more on that below.) Keep it in your pocket; we’ll follow it the whole way. ...

July 7, 2026 · 13 min · 2713 words · Kemal Akkoyun