We’re going to take one benchmark and stop looking at it as a pair of numbers. First we’ll see why a pair misleads, then we’ll build a small time series and read it.
Start with a pair that misled me. On dd-trace-go #4891, the benchmark bot’s comment now says BenchmarkOTLPProtoSize/1span is 7.6% to 8.0% faster than its baseline. Eleven days after #4891 was opened, on #4926, the bot said the same sub-benchmark was 8.1% to 8.5% slower. Both comments are correct arithmetic over two sets of runs, and the baselines differ: 05d9b8a on the first, 1e830f6 on the second. Same benchmark, opposite verdicts. 🔀
This is a companion to the Why Your Go Benchmarks Are Lying series, between part 4 and part 5. Part 5 tells the #4891 story in full; here we ask what it looks like over time. Disclosure: I work at Datadog, which maintains dd-trace-go.
A pair carries an assumption
The bot says what it does in its own comment: “This is an A/B test comparing a candidate commit’s performance against that of a baseline commit.” For a question about one change, that is fair, and A PR gate that actually fails builds one.
The pair quietly assumes the baseline is a fixed, neutral reference. It isn’t. main moves, and the binary’s layout and the machine’s mood move with it. Change the baseline and the same benchmark gets a different verdict, as above; on #4891 alone the sign changed between bot updates (part 5). A pair can’t show that, because it has one point on each side.
What does the same benchmark look like when we keep every night instead of comparing two? Let’s make one up.
Sixty nights of one benchmark
Everything in this section is synthetic: a Go program generated it, and no real benchmark did. The benchmark is a pretend 1000 ns/op function measured once a night, with 0.5% Gaussian noise. Three times, at nights 20, 35 and 50, a “PR” lands and adds 1.5%. On night 42 a noisy neighbour adds 8% for one night only. We’ll read it two ways: a 5% gate comparing each night with the one before, standing in for a PR check against one baseline run, and a change-point detector we’ll write after that.
| |
The first loop is the per-PR model. Here is the first half of the output, from Go 1.27.1:
| |
The gate fires twice, both times on the noisy neighbour: night 42 looks like a regression, night 43 like a recovery. It never fires on the three real steps, because 1.5% is well under its threshold, yet together they made the benchmark about 4.6% slower. The gate blames the innocent night and lets the guilty ones through. Part 4 said a change costing 1% per PR won’t trip a per-PR gate. Here are the numbers.
Reading the series
Change-point detection scans the whole history for places where the results shift, and stays quiet about everything else. MongoDB engineers moved to an E-Divisive means detector after threshold-based detection, and report that it “dramatically dropped our false positive rate”.
Our toy reader is much cruder and does not use E-Divisive. It does binary segmentation: find the night with the largest median shift, and if that beats a threshold, cut there and repeat on both halves. Here is the rest of the program:
| |
split tries every cut at least minLen nights from either end and keeps the strongest. The threshold of 6 is a knob I picked by eye for this data, and a real deployment has to set its own. The last part of the output:
| |
It finds three level changes and ignores the night-42 spike. The planted steps were at 20, 35 and 50, so it lands within two nights of each. On a real series the output is a list of candidate nights, and the next job is git log between them. Part 4 lists tools that ship change-point detection.
How the Go project reads its own
The Go wiki’s PerformanceMonitoring page states the principle plainly: “We never report performance numbers in isolation, and only relative to some baseline”, because “comparing performance data taken far apart in time, even on the same hardware, can result in a lot of noise that goes unaccounted for”. Before merge, SlowBot perf_vs_parent and perf_vs_tip run a pair on purpose. After merge, the performance dashboard gives “continuous monitoring of benchmark performance for every commit”, and because the tip-of-tree baseline “is always the latest overall release of Go”, “on every minor release of Go, the baseline shifts”. Read that next to our opening pair: the sign on #4891 flipped when main moved, and Go’s dashboard puts baseline changes where you can see them.
A ledger for alarms that weren’t
Bot comments will keep arriving, and some will be old friends. A mute regex silences a benchmark and forgets why. I’d add a ledger entry per dismissed alarm, in the repository, next to the benchmark. This one is illustrative, written from the public bot comments and part 5:
| |
seen turns one surprise into a count, cause says how sure we are, and check names the decisive test, so the next person doesn’t improvise one.
Back to the one benchmark
The pair told us −8% on one PR and +8% on another. Our sixty nights told us which night was a neighbour and which were the code. Keep the nightly series from part 4, use pairs for “did this change do it”, and write down the alarms that turned out to be nothing.
Versions, links and commands checked on 2 October 2026.
Try it
The two Go blocks above are one program: paste the second under the first in a main.go, run go mod init cp and go run .. I ran that with Go 1.27.1 and got the output shown. Change the seed or threshold and watch the detector miss a step or invent one. Then ask what the baseline was in your own bot’s last three comments.
A benchmark is a story told over time. Stop reading only the last sentence. 📈