Part 4 ended with a promise: the PR gate blocks only when benchstat says the delta is real and bigger than a floor we chose. A fair reader question is how. benchstat prints a table, and I checked that it exits 0 even when the table says +199.00% (two made-up result files, the pinned version, Go 1.27.1). Who, then, turns that table into a red check?
We’ll build that job together. We’ll write one deliberately slow commit, let the job judge it, revert the commit, and let the job judge the revert. Two runs, two verdicts, both on a GitHub-hosted runner:
| |
The first line is the slow commit, and the second is its revert. This is a companion to the Why Your Go Benchmarks Are Lying series, between part 4 and part 5, the written companion of my GopherCon UK talk. GitHub deletes Actions logs after a retention period, so this post carries the key log lines itself. The benchgate tool is kakkoyun/benchlab at v0.1.0, and the workflow and both runs live on PR #5, a closed demo.
The slow commit
The patient is a small benchmark, BenchmarkInts, which sums a slice of integers. The slow commit routes every addition through a helper marked //go:noinline, so each total = add(total, x) becomes a real function call. I made it slow on purpose, and by a lot: roughly three times.
A PR against main has no benchmark on its base side, so the base binary couldn’t even build. The demo PR targets a branch that already holds the workflow and the benchmark, and the diff under test is the slow commit alone. First run, slow commit. Second run, I push the revert to the same PR.
Now the judge. Where should it run, and against what?
Why base and head share a job
Part 4 keeps the baseline from main in a cache, which suits a pinned runner. On a hosted runner, a cached baseline was measured on whatever VM ran the earlier job, and nothing promises that VM is the one judging your PR. This gate builds both binaries in one job and runs them on the same VM, taking turns: base, head, base, head. That interleaving (ABAB) spreads slow drift evenly over both sides. It does not remove noisy neighbours, and we’ll come back to that.
Here is the whole workflow. Read the comments, then we’ll walk through it step by step.
| |
Walking through it
The env block holds every number we might argue about later. ROUNDS is both the number of ABAB rounds and the sample count per side, so n=10. MIN_DELTA_PCT and CV_THRESHOLD are the two thresholds: a regression floor and a noise ceiling.
The checkout step uses fetch-depth: 2. On pull_request, GitHub checks out a merge commit, so HEAD^1 is the tip of the base branch and HEAD is what would land. Depth 2 reaches both. The build step refuses to continue if HEAD doesn’t have two parents, and it prints both short hashes, which is how the logs told me the comparison was the right one (base: 021b688 head: fb2ae9c in run 1). It then runs go test -c once per side, so we have two frozen test binaries and no rebuild between rounds.
The ABAB loop runs each binary ten times with -test.count=1, appending to base.txt and head.txt. The flags are a budget. Ten samples per side is the floor the benchstat docs recommend, and 200 ms is plenty for a benchmark that takes a few microseconds, so the whole job took about 35 seconds. Part 4’s nightly suite can afford -count=20 -benchtime=5s on controlled hardware, and this gate cannot afford that on every push.
The compare step is where the exit code gets made. It runs benchstat twice, once for the human-readable table and once with -format csv. The awk script reads rows shaped like name,base,CI,head,CI,delta,P, looks only at the sec/op table, and keeps a row only if its delta is a signed +N% with N at or above the floor. When benchstat sees noise, the delta column holds ~, so noise can never match. A faster result starts with a minus and is ignored too. The gate counts compared rows and writes the survivors to regressions.txt.
Then benchgate. At tag v0.1.0, -baseline only prints the benchstat output and the exit code never looks at it. The exit code answers a different question: is the sample stable enough to trust? It exits 1 when any benchmark’s coefficient of variation is above the threshold. That makes benchgate the noise gate and benchstat the regression gate. Also, benchgate runs its own go test on the head tree, one after another, so it measures how noisy the runner was, not the interleaved samples. The ± column in the benchstat table is the better noise indicator for the comparison itself.
Last, the verdict step folds everything into one exit code: 0 pass, 1 regression, 2 tool or setup error (nothing compared, benchgate failed, not a merge commit), 3 unstable sample. A regression wins over “unstable”, and each finding becomes an ::error annotation on the PR. Enough reading. Let’s watch it judge.
Run 1: the slow commit
Run 1 failed, and it failed in the right step. The benchstat step printed this (the log also names the CPU, an AMD EPYC 7763):
| |
Ints went from 2.579 µs to 7.700 µs, and the ± columns are small. Then the stability check:
| |
A CV of 3.1% is under the 10% ceiling, so the sample was trusted. The verdict step did the rest:
| |
The ##[error] line is the annotation from the awk output, and the last line is the red check. Was it a fluke? Let’s judge the revert.
Run 2: the revert
The revert makes the head tree identical to the base, so run 2 is a null test: there is nothing to find, and the gate should find nothing.
| |
~ means benchstat could not tell the sides apart, and benchgate reported cv= 3.4% ✓. The verdict step closed with:
| |
The gate went red for the slow commit and green for the revert, on the same kind of hosted runner, a minute apart.
What this does not prove
Part 4’s argument stands: ubuntu-latest is a shared machine, and I can’t turn off SMT or frequency scaling from inside it. These runs don’t show that the runner is quiet. They show that a 3x change is far outside whatever noise was there. I’d call a hosted gate weaker than one on bare metal, not worthless. It catches the obvious regressions, and it can’t promise to catch a 5–10% one. With n=10 on a shared runner, a change that small can flap: pass once, fail the next time. A/B is the wrong model for CI looks at what one pair per PR misses.
The CV numbers give a partial answer to a question readers asked, whether -cv-threshold 5 always fails on hosted runners. In these two runs it would not have (3.1% and 3.4%). I set 10, and two runs are a sample of two. If a noisy run does trip the ceiling, the job fails closed with exit 3, a red check that is not a regression. I saw that path only on my laptop, where noise tripped it. I never saw exit 2 on GitHub.
The runs are also a single benchmark on one machine type, captured on 4 October 2026. If you adopt this, calibrate BENCHTIME, ROUNDS and the two thresholds on your own benchmarks first. The gate is only worth what that calibration says.
Versions, links and commands checked on 2 October 2026; the CI runs shown were captured on 4 October 2026.
Try it
Copy the workflow into a repository with a benchmark package, point BENCH_PKG at it, and push a branch holding the workflow and the benchmark. Open a PR against that branch with a deliberately slow commit, watch the check go red, then push a revert and watch it go green. That is exactly what I did in the demo, and actionlint had no complaints about the file.
A red check only means something if we trust what it compares. The slow commit was easy to catch: at about 3x, nothing about the hardware could hide it. The dangerous deltas are small ones that look real, and a green or red check that is really measuring code layout. Part 5 opens with a CI regression that turned out to be a speedup. 🎟️