Engineering · · 8 min

A/B is the wrong model for CI

We’re going to take one benchmark and stop looking at it as a pair of numbers. First we’ll see why a pair misleads, then we’ll build a small time series and read it. Start with a pair that misled me. On dd-trace-go #4891, the benchmark bot’s comment now says BenchmarkOTLPProtoSize/1span is 7.6% to 8.0% faster than its baseline. Eleven days after #4891 was opened, on #4926, the bot said the same sub-benchmark was 8.1% to 8.5% slower. Both comments are correct arithmetic over two sets of runs, and the baselines differ: 05d9b8a on the first, 1e830f6 on the second. Same benchmark, opposite verdicts. 🔀 ...

August 1, 2026 · 8 min · 1488 words · Kemal Akkoyun
Engineering · · 11 min

A PR gate that actually fails

Part 4 ended with a promise: the PR gate blocks only when benchstat says the delta is real and bigger than a floor we chose. A fair reader question is how. benchstat prints a table, and I checked that it exits 0 even when the table says +199.00% (two made-up result files, the pinned version, Go 1.27.1). Who, then, turns that table into a red check? We’ll build that job together. We’ll write one deliberately slow commit, let the job judge it, revert the commit, and let the job judge the revert. Two runs, two verdicts, both on a GitHub-hosted runner: ...

July 31, 2026 · 11 min · 2277 words · Kemal Akkoyun
Engineering · · 14 min

Benchmark CI That Doesn't Lie

Say we put a benchmark gate on every pull request: if performance regresses, the check fails. After a week, something strange shows up. A PR that does nothing performance-sensitive fails the gate. Another PR, one that rewrites a hot path, sails through with a green check. Nobody is imagining it. Both outcomes are correct given the data, and the data is just wrong. A shared CI runner is a multi-tenant machine, so our benchmark lands on a host that also runs other teams’ compile jobs, Docker builds and test suites. A real 10% regression can vanish into that noise, and a phantom one can appear from nowhere. The gate isn’t measuring our code. It’s measuring the lottery of what happens to be running next door. 🎰 ...

July 28, 2026 · 14 min · 2874 words · Kemal Akkoyun
Engineering · · 6 min

The laptop said 230%

Here is a pull request with two numbers for the same change. My machine said the multi scenario cost 230% over a plain build. CI said 81%. Both numbers were honest readings of a stopwatch. I think only one of them was measuring the pull request. We’ll work out which, and find the one question that would have told me before CI did. This is a companion to the Why Your Go Benchmarks Are Lying series, between Part 3 (the laptop) and Part 4 (CI). Those parts use microbenchmarks. This story is about a macrobenchmark, a whole build timed end to end, and it didn’t fit my GopherCon UK talk. Disclosure: I work at Datadog and I’m one of otelc’s maintainers. ...

July 24, 2026 · 6 min · 1134 words · Kemal Akkoyun
Blogmentation · · 7 min

Making Drone Builds 10 Times Faster!

We open sourced drone-cache, a plugin for the popular Continuous Delivery platform Drone. It allows you to cache dependencies and interim files between builds to reduce your build times. This post explains why we are using Drone, why we needed a cache plugin, and what I learned while trying to release drone-cache as open source software. Read on for the story behind drone-cache or if you want to jump into action directly, go to the github.com/meltwater/drone-cache, and try it for yourself. ...

April 10, 2019 · 7 min · 1291 words · Kemal Akkoyun