Engineering · · 14 min

Benchmark CI That Doesn't Lie

Say we put a benchmark gate on every pull request: if performance regresses, the check fails. After a week, something strange shows up. A PR that does nothing performance-sensitive fails the gate. Another PR, one that rewrites a hot path, sails through with a green check. Nobody is imagining it. Both outcomes are correct given the data, and the data is just wrong. A shared CI runner is a multi-tenant machine, so our benchmark lands on a host that also runs other teams’ compile jobs, Docker builds and test suites. A real 10% regression can vanish into that noise, and a phantom one can appear from nowhere. The gate isn’t measuring our code. It’s measuring the lottery of what happens to be running next door. 🎰 ...

July 28, 2026 · 14 min · 2874 words · Kemal Akkoyun
Engineering · · 7 min

Context across goroutines and connections

We’re going to follow one request, GET /orders/42, and ask what “propagation” means at each hop. Three different jobs hide behind that word, the tools in this series do different ones, and the request’s own data falls into the gaps between them. The cheapest way to see a gap is a short Go program. This is a companion to the How to Instrument Go Without Changing a Single Line of Code series, between Part 4 and Part 5. The series accompanies my GopherCon UK talk. Part 1 showed tools hacking a field into g. Here we look at what that field holds, what it can’t, and who does the rest. ...

July 27, 2026 · 7 min · 1409 words · Kemal Akkoyun
Engineering · · 8 min

See it run: OBI and the eBPF profiler without Kubernetes

Parts 2 and 4 of this series describe what OBI (OpenTelemetry eBPF Instrumentation) and the OpenTelemetry eBPF profiler can do. Here we make them do it with Docker Compose and no Kubernetes: one Go server nobody touches, OBI watching its requests, the profiler watching its CPU. We’ll follow one request, GET /orders/42, and see it twice, first as a trace and then as CPU samples. This is a companion to the How to Instrument Go Without Changing a Single Line of Code series, between Part 4 and Part 5, and the hands-on side of my GopherCon UK talk. It all ran on one machine: a Colima VM (Ubuntu 24.04, kernel 6.8.0, linux/arm64, 2 CPUs, 2 GB) on an M4 Max laptop, not Docker Desktop. It’s one run, a demo and not a benchmark. ...

July 25, 2026 · 8 min · 1548 words · Kemal Akkoyun
Engineering · · 6 min

The laptop said 230%

Here is a pull request with two numbers for the same change. My machine said the multi scenario cost 230% over a plain build. CI said 81%. Both numbers were honest readings of a stopwatch. I think only one of them was measuring the pull request. We’ll work out which, and find the one question that would have told me before CI did. This is a companion to the Why Your Go Benchmarks Are Lying series, between Part 3 (the laptop) and Part 4 (CI). Those parts use microbenchmarks. This story is about a macrobenchmark, a whole build timed end to end, and it didn’t fit my GopherCon UK talk. Disclosure: I work at Datadog and I’m one of otelc’s maintainers. ...

July 24, 2026 · 6 min · 1134 words · Kemal Akkoyun
Engineering · · 7 min

The fourth signal: continuous profiling without code changes

Many tools that profile Go binaries resolve function names from the ELF symbol table or DWARF debug info, and production builds strip both. Point one of those profilers at a stripped binary and you get addresses, not names. The opentelemetry-ebpf-profiler doesn’t have this problem, and we’re going to find out why. Let’s start with a binary rather than a profiler. Here is the whole program, with //go:noinline so the compiler can’t fold checkout into main: ...

July 23, 2026 · 7 min · 1385 words · Kemal Akkoyun