We’re going to write one small benchmark, run it, and read every word of what comes back. By the end you’ll know where ns/op comes from, what -count and -benchtime change, what the -16 is, and how to ask the compiler and the profiler what the benchmark really did. Here is the line we’ll be decoding:
| |
This is a primer before part 1 of Why Your Go Benchmarks Are Lying, the written companion to my GopherCon UK talk, Why Your Go Benchmarks Are Lying. The series argues about numbers like the one above, so it helps to know how they are made first. If you have run go test -bench before, skim.
Write it, run it
Our benchmark allocates a 64-byte buffer and keeps it. It’s the one the series uses as its running suspect, and it lives in a scratch module (go mod init tiny, one file, tiny_test.go):
| |
A benchmark is a function named BenchmarkXxx that takes a *testing.B and does its work b.N times. The package-level sink is there so the compiler can’t throw the work away; part 1 is about what happens without it. Let’s run it:
| |
Every number in this post is from one run on one machine: an Apple M4 Max, Go 1.27.1, with a load average near 28 when I checked. A busy laptop is a fine teacher here, as you’ll see. I’m recapturing these outputs on a quiet machine with a reproduction kit, and I’ll update them when the new numbers are in.
Now the flags. -bench . selects benchmarks by regular expression, and by default none run. -run XXX does the same for tests; nothing is named XXX, so no test runs and the timing isn’t muddied by them. -benchmem adds the last two columns: bytes and heap allocations per iteration. The first number is how many times the loop ran, and the second, ns/op, is the time per iteration. We’ll start with the first.
Where does 99,632,604 come from?
Nobody chose it. The testing package runs the function once with b.N = 1, then predicts how many iterations would fill the benchmark time, adds 20% headroom, grows by at most 100× per step and stops at a billion. To watch it happen, I temporarily added a fmt.Fprintln(os.Stderr, "b.N =", b.N) to a copy of the benchmark:
| |
The function ran six times, and the printed line comes from the last run alone. The result is its total time divided by b.N, which is why ns/op is a mean. The ramp isn’t a fixed ladder: on the earlier run the last b.N was 99,632,604. The target is time, not a count, and -benchtime sets it. The default is 1s; a value like 100x means a fixed number of iterations.
Every -count repetition is its own ramp and its own mean:
| |
Same binary, same code, and the slowest sample is about 35% slower than the fastest. Each line is one mean, so each is one sample. Part 2 is about what to do with a handful of them.
Go is compiled ahead of time, so there is no JIT warmup for the ramp to wait out. One iteration still isn’t representative. With -benchtime=1x the same benchmark reported 2500, 1667 and 2000 ns/op. I haven’t chased down why (cold caches, the first allocation and timer granularity are my suspects).
That leaves the suffix. The -16 is GOMAXPROCS, and the suffix is dropped when the value is 1. With -cpu 1,4 we get both forms:
| |
A missing suffix in someone else’s output means they ran with GOMAXPROCS=1. We can read the whole line now. What we can’t tell yet is whether the loop did what we think it did.
The newer loop, and a suspicious zero
Go 1.24 added testing.B.Loop as a replacement for the b.N loop. We add a second benchmark that uses it, with no sink:
| |
Loop resets the timer on its first call and stops it when it returns false, so setup before it isn’t measured. The benchmark function runs once per -count, and the compiler keeps the arguments and results of calls in the body alive, so it can’t optimize away the whole loop body. (It has had bugs; part 1 lists them.) Go 1.26 changed how it keeps them alive: inlining in the body is no longer blocked. Here are both benchmarks together:
| |
Same function, yet the loop version reports about 3 ns against about 20, and zero allocations. The series spends a post on numbers like that. Let’s ask the compiler directly.
Ask the compiler
-gcflags=-m prints the compiler’s inlining and escape-analysis decisions. The code lives in a _test.go file, so it has to be go test, and -run XXX -bench XXX compiles everything without running anything. I kept the two lines about our buffer:
| |
Line 14 is the b.N benchmark: the buffer goes to the heap, hence 1 allocs/op. Line 21 is the b.Loop one: the buffer stays on the stack, so the zero is honest. -S goes one level lower and prints the assembly. The b.N benchmark calls the allocator (I trimmed the file path; the same grep on BenchmarkMakeBufferLoop STEXT prints nothing):
| |
That is enough compiler for now; part 1 reads these outputs in anger.
Where the profiler fits
A benchmark tells you how long; a profile tells you where. Our buffer is a poor subject, since in my run its profile was dominated by runtime frames, so let’s use the series demo’s hashing benchmark, BenchmarkHash_BLoop, from the companion repository. We add -cpuprofile to the run and read it with go tool pprof -top, keeping three rows:
| |
The benchmark spent 92.86% of its CPU in blockSHA2, the SHA-256 block function, which is what a hash benchmark should be doing. If the profile of your own benchmark is dominated by runtime and testing frames instead, you may be timing the test framework.
Versions, links and commands checked on 2 October 2026.
Try it
b.Loop needs Go 1.24, but the zero allocs/op result needs Go 1.26. On 1.24 and 1.25 the loop version shows 1 alloc/op too, which is the inlining change from earlier. Make a directory, run go mod init tiny, put the two benchmarks in tiny_test.go and run go test -run XXX -bench . -benchmem -count=3. Your numbers will differ from mine; on Go 1.26 or later your allocs/op column should not. Then add -benchtime=1x, -cpu 1,4 and -gcflags=-m one at a time and watch which column moves.
We wrote a dozen lines of Go, got a line of output, and know where every field came from: the ramp, the mean, the sample, the suffix and the allocator. What we can’t yet say is whether a given ns/op deserves trust. That’s the question the series asks, and part 1 starts by catching the compiler emptying a loop. 🔎