A/B is the wrong model for CI
We’re going to take one benchmark and stop looking at it as a pair of numbers. First we’ll see why a pair misleads, then we’ll build a small time series and read it. Start with a pair that misled me. On dd-trace-go #4891, the benchmark bot’s comment now says BenchmarkOTLPProtoSize/1span is 7.6% to 8.0% faster than its baseline. Eleven days after #4891 was opened, on #4926, the bot said the same sub-benchmark was 8.1% to 8.5% slower. Both comments are correct arithmetic over two sets of runs, and the baselines differ: 05d9b8a on the first, 1e830f6 on the second. Same benchmark, opposite verdicts. 🔀 ...