Rebalance CodSpeed shards and build CPU benchmarks off metal - #9899
Conversation
Merging this PR will degrade performance by 19.88%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | WallTime | arrow_checked_add_u32_neon[16384] |
12.3 µs | 20.5 µs | -39.83% |
| ❌ | Simulation | random_i8[0.5] |
70.6 µs | 94.5 µs | -25.31% |
| ❌ | Simulation | decompress[u64, (4000, 1024)] |
72 µs | 86.7 µs | -16.97% |
| ❌ | WallTime | words_gather_scalar_avx2[65536] |
8.3 µs | 9.4 µs | -11.83% |
| ❌ | WallTime | mul_u32_nonnull_avx512 |
5.7 µs | 6.4 µs | -10.64% |
| ❌ | WallTime | mul_i32_nonnull_avx512 |
7.2 µs | 8 µs | -10.05% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing rk/codspeedbuild (f028bb0) with develop (313bc4f)
Footnotes
-
176 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
-
1 benchmark was run, but is now archived. If it was deleted in another branch, consider rebasing to remove it from the report. Instead if it was added back, click here to restore it. ↩
c9f0adc to
1039328
Compare
Signed-off-by: "Robert Kruszewski" <github@robertk.io> Signed-off-by: Robert Kruszewski <github@robertk.io>
Signed-off-by: "Robert Kruszewski" <github@robertk.io> Signed-off-by: Robert Kruszewski <github@robertk.io>
1039328 to
f028bb0
Compare
Just like macro benchmarks build binaries for codspeed on smaller machines. We also rebalance benchmark targets across machines to have more uniform run duration across all of them.
We avoid building all benchmark targets every time by enumerating the actual binaries we are benchmarking
First-run timings
The four array shards each take 75–100s to build and run, versus 436s for the previous single array shard in the baseline run. The slowest simulation shard is now 271s, but build times still vary substantially: storage and strings each take 189s to compile. This first run demonstrates improvement, not an evenly balanced steady state; cache state and changed package feature combinations can affect the comparison.
Combined metal job duration fell from 678s to 200s (about 70%). VM artifact uploads took another 30–47s, with downloads on metal taking 6–8s. These are observed job durations, not billing measurements.