Repository navigation
[perf] faster sparse bitpacked filter & take - #9723
lwwmanning wants to merge 2 commits into
Performance Gate Passed
⚠️ Unknown Walltime execution environment detected
Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.
For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.
⚠️ Different runtime environments detected
Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.
⚡ 28 improved benchmarks
✅ 2246 untouched benchmarks
⏩ 218 skipped benchmarks1
Performance Changes
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ⚡ | WallTime | threshold_neon[8] |
4.6 µs | 1.7 µs | ×2.7 |
| ⚡ | WallTime | threshold_avx2[8] |
3.4 µs | 1.7 µs | +94.7% |
| ⚡ | WallTime | threshold_neon[8] |
3.2 µs | 1.8 µs | +79.55% |
| ⚡ | WallTime | threshold_neon[160] |
8.5 µs | 4.9 µs | +71.94% |
| ⚡ | WallTime | threshold_avx512[8] |
3 µs | 1.8 µs | +70.39% |
| ⚡ | WallTime | threshold_neon[64] |
4.7 µs | 2.9 µs | +63.39% |
| ⚡ | WallTime | threshold_avx2[8] |
2.5 µs | 1.7 µs | +44.48% |
| ⚡ | WallTime | threshold_avx2[160] |
6.3 µs | 4.6 µs | +37.63% |
| ⚡ | WallTime | threshold_avx512[192] |
6.5 µs | 4.7 µs | +36.6% |
| ⚡ | WallTime | threshold_avx512[160] |
5.9 µs | 4.4 µs | +34.06% |
| ⚡ | WallTime | threshold_avx512[8] |
2.3 µs | 1.7 µs | +33.28% |
| ⚡ | WallTime | filter_neon[64] |
3.2 µs | 2.4 µs | +31.21% |
| ⚡ | WallTime | threshold_neon[192] |
9.2 µs | 7.1 µs | +30.43% |
| ⚡ | WallTime | threshold_neon[32] |
2.9 µs | 2.3 µs | +28.18% |
| ⚡ | WallTime | threshold_neon[8] |
2.2 µs | 1.7 µs | +27.7% |
| ⚡ | WallTime | threshold_avx2[192] |
7 µs | 5.5 µs | +27.15% |
| ⚡ | WallTime | filter_neon[160] |
5.7 µs | 4.5 µs | +26.45% |
| ⚡ | WallTime | threshold_neon[80] |
5.1 µs | 4.1 µs | +24.08% |
| ⚡ | WallTime | threshold_avx2[80] |
3.9 µs | 3.1 µs | +23.73% |
| ⚡ | WallTime | threshold_avx512[64] |
3.3 µs | 2.7 µs | +22.82% |
| ... | ... | ... | ... | ... | ... |
ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.
Tip
Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.
Comparing wm/fastlanes-sparse-extraction (895c184) with rk/fastlanes-sparse-benchmarks (f17daff)
Footnotes
-
218 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩