Skip to content

ci: Benchmark the walltime on CodSpeed - #1743

Draft
chfast wants to merge 4 commits into
masterfrom
ci/codspeed-walltime
Draft

chfast wants to merge 4 commits into
masterfrom
ci/codspeed-walltime

Conversation

@chfast

@chfast chfast commented Sep 28, 2026

Copy link
Copy Markdown
Member

The CodSpeed simulation counts the instructions, so it misses what only
shows in the time: the cache and branch prediction effects. This adds a
CircleCI job running the evmone and precompiles benchmarks with the
CodSpeed walltime instrument on a dedicated VM, which measures stably
enough without CodSpeed's own bare-metal runners.

🤖 Generated with Claude Code

The CodSpeed simulation counts the instructions, so it misses what only
shows in the time: the cache and branch prediction effects. This adds a
CircleCI job running the evmone and precompiles benchmarks with the
CodSpeed walltime instrument on a dedicated VM, which measures stably
enough without CodSpeed's own bare-metal runners.
@codspeed

codspeed Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will improve performance by 2.71%

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 1 improved benchmark
✅ 898 untouched benchmarks

Performance Changes

Benchmark BASE HEAD Efficiency
⚡ baseline/analyse/main/sha1_shifts 11.5 µs 11.2 µs +2.71%

Tip

Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.


Comparing ci/codspeed-walltime (f2227e0) with master (d9e925a)

Open in CodSpeed

@codecov

codecov Bot commented Sep 28, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 97.98%. Comparing base (d9e925a) to head (737bee3).

Additional details and impacted files
@@           Coverage Diff           @@
##           master    #1743   +/-   ##
=======================================
  Coverage   97.98%   97.98%           
=======================================
  Files         183      183           
  Lines       16856    16856           
  Branches     3856     3856           
=======================================
  Hits        16516    16516           
  Misses        250      250           
  Partials       90       90           
Flag Coverage Δ
eest-develop 81.77% <ø> (ø)
eest-develop-gmp 25.88% <ø> (ø)
eest-legacy 17.11% <ø> (ø)
eest-libsecp256k1 28.08% <ø> (ø)
eest-stable 81.77% <ø> (ø)
evmone-unittests 94.41% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

Components Coverage Δ
core 95.95% <ø> (ø)
tooling 94.29% <ø> (ø)
tests 99.81% <ø> (ø)
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@codspeed

codspeed Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Too many benchmarks in a single upload

The performance report could not be generated because there were too many benchmarks in a single upload to CodSpeed. We recommend sharding your benchmarks into smaller uploads, with max 1000 benchmarks per upload. See the documentation for more information.

@chfast

chfast commented Sep 28, 2026

Copy link
Copy Markdown
Member Author

Walltime trial on CodSpeed macro runners (run 36396100833)

The walltime jobs ran on codspeed-macro-arm64-graviton-ubuntu-22-04 (the free 600 min/month tier) and all passed. evmone test now reports the execution-specs tests' walltime too (5 rounds of ≥0.1 s each per test). But CodSpeed produced no report:

Too many benchmarks in a single upload … max 1000 benchmarks per upload

No single job uploads more than 429 benchmarks (instruction). The whole run has 1798: 899 per mode. Simulation alone (899) has always been accepted. So the limit seems to apply to the whole run, not to each job, although the docs don't say this. We need to curate the benchmarks to stay under 1000 in total (or get CodSpeed to raise the limit).

Cost per push

The 6 walltime jobs took ~56 min in total, so ~10 pushes/month fit in the free tier. Most of it is per-job setup:

Step (each walltime job) Time
Install GMP, CMake, GCC ~1.4 min
Configure ~1.5 min
Build ~2.9 min
Download fixtures ~0.7 min
Benchmark 0.7–8.3 min (16 min across all 6 jobs)

Building once (e.g. on the free ubuntu-22.04-arm runner) and running everything in one walltime job would bring this down to ~20 min per push.

Stability

This is the spread between the 5 repetitions of the internal benchmarks (129 benchmarks, stdev/mean):

median p90 max
Graviton macro runner 0.02% 0.71% 4.36%
CircleCI VM (ubuntu-2604, large) 0.04% 0.18% 1.48%

The Graviton tail comes from the interpreter execute benchmarks (memory_grow_mstore/by1 4.4%, sha1_shifts/5311 3.5%).

Other findings

  • Walltime and simulation must upload from the same workflow run. When they came from separate workflows (the earlier CircleCI job), the later report replaced the earlier one and marked the other mode's 899 benchmarks as skipped.
  • Only 5 repetitions come from CodSpeed's google/benchmark fork (walltime builds only). The 0.5 s min time is the google/benchmark default.

Before un-drafting

  • Curate the benchmarks to stay under 1000 in total.
  • Drop the "DROP: Disable CircleCI" commit and squash the rest.

Parked for now.

🤖 Generated with Claude Code

@chfast
chfast marked this pull request as draft September 28, 2026 08:44

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant