Optimize IB Index Lookup - #1916
danieljvickers wants to merge 8 commits into
Conversation
|
Claude Code Review Head SHA: b080887 Files changed:
Findings:
|
Its sandiego.yaml mechanism is the only non-h2o2 one in the suite, so it forces a second chemistry build of every target just for this case, which takes too long to compile. It stays in the example skip list.
Conflict in m_ibm.fpp: this branch moves s_get_neighborhood_idx and s_update_ib_lookup to m_ib_patches; master (MFlowCode#1914) added s_check_every_patch_marked between them. Kept the new check and this branch's removal.
|
Merged master into this branch ( |
|
@sbryngelson Would I be able to request that we not do upstream merges into draft PRs until they are opened as a full PR? Upstream branches that I am not aware of cause me to keep diverging from main |
Lines of Code
|
|
@danieljvickers yes good idea. sometimes it's me trying to save people trouble but i know what you mean |
Contribution Policy
The IBM weak scaling runs performed previously achieved 64% scaling efficiency on all of Frontier, with the major slow down coming from the IB ownership handoff. A closer look at this subroutine reveals that the issue is a growing cost to rebuilt the IB lookup array as the number of IBs approaches 100s of millions of entries.
This PR optimizes that subroutine by refactoring the IB lookup array.
Old Lookup Behavior
The lookup was a sparce array of length
num_gbl_ibswhich held null values except for a few entries where global IBs were being tracked by the local neighborhood rank. At the largest simulation of 500 million IBs, this is about 2 GB of integers where only a few kilobytes hold required information.New Behavior
I have replaced the old lookup with a set of key-value-pair arrays that are sorted by key. Because we expect few IBs to be crossing boundaries, I opted to do an insertion sort where new IBs are inserted into the previously sorted array of elements. This reduces the size to track particles to the order of 100s of kilobytes. The lookup time slows slightly from O(1) to O(log(num_ibs)), but that time is still very negligible compared to the performance gain, which reduces the array rebuilding from O(num_gbl_ibs) to O(log(num_ibs)).
Performance
I have run this PR on Frontier using the old weak scaling case. I skipped the final data point because the trend was clear. The refactor shows significant improvement in performance.