Repository navigation
Default allocator speedup - #6105
Open
thrasibule wants to merge 5 commits into
Open
thrasibule wants to merge 5 commits into
thrasibule wants to merge 5 commits into
Conversation
When more buffers are in use at once than the main table holds (2 * NUM_THREADS * NUM_PARALLEL), blas_memory_alloc() falls back to an overflow area of NEW_BUFFERS slots. Three problems there: - Unlike the main table, an overflow slot mapped a new buffer on every allocation instead of reusing the one it already had, leaking 128 MB of address space per call. Every mapping is also recorded for release at shutdown, and after NUM_BUFFERS + NEW_BUFFERS mappings that record was written past the end of new_release_info: the program crashed. Map a slot's buffer only once, as the main table does. - A thread that found the main table full while another thread was creating the overflow area then saw memory_overflowed set and terminated with "too many memory regions" and a NULL buffer, without looking in the overflow area. It now searches it first. memory_overflowed is also only set once the overflow arrays exist. - In OpenMP builds the error path took no lock, so several threads that found the main table full at the same time each created a new overflow area, replacing one another's. Threads still holding slots in a replaced array then crashed or hung. Take alloc_lock on that path in every build, and make it the only way into the overflow area: the search that followed the scan of the main table duplicated the one added above, and goes (in OpenMP builds it ran without alloc_lock, locking each overflow slot instead). With alloc_lock covering both creating the area and claiming its slots, no build locks individual overflow slots any more. The path's label, error:, becomes overflow:, as that is all it handles. The stress test has 40 threads holding up to 3 buffers each against a 50-slot table (MAX_CPU_NUMBER=2). Before, it crashed in every run, and in OpenMP builds the overflow area was created up to 21 times. It now passes 20 runs out of 20 in both builds, with the area created once and every buffer owned by one thread at a time. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
In pthreads builds, blas_memory_alloc() took the global alloc_lock while scanning the table for a free slot, and blas_memory_free() took it again while searching for the buffer, so every BLAS call that needs a buffer serialized all calling threads twice. OpenMP builds already claim a slot by locking only that slot; use the same scheme in all builds. Freeing a buffer from the main table needs no lock at all, since only its owner frees it. The rarely used overflow area keeps alloc_lock. With the scan of the main table lock-free, a full table goes straight to overflow:, which takes alloc_lock itself. In blas_memory_free(), the branch that cleared a main-table slot under alloc_lock becomes unreachable and goes, as does a DEBUG-only check that read memory[] past its end. Many threads each calling dsymv (n=15) with 1 BLAS thread, 8-core Zen+, M calls/s in total: 1 thread 5.7 -> 6.3, 4 threads 2.2 -> 7.3, 16 threads 1.5 -> 5.5. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every buffer search started at slot 0, so all threads calling BLAS
probed, and locked, the same few slots and cache lines, and each one
first stepped over the thread pool's buffers, which are allocated at
startup and sit at the start of the table.
Remember in a thread-local variable the slot each thread used last and
try it first, both when allocating and when freeing. A thread that keeps
calling BLAS keeps reusing its own slot and touches nothing shared.
When that slot is busy, the search falls back to the lowest free slot as
before, so threads that take turns still share a buffer and the number
of buffers in use still follows the number of concurrent callers. The
pool's buffers (procpos 2) are taken from the top of the table. Without
compiler thread-local storage, searches start at slot 0 as before.
Many threads each calling dsymv (n=15) or dtrsv (n=20) with 1 BLAS
thread, 8-core Zen+, M calls/s in total (median of 3 runs, all three
commits of this series against upstream):
1 thread 4 threads 16 threads
dsymv upstream 5.6 2.3 1.4
this series 6.5 21.8 39.9
dtrsv upstream 3.7 2.0 1.3
this series 4.0 14.3 32.4
Memory use is unchanged: 32 threads taking turns on dgemm (n=1024) or
dsymv map the same buffers and use the same RSS as before.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claiming a slot of memory[] took the slot's lock, set used, and dropped the lock, and freeing set used again; with the _Atomic members each of those is a locked instruction on x86 (four per BLAS call, each a full barrier), against none in the TLS allocator. That costs most where two hyperthreads share a core. Claim a slot by switching used from 0 to 1 with a single compare-and-swap, and release it with a release store (a plain store on x86). The lock member is no longer used for the main table; the overflow area keeps its locking. On an 8-core Zen+, one thread calling dsymv (n=15): 7.1 -> 8.0 M calls/s, the same as the TLS allocator; no change in scaling there. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Threads calling BLAS concurrently now each keep to their own slot of memory[] and write its used flag on every call. Two things let those writes slow down the other threads: - The slots were 64 bytes with no alignment, and in a pthreads build on x86-64 the array started 32 bytes into a cache line, so every slot straddled two lines shared with its neighbours. Align the slots to 128 bytes (the line, plus the one Intel's adjacent-line prefetcher pairs with it). The table grows from 64 to 128 bytes per slot. - The table also shared a 4 KB page with the library's GOT and with memory_initialized. Intel's L2 streamer prefetcher watches the lines a core misses within each 4 KB page and then fetches more lines of that page, never crossing a page boundary. Every call touched several lines of that shared page (two GOT entries, memory_initialized, its own slot), so the streamer kept pulling in lines that other cores were writing, and the cores kept taking them back from each other. perf c2c showed it as the GOT entries of __tls_get_addr and memset and the line of memory_initialized, which are only ever read, being found modified in another core's cache. Align the table to 4 KB and round it up to a whole number of pages (the trailing slots are never used), so that each call touches one line on the table's pages, its own slot, and the streamer has no pattern to follow. Aligning the slots alone was not enough: dsymv (n=15) from 28 threads on an i9-7940X (Skylake-X) still ran at a third of the TLS allocator's speed. Measured against this series with the slots aligned but not the pages, turning the four hardware prefetchers off one at a time (MSR 0x1A4) showed that the streamer alone was responsible; the other three made no difference: dsymv, M calls/s, slots aligned only: all on L2 streamer off i9-7940X (Skylake-X), 28 threads 31.3 100.2 2x Xeon E5-2680 v2 (Ivy Bridge-EP), 40 29.4 82.3 In the slow runs, the L2 prefetches went from 0.4-0.8 M to 22-89 M, and with them the memory ordering machine clears (from a few thousand to 1.9-6.8 M) and the loads that found their line modified in another core's cache; on the two-socket machine, mostly in the other socket's. With the pages aligned too, dsymv from 28 threads on the i9-7940X went from 52 to 92 M calls/s, the same as the TLS allocator, and dtrsv from 67 to 68 (TLS: 72). On the Ivy Bridge-EP, dsymv from 40 threads went from 26-40 to 62-79 over six runs, level with the TLS allocator (70.1 against 75.6 in the same session). The page alignment costs at most 3 KB of padding. This only holds while a call touches nothing else on these pages; the comment at the table says so. On the i9-7940X, one extra read per call of an unused slot on the same page as the threads' slots ran dsymv at 87 M calls/s, against 100 with the same read on another page, and with five times the machine clears. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
These changes are to the default (non-TLS) buffer allocator in driver/others/memory.c. The first commit fixes crashes once more buffers are in use than the main table holds. The other four remove the contention that serialized threads calling BLAS concurrently. This is intended to fix #5589.
Performance
Many threads each calling a small level-2 routine with 1 BLAS thread, in total M calls/s (8-core Zen+, median of 3 runs, upstream vs. the first three commits of the branch):