Repository navigation
Split gotoblas dispatch table by group / operation (2/2) - #6088
topolarity wants to merge 7 commits into
Conversation
No functional change. With DYNAMIC_ARCH the small-matrix kernel tables in interface/gemm*.c hold FUNC_OFFSET()s, which are offsets into gotoblas_t, and add them to gotoblas. Spell that base FUNC_BASE(func), next to FUNC_OFFSET(func), and reach it through GEMM_SMALL_KERNEL_BASE in the same way as the offsets themselves. This lets the kernels move out of gotoblas_t without interface/gemm*.c knowing where they went. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Preparation for numbering the DYNAMIC_ARCH cores; nothing includes the header yet. dyn_cores.h holds one X-macro, OPENBLAS_CORE_LIST(X, arg), which expands to X(CORE, arg) for each core of the build, in sorted order and without duplicates. Makefile.system and cmake/arch.cmake each write it where they finish computing DYNAMIC_CORE. dyn_cores.h is written by the top-level make only - sub-makes inherit OPENBLAS_DYN_CORES - so that parallel sub-makes never rewrite it under a compile. It is emitted through $(HASH) because GNU Make 3.81 reads a literal '#' inside $(shell ...) as a comment. cmake rewrites it only when the list changes, and puts it in the build directory, which the kernel targets now have on their include path. A stale copy left in the source tree by make would shadow cmake's, so cmake refuses to configure with one there, as it already does for config_kernel.h. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
1508a72 to
4d4170c
Compare
Preparation for splitting the kernels out of gotoblas_t; nothing uses the
new macros yet.
* common_param.h numbers the cores of the build, OPENBLAS_CORE_<CORE>,
from OPENBLAS_CORE_LIST() in dyn_cores.h, and every core's gotoblas_t
records its number in the new "core" field.
* OPENBLAS_DISPATCH(group) is the running core's table for a group of
kernels, from the array openblas_<group>_dispatch[], indexed by
gotoblas->core. driver/others/dispatch.c will define those arrays,
with DEFINE_DISPATCH_TABLE(group). OPENBLAS_DISPATCH_OFFSET/_BASE do
for the offset tables of interface/gemm*.c what FUNC_OFFSET/FUNC_BASE
do today.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Preparation for splitting the kernels out of gotoblas_t, later in this series. After that split, each core's kernel tables are defined together in setparam_<CORE>.o, and all the per-group arrays together in dispatch.o; a static link would keep all of them as soon as it kept any, and with them every kernel. Separate sections let --gc-sections keep only the groups a program calls. On its own this changes no static link size: every kernel is still reachable from gotoblas_t. Only for compilers known to accept the flags, and not for MSVC-style frontends. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
No functional change. The gotoblas_t initializer was positional across roughly 1300 entries and two dozen nested #if blocks, so which value lands in which field could only be worked out by lining it up against the struct definition by hand. Label every entry with its field, one entry per line. Generated by dispatch-split/add_explicit_inits.py, which pairs the entries with the fields of gotoblas_t for every combination of the macros that guard those fields (BUILD_*, SMALL_MATRIX_OPT, EXPRECISION, ARCH_*: 1024 configurations) and stops if any entry would get a different field in any of them. setparam_<CORE>.o is byte-identical before and after for all 14 x86_64 DYNAMIC_ARCH cores, both as built and with BFLOAT16/HFLOAT16 enabled, EXPRECISION and SMALL_MATRIX_OPT disabled, or only one of BUILD_SINGLE, BUILD_DOUBLE, BUILD_COMPLEX and BUILD_COMPLEX16 enabled. Generated-by: dispatch-split/add_explicit_inits.py Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With DYNAMIC_ARCH, gotoblas_t holds every kernel of a core, and the cpu detection code in driver/others/dynamic*.c takes the address of every core's gotoblas_t. One reference to a core's table thus keeps all of that core's kernels, and a static link keeps every kernel of every core, of which --gc-sections can remove almost nothing. Move the 904 kernel pointers into 233 small tables, one per kernel group (the kernels whose names share a leading token: dgemm_kernel, dgemm_beta, dgemm_incopy, ... are "dgemm"), each instantiated per core in setparam-ref.c and reached through OPENBLAS_DISPATCH(group), from an array of them that driver/others/dispatch.c defines. gotoblas_t keeps only the tuning parameters and init(), so cpu detection no longer references any kernel, and a program links only the groups it calls: text of a static executable before after dgemm only 19.5 MB 0.7 MB dgesv + daxpy + dgemv + dnrm2 19.5 MB 1.6 MB every cblas_d* routine 19.8 MB 3.5 MB (x86_64, 14 cores, linked with --gc-sections; dispatch-split/measure-size.sh.) This commit is the output of dispatch-split/split_dispatch_groups.py and nothing else; see there for the exact rules. "gotoblas -> dgemm_kernel" becomes "OPENBLAS_DISPATCH(dgemm) -> dgemm_kernel", that is openblas_dgemm_dispatch[gotoblas->core]->dgemm_kernel. That costs nothing on real work (single-threaded dgemm, n=2048: 57.1 GFLOP/s before, 57.2 after) but shows on the smallest calls (ddot, n=8: 5.05 ns per call before, 5.6-5.8 ns after). For all 14 x86_64 cores, every kernel pointer (resolved to its symbol) and every tuning parameter (after init()) is the same as before, 11830 entries in all (dispatch-split/compare-tables.py). Generated-by: dispatch-split/split_dispatch_groups.py Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Nothing uses them since the kernels moved out of gotoblas_t: the small-matrix kernel tables use OPENBLAS_DISPATCH_OFFSET/_BASE instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
4d4170c to
37da176
Compare
|
@martin-frbg Do you have any thoughts on these patches? They are invasive, but I do not know of any way to make them smaller and achieve the same ~80-95%+ size savings. FWIW, we'd like to use this We can carry this patch downstream and maintain it there, but it'd be preferable to find a way to get these improvements upstream. Any interest here, or should I close and maintain downstream? |
|
I'd like to get 0.3.35 released first (at the very least). Not sure how common the case of statically linking OpenBLAS to some widely distributed code is - but if I understand you correctly, the potential benefits from your change would not be platform-specific to Macs ? |
Sure thing - no rush here
That's right. The size numbers above are on Linux x86-64 (with |
Dependent on #6087. This is PR 2 of 2 to split the
gotoblasdispatch table by "group" (mostly by operation / kernel). The primary motivation is to make OpenBLAS amenable to pruning by the linker via--Wl,--gc-sectionseven in the presence ofDYNAMIC_ARCH.For applications linking statically against OpenBLAS, this can make a big difference to binary size:
dgemmonlydgesv+daxpy+dgemv+dnrm2cblas_d*routineThis is a very big diff, but it is a very mechanical change.
The changes are:
setparam-ref.cto use designated initializers. This change is not required, but it makes the remaining change more mechanical and less dangerous. This is done by add_explicit_inits.py.gotoblas->group_suffixtoOPENBLAS_DISPATCH(group)->group_suffix, so that all dispatches use the group-specific dispatch table. It also updatessetparam-ref.cto initialize the group-specific dispatch tables and defines them conditionally indispatch.c. This is done by split_dispatch_groups.py.I have tried to make this reviewable, to the extent that I can.
This was generated with the assistance of Claude Opus 5.5 🤖, but the change was done entirely (except for the last 5-line commit) by applying the above scripts to the repo to make the changes in bulk. Hopefully this avoids being at the whim / vigilance of the AI, although care still needs to be taken to sanity check the result.
I have reviewed the core changes to the extent that I can, and I am still running some checks on other platforms.
This is only tested thoroughly on
x86_64-linuxat the moment. I will test soon onaarch64-darwin.