Repository navigation
Fall back to the cuDNN heuristics when benchmarking runs out of memory - #3166
Merged
davisking merged 1 commit intoOct 3, 2026
Conversation
The cudnnFind*Algorithm functions allocate GPU memory to benchmark the candidate algorithms. Now that each algorithm is chosen the first time it is needed, and optionally for each input shape, that can happen while the rest of the program is using the GPU, and then the allocation can fail (CUDNN_STATUS_INTERNAL_ERROR_DEVICE_ALLOCATION_FAILED) and take the forward or backward pass down with it. Pick the algorithm with the cudnnGet*Algorithm_v7 heuristics instead then, and leave it out of the cache, so that the configuration gets benchmarked the next time.
Owner
|
Nice, thanks for another PR :D |
reunanen
deleted the
fall-back-to-heuristics-when-benchmarking-runs-out-of-memory
branch
October 4, 2026 05:55
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
When
cudnnFind*Algorithmcannot allocate the GPU memory it needs for benchmarking, it fails withCUDNN_STATUS_INTERNAL_ERROR_DEVICE_ALLOCATION_FAILED(CUDNN_STATUS_ALLOC_FAILEDbefore cuDNN 9), and the forward or backward pass fails with it.Since #3162 each algorithm is chosen the first time it is needed, and with
set_dnn_choose_algorithms_per_input_shape(true)again for each new input shape, so benchmarking can happen while the rest of the program is using the GPU. We hit this in CI, with two inference threads sharing a GPU.In that case this falls back to the
cudnnGet*Algorithm_v7heuristics, and leaves the result out of the cache, so that the configuration is benchmarked properly the next time it comes up.