Repository navigation
[Vulkan] Add aten.any.dim - #23214
Closed
mergennachin wants to merge 1 commit into
Closed
mergennachin wants to merge 1 commit into
mergennachin wants to merge 1 commit into
Conversation
Padding masks in transformer blocks reduce bool tensors with any(dim), which forced a CPU fallback in the middle of the block. This adds uint8 any variants to the reduce and per-row reduce shaders, accumulating with max over uint8-backed bool storage, and registers the op with the existing reduce support rules. Authored with OpenAI Codex; split planned with Claude Code.
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23214
Note: Links to docs will display an error until the docs builds have been completed. ❌ 1 New FailureAs of commit eb2400e with merge base a318382 ( NEW FAILURE - The following job has failed:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
mergennachin
added a commit
that referenced
this pull request
Oct 8, 2026
Boolean any(dim) reductions in attention masks fell back to CPU. This adds uint8 any variants to the texture and per-row buffer reduction shaders and supports both keepdim settings on textures. For keepdim=False, including the default when keepdim is omitted, the runtime reduces into a temporary texture with the reduced axis retained, then uses the existing texture view shader to repack into the requested output shape. Both nodes resize with the input. This adds one temporary texture and one GPU repack dispatch. Texture execution requires no 8-bit storage buffers, so surrounding texture operators can remain in the same delegate. The reduction shader writes the false identity when the reduced axis is empty. Delegation is limited to bool tensors and supported reduction axes. Rank-zero inputs and unsupported 4D batch/channel reductions fall back to CPU. Buffer execution is available for the last axis on devices with 8-bit storage support. Part 14/15 of the Vulkan transformer and operator-conformance stack. Depends on #23252; review against the selected base branch. Integration PR: #23254. Validation: The Release Vulkan/portable runtime builds and 11 focused native tests pass on Apple M1 Pro / MoltenVK. Tests assert actual texture storage and exact ATen results for omitted and explicit keepdim=False, chained reductions, scalar outputs, singleton and empty dimensions, supported 4D shapes, and dynamic dimensions shrinking to zero and growing again. Existing keepdim=True texture/buffer execution, transformer integration, reduction special values, and FP16 rounding tests pass. Through the shared backends/test harness, 27 FACTO-generated bool any.dim cases and six boundary cases match both ATen and portable kernels exactly: 23 execute on Vulkan textures and 10 use expected CPU fallback. Lintrunner and git diff --check pass. SwiftShader and hardware CI are pending on the updated heads. Recreates #23214 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: 148791b ghstack-comment-id: 5892898410 Pull-Request: #23253
mergennachin
added a commit
that referenced
this pull request
Oct 8, 2026
Boolean any(dim) reductions in attention masks fell back to CPU. This adds uint8 any variants to the texture and per-row buffer reduction shaders and supports both keepdim settings on textures. For keepdim=False, including the default when keepdim is omitted, the runtime reduces into a temporary texture with the reduced axis retained, then uses the existing texture view shader to repack into the requested output shape. Both nodes resize with the input. This adds one temporary texture and one GPU repack dispatch. Texture execution requires no 8-bit storage buffers, so surrounding texture operators can remain in the same delegate. The reduction shader writes the false identity when the reduced axis is empty. Delegation is limited to bool tensors and supported reduction axes. Rank-zero inputs and unsupported 4D batch/channel reductions fall back to CPU. Buffer execution is available for the last axis on devices with 8-bit storage support. Part 14/15 of the Vulkan transformer and operator-conformance stack. #23252 has landed; this PR is rebased directly onto `origin/main` at `9e47127d28`. Integration PR: #23254. Test Plan: Rebuilt ExecuTorch with Vulkan enabled from the rebased source on macOS arm64. Ran eight focused tests from `test_vulkan_dynamic.py` through the public Buck Python test adapter, explicitly selecting each driver ICD. SwiftShader: 8/8 passed in 145.877s. MoltenVK: 8/8 passed in 151.952s. Both runs had zero failures, errors, or skips. Coverage includes all five added or modified tests plus reduction special values and FP16 halfway/overflow rounding: unsupported-input partitioning, omitted/false keepdim, chained texture reductions, scalar and empty outputs, 4D axes, and dynamic dimensions shrinking to zero and growing again. Lintrunner on all six changed files, `git diff --check`, and `pip check` passed. The unchanged `test_dynamic_any_dim` reproduced a Buck adapter assertion when skipping an unsupported SwiftShader buffer subtest. The test now selects supported storages before entering subtests, matching the existing SwiftShader guards. SwiftShader retains all texture cases; MoltenVK retains both texture and buffer cases. Recreates #23214 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Superseded by #23253, part 14/15 of the ghstack replacement series. The current integration PR is #23254. This PR is closed in favor of the replacement; its previous description and review history are retained.
Padding masks in transformer blocks reduce bool tensors with any(dim), which forced a CPU fallback in the middle of the block. This adds uint8 any variants to the reduce and per-row reduce shaders, accumulating with max over uint8-backed bool storage, and registers the op with the existing reduce support rules.
Part 14/15 of the Vulkan transformer and operator-conformance stack. Depends on #23213; review against the selected base branch. Integration PR: #23162.
Validation: 3 passed, 2 warnings in 57.92s. Lintrunner and git diff --check pass. Native tests use MoltenVK with portable CPU kernels; hardware and SwiftShader CI are pending.
Authored with OpenAI Codex; split planned with Claude Code.
cc @SS-JIA @manuelcandales @digantdesai @cbilgin