Skip to content

[Vulkan] Add aten.any.dim - #23214

Closed
mergennachin wants to merge 1 commit into
mergennachin/vulkan-23156-13-mul-scalarfrom
mergennachin/vulkan-23156-14-any
Closed

mergennachin wants to merge 1 commit into
mergennachin/vulkan-23156-13-mul-scalarfrom
mergennachin/vulkan-23156-14-any

Conversation

@mergennachin

@mergennachin mergennachin commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Superseded by #23253, part 14/15 of the ghstack replacement series. The current integration PR is #23254. This PR is closed in favor of the replacement; its previous description and review history are retained.


Padding masks in transformer blocks reduce bool tensors with any(dim), which forced a CPU fallback in the middle of the block. This adds uint8 any variants to the reduce and per-row reduce shaders, accumulating with max over uint8-backed bool storage, and registers the op with the existing reduce support rules.

Part 14/15 of the Vulkan transformer and operator-conformance stack. Depends on #23213; review against the selected base branch. Integration PR: #23162.

Validation: 3 passed, 2 warnings in 57.92s. Lintrunner and git diff --check pass. Native tests use MoltenVK with portable CPU kernels; hardware and SwiftShader CI are pending.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

Padding masks in transformer blocks reduce bool tensors with any(dim), which forced a CPU fallback in the middle of the block. This adds uint8 any variants to the reduce and per-row reduce shaders, accumulating with max over uint8-backed bool storage, and registers the op with the existing reduce support rules.

Authored with OpenAI Codex; split planned with Claude Code.
@pytorch-bot pytorch-bot Bot added the module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/ label Sep 28, 2026
@pytorch-bot

pytorch-bot Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23214

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit eb2400e with merge base a318382 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 28, 2026
mergennachin added a commit that referenced this pull request Oct 8, 2026
Boolean any(dim) reductions in attention masks fell back to CPU. This adds uint8 any variants to the texture and per-row buffer reduction shaders and supports both keepdim settings on textures.

For keepdim=False, including the default when keepdim is omitted, the runtime reduces into a temporary texture with the reduced axis retained, then uses the existing texture view shader to repack into the requested output shape. Both nodes resize with the input. This adds one temporary texture and one GPU repack dispatch. Texture execution requires no 8-bit storage buffers, so surrounding texture operators can remain in the same delegate. The reduction shader writes the false identity when the reduced axis is empty.

Delegation is limited to bool tensors and supported reduction axes. Rank-zero inputs and unsupported 4D batch/channel reductions fall back to CPU. Buffer execution is available for the last axis on devices with 8-bit storage support.

Part 14/15 of the Vulkan transformer and operator-conformance stack. Depends on #23252; review against the selected base branch. Integration PR: #23254.

Validation: The Release Vulkan/portable runtime builds and 11 focused native tests pass on Apple M1 Pro / MoltenVK. Tests assert actual texture storage and exact ATen results for omitted and explicit keepdim=False, chained reductions, scalar outputs, singleton and empty dimensions, supported 4D shapes, and dynamic dimensions shrinking to zero and growing again. Existing keepdim=True texture/buffer execution, transformer integration, reduction special values, and FP16 rounding tests pass. Through the shared backends/test harness, 27 FACTO-generated bool any.dim cases and six boundary cases match both ATen and portable kernels exactly: 23 execute on Vulkan textures and 10 use expected CPU fallback. Lintrunner and git diff --check pass. SwiftShader and hardware CI are pending on the updated heads.

Recreates #23214 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: 148791b
ghstack-comment-id: 5892898410
Pull-Request: #23253
mergennachin added a commit that referenced this pull request Oct 8, 2026
Boolean any(dim) reductions in attention masks fell back to CPU. This
adds uint8 any variants to the texture and per-row buffer reduction
shaders and supports both keepdim settings on textures.

For keepdim=False, including the default when keepdim is omitted, the
runtime reduces into a temporary texture with the reduced axis retained,
then uses the existing texture view shader to repack into the requested
output shape. Both nodes resize with the input. This adds one temporary
texture and one GPU repack dispatch. Texture execution requires no 8-bit
storage buffers, so surrounding texture operators can remain in the same
delegate. The reduction shader writes the false identity when the
reduced axis is empty.

Delegation is limited to bool tensors and supported reduction axes.
Rank-zero inputs and unsupported 4D batch/channel reductions fall back
to CPU. Buffer execution is available for the last axis on devices with
8-bit storage support.

Part 14/15 of the Vulkan transformer and operator-conformance stack.
#23252 has landed; this PR is rebased directly onto `origin/main` at
`9e47127d28`. Integration PR: #23254.

Test Plan: Rebuilt ExecuTorch with Vulkan enabled from the rebased
source on macOS arm64. Ran eight focused tests from
`test_vulkan_dynamic.py` through the public Buck Python test adapter,
explicitly selecting each driver ICD. SwiftShader: 8/8 passed in
145.877s. MoltenVK: 8/8 passed in 151.952s. Both runs had zero failures,
errors, or skips. Coverage includes all five added or modified tests
plus reduction special values and FP16 halfway/overflow rounding:
unsupported-input partitioning, omitted/false keepdim, chained texture
reductions, scalar and empty outputs, 4D axes, and dynamic dimensions
shrinking to zero and growing again. Lintrunner on all six changed
files, `git diff --check`, and `pip check` passed.

The unchanged `test_dynamic_any_dim` reproduced a Buck adapter assertion
when skipping an unsupported SwiftShader buffer subtest. The test now
selects supported storages before entering subtests, matching the
existing SwiftShader guards. SwiftShader retains all texture cases;
MoltenVK retains both texture and buffer cases.

Recreates #23214 through ghstack. Prior review discussion remains on
that PR.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

This branch was successfully deployed

1 active deployment
cadence — eb2400ea Deployed Sep 28, 2026 by mergennachin via hifi-op-test / hifi4 #30218
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant