Skip to content

[Vulkan] Stage bool tensors without 8-bit storage and register logical_not - #23246

Merged
mergennachin merged 26 commits into
mainfrom
gh/mergennachin/30/head
Oct 5, 2026
Merged

mergennachin merged 26 commits into
mainfrom
gh/mergennachin/30/head

Conversation

@mergennachin

@mergennachin mergennachin commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

is_bitw8 did not include kBool, so bool staging required 8-bit storage buffers even for texture tensors. It now takes the bitw8 path. VulkanBackend::init rejects graphs with non-constant bool buffers with NotSupported on devices without 8-bit storage buffers, and destroys the partially built graph on any init error. aten.logical_not was already accepted by the partitioner but had no runtime registration; it now maps to the uint8 bitwise_not kernel.

Part 7/15 of the Vulkan transformer and operator-conformance stack. #23245 has landed; this is now the bottom PR, reviewed directly against main. Integration PR: #23254. Related logical_not registration: #22786.

Prior native validation: 2 passed and one expected skip on MoltenVK with portable CPU kernels.

Rebased onto main at 91d26b314053, including the upstream Adreno UBO indexing fix. The code patch is unchanged. The complete shader set compiles, all 15 graph-builder/serialization tests pass, and lintrunner and git diff --check pass. Hardware and SwiftShader CI will rerun on the rebased stack.

Recreates #23207 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

@mergennachin mergennachin added the module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/ label Sep 29, 2026
@pytorch-bot

pytorch-bot Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23246

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit bcd89c3 with merge base 91d26b3 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 29, 2026
mergennachin added a commit that referenced this pull request Sep 29, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers.

Fixes #23156.

This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 | Hardware Vulkan CI |
| 2 | #23241 | Scalar cache type and signed-zero keys |
| 3 | #23242 | Vulkan-local signed-zero serialization |
| 4 | #23243 | GELU modes and view kwargs |
| 5 | #23244 | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 | Reduction dimension guards |
| 7 | #23246 | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename.

Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons.

Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed.

NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.


cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: c6a7e94
ghstack-comment-id: 5892899249
Pull-Request: #23254
mergennachin added a commit that referenced this pull request Oct 3, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers.

Fixes #23156.

This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 | Hardware Vulkan CI |
| 2 | #23241 | Scalar cache type and signed-zero keys |
| 3 | #23242 | Vulkan-local signed-zero serialization |
| 4 | #23243 | GELU modes and view kwargs |
| 5 | #23244 | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 | Reduction dimension guards |
| 7 | #23246 | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename.

Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons.

Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed.

NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.


cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: c9b5269
ghstack-comment-id: 5892899249
Pull-Request: #23254
mergennachin added a commit that referenced this pull request Oct 3, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers.

Fixes #23156.

This is part 15 of the original ghstack series. Parts 1–4 (#23240 through #23243) have landed in main; the remaining eleven PRs run from #23244 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 (landed) | Hardware Vulkan CI |
| 2 | #23241 (landed) | Scalar cache type and signed-zero keys |
| 3 | #23242 (landed) | Vulkan-local signed-zero serialization |
| 4 | #23243 (landed) | GELU modes and view kwargs |
| 5 | #23244 | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 | Reduction and arg-reduction dimension/storage guards |
| 7 | #23246 | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim with keepdim=True |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references.

Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons.

Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

Validation: Rebased onto main at `2c8510343e47` after parts 1–4 landed. `git range-diff` confirms all eleven remaining patches are unchanged by this rebase. On the previous base (`0b3d26d8c1f9`), the Release Vulkan/portable runtime rebuilt successfully and the native suite passed **51 tests with one expected SwiftShader-only skip**, before subsequent review tests were added. Coverage includes the dynamic tests, graph builder, serialization, and PT2E quantized linear with downcasting disabled. The two FP16 texture rounding tests pass 18 exact ATen cases covering both signs, signed zero, even/odd ties, exponent carry, the normal/subnormal and zero/subnormal boundaries, and values around the ±65520 overflow threshold. A separate test directly lowers int32 amax/amin to Vulkan buffers and checks exact ATen agreement beyond ±65504; the same inputs also match portable CPU execution. The arg-reduction guard update then passed four focused native tests, including 28 argmax/argmin cases checking Vulkan buffer execution and portable CPU fallback against ATen. The final rebase preserved these tested Vulkan sources; only upstream CI configuration and documentation changed. Lintrunner and `git diff --check` passed.

Delegation of `any.dim` requires `keepdim=True`. Omitting `keepdim` or setting it to `False` keeps the reduction on CPU, avoiding a new bool-buffer requirement for surrounding texture delegates. The default/explicit-false regression test verifies texture storage and exact ATen results across changing shapes. Five focused native tests pass on MoltenVK, including existing `keepdim=True` execution and the full dynamic transformer tests; lintrunner and `git diff --check` also pass. SwiftShader CI will verify the updated heads.

Vulkan CI will run again on the rebased ghstack heads.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Regression coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain. The power and chained-multiplication tests compare with ATen in both texture and buffer storage.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: afcd265
ghstack-comment-id: 5892899249
Pull-Request: #23254
mergennachin added a commit that referenced this pull request Oct 3, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows.

Fixes #23156.

This is part 15 of the original ghstack series. Parts 1–5 (#23240 through #23244) have landed in main; the remaining ten PRs run from #23245 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 (landed) | Hardware Vulkan CI |
| 2 | #23241 (landed) | Scalar cache type and signed-zero keys |
| 3 | #23242 (landed) | Vulkan-local signed-zero serialization |
| 4 | #23243 (landed) | GELU modes and view kwargs |
| 5 | #23244 (landed) | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 | Reduction and arg-reduction dimension/storage guards |
| 7 | #23246 | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim textures with either keepdim setting |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. Coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain in both texture and buffer storage. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references.

In #23253, `any.dim` now supports both `keepdim` settings on textures. With `keepdim=False`, a reduction into a temporary texture is followed by a GPU view/repack into the squeezed output shape. Both nodes resize dynamically, and empty reduced axes produce false. This adds one temporary texture and one GPU dispatch without introducing a bool-buffer requirement for texture models. Unsupported scalar inputs and 4D batch/channel axes use CPU fallback.

Current validation: Rebased onto main at `903cef063774` after parts 1–5 landed, preserving the ten existing code patches before adding the texture implementation in #23253. The Release Vulkan/portable runtime builds. Eleven focused native tests pass on Apple M1 Pro / MoltenVK, covering dynamic and chained any reductions, actual texture storage, scalar outputs, singleton and empty dimensions, growth after a zero-length dimension, transformer integration, supported and unsupported 4D reductions, special values, and FP16 rounding. An additional 27 FACTO-generated bool any.dim cases and six boundary cases match ATen and portable kernels exactly through the shared `backends/test/` harness: 23 run on Vulkan textures and 10 exercise expected CPU fallback. Lintrunner and `git diff --check` pass. SwiftShader and hardware CI will run on the updated ghstack heads.

Earlier validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covered all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. The prior full native run passed 51 tests with one expected SwiftShader-only skip. Subsequent review coverage added exact FP16 rounding boundaries, int32 buffer amax/amin range checks against both ATen and portable kernels, and 28 argmax/argmin execution and fallback cases. These are prior validation results, not a repeat of the full suite for this update.

Validation uses Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: 980827d
ghstack-comment-id: 5892899249
Pull-Request: #23254
Base automatically changed from gh/mergennachin/29/head to main October 4, 2026 14:57
mergennachin added a commit that referenced this pull request Oct 5, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows.

Fixes #23156.

This is part 15 of the original ghstack series. Parts 1–6 (#23240 through #23245) have landed in main; the remaining nine PRs run from #23246 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 (landed) | Hardware Vulkan CI |
| 2 | #23241 (landed) | Scalar cache type and signed-zero keys |
| 3 | #23242 (landed) | Vulkan-local signed-zero serialization |
| 4 | #23243 (landed) | GELU modes and view kwargs |
| 5 | #23244 (landed) | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 (landed) | Reduction and arg-reduction dimension/storage guards |
| 7 | #23246 | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim textures with either keepdim setting |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. Coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain in both texture and buffer storage. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references.

In #23253, `any.dim` now supports both `keepdim` settings on textures. With `keepdim=False`, a reduction into a temporary texture is followed by a GPU view/repack into the squeezed output shape. Both nodes resize dynamically, and empty reduced axes produce false. This adds one temporary texture and one GPU dispatch without introducing a bool-buffer requirement for texture models. Unsupported scalar inputs and 4D batch/channel axes use CPU fallback.

Current rebase: Rebased onto main at `91d26b314053` after #23245 landed. All nine remaining code patches are preserved; the only change to the combined source tree is main's Adreno UBO vector-indexing hardening (`b97239a0e4`). All 1153 shader variants compile with glslc, the 15 graph-builder/serialization tests pass, and lintrunner and `git diff --check` pass. Hardware and SwiftShader CI will rerun on the updated ghstack heads.

Prior texture-path validation on main at `903cef063774`: the Release Vulkan/portable runtime built and eleven focused native tests passed on Apple M1 Pro / MoltenVK. Coverage included dynamic and chained any reductions, actual texture storage, scalar outputs, singleton and empty dimensions, growth after a zero-length dimension, transformer integration, supported and unsupported 4D reductions, special values, and FP16 rounding. An additional 27 FACTO-generated bool any.dim cases and six boundary cases matched ATen and portable kernels exactly through the shared `backends/test/` harness: 23 ran on Vulkan textures and 10 exercised expected CPU fallback. The subsequent review of #23245 with #23253 verified 784 support/storage combinations and eight mixed dynamic models; all nine review tests passed, with execution results matching ATen and portable kernels. These GPU execution results precede the latest upstream shader-indexing change.

Earlier validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covered all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. The prior full native run passed 51 tests with one expected SwiftShader-only skip. Subsequent review coverage added exact FP16 rounding boundaries, int32 buffer amax/amin range checks against both ATen and portable kernels, and 28 argmax/argmin execution and fallback cases. These are prior validation results, not a repeat of the full suite for this update.

Validation uses Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: 4c0dff7
ghstack-comment-id: 5892899249
Pull-Request: #23254
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@mergennachin
mergennachin merged commit f0c57bb into main Oct 5, 2026
@mergennachin
mergennachin deleted the gh/mergennachin/30/head branch October 5, 2026 14:30
mergennachin added a commit that referenced this pull request Oct 7, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows.

Fixes #23156.

This is part 15 of the original ghstack series. Parts 1–7 (#23240 through #23246) have landed in main; the remaining eight PRs run from #23247 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 (landed) | Hardware Vulkan CI |
| 2 | #23241 (landed) | Scalar cache type and signed-zero keys |
| 3 | #23242 (landed) | Vulkan-local signed-zero serialization |
| 4 | #23243 (landed) | GELU modes and view kwargs |
| 5 | #23244 (landed) | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 (landed) | Reduction and arg-reduction dimension/storage guards |
| 7 | #23246 (landed) | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim textures with either keepdim setting |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. Coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain in both texture and buffer storage. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references.

In #23253, `any.dim` now supports both `keepdim` settings on textures. With `keepdim=False`, a reduction into a temporary texture is followed by a GPU view/repack into the squeezed output shape. Both nodes resize dynamically, and empty reduced axes produce false. This adds one temporary texture and one GPU dispatch without introducing a bool-buffer requirement for texture models. Unsupported scalar inputs and 4D batch/channel axes use CPU fallback.

Current rebase: Rebased onto main at `5f72739cf26e` after #23246 landed. All eight remaining code patches are unchanged, and the rebase had no conflicts. All 1153 shader variants compile with glslc, the 15 graph-builder/serialization tests pass, and lintrunner and `git diff --check` pass. Native GPU execution was validated on an earlier base; hardware and SwiftShader CI will rerun on the updated ghstack heads.

Prior texture-path validation on main at `903cef063774`: the Release Vulkan/portable runtime built and eleven focused native tests passed on Apple M1 Pro / MoltenVK. Coverage included dynamic and chained any reductions, actual texture storage, scalar outputs, singleton and empty dimensions, growth after a zero-length dimension, transformer integration, supported and unsupported 4D reductions, special values, and FP16 rounding. An additional 27 FACTO-generated bool any.dim cases and six boundary cases matched ATen and portable kernels exactly through the shared `backends/test/` harness: 23 ran on Vulkan textures and 10 exercised expected CPU fallback. The subsequent review of #23245 with #23253 verified 784 support/storage combinations and eight mixed dynamic models; all nine review tests passed, with execution results matching ATen and portable kernels. These GPU execution results precede the latest upstream shader-indexing change.

Earlier validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covered all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. The prior full native run passed 51 tests with one expected SwiftShader-only skip. Subsequent review coverage added exact FP16 rounding boundaries, int32 buffer amax/amin range checks against both ATen and portable kernels, and 28 argmax/argmin execution and fallback cases. These are prior validation results, not a repeat of the full suite for this update.

Validation uses Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: e508cde
ghstack-comment-id: 5892899249
Pull-Request: #23254
mergennachin added a commit that referenced this pull request Oct 7, 2026
#23247)

The binary-scalar and comparison kernels receive scalars as int32 or
float, yet the partitioner accepted NaN, integers outside int32 and
symbolic scalars, which lowered with wrong values or failed in the graph
builder. is_scalar_value_supported rejects them for pow.Tensor_Scalar
and the comparison scalar ops. The partitioner also rejects SymFloat and
SymBool values and keeps SymInt producers on CPU when a consumer is
unsupported; node support is memoized so that check stays linear.

Part 8/15 of the Vulkan transformer and operator-conformance stack.
Depends on #23246; review against the selected base branch. Integration
PR: #23254.

Validation: 5 passed, 2 warnings in 150.36s (0:02:30). Lintrunner and
git diff --check pass. Native tests use MoltenVK with portable CPU
kernels; hardware and SwiftShader CI are pending.

Recreates #23208 through ghstack. Prior review discussion remains on
that PR.

Authored with OpenAI Codex; split planned with Claude Code.


cc @SS-JIA @manuelcandales @digantdesai @cbilgin
mergennachin added a commit that referenced this pull request Oct 7, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows.

Fixes #23156.

This is part 15 of the original ghstack series. Parts 1–7 (#23240 through #23246) have landed in main; the remaining eight PRs run from #23247 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 (landed) | Hardware Vulkan CI |
| 2 | #23241 (landed) | Scalar cache type and signed-zero keys |
| 3 | #23242 (landed) | Vulkan-local signed-zero serialization |
| 4 | #23243 (landed) | GELU modes and view kwargs |
| 5 | #23244 (landed) | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 (landed) | Reduction and arg-reduction dimension/storage guards |
| 7 | #23246 (landed) | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim textures with either keepdim setting |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. Coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain in both texture and buffer storage. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references.

In #23253, `any.dim` now supports both `keepdim` settings on textures. With `keepdim=False`, a reduction into a temporary texture is followed by a GPU view/repack into the squeezed output shape. Both nodes resize dynamically, and empty reduced axes produce false. This adds one temporary texture and one GPU dispatch without introducing a bool-buffer requirement for texture models. Unsupported scalar inputs and 4D batch/channel axes use CPU fallback.

Current rebase: Rebased onto main at `5f72739cf26e` after #23246 landed. All eight remaining code patches are unchanged, and the rebase had no conflicts. All 1153 shader variants compile with glslc, the 15 graph-builder/serialization tests pass, and lintrunner and `git diff --check` pass. Native GPU execution was validated on an earlier base; hardware and SwiftShader CI will rerun on the updated ghstack heads.

Prior texture-path validation on main at `903cef063774`: the Release Vulkan/portable runtime built and eleven focused native tests passed on Apple M1 Pro / MoltenVK. Coverage included dynamic and chained any reductions, actual texture storage, scalar outputs, singleton and empty dimensions, growth after a zero-length dimension, transformer integration, supported and unsupported 4D reductions, special values, and FP16 rounding. An additional 27 FACTO-generated bool any.dim cases and six boundary cases matched ATen and portable kernels exactly through the shared `backends/test/` harness: 23 ran on Vulkan textures and 10 exercised expected CPU fallback. The subsequent review of #23245 with #23253 verified 784 support/storage combinations and eight mixed dynamic models; all nine review tests passed, with execution results matching ATen and portable kernels. These GPU execution results precede the latest upstream shader-indexing change.

Earlier validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covered all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. The prior full native run passed 51 tests with one expected SwiftShader-only skip. Subsequent review coverage added exact FP16 rounding boundaries, int32 buffer amax/amin range checks against both ATen and portable kernels, and 28 argmax/argmin execution and fallback cases. These are prior validation results, not a repeat of the full suite for this update.

Validation uses Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: fe79288
ghstack-comment-id: 5892899249
Pull-Request: #23254
mergennachin added a commit that referenced this pull request Oct 7, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows.

Fixes #23156.

This is part 15 of the original ghstack series. Parts 1–7 (#23240 through #23246) have landed in main; the remaining eight PRs run from #23247 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 (landed) | Hardware Vulkan CI |
| 2 | #23241 (landed) | Scalar cache type and signed-zero keys |
| 3 | #23242 (landed) | Vulkan-local signed-zero serialization |
| 4 | #23243 (landed) | GELU modes and view kwargs |
| 5 | #23244 (landed) | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 (landed) | Reduction and arg-reduction dimension/storage guards |
| 7 | #23246 (landed) | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim textures with either keepdim setting |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. Coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain in both texture and buffer storage. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references.

In #23253, `any.dim` now supports both `keepdim` settings on textures. With `keepdim=False`, a reduction into a temporary texture is followed by a GPU view/repack into the squeezed output shape. Both nodes resize dynamically, and empty reduced axes produce false. This adds one temporary texture and one GPU dispatch without introducing a bool-buffer requirement for texture models. Unsupported scalar inputs and 4D batch/channel axes use CPU fallback.

Current rebase: Rebased onto main at `5f72739cf26e` after #23246 landed. All eight remaining code patches are unchanged, and the rebase had no conflicts. All 1153 shader variants compile with glslc, the 15 graph-builder/serialization tests pass, and lintrunner and `git diff --check` pass. Native GPU execution was validated on an earlier base; hardware and SwiftShader CI will rerun on the updated ghstack heads.

Prior texture-path validation on main at `903cef063774`: the Release Vulkan/portable runtime built and eleven focused native tests passed on Apple M1 Pro / MoltenVK. Coverage included dynamic and chained any reductions, actual texture storage, scalar outputs, singleton and empty dimensions, growth after a zero-length dimension, transformer integration, supported and unsupported 4D reductions, special values, and FP16 rounding. An additional 27 FACTO-generated bool any.dim cases and six boundary cases matched ATen and portable kernels exactly through the shared `backends/test/` harness: 23 ran on Vulkan textures and 10 exercised expected CPU fallback. The subsequent review of #23245 with #23253 verified 784 support/storage combinations and eight mixed dynamic models; all nine review tests passed, with execution results matching ATen and portable kernels. These GPU execution results precede the latest upstream shader-indexing change.

Earlier validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covered all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. The prior full native run passed 51 tests with one expected SwiftShader-only skip. Subsequent review coverage added exact FP16 rounding boundaries, int32 buffer amax/amin range checks against both ATen and portable kernels, and 28 argmax/argmin execution and fallback cases. These are prior validation results, not a repeat of the full suite for this update.

Validation uses Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: a7c2766
ghstack-comment-id: 5892899249
Pull-Request: #23254
mergennachin added a commit that referenced this pull request Oct 7, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows.

Fixes #23156.

This is part 15 of the original ghstack series. Parts 1–7 (#23240 through #23246) have landed in main; the remaining eight PRs run from #23247 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 (landed) | Hardware Vulkan CI |
| 2 | #23241 (landed) | Scalar cache type and signed-zero keys |
| 3 | #23242 (landed) | Vulkan-local signed-zero serialization |
| 4 | #23243 (landed) | GELU modes and view kwargs |
| 5 | #23244 (landed) | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 (landed) | Reduction and arg-reduction dimension/storage guards |
| 7 | #23246 (landed) | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim textures with either keepdim setting |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. Coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain in both texture and buffer storage. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references.

In #23253, `any.dim` now supports both `keepdim` settings on textures. With `keepdim=False`, a reduction into a temporary texture is followed by a GPU view/repack into the squeezed output shape. Both nodes resize dynamically, and empty reduced axes produce false. This adds one temporary texture and one GPU dispatch without introducing a bool-buffer requirement for texture models. Unsupported scalar inputs and 4D batch/channel axes use CPU fallback.

Current rebase: Rebased onto main at `5f72739cf26e` after #23246 landed. All eight remaining code patches are unchanged, and the rebase had no conflicts. All 1153 shader variants compile with glslc, the 15 graph-builder/serialization tests pass, and lintrunner and `git diff --check` pass. Native GPU execution was validated on an earlier base; hardware and SwiftShader CI will rerun on the updated ghstack heads.

Prior texture-path validation on main at `903cef063774`: the Release Vulkan/portable runtime built and eleven focused native tests passed on Apple M1 Pro / MoltenVK. Coverage included dynamic and chained any reductions, actual texture storage, scalar outputs, singleton and empty dimensions, growth after a zero-length dimension, transformer integration, supported and unsupported 4D reductions, special values, and FP16 rounding. An additional 27 FACTO-generated bool any.dim cases and six boundary cases matched ATen and portable kernels exactly through the shared `backends/test/` harness: 23 ran on Vulkan textures and 10 exercised expected CPU fallback. The subsequent review of #23245 with #23253 verified 784 support/storage combinations and eight mixed dynamic models; all nine review tests passed, with execution results matching ATen and portable kernels. These GPU execution results precede the latest upstream shader-indexing change.

Earlier validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covered all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. The prior full native run passed 51 tests with one expected SwiftShader-only skip. Subsequent review coverage added exact FP16 rounding boundaries, int32 buffer amax/amin range checks against both ATen and portable kernels, and 28 argmax/argmin execution and fallback cases. These are prior validation results, not a repeat of the full suite for this update.

Validation uses Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: 6b879b2
ghstack-comment-id: 5892899249
Pull-Request: #23254
mergennachin added a commit that referenced this pull request Oct 7, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows.

Fixes #23156.

This is part 15 of the original ghstack series. Parts 1–7 (#23240 through #23246) have landed in main; the remaining eight PRs run from #23247 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 (landed) | Hardware Vulkan CI |
| 2 | #23241 (landed) | Scalar cache type and signed-zero keys |
| 3 | #23242 (landed) | Vulkan-local signed-zero serialization |
| 4 | #23243 (landed) | GELU modes and view kwargs |
| 5 | #23244 (landed) | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 (landed) | Reduction and arg-reduction dimension/storage guards |
| 7 | #23246 (landed) | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim textures with either keepdim setting |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. Coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain in both texture and buffer storage. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references.

In #23253, `any.dim` now supports both `keepdim` settings on textures. With `keepdim=False`, a reduction into a temporary texture is followed by a GPU view/repack into the squeezed output shape. Both nodes resize dynamically, and empty reduced axes produce false. This adds one temporary texture and one GPU dispatch without introducing a bool-buffer requirement for texture models. Unsupported scalar inputs and 4D batch/channel axes use CPU fallback.

Current rebase: Rebased onto main at `5f72739cf26e` after #23246 landed. All eight remaining code patches are unchanged, and the rebase had no conflicts. All 1153 shader variants compile with glslc, the 15 graph-builder/serialization tests pass, and lintrunner and `git diff --check` pass. Native GPU execution was validated on an earlier base; hardware and SwiftShader CI will rerun on the updated ghstack heads.

Prior texture-path validation on main at `903cef063774`: the Release Vulkan/portable runtime built and eleven focused native tests passed on Apple M1 Pro / MoltenVK. Coverage included dynamic and chained any reductions, actual texture storage, scalar outputs, singleton and empty dimensions, growth after a zero-length dimension, transformer integration, supported and unsupported 4D reductions, special values, and FP16 rounding. An additional 27 FACTO-generated bool any.dim cases and six boundary cases matched ATen and portable kernels exactly through the shared `backends/test/` harness: 23 ran on Vulkan textures and 10 exercised expected CPU fallback. The subsequent review of #23245 with #23253 verified 784 support/storage combinations and eight mixed dynamic models; all nine review tests passed, with execution results matching ATen and portable kernels. These GPU execution results precede the latest upstream shader-indexing change.

Earlier validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covered all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. The prior full native run passed 51 tests with one expected SwiftShader-only skip. Subsequent review coverage added exact FP16 rounding boundaries, int32 buffer amax/amin range checks against both ATen and portable kernels, and 28 argmax/argmin execution and fallback cases. These are prior validation results, not a repeat of the full suite for this update.

Validation uses Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: 0705e5f
ghstack-comment-id: 5892899249
Pull-Request: #23254

This branch was successfully deployed

1 active deployment
cadence — bcd89c39 Deployed Oct 5, 2026 by mergennachin
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants