Repository navigation
[Vulkan] Delegate aten.scalar_tensor with exact integer values - #23249
Merged
Merged
Conversation
…ghstack [ghstack-poisoned]
…ghstack [ghstack-poisoned]
…ghstack [ghstack-poisoned]
…ghstack [ghstack-poisoned]
…ghstack [ghstack-poisoned]
…ghstack [ghstack-poisoned]
…ghstack [ghstack-poisoned]
…ghstack [ghstack-poisoned]
…ghstack [ghstack-poisoned]
…ghstack [ghstack-poisoned]
Contributor
Author
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23249
Note: Links to docs will display an error until the docs builds have been completed. ❌ 2 Cancelled JobsAs of commit a7e71de with merge base 9cced90 ( CANCELLED JOBS - The following jobs were cancelled. Please retry:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This was referenced Sep 29, 2026
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
mergennachin
added a commit
that referenced
this pull request
Oct 7, 2026
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows. Fixes #23156. This is part 15 of the original ghstack series. Parts 1–7 (#23240 through #23246) have landed in main; the remaining eight PRs run from #23247 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 (landed) | Hardware Vulkan CI | | 2 | #23241 (landed) | Scalar cache type and signed-zero keys | | 3 | #23242 (landed) | Vulkan-local signed-zero serialization | | 4 | #23243 (landed) | GELU modes and view kwargs | | 5 | #23244 (landed) | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 (landed) | Reduction and arg-reduction dimension/storage guards | | 7 | #23246 (landed) | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim textures with either keepdim setting | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. Coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain in both texture and buffer storage. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references. In #23253, `any.dim` now supports both `keepdim` settings on textures. With `keepdim=False`, a reduction into a temporary texture is followed by a GPU view/repack into the squeezed output shape. Both nodes resize dynamically, and empty reduced axes produce false. This adds one temporary texture and one GPU dispatch without introducing a bool-buffer requirement for texture models. Unsupported scalar inputs and 4D batch/channel axes use CPU fallback. Current rebase: Rebased onto main at `5f72739cf26e` after #23246 landed. All eight remaining code patches are unchanged, and the rebase had no conflicts. All 1153 shader variants compile with glslc, the 15 graph-builder/serialization tests pass, and lintrunner and `git diff --check` pass. Native GPU execution was validated on an earlier base; hardware and SwiftShader CI will rerun on the updated ghstack heads. Prior texture-path validation on main at `903cef063774`: the Release Vulkan/portable runtime built and eleven focused native tests passed on Apple M1 Pro / MoltenVK. Coverage included dynamic and chained any reductions, actual texture storage, scalar outputs, singleton and empty dimensions, growth after a zero-length dimension, transformer integration, supported and unsupported 4D reductions, special values, and FP16 rounding. An additional 27 FACTO-generated bool any.dim cases and six boundary cases matched ATen and portable kernels exactly through the shared `backends/test/` harness: 23 ran on Vulkan textures and 10 exercised expected CPU fallback. The subsequent review of #23245 with #23253 verified 784 support/storage combinations and eight mixed dynamic models; all nine review tests passed, with execution results matching ATen and portable kernels. These GPU execution results precede the latest upstream shader-indexing change. Earlier validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covered all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. The prior full native run passed 51 tests with one expected SwiftShader-only skip. Subsequent review coverage added exact FP16 rounding boundaries, int32 buffer amax/amin range checks against both ATen and portable kernels, and 28 argmax/argmin execution and fallback cases. These are prior validation results, not a repeat of the full suite for this update. Validation uses Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: e508cde ghstack-comment-id: 5892899249 Pull-Request: #23254
mergennachin
added a commit
that referenced
this pull request
Oct 7, 2026
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows. Fixes #23156. This is part 15 of the original ghstack series. Parts 1–7 (#23240 through #23246) have landed in main; the remaining eight PRs run from #23247 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 (landed) | Hardware Vulkan CI | | 2 | #23241 (landed) | Scalar cache type and signed-zero keys | | 3 | #23242 (landed) | Vulkan-local signed-zero serialization | | 4 | #23243 (landed) | GELU modes and view kwargs | | 5 | #23244 (landed) | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 (landed) | Reduction and arg-reduction dimension/storage guards | | 7 | #23246 (landed) | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim textures with either keepdim setting | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. Coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain in both texture and buffer storage. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references. In #23253, `any.dim` now supports both `keepdim` settings on textures. With `keepdim=False`, a reduction into a temporary texture is followed by a GPU view/repack into the squeezed output shape. Both nodes resize dynamically, and empty reduced axes produce false. This adds one temporary texture and one GPU dispatch without introducing a bool-buffer requirement for texture models. Unsupported scalar inputs and 4D batch/channel axes use CPU fallback. Current rebase: Rebased onto main at `5f72739cf26e` after #23246 landed. All eight remaining code patches are unchanged, and the rebase had no conflicts. All 1153 shader variants compile with glslc, the 15 graph-builder/serialization tests pass, and lintrunner and `git diff --check` pass. Native GPU execution was validated on an earlier base; hardware and SwiftShader CI will rerun on the updated ghstack heads. Prior texture-path validation on main at `903cef063774`: the Release Vulkan/portable runtime built and eleven focused native tests passed on Apple M1 Pro / MoltenVK. Coverage included dynamic and chained any reductions, actual texture storage, scalar outputs, singleton and empty dimensions, growth after a zero-length dimension, transformer integration, supported and unsupported 4D reductions, special values, and FP16 rounding. An additional 27 FACTO-generated bool any.dim cases and six boundary cases matched ATen and portable kernels exactly through the shared `backends/test/` harness: 23 ran on Vulkan textures and 10 exercised expected CPU fallback. The subsequent review of #23245 with #23253 verified 784 support/storage combinations and eight mixed dynamic models; all nine review tests passed, with execution results matching ATen and portable kernels. These GPU execution results precede the latest upstream shader-indexing change. Earlier validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covered all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. The prior full native run passed 51 tests with one expected SwiftShader-only skip. Subsequent review coverage added exact FP16 rounding boundaries, int32 buffer amax/amin range checks against both ATen and portable kernels, and 28 argmax/argmin execution and fallback cases. These are prior validation results, not a repeat of the full suite for this update. Validation uses Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: fe79288 ghstack-comment-id: 5892899249 Pull-Request: #23254
[ghstack-poisoned]
[ghstack-poisoned]
mergennachin
added a commit
that referenced
this pull request
Oct 7, 2026
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows. Fixes #23156. This is part 15 of the original ghstack series. Parts 1–7 (#23240 through #23246) have landed in main; the remaining eight PRs run from #23247 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 (landed) | Hardware Vulkan CI | | 2 | #23241 (landed) | Scalar cache type and signed-zero keys | | 3 | #23242 (landed) | Vulkan-local signed-zero serialization | | 4 | #23243 (landed) | GELU modes and view kwargs | | 5 | #23244 (landed) | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 (landed) | Reduction and arg-reduction dimension/storage guards | | 7 | #23246 (landed) | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim textures with either keepdim setting | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. Coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain in both texture and buffer storage. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references. In #23253, `any.dim` now supports both `keepdim` settings on textures. With `keepdim=False`, a reduction into a temporary texture is followed by a GPU view/repack into the squeezed output shape. Both nodes resize dynamically, and empty reduced axes produce false. This adds one temporary texture and one GPU dispatch without introducing a bool-buffer requirement for texture models. Unsupported scalar inputs and 4D batch/channel axes use CPU fallback. Current rebase: Rebased onto main at `5f72739cf26e` after #23246 landed. All eight remaining code patches are unchanged, and the rebase had no conflicts. All 1153 shader variants compile with glslc, the 15 graph-builder/serialization tests pass, and lintrunner and `git diff --check` pass. Native GPU execution was validated on an earlier base; hardware and SwiftShader CI will rerun on the updated ghstack heads. Prior texture-path validation on main at `903cef063774`: the Release Vulkan/portable runtime built and eleven focused native tests passed on Apple M1 Pro / MoltenVK. Coverage included dynamic and chained any reductions, actual texture storage, scalar outputs, singleton and empty dimensions, growth after a zero-length dimension, transformer integration, supported and unsupported 4D reductions, special values, and FP16 rounding. An additional 27 FACTO-generated bool any.dim cases and six boundary cases matched ATen and portable kernels exactly through the shared `backends/test/` harness: 23 ran on Vulkan textures and 10 exercised expected CPU fallback. The subsequent review of #23245 with #23253 verified 784 support/storage combinations and eight mixed dynamic models; all nine review tests passed, with execution results matching ATen and portable kernels. These GPU execution results precede the latest upstream shader-indexing change. Earlier validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covered all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. The prior full native run passed 51 tests with one expected SwiftShader-only skip. Subsequent review coverage added exact FP16 rounding boundaries, int32 buffer amax/amin range checks against both ATen and portable kernels, and 28 argmax/argmin execution and fallback cases. These are prior validation results, not a repeat of the full suite for this update. Validation uses Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: a7c2766 ghstack-comment-id: 5892899249 Pull-Request: #23254
[ghstack-poisoned]
mergennachin
added a commit
that referenced
this pull request
Oct 7, 2026
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows. Fixes #23156. This is part 15 of the original ghstack series. Parts 1–7 (#23240 through #23246) have landed in main; the remaining eight PRs run from #23247 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 (landed) | Hardware Vulkan CI | | 2 | #23241 (landed) | Scalar cache type and signed-zero keys | | 3 | #23242 (landed) | Vulkan-local signed-zero serialization | | 4 | #23243 (landed) | GELU modes and view kwargs | | 5 | #23244 (landed) | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 (landed) | Reduction and arg-reduction dimension/storage guards | | 7 | #23246 (landed) | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim textures with either keepdim setting | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. Coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain in both texture and buffer storage. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references. In #23253, `any.dim` now supports both `keepdim` settings on textures. With `keepdim=False`, a reduction into a temporary texture is followed by a GPU view/repack into the squeezed output shape. Both nodes resize dynamically, and empty reduced axes produce false. This adds one temporary texture and one GPU dispatch without introducing a bool-buffer requirement for texture models. Unsupported scalar inputs and 4D batch/channel axes use CPU fallback. Current rebase: Rebased onto main at `5f72739cf26e` after #23246 landed. All eight remaining code patches are unchanged, and the rebase had no conflicts. All 1153 shader variants compile with glslc, the 15 graph-builder/serialization tests pass, and lintrunner and `git diff --check` pass. Native GPU execution was validated on an earlier base; hardware and SwiftShader CI will rerun on the updated ghstack heads. Prior texture-path validation on main at `903cef063774`: the Release Vulkan/portable runtime built and eleven focused native tests passed on Apple M1 Pro / MoltenVK. Coverage included dynamic and chained any reductions, actual texture storage, scalar outputs, singleton and empty dimensions, growth after a zero-length dimension, transformer integration, supported and unsupported 4D reductions, special values, and FP16 rounding. An additional 27 FACTO-generated bool any.dim cases and six boundary cases matched ATen and portable kernels exactly through the shared `backends/test/` harness: 23 ran on Vulkan textures and 10 exercised expected CPU fallback. The subsequent review of #23245 with #23253 verified 784 support/storage combinations and eight mixed dynamic models; all nine review tests passed, with execution results matching ATen and portable kernels. These GPU execution results precede the latest upstream shader-indexing change. Earlier validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covered all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. The prior full native run passed 51 tests with one expected SwiftShader-only skip. Subsequent review coverage added exact FP16 rounding boundaries, int32 buffer amax/amin range checks against both ATen and portable kernels, and 28 argmax/argmin execution and fallback cases. These are prior validation results, not a repeat of the full suite for this update. Validation uses Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: 6b879b2 ghstack-comment-id: 5892899249 Pull-Request: #23254
This PR needs a
|
mergennachin
added a commit
that referenced
this pull request
Oct 7, 2026
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows. Fixes #23156. This is part 15 of the original ghstack series. Parts 1–7 (#23240 through #23246) have landed in main; the remaining eight PRs run from #23247 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 (landed) | Hardware Vulkan CI | | 2 | #23241 (landed) | Scalar cache type and signed-zero keys | | 3 | #23242 (landed) | Vulkan-local signed-zero serialization | | 4 | #23243 (landed) | GELU modes and view kwargs | | 5 | #23244 (landed) | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 (landed) | Reduction and arg-reduction dimension/storage guards | | 7 | #23246 (landed) | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim textures with either keepdim setting | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. Coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain in both texture and buffer storage. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references. In #23253, `any.dim` now supports both `keepdim` settings on textures. With `keepdim=False`, a reduction into a temporary texture is followed by a GPU view/repack into the squeezed output shape. Both nodes resize dynamically, and empty reduced axes produce false. This adds one temporary texture and one GPU dispatch without introducing a bool-buffer requirement for texture models. Unsupported scalar inputs and 4D batch/channel axes use CPU fallback. Current rebase: Rebased onto main at `5f72739cf26e` after #23246 landed. All eight remaining code patches are unchanged, and the rebase had no conflicts. All 1153 shader variants compile with glslc, the 15 graph-builder/serialization tests pass, and lintrunner and `git diff --check` pass. Native GPU execution was validated on an earlier base; hardware and SwiftShader CI will rerun on the updated ghstack heads. Prior texture-path validation on main at `903cef063774`: the Release Vulkan/portable runtime built and eleven focused native tests passed on Apple M1 Pro / MoltenVK. Coverage included dynamic and chained any reductions, actual texture storage, scalar outputs, singleton and empty dimensions, growth after a zero-length dimension, transformer integration, supported and unsupported 4D reductions, special values, and FP16 rounding. An additional 27 FACTO-generated bool any.dim cases and six boundary cases matched ATen and portable kernels exactly through the shared `backends/test/` harness: 23 ran on Vulkan textures and 10 exercised expected CPU fallback. The subsequent review of #23245 with #23253 verified 784 support/storage combinations and eight mixed dynamic models; all nine review tests passed, with execution results matching ATen and portable kernels. These GPU execution results precede the latest upstream shader-indexing change. Earlier validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covered all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. The prior full native run passed 51 tests with one expected SwiftShader-only skip. Subsequent review coverage added exact FP16 rounding boundaries, int32 buffer amax/amin range checks against both ATen and portable kernels, and 28 argmax/argmin execution and fallback cases. These are prior validation results, not a repeat of the full suite for this update. Validation uses Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: 0705e5f ghstack-comment-id: 5892899249 Pull-Request: #23254
mergennachin
added a commit
that referenced
this pull request
Oct 7, 2026
The full shaders always received the fill value as a float, which loses int32 fills above 2^24 and is ill-defined for bools. Full.cpp now passes the fill value in the output's accumulator type. Both full shaders explicitly declare the matching fill UBO type for each output variant, independent of the accumulator-type helper. full, full_like, zeros, ones and their _like variants are marked resizable so they can live inside dynamic-shape delegates, and fill values the output dtype cannot represent stay on CPU. Part 11/15 of the Vulkan transformer and operator-conformance stack. Depends on #23249; review against the selected base branch. Integration PR: #23254. Validation: The earlier MoltenVK/portable-kernel run had 6 passes and 2 warnings in 113.26s (0:01:53). For this follow-up, all 8 full shader variants compiled with glslc; commit-hook lint and git diff --check pass. Native tests have not been rerun since this edit; hardware and SwiftShader CI are pending. Recreates #23211 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
EXIR leaves scalar_tensor in the ATen dialect, so the partitioner never matched the edge-dialect registration, and the graph builder emitted a name the runtime does not recognize. The op is now registered for both targets, serialized as aten.scalar_tensor.default, and the scalar is uploaded in the output's type so integer values are exact. Scalars that cannot be represented stay on CPU.
Fixes #23158
Part 10/15 of the Vulkan transformer and operator-conformance stack. Depends on #23248; review against the selected base branch. Integration PR: #23254.
Validation: 8 passed, 2 warnings in 45.35s. Lintrunner and git diff --check pass. Native tests use MoltenVK with portable CPU kernels; hardware and SwiftShader CI are pending.
Recreates #23210 through ghstack. Prior review discussion remains on that PR.
Authored with OpenAI Codex; split planned with Claude Code.
cc @SS-JIA @manuelcandales @digantdesai @cbilgin