Skip to content

[Vulkan] Delegate aten.mul.Scalar - #23213

Closed
mergennachin wants to merge 1 commit into
mergennachin/vulkan-23156-12-power-contractfrom
mergennachin/vulkan-23156-13-mul-scalar
Closed

mergennachin wants to merge 1 commit into
mergennachin/vulkan-23156-12-power-contractfrom
mergennachin/vulkan-23156-13-mul-scalar

Conversation

@mergennachin

@mergennachin mergennachin commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Superseded by #23252, part 13/15 of the ghstack replacement series. The current integration PR is #23254. This PR is closed in favor of the replacement; its previous description and review history are retained.


Tensor-times-scalar products appear in attention scaling and were kept on CPU, splitting otherwise fully delegated blocks. This adds a mul variant to the binary scalar shaders and registers it with pow.Tensor_Scalar under a shared op_registry entry that applies the same scalar-value guard.

Part 13/15 of the Vulkan transformer and operator-conformance stack. Depends on #23212; review against the selected base branch. Integration PR: #23162.

Validation: 1 passed, 2 warnings in 10.21s. Lintrunner and git diff --check pass. Native tests use MoltenVK with portable CPU kernels; hardware and SwiftShader CI are pending.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

Tensor-times-scalar products appear in attention scaling and were kept on CPU, splitting otherwise fully delegated blocks. This adds a mul variant to the binary scalar shaders and registers it with pow.Tensor_Scalar under a shared op_registry entry that applies the same scalar-value guard.

Authored with OpenAI Codex; split planned with Claude Code.
@pytorch-bot pytorch-bot Bot added the module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/ label Sep 28, 2026
@pytorch-bot

pytorch-bot Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23213

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 35ddf0b with merge base a318382 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 28, 2026
mergennachin added a commit that referenced this pull request Oct 8, 2026
Tensor-times-scalar products appear in attention scaling and were kept
on CPU, splitting otherwise fully delegated blocks. Add a mul variant to
the binary scalar shaders and register it with pow.Tensor_Scalar under
the shared scalar-value guard.

FP16 mul.Scalar rounds each result to half on every device, including
devices that emulate FP16 storage with FP32. A regression test keeps two
mul.Scalar nodes inside one delegate and checks intermediate rounding
with exact ATen output comparison for texture and buffer storage.

Part 13/15 of the Vulkan transformer and operator-conformance stack.
#23251 has landed; this PR is rebased directly onto `origin/main` at
`286643dbb2`. Integration PR: #23254.

Test Plan: Rebuilt ExecuTorch with Vulkan enabled from the rebased
source on macOS arm64. Ran `TestVulkanDynamic.test_signed_zero_scalars`
and `TestVulkanDynamic.test_fp16_chained_mul_scalar` through the public
Buck Python test adapter with each driver ICD explicitly selected.
SwiftShader passed both tests and all 6 subcases in 15.542s; MoltenVK
passed both tests and all 6 subcases in 15.620s. Both runs had zero
failures, errors, or skips. Coverage includes signed-zero preservation
for FP32/FP16 and exact intermediate FP16 rounding with two mul.Scalar
nodes in one delegate, using both texture and buffer storage.
SwiftShader exercises FP16 storage emulation. Lintrunner on all 6
changed files, `git diff --check`, and `pip check` passed.

Recreates #23213 through ghstack. Prior review discussion remains on
that PR.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

This branch was successfully deployed

1 active deployment
cadence — 35ddf0b1 Deployed Sep 28, 2026 by mergennachin via hifi-op-test / hifi4 #30215
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant