Skip to content

fix(warp): properly interpret floating point numbers in match_any - #407

Merged
LegNeato merged 1 commit into
Rust-GPU:mainfrom
Snehal-Reddy:fix-warp-match-float
Sep 30, 2026
Merged

LegNeato merged 1 commit into
Rust-GPU:mainfrom
Snehal-Reddy:fix-warp-match-float

Conversation

@Snehal-Reddy

Copy link
Copy Markdown
Contributor

Summary

This PR fixes a critical logic bug in WarpMatchValue where f32 and f64 types were being incorrectly processed by the generic impl_match! macro.

The macro dynamically applies integer casting (value as u32 and value as u64) to cast inputs to bitmasks before executing the underlying SIMT PTX intrinsic. While this acts as a safe bitwise transmute for values like i32, in Rust, casting a float via as u32 executes a truncating numeric conversion. For example, both 1.1f32 and 1.9f32 truncate to 1u32, and all negative floats cast identically to 0. Consequently, operations like warp_match_any between completely distinct floats would silently evaluate as a perfect match, catastrophically corrupting thread convergence checks and logic algorithms expecting structural equality.
Closes #406

Changes

  • Removed the f32 and f64 types from the impl_match! macro in crates/cuda_std/src/warp.rs to prevent integer cascading behavior.
  • Implemented WarpMatchValue manually for f32, strictly leveraging value.to_bits() to route structural memory bits into match_any_32 and match_all_32.
  • Implemented WarpMatchValue manually for f64, strictly leveraging value.to_bits() to route structural memory bits into match_any_64 and match_all_64.

Testing

  • cargo build passes
  • cargo clippy --workspace passes
  • Tested on: Ubuntu 24.04.1 LTS, NVIDIA GeForce RTX 5060 Ti, CUDA 12.8.93

@Snehal-Reddy
Snehal-Reddy force-pushed the fix-warp-match-float branch 2 times, most recently from 1b3e113 to d3a35a7 Compare August 8, 2026 16:28
This commit separates the f32 and f64 types from the impl_match! macro and correctly handles float equality matching using .to_bits(), preventing catastrophic truncation precision loss during implicit integer casts.

TAG=agy
@LegNeato
LegNeato force-pushed the fix-warp-match-float branch from d3a35a7 to 2649b95 Compare September 29, 2026 03:22
@LegNeato
LegNeato added this pull request to the merge queue Sep 30, 2026
Merged via the queue into Rust-GPU:main with commit 3a82c35 Sep 30, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

WarpMatchValue for f32/f64 uses truncating casts instead of to_bits()

2 participants