[llvm] [AMDGPU] Fix v_mov_b16 pair merging when the second mov reads the first (PR #227502)

via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 29 15:46:06 PDT 2026


llvmorg-github-actions[bot] wrote:


<!--LLVM PR SUMMARY COMMENT-->

@llvm/pr-subscribers-backend-amdgpu

Author: Larry Meadows (lfmeadow)

<details>
<summary>Changes</summary>

`SIPostRA16BitMovFolding` (added in #<!-- -->208625) merges a lo16/hi16 `v_mov_b16` pair into one 32-bit instruction. The merged instruction reads both sources before it writes the destination. When the second mov reads the half written by the first, the merge reads stale data:

```
v_mov_b16 v1.h, 0           ; v1 = 0
v_mov_b16 v1.l, v1.h
  => v_lshrrev_b32 v1, 16, v1   ; v1.l = old v1.h, not 0
```

The `v_lshlrev_b32` and `v_perm_b32` forms have the same problem. This miscompiles real code on True16 targets (gfx11, gfx12). Two examples:

- rocPRIM's `int8` radix-sort histogram kernel zeroes a register this way.
- MIOpen's vectorized BF16/FP16 LayerNorm forward kernel zeroes accumulators this way, which produces NaNs.

rocPRIM radix sort and scan-by-key, rocThrust pair/tuple scans and reductions, and MIOpen LayerNorm tests all fail on gfx110x and gfx120x in ROCm CI after this pass landed.

The fix skips the merge when the second mov's source overlaps the first mov's destination. It also turns the "no EXEC write" assert in the scan loop into a bail-out. In release builds that assert does nothing, so the pass could move a mov across an EXEC change.

Testing:
- New MIR cases in `si-post-ra-merge-v-mov-b16.mir`: three negative cases for the dependent-pair bug (lo first, hi first, zero then copy), one positive case where the first mov reads the second's destination (still correct to merge), and one EXEC-write negative case.
- `check-llvm-codegen-amdgpu` and `check-llvm-mc-amdgpu` pass. No existing CHECK lines change.
- On an RX 9070 XT (gfx1201), the failing rocPRIM and rocThrust unit tests pass with this patch and fail without it. The full rocPRIM test suite passes with it.

Assisted-by: Cursor


---
Full diff: https://github.com/llvm/llvm-project/pull/227502.diff


2 Files Affected:

- (modified) llvm/lib/Target/AMDGPU/SIPostRA16BitMovFolding.cpp (+8-3) 
- (modified) llvm/test/CodeGen/AMDGPU/si-post-ra-merge-v-mov-b16.mir (+93) 


``````````diff
The server is unavailable at this time. Please wait a few minutes before you try again.
``````````

</details>


https://github.com/llvm/llvm-project/pull/227502


More information about the llvm-commits mailing list