[llvm] [AMDGPU][CodeGen] Do not rematerialize registers with convergent users (PR #222322)
via llvm-commits
llvm-commits at lists.llvm.org
Wed Sep 9 05:55:59 PDT 2026
llvmorg-github-actions[bot] wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-backend-amdgpu
Author: Lucas Ramirez (lucas-rami)
<details>
<summary>Changes</summary>
Convergent users of rematerializable registers may observe register lanes that are active at the original definition point but inactive at the rematerialization point, and will thus read incorrect values if we allow the rematerialization to occur. This prevents such rematerializations.
---
Full diff: https://github.com/llvm/llvm-project/pull/222322.diff
2 Files Affected:
- (modified) llvm/lib/Target/AMDGPU/GCNSchedStrategy.cpp (+9)
- (modified) llvm/test/CodeGen/AMDGPU/machine-scheduler-sink-trivial-remats.mir (+52)
``````````diff
diff --git a/llvm/lib/Target/AMDGPU/GCNSchedStrategy.cpp b/llvm/lib/Target/AMDGPU/GCNSchedStrategy.cpp
index aa85be758559a..f16f42f77f421 100644
--- a/llvm/lib/Target/AMDGPU/GCNSchedStrategy.cpp
+++ b/llvm/lib/Target/AMDGPU/GCNSchedStrategy.cpp
@@ -1593,6 +1593,15 @@ bool PreRARematStage::initGCNSchedStage() {
[](const MachineInstr *DefMI) { return DefMI->isConvergent(); }))
continue;
+ // A convergent user (e.g., V_READLANE*) may observe the definition's lanes
+ // whose contents depend on the EXEC mask in effect at the def. Moving the
+ // def into the use's region can change EXEC across the def and thus alter
+ // those lanes, so prevent rematerialization in that case.
+ if (any_of(Users, [](const MachineInstr *UserMI) {
+ return UserMI->isConvergent();
+ }))
+ continue;
+
// We further filter the registers that we can rematerialize based on our
// current tracking capabilities in the stage. Users cannot themselves be
// marked rematerializable, and no register operand of the defining MI can
diff --git a/llvm/test/CodeGen/AMDGPU/machine-scheduler-sink-trivial-remats.mir b/llvm/test/CodeGen/AMDGPU/machine-scheduler-sink-trivial-remats.mir
index 8c9e4a5a26c83..fff3be77d3244 100644
--- a/llvm/test/CodeGen/AMDGPU/machine-scheduler-sink-trivial-remats.mir
+++ b/llvm/test/CodeGen/AMDGPU/machine-scheduler-sink-trivial-remats.mir
@@ -13563,3 +13563,55 @@ body: |
S_ENDPGM 0
...
+---
+name: dont_remat_with_convergent_user
+tracksRegLiveness: true
+machineFunctionInfo:
+ isEntryFunction: true
+body: |
+ bb.0:
+ ; GFX908-LABEL: name: dont_remat_with_convergent_user
+ ; GFX908: %full_exec:sreg_64 = S_MOV_B64 -1
+ ; GFX908-NEXT: %saved_exec:sreg_64 = S_MOV_B64 $exec
+ ; GFX908-NEXT: $exec = S_MOV_B64 %full_exec
+ ; GFX908-NEXT: %pressure:vreg_1024 = IMPLICIT_DEF
+ ; GFX908-NEXT: %narrow_exec:sreg_64 = S_MOV_B64 1
+ ; GFX908-NEXT: %candidate:vgpr_32 = V_MOV_B32_e32 42, implicit $exec
+ ; GFX908-NEXT: $exec = S_MOV_B64 %narrow_exec
+ ; GFX908-NEXT: S_NOP 0, implicit %pressure
+ ; GFX908-NEXT: %inactive_lane:sreg_32_xm0 = V_READLANE_B32 %candidate, 1
+ ; GFX908-NEXT: S_NOP 0, implicit %inactive_lane
+ ; GFX908-NEXT: $exec = S_MOV_B64 %saved_exec
+ ; GFX908-NEXT: S_ENDPGM 0
+ ;
+ ; GFX908-GCNTRACKERS-LABEL: name: dont_remat_with_convergent_user
+ ; GFX908-GCNTRACKERS: %full_exec:sreg_64 = S_MOV_B64 -1
+ ; GFX908-GCNTRACKERS-NEXT: %saved_exec:sreg_64 = S_MOV_B64 $exec
+ ; GFX908-GCNTRACKERS-NEXT: $exec = S_MOV_B64 %full_exec
+ ; GFX908-GCNTRACKERS-NEXT: %pressure:vreg_1024 = IMPLICIT_DEF
+ ; GFX908-GCNTRACKERS-NEXT: %narrow_exec:sreg_64 = S_MOV_B64 1
+ ; GFX908-GCNTRACKERS-NEXT: %candidate:vgpr_32 = V_MOV_B32_e32 42, implicit $exec
+ ; GFX908-GCNTRACKERS-NEXT: $exec = S_MOV_B64 %narrow_exec
+ ; GFX908-GCNTRACKERS-NEXT: S_NOP 0, implicit %pressure
+ ; GFX908-GCNTRACKERS-NEXT: %inactive_lane:sreg_32_xm0 = V_READLANE_B32 %candidate, 1
+ ; GFX908-GCNTRACKERS-NEXT: S_NOP 0, implicit %inactive_lane
+ ; GFX908-GCNTRACKERS-NEXT: $exec = S_MOV_B64 %saved_exec
+ ; GFX908-GCNTRACKERS-NEXT: S_ENDPGM 0
+ %saved_exec:sreg_64 = S_MOV_B64 $exec
+ %full_exec:sreg_64 = S_MOV_B64 -1
+ $exec = S_MOV_B64 %full_exec
+
+ %pressure:vreg_1024 = IMPLICIT_DEF
+ %candidate:vgpr_32 = V_MOV_B32_e32 42, implicit $exec
+
+ %narrow_exec:sreg_64 = S_MOV_B64 1
+ $exec = S_MOV_B64 %narrow_exec
+
+ S_NOP 0, implicit %pressure
+
+ %inactive_lane:sreg_32_xm0 = V_READLANE_B32 %candidate, 1
+ S_NOP 0, implicit %inactive_lane
+
+ $exec = S_MOV_B64 %saved_exec
+ S_ENDPGM 0
+...
``````````
</details>
https://github.com/llvm/llvm-project/pull/222322
More information about the llvm-commits
mailing list