[llvm] [AMDGPU] Constant folding for wave-reduce intrinsics (PR #212755)

via llvm-commits llvm-commits at lists.llvm.org
Thu Jul 30 02:02:50 PDT 2026


================
@@ -6013,6 +6044,20 @@ static MachineBasicBlock *lowerWaveReduce(MachineInstr &MI,
         break;
       }
       case AMDGPU::S_ADD_I32: {
+        // Check if Src is a known identity constant.
+        MachineInstr *SrcDef = MRI.getVRegDef(SrcReg);
+        if (SrcDef && SrcDef->isMoveImmediate()) {
+          int64_t Imm = SrcDef->getOperand(1).getImm();
+          if (Imm == 0) { // 0 * bitcount(exec) = 0
+            BuildMI(BB, MI, DL, TII->get(AMDGPU::S_MOV_B32), DstReg).addImm(0);
+            break;
+          }
----------------
easyonaadit wrote:

Not really. The `add`, `sub`, and `xor` have not been marked for instcombine. Not all constants can be folded out, as the final value has to be scaled by wave size. `0`,`1` are special cases.

https://github.com/llvm/llvm-project/pull/212755


More information about the llvm-commits mailing list