[llvm] [AMDGPU] Constant folding for wave-reduce intrinsics (PR #212755)
via llvm-commits
llvm-commits at lists.llvm.org
Thu Jul 30 02:02:50 PDT 2026
================
@@ -6013,6 +6044,20 @@ static MachineBasicBlock *lowerWaveReduce(MachineInstr &MI,
break;
}
case AMDGPU::S_ADD_I32: {
+ // Check if Src is a known identity constant.
+ MachineInstr *SrcDef = MRI.getVRegDef(SrcReg);
+ if (SrcDef && SrcDef->isMoveImmediate()) {
+ int64_t Imm = SrcDef->getOperand(1).getImm();
+ if (Imm == 0) { // 0 * bitcount(exec) = 0
+ BuildMI(BB, MI, DL, TII->get(AMDGPU::S_MOV_B32), DstReg).addImm(0);
+ break;
+ }
----------------
easyonaadit wrote:
Not really. The `add`, `sub`, and `xor` have not been marked for instcombine. Not all constants can be folded out, as the final value has to be scaled by wave size. `0`,`1` are special cases.
https://github.com/llvm/llvm-project/pull/212755
More information about the llvm-commits
mailing list