[llvm] [AMDGPU] Keep divergent i64 mul feeding an add for the mad64 fold (PR #226963)
Pankaj Dwivedi via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 29 05:10:36 PDT 2026
================
@@ -1446,14 +1450,26 @@ bool AMDGPUCodeGenPrepareImpl::tryNarrowMathIfNoOverflow(Instruction *I) {
return true;
}
+// Mul24 or narrowing would hide the ISD::MUL that tryFoldToMad64_32 matches.
+bool AMDGPUCodeGenPrepareImpl::shouldKeepMulForMad64(
+ const BinaryOperator &I) const {
+ if (I.getOpcode() != Instruction::Mul || !I.getType()->isIntegerTy(64))
+ return false;
+ if (ST.getGeneration() < AMDGPUSubtarget::GFX9 || UA.isUniformAtDef(&I))
+ return false;
+ return I.hasOneUse() && match(I.user_back(), m_Add(m_Value(), m_Value()));
+}
+
bool AMDGPUCodeGenPrepareImpl::visitBinaryOperator(BinaryOperator &I) {
if (foldBinOpIntoSelect(I))
return true;
- if (UseMul24Intrin && replaceMulWithMul24(I))
- return true;
- if (tryNarrowMathIfNoOverflow(&I))
- return true;
+ if (!shouldKeepMulForMad64(I)) {
----------------
PankajDwivedi-25 wrote:
This only suppresses the CGP mul24/narrowing rewrite. performMulCombine in SIISelLowering can still replace a divergent i64 ISD::MUL with build_pair(mul24, mulhi24) before tryFoldToMad64_32 runs?
https://github.com/llvm/llvm-project/pull/226963
More information about the llvm-commits
mailing list