[llvm] [AMDGPU] Keep divergent i64 mul feeding an add for the mad64 fold (PR #226963)

Pankaj Dwivedi via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 29 05:10:36 PDT 2026


================
@@ -1446,14 +1450,26 @@ bool AMDGPUCodeGenPrepareImpl::tryNarrowMathIfNoOverflow(Instruction *I) {
   return true;
 }
 
+// Mul24 or narrowing would hide the ISD::MUL that tryFoldToMad64_32 matches.
+bool AMDGPUCodeGenPrepareImpl::shouldKeepMulForMad64(
+    const BinaryOperator &I) const {
+  if (I.getOpcode() != Instruction::Mul || !I.getType()->isIntegerTy(64))
+    return false;
+  if (ST.getGeneration() < AMDGPUSubtarget::GFX9 || UA.isUniformAtDef(&I))
+    return false;
+  return I.hasOneUse() && match(I.user_back(), m_Add(m_Value(), m_Value()));
+}
+
 bool AMDGPUCodeGenPrepareImpl::visitBinaryOperator(BinaryOperator &I) {
   if (foldBinOpIntoSelect(I))
     return true;
 
-  if (UseMul24Intrin && replaceMulWithMul24(I))
-    return true;
-  if (tryNarrowMathIfNoOverflow(&I))
-    return true;
+  if (!shouldKeepMulForMad64(I)) {
----------------
PankajDwivedi-25 wrote:

This only suppresses the CGP mul24/narrowing rewrite. performMulCombine in SIISelLowering can still replace a divergent i64 ISD::MUL with build_pair(mul24, mulhi24) before tryFoldToMad64_32 runs? 

https://github.com/llvm/llvm-project/pull/226963


More information about the llvm-commits mailing list