[llvm] [AMDGPU] Fuse mad64_32 from 24-bit multiply-add (PR #225114)

Pankaj Dwivedi via llvm-commits llvm-commits at lists.llvm.org
Thu Sep 24 02:52:34 PDT 2026


================
@@ -17440,6 +17440,103 @@ static SDValue tryFoldMADwithSRL(SelectionDAG &DAG, const SDLoc &SL,
                      DAG.getZeroExtendInReg(AddRHS, SL, MVT::i32), false);
 }
 
+// performMulCombine may lower an i64 ISD::MUL whose operands are known to fit
+// in 24 bits into 24-bit multiply nodes before the enclosing add is combined:
+//
+//   mul i64 x, y  -->  build_pair (mul_u24 x, y), (mulhi_u24 x, y)
+//
+// That rewrite removes the ISD::MUL that tryFoldToMad64_32 keys on, so the
+// fused mad is never formed. Recover it by matching the 24-bit multiply forms
+// directly and folding (add mul24(x, y), z) --> mad_[iu]64_[iu]32 x, y, z.
+static SDValue tryFoldMul24ToMad64_32(SDNode *N, SelectionDAG &DAG,
----------------
PankajDwivedi-25 wrote:

Looks like something similar attempted here #72393 folds a related add (i64 MUL_U24), z pattern, but only on GFX9+ when both operands are known 24-bit (// exclude pre-GFX9 where it was slow). This DAG combine should follow the same rule? only fold when the mad is actually cheaper, and only when the operands already fit in 24 bits so you don’t need the ANDs. Otherwise leave the mul24 pair?

https://github.com/llvm/llvm-project/pull/225114


More information about the llvm-commits mailing list