[llvm] [AMDGPU] Add DPP16 Row Share optimization for llvm.amdgcn.wave.shuffle (PR #177470)

Matt Arsenault via llvm-commits llvm-commits at lists.llvm.org
Tue Jan 27 02:21:24 PST 2026


================
@@ -553,6 +553,103 @@ static CallInst *rewriteCall(IRBuilderBase &B, CallInst &Old,
   return NewCall;
 }
 
+// Return true for sequences of instructions that effectively assign
+// each lane to its thread ID
+bool isThreadID(const GCNSubtarget *ST, Value *V) {
+  // Case 1:
+  //   wave32: mbcnt_lo(-1, 0)
+  //   wave64: mbcnt_hi(-1, mbcnt_lo(-1, 0))
+  ConstantInt *HiMask, *LoMask, *Input;
+  auto W32Pred = m_Intrinsic<Intrinsic::amdgcn_mbcnt_lo>(m_ConstantInt(LoMask),
+                                                         m_ConstantInt(Input));
+  auto W64Pred = m_Intrinsic<Intrinsic::amdgcn_mbcnt_hi>(
+      m_ConstantInt(HiMask), m_Intrinsic<Intrinsic::amdgcn_mbcnt_lo>(
+                                 m_ConstantInt(LoMask), m_ConstantInt(Input)));
+  if (ST->isWave32() && match(V, W32Pred) && LoMask->getSExtValue() == -1 &&
+      Input->getZExtValue() == 0)
----------------
arsenm wrote:

Can directly match the specific constant values in the match()

https://github.com/llvm/llvm-project/pull/177470


More information about the llvm-commits mailing list