[llvm] [AMDGPU] Add DPP16 Row Share optimization for llvm.amdgcn.wave.shuffle (PR #177470)
Matt Arsenault via llvm-commits
llvm-commits at lists.llvm.org
Tue Jan 27 02:21:24 PST 2026
================
@@ -553,6 +553,103 @@ static CallInst *rewriteCall(IRBuilderBase &B, CallInst &Old,
return NewCall;
}
+// Return true for sequences of instructions that effectively assign
+// each lane to its thread ID
+bool isThreadID(const GCNSubtarget *ST, Value *V) {
+ // Case 1:
+ // wave32: mbcnt_lo(-1, 0)
+ // wave64: mbcnt_hi(-1, mbcnt_lo(-1, 0))
+ ConstantInt *HiMask, *LoMask, *Input;
+ auto W32Pred = m_Intrinsic<Intrinsic::amdgcn_mbcnt_lo>(m_ConstantInt(LoMask),
+ m_ConstantInt(Input));
+ auto W64Pred = m_Intrinsic<Intrinsic::amdgcn_mbcnt_hi>(
+ m_ConstantInt(HiMask), m_Intrinsic<Intrinsic::amdgcn_mbcnt_lo>(
+ m_ConstantInt(LoMask), m_ConstantInt(Input)));
+ if (ST->isWave32() && match(V, W32Pred) && LoMask->getSExtValue() == -1 &&
+ Input->getZExtValue() == 0)
----------------
arsenm wrote:
Can directly match the specific constant values in the match()
https://github.com/llvm/llvm-project/pull/177470
More information about the llvm-commits
mailing list