[llvm] [AMDGPU] Reduce unnecessary src2 bridge copies in RewriteMFMAFormStage (PR #215806)
Petr Kurapov via llvm-commits
llvm-commits at lists.llvm.org
Wed Sep 9 04:05:05 PDT 2026
================
@@ -2317,19 +2321,21 @@ void GCNSchedStage::modifyRegionSchedule(unsigned RegionIdx,
}
/// Returns true if reaching def \p RD will be in AGPR form after the rewrite
-/// and so needs no bridge copy: a candidate MFMA in \p RewriteSet, an
-/// AV_MOV_*_IMM_PSEUDO, or a copy from a candidate src2 reg in \p CandSrc2Regs.
-/// A non-candidate MFMA stays in VGPR form and still needs a bridge.
+/// and so needs no bridge copy.
static bool isReachingDefAGPRForm(
MachineInstr *RD, const SmallPtrSetImpl<MachineInstr *> &RewriteSet,
const DenseSet<Register> &CandSrc2Regs, const SIInstrInfo &TII) {
if (TII.isMAI(*RD))
return RewriteSet.contains(RD);
- if (RD->getOpcode() == AMDGPU::AV_MOV_B32_IMM_PSEUDO ||
- RD->getOpcode() == AMDGPU::AV_MOV_B64_IMM_PSEUDO)
- return true;
if (RD->isCopy() && CandSrc2Regs.contains(RD->getOperand(1).getReg()))
return true;
+ // Instructions whose def operand accepts AGPR (DS_READ, AV_MOV, etc.)
+ // will produce AGPR after reclassification - no bridge copy needed.
+ const SIRegisterInfo &SRI =
+ static_cast<const SIRegisterInfo &>(TII.getRegisterInfo());
+ const TargetRegisterClass *DefRC = TII.getRegClass(RD->getDesc(), 0);
+ if (DefRC && (SRI.isAGPRClass(DefRC) || SRI.isVectorSuperClass(DefRC)))
+ return true;
----------------
kurapov-peter wrote:
This change seems to unmask (by skipping the cost check) a behavior with multiple reaching defs. The cost estimation doesn't consider all the bridge copies that need to be inserted on all the paths reaching a rewritten mfma and thus accepts a "disadvantageous" rewrite. Well, that's one way of seeing it.
Something like:
```
acc = COPY src
loop:
IF cond:
acc = DS_READ(ptr)
t0 = MFMA(a, b, acc) # reached by both defs
t1 = MFMA(a, b, t0)
consume(t1.sub1) # a vgpr use
```
would emit two bridge copies because of the first copy, although, in principle, could have retagged it and convert everything to agprs without any bridge copies (avoiding the one on ds read). This is pretty-much what's happening after post-ra copy propagation, but the stage itself would introduce one after ds_read. I'd expect there are some situations when the clean-up doesn't happen and the estimation fails.
https://github.com/llvm/llvm-project/pull/215806
More information about the llvm-commits
mailing list