[llvm] AMDGPU: VALU data fast-forwarding needs no s_delay (PR #205481)
Jay Foad via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 29 03:48:18 PDT 2026
================
@@ -392,19 +552,42 @@ class AMDGPUInsertDelayAlu {
State = DelayState();
} else if (Type != OTHER) {
DelayInfo Delay;
+ Register VccReg = AMDGPU::LaneMaskConstants::get(*ST).VccReg;
+ Register ExecReg = AMDGPU::LaneMaskConstants::get(*ST).ExecReg;
+
// C-reuse: back-to-back WMMAs into the same C register forward the
// accumulator in place, so the tied srcC read has no dependency. WMMA
// implies GFX11+, so no explicit subtarget check is needed.
bool IsWMMACReuse =
PrevWMMAVDst.isValid() && (SII->isWMMA(MI) || SII->isSWMMAC(MI));
- // TODO: Scan implicit uses too?
- for (const auto &Op : MI.explicit_uses()) {
+
+ for (const auto &Op : MI.all_uses()) {
if (Op.isReg()) {
// One of the operands of the writelane is also the output operand.
// This creates the insertion of redundant delays. Hence, we have to
// ignore this operand.
if (MI.getOpcode() == AMDGPU::V_WRITELANE_B32 && Op.isTied())
continue;
+
+ Register Reg = Op.getReg();
+ unsigned OperandNo = MI.getOperandNo(&Op);
+ // Skip operands that are not part of the instruction definition
+ if (OperandNo >= MI.getDesc().getNumOperands() &&
+ !MI.getDesc().hasImplicitUseOfPhysReg(
+ Reg == VccReg ? AMDGPU::VCC : Reg))
----------------
jayfoad wrote:
Why do you handle VccReg specially here? (This is related to what the AI code review said about "Wave32 inconsistency in the def-side filter".)
https://github.com/llvm/llvm-project/pull/205481
More information about the llvm-commits
mailing list