[llvm] AMDGPU: VALU data fast-forwarding needs no s_delay (PR #205481)

Jay Foad via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 29 03:48:18 PDT 2026


================
@@ -392,19 +552,42 @@ class AMDGPUInsertDelayAlu {
         State = DelayState();
       } else if (Type != OTHER) {
         DelayInfo Delay;
+        Register VccReg = AMDGPU::LaneMaskConstants::get(*ST).VccReg;
+        Register ExecReg = AMDGPU::LaneMaskConstants::get(*ST).ExecReg;
+
         // C-reuse: back-to-back WMMAs into the same C register forward the
         // accumulator in place, so the tied srcC read has no dependency. WMMA
         // implies GFX11+, so no explicit subtarget check is needed.
         bool IsWMMACReuse =
             PrevWMMAVDst.isValid() && (SII->isWMMA(MI) || SII->isSWMMAC(MI));
-        // TODO: Scan implicit uses too?
-        for (const auto &Op : MI.explicit_uses()) {
+
+        for (const auto &Op : MI.all_uses()) {
           if (Op.isReg()) {
             // One of the operands of the writelane is also the output operand.
             // This creates the insertion of redundant delays. Hence, we have to
             // ignore this operand.
             if (MI.getOpcode() == AMDGPU::V_WRITELANE_B32 && Op.isTied())
               continue;
+
+            Register Reg = Op.getReg();
+            unsigned OperandNo = MI.getOperandNo(&Op);
+            // Skip operands that are not part of the instruction definition
+            if (OperandNo >= MI.getDesc().getNumOperands() &&
+                !MI.getDesc().hasImplicitUseOfPhysReg(
+                    Reg == VccReg ? AMDGPU::VCC : Reg))
----------------
jayfoad wrote:

Why do you handle VccReg specially here? (This is related to what the AI code review said about "Wave32 inconsistency in the def-side filter".)

https://github.com/llvm/llvm-project/pull/205481


More information about the llvm-commits mailing list