[llvm] [AMDGPU] Post-RA Peephole for Two-Address Instructions (PR #207731)

Carl Ritson via llvm-commits llvm-commits at lists.llvm.org
Wed Jul 15 00:04:43 PDT 2026


================
@@ -803,6 +1016,36 @@ bool SIPreEmitPeephole::run(MachineFunction &MF, MachineLoopInfo *LoopInfo) {
       }
     }
 
+    // Delayed optimization to replace two-address instructions followed by
+    // copies with a three-address instruction that directly writes to the
+    // copied registers instead. The pattern we seek to optimize is:
+    // $vgpr2_vgpr3 = V_FMAC_F64 ..., $vgpr_2_vgpr3
+    // $vgpr0 = V_MOV_B32 $vgrp2
+    // $vgpr1 = V_MOV_B32 $vgrp3
+    //
+    // This can be optimized to:
+    // $vgpr0_vgpr1 = V_FMA_F64 ..., $vgpr2_vgpr3
+    //
+    // The TwoAddressInstruction pass handles some cases, but bails in more
+    // complex cases.
+    SmallVector<ThreeAddressCandidate> Candidates;
+    // First, collect the candidates. We can't perform the transformation right
----------------
perlfu wrote:

See comment below, but this is not accurate?

https://github.com/llvm/llvm-project/pull/207731


More information about the llvm-commits mailing list