[llvm] [AMDGPU] Fold v_perm pair into v_swap (PR #181966)

Matt Arsenault via llvm-commits llvm-commits at lists.llvm.org
Wed Apr 22 07:31:21 PDT 2026


================
@@ -843,6 +850,181 @@ MachineInstr *SIShrinkInstructions::matchSwap(MachineInstr &MovT) const {
   return nullptr;
 }
 
+// Matches two v_perms that together swap 16-bit halves between two inputs. For
+// example:
+//
+// v_perm_b32 v2, v0, v1, 0x5040100
+// v_perm_b32 v3, v0, v1, 0x7060302
+// =>
+// v_swap_b16 v0.l, v1.h
+MachineInstr *
+SIShrinkInstructions::matchSwapB16(MachineInstr &Perm,
+                                   PendingSwapMap &SwapCandidates) const {
+  assert(Perm.getOpcode() == AMDGPU::V_PERM_B32_e64);
+  if (IsPostRA)
+    return nullptr;
+
+  if (!ST->useRealTrue16Insts() ||
+      TII->pseudoToMCOpcode(AMDGPU::V_SWAP_B16) == -1)
----------------
arsenm wrote:

v_swap_b16 always existed for true16 targets, so this pseudoToMCOpcode check is redundant? 

https://github.com/llvm/llvm-project/pull/181966


More information about the llvm-commits mailing list