[llvm] [AMDGPU][DAG] Allow extract_subvector for free register copy (PR #214152)

via llvm-commits llvm-commits at lists.llvm.org
Mon Sep 7 02:23:56 PDT 2026


================
@@ -1456,8 +1456,8 @@ define amdgpu_kernel void @simple_write2_v4f32_superreg_align4(ptr addrspace(3)
 ; GFX1250S-UNALIGNED-NEXT:    s_delay_alu instid0(VALU_DEP_1)
 ; GFX1250S-UNALIGNED-NEXT:    v_lshl_add_u32 v4, v0, 4, s8
 ; GFX1250S-UNALIGNED-NEXT:    s_wait_kmcnt 0x0
-; GFX1250S-UNALIGNED-NEXT:    v_dual_mov_b32 v0, s2 :: v_dual_mov_b32 v1, s3
-; GFX1250S-UNALIGNED-NEXT:    v_dual_mov_b32 v2, s0 :: v_dual_mov_b32 v3, s1
+; GFX1250S-UNALIGNED-NEXT:    v_mov_b64_e32 v[0:1], s[2:3]
----------------
Shoreshen wrote:

Hi @arsenm from the latency model I notice that dual move is 1 cycle faster then mov_b64, but double the throughput.

The current regression is due to converting 64bit `COPY` into `V_MOV_B64_e32` in `postrapseudos` pass, a proper change I think would be lower `V_MOV_B64_e32` in `gcn-create-vopd` pass. This will convert all `V_MOV_B64_e32` to dual move. I'm not sure if this is prefered

https://github.com/llvm/llvm-project/pull/214152


More information about the llvm-commits mailing list