[llvm] [AMDGPU][DAG] Allow extract_subvector for free register copy (PR #214152)
via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 7 02:23:56 PDT 2026
================
@@ -1456,8 +1456,8 @@ define amdgpu_kernel void @simple_write2_v4f32_superreg_align4(ptr addrspace(3)
; GFX1250S-UNALIGNED-NEXT: s_delay_alu instid0(VALU_DEP_1)
; GFX1250S-UNALIGNED-NEXT: v_lshl_add_u32 v4, v0, 4, s8
; GFX1250S-UNALIGNED-NEXT: s_wait_kmcnt 0x0
-; GFX1250S-UNALIGNED-NEXT: v_dual_mov_b32 v0, s2 :: v_dual_mov_b32 v1, s3
-; GFX1250S-UNALIGNED-NEXT: v_dual_mov_b32 v2, s0 :: v_dual_mov_b32 v3, s1
+; GFX1250S-UNALIGNED-NEXT: v_mov_b64_e32 v[0:1], s[2:3]
----------------
Shoreshen wrote:
Hi @arsenm from the latency model I notice that dual move is 1 cycle faster then mov_b64, but double the throughput.
The current regression is due to converting 64bit `COPY` into `V_MOV_B64_e32` in `postrapseudos` pass, a proper change I think would be lower `V_MOV_B64_e32` in `gcn-create-vopd` pass. This will convert all `V_MOV_B64_e32` to dual move. I'm not sure if this is prefered
https://github.com/llvm/llvm-project/pull/214152
More information about the llvm-commits
mailing list