[llvm] [AMDGPU] Do not commute DPP instructions with a non-identity dpp_ctrl (PR #218393)

Arseniy Obolenskiy via llvm-commits llvm-commits at lists.llvm.org
Mon Aug 24 22:05:25 PDT 2026


================
@@ -228,6 +228,22 @@ define amdgpu_kernel void @dpp_src1_sgpr(ptr addrspace(1) %out, i16 %in) {
   ret void
 }
 
+; clamp forces VOP3 encoding so src1 takes an inline constant, but folding
+; 1.0 there needs a commute that would move row_shl:1 onto %x instead.
+; GCN-LABEL: {{^}}dpp_src1_imm_no_commute:
+; GFX9GFX10: v_mov_b32_dpp {{v[0-9]+}}, {{v[0-9]+}} row_shl:1 row_mask:0xf bank_mask:0xf bound_ctrl:1
+; GFX9GFX10: v_add_f32_e64 v0, {{v[0-9]+}}, v0 clamp
+; GFX11-TRUE16: v_add_f32_e64_dpp v0, {{v[0-9]+}}, v0 clamp row_shl:1 row_mask:0xf bank_mask:0xf bound_ctrl:1
+; GFX11-FAKE16: v_add_f32_e64_dpp v0, {{v[0-9]+}}, v0 clamp row_shl:1 row_mask:0xf bank_mask:0xf bound_ctrl:1
----------------
aobolensk wrote:

This is outside of the scope of current fix, on gfx11 DPP operand gets fused directly into src0 of the VOP3 form. I have removed the CHECKs for that

https://github.com/llvm/llvm-project/pull/218393


More information about the llvm-commits mailing list