[llvm] [AMDGPU] Break v2i32 and/or/xor ops into i32 pairs (PR #191422)
Matt Arsenault via llvm-commits
llvm-commits at lists.llvm.org
Thu May 7 02:21:46 PDT 2026
================
@@ -741,6 +741,329 @@ define amdgpu_ps half @test_and_or_b16(i16 %a, i16 %b, i16 %c) {
%ret_cast = bitcast i16 %or1 to half
ret half %ret_cast
}
+
+; ========= v2i32-i32 pattern tests =========
+
+define i32 @v_v2i32_or_i32_xor(<2 x i32> %a, <2 x i32> %b) {
+; GFX950-LABEL: v_v2i32_or_i32_xor:
+; GFX950: ; %bb.0:
+; GFX950-NEXT: s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; GFX950-NEXT: v_or_b32_e32 v1, v1, v3
+; GFX950-NEXT: v_bitop3_b32 v0, v0, v1, v2 bitop3:0x36
+; GFX950-NEXT: s_setpc_b64 s[30:31]
+;
+; GFX1250-LABEL: v_v2i32_or_i32_xor:
+; GFX1250: ; %bb.0:
+; GFX1250-NEXT: s_wait_loadcnt_dscnt 0x0
+; GFX1250-NEXT: s_wait_kmcnt 0x0
+; GFX1250-NEXT: v_or_b32_e32 v1, v1, v3
+; GFX1250-NEXT: s_delay_alu instid0(VALU_DEP_1)
+; GFX1250-NEXT: v_bitop3_b32 v0, v0, v1, v2 bitop3:0x36
----------------
arsenm wrote:
I still think it would be worth trying to directly match vector bitop3 in the selector
https://github.com/llvm/llvm-project/pull/191422
More information about the llvm-commits
mailing list