[llvm] [AMDGPU] Break v2i32 and/or/xor ops into i32 pairs (PR #191422)

Matt Arsenault via llvm-commits llvm-commits at lists.llvm.org
Thu May 7 02:21:46 PDT 2026


================
@@ -741,6 +741,329 @@ define amdgpu_ps half @test_and_or_b16(i16 %a, i16 %b, i16 %c) {
   %ret_cast = bitcast i16 %or1 to half
   ret half %ret_cast
 }
+
+; ========= v2i32-i32 pattern tests =========
+
+define i32 @v_v2i32_or_i32_xor(<2 x i32> %a, <2 x i32> %b) {
+; GFX950-LABEL: v_v2i32_or_i32_xor:
+; GFX950:       ; %bb.0:
+; GFX950-NEXT:    s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; GFX950-NEXT:    v_or_b32_e32 v1, v1, v3
+; GFX950-NEXT:    v_bitop3_b32 v0, v0, v1, v2 bitop3:0x36
+; GFX950-NEXT:    s_setpc_b64 s[30:31]
+;
+; GFX1250-LABEL: v_v2i32_or_i32_xor:
+; GFX1250:       ; %bb.0:
+; GFX1250-NEXT:    s_wait_loadcnt_dscnt 0x0
+; GFX1250-NEXT:    s_wait_kmcnt 0x0
+; GFX1250-NEXT:    v_or_b32_e32 v1, v1, v3
+; GFX1250-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1250-NEXT:    v_bitop3_b32 v0, v0, v1, v2 bitop3:0x36
----------------
arsenm wrote:

I still think it would be worth trying to directly match vector bitop3 in the selector 

https://github.com/llvm/llvm-project/pull/191422


More information about the llvm-commits mailing list