[llvm] [AMDGPU] Select the high-half 16-bit packing idiom to v_lshl_or_b32 (PR #206058)

Barbara Mitic via llvm-commits llvm-commits at lists.llvm.org
Fri Jul 3 02:44:19 PDT 2026


barbara-amd wrote:

> Is there any delta in the tests on the true16 path with the new pattern? I didn't see any. Maybe the pattern is not being used.

Yes, it's used - though on the true16 path its footprint is small: across the changed tests it fires in exactly one existing test, global-load-xcnt.ll (test_i8load_v4i8store), on the SelectionDAG +real-true16 run (gfx1250 -mattr=+real-true16). 

The delta there is the fusion of the (hi << 16) | z idiom. Without the pattern the pack is two ops:
v_lshlrev_b32_e32 v0, 16, v0
v_or_b32_e32 v0, v1, v0
global_store_b32 v[8:9], v0, off

With the pattern it's one:
v_or_b16 v1.h, v0.l, v1.h
global_store_b32 v[8:9], v1, off

I also compared the v_or_b16 sequence against simply emitting v_lshl_or_b32 for true16:
v_lshl_or_b32 v0, v0, 16, v1
global_store_b32 v[8:9], v0, off

Since Jay has already approved this, my suggestion is, if there are no other objections, to keep it as-is for now, and later  we can think of a more optimal way of doing this for true16, if needed.

https://github.com/llvm/llvm-project/pull/206058


More information about the llvm-commits mailing list