[llvm] [AMDGPU] merge 16bit mov pairs in post-RA peephole (PR #208625)

Guo Chen via llvm-commits llvm-commits at lists.llvm.org
Wed Aug 26 07:20:20 PDT 2026


broxigarchen wrote:

Sorry for the delay. Updated the patch including:
1. As discussed previously, created a new pass before waitcnt pass so that the waitcnt is set correctly
2. Found a problem in VK cts testing, that using v_pack_b32_f16 could cause problem in lane shuffle tests. The register value could be a denorm and get flushed away. Replaced with v_perm_b32 solves this problem, but the caveats is that we have the constant bus restriction

https://github.com/llvm/llvm-project/pull/208625


More information about the llvm-commits mailing list