[llvm] [AMDGPU] merge 16bit mov pairs in post-RA peephole (PR #208625)
Guo Chen via llvm-commits
llvm-commits at lists.llvm.org
Wed Aug 26 07:20:20 PDT 2026
broxigarchen wrote:
Sorry for the delay. Updated the patch including:
1. As discussed previously, created a new pass before waitcnt pass so that the waitcnt is set correctly
2. Found a problem in VK cts testing, that using v_pack_b32_f16 could cause problem in lane shuffle tests. The register value could be a denorm and get flushed away. Replaced with v_perm_b32 solves this problem, but the caveats is that we have the constant bus restriction
https://github.com/llvm/llvm-project/pull/208625
More information about the llvm-commits
mailing list