[llvm] [AMDGPU] Fuse dword load + sign/zero-extension into subword load (PR #219019)
via llvm-commits
llvm-commits at lists.llvm.org
Thu Aug 27 07:17:39 PDT 2026
LU-JOHN wrote:
> SILoadStoreOptimizer is surely not the right place to implement something like this. It should be done as an IR optimization, or at instruction selection time. How do other targets handle this?
>
> Edit: or if this is completely specific to kernel argument loading and the sub-dword scalar loads that were new in GFX12, then the kernel argument loading code should be taught how to use them.
I tried implementing this earlier in AMDGPULowerKernelArguments.cpp, but I found that it inhibits fusing adjacent loads into a wider load. Thus, I implemented this immediately after load/store fusion has been done in SILoadStoreOptimizer.
https://github.com/llvm/llvm-project/pull/219019
More information about the llvm-commits
mailing list