[llvm] [AMDGPU/GISEL] Legalizing sub-dword scalar load for pre-gfx12 arch (PR #195800)
Steffen Larsen via llvm-commits
llvm-commits at lists.llvm.org
Tue May 5 00:44:54 PDT 2026
================
@@ -7515,6 +7515,13 @@ bool AMDGPULegalizerInfo::legalizeSBufferLoad(LegalizerHelper &Helper,
// The 8-bit and 16-bit scalar buffer load instructions have 32-bit
// destination register.
Dst = B.getMRI()->createGenericVirtualRegister(LLT::scalar(32));
+ } else if (Size < 32) {
+ // No native sub-dword scalar buffer load on this subtarget.
+ // Widen to a 32-bit load; a G_TRUNC is inserted after to recover the
+ // original narrow type. s8/s16 are not valid SGPR register types.
+ assert(Size == 8 || Size == 16);
+ Opc = AMDGPU::G_AMDGPU_S_BUFFER_LOAD;
+ Dst = B.getMRI()->createGenericVirtualRegister(LLT::scalar(32));
} else {
----------------
steffenlarsen wrote:
Minor preference since the two branches overlap. You could also make it a larger ternary.
```suggestion
if (Size < 32) {
assert(Size == 8 || Size == 16);
if (ST.hasScalarSubwordLoads()) {
// Native sub-dword scalar buffer load is available on this subtarget.
Opc = Size == 8 ? AMDGPU::G_AMDGPU_S_BUFFER_LOAD_UBYTE
: AMDGPU::G_AMDGPU_S_BUFFER_LOAD_USHORT;
} else {
// No native sub-dword scalar buffer load on this subtarget.
// Widen to a 32-bit load; a G_TRUNC is inserted after to recover the
// original narrow type. s8/s16 are not valid SGPR register types.
Opc = AMDGPU::G_AMDGPU_S_BUFFER_LOAD;
}
// The 8-bit and 16-bit scalar buffer load instructions have 32-bit
// destination register.
Dst = B.getMRI()->createGenericVirtualRegister(LLT::scalar(32));
} else {
```
Alternatively, you could change the initial value of `Opt` before the conditional to `AMDGPU::G_AMDGPU_S_BUFFER_LOAD` and only keep the `if (ST.hasScalarSubwordLoads())`. In this case you could also remove the assignment in the else-branch below this.
https://github.com/llvm/llvm-project/pull/195800
More information about the llvm-commits
mailing list