[llvm] [AMDGPU] Fuse dword load + sign/zero-extension into subword load (PR #219019)

via llvm-commits llvm-commits at lists.llvm.org
Wed Aug 26 16:55:27 PDT 2026


================
@@ -2861,6 +2881,122 @@ bool SILoadStoreOptimizerLegacy::runOnMachineFunction(MachineFunction &MF) {
       .run(MF);
 }
 
+enum SubwordExtensionType {
+  EXT_U8 = 0,
+  EXT_U16 = 1,
+  EXT_I8 = 2,
+  EXT_I16 = 3,
+  EXT_NONE = -1
+};
+
+// Determine the extension type based on the use instruction.
+// Returns EXT_NONE if the use pattern doesn't match a subword load.
+static SubwordExtensionType getExtensionType(const MachineInstr *UseInst) {
+  switch (UseInst->getOpcode()) {
+  case AMDGPU::S_AND_B32: {
+    // Check if it's masking with 0xff or 0xffff
+    const MachineOperand &MaskOp = UseInst->getOperand(2);
----------------
LU-JOHN wrote:

> Fair - this also gives me the naive "can we tablegen this" question

This transformation should be done after loads have been fused into wider loads.  Loads that zero/sign-extend can't be fused with other loads.  I think, tablegen is too early.

https://github.com/llvm/llvm-project/pull/219019


More information about the llvm-commits mailing list