[llvm] [AMDGPU] Fuse dword load + sign/zero-extension into subword load (PR #219019)
via llvm-commits
llvm-commits at lists.llvm.org
Wed Aug 26 16:55:27 PDT 2026
================
@@ -2861,6 +2881,122 @@ bool SILoadStoreOptimizerLegacy::runOnMachineFunction(MachineFunction &MF) {
.run(MF);
}
+enum SubwordExtensionType {
+ EXT_U8 = 0,
+ EXT_U16 = 1,
+ EXT_I8 = 2,
+ EXT_I16 = 3,
+ EXT_NONE = -1
+};
+
+// Determine the extension type based on the use instruction.
+// Returns EXT_NONE if the use pattern doesn't match a subword load.
+static SubwordExtensionType getExtensionType(const MachineInstr *UseInst) {
+ switch (UseInst->getOpcode()) {
+ case AMDGPU::S_AND_B32: {
+ // Check if it's masking with 0xff or 0xffff
+ const MachineOperand &MaskOp = UseInst->getOperand(2);
----------------
LU-JOHN wrote:
> Fair - this also gives me the naive "can we tablegen this" question
This transformation should be done after loads have been fused into wider loads. Loads that zero/sign-extend can't be fused with other loads. I think, tablegen is too early.
https://github.com/llvm/llvm-project/pull/219019
More information about the llvm-commits
mailing list