[llvm] 4920eea - [AMDGPU] Reject image-load merges when either op has TFE/LWE (#210583)

via llvm-commits llvm-commits at lists.llvm.org
Thu Aug 27 07:02:53 PDT 2026


Author: Yudistira Putra
Date: 2026-08-27T16:02:47+02:00
New Revision: 4920eeabdfc4b41ae4174156e203953f1fdb1b32

URL: https://github.com/llvm/llvm-project/commit/4920eeabdfc4b41ae4174156e203953f1fdb1b32
DIFF: https://github.com/llvm/llvm-project/commit/4920eeabdfc4b41ae4174156e203953f1fdb1b32.diff

LOG: [AMDGPU] Reject image-load merges when either op has TFE/LWE (#210583)

## Summary

MIMG instructions with TFE or LWE enabled are now excluded during
mergeable candidate collection in `collectMergeableInsts()`, preventing
them from ever entering the merge candidate list in
`SILoadStoreOptimizer`.

## Problem

TFE (Texture Fetch Error) and LWE (Load Word Error) add status-result
semantics to image load instructions — the result register includes an
additional status word alongside the texture data. The image merge path
in `SILoadStoreOptimizer` does not reconstruct or represent these status
lanes. The previous eligibility logic in `dmasksCanBeCombined()` only
checked the leading instruction (`CI.I`) for TFE/LWE, which allowed the
asymmetric ordinary→TFE/LWE ordering to enter the merge path.

## Fix

- `collectMergeableInsts()` now skips MIMG instructions with enabled TFE
or LWE before they are added to the mergeable candidate list.
- Because TFE/LWE instructions never enter the candidate list, both
orderings (ordinary→TFE/LWE and TFE/LWE→ordinary) are excluded from pair
combination.
- No changes were made to `dmasksCanBeCombined()` relative to upstream —
its existing `CI.I`-only TFE/LWE check remains as defense-in-depth.
- `AMDGPUImageIntrinsicOptimizer` is not part of this change.

## Testing

Focused lit tests run with `llvm-lit -sv`:
- `llvm/test/CodeGen/AMDGPU/merge-image-load.mir`
- `llvm/test/CodeGen/AMDGPU/merge-image-load-gfx10.mir`
- `llvm/test/CodeGen/AMDGPU/merge-image-load-gfx11.mir`
- `llvm/test/CodeGen/AMDGPU/merge-image-load-gfx12.mir`

Coverage matrix (in `merge-image-load.mir`):
- **Newly added regressions**:
- `image_load_tfe_second_not_merged` — ordinary → TFE: **not merged**
(Passed)
- `image_load_lwe_second_not_merged` — ordinary → LWE: **not merged**
(Passed)
- **Pre-existing reverse-order tests**:
  - `image_load_not_merged_7` — TFE → ordinary: **not merged** (Passed)
  - `image_load_not_merged_8` — LWE → ordinary: **not merged** (Passed)
- **Pre-existing status-free positive merge tests**:
- `image_load_merged_v1v3` (and others) — ordinary → ordinary:
**merged** (Passed)

## AI tool disclosure

AI-assisted tools were used during the investigation, implementation
support, and drafting of this change. I personally reviewed and
understand the final code and tests, ran the validation listed above,
and remain responsible for the contribution.

Fixes #187335

Signed-off-by: Yudistira Putra <85178972+Yudis-bit at users.noreply.github.com>

Added: 
    

Modified: 
    llvm/lib/Target/AMDGPU/SILoadStoreOptimizer.cpp
    llvm/test/CodeGen/AMDGPU/merge-image-load.mir

Removed: 
    


################################################################################
diff  --git a/llvm/lib/Target/AMDGPU/SILoadStoreOptimizer.cpp b/llvm/lib/Target/AMDGPU/SILoadStoreOptimizer.cpp
index 57228e9a72b7e..4bd6a9a0fe867 100644
--- a/llvm/lib/Target/AMDGPU/SILoadStoreOptimizer.cpp
+++ b/llvm/lib/Target/AMDGPU/SILoadStoreOptimizer.cpp
@@ -2686,6 +2686,16 @@ SILoadStoreOptimizer::collectMergeableInsts(
         LLVM_DEBUG(dbgs() << "Skip tbuffer with unknown format: " << MI);
         continue;
       }
+    } else if (InstClass == MIMG) {
+      // Do not merge MIMG instructions with tfe or lwe enabled.
+      // TFE/LWE add a status result that the image merge path does not model.
+      const auto *TFEOp = TII->getNamedOperand(MI, AMDGPU::OpName::tfe);
+      if (TFEOp && TFEOp->getImm())
+        continue;
+
+      const auto *LWEOp = TII->getNamedOperand(MI, AMDGPU::OpName::lwe);
+      if (LWEOp && LWEOp->getImm())
+        continue;
     }
 
     CombineInfo CI;

diff  --git a/llvm/test/CodeGen/AMDGPU/merge-image-load.mir b/llvm/test/CodeGen/AMDGPU/merge-image-load.mir
index 2d3f9b0f482a7..dfebc1cc4a608 100644
--- a/llvm/test/CodeGen/AMDGPU/merge-image-load.mir
+++ b/llvm/test/CodeGen/AMDGPU/merge-image-load.mir
@@ -485,3 +485,38 @@ body:             |
 ...
 ---
 
+# GFX9-LABEL: name: image_load_tfe_second_not_merged
+# GFX9: %{{[0-9]+}}:vgpr_32 = IMAGE_LOAD_V1_V4 %5, %3, 8, 0, 0, 0, 0, 0, -1, 0, implicit $exec :: (dereferenceable load (s32), addrspace 4)
+# GFX9: %{{[0-9]+}}:vreg_128 = IMAGE_LOAD_V4_V4 %5, %3, 7, 0, 0, 0, 1, 0, -1, 0, implicit $exec :: (dereferenceable load (s96), align 16, addrspace 4)
+
+name: image_load_tfe_second_not_merged
+body:             |
+  bb.0.entry:
+    %0:sgpr_64 = COPY $sgpr0_sgpr1
+    %1:sreg_64_xexec = S_LOAD_DWORDX2_IMM %0, 36, 0
+    %2:sgpr_128 = COPY $sgpr96_sgpr97_sgpr98_sgpr99
+    %3:sgpr_256 = S_LOAD_DWORDX8_IMM %1, 208, 0
+    %4:vgpr_32 = COPY %2.sub3
+    %5:vreg_128 = BUFFER_LOAD_DWORDX4_OFFSET %2, 0, 0, 0, 0, implicit $exec :: (dereferenceable invariant load (s128))
+    %6:vgpr_32 = IMAGE_LOAD_V1_V4 %5, %3, 8, 0, 0, 0, 0, 0, -1, 0, implicit $exec :: (dereferenceable load (s32), addrspace 4)
+    %7:vreg_128 = IMAGE_LOAD_V4_V4 %5, %3, 7, 0, 0, 0, 1, 0, -1, 0, implicit $exec :: (dereferenceable load (s96), align 16, addrspace 4)
+...
+---
+
+# GFX9-LABEL: name: image_load_lwe_second_not_merged
+# GFX9: %{{[0-9]+}}:vgpr_32 = IMAGE_LOAD_V1_V4 %5, %3, 8, 0, 0, 0, 0, 0, -1, 0, implicit $exec :: (dereferenceable load (s32), addrspace 4)
+# GFX9: %{{[0-9]+}}:vreg_128 = IMAGE_LOAD_V4_V4 %5, %3, 7, 0, 0, 0, 0, 1, -1, 0, implicit $exec :: (dereferenceable load (s96), align 16, addrspace 4)
+
+name: image_load_lwe_second_not_merged
+body:             |
+  bb.0.entry:
+    %0:sgpr_64 = COPY $sgpr0_sgpr1
+    %1:sreg_64_xexec = S_LOAD_DWORDX2_IMM %0, 36, 0
+    %2:sgpr_128 = COPY $sgpr96_sgpr97_sgpr98_sgpr99
+    %3:sgpr_256 = S_LOAD_DWORDX8_IMM %1, 208, 0
+    %4:vgpr_32 = COPY %2.sub3
+    %5:vreg_128 = BUFFER_LOAD_DWORDX4_OFFSET %2, 0, 0, 0, 0, implicit $exec :: (dereferenceable invariant load (s128))
+    %6:vgpr_32 = IMAGE_LOAD_V1_V4 %5, %3, 8, 0, 0, 0, 0, 0, -1, 0, implicit $exec :: (dereferenceable load (s32), addrspace 4)
+    %7:vreg_128 = IMAGE_LOAD_V4_V4 %5, %3, 7, 0, 0, 0, 0, 1, -1, 0, implicit $exec :: (dereferenceable load (s96), align 16, addrspace 4)
+...
+---


        


More information about the llvm-commits mailing list