[llvm] [DAGCombiner] Refine an extload to a zextload under zext(trunc(...)) (PR #225053)

Paweł Bylica via llvm-commits llvm-commits at lists.llvm.org
Mon Sep 21 02:51:04 PDT 2026


https://github.com/chfast created https://github.com/llvm/llvm-project/pull/225053

When a narrow load also has a user at its extended width, the load is retyped
into an any-extending load early and the remaining narrow use becomes a
truncate. A later `zext` of that truncate then has nothing to fold against:
the extload's high bits are undefined, so the mask cannot be removed, and
`reduceLoadWidth` refuses to narrow a load with several users.

Refine the extload to a zextload in place instead. The truncate discards only
the load's extension bits, which are undefined, so defining them as zero is
valid for every existing user and it makes the `zext` a no-op. The load keeps
its value type and only its extension kind changes, unlike
`tryToFoldExtOfExtload`, which widens the load and therefore requires a single
use.

`EXTLOAD` is the only extension kind this applies to:

- a zextload already has those bits known zero, so the known-bits
  `(zext (truncate x))` fold collapses that case before this one is reached;
- a sextload has them defined, so zeroing them would miscompile;
- undefined bits are precisely the case analysis cannot settle, because they
  become zero only by performing the rewrite.

On AArch64 this turns `and x1, x0, #0xff` into `mov x1, x0`. It also fixes
`llvm/test/CodeGen/BPF/remove_truncate_9.ll` under
`-combiner-topological-sorting` for `-mcpu=v4`, where the producer-first
worklist order reaches this shape on a target whose tests assert that no
zero-extension instruction is emitted.

### Notes for review

- `N` is replaced before the load is retyped, as `tryToFoldExtOfLoad` and
  `tryToFoldExtOfExtload` do. `getExtLoad` can return an existing zextload of
  the same address, in which case replacing the load first CSEs the truncate,
  and `N` along with it, into pre-existing nodes and deletes `N`. There is a
  regression test for that.
- The truncate's debug values are transferred to the replacement, since the
  truncate dies with `N`. There is a regression test for that too.
- The legality condition matches `tryToFoldExtOfExtload`. Volatile and atomic
  loads are not excluded, since the access is unchanged and only the
  in-register extension differs, which is how the in-place retype in the
  `(and (load), mask)` fold already treats them. Happy to tighten this if
  reviewers prefer.
- The sign-extending analogue is deliberately not done: refining to a
  `SEXTLOAD` forces users that want the raw narrow value to materialize a
  copy, costing an instruction on x86-64 with no measured win elsewhere.

### Testing

`llvm/test/CodeGen` and `llvm/test/DebugInfo` are clean (33715 tests). The fold
fires 23 times across the in-tree suite. Each added test was checked to
discriminate: reverting the fold, or the specific behaviour it names, changes
its output.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

>From 17e2f363a1db1eddf590c09fad2b7fc5370ead50 Mon Sep 17 00:00:00 2001
From: =?UTF-8?q?Pawe=C5=82=20Bylica?= <pawel at hepcolgum.band>
Date: Fri, 18 Sep 2026 16:21:28 +0200
Subject: [PATCH] [DAGCombiner] Refine an extload to a zextload under
 zext(trunc(...))

When a narrow load also has a user at its extended width, the load is
retyped into an any-extending load early and the remaining narrow use
becomes a truncate. A later zext of that truncate then has nothing to
fold against: the extload's high bits are undefined, so the mask cannot
be removed, and reduceLoadWidth refuses to narrow a load with several
users.

Refine the extload to a zextload in place instead. The truncate discards
only the load's extension bits, which are undefined, so defining them as
zero is valid for every existing user and it makes the zext a no-op. The
load keeps its value type and only its extension kind changes, unlike
tryToFoldExtOfExtload, which widens the load and therefore requires a
single use.

EXTLOAD is the only extension kind this applies to. A zextload already
has those bits known zero, so the known-bits (zext (truncate x)) fold
collapses that case before this one is reached, and a sextload has them
defined, so zeroing them would miscompile. Undefined bits are precisely
the case analysis cannot settle, because they become zero only by
performing the rewrite.

The legality condition matches tryToFoldExtOfExtload: a simple load can
be retyped before legalization without proving the result legal. Volatile
and atomic loads are not excluded, since the access itself is unchanged
and only the in-register extension differs, which is also how the
in-place retype in the (and (load), mask) fold treats them.

N is replaced before the load is retyped, as tryToFoldExtOfLoad and
tryToFoldExtOfExtload do. getExtLoad can return an existing zextload of
the same address, in which case replacing the load first CSEs the
truncate, and N along with it, into pre-existing nodes and deletes N.
The truncate's debug values are transferred to the replacement for the
same reason: it dies with N.

On AArch64 this turns "and x1, x0, #0xff" into "mov x1, x0". It also
fixes llvm/test/CodeGen/BPF/remove_truncate_9.ll under
-combiner-topological-sorting for -mcpu=v4, where the producer-first
worklist order reaches this shape on a target whose tests assert that no
zero-extension instruction is emitted.

The sign-extending analogue is deliberately not done: refining to a
SEXTLOAD forces users that want the raw narrow value to materialize a
copy, which costs an instruction on x86-64 with no measured win
elsewhere.

Assisted-by: Claude Code
---
 llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp |  63 ++++++++++
 .../CodeGen/AArch64/zext-trunc-extload-dbg.ll |  42 +++++++
 .../CodeGen/AArch64/zext-trunc-extload.ll     | 117 ++++++++++++++++++
 3 files changed, 222 insertions(+)
 create mode 100644 llvm/test/CodeGen/AArch64/zext-trunc-extload-dbg.ll
 create mode 100644 llvm/test/CodeGen/AArch64/zext-trunc-extload.ll

diff --git a/llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp b/llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
index d4a7403a6b25f..a4f1a1d8b2aeb 100644
--- a/llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
@@ -15627,6 +15627,65 @@ static SDValue tryToFoldExtOfExtload(SelectionDAG &DAG, DAGCombiner &Combiner,
   return SDValue(N, 0); // Return N so it doesn't get rechecked!
 }
 
+// fold (zext (truncate (extload x))) -> (zext (zextload x)) when the truncate
+// only discards the load's extension bits.
+//
+// At heart this retypes the load: EXTLOAD becomes ZEXTLOAD, keeping the same
+// value type, and the zext of the truncate is then redundant. It is the one
+// extension kind for which that is both needed and allowed:
+//  - ZEXTLOAD already has its extension bits known zero, so the known-bits
+//    (zext (truncate x)) fold above collapses that case and this is never
+//    reached.
+//  - SEXTLOAD has *defined* sign bits, so zeroing them would miscompile.
+//  - EXTLOAD has undefined extension bits, which computeKnownBits cannot
+//    report as zero (they only become zero by performing this rewrite), so
+//    no amount of analysis reaches this case; the node has to change.
+//
+// Defining undefined bits refines every existing user, which is what makes it
+// valid for a load with several users, unlike tryToFoldExtOfExtload, which
+// widens the load itself and so needs a single use, and unlike
+// reduceLoadWidth, which bails out on multiple uses.
+static SDValue tryToFoldZExtOfTruncOfExtLoad(SelectionDAG &DAG,
+                                             DAGCombiner &Combiner,
+                                             const TargetLowering &TLI, EVT VT,
+                                             bool LegalOperations, SDNode *N,
+                                             SDValue N0) {
+  if (VT.isVector())
+    return SDValue();
+
+  auto *Load = dyn_cast<LoadSDNode>(N0.getOperand(0));
+  if (!Load || !ISD::isEXTLoad(Load) || !ISD::isUNINDEXEDLoad(Load))
+    return SDValue();
+
+  EVT MemVT = Load->getMemoryVT();
+  if (!MemVT.bitsLE(N0.getValueType()))
+    return SDValue();
+
+  // Same condition as tryToFoldExtOfExtload above: only a simple load can be
+  // retyped without first proving the result legal. A volatile or atomic load
+  // is not excluded, since the access itself is unchanged and only the
+  // in-register extension differs.
+  if ((LegalOperations || !Load->isSimple()) &&
+      !TLI.isLoadLegal(Load->getValueType(0), MemVT, Load->getAlign(),
+                       Load->getAddressSpace(), ISD::ZEXTLOAD, false))
+    return SDValue();
+
+  SDValue ZExtLoad =
+      DAG.getExtLoad(ISD::ZEXTLOAD, SDLoc(Load), Load->getValueType(0),
+                     Load->getChain(), Load->getBasePtr(), MemVT,
+                     Load->getMemOperand());
+  // Replace N before retyping the load. getExtLoad may have returned an
+  // existing zextload of the same address, in which case replacing the load
+  // CSEs the truncate, and N along with it, into pre-existing nodes and
+  // deletes N.
+  SDValue Res = DAG.getZExtOrTrunc(ZExtLoad, SDLoc(N), VT);
+  // The truncate dies with N, and its low bits are the low bits of Res.
+  DAG.transferDbgValues(N0, Res);
+  Combiner.CombineTo(N, Res);
+  Combiner.CombineTo(Load, ZExtLoad, ZExtLoad.getValue(1));
+  return SDValue(N, 0); // Return N so it doesn't get rechecked!
+}
+
 // fold ([s|z]ext (load x)) -> ([s|z]ext (truncate ([s|z]extload x)))
 // Only generate vector extloads when 1) they're legal, and 2) they are
 // deemed desirable by the target. NonNegZExt can be set to true if a zero
@@ -16375,6 +16434,10 @@ SDValue DAGCombiner::visitZERO_EXTEND(SDNode *N) {
       return SDValue(N, 0); // Return N so it doesn't get rechecked!
     }
 
+    if (SDValue ZExt = tryToFoldZExtOfTruncOfExtLoad(DAG, *this, TLI, VT,
+                                                  LegalOperations, N, N0))
+      return ZExt;
+
     EVT SrcVT = N0.getOperand(0).getValueType();
     EVT MinVT = N0.getValueType();
 
diff --git a/llvm/test/CodeGen/AArch64/zext-trunc-extload-dbg.ll b/llvm/test/CodeGen/AArch64/zext-trunc-extload-dbg.ll
new file mode 100644
index 0000000000000..5ea980e2cee29
--- /dev/null
+++ b/llvm/test/CodeGen/AArch64/zext-trunc-extload-dbg.ll
@@ -0,0 +1,42 @@
+; RUN: llc -mtriple=aarch64 -stop-after=finalize-isel < %s | FileCheck %s
+
+; Refining the extload to a zextload leaves the truncate dead, and the truncate
+; is where the location of %a had been parked. Its debug values have to move to
+; the replacement, or "a" loses its location.
+
+; The metadata block precedes the function body in MIR, so bind the variables
+; first and then require both to have a register location.
+; CHECK-DAG: ![[A:[0-9]+]] = !DILocalVariable(name: "a"
+; CHECK-DAG: ![[B:[0-9]+]] = !DILocalVariable(name: "b"
+; CHECK-LABEL: name: f
+; CHECK-DAG: DBG_VALUE %[[R:[0-9]+]], $noreg, ![[A]], !DIExpression()
+; CHECK-DAG: DBG_VALUE %[[R]], $noreg, ![[B]], !DIExpression()
+
+declare void @sink(i8, i64)
+
+define void @f(ptr %p) nounwind !dbg !5 {
+entry:
+  %a = load i8, ptr %p, align 1, !dbg !9
+    #dbg_value(i8 %a, !10, !DIExpression(), !9)
+  %b = zext i8 %a to i64, !dbg !9
+    #dbg_value(i64 %b, !12, !DIExpression(), !9)
+  call void @sink(i8 %a, i64 %b), !dbg !9
+  ret void, !dbg !9
+}
+
+!llvm.dbg.cu = !{!0}
+!llvm.module.flags = !{!3, !4}
+
+!0 = distinct !DICompileUnit(language: DW_LANG_C99, file: !1, producer: "clang", isOptimized: true, runtimeVersion: 0, emissionKind: FullDebug)
+!1 = !DIFile(filename: "t.c", directory: "/")
+!3 = !{i32 7, !"Dwarf Version", i32 5}
+!4 = !{i32 2, !"Debug Info Version", i32 3}
+!5 = distinct !DISubprogram(name: "f", scope: !1, file: !1, line: 1, type: !6, scopeLine: 1, spFlags: DISPFlagDefinition, unit: !0, retainedNodes: !8)
+!6 = !DISubroutineType(types: !7)
+!7 = !{null}
+!8 = !{!10, !12}
+!9 = !DILocation(line: 2, column: 1, scope: !5)
+!10 = !DILocalVariable(name: "a", scope: !5, file: !1, line: 2, type: !11)
+!11 = !DIBasicType(name: "char", size: 8, encoding: DW_ATE_unsigned_char)
+!12 = !DILocalVariable(name: "b", scope: !5, file: !1, line: 3, type: !13)
+!13 = !DIBasicType(name: "long", size: 64, encoding: DW_ATE_unsigned)
diff --git a/llvm/test/CodeGen/AArch64/zext-trunc-extload.ll b/llvm/test/CodeGen/AArch64/zext-trunc-extload.ll
new file mode 100644
index 0000000000000..df0a3bc9b832c
--- /dev/null
+++ b/llvm/test/CodeGen/AArch64/zext-trunc-extload.ll
@@ -0,0 +1,117 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 6
+; RUN: llc -mtriple=aarch64 < %s | FileCheck %s
+
+; The i8 argument makes the narrow load an any-extending i32 load early, which
+; leaves (zext (trunc (extload))) behind for %b. reduceLoadWidth cannot narrow
+; the load because it has several users, so the extension bits must instead be
+; defined in place by refining the extload to a zextload. Otherwise the zext
+; survives as an explicit mask.
+;
+; Each function below is checked to discriminate: reverting the fold changes
+; its output. Note that the guard on the truncate width is not covered here,
+; because on AArch64 a load narrow enough to reach it is already a zextload and
+; the fold is never entered; the AMDGPU load-constant-i1 and load-global-i1
+; tests cover it instead.
+
+declare void @sink(i8, i64, i64, i64, i1)
+declare void @sink16(i16, i64)
+declare void @sink8(i8)
+
+; The motivating shape: without the fold this ends in "and x1, x0, #0xff".
+define void @zext_of_trunc_of_multiuse_extload(ptr %p) nounwind {
+; CHECK-LABEL: zext_of_trunc_of_multiuse_extload:
+; CHECK:       // %bb.0: // %entry
+; CHECK-NEXT:    str x30, [sp, #-16]! // 8-byte Folded Spill
+; CHECK-NEXT:    ldrb w0, [x0]
+; CHECK-NEXT:    ubfx x3, x0, #0, #8
+; CHECK-NEXT:    lsl x2, x0, #56
+; CHECK-NEXT:    mov x1, x0
+; CHECK-NEXT:    cmp x3, #0
+; CHECK-NEXT:    cset w4, eq
+; CHECK-NEXT:    bl sink
+; CHECK-NEXT:    ldr x30, [sp], #16 // 8-byte Folded Reload
+; CHECK-NEXT:    ret
+entry:
+  %a = load i8, ptr %p, align 1
+  %b = zext i8 %a to i64
+  %c = shl i64 %b, 56
+  %d = lshr i64 %c, 56
+  %e = icmp eq i64 %d, 0
+  call void @sink(i8 %a, i64 %b, i64 %c, i64 %d, i1 %e)
+  ret void
+}
+
+; The same with an i16 memory type, so the fold is not tied to one width.
+define void @zext_of_trunc_of_multiuse_extload_i16(ptr %p) nounwind {
+; CHECK-LABEL: zext_of_trunc_of_multiuse_extload_i16:
+; CHECK:       // %bb.0: // %entry
+; CHECK-NEXT:    str x30, [sp, #-16]! // 8-byte Folded Spill
+; CHECK-NEXT:    ldrh w0, [x0]
+; CHECK-NEXT:    mov x1, x0
+; CHECK-NEXT:    bl sink16
+; CHECK-NEXT:    ldr x30, [sp], #16 // 8-byte Folded Reload
+; CHECK-NEXT:    ret
+entry:
+  %a = load i16, ptr %p, align 2
+  %b = zext i16 %a to i64
+  call void @sink16(i16 %a, i64 %b)
+  ret void
+}
+
+; A volatile load is refined too: the access itself is unchanged and only the
+; in-register extension differs. This matches the in-place retype in the
+; (and (load), mask) fold, which also does not exclude volatile loads.
+define void @volatile_load_is_refined(ptr %p) nounwind {
+; CHECK-LABEL: volatile_load_is_refined:
+; CHECK:       // %bb.0: // %entry
+; CHECK-NEXT:    stp x30, x19, [sp, #-16]! // 16-byte Folded Spill
+; CHECK-NEXT:    ldrb w19, [x0]
+; CHECK-NEXT:    mov w0, wzr
+; CHECK-NEXT:    mov x1, x19
+; CHECK-NEXT:    bl sink16
+; CHECK-NEXT:    mov w0, w19
+; CHECK-NEXT:    bl sink8
+; CHECK-NEXT:    ldp x30, x19, [sp], #16 // 16-byte Folded Reload
+; CHECK-NEXT:    ret
+entry:
+  %a = load volatile i8, ptr %p, align 1
+  %b = zext i8 %a to i64
+  call void @sink16(i16 0, i64 %b)
+  call void @sink8(i8 %a)
+  ret void
+}
+
+; Regression test for the replacement order: getExtLoad CSEs into the zextload
+; already built for %b once the load for %a is rechained past the non-aliasing
+; store. Replacing the load before N then folds the truncate, and N with it,
+; into pre-existing nodes and deletes N, which used to leave the combiner
+; holding a deleted node and assert.
+ at g1 = global i8 0
+ at g2 = global i8 0
+declare void @sink_cse(i8, i64, i64, i8 zeroext)
+
+define void @zext_of_trunc_cse_with_existing_zextload() nounwind {
+; CHECK-LABEL: zext_of_trunc_cse_with_existing_zextload:
+; CHECK:       // %bb.0: // %entry
+; CHECK-NEXT:    str x30, [sp, #-16]! // 8-byte Folded Spill
+; CHECK-NEXT:    adrp x8, :got:g1
+; CHECK-NEXT:    adrp x9, :got:g2
+; CHECK-NEXT:    ldr x8, [x8, :got_lo12:g1]
+; CHECK-NEXT:    ldr x9, [x9, :got_lo12:g2]
+; CHECK-NEXT:    ldrb w0, [x8]
+; CHECK-NEXT:    strb wzr, [x9]
+; CHECK-NEXT:    mov x1, x0
+; CHECK-NEXT:    mov x2, x0
+; CHECK-NEXT:    mov w3, w0
+; CHECK-NEXT:    bl sink_cse
+; CHECK-NEXT:    ldr x30, [sp], #16 // 8-byte Folded Reload
+; CHECK-NEXT:    ret
+entry:
+  %b = load i8, ptr @g1
+  %zb = zext i8 %b to i64
+  store i8 0, ptr @g2
+  %a = load i8, ptr @g1
+  %za = zext i8 %a to i64
+  call void @sink_cse(i8 %a, i64 %za, i64 %zb, i8 zeroext %b)
+  ret void
+}



More information about the llvm-commits mailing list