[llvm] [IndVarSimplify] Widen select-controlled recurrences (PR #206024)
Madhur Amilkanthwar via llvm-commits
llvm-commits at lists.llvm.org
Fri Jun 26 03:21:12 PDT 2026
https://github.com/madhur13490 created https://github.com/llvm/llvm-project/pull/206024
Recognize a narrow header phi whose only wide-cast use is a sign-extend to a strictly wider legal integer iM and whose latch incoming is a select whose new-value arm is a one-use `trunc nsw` of an iM value (arms in either order). Rewrite it into an iM-typed recurrence, sinking a truncation onto the narrow uses.
%wide.next is any loop-defined iM value (typically the next-value of a wider companion IV); %c is the i1 select condition.
Before:
```
loop:
%iv = phi iN [ %init, ph ], [ %sel, latch ]
%ext = sext iN %iv to iM ; only wide-cast use
%narrow = trunc nsw iM %wide.next to iN ; one-use
%sel = select i1 %c, iN %iv, iN %narrow ; %iv's other use
```
After:
```
preheader:
%init.wide = sext iN %init to iM ; one-time widen
loop:
%iv.wide = phi iM [ %init.wide, ph ], [ %sel.wide, latch ]
%sel.wide = select i1 %c, iM %iv.wide, iM %wide.next
%sel.trunc = trunc nsw iM %sel.wide to iN ; for narrow uses
```
What changes:
- %ext's uses are rewired to %iv.wide, eliminating the per-iteration sign-extend on the wide use.
- The trunc feeding the select is gone; %sel.trunc replaces it on the output side so narrow consumers still see an iN value. Net trunc count per iteration is unchanged.
- %init is sign-extended once in the preheader, never inside the loop.
This patch is a good-to-have patch for the GVN patch #162259 as it simplifies the IR fed to GVN. If this is commited, the matching part in GVN code can be simplified a bit (not by a huge amount). This also goes back to what Nikita suggested in https://github.com/llvm/llvm-project/pull/144987#issuecomment-3077906788
Thanks to @rj-jesus for the baseline version of this patch.
Assisted by Cursor.
>From c8138d08633eeefa68242b836cdc1279d889e42b Mon Sep 17 00:00:00 2001
From: Madhur Amilkanthwar <madhura at nvidia.com>
Date: Fri, 26 Jun 2026 01:03:36 -0700
Subject: [PATCH] [IndVarSimplify] Widen select-controlled recurrences
Recognize a narrow header phi whose only wide-cast use is a sign-extend
to a strictly wider legal integer iM and whose latch incoming is a
select whose new-value arm is a one-use `trunc nsw` of an iM value
(arms in either order). Rewrite it into an iM-typed recurrence, sinking
a truncation onto the narrow uses.
%wide.next is any loop-defined iM value (typically the next-value of a
wider companion IV); %c is the i1 select condition.
Before:
loop:
%iv = phi iN [ %init, ph ], [ %sel, latch ]
%ext = sext iN %iv to iM ; only wide-cast use
%narrow = trunc nsw iM %wide.next to iN ; one-use
%sel = select i1 %c, iN %iv, iN %narrow ; %iv's other use
After:
preheader:
%init.wide = sext iN %init to iM ; one-time widen
loop:
%iv.wide = phi iM [ %init.wide, ph ], [ %sel.wide, latch ]
%sel.wide = select i1 %c, iM %iv.wide, iM %wide.next
%sel.trunc = trunc nsw iM %sel.wide to iN ; for narrow uses
What changes:
- %ext's uses are rewired to %iv.wide, eliminating the per-iteration
sign-extend on the wide use.
- The trunc feeding the select is gone; %sel.trunc replaces it on the
output side so narrow consumers still see an iN value. Net trunc
count per iteration is unchanged.
- %init is sign-extended once in the preheader, never inside the loop.
This patch is a good-to-have patch for the GVN patch #162259
as it simplifies the IR fed to GVN. If this is commited, the matching
part in GVN code can be simplified a bit (not by a huge amount).
---
llvm/lib/Transforms/Scalar/IndVarSimplify.cpp | 116 +++++++++++++++
.../IndVarSimplify/widen-select-recurrence.ll | 135 ++++++++++++++++++
2 files changed, 251 insertions(+)
create mode 100644 llvm/test/Transforms/IndVarSimplify/widen-select-recurrence.ll
diff --git a/llvm/lib/Transforms/Scalar/IndVarSimplify.cpp b/llvm/lib/Transforms/Scalar/IndVarSimplify.cpp
index c92efadded635..0201a6ce680ae 100644
--- a/llvm/lib/Transforms/Scalar/IndVarSimplify.cpp
+++ b/llvm/lib/Transforms/Scalar/IndVarSimplify.cpp
@@ -663,6 +663,115 @@ static void visitIVCast(CastInst *Cast, WideIVInfo &WI,
WI.IsSigned |= IsSigned;
}
+/// Widen a select-controlled recurrence so it runs in the wider type used by
+/// its sole wide use, sinking a truncation onto the narrow uses.
+///
+/// %wide.next is any loop-defined iM value (typically the next-value of a
+/// wider companion IV); %c is the i1 select condition. The matcher only
+/// requires the trunc to be one-use so dropping it is safe.
+///
+/// Before:
+/// loop:
+/// %iv = phi iN [ %init, ph ], [ %sel, latch ]
+/// %ext = sext iN %iv to iM ; %iv's only wide-cast use
+/// %narrow = trunc nsw iM %wide.next to iN ; one-use
+/// %sel = select i1 %c, iN %iv, iN %narrow ; %iv's other use;
+/// ; arms in any order
+///
+/// After:
+/// preheader:
+/// %init.wide = sext iN %init to iM ; one-time widen of init
+/// loop:
+/// %iv.wide = phi iM [ %init.wide, ph ], [ %sel.wide, latch ]
+/// %sel.wide = select i1 %c, iM %iv.wide, iM %wide.next ; now in iM
+/// %sel.trunc = trunc nsw iM %sel.wide to iN ; for narrow consumers
+///
+/// What changed and why:
+/// - The recurrence runs in iM instead of iN. %ext's uses are rewired to
+/// %iv.wide, eliminating the per-iteration sign-extend on the wide use.
+/// - %narrow (trunc feeding the select) is gone; the select consumes
+/// %wide.next directly. %sel.trunc replaces it on the output side, so
+/// any narrow consumer of the recurrence still gets an iN value. Net
+/// trunc count per iteration is unchanged, but the recurrence itself
+/// is now wide.
+/// - %init is sign-extended once in the preheader, never inside the loop.
+///
+/// SCEV does not classify select-controlled phis as induction variables, so
+/// simplifyUsersOfIV's IV walker never visits them and createWideIV never
+/// runs. The structural preconditions are therefore checked here directly.
+static PHINode *
+widenSelectRecurrence(PHINode *PN, LoopInfo *LI,
+ SmallVectorImpl<WeakTrackingVH> &DeadInsts,
+ unsigned &NumWidened) {
+ Type *SrcTy = PN->getType();
+ if (!SrcTy->isIntegerTy() || PN->getNumIncomingValues() != 2 ||
+ !PN->hasNUses(2))
+ return nullptr;
+
+ Loop *L = LI->getLoopFor(PN->getParent());
+ if (!L || PN->getParent() != L->getHeader())
+ return nullptr;
+ BasicBlock *PH = L->getLoopPreheader();
+ BasicBlock *Latch = L->getLoopLatch();
+ if (!PH || !Latch)
+ return nullptr;
+
+ auto *Sel = dyn_cast<SelectInst>(PN->getIncomingValueForBlock(Latch));
+ if (!Sel)
+ return nullptr;
+
+ Value *V;
+ if (!match(Sel, m_c_Select(m_Specific(PN), m_OneUse(m_NSWTrunc(m_Value(V))))))
+ return nullptr;
+
+ // Widen to the truncated arm's source type, which must be a strictly wider
+ // legal integer.
+ Type *DestTy = V->getType();
+ if (!DestTy->isIntegerTy() ||
+ DestTy->getIntegerBitWidth() <= SrcTy->getIntegerBitWidth() ||
+ !PN->getDataLayout().isLegalInteger(DestTy->getIntegerBitWidth()))
+ return nullptr;
+
+ // The non-select use must be a sign-extend to DestTy.
+ SExtInst *Ext = nullptr;
+ for (User *U : PN->users()) {
+ if (U == Sel)
+ continue;
+ Ext = dyn_cast<SExtInst>(U);
+ if (!Ext || Ext->getType() != DestTy)
+ return nullptr;
+ }
+ if (!Ext)
+ return nullptr;
+
+ auto *WidePN =
+ PHINode::Create(DestTy, 2, PN->getName() + ".wide", PN->getIterator());
+ auto *WideInit =
+ CastInst::CreateSExtOrBitCast(PN->getIncomingValueForBlock(PH), DestTy,
+ "", PH->getTerminator()->getIterator());
+ WidePN->addIncoming(WideInit, PH);
+
+ Value *TVal = WidePN, *FVal = V;
+ if (Sel->getTrueValue() != PN)
+ std::swap(TVal, FVal);
+ auto *WideSel =
+ SelectInst::Create(Sel->getCondition(), TVal, FVal,
+ Sel->getName() + ".wide", Sel->getIterator());
+ WidePN->addIncoming(WideSel, Latch);
+
+ auto *Trunc =
+ CastInst::CreateTruncOrBitCast(WideSel, SrcTy, "", Sel->getIterator());
+ Trunc->setHasNoSignedWrap(true);
+
+ Ext->replaceAllUsesWith(WidePN);
+ Sel->replaceAllUsesWith(Trunc);
+ DeadInsts.emplace_back(Ext);
+ DeadInsts.emplace_back(Sel);
+ DeadInsts.emplace_back(PN);
+ ++NumWidened;
+ return WidePN;
+}
+
//===----------------------------------------------------------------------===//
// Live IV Reduction - Minimize IVs live across the loop.
//===----------------------------------------------------------------------===//
@@ -754,6 +863,13 @@ bool IndVarSimplify::simplifyAndExtend(Loop *L,
NumWidened += Widened;
Changed = true;
LoopPhis.push_back(WidePhi);
+ } else if (PHINode *WidePhi = widenSelectRecurrence(
+ WideIVs.back().NarrowIV, LI, DeadInsts, Widened)) {
+ // createWideIV did not fire (the phi is not an AddRec). Try the
+ // structural select-recurrence widener.
+ NumWidened += Widened;
+ Changed = true;
+ LoopPhis.push_back(WidePhi);
}
}
}
diff --git a/llvm/test/Transforms/IndVarSimplify/widen-select-recurrence.ll b/llvm/test/Transforms/IndVarSimplify/widen-select-recurrence.ll
new file mode 100644
index 0000000000000..ca4d9d29ab4de
--- /dev/null
+++ b/llvm/test/Transforms/IndVarSimplify/widen-select-recurrence.ll
@@ -0,0 +1,135 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
+; RUN: opt < %s -S -passes='indvars' -verify-loop-info -verify-dom-info -verify-scev | FileCheck %s
+
+target datalayout = "e-m:e-i8:8:32-i16:16:32-i64:64-i128:128-n32:64-S128"
+
+; A narrow header phi with exactly two uses (one sext to a strictly wider legal
+; integer; one latch select whose new-value arm is `trunc nsw` of a wider value)
+; should be widened into the destination type, sinking a truncation onto the
+; narrow uses. SCEV does not classify select-controlled recurrences as
+; induction variables, so the widening relies on a structural match rather
+; than the SCEV-driven IV walker.
+define i32 @widen_select_recurrence_positive(ptr %arr, i32 %n) {
+; CHECK-LABEL: @widen_select_recurrence_positive(
+; CHECK-NEXT: entry:
+; CHECK-NEXT: [[TMP0:%.*]] = sext i32 0 to i64
+; CHECK-NEXT: br label [[LOOP:%.*]]
+; CHECK: loop:
+; CHECK-NEXT: [[IV:%.*]] = phi i64 [ 0, [[ENTRY:%.*]] ], [ [[IV_NEXT:%.*]], [[LOOP]] ]
+; CHECK-NEXT: [[MIN_IDX_WIDE:%.*]] = phi i64 [ [[TMP0]], [[ENTRY]] ], [ [[MIN_IDX_NEXT_WIDE:%.*]], [[LOOP]] ]
+; CHECK-NEXT: [[GEP:%.*]] = getelementptr float, ptr [[ARR:%.*]], i64 [[MIN_IDX_WIDE]]
+; CHECK-NEXT: [[VAL:%.*]] = load float, ptr [[GEP]], align 4
+; CHECK-NEXT: [[CMP:%.*]] = fcmp olt float [[VAL]], 0.000000e+00
+; CHECK-NEXT: [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; CHECK-NEXT: [[MIN_IDX_NEXT_WIDE]] = select i1 [[CMP]], i64 [[IV_NEXT]], i64 [[MIN_IDX_WIDE]]
+; CHECK-NEXT: [[TMP1:%.*]] = trunc nsw i64 [[MIN_IDX_NEXT_WIDE]] to i32
+; CHECK-NEXT: [[EXITCOND:%.*]] = icmp ne i64 [[IV_NEXT]], 100
+; CHECK-NEXT: br i1 [[EXITCOND]], label [[LOOP]], label [[EXIT:%.*]]
+; CHECK: exit:
+; CHECK-NEXT: [[MIN_IDX_NEXT_LCSSA:%.*]] = phi i32 [ [[TMP1]], [[LOOP]] ]
+; CHECK-NEXT: ret i32 [[MIN_IDX_NEXT_LCSSA]]
+;
+entry:
+ br label %loop
+
+loop:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %loop ]
+ %min.idx = phi i32 [ 0, %entry ], [ %min.idx.next, %loop ]
+ %min.idx.ext = sext i32 %min.idx to i64
+ %gep = getelementptr float, ptr %arr, i64 %min.idx.ext
+ %val = load float, ptr %gep, align 4
+ %cmp = fcmp olt float %val, 0.0
+ %iv.next = add nsw i64 %iv, 1
+ %new.idx = trunc nsw i64 %iv.next to i32
+ %min.idx.next = select i1 %cmp, i32 %new.idx, i32 %min.idx
+ %trip = icmp slt i64 %iv.next, 100
+ br i1 %trip, label %loop, label %exit
+
+exit:
+ ret i32 %min.idx.next
+}
+
+; Negative: latch select's new-value arm has multiple uses, so m_OneUse on the
+; trunc fails; the recurrence is not widened.
+define i32 @widen_select_recurrence_multi_use_trunc(ptr %arr) {
+; CHECK-LABEL: @widen_select_recurrence_multi_use_trunc(
+; CHECK-NEXT: entry:
+; CHECK-NEXT: br label [[LOOP:%.*]]
+; CHECK: loop:
+; CHECK-NEXT: [[IV:%.*]] = phi i64 [ 0, [[ENTRY:%.*]] ], [ [[IV_NEXT:%.*]], [[LOOP]] ]
+; CHECK-NEXT: [[MIN_IDX:%.*]] = phi i32 [ 0, [[ENTRY]] ], [ [[MIN_IDX_NEXT:%.*]], [[LOOP]] ]
+; CHECK-NEXT: [[MIN_IDX_EXT:%.*]] = sext i32 [[MIN_IDX]] to i64
+; CHECK-NEXT: [[GEP:%.*]] = getelementptr float, ptr [[ARR:%.*]], i64 [[MIN_IDX_EXT]]
+; CHECK-NEXT: [[VAL:%.*]] = load float, ptr [[GEP]], align 4
+; CHECK-NEXT: [[CMP:%.*]] = fcmp olt float [[VAL]], 0.000000e+00
+; CHECK-NEXT: [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; CHECK-NEXT: [[NEW_IDX:%.*]] = trunc nsw i64 [[IV_NEXT]] to i32
+; CHECK-NEXT: store i32 [[NEW_IDX]], ptr [[ARR]], align 4
+; CHECK-NEXT: [[MIN_IDX_NEXT]] = select i1 [[CMP]], i32 [[NEW_IDX]], i32 [[MIN_IDX]]
+; CHECK-NEXT: [[EXITCOND:%.*]] = icmp ne i64 [[IV_NEXT]], 100
+; CHECK-NEXT: br i1 [[EXITCOND]], label [[LOOP]], label [[EXIT:%.*]]
+; CHECK: exit:
+; CHECK-NEXT: [[MIN_IDX_NEXT_LCSSA:%.*]] = phi i32 [ [[MIN_IDX_NEXT]], [[LOOP]] ]
+; CHECK-NEXT: ret i32 [[MIN_IDX_NEXT_LCSSA]]
+;
+entry:
+ br label %loop
+
+loop:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %loop ]
+ %min.idx = phi i32 [ 0, %entry ], [ %min.idx.next, %loop ]
+ %min.idx.ext = sext i32 %min.idx to i64
+ %gep = getelementptr float, ptr %arr, i64 %min.idx.ext
+ %val = load float, ptr %gep, align 4
+ %cmp = fcmp olt float %val, 0.0
+ %iv.next = add nsw i64 %iv, 1
+ %new.idx = trunc nsw i64 %iv.next to i32
+ store i32 %new.idx, ptr %arr, align 4
+ %min.idx.next = select i1 %cmp, i32 %new.idx, i32 %min.idx
+ %trip = icmp slt i64 %iv.next, 100
+ br i1 %trip, label %loop, label %exit
+
+exit:
+ ret i32 %min.idx.next
+}
+
+; Negative: the wide cast user is a zext, not a sext; the widener requires sext.
+define i32 @widen_select_recurrence_zext_user(ptr %arr) {
+; CHECK-LABEL: @widen_select_recurrence_zext_user(
+; CHECK-NEXT: entry:
+; CHECK-NEXT: br label [[LOOP:%.*]]
+; CHECK: loop:
+; CHECK-NEXT: [[IV:%.*]] = phi i64 [ 0, [[ENTRY:%.*]] ], [ [[IV_NEXT:%.*]], [[LOOP]] ]
+; CHECK-NEXT: [[MIN_IDX:%.*]] = phi i32 [ 0, [[ENTRY]] ], [ [[MIN_IDX_NEXT:%.*]], [[LOOP]] ]
+; CHECK-NEXT: [[MIN_IDX_EXT:%.*]] = zext i32 [[MIN_IDX]] to i64
+; CHECK-NEXT: [[GEP:%.*]] = getelementptr float, ptr [[ARR:%.*]], i64 [[MIN_IDX_EXT]]
+; CHECK-NEXT: [[VAL:%.*]] = load float, ptr [[GEP]], align 4
+; CHECK-NEXT: [[CMP:%.*]] = fcmp olt float [[VAL]], 0.000000e+00
+; CHECK-NEXT: [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; CHECK-NEXT: [[NEW_IDX:%.*]] = trunc nsw i64 [[IV_NEXT]] to i32
+; CHECK-NEXT: [[MIN_IDX_NEXT]] = select i1 [[CMP]], i32 [[NEW_IDX]], i32 [[MIN_IDX]]
+; CHECK-NEXT: [[EXITCOND:%.*]] = icmp ne i64 [[IV_NEXT]], 100
+; CHECK-NEXT: br i1 [[EXITCOND]], label [[LOOP]], label [[EXIT:%.*]]
+; CHECK: exit:
+; CHECK-NEXT: [[MIN_IDX_NEXT_LCSSA:%.*]] = phi i32 [ [[MIN_IDX_NEXT]], [[LOOP]] ]
+; CHECK-NEXT: ret i32 [[MIN_IDX_NEXT_LCSSA]]
+;
+entry:
+ br label %loop
+
+loop:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %loop ]
+ %min.idx = phi i32 [ 0, %entry ], [ %min.idx.next, %loop ]
+ %min.idx.ext = zext i32 %min.idx to i64
+ %gep = getelementptr float, ptr %arr, i64 %min.idx.ext
+ %val = load float, ptr %gep, align 4
+ %cmp = fcmp olt float %val, 0.0
+ %iv.next = add nsw i64 %iv, 1
+ %new.idx = trunc nsw i64 %iv.next to i32
+ %min.idx.next = select i1 %cmp, i32 %new.idx, i32 %min.idx
+ %trip = icmp slt i64 %iv.next, 100
+ br i1 %trip, label %loop, label %exit
+
+exit:
+ ret i32 %min.idx.next
+}
More information about the llvm-commits
mailing list