[llvm] [IndVarSimplify] Widen select-controlled recurrences (PR #206024)

Madhur Amilkanthwar via llvm-commits llvm-commits at lists.llvm.org
Fri Jun 26 03:21:12 PDT 2026


https://github.com/madhur13490 created https://github.com/llvm/llvm-project/pull/206024

Recognize a narrow header phi whose only wide-cast use is a sign-extend to a strictly wider legal integer iM and whose latch incoming is a select whose new-value arm is a one-use `trunc nsw` of an iM value (arms in either order). Rewrite it into an iM-typed recurrence, sinking a truncation onto the narrow uses.

%wide.next is any loop-defined iM value (typically the next-value of a wider companion IV); %c is the i1 select condition.

Before:
```
  loop:
    %iv     = phi iN [ %init, ph ], [ %sel, latch ]
    %ext    = sext iN %iv to iM                  ; only wide-cast use
    %narrow = trunc nsw iM %wide.next to iN      ; one-use
    %sel    = select i1 %c, iN %iv, iN %narrow   ; %iv's other use
```

After:
 ```
 preheader:
    %init.wide = sext iN %init to iM             ; one-time widen
  loop:
    %iv.wide   = phi iM [ %init.wide, ph ], [ %sel.wide, latch ]
    %sel.wide  = select i1 %c, iM %iv.wide, iM %wide.next
    %sel.trunc = trunc nsw iM %sel.wide to iN    ; for narrow uses

```
What changes:
- %ext's uses are rewired to %iv.wide, eliminating the per-iteration sign-extend on the wide use.
- The trunc feeding the select is gone; %sel.trunc replaces it on the output side so narrow consumers still see an iN value. Net trunc count per iteration is unchanged.
- %init is sign-extended once in the preheader, never inside the loop.

This patch is a good-to-have patch for the GVN patch #162259 as it simplifies the IR fed to GVN. If this is commited, the matching part in GVN code can be simplified a bit (not by a huge amount). This also goes back to what Nikita suggested in https://github.com/llvm/llvm-project/pull/144987#issuecomment-3077906788

Thanks to @rj-jesus for the baseline version of this patch. 

Assisted by Cursor.


>From c8138d08633eeefa68242b836cdc1279d889e42b Mon Sep 17 00:00:00 2001
From: Madhur Amilkanthwar <madhura at nvidia.com>
Date: Fri, 26 Jun 2026 01:03:36 -0700
Subject: [PATCH] [IndVarSimplify] Widen select-controlled recurrences

Recognize a narrow header phi whose only wide-cast use is a sign-extend
to a strictly wider legal integer iM and whose latch incoming is a
select whose new-value arm is a one-use `trunc nsw` of an iM value
(arms in either order). Rewrite it into an iM-typed recurrence, sinking
a truncation onto the narrow uses.

%wide.next is any loop-defined iM value (typically the next-value of a
wider companion IV); %c is the i1 select condition.

Before:
  loop:
    %iv     = phi iN [ %init, ph ], [ %sel, latch ]
    %ext    = sext iN %iv to iM                  ; only wide-cast use
    %narrow = trunc nsw iM %wide.next to iN      ; one-use
    %sel    = select i1 %c, iN %iv, iN %narrow   ; %iv's other use

After:
  preheader:
    %init.wide = sext iN %init to iM             ; one-time widen
  loop:
    %iv.wide   = phi iM [ %init.wide, ph ], [ %sel.wide, latch ]
    %sel.wide  = select i1 %c, iM %iv.wide, iM %wide.next
    %sel.trunc = trunc nsw iM %sel.wide to iN    ; for narrow uses

What changes:
- %ext's uses are rewired to %iv.wide, eliminating the per-iteration
  sign-extend on the wide use.
- The trunc feeding the select is gone; %sel.trunc replaces it on the
  output side so narrow consumers still see an iN value. Net trunc
  count per iteration is unchanged.
- %init is sign-extended once in the preheader, never inside the loop.

This patch is a good-to-have patch for the GVN patch #162259
as it simplifies the IR fed to GVN. If this is commited, the matching
part in GVN code can be simplified a bit (not by a huge amount).
---
 llvm/lib/Transforms/Scalar/IndVarSimplify.cpp | 116 +++++++++++++++
 .../IndVarSimplify/widen-select-recurrence.ll | 135 ++++++++++++++++++
 2 files changed, 251 insertions(+)
 create mode 100644 llvm/test/Transforms/IndVarSimplify/widen-select-recurrence.ll

diff --git a/llvm/lib/Transforms/Scalar/IndVarSimplify.cpp b/llvm/lib/Transforms/Scalar/IndVarSimplify.cpp
index c92efadded635..0201a6ce680ae 100644
--- a/llvm/lib/Transforms/Scalar/IndVarSimplify.cpp
+++ b/llvm/lib/Transforms/Scalar/IndVarSimplify.cpp
@@ -663,6 +663,115 @@ static void visitIVCast(CastInst *Cast, WideIVInfo &WI,
   WI.IsSigned |= IsSigned;
 }
 
+/// Widen a select-controlled recurrence so it runs in the wider type used by
+/// its sole wide use, sinking a truncation onto the narrow uses.
+///
+/// %wide.next is any loop-defined iM value (typically the next-value of a
+/// wider companion IV); %c is the i1 select condition. The matcher only
+/// requires the trunc to be one-use so dropping it is safe.
+///
+/// Before:
+///   loop:
+///     %iv     = phi iN [ %init, ph ], [ %sel, latch ]
+///     %ext    = sext iN %iv to iM                  ; %iv's only wide-cast use
+///     %narrow = trunc nsw iM %wide.next to iN      ; one-use
+///     %sel    = select i1 %c, iN %iv, iN %narrow   ; %iv's other use;
+///                                                  ;   arms in any order
+///
+/// After:
+///   preheader:
+///     %init.wide = sext iN %init to iM             ; one-time widen of init
+///   loop:
+///     %iv.wide   = phi iM [ %init.wide, ph ], [ %sel.wide, latch ]
+///     %sel.wide  = select i1 %c, iM %iv.wide, iM %wide.next  ; now in iM
+///     %sel.trunc = trunc nsw iM %sel.wide to iN    ; for narrow consumers
+///
+/// What changed and why:
+///   - The recurrence runs in iM instead of iN. %ext's uses are rewired to
+///     %iv.wide, eliminating the per-iteration sign-extend on the wide use.
+///   - %narrow (trunc feeding the select) is gone; the select consumes
+///     %wide.next directly. %sel.trunc replaces it on the output side, so
+///     any narrow consumer of the recurrence still gets an iN value. Net
+///     trunc count per iteration is unchanged, but the recurrence itself
+///     is now wide.
+///   - %init is sign-extended once in the preheader, never inside the loop.
+///
+/// SCEV does not classify select-controlled phis as induction variables, so
+/// simplifyUsersOfIV's IV walker never visits them and createWideIV never
+/// runs. The structural preconditions are therefore checked here directly.
+static PHINode *
+widenSelectRecurrence(PHINode *PN, LoopInfo *LI,
+                      SmallVectorImpl<WeakTrackingVH> &DeadInsts,
+                      unsigned &NumWidened) {
+  Type *SrcTy = PN->getType();
+  if (!SrcTy->isIntegerTy() || PN->getNumIncomingValues() != 2 ||
+      !PN->hasNUses(2))
+    return nullptr;
+
+  Loop *L = LI->getLoopFor(PN->getParent());
+  if (!L || PN->getParent() != L->getHeader())
+    return nullptr;
+  BasicBlock *PH = L->getLoopPreheader();
+  BasicBlock *Latch = L->getLoopLatch();
+  if (!PH || !Latch)
+    return nullptr;
+
+  auto *Sel = dyn_cast<SelectInst>(PN->getIncomingValueForBlock(Latch));
+  if (!Sel)
+    return nullptr;
+
+  Value *V;
+  if (!match(Sel, m_c_Select(m_Specific(PN), m_OneUse(m_NSWTrunc(m_Value(V))))))
+    return nullptr;
+
+  // Widen to the truncated arm's source type, which must be a strictly wider
+  // legal integer.
+  Type *DestTy = V->getType();
+  if (!DestTy->isIntegerTy() ||
+      DestTy->getIntegerBitWidth() <= SrcTy->getIntegerBitWidth() ||
+      !PN->getDataLayout().isLegalInteger(DestTy->getIntegerBitWidth()))
+    return nullptr;
+
+  // The non-select use must be a sign-extend to DestTy.
+  SExtInst *Ext = nullptr;
+  for (User *U : PN->users()) {
+    if (U == Sel)
+      continue;
+    Ext = dyn_cast<SExtInst>(U);
+    if (!Ext || Ext->getType() != DestTy)
+      return nullptr;
+  }
+  if (!Ext)
+    return nullptr;
+
+  auto *WidePN =
+      PHINode::Create(DestTy, 2, PN->getName() + ".wide", PN->getIterator());
+  auto *WideInit =
+      CastInst::CreateSExtOrBitCast(PN->getIncomingValueForBlock(PH), DestTy,
+                                    "", PH->getTerminator()->getIterator());
+  WidePN->addIncoming(WideInit, PH);
+
+  Value *TVal = WidePN, *FVal = V;
+  if (Sel->getTrueValue() != PN)
+    std::swap(TVal, FVal);
+  auto *WideSel =
+      SelectInst::Create(Sel->getCondition(), TVal, FVal,
+                         Sel->getName() + ".wide", Sel->getIterator());
+  WidePN->addIncoming(WideSel, Latch);
+
+  auto *Trunc =
+      CastInst::CreateTruncOrBitCast(WideSel, SrcTy, "", Sel->getIterator());
+  Trunc->setHasNoSignedWrap(true);
+
+  Ext->replaceAllUsesWith(WidePN);
+  Sel->replaceAllUsesWith(Trunc);
+  DeadInsts.emplace_back(Ext);
+  DeadInsts.emplace_back(Sel);
+  DeadInsts.emplace_back(PN);
+  ++NumWidened;
+  return WidePN;
+}
+
 //===----------------------------------------------------------------------===//
 //  Live IV Reduction - Minimize IVs live across the loop.
 //===----------------------------------------------------------------------===//
@@ -754,6 +863,13 @@ bool IndVarSimplify::simplifyAndExtend(Loop *L,
         NumWidened += Widened;
         Changed = true;
         LoopPhis.push_back(WidePhi);
+      } else if (PHINode *WidePhi = widenSelectRecurrence(
+                     WideIVs.back().NarrowIV, LI, DeadInsts, Widened)) {
+        // createWideIV did not fire (the phi is not an AddRec). Try the
+        // structural select-recurrence widener.
+        NumWidened += Widened;
+        Changed = true;
+        LoopPhis.push_back(WidePhi);
       }
     }
   }
diff --git a/llvm/test/Transforms/IndVarSimplify/widen-select-recurrence.ll b/llvm/test/Transforms/IndVarSimplify/widen-select-recurrence.ll
new file mode 100644
index 0000000000000..ca4d9d29ab4de
--- /dev/null
+++ b/llvm/test/Transforms/IndVarSimplify/widen-select-recurrence.ll
@@ -0,0 +1,135 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
+; RUN: opt < %s -S -passes='indvars' -verify-loop-info -verify-dom-info -verify-scev | FileCheck %s
+
+target datalayout = "e-m:e-i8:8:32-i16:16:32-i64:64-i128:128-n32:64-S128"
+
+; A narrow header phi with exactly two uses (one sext to a strictly wider legal
+; integer; one latch select whose new-value arm is `trunc nsw` of a wider value)
+; should be widened into the destination type, sinking a truncation onto the
+; narrow uses. SCEV does not classify select-controlled recurrences as
+; induction variables, so the widening relies on a structural match rather
+; than the SCEV-driven IV walker.
+define i32 @widen_select_recurrence_positive(ptr %arr, i32 %n) {
+; CHECK-LABEL: @widen_select_recurrence_positive(
+; CHECK-NEXT:  entry:
+; CHECK-NEXT:    [[TMP0:%.*]] = sext i32 0 to i64
+; CHECK-NEXT:    br label [[LOOP:%.*]]
+; CHECK:       loop:
+; CHECK-NEXT:    [[IV:%.*]] = phi i64 [ 0, [[ENTRY:%.*]] ], [ [[IV_NEXT:%.*]], [[LOOP]] ]
+; CHECK-NEXT:    [[MIN_IDX_WIDE:%.*]] = phi i64 [ [[TMP0]], [[ENTRY]] ], [ [[MIN_IDX_NEXT_WIDE:%.*]], [[LOOP]] ]
+; CHECK-NEXT:    [[GEP:%.*]] = getelementptr float, ptr [[ARR:%.*]], i64 [[MIN_IDX_WIDE]]
+; CHECK-NEXT:    [[VAL:%.*]] = load float, ptr [[GEP]], align 4
+; CHECK-NEXT:    [[CMP:%.*]] = fcmp olt float [[VAL]], 0.000000e+00
+; CHECK-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; CHECK-NEXT:    [[MIN_IDX_NEXT_WIDE]] = select i1 [[CMP]], i64 [[IV_NEXT]], i64 [[MIN_IDX_WIDE]]
+; CHECK-NEXT:    [[TMP1:%.*]] = trunc nsw i64 [[MIN_IDX_NEXT_WIDE]] to i32
+; CHECK-NEXT:    [[EXITCOND:%.*]] = icmp ne i64 [[IV_NEXT]], 100
+; CHECK-NEXT:    br i1 [[EXITCOND]], label [[LOOP]], label [[EXIT:%.*]]
+; CHECK:       exit:
+; CHECK-NEXT:    [[MIN_IDX_NEXT_LCSSA:%.*]] = phi i32 [ [[TMP1]], [[LOOP]] ]
+; CHECK-NEXT:    ret i32 [[MIN_IDX_NEXT_LCSSA]]
+;
+entry:
+  br label %loop
+
+loop:
+  %iv = phi i64 [ 0, %entry ], [ %iv.next, %loop ]
+  %min.idx = phi i32 [ 0, %entry ], [ %min.idx.next, %loop ]
+  %min.idx.ext = sext i32 %min.idx to i64
+  %gep = getelementptr float, ptr %arr, i64 %min.idx.ext
+  %val = load float, ptr %gep, align 4
+  %cmp = fcmp olt float %val, 0.0
+  %iv.next = add nsw i64 %iv, 1
+  %new.idx = trunc nsw i64 %iv.next to i32
+  %min.idx.next = select i1 %cmp, i32 %new.idx, i32 %min.idx
+  %trip = icmp slt i64 %iv.next, 100
+  br i1 %trip, label %loop, label %exit
+
+exit:
+  ret i32 %min.idx.next
+}
+
+; Negative: latch select's new-value arm has multiple uses, so m_OneUse on the
+; trunc fails; the recurrence is not widened.
+define i32 @widen_select_recurrence_multi_use_trunc(ptr %arr) {
+; CHECK-LABEL: @widen_select_recurrence_multi_use_trunc(
+; CHECK-NEXT:  entry:
+; CHECK-NEXT:    br label [[LOOP:%.*]]
+; CHECK:       loop:
+; CHECK-NEXT:    [[IV:%.*]] = phi i64 [ 0, [[ENTRY:%.*]] ], [ [[IV_NEXT:%.*]], [[LOOP]] ]
+; CHECK-NEXT:    [[MIN_IDX:%.*]] = phi i32 [ 0, [[ENTRY]] ], [ [[MIN_IDX_NEXT:%.*]], [[LOOP]] ]
+; CHECK-NEXT:    [[MIN_IDX_EXT:%.*]] = sext i32 [[MIN_IDX]] to i64
+; CHECK-NEXT:    [[GEP:%.*]] = getelementptr float, ptr [[ARR:%.*]], i64 [[MIN_IDX_EXT]]
+; CHECK-NEXT:    [[VAL:%.*]] = load float, ptr [[GEP]], align 4
+; CHECK-NEXT:    [[CMP:%.*]] = fcmp olt float [[VAL]], 0.000000e+00
+; CHECK-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; CHECK-NEXT:    [[NEW_IDX:%.*]] = trunc nsw i64 [[IV_NEXT]] to i32
+; CHECK-NEXT:    store i32 [[NEW_IDX]], ptr [[ARR]], align 4
+; CHECK-NEXT:    [[MIN_IDX_NEXT]] = select i1 [[CMP]], i32 [[NEW_IDX]], i32 [[MIN_IDX]]
+; CHECK-NEXT:    [[EXITCOND:%.*]] = icmp ne i64 [[IV_NEXT]], 100
+; CHECK-NEXT:    br i1 [[EXITCOND]], label [[LOOP]], label [[EXIT:%.*]]
+; CHECK:       exit:
+; CHECK-NEXT:    [[MIN_IDX_NEXT_LCSSA:%.*]] = phi i32 [ [[MIN_IDX_NEXT]], [[LOOP]] ]
+; CHECK-NEXT:    ret i32 [[MIN_IDX_NEXT_LCSSA]]
+;
+entry:
+  br label %loop
+
+loop:
+  %iv = phi i64 [ 0, %entry ], [ %iv.next, %loop ]
+  %min.idx = phi i32 [ 0, %entry ], [ %min.idx.next, %loop ]
+  %min.idx.ext = sext i32 %min.idx to i64
+  %gep = getelementptr float, ptr %arr, i64 %min.idx.ext
+  %val = load float, ptr %gep, align 4
+  %cmp = fcmp olt float %val, 0.0
+  %iv.next = add nsw i64 %iv, 1
+  %new.idx = trunc nsw i64 %iv.next to i32
+  store i32 %new.idx, ptr %arr, align 4
+  %min.idx.next = select i1 %cmp, i32 %new.idx, i32 %min.idx
+  %trip = icmp slt i64 %iv.next, 100
+  br i1 %trip, label %loop, label %exit
+
+exit:
+  ret i32 %min.idx.next
+}
+
+; Negative: the wide cast user is a zext, not a sext; the widener requires sext.
+define i32 @widen_select_recurrence_zext_user(ptr %arr) {
+; CHECK-LABEL: @widen_select_recurrence_zext_user(
+; CHECK-NEXT:  entry:
+; CHECK-NEXT:    br label [[LOOP:%.*]]
+; CHECK:       loop:
+; CHECK-NEXT:    [[IV:%.*]] = phi i64 [ 0, [[ENTRY:%.*]] ], [ [[IV_NEXT:%.*]], [[LOOP]] ]
+; CHECK-NEXT:    [[MIN_IDX:%.*]] = phi i32 [ 0, [[ENTRY]] ], [ [[MIN_IDX_NEXT:%.*]], [[LOOP]] ]
+; CHECK-NEXT:    [[MIN_IDX_EXT:%.*]] = zext i32 [[MIN_IDX]] to i64
+; CHECK-NEXT:    [[GEP:%.*]] = getelementptr float, ptr [[ARR:%.*]], i64 [[MIN_IDX_EXT]]
+; CHECK-NEXT:    [[VAL:%.*]] = load float, ptr [[GEP]], align 4
+; CHECK-NEXT:    [[CMP:%.*]] = fcmp olt float [[VAL]], 0.000000e+00
+; CHECK-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; CHECK-NEXT:    [[NEW_IDX:%.*]] = trunc nsw i64 [[IV_NEXT]] to i32
+; CHECK-NEXT:    [[MIN_IDX_NEXT]] = select i1 [[CMP]], i32 [[NEW_IDX]], i32 [[MIN_IDX]]
+; CHECK-NEXT:    [[EXITCOND:%.*]] = icmp ne i64 [[IV_NEXT]], 100
+; CHECK-NEXT:    br i1 [[EXITCOND]], label [[LOOP]], label [[EXIT:%.*]]
+; CHECK:       exit:
+; CHECK-NEXT:    [[MIN_IDX_NEXT_LCSSA:%.*]] = phi i32 [ [[MIN_IDX_NEXT]], [[LOOP]] ]
+; CHECK-NEXT:    ret i32 [[MIN_IDX_NEXT_LCSSA]]
+;
+entry:
+  br label %loop
+
+loop:
+  %iv = phi i64 [ 0, %entry ], [ %iv.next, %loop ]
+  %min.idx = phi i32 [ 0, %entry ], [ %min.idx.next, %loop ]
+  %min.idx.ext = zext i32 %min.idx to i64
+  %gep = getelementptr float, ptr %arr, i64 %min.idx.ext
+  %val = load float, ptr %gep, align 4
+  %cmp = fcmp olt float %val, 0.0
+  %iv.next = add nsw i64 %iv, 1
+  %new.idx = trunc nsw i64 %iv.next to i32
+  %min.idx.next = select i1 %cmp, i32 %new.idx, i32 %min.idx
+  %trip = icmp slt i64 %iv.next, 100
+  br i1 %trip, label %loop, label %exit
+
+exit:
+  ret i32 %min.idx.next
+}



More information about the llvm-commits mailing list