[llvm] [QTOOL-145035] - Fix hwloop trip count for post-increment load induction (PR #225411)

Santanu Das via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 22 07:36:16 PDT 2026


https://github.com/quic-santdas created https://github.com/llvm/llvm-project/pull/225411

A pointer-chasing loop whose exit test is a null check was miscompiled into a hardware loop with a bogus trip count.

A post-increment load defines two registers:
%3, %11 = L2_loadri_pi %1(tied-def 1), 4
%11 is the incremented address (tied to the base), but %3 is the value loaded from memory, which bears no relation to the base.  The helper only checked that the base register is defined by the PHI, never that the register feeding the PHI back from the latch is the incremented address. Here the PHI is fed by %3 -- so the loop was treated as bumping its induction variable by a constant 4 per iteration, and getLoopTripCount derived a count from the numeric value of head.

Require that the register feeding the PHI is the post-incremented address, located through the base operand's tie rather than a fixed operand index: the tied def is operand 1 for loads (L2_loadri_pi, V6_vL32b_pi) but operand 0 for stores (S2_storeri_pi, V6_vS32b_pi), and predicated forms shift both positions.

Change-Id: I99b13af1ee05fa6dccabfa35884fc2c8f4b932d3

>From dc6059f54b79839ae1ac9fed79e799b2e5984a6a Mon Sep 17 00:00:00 2001
From: quic-santdas <quic_santdas at qti.qualcomm.com>
Date: Thu, 27 Aug 2026 04:33:49 -0700
Subject: [PATCH] [QTOOL-145035] - Fix hwloop trip count for post-increment
 load induction

A pointer-chasing loop whose exit test is a null check was miscompiled
into a hardware loop with a bogus trip count.

A post-increment load defines two registers:
%3, %11 = L2_loadri_pi %1(tied-def 1), 4
%11 is the incremented address (tied to the base), but %3 is the value
loaded from memory, which bears no relation to the base.  The helper only
checked that the base register is defined by the PHI, never that the
register feeding the PHI back from the latch is the incremented address.
Here the PHI is fed by %3 -- so the loop was treated as
bumping its induction variable by a constant 4 per iteration, and
getLoopTripCount derived a count from the numeric value of head.

Require that the register feeding the PHI is the post-incremented
address, located through the base operand's tie rather than a fixed
operand index: the tied def is operand 1 for loads (L2_loadri_pi,
V6_vL32b_pi) but operand 0 for stores (S2_storeri_pi, V6_vS32b_pi), and
predicated forms shift both positions.

Change-Id: I99b13af1ee05fa6dccabfa35884fc2c8f4b932d3
---
 .../Target/Hexagon/HexagonHardwareLoops.cpp   | 47 ++++++++++++++
 .../Hexagon/hwloop-postinc-iv-ptrchase.ll     | 41 ++++++++++++
 .../hwloop-postinc-iv-tied-operands-hvx.ll    | 55 ++++++++++++++++
 .../hwloop-postinc-iv-tied-operands.ll        | 63 +++++++++++++++++++
 4 files changed, 206 insertions(+)
 create mode 100644 llvm/test/CodeGen/Hexagon/hwloop-postinc-iv-ptrchase.ll
 create mode 100644 llvm/test/CodeGen/Hexagon/hwloop-postinc-iv-tied-operands-hvx.ll
 create mode 100644 llvm/test/CodeGen/Hexagon/hwloop-postinc-iv-tied-operands.ll

diff --git a/llvm/lib/Target/Hexagon/HexagonHardwareLoops.cpp b/llvm/lib/Target/Hexagon/HexagonHardwareLoops.cpp
index 9ec0b1810b0ba..b1b02907b2fcb 100644
--- a/llvm/lib/Target/Hexagon/HexagonHardwareLoops.cpp
+++ b/llvm/lib/Target/Hexagon/HexagonHardwareLoops.cpp
@@ -272,6 +272,14 @@ namespace {
     /// value, either directly, or via a register.
     void setImmediate(MachineOperand &MO, int64_t Val);
 
+    /// If DI is a post-increment instruction whose base register is defined
+    /// by Phi and whose incremented address is PhiOpReg (the register feeding
+    /// Phi from the latch), extract the induction register and immediate bump
+    /// into IndReg and IVBump and return true.  Returns false otherwise.
+    bool tryExtractPostIncInduction(MachineInstr *DI, MachineInstr *Phi,
+                                    Register PhiOpReg, Register &IndReg,
+                                    int64_t &IVBump) const;
+
     /// Fix the data flow of the induction variable.
     /// The desired flow is: phi ---> bump -+-> comparison-in-latch.
     ///                                     |
@@ -399,6 +407,35 @@ bool HexagonHardwareLoops::runOnMachineFunction(MachineFunction &MF) {
   return Changed;
 }
 
+bool HexagonHardwareLoops::tryExtractPostIncInduction(MachineInstr *DI,
+                                                      MachineInstr *Phi,
+                                                      Register PhiOpReg,
+                                                      Register &IndReg,
+                                                      int64_t &IVBump) const {
+  if (!TII->isPostIncWithImmOffset(*DI))
+    return false;
+
+  unsigned BasePos, OffsetPos;
+  if (!TII->getBaseAndOffsetPosition(*DI, BasePos, OffsetPos))
+    return false;
+
+  if (BasePos >= DI->getNumOperands() || OffsetPos >= DI->getNumOperands())
+    return false;
+
+  // A post-increment load also defines the loaded value, which is unrelated
+  // to the base.  Only the incremented address, tied to the base operand, is
+  // "base + offset", so require that it is what feeds the PHI.
+  const MachineOperand &BaseOp = DI->getOperand(BasePos);
+  if (!BaseOp.isReg() || !BaseOp.isTied())
+    return false;
+  if (DI->getOperand(DI->findTiedOperandIdx(BasePos)).getReg() != PhiOpReg)
+    return false;
+
+  IndReg = BaseOp.getReg();
+  IVBump = DI->getOperand(OffsetPos).getImm();
+  return MRI->getVRegDef(IndReg) == Phi;
+}
+
 bool HexagonHardwareLoops::findInductionRegister(MachineLoop *L,
                                                  Register &Reg,
                                                  int64_t &IVBump,
@@ -449,6 +486,11 @@ bool HexagonHardwareLoops::findInductionRegister(MachineLoop *L,
           Register UpdReg = DI->getOperand(0).getReg();
           IndMap.insert(std::make_pair(UpdReg, std::make_pair(IndReg, V)));
         }
+      } else {
+        Register IndReg;
+        int64_t V;
+        if (tryExtractPostIncInduction(DI, Phi, PhiOpReg, IndReg, V))
+          IndMap.insert(std::make_pair(PhiOpReg, std::make_pair(IndReg, V)));
       }
     }  // for (i)
   }  // for (instr)
@@ -1713,6 +1755,11 @@ bool HexagonHardwareLoops::fixupInductionVariable(MachineLoop *L) {
           Register UpdReg = DI->getOperand(0).getReg();
           IndRegs.insert(std::make_pair(UpdReg, std::make_pair(IndReg, V)));
         }
+      } else {
+        Register IndReg;
+        int64_t V;
+        if (tryExtractPostIncInduction(DI, Phi, PhiReg, IndReg, V))
+          IndRegs.insert(std::make_pair(PhiReg, std::make_pair(IndReg, V)));
       }
     }  // for (i)
   }  // for (instr)
diff --git a/llvm/test/CodeGen/Hexagon/hwloop-postinc-iv-ptrchase.ll b/llvm/test/CodeGen/Hexagon/hwloop-postinc-iv-ptrchase.ll
new file mode 100644
index 0000000000000..d51005fb00e38
--- /dev/null
+++ b/llvm/test/CodeGen/Hexagon/hwloop-postinc-iv-ptrchase.ll
@@ -0,0 +1,41 @@
+; RUN: llc -mtriple=hexagon -O2 < %s | FileCheck %s
+;
+; A pointer-chasing loop whose trip count is not computable: the traversal
+; pointer is reloaded from memory each iteration (n = n->next) and the loop
+; exits when it becomes null.  The post-increment load in this loop defines
+; two registers -- the loaded value and the incremented address -- and only
+; the incremented address is a "base + constant" induction variable.
+;
+; HexagonHardwareLoops must not mistake the loaded value for the bumped
+; address; doing so makes it derive a trip count from the numeric value of
+; the initial head pointer (sub(#0,rN) / lsr(rN,#2)) and drop the null test
+; from the loop body, so the loop runs past the end of the list.
+
+; CHECK-LABEL: find:
+; CHECK-NOT: loop0
+; CHECK-NOT: endloop0
+
+%struct.S = type { i32, i32 }
+%struct.N = type { ptr, %struct.S }
+
+define ptr @find(ptr %head, i32 %t) {
+entry:
+  %cmp.not6 = icmp eq ptr %head, null
+  br i1 %cmp.not6, label %for.end, label %for.body
+
+for.body:
+  %r.08 = phi ptr [ %spec.select, %for.body ], [ null, %entry ]
+  %n.07 = phi ptr [ %next, %for.body ], [ %head, %entry ]
+  %s = getelementptr inbounds nuw i8, ptr %n.07, i32 4
+  %valp = getelementptr inbounds nuw i8, ptr %n.07, i32 8
+  %val = load i32, ptr %valp, align 4
+  %cmp1 = icmp eq i32 %val, %t
+  %next = load ptr, ptr %n.07, align 4
+  %spec.select = select i1 %cmp1, ptr %s, ptr %r.08
+  %cmp.not = icmp eq ptr %next, null
+  br i1 %cmp.not, label %for.end, label %for.body
+
+for.end:
+  %r.0.lcssa = phi ptr [ null, %entry ], [ %spec.select, %for.body ]
+  ret ptr %r.0.lcssa
+}
diff --git a/llvm/test/CodeGen/Hexagon/hwloop-postinc-iv-tied-operands-hvx.ll b/llvm/test/CodeGen/Hexagon/hwloop-postinc-iv-tied-operands-hvx.ll
new file mode 100644
index 0000000000000..73ac1dc80b3a3
--- /dev/null
+++ b/llvm/test/CodeGen/Hexagon/hwloop-postinc-iv-tied-operands-hvx.ll
@@ -0,0 +1,55 @@
+; RUN: llc -mtriple=hexagon -mattr=+hvxv68,+hvx-length128b -O2 < %s \
+; RUN:   | FileCheck %s
+;
+; HVX counterparts of hwloop-postinc-iv-tied-operands.ll.  The vector
+; post-increment forms have the same two operand shapes as the scalar ones,
+; so the tied def index likewise differs by opcode:
+;
+;   V6_vS32b_pi    (outs $Rx32),         (ins $Rx32in, $Ii, $Vs32)  -> tied def 0
+;   V6_vL32b_pi    (outs $Vd32, $Rx32),  (ins $Rx32in, $Ii)         -> tied def 1
+;
+; Both loops have a genuine "base + constant" induction variable (the PHI is
+; fed by the incremented address), so a hardware loop must be formed.
+
+; --- V6_vS32b_pi: single def, tied def index 0 -----------------------------
+; CHECK-LABEL: vstore_pi_tied0:
+; CHECK: loop0
+define void @vstore_pi_tied0(ptr %begin, ptr %end, <32 x i32> %v) {
+entry:
+  %empty = icmp eq ptr %begin, %end
+  br i1 %empty, label %exit, label %loop
+
+loop:
+  %ptr = phi ptr [ %begin, %entry ], [ %ptr.next, %loop ]
+  %old = load <32 x i32>, ptr %ptr, align 128
+  %sum = add <32 x i32> %old, %v
+  store <32 x i32> %sum, ptr %ptr, align 128
+  %ptr.next = getelementptr inbounds <32 x i32>, ptr %ptr, i32 1
+  %done = icmp eq ptr %ptr.next, %end
+  br i1 %done, label %exit, label %loop
+
+exit:
+  ret void
+}
+
+; --- V6_vL32b_pi: two defs, tied def index 1 ------------------------------
+; CHECK-LABEL: vload_pi_tied1:
+; CHECK: loop0
+define <32 x i32> @vload_pi_tied1(ptr %begin, ptr %end) {
+entry:
+  %empty = icmp eq ptr %begin, %end
+  br i1 %empty, label %exit, label %loop
+
+loop:
+  %ptr = phi ptr [ %begin, %entry ], [ %ptr.next, %loop ]
+  %acc = phi <32 x i32> [ zeroinitializer, %entry ], [ %acc.next, %loop ]
+  %val = load <32 x i32>, ptr %ptr, align 128
+  %acc.next = add <32 x i32> %acc, %val
+  %ptr.next = getelementptr inbounds <32 x i32>, ptr %ptr, i32 1
+  %done = icmp eq ptr %ptr.next, %end
+  br i1 %done, label %exit, label %loop
+
+exit:
+  %result = phi <32 x i32> [ zeroinitializer, %entry ], [ %acc.next, %loop ]
+  ret <32 x i32> %result
+}
diff --git a/llvm/test/CodeGen/Hexagon/hwloop-postinc-iv-tied-operands.ll b/llvm/test/CodeGen/Hexagon/hwloop-postinc-iv-tied-operands.ll
new file mode 100644
index 0000000000000..b10cbff57dcbf
--- /dev/null
+++ b/llvm/test/CodeGen/Hexagon/hwloop-postinc-iv-tied-operands.ll
@@ -0,0 +1,63 @@
+; RUN: llc -mtriple=hexagon -O2 < %s | FileCheck %s
+;
+; HexagonHardwareLoops recognizes a post-increment memory instruction as the
+; induction-variable bump.  Such instructions tie their base-address use to
+; the def that holds the incremented address, but the index of that tied def
+; depends on the opcode shape:
+;
+;   S2_storeri_pi  (outs $Rx32),         (ins $Rx32in, $Ii, $Rt32)  -> tied def 0
+;   L2_loadri_pi   (outs $Rd32, $Rx32),  (ins $Rx32in, $Ii)         -> tied def 1
+;
+; The pass must locate the incremented address through the base operand's tie
+; rather than a fixed operand index, and must require that the register fed
+; back into the loop PHI is that incremented address.  A post-increment load
+; also defines the *loaded value*, which bears no relation to the base; see
+; hwloop-postinc-iv-ptrchase.ll for the miscompile that results from
+; confusing the two.  The HVX vector forms are covered by
+; hwloop-postinc-iv-tied-operands-hvx.ll.
+;
+; These are the positive cases: each loop has a genuine "base + constant"
+; induction variable, so a hardware loop must still be formed.
+
+; --- S2_storeri_pi: single def, tied def index 0 ---------------------------
+; CHECK-LABEL: store_pi_tied0:
+; CHECK: loop0
+define void @store_pi_tied0(ptr %begin, ptr %end, i32 %v) {
+entry:
+  %empty = icmp eq ptr %begin, %end
+  br i1 %empty, label %exit, label %loop
+
+loop:
+  %ptr = phi ptr [ %begin, %entry ], [ %ptr.next, %loop ]
+  store i32 %v, ptr %ptr, align 4
+  %ptr.next = getelementptr inbounds i32, ptr %ptr, i32 1
+  %done = icmp eq ptr %ptr.next, %end
+  br i1 %done, label %exit, label %loop
+
+exit:
+  ret void
+}
+
+; --- L2_loadri_pi: two defs, tied def index 1 -----------------------------
+; The PHI is fed by the incremented address, not the loaded value, so this
+; is a real induction variable.
+; CHECK-LABEL: load_pi_tied1:
+; CHECK: loop0
+define i32 @load_pi_tied1(ptr %begin, ptr %end) {
+entry:
+  %empty = icmp eq ptr %begin, %end
+  br i1 %empty, label %exit, label %loop
+
+loop:
+  %ptr = phi ptr [ %begin, %entry ], [ %ptr.next, %loop ]
+  %acc = phi i32 [ 0, %entry ], [ %acc.next, %loop ]
+  %val = load i32, ptr %ptr, align 4
+  %acc.next = add i32 %acc, %val
+  %ptr.next = getelementptr inbounds i32, ptr %ptr, i32 1
+  %done = icmp eq ptr %ptr.next, %end
+  br i1 %done, label %exit, label %loop
+
+exit:
+  %result = phi i32 [ 0, %entry ], [ %acc.next, %loop ]
+  ret i32 %result
+}



More information about the llvm-commits mailing list