[polly] [Polly] Distribute the innermost point loop of tile over its statements (PR #225844)
Timur Baidusenov via llvm-commits
llvm-commits at lists.llvm.org
Sat Sep 26 10:41:37 PDT 2026
https://github.com/bai-tim updated https://github.com/llvm/llvm-project/pull/225844
>From 07848270a736dbc32b6d6d9bf33b1a51d241ecef Mon Sep 17 00:00:00 2001
From: Timur Baidusenov <timurbaidusenov at gmail.com>
Date: Wed, 23 Sep 2026 19:00:15 +0300
Subject: [PATCH 1/3] [Polly] Distribute the innermost point loop of a tile
over its statements
When the scheduler time-tiles a stencil such as jacobi-2d or heat-3d, it skews the spatial loops and shifts the statements of a time step against each other so that they form a single permutable band; after tiling, the innermost point loop runs all statements under conditions and carries the dependences between them, so it is not vectorized and the tiled code is up to 2.2 times slower than plain -O3. This patch distributes the innermost loop of a tiled innermost band over its statements when that is legal and makes every resulting loop parallel: considering only the dependences between instances that share the iterations of the outer loops, statements that depend on each other stay in one group, no group may carry a dependence in the innermost loop, and the groups are emitted in dependence order, while a loop that is parallel already is only distributed over statements that do not exchange data within it, which keeps fusion that provides reuse; the transformation is skipped with prevectorization or register tiling, leaves the band unchanged if its analysis exceeds the isl operation quota, and can be disabled with -polly-distribute-point-loops=false. On PolyBench/C 4.2.1 LARGE (x86-64, -O3 -march=native -polly) jacobi-1d, jacobi-2d and heat-3d become 1.97, 2.18 and 2.20 times faster, which puts jacobi-2d 22% ahead of plain -O3, and no other kernel changes; in llvm-test-suite only these three programs change (jacobi-2d 1.87 s to 1.07 s, heat-3d 1.88 s to 0.85 s), all tests pass, CTMark generates identical code, and the compile time of the three files grows by 28 to 73 ms.
Assisted-by: Claude (Anthropic)
---
polly/lib/Transform/ScheduleOptimizer.cpp | 155 +++++++++++++++++
.../distribute_point_loops.ll | 162 ++++++++++++++++++
.../distribute_point_loops_cycle.ll | 70 ++++++++
3 files changed, 387 insertions(+)
create mode 100644 polly/test/ScheduleOptimizer/distribute_point_loops.ll
create mode 100644 polly/test/ScheduleOptimizer/distribute_point_loops_cycle.ll
diff --git a/polly/lib/Transform/ScheduleOptimizer.cpp b/polly/lib/Transform/ScheduleOptimizer.cpp
index b9b9abbd85ae4..c1611be5c296c 100644
--- a/polly/lib/Transform/ScheduleOptimizer.cpp
+++ b/polly/lib/Transform/ScheduleOptimizer.cpp
@@ -55,6 +55,7 @@
#include "polly/ScopInfo.h"
#include "polly/Support/ISLOStream.h"
#include "polly/Support/ISLTools.h"
+#include "llvm/ADT/BitVector.h"
#include "llvm/ADT/Sequence.h"
#include "llvm/ADT/Statistic.h"
#include "llvm/Analysis/OptimizationRemarkEmitter.h"
@@ -159,6 +160,12 @@ static cl::opt<bool> RegisterTiling("polly-register-tiling",
cl::desc("Enable register tiling"),
cl::cat(PollyCategory));
+static cl::opt<bool> DistributePointLoops(
+ "polly-distribute-point-loops",
+ cl::desc("Distribute the innermost point loop of a tile over its "
+ "statements if this makes the loop of every statement parallel"),
+ cl::init(true), cl::cat(PollyCategory));
+
static cl::opt<int> RegisterDefaultTileSize(
"polly-register-tiling-default-tile-size",
cl::desc("The default register tile size (if not enough were provided by"
@@ -226,6 +233,8 @@ STATISTIC(FirstLevelTileOpts, "Number of first level tiling applied");
STATISTIC(SecondLevelTileOpts, "Number of second level tiling applied");
STATISTIC(RegisterTileOpts, "Number of register tiling applied");
STATISTIC(PrevectOpts, "Number of strip-mining for prevectorization applied");
+STATISTIC(PointLoopDistributions,
+ "Number of innermost point loops distributed over statements");
STATISTIC(MatMulOpts,
"Number of matrix multiplication patterns detected and optimized");
@@ -377,6 +386,25 @@ class ScheduleTreeOptimizer final {
/// @param Node The schedule node to (possibly) optimize.
static isl::schedule_node applyTileBandOpt(isl::schedule_node Node);
+ /// Distribute the innermost loop of a band over its statements.
+ ///
+ /// Time tiling of a stencil shifts the statements of a time step against
+ /// each other and fuses them in the innermost point loop. The loop then runs
+ /// each statement under a condition and carries the dependences between
+ /// them, which keeps it from being vectorized. Run the innermost loop once
+ /// for each group of statements instead, where statements that depend on
+ /// each other form a group, if the groups can be ordered and none of them
+ /// carries a dependence in the innermost loop. A loop that is parallel
+ /// already is only distributed over statements that do not exchange any
+ /// data within it.
+ ///
+ /// @param Node The innermost band to distribute, typically the point band of
+ /// a tiling.
+ /// @param D The dependences of the SCoP.
+ /// @return The node at the position of @p Node.
+ static isl::schedule_node distributeInnermostLoop(isl::schedule_node Node,
+ const Dependences *D);
+
/// Apply prevectorization on the bands in the schedule tree.
///
/// @param Node The schedule node to (possibly) prevectorize.
@@ -548,6 +576,114 @@ ScheduleTreeOptimizer::applyTileBandOpt(isl::schedule_node Node) {
return Node;
}
+isl::schedule_node
+ScheduleTreeOptimizer::distributeInnermostLoop(isl::schedule_node Node,
+ const Dependences *D) {
+ if (!isSimpleInnermostBand(Node))
+ return Node;
+
+ // Number the statements in the order of their names, so that the result does
+ // not depend on the order in which isl lists them.
+ isl::union_set Domain = Node.get_domain();
+ SmallVector<isl::set, 8> Stmts;
+ for (isl::set Stmt : Domain.get_set_list())
+ Stmts.push_back(Stmt);
+ unsigned NumStmts = Stmts.size();
+ if (NumStmts < 2)
+ return Node;
+ llvm::sort(Stmts, [](const isl::set &A, const isl::set &B) {
+ return A.get_tuple_name() < B.get_tuple_name();
+ });
+ DenseMap<isl_id *, unsigned> StmtIndex;
+ for (auto [Idx, Stmt] : enumerate(Stmts))
+ StmtIndex[Stmt.get_tuple_id().get()] = Idx;
+
+ isl::schedule_node_band Band = Node.as<isl::schedule_node_band>();
+ unsigned NumMembers = unsignedFromIslSize(Band.n_member());
+ isl::schedule_node_band Inner = Band;
+ if (NumMembers > 1)
+ Inner = Band.split(NumMembers - 1).child(0).as<isl::schedule_node_band>();
+
+ // The dependences between the instances that share the iterations of all
+ // loops around the innermost one, split into those that the innermost loop
+ // carries and those within one of its iterations.
+ isl::union_map Deps = D->getDependences(
+ Dependences::TYPE_RAW | Dependences::TYPE_WAR | Dependences::TYPE_WAW);
+ Deps = Deps.intersect_domain(Domain).intersect_range(Domain);
+ Deps = Deps.eq_at(Inner.get_prefix_schedule_multi_union_pw_aff());
+ isl::union_map SameIteration = Deps.eq_at(Inner.get_partial_schedule());
+ isl::union_map Carried = Deps.subtract(SameIteration);
+ if (Carried.is_null())
+ return Node;
+
+ auto getIndex = [&](const isl::map &Dep, isl::dim Dim) {
+ return StmtIndex.lookup(Dep.get_tuple_id(Dim).get());
+ };
+
+ // If the fused loop is parallel already, only separate statements that do
+ // not exchange any data within it, to keep the reuse between the others.
+ if (Carried.is_empty().is_true())
+ for (isl::map Dep : Deps.get_map_list())
+ if (getIndex(Dep, isl::dim::in) != getIndex(Dep, isl::dim::out))
+ return Node;
+
+ // Reaches[I][J] is set if statement J depends on statement I, possibly
+ // through other statements. Statements that depend on each other form a
+ // group.
+ SmallVector<BitVector, 8> Reaches(NumStmts, BitVector(NumStmts));
+ for (isl::map Dep : Deps.get_map_list())
+ Reaches[getIndex(Dep, isl::dim::in)].set(getIndex(Dep, isl::dim::out));
+ for (unsigned K : seq(NumStmts))
+ for (unsigned I : seq(NumStmts))
+ if (Reaches[I][K])
+ Reaches[I] |= Reaches[K];
+
+ // The loop of every group must be parallel.
+ for (isl::map Dep : Carried.get_map_list())
+ if (Reaches[getIndex(Dep, isl::dim::out)][getIndex(Dep, isl::dim::in)])
+ return Node;
+
+ SmallVector<unsigned, 8> GroupOf(NumStmts);
+ unsigned NumGroups = 0;
+ for (unsigned I : seq(NumStmts)) {
+ auto Earlier = find_if(
+ seq(I), [&](unsigned J) { return Reaches[I][J] && Reaches[J][I]; });
+ GroupOf[I] = Earlier != seq(I).end() ? GroupOf[*Earlier] : NumGroups++;
+ }
+ if (NumGroups < 2)
+ return Node;
+
+ // A group depends on more statements than every group it depends on, so
+ // ordering by this number respects the dependences.
+ SmallVector<unsigned, 8> Leader(NumGroups);
+ for (unsigned Stmt : reverse(seq(NumStmts)))
+ Leader[GroupOf[Stmt]] = Stmt;
+ auto NumPredecessors = [&](unsigned Group) {
+ return count_if(seq(NumStmts), [&](unsigned Pred) {
+ return GroupOf[Pred] != Group && Reaches[Pred][Leader[Group]];
+ });
+ };
+ SmallVector<unsigned, 8> Order = to_vector<8>(seq(NumGroups));
+ stable_sort(Order, [&](unsigned A, unsigned B) {
+ return NumPredecessors(A) < NumPredecessors(B);
+ });
+
+ isl::union_set_list Filters(Node.ctx(), NumGroups);
+ for (unsigned Group : Order) {
+ isl::union_set Filter = isl::union_set::empty(Node.ctx());
+ for (unsigned Stmt : seq(NumStmts))
+ if (GroupOf[Stmt] == Group)
+ Filter = Filter.unite(Stmts[Stmt]);
+ Filters = Filters.add(Filter);
+ }
+
+ isl::schedule_node Sequence = Inner.insert_sequence(Filters);
+ if (Sequence.is_null())
+ return Node;
+ PointLoopDistributions++;
+ return NumMembers > 1 ? Sequence.parent() : Sequence;
+}
+
isl::schedule_node
ScheduleTreeOptimizer::applyPrevectBandOpt(isl::schedule_node Node) {
auto Space = isl::manage(isl_schedule_node_band_get_space(Node.get()));
@@ -589,6 +725,25 @@ ScheduleTreeOptimizer::optimizeBand(__isl_take isl_schedule_node *NodeArg,
if (OAI->Postopts)
Node = applyTileBandOpt(Node);
+ // Prevectorization strip-mines a parallel point loop and register tiling
+ // unrolls the point loops, leave those to them.
+ if (OAI->Postopts && DistributePointLoops && !OAI->Prevect &&
+ !RegisterTiling) {
+ isl::schedule_node Distributed;
+ {
+ IslQuotaScope MaxScope = OAI->MaxOpGuard.enter();
+ Distributed = distributeInnermostLoop(Node, OAI->D);
+ // Leave the band as it is if the analysis exceeds the quota, rather than
+ // discarding the other optimizations of the SCoP.
+ if (OAI->MaxOpGuard.hasQuotaExceeded()) {
+ Distributed = {};
+ isl_ctx_reset_error(Node.ctx().get());
+ }
+ }
+ if (!Distributed.is_null())
+ Node = Distributed;
+ }
+
if (OAI->Prevect) {
IslQuotaScope MaxScope = OAI->MaxOpGuard.enter();
diff --git a/polly/test/ScheduleOptimizer/distribute_point_loops.ll b/polly/test/ScheduleOptimizer/distribute_point_loops.ll
new file mode 100644
index 0000000000000..bd6ca33eaedd1
--- /dev/null
+++ b/polly/test/ScheduleOptimizer/distribute_point_loops.ll
@@ -0,0 +1,162 @@
+; RUN: opt %loadNPMPolly '-passes=polly-custom<opt-isl;ast>' -polly-print-ast -disable-output < %s | FileCheck %s
+; RUN: opt %loadNPMPolly '-passes=polly-custom<opt-isl;ast>' -polly-print-ast -polly-distribute-point-loops=false -disable-output < %s | FileCheck %s --check-prefix=FUSED
+;
+; Distribute the innermost point loop of a time tiled stencil over its
+; statements.
+;
+; void jacobi(int tsteps, int n, double A[restrict n][n],
+; double B[restrict n][n]) {
+; for (int t = 0; t < tsteps; t++) {
+; for (int i = 1; i < n - 1; i++)
+; for (int j = 1; j < n - 1; j++)
+; B[i][j] = 0.2 * (A[i][j] + A[i][j - 1] + A[i][j + 1] +
+; A[i + 1][j] + A[i - 1][j]);
+; for (int i = 1; i < n - 1; i++)
+; for (int j = 1; j < n - 1; j++)
+; A[i][j] = 0.2 * (B[i][j] + B[i][j - 1] + B[i][j + 1] +
+; B[i + 1][j] + B[i - 1][j]);
+; }
+; }
+;
+; The scheduler skews the loops over i and j by 2t and shifts the second
+; statement by one iteration against the first, so that both form a single
+; permutable band. The innermost point loop then runs both statements under
+; conditions and carries the dependence of the second statement on the first.
+; Within an iteration of the loops around it, however, only the second
+; statement depends on the first, and neither loop carries a dependence on
+; its own, so the loop is run once for each statement.
+
+; CHECK: // 1st level tiling - Points
+; CHECK-NEXT: for (int c3 = {{.*}})
+; CHECK-NEXT: for (int c4 = {{.*}}) {
+; CHECK-NEXT: if ({{.*}})
+; CHECK-NEXT: for (int c5 = {{.*}})
+; CHECK-NEXT: Stmt_for_body9(
+; CHECK-NEXT: if ({{.*}})
+; CHECK-NEXT: for (int c5 = {{.*}})
+; CHECK-NEXT: Stmt_for_body53(
+; CHECK-NEXT: }
+
+; FUSED: // 1st level tiling - Points
+; FUSED-NEXT: for (int c3 = {{.*}})
+; FUSED-NEXT: for (int c4 = {{.*}})
+; FUSED-NEXT: for (int c5 = {{.*}}) {
+; FUSED-NEXT: if ({{.*}})
+; FUSED-NEXT: Stmt_for_body9(
+; FUSED-NEXT: if ({{.*}})
+; FUSED-NEXT: Stmt_for_body53(
+; FUSED-NEXT: }
+
+target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i64:64-i128:128-f80:128-n8:16:32:64-S128"
+
+define void @jacobi(i32 %tsteps, i32 %n, ptr noalias %A, ptr noalias %B) {
+entry:
+ %0 = zext i32 %n to i64
+ %cmp150 = icmp sgt i32 %tsteps, 0
+ br i1 %cmp150, label %for.cond1.preheader.lr.ph, label %for.cond.cleanup
+
+for.cond1.preheader.lr.ph:
+ %sub = add i32 %n, -1
+ %cmp2144 = icmp sgt i32 %n, 2
+ %cmp45148 = icmp sgt i32 %n, 2
+ %wide.trip.count157 = zext i32 %sub to i64
+ %wide.trip.count = zext i32 %sub to i64
+ %wide.trip.count168 = zext i32 %sub to i64
+ %wide.trip.count162 = zext i32 %sub to i64
+ br label %for.cond1.preheader
+
+for.cond1.preheader:
+ %t.0151 = phi i32 [ 0, %for.cond1.preheader.lr.ph ], [ %inc94, %for.cond.cleanup46 ]
+ br i1 %cmp2144, label %for.cond5.preheader, label %for.cond43.preheader
+
+for.cond.cleanup:
+ ret void
+
+for.cond43.preheader:
+ br i1 %cmp45148, label %for.cond49.preheader, label %for.cond.cleanup46
+
+for.cond5.preheader:
+ %indvars.iv153 = phi i64 [ %indvars.iv.next154, %for.cond5.for.cond.cleanup8_crit_edge ], [ 1, %for.cond1.preheader ]
+ %1 = mul nuw nsw i64 %indvars.iv153, %0
+ %arrayidx = getelementptr inbounds nuw [8 x i8], ptr %A, i64 %1
+ %indvars.iv.next154 = add nuw nsw i64 %indvars.iv153, 1
+ %2 = mul nuw nsw i64 %indvars.iv.next154, %0
+ %arrayidx25 = getelementptr inbounds nuw [8 x i8], ptr %A, i64 %2
+ %3 = add nsw i64 %indvars.iv153, -1
+ %4 = mul nuw nsw i64 %3, %0
+ %arrayidx31 = getelementptr inbounds [8 x i8], ptr %A, i64 %4
+ %arrayidx36 = getelementptr inbounds nuw [8 x i8], ptr %B, i64 %1
+ br label %for.body9
+
+for.cond5.for.cond.cleanup8_crit_edge:
+ %exitcond158.not = icmp eq i64 %indvars.iv.next154, %wide.trip.count157
+ br i1 %exitcond158.not, label %for.cond43.preheader, label %for.cond5.preheader
+
+for.body9:
+ %indvars.iv = phi i64 [ 1, %for.cond5.preheader ], [ %indvars.iv.next, %for.body9 ]
+ %arrayidx11 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx, i64 %indvars.iv
+ %5 = load double, ptr %arrayidx11, align 8
+ %arrayidx16 = getelementptr i8, ptr %arrayidx11, i64 -8
+ %6 = load double, ptr %arrayidx16, align 8
+ %add = fadd double %5, %6
+ %indvars.iv.next = add nuw nsw i64 %indvars.iv, 1
+ %arrayidx21 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx, i64 %indvars.iv.next
+ %7 = load double, ptr %arrayidx21, align 8
+ %add22 = fadd double %add, %7
+ %arrayidx27 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx25, i64 %indvars.iv
+ %8 = load double, ptr %arrayidx27, align 8
+ %add28 = fadd double %add22, %8
+ %arrayidx33 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx31, i64 %indvars.iv
+ %9 = load double, ptr %arrayidx33, align 8
+ %add34 = fadd double %add28, %9
+ %mul = fmul double %add34, 2.000000e-01
+ %arrayidx38 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx36, i64 %indvars.iv
+ store double %mul, ptr %arrayidx38, align 8
+ %exitcond.not = icmp eq i64 %indvars.iv.next, %wide.trip.count
+ br i1 %exitcond.not, label %for.cond5.for.cond.cleanup8_crit_edge, label %for.body9
+
+for.cond49.preheader:
+ %indvars.iv164 = phi i64 [ %indvars.iv.next165, %for.cond49.for.cond.cleanup52_crit_edge ], [ 1, %for.cond43.preheader ]
+ %10 = mul nuw nsw i64 %indvars.iv164, %0
+ %arrayidx55 = getelementptr inbounds nuw [8 x i8], ptr %B, i64 %10
+ %indvars.iv.next165 = add nuw nsw i64 %indvars.iv164, 1
+ %11 = mul nuw nsw i64 %indvars.iv.next165, %0
+ %arrayidx72 = getelementptr inbounds nuw [8 x i8], ptr %B, i64 %11
+ %12 = add nsw i64 %indvars.iv164, -1
+ %13 = mul nuw nsw i64 %12, %0
+ %arrayidx78 = getelementptr inbounds [8 x i8], ptr %B, i64 %13
+ %arrayidx84 = getelementptr inbounds nuw [8 x i8], ptr %A, i64 %10
+ br label %for.body53
+
+for.cond.cleanup46:
+ %inc94 = add nuw nsw i32 %t.0151, 1
+ %exitcond170.not = icmp eq i32 %inc94, %tsteps
+ br i1 %exitcond170.not, label %for.cond.cleanup, label %for.cond1.preheader
+
+for.cond49.for.cond.cleanup52_crit_edge:
+ %exitcond169.not = icmp eq i64 %indvars.iv.next165, %wide.trip.count168
+ br i1 %exitcond169.not, label %for.cond.cleanup46, label %for.cond49.preheader
+
+for.body53:
+ %indvars.iv159 = phi i64 [ 1, %for.cond49.preheader ], [ %indvars.iv.next160, %for.body53 ]
+ %arrayidx57 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx55, i64 %indvars.iv159
+ %14 = load double, ptr %arrayidx57, align 8
+ %arrayidx62 = getelementptr i8, ptr %arrayidx57, i64 -8
+ %15 = load double, ptr %arrayidx62, align 8
+ %add63 = fadd double %14, %15
+ %indvars.iv.next160 = add nuw nsw i64 %indvars.iv159, 1
+ %arrayidx68 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx55, i64 %indvars.iv.next160
+ %16 = load double, ptr %arrayidx68, align 8
+ %add69 = fadd double %add63, %16
+ %arrayidx74 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx72, i64 %indvars.iv159
+ %17 = load double, ptr %arrayidx74, align 8
+ %add75 = fadd double %add69, %17
+ %arrayidx80 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx78, i64 %indvars.iv159
+ %18 = load double, ptr %arrayidx80, align 8
+ %add81 = fadd double %add75, %18
+ %mul82 = fmul double %add81, 2.000000e-01
+ %arrayidx86 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx84, i64 %indvars.iv159
+ store double %mul82, ptr %arrayidx86, align 8
+ %exitcond163.not = icmp eq i64 %indvars.iv.next160, %wide.trip.count162
+ br i1 %exitcond163.not, label %for.cond49.for.cond.cleanup52_crit_edge, label %for.body53
+}
diff --git a/polly/test/ScheduleOptimizer/distribute_point_loops_cycle.ll b/polly/test/ScheduleOptimizer/distribute_point_loops_cycle.ll
new file mode 100644
index 0000000000000..964ed275ac287
--- /dev/null
+++ b/polly/test/ScheduleOptimizer/distribute_point_loops_cycle.ll
@@ -0,0 +1,70 @@
+; RUN: opt %loadNPMPolly -polly-stmt-granularity=store '-passes=polly-custom<simplify-0;optree;delicm;simplify-1;opt-isl;ast>' -polly-print-ast -disable-output < %s | FileCheck %s
+;
+; Do not distribute the innermost point loop if the statements in it depend on
+; each other in both directions.
+;
+; void cyc(int n, double A[restrict n][n], double B[restrict n][n]) {
+; for (int i = 1; i < n; i++)
+; for (int j = 1; j < n; j++) {
+; A[i][j] = B[i][j - 1] + A[i - 1][j];
+; B[i][j] = 2 * A[i][j];
+; }
+; }
+;
+; The second statement depends on the first within an iteration of j, and the
+; first on the second from the previous iteration of j.
+
+; CHECK: // 1st level tiling - Points
+; CHECK-NEXT: for (int c2 = {{.*}})
+; CHECK-NEXT: for (int c3 = {{.*}}) {
+; CHECK-NEXT: Stmt_for_body4(
+; CHECK-NEXT: Stmt_for_body4_b(
+; CHECK-NEXT: }
+
+target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i64:64-i128:128-f80:128-n8:16:32:64-S128"
+
+define void @cyc(i32 %n, ptr noalias %A, ptr noalias %B) {
+entry:
+ %0 = zext i32 %n to i64
+ %cmp49 = icmp sgt i32 %n, 1
+ br i1 %cmp49, label %for.cond1.preheader.preheader, label %for.cond.cleanup
+
+for.cond1.preheader.preheader:
+ %wide.trip.count56 = zext nneg i32 %n to i64
+ %wide.trip.count = zext nneg i32 %n to i64
+ br label %for.cond1.preheader
+
+for.cond1.preheader:
+ %indvars.iv52 = phi i64 [ 1, %for.cond1.preheader.preheader ], [ %indvars.iv.next53, %for.cond1.for.cond.cleanup3_crit_edge ]
+ %1 = mul nuw nsw i64 %indvars.iv52, %0
+ %arrayidx = getelementptr inbounds nuw [8 x i8], ptr %B, i64 %1
+ %2 = add nsw i64 %indvars.iv52, -1
+ %3 = mul nuw nsw i64 %2, %0
+ %arrayidx9 = getelementptr inbounds [8 x i8], ptr %A, i64 %3
+ %arrayidx13 = getelementptr inbounds nuw [8 x i8], ptr %A, i64 %1
+ %load_initial = load double, ptr %arrayidx, align 8
+ br label %for.body4
+
+for.cond.cleanup:
+ ret void
+
+for.cond1.for.cond.cleanup3_crit_edge:
+ %indvars.iv.next53 = add nuw nsw i64 %indvars.iv52, 1
+ %exitcond57.not = icmp eq i64 %indvars.iv.next53, %wide.trip.count56
+ br i1 %exitcond57.not, label %for.cond.cleanup, label %for.cond1.preheader
+
+for.body4:
+ %store_forwarded = phi double [ %load_initial, %for.cond1.preheader ], [ %mul, %for.body4 ]
+ %indvars.iv = phi i64 [ 1, %for.cond1.preheader ], [ %indvars.iv.next, %for.body4 ]
+ %4 = getelementptr [8 x i8], ptr %arrayidx, i64 %indvars.iv
+ %arrayidx11 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx9, i64 %indvars.iv
+ %5 = load double, ptr %arrayidx11, align 8
+ %add = fadd double %store_forwarded, %5
+ %arrayidx15 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx13, i64 %indvars.iv
+ store double %add, ptr %arrayidx15, align 8
+ %mul = fmul double %add, 2.000000e+00
+ store double %mul, ptr %4, align 8
+ %indvars.iv.next = add nuw nsw i64 %indvars.iv, 1
+ %exitcond.not = icmp eq i64 %indvars.iv.next, %wide.trip.count
+ br i1 %exitcond.not, label %for.cond1.for.cond.cleanup3_crit_edge, label %for.body4
+}
>From 7b121ac8d5ff213de6acac1500577066f37d0e62 Mon Sep 17 00:00:00 2001
From: Timur Baidusenov <timurbaidusenov at gmail.com>
Date: Thu, 24 Sep 2026 19:05:00 +0300
Subject: [PATCH 2/3] Address review: order statements as in the SCoP, check
isl results for errors
Number the statements in the order of Scop::Stmts instead of by their names, which depend on -polly-use-llvm-names and the build type. Since the analysis runs within a quota of isl operations, check the domain, its statement ids, the number of band members and the dependences for errors before using them in conditions, iterations or lookups, and leave the band unchanged if any of them failed. Explain why the error is reset after exceeding the quota.
Assisted-by: Claude (Anthropic)
---
polly/lib/Transform/ScheduleOptimizer.cpp | 103 ++++++++++++++++------
1 file changed, 74 insertions(+), 29 deletions(-)
diff --git a/polly/lib/Transform/ScheduleOptimizer.cpp b/polly/lib/Transform/ScheduleOptimizer.cpp
index c1611be5c296c..ce34a1d6097be 100644
--- a/polly/lib/Transform/ScheduleOptimizer.cpp
+++ b/polly/lib/Transform/ScheduleOptimizer.cpp
@@ -579,30 +579,50 @@ ScheduleTreeOptimizer::applyTileBandOpt(isl::schedule_node Node) {
isl::schedule_node
ScheduleTreeOptimizer::distributeInnermostLoop(isl::schedule_node Node,
const Dependences *D) {
- if (!isSimpleInnermostBand(Node))
+ // This runs within a quota of isl operations. Every isl result that is used
+ // in a condition, iterated over or dereferenced is checked for an error
+ // first; in that case the band is left as it is.
+ if (Node.is_null() || !isSimpleInnermostBand(Node))
return Node;
- // Number the statements in the order of their names, so that the result does
- // not depend on the order in which isl lists them.
isl::union_set Domain = Node.get_domain();
- SmallVector<isl::set, 8> Stmts;
- for (isl::set Stmt : Domain.get_set_list())
- Stmts.push_back(Stmt);
+ if (Domain.is_null())
+ return Node;
+ isl::set_list DomainList = Domain.get_set_list();
+ if (DomainList.is_null())
+ return Node;
+ SmallVector<std::pair<ScopStmt *, isl::set>, 8> Stmts;
+ for (isl::set Set : DomainList) {
+ isl::id Id = Set.get_tuple_id();
+ if (Id.is_null())
+ return Node;
+ Stmts.push_back({static_cast<ScopStmt *>(Id.get_user()), Set});
+ }
unsigned NumStmts = Stmts.size();
if (NumStmts < 2)
return Node;
- llvm::sort(Stmts, [](const isl::set &A, const isl::set &B) {
- return A.get_tuple_name() < B.get_tuple_name();
+
+ // Number the statements in the order of Scop::Stmts, so that the result does
+ // not depend on the order in which isl lists them.
+ DenseMap<const ScopStmt *, unsigned> ScopOrder;
+ for (const ScopStmt &Stmt : *Stmts.front().first->getParent())
+ ScopOrder.insert({&Stmt, ScopOrder.size()});
+ llvm::sort(Stmts, [&](const auto &A, const auto &B) {
+ return ScopOrder.lookup(A.first) < ScopOrder.lookup(B.first);
});
- DenseMap<isl_id *, unsigned> StmtIndex;
+ DenseMap<const ScopStmt *, unsigned> StmtIndex;
for (auto [Idx, Stmt] : enumerate(Stmts))
- StmtIndex[Stmt.get_tuple_id().get()] = Idx;
+ StmtIndex[Stmt.first] = Idx;
isl::schedule_node_band Band = Node.as<isl::schedule_node_band>();
- unsigned NumMembers = unsignedFromIslSize(Band.n_member());
+ isl::size NumMembers = Band.n_member();
+ if (NumMembers.is_error())
+ return Node;
isl::schedule_node_band Inner = Band;
- if (NumMembers > 1)
- Inner = Band.split(NumMembers - 1).child(0).as<isl::schedule_node_band>();
+ if (unsigned(NumMembers) > 1)
+ Inner = Band.split(unsigned(NumMembers) - 1)
+ .child(0)
+ .as<isl::schedule_node_band>();
// The dependences between the instances that share the iterations of all
// loops around the innermost one, split into those that the innermost loop
@@ -613,34 +633,58 @@ ScheduleTreeOptimizer::distributeInnermostLoop(isl::schedule_node Node,
Deps = Deps.eq_at(Inner.get_prefix_schedule_multi_union_pw_aff());
isl::union_map SameIteration = Deps.eq_at(Inner.get_partial_schedule());
isl::union_map Carried = Deps.subtract(SameIteration);
- if (Carried.is_null())
- return Node;
- auto getIndex = [&](const isl::map &Dep, isl::dim Dim) {
- return StmtIndex.lookup(Dep.get_tuple_id(Dim).get());
+ // The pairs of the indices of the source and target statements of the
+ // dependences in UMap; false if an isl operation failed.
+ using EdgeList = SmallVector<std::pair<unsigned, unsigned>, 16>;
+ auto collectEdges = [&](const isl::union_map &UMap, EdgeList &Edges) {
+ if (UMap.is_null())
+ return false;
+ isl::map_list List = UMap.get_map_list();
+ if (List.is_null())
+ return false;
+ for (isl::map Dep : List) {
+ isl::boolean IsEmpty = Dep.is_empty();
+ if (IsEmpty.is_error())
+ return false;
+ if (IsEmpty.is_true())
+ continue;
+ isl::id Src = Dep.get_tuple_id(isl::dim::in);
+ isl::id Dst = Dep.get_tuple_id(isl::dim::out);
+ if (Src.is_null() || Dst.is_null())
+ return false;
+ auto SrcIt = StmtIndex.find(static_cast<ScopStmt *>(Src.get_user()));
+ auto DstIt = StmtIndex.find(static_cast<ScopStmt *>(Dst.get_user()));
+ if (SrcIt == StmtIndex.end() || DstIt == StmtIndex.end())
+ return false;
+ Edges.push_back({SrcIt->second, DstIt->second});
+ }
+ return true;
};
+ EdgeList DepEdges, CarriedEdges;
+ if (!collectEdges(Deps, DepEdges) || !collectEdges(Carried, CarriedEdges))
+ return Node;
// If the fused loop is parallel already, only separate statements that do
// not exchange any data within it, to keep the reuse between the others.
- if (Carried.is_empty().is_true())
- for (isl::map Dep : Deps.get_map_list())
- if (getIndex(Dep, isl::dim::in) != getIndex(Dep, isl::dim::out))
- return Node;
+ if (CarriedEdges.empty() &&
+ any_of(DepEdges, [](auto Edge) { return Edge.first != Edge.second; }))
+ return Node;
// Reaches[I][J] is set if statement J depends on statement I, possibly
// through other statements. Statements that depend on each other form a
// group.
SmallVector<BitVector, 8> Reaches(NumStmts, BitVector(NumStmts));
- for (isl::map Dep : Deps.get_map_list())
- Reaches[getIndex(Dep, isl::dim::in)].set(getIndex(Dep, isl::dim::out));
+ for (auto [Src, Dst] : DepEdges)
+ Reaches[Src].set(Dst);
for (unsigned K : seq(NumStmts))
for (unsigned I : seq(NumStmts))
if (Reaches[I][K])
Reaches[I] |= Reaches[K];
// The loop of every group must be parallel.
- for (isl::map Dep : Carried.get_map_list())
- if (Reaches[getIndex(Dep, isl::dim::out)][getIndex(Dep, isl::dim::in)])
+ for (auto [Src, Dst] : CarriedEdges)
+ if (Reaches[Dst][Src])
return Node;
SmallVector<unsigned, 8> GroupOf(NumStmts);
@@ -673,7 +717,7 @@ ScheduleTreeOptimizer::distributeInnermostLoop(isl::schedule_node Node,
isl::union_set Filter = isl::union_set::empty(Node.ctx());
for (unsigned Stmt : seq(NumStmts))
if (GroupOf[Stmt] == Group)
- Filter = Filter.unite(Stmts[Stmt]);
+ Filter = Filter.unite(Stmts[Stmt].second);
Filters = Filters.add(Filter);
}
@@ -681,7 +725,7 @@ ScheduleTreeOptimizer::distributeInnermostLoop(isl::schedule_node Node,
if (Sequence.is_null())
return Node;
PointLoopDistributions++;
- return NumMembers > 1 ? Sequence.parent() : Sequence;
+ return unsigned(NumMembers) > 1 ? Sequence.parent() : Sequence;
}
isl::schedule_node
@@ -733,8 +777,9 @@ ScheduleTreeOptimizer::optimizeBand(__isl_take isl_schedule_node *NodeArg,
{
IslQuotaScope MaxScope = OAI->MaxOpGuard.enter();
Distributed = distributeInnermostLoop(Node, OAI->D);
- // Leave the band as it is if the analysis exceeds the quota, rather than
- // discarding the other optimizations of the SCoP.
+ // Leave the band as it is if the analysis exceeds the quota. Also reset
+ // the error: if no other quota scope follows, runIslScheduleOptimizer
+ // would otherwise see it and discard all optimizations of the SCoP.
if (OAI->MaxOpGuard.hasQuotaExceeded()) {
Distributed = {};
isl_ctx_reset_error(Node.ctx().get());
>From b3ac7482f63824a4ccffdc22526e617b063ef705 Mon Sep 17 00:00:00 2001
From: Timur Baidusenov <timurbaidusenov at gmail.com>
Date: Sat, 26 Sep 2026 20:28:36 +0300
Subject: [PATCH 3/3] Address review: also check the statement of a domain id
for null
The statement is dereferenced to number the statements in the order of Scop::Stmts, so leave the band unchanged if a domain id has no statement, as for the other isl results that are checked.
Assisted-by: Claude (Anthropic)
---
polly/lib/Transform/ScheduleOptimizer.cpp | 5 ++++-
1 file changed, 4 insertions(+), 1 deletion(-)
diff --git a/polly/lib/Transform/ScheduleOptimizer.cpp b/polly/lib/Transform/ScheduleOptimizer.cpp
index ce34a1d6097be..d750b2981f4c8 100644
--- a/polly/lib/Transform/ScheduleOptimizer.cpp
+++ b/polly/lib/Transform/ScheduleOptimizer.cpp
@@ -596,7 +596,10 @@ ScheduleTreeOptimizer::distributeInnermostLoop(isl::schedule_node Node,
isl::id Id = Set.get_tuple_id();
if (Id.is_null())
return Node;
- Stmts.push_back({static_cast<ScopStmt *>(Id.get_user()), Set});
+ auto *Stmt = static_cast<ScopStmt *>(Id.get_user());
+ if (!Stmt)
+ return Node;
+ Stmts.push_back({Stmt, Set});
}
unsigned NumStmts = Stmts.size();
if (NumStmts < 2)
More information about the llvm-commits
mailing list