[polly] [Polly] Distribute the innermost point loop of tile over its statements (PR #225844)
Timur Baidusenov via llvm-commits
llvm-commits at lists.llvm.org
Wed Sep 23 09:46:11 PDT 2026
https://github.com/bai-tim created https://github.com/llvm/llvm-project/pull/225844
For time-iterated stencils (jacobi-1d, jacobi-2d, heat-3d), the isl scheduler
skews the spatial loops by the time loop and shifts the second statement of a
time step by one iteration against the first. Both statements then form a single
permutable band, which Polly tiles. In the innermost point loop, both statements
run under conditions, and the loop carries the dependence of one statement on
the other:
for (c5 ...) {
if (...) Stmt_B(...); // B[i][j] = f(A)
if (...) Stmt_A(...); // A[i][j] = f(B), reads B[i][j-1] written one iteration before
}
This loop is not vectorized. Tiling does not recover the loss: the tiled code
is as slow as the skewed schedule without tiling, and up to 2.2x slower than
plain -O3.
Within one iteration of the loops around the innermost one, the only
dependences go from `Stmt_B` to `Stmt_A`, and neither statement's own loop
carries a dependence. The patch therefore runs the innermost loop once per
statement group:
for (c5 ...) if (...) Stmt_B(...);
for (c5 ...) if (...) Stmt_A(...);
### Rule
The patch considers only the dependences between instances that share the
iterations of all loops except the innermost one (`eq_at` of the prefix
schedule):
- Statements that depend on each other stay in one group, i.e. groups are the
SCCs of the dependence graph.
- No group may carry a dependence in the innermost loop, so every resulting loop
is parallel.
- Groups are emitted in dependence order.
- A loop that is parallel already is only distributed over statements that
exchange no data within it. This keeps fusion that provides reuse; see
`tile_after_fusion.ll`, which stays fused.
- Only innermost bands are considered (child is a leaf or a sequence of leaves).
- The transformation is skipped with prevectorization or register tiling.
- If the analysis exceeds the isl quota, the band is left unchanged instead of
dropping the whole SCoP's optimization.
- `-polly-distribute-point-loops=false` disables it.
### Measurements
x86-64 (Ryzen 7 9700X), `-O3 -march=native -mllvm -polly`. The baseline is
`-polly-distribute-point-loops=false`.
PolyBench 4.2.1 LARGE, min of 3 to 21 runs, pinned:
| kernel | before | after | speedup | vs plain -O3 after |
|---|---|---|---|---|
| jacobi-1d | 0.904 ms | 0.458 ms | 1.97x | 3.7x slower (tiny kernel, data fits in L2) |
| jacobi-2d | 0.897 s | 0.412 s | 2.18x | 1.28x faster |
| heat-3d | 1.214 s | 0.552 s | 2.20x | on par |
llvm-test-suite, min of 5, pinned:
| test | before | after |
|---|---|---|
| jacobi-2d | 1.868 s | 1.065 s |
| heat-3d | 1.876 s | 0.845 s |
| jacobi-1d (whole process, 300 runs) | 3.27 ms | 2.18 ms |
Compile time and code size of the changed files (min of 5, sequential):
| file | compile before | compile after | .text before | .text after |
|---|---|---|---|---|
| jacobi-1d.c | 109 ms | 137 ms | 2099 | 3440 |
| jacobi-2d.c | 535 ms | 608 ms | 4879 | 6559 |
| heat-3d.c | 667 ms | 704 ms | 7170 | 7589 |
- Among the 30 PolyBench kernels, only these three change.
- In llvm-test-suite (SingleSource + MultiSource), the transformation triggers
only in these three programs. All 2137 tests pass in both configurations. Two
other objects differ between the builds by Polly's pre-existing
nondeterminism.
- Total .text grows by 5.7 KB, +0.025%.
- CTMark (`-j1`, pinned, 2 repetitions): identical generated code, total compile
time within noise.
- Correctness: array dumps of all 30 PolyBench kernels are bit-identical to
-O3 without Polly. This was checked on SMALL and MEDIUM, with default options,
`-polly-parallel`, 2nd-level tiling, and tile sizes 7,5,3 and 4,3,5
(300 comparisons).
### Tests
- `distribute_point_loops.ll`: jacobi-2d with parametric sizes. With the patch,
the point band ends in two loops; with `=false`, it ends in one fused loop
with both statements.
- `distribute_point_loops_cycle.ll`: two statements that depend on each other
in both directions within a row. They stay fused.
- `tile_after_fusion.ll` (existing): stays fused, because the fused loop is
parallel and the statements exchange data within an iteration.
Assisted-by: Claude (Anthropic)
>From 07848270a736dbc32b6d6d9bf33b1a51d241ecef Mon Sep 17 00:00:00 2001
From: Timur Baidusenov <timurbaidusenov at gmail.com>
Date: Wed, 23 Sep 2026 19:00:15 +0300
Subject: [PATCH] [Polly] Distribute the innermost point loop of a tile over
its statements
When the scheduler time-tiles a stencil such as jacobi-2d or heat-3d, it skews the spatial loops and shifts the statements of a time step against each other so that they form a single permutable band; after tiling, the innermost point loop runs all statements under conditions and carries the dependences between them, so it is not vectorized and the tiled code is up to 2.2 times slower than plain -O3. This patch distributes the innermost loop of a tiled innermost band over its statements when that is legal and makes every resulting loop parallel: considering only the dependences between instances that share the iterations of the outer loops, statements that depend on each other stay in one group, no group may carry a dependence in the innermost loop, and the groups are emitted in dependence order, while a loop that is parallel already is only distributed over statements that do not exchange data within it, which keeps fusion that provides reuse; the transformation is skipped with prevectorization or register tiling, leaves the band unchanged if its analysis exceeds the isl operation quota, and can be disabled with -polly-distribute-point-loops=false. On PolyBench/C 4.2.1 LARGE (x86-64, -O3 -march=native -polly) jacobi-1d, jacobi-2d and heat-3d become 1.97, 2.18 and 2.20 times faster, which puts jacobi-2d 22% ahead of plain -O3, and no other kernel changes; in llvm-test-suite only these three programs change (jacobi-2d 1.87 s to 1.07 s, heat-3d 1.88 s to 0.85 s), all tests pass, CTMark generates identical code, and the compile time of the three files grows by 28 to 73 ms.
Assisted-by: Claude (Anthropic)
---
polly/lib/Transform/ScheduleOptimizer.cpp | 155 +++++++++++++++++
.../distribute_point_loops.ll | 162 ++++++++++++++++++
.../distribute_point_loops_cycle.ll | 70 ++++++++
3 files changed, 387 insertions(+)
create mode 100644 polly/test/ScheduleOptimizer/distribute_point_loops.ll
create mode 100644 polly/test/ScheduleOptimizer/distribute_point_loops_cycle.ll
diff --git a/polly/lib/Transform/ScheduleOptimizer.cpp b/polly/lib/Transform/ScheduleOptimizer.cpp
index b9b9abbd85ae41..c1611be5c296ca 100644
--- a/polly/lib/Transform/ScheduleOptimizer.cpp
+++ b/polly/lib/Transform/ScheduleOptimizer.cpp
@@ -55,6 +55,7 @@
#include "polly/ScopInfo.h"
#include "polly/Support/ISLOStream.h"
#include "polly/Support/ISLTools.h"
+#include "llvm/ADT/BitVector.h"
#include "llvm/ADT/Sequence.h"
#include "llvm/ADT/Statistic.h"
#include "llvm/Analysis/OptimizationRemarkEmitter.h"
@@ -159,6 +160,12 @@ static cl::opt<bool> RegisterTiling("polly-register-tiling",
cl::desc("Enable register tiling"),
cl::cat(PollyCategory));
+static cl::opt<bool> DistributePointLoops(
+ "polly-distribute-point-loops",
+ cl::desc("Distribute the innermost point loop of a tile over its "
+ "statements if this makes the loop of every statement parallel"),
+ cl::init(true), cl::cat(PollyCategory));
+
static cl::opt<int> RegisterDefaultTileSize(
"polly-register-tiling-default-tile-size",
cl::desc("The default register tile size (if not enough were provided by"
@@ -226,6 +233,8 @@ STATISTIC(FirstLevelTileOpts, "Number of first level tiling applied");
STATISTIC(SecondLevelTileOpts, "Number of second level tiling applied");
STATISTIC(RegisterTileOpts, "Number of register tiling applied");
STATISTIC(PrevectOpts, "Number of strip-mining for prevectorization applied");
+STATISTIC(PointLoopDistributions,
+ "Number of innermost point loops distributed over statements");
STATISTIC(MatMulOpts,
"Number of matrix multiplication patterns detected and optimized");
@@ -377,6 +386,25 @@ class ScheduleTreeOptimizer final {
/// @param Node The schedule node to (possibly) optimize.
static isl::schedule_node applyTileBandOpt(isl::schedule_node Node);
+ /// Distribute the innermost loop of a band over its statements.
+ ///
+ /// Time tiling of a stencil shifts the statements of a time step against
+ /// each other and fuses them in the innermost point loop. The loop then runs
+ /// each statement under a condition and carries the dependences between
+ /// them, which keeps it from being vectorized. Run the innermost loop once
+ /// for each group of statements instead, where statements that depend on
+ /// each other form a group, if the groups can be ordered and none of them
+ /// carries a dependence in the innermost loop. A loop that is parallel
+ /// already is only distributed over statements that do not exchange any
+ /// data within it.
+ ///
+ /// @param Node The innermost band to distribute, typically the point band of
+ /// a tiling.
+ /// @param D The dependences of the SCoP.
+ /// @return The node at the position of @p Node.
+ static isl::schedule_node distributeInnermostLoop(isl::schedule_node Node,
+ const Dependences *D);
+
/// Apply prevectorization on the bands in the schedule tree.
///
/// @param Node The schedule node to (possibly) prevectorize.
@@ -548,6 +576,114 @@ ScheduleTreeOptimizer::applyTileBandOpt(isl::schedule_node Node) {
return Node;
}
+isl::schedule_node
+ScheduleTreeOptimizer::distributeInnermostLoop(isl::schedule_node Node,
+ const Dependences *D) {
+ if (!isSimpleInnermostBand(Node))
+ return Node;
+
+ // Number the statements in the order of their names, so that the result does
+ // not depend on the order in which isl lists them.
+ isl::union_set Domain = Node.get_domain();
+ SmallVector<isl::set, 8> Stmts;
+ for (isl::set Stmt : Domain.get_set_list())
+ Stmts.push_back(Stmt);
+ unsigned NumStmts = Stmts.size();
+ if (NumStmts < 2)
+ return Node;
+ llvm::sort(Stmts, [](const isl::set &A, const isl::set &B) {
+ return A.get_tuple_name() < B.get_tuple_name();
+ });
+ DenseMap<isl_id *, unsigned> StmtIndex;
+ for (auto [Idx, Stmt] : enumerate(Stmts))
+ StmtIndex[Stmt.get_tuple_id().get()] = Idx;
+
+ isl::schedule_node_band Band = Node.as<isl::schedule_node_band>();
+ unsigned NumMembers = unsignedFromIslSize(Band.n_member());
+ isl::schedule_node_band Inner = Band;
+ if (NumMembers > 1)
+ Inner = Band.split(NumMembers - 1).child(0).as<isl::schedule_node_band>();
+
+ // The dependences between the instances that share the iterations of all
+ // loops around the innermost one, split into those that the innermost loop
+ // carries and those within one of its iterations.
+ isl::union_map Deps = D->getDependences(
+ Dependences::TYPE_RAW | Dependences::TYPE_WAR | Dependences::TYPE_WAW);
+ Deps = Deps.intersect_domain(Domain).intersect_range(Domain);
+ Deps = Deps.eq_at(Inner.get_prefix_schedule_multi_union_pw_aff());
+ isl::union_map SameIteration = Deps.eq_at(Inner.get_partial_schedule());
+ isl::union_map Carried = Deps.subtract(SameIteration);
+ if (Carried.is_null())
+ return Node;
+
+ auto getIndex = [&](const isl::map &Dep, isl::dim Dim) {
+ return StmtIndex.lookup(Dep.get_tuple_id(Dim).get());
+ };
+
+ // If the fused loop is parallel already, only separate statements that do
+ // not exchange any data within it, to keep the reuse between the others.
+ if (Carried.is_empty().is_true())
+ for (isl::map Dep : Deps.get_map_list())
+ if (getIndex(Dep, isl::dim::in) != getIndex(Dep, isl::dim::out))
+ return Node;
+
+ // Reaches[I][J] is set if statement J depends on statement I, possibly
+ // through other statements. Statements that depend on each other form a
+ // group.
+ SmallVector<BitVector, 8> Reaches(NumStmts, BitVector(NumStmts));
+ for (isl::map Dep : Deps.get_map_list())
+ Reaches[getIndex(Dep, isl::dim::in)].set(getIndex(Dep, isl::dim::out));
+ for (unsigned K : seq(NumStmts))
+ for (unsigned I : seq(NumStmts))
+ if (Reaches[I][K])
+ Reaches[I] |= Reaches[K];
+
+ // The loop of every group must be parallel.
+ for (isl::map Dep : Carried.get_map_list())
+ if (Reaches[getIndex(Dep, isl::dim::out)][getIndex(Dep, isl::dim::in)])
+ return Node;
+
+ SmallVector<unsigned, 8> GroupOf(NumStmts);
+ unsigned NumGroups = 0;
+ for (unsigned I : seq(NumStmts)) {
+ auto Earlier = find_if(
+ seq(I), [&](unsigned J) { return Reaches[I][J] && Reaches[J][I]; });
+ GroupOf[I] = Earlier != seq(I).end() ? GroupOf[*Earlier] : NumGroups++;
+ }
+ if (NumGroups < 2)
+ return Node;
+
+ // A group depends on more statements than every group it depends on, so
+ // ordering by this number respects the dependences.
+ SmallVector<unsigned, 8> Leader(NumGroups);
+ for (unsigned Stmt : reverse(seq(NumStmts)))
+ Leader[GroupOf[Stmt]] = Stmt;
+ auto NumPredecessors = [&](unsigned Group) {
+ return count_if(seq(NumStmts), [&](unsigned Pred) {
+ return GroupOf[Pred] != Group && Reaches[Pred][Leader[Group]];
+ });
+ };
+ SmallVector<unsigned, 8> Order = to_vector<8>(seq(NumGroups));
+ stable_sort(Order, [&](unsigned A, unsigned B) {
+ return NumPredecessors(A) < NumPredecessors(B);
+ });
+
+ isl::union_set_list Filters(Node.ctx(), NumGroups);
+ for (unsigned Group : Order) {
+ isl::union_set Filter = isl::union_set::empty(Node.ctx());
+ for (unsigned Stmt : seq(NumStmts))
+ if (GroupOf[Stmt] == Group)
+ Filter = Filter.unite(Stmts[Stmt]);
+ Filters = Filters.add(Filter);
+ }
+
+ isl::schedule_node Sequence = Inner.insert_sequence(Filters);
+ if (Sequence.is_null())
+ return Node;
+ PointLoopDistributions++;
+ return NumMembers > 1 ? Sequence.parent() : Sequence;
+}
+
isl::schedule_node
ScheduleTreeOptimizer::applyPrevectBandOpt(isl::schedule_node Node) {
auto Space = isl::manage(isl_schedule_node_band_get_space(Node.get()));
@@ -589,6 +725,25 @@ ScheduleTreeOptimizer::optimizeBand(__isl_take isl_schedule_node *NodeArg,
if (OAI->Postopts)
Node = applyTileBandOpt(Node);
+ // Prevectorization strip-mines a parallel point loop and register tiling
+ // unrolls the point loops, leave those to them.
+ if (OAI->Postopts && DistributePointLoops && !OAI->Prevect &&
+ !RegisterTiling) {
+ isl::schedule_node Distributed;
+ {
+ IslQuotaScope MaxScope = OAI->MaxOpGuard.enter();
+ Distributed = distributeInnermostLoop(Node, OAI->D);
+ // Leave the band as it is if the analysis exceeds the quota, rather than
+ // discarding the other optimizations of the SCoP.
+ if (OAI->MaxOpGuard.hasQuotaExceeded()) {
+ Distributed = {};
+ isl_ctx_reset_error(Node.ctx().get());
+ }
+ }
+ if (!Distributed.is_null())
+ Node = Distributed;
+ }
+
if (OAI->Prevect) {
IslQuotaScope MaxScope = OAI->MaxOpGuard.enter();
diff --git a/polly/test/ScheduleOptimizer/distribute_point_loops.ll b/polly/test/ScheduleOptimizer/distribute_point_loops.ll
new file mode 100644
index 00000000000000..bd6ca33eaedd1c
--- /dev/null
+++ b/polly/test/ScheduleOptimizer/distribute_point_loops.ll
@@ -0,0 +1,162 @@
+; RUN: opt %loadNPMPolly '-passes=polly-custom<opt-isl;ast>' -polly-print-ast -disable-output < %s | FileCheck %s
+; RUN: opt %loadNPMPolly '-passes=polly-custom<opt-isl;ast>' -polly-print-ast -polly-distribute-point-loops=false -disable-output < %s | FileCheck %s --check-prefix=FUSED
+;
+; Distribute the innermost point loop of a time tiled stencil over its
+; statements.
+;
+; void jacobi(int tsteps, int n, double A[restrict n][n],
+; double B[restrict n][n]) {
+; for (int t = 0; t < tsteps; t++) {
+; for (int i = 1; i < n - 1; i++)
+; for (int j = 1; j < n - 1; j++)
+; B[i][j] = 0.2 * (A[i][j] + A[i][j - 1] + A[i][j + 1] +
+; A[i + 1][j] + A[i - 1][j]);
+; for (int i = 1; i < n - 1; i++)
+; for (int j = 1; j < n - 1; j++)
+; A[i][j] = 0.2 * (B[i][j] + B[i][j - 1] + B[i][j + 1] +
+; B[i + 1][j] + B[i - 1][j]);
+; }
+; }
+;
+; The scheduler skews the loops over i and j by 2t and shifts the second
+; statement by one iteration against the first, so that both form a single
+; permutable band. The innermost point loop then runs both statements under
+; conditions and carries the dependence of the second statement on the first.
+; Within an iteration of the loops around it, however, only the second
+; statement depends on the first, and neither loop carries a dependence on
+; its own, so the loop is run once for each statement.
+
+; CHECK: // 1st level tiling - Points
+; CHECK-NEXT: for (int c3 = {{.*}})
+; CHECK-NEXT: for (int c4 = {{.*}}) {
+; CHECK-NEXT: if ({{.*}})
+; CHECK-NEXT: for (int c5 = {{.*}})
+; CHECK-NEXT: Stmt_for_body9(
+; CHECK-NEXT: if ({{.*}})
+; CHECK-NEXT: for (int c5 = {{.*}})
+; CHECK-NEXT: Stmt_for_body53(
+; CHECK-NEXT: }
+
+; FUSED: // 1st level tiling - Points
+; FUSED-NEXT: for (int c3 = {{.*}})
+; FUSED-NEXT: for (int c4 = {{.*}})
+; FUSED-NEXT: for (int c5 = {{.*}}) {
+; FUSED-NEXT: if ({{.*}})
+; FUSED-NEXT: Stmt_for_body9(
+; FUSED-NEXT: if ({{.*}})
+; FUSED-NEXT: Stmt_for_body53(
+; FUSED-NEXT: }
+
+target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i64:64-i128:128-f80:128-n8:16:32:64-S128"
+
+define void @jacobi(i32 %tsteps, i32 %n, ptr noalias %A, ptr noalias %B) {
+entry:
+ %0 = zext i32 %n to i64
+ %cmp150 = icmp sgt i32 %tsteps, 0
+ br i1 %cmp150, label %for.cond1.preheader.lr.ph, label %for.cond.cleanup
+
+for.cond1.preheader.lr.ph:
+ %sub = add i32 %n, -1
+ %cmp2144 = icmp sgt i32 %n, 2
+ %cmp45148 = icmp sgt i32 %n, 2
+ %wide.trip.count157 = zext i32 %sub to i64
+ %wide.trip.count = zext i32 %sub to i64
+ %wide.trip.count168 = zext i32 %sub to i64
+ %wide.trip.count162 = zext i32 %sub to i64
+ br label %for.cond1.preheader
+
+for.cond1.preheader:
+ %t.0151 = phi i32 [ 0, %for.cond1.preheader.lr.ph ], [ %inc94, %for.cond.cleanup46 ]
+ br i1 %cmp2144, label %for.cond5.preheader, label %for.cond43.preheader
+
+for.cond.cleanup:
+ ret void
+
+for.cond43.preheader:
+ br i1 %cmp45148, label %for.cond49.preheader, label %for.cond.cleanup46
+
+for.cond5.preheader:
+ %indvars.iv153 = phi i64 [ %indvars.iv.next154, %for.cond5.for.cond.cleanup8_crit_edge ], [ 1, %for.cond1.preheader ]
+ %1 = mul nuw nsw i64 %indvars.iv153, %0
+ %arrayidx = getelementptr inbounds nuw [8 x i8], ptr %A, i64 %1
+ %indvars.iv.next154 = add nuw nsw i64 %indvars.iv153, 1
+ %2 = mul nuw nsw i64 %indvars.iv.next154, %0
+ %arrayidx25 = getelementptr inbounds nuw [8 x i8], ptr %A, i64 %2
+ %3 = add nsw i64 %indvars.iv153, -1
+ %4 = mul nuw nsw i64 %3, %0
+ %arrayidx31 = getelementptr inbounds [8 x i8], ptr %A, i64 %4
+ %arrayidx36 = getelementptr inbounds nuw [8 x i8], ptr %B, i64 %1
+ br label %for.body9
+
+for.cond5.for.cond.cleanup8_crit_edge:
+ %exitcond158.not = icmp eq i64 %indvars.iv.next154, %wide.trip.count157
+ br i1 %exitcond158.not, label %for.cond43.preheader, label %for.cond5.preheader
+
+for.body9:
+ %indvars.iv = phi i64 [ 1, %for.cond5.preheader ], [ %indvars.iv.next, %for.body9 ]
+ %arrayidx11 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx, i64 %indvars.iv
+ %5 = load double, ptr %arrayidx11, align 8
+ %arrayidx16 = getelementptr i8, ptr %arrayidx11, i64 -8
+ %6 = load double, ptr %arrayidx16, align 8
+ %add = fadd double %5, %6
+ %indvars.iv.next = add nuw nsw i64 %indvars.iv, 1
+ %arrayidx21 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx, i64 %indvars.iv.next
+ %7 = load double, ptr %arrayidx21, align 8
+ %add22 = fadd double %add, %7
+ %arrayidx27 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx25, i64 %indvars.iv
+ %8 = load double, ptr %arrayidx27, align 8
+ %add28 = fadd double %add22, %8
+ %arrayidx33 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx31, i64 %indvars.iv
+ %9 = load double, ptr %arrayidx33, align 8
+ %add34 = fadd double %add28, %9
+ %mul = fmul double %add34, 2.000000e-01
+ %arrayidx38 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx36, i64 %indvars.iv
+ store double %mul, ptr %arrayidx38, align 8
+ %exitcond.not = icmp eq i64 %indvars.iv.next, %wide.trip.count
+ br i1 %exitcond.not, label %for.cond5.for.cond.cleanup8_crit_edge, label %for.body9
+
+for.cond49.preheader:
+ %indvars.iv164 = phi i64 [ %indvars.iv.next165, %for.cond49.for.cond.cleanup52_crit_edge ], [ 1, %for.cond43.preheader ]
+ %10 = mul nuw nsw i64 %indvars.iv164, %0
+ %arrayidx55 = getelementptr inbounds nuw [8 x i8], ptr %B, i64 %10
+ %indvars.iv.next165 = add nuw nsw i64 %indvars.iv164, 1
+ %11 = mul nuw nsw i64 %indvars.iv.next165, %0
+ %arrayidx72 = getelementptr inbounds nuw [8 x i8], ptr %B, i64 %11
+ %12 = add nsw i64 %indvars.iv164, -1
+ %13 = mul nuw nsw i64 %12, %0
+ %arrayidx78 = getelementptr inbounds [8 x i8], ptr %B, i64 %13
+ %arrayidx84 = getelementptr inbounds nuw [8 x i8], ptr %A, i64 %10
+ br label %for.body53
+
+for.cond.cleanup46:
+ %inc94 = add nuw nsw i32 %t.0151, 1
+ %exitcond170.not = icmp eq i32 %inc94, %tsteps
+ br i1 %exitcond170.not, label %for.cond.cleanup, label %for.cond1.preheader
+
+for.cond49.for.cond.cleanup52_crit_edge:
+ %exitcond169.not = icmp eq i64 %indvars.iv.next165, %wide.trip.count168
+ br i1 %exitcond169.not, label %for.cond.cleanup46, label %for.cond49.preheader
+
+for.body53:
+ %indvars.iv159 = phi i64 [ 1, %for.cond49.preheader ], [ %indvars.iv.next160, %for.body53 ]
+ %arrayidx57 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx55, i64 %indvars.iv159
+ %14 = load double, ptr %arrayidx57, align 8
+ %arrayidx62 = getelementptr i8, ptr %arrayidx57, i64 -8
+ %15 = load double, ptr %arrayidx62, align 8
+ %add63 = fadd double %14, %15
+ %indvars.iv.next160 = add nuw nsw i64 %indvars.iv159, 1
+ %arrayidx68 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx55, i64 %indvars.iv.next160
+ %16 = load double, ptr %arrayidx68, align 8
+ %add69 = fadd double %add63, %16
+ %arrayidx74 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx72, i64 %indvars.iv159
+ %17 = load double, ptr %arrayidx74, align 8
+ %add75 = fadd double %add69, %17
+ %arrayidx80 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx78, i64 %indvars.iv159
+ %18 = load double, ptr %arrayidx80, align 8
+ %add81 = fadd double %add75, %18
+ %mul82 = fmul double %add81, 2.000000e-01
+ %arrayidx86 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx84, i64 %indvars.iv159
+ store double %mul82, ptr %arrayidx86, align 8
+ %exitcond163.not = icmp eq i64 %indvars.iv.next160, %wide.trip.count162
+ br i1 %exitcond163.not, label %for.cond49.for.cond.cleanup52_crit_edge, label %for.body53
+}
diff --git a/polly/test/ScheduleOptimizer/distribute_point_loops_cycle.ll b/polly/test/ScheduleOptimizer/distribute_point_loops_cycle.ll
new file mode 100644
index 00000000000000..964ed275ac2876
--- /dev/null
+++ b/polly/test/ScheduleOptimizer/distribute_point_loops_cycle.ll
@@ -0,0 +1,70 @@
+; RUN: opt %loadNPMPolly -polly-stmt-granularity=store '-passes=polly-custom<simplify-0;optree;delicm;simplify-1;opt-isl;ast>' -polly-print-ast -disable-output < %s | FileCheck %s
+;
+; Do not distribute the innermost point loop if the statements in it depend on
+; each other in both directions.
+;
+; void cyc(int n, double A[restrict n][n], double B[restrict n][n]) {
+; for (int i = 1; i < n; i++)
+; for (int j = 1; j < n; j++) {
+; A[i][j] = B[i][j - 1] + A[i - 1][j];
+; B[i][j] = 2 * A[i][j];
+; }
+; }
+;
+; The second statement depends on the first within an iteration of j, and the
+; first on the second from the previous iteration of j.
+
+; CHECK: // 1st level tiling - Points
+; CHECK-NEXT: for (int c2 = {{.*}})
+; CHECK-NEXT: for (int c3 = {{.*}}) {
+; CHECK-NEXT: Stmt_for_body4(
+; CHECK-NEXT: Stmt_for_body4_b(
+; CHECK-NEXT: }
+
+target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i64:64-i128:128-f80:128-n8:16:32:64-S128"
+
+define void @cyc(i32 %n, ptr noalias %A, ptr noalias %B) {
+entry:
+ %0 = zext i32 %n to i64
+ %cmp49 = icmp sgt i32 %n, 1
+ br i1 %cmp49, label %for.cond1.preheader.preheader, label %for.cond.cleanup
+
+for.cond1.preheader.preheader:
+ %wide.trip.count56 = zext nneg i32 %n to i64
+ %wide.trip.count = zext nneg i32 %n to i64
+ br label %for.cond1.preheader
+
+for.cond1.preheader:
+ %indvars.iv52 = phi i64 [ 1, %for.cond1.preheader.preheader ], [ %indvars.iv.next53, %for.cond1.for.cond.cleanup3_crit_edge ]
+ %1 = mul nuw nsw i64 %indvars.iv52, %0
+ %arrayidx = getelementptr inbounds nuw [8 x i8], ptr %B, i64 %1
+ %2 = add nsw i64 %indvars.iv52, -1
+ %3 = mul nuw nsw i64 %2, %0
+ %arrayidx9 = getelementptr inbounds [8 x i8], ptr %A, i64 %3
+ %arrayidx13 = getelementptr inbounds nuw [8 x i8], ptr %A, i64 %1
+ %load_initial = load double, ptr %arrayidx, align 8
+ br label %for.body4
+
+for.cond.cleanup:
+ ret void
+
+for.cond1.for.cond.cleanup3_crit_edge:
+ %indvars.iv.next53 = add nuw nsw i64 %indvars.iv52, 1
+ %exitcond57.not = icmp eq i64 %indvars.iv.next53, %wide.trip.count56
+ br i1 %exitcond57.not, label %for.cond.cleanup, label %for.cond1.preheader
+
+for.body4:
+ %store_forwarded = phi double [ %load_initial, %for.cond1.preheader ], [ %mul, %for.body4 ]
+ %indvars.iv = phi i64 [ 1, %for.cond1.preheader ], [ %indvars.iv.next, %for.body4 ]
+ %4 = getelementptr [8 x i8], ptr %arrayidx, i64 %indvars.iv
+ %arrayidx11 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx9, i64 %indvars.iv
+ %5 = load double, ptr %arrayidx11, align 8
+ %add = fadd double %store_forwarded, %5
+ %arrayidx15 = getelementptr inbounds nuw [8 x i8], ptr %arrayidx13, i64 %indvars.iv
+ store double %add, ptr %arrayidx15, align 8
+ %mul = fmul double %add, 2.000000e+00
+ store double %mul, ptr %4, align 8
+ %indvars.iv.next = add nuw nsw i64 %indvars.iv, 1
+ %exitcond.not = icmp eq i64 %indvars.iv.next, %wide.trip.count
+ br i1 %exitcond.not, label %for.cond1.for.cond.cleanup3_crit_edge, label %for.body4
+}
More information about the llvm-commits
mailing list