[polly] 8d79b74 - [Polly] Keep proximity dependence exact if simplifying it unbounds its distance (#225840)
via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 29 00:09:30 PDT 2026
Author: Timur Baidusenov
Date: 2026-09-29T09:09:23+02:00
New Revision: 8d79b74443de46e78b4f2c90aa65bd42c491e909
URL: https://github.com/llvm/llvm-project/commit/8d79b74443de46e78b4f2c90aa65bd42c491e909
DIFF: https://github.com/llvm/llvm-project/commit/8d79b74443de46e78b4f2c90aa65bd42c491e909.diff
LOG: [Polly] Keep proximity dependence exact if simplifying it unbounds its distance (#225840)
Before calling the isl scheduler, `runIslScheduleOptimizer` gists the
dependences
with the iteration domains (`-polly-opt-simplify-deps=yes`, the default
since
a26db470834a, 2012). For a statement that reads a value written in one
iteration
of a loop in all later iterations, the gist also drops the upper bound
of that
loop:
{ S[k, i, k] -> S[k, i, j] : k < j <= n - 1 } becomes { S[k, i, k] ->
S[k, i, j] : j > k }
The proximity distance along j is then unbounded. isl's Pluto-like step
needs a
bound `u·n + w` on every proximity distance, finds no row that satisfies
it, and
falls back to carrying dependences one row at a time. For floyd-warshall
this
yields a wavefront over `i + j` below the `k` loop instead of a
permutable band
`(i, j)`. That band is never tiled, and the innermost loop walks the
array
diagonally.
The patch keeps the exact proximity dependence of a statement on itself
when
the gisted dependence has unbounded distances and the exact one has
bounded
distances. Validity dependences, and dependences between different
statements,
are still gisted. The cost of the exact dependences therefore only
applies to
the few maps where the gist removed the bound; heat-3d, where keeping
all
proximity dependences exact makes scheduling take minutes, is
unaffected.
### Measurements
x86-64 (Ryzen 7 9700X), `-O3 -march=native -mllvm -polly`, baseline is
the same
compiler without the change.
| | before | after | |
|---|---|---|---|
| PolyBench 4.2.1 LARGE floyd-warshall (min of 3, pinned) | 26.7 s | 6.8
s | 3.9x |
| llvm-test-suite floyd-warshall (min of 5, pinned) | 60.9 s | 16.4 s |
3.7x |
| compile time floyd-warshall.c (min of 5) | 131 ms | 134 ms | |
| .text floyd-warshall | 4775 | 6298 | |
- Among the 30 PolyBench kernels, only floyd-warshall generates
different code.
- In llvm-test-suite (SingleSource + MultiSource), only floyd-warshall
and two
small regression tests (Regression-C-test_indvars,
GCC-C-execute-20000422-1)
change. All 2137 tests pass in both configurations. The other objects
whose
code differed between the two builds (oggenc, two TSVC programs) differ
by
Polly's pre-existing run-to-run nondeterminism; the patch did not
trigger
there.
- CTMark (`-j1`, pinned, 2 repetitions): identical generated code, total
compile
time within noise.
- With the default 32x32 tiles, floyd-warshall is still 1.2x slower than
plain
-O3 on this machine. That is a tile size effect: with
`-polly-tile-sizes=16,32`
the kernel runs in 1.6 s, 3.5x faster than -O3.
### Test
`polly/test/ScheduleOptimizer/simplify_deps_bounded_proximity.ll`:
floyd-warshall
with a parametric size; checks for the `(k)` band followed by the tiled
`(i, j)`
band. It fails without the patch, which produces the `(i + j)`
wavefront.
Assisted-by: Claude (Anthropic)
Added:
polly/test/ScheduleOptimizer/simplify_deps_bounded_proximity.ll
Modified:
polly/lib/Transform/ScheduleOptimizer.cpp
Removed:
################################################################################
diff --git a/polly/lib/Transform/ScheduleOptimizer.cpp b/polly/lib/Transform/ScheduleOptimizer.cpp
index 68de330729955..eeb5116ee85ba 100644
--- a/polly/lib/Transform/ScheduleOptimizer.cpp
+++ b/polly/lib/Transform/ScheduleOptimizer.cpp
@@ -655,6 +655,48 @@ static void printSchedule(llvm::raw_ostream &OS, const isl::schedule &Schedule,
}
#endif
+/// Return whether the dependence distances of @p Map, which relates instances
+/// of the same statement, are bounded.
+static bool hasBoundedDistances(const isl::map &Map) {
+ isl::set Deltas = Map.deltas();
+ return !Deltas.is_null() && Deltas.is_bounded().is_true();
+}
+
+/// Undo the simplification of the proximity dependences of a statement on
+/// itself where it made their distances unbounded.
+///
+/// The scheduler looks for schedule rows that bound the distance of every
+/// proximity dependence. If the simplification drops the constraints of the
+/// domain that bound the distance of a dependence, such as a value that is
+/// read by all later iterations of a loop, then every row that advances along
+/// that loop has an unbounded distance, and the scheduler falls back to
+/// carrying dependences one row at a time instead of forming a permutable
+/// band.
+///
+/// @param Simplified The simplified proximity dependences.
+/// @param Exact The proximity dependences before simplification.
+static isl::union_map keepBoundedDistances(const isl::union_map &Simplified,
+ const isl::union_map &Exact) {
+ isl::union_map Result = isl::union_map::empty(Simplified.ctx());
+ for (isl::map Map : Simplified.get_map_list()) {
+ isl::space Space = Map.get_space();
+ if (Space.domain().is_equal(Space.range()) && !hasBoundedDistances(Map)) {
+ // Only add the constraints that bound the distances before the
+ // simplification rather than restoring all constraints of the exact
+ // dependence: preferably the hull of the exact distances, which is a
+ // single convex set, otherwise the exact distances themselves.
+ isl::set ExactDeltas = Exact.extract_map(Space).deltas();
+ isl::map Bounded = Map.intersect(ExactDeltas.simple_hull().translation());
+ if (!hasBoundedDistances(Bounded))
+ Bounded = Map.intersect(ExactDeltas.translation());
+ if (hasBoundedDistances(Bounded))
+ Map = Bounded;
+ }
+ Result = Result.unite(isl::union_map(Map));
+ }
+ return Result;
+}
+
/// Collect statistics for the schedule tree.
///
/// @param Schedule The schedule tree to analyze. If not a schedule tree it is
@@ -813,10 +855,12 @@ static void runIslScheduleOptimizerImpl(
// interesting anyway. In some cases this option may stop the scheduler to
// find any schedule.
if (SimplifyDeps == "yes") {
+ isl::union_map ExactProximity = Proximity;
Validity = Validity.gist_domain(Domain);
Validity = Validity.gist_range(Domain);
Proximity = Proximity.gist_domain(Domain);
Proximity = Proximity.gist_range(Domain);
+ Proximity = keepBoundedDistances(Proximity, ExactProximity);
} else if (SimplifyDeps != "no") {
errs()
<< "warning: Option -polly-opt-simplify-deps should either be 'yes' "
diff --git a/polly/test/ScheduleOptimizer/simplify_deps_bounded_proximity.ll b/polly/test/ScheduleOptimizer/simplify_deps_bounded_proximity.ll
new file mode 100644
index 0000000000000..ccf2a4e405a74
--- /dev/null
+++ b/polly/test/ScheduleOptimizer/simplify_deps_bounded_proximity.ll
@@ -0,0 +1,84 @@
+; RUN: opt %loadNPMPolly '-passes=polly-custom<opt-isl>' -polly-print-opt-isl -disable-output < %s | FileCheck %s
+;
+; Keep the exact proximity dependence of a statement on itself if simplifying
+; it makes its distances unbounded.
+;
+; void floyd(int n, int path[restrict n][n]) {
+; for (int k = 0; k < n; k++)
+; for (int i = 0; i < n; i++)
+; for (int j = 0; j < n; j++)
+; path[i][j] = path[i][j] < path[i][k] + path[k][j]
+; ? path[i][j]
+; : path[i][k] + path[k][j];
+; }
+;
+; Within an iteration of k, path[i][k] is read by all later iterations of j
+; and path[k][j] by all later iterations of i. The simplified dependences lose
+; the upper bound of j and i, so no schedule row along i or j bounds their
+; distance. The scheduler then carried the dependences with a wavefront over
+; i + j instead of forming a permutable band of i and j.
+
+; CHECK: schedule: "[n] -> [{ Stmt_for_body8[i0, i1, i2] -> [(i0)] }]"
+; CHECK-NEXT: permutable: 1
+; CHECK-NEXT: child:
+; CHECK-NEXT: mark: "1st level tiling - Tiles"
+; CHECK-NEXT: child:
+; CHECK-NEXT: schedule: "[n] -> [{ Stmt_for_body8[i0, i1, i2] -> [(floor((i1)/32))] }, { Stmt_for_body8[i0, i1, i2] -> [(floor((i2)/32))] }]"
+; CHECK-NEXT: permutable: 1
+
+target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i64:64-i128:128-f80:128-n8:16:32:64-S128"
+
+define void @floyd(i32 %n, ptr noalias %path) {
+entry:
+ %0 = zext i32 %n to i64
+ %cmp74 = icmp sgt i32 %n, 0
+ br i1 %cmp74, label %for.cond1.preheader.preheader, label %for.cond.cleanup
+
+for.cond1.preheader.preheader:
+ %wide.trip.count85 = zext nneg i32 %n to i64
+ %wide.trip.count80 = zext nneg i32 %n to i64
+ %wide.trip.count = zext nneg i32 %n to i64
+ br label %for.cond1.preheader
+
+for.cond1.preheader:
+ %indvars.iv82 = phi i64 [ 0, %for.cond1.preheader.preheader ], [ %indvars.iv.next83, %for.cond1.for.cond.cleanup3_crit_edge ]
+ %1 = mul nuw nsw i64 %indvars.iv82, %0
+ %arrayidx16 = getelementptr inbounds nuw [4 x i8], ptr %path, i64 %1
+ br label %for.cond5.preheader
+
+for.cond.cleanup:
+ ret void
+
+for.cond5.preheader:
+ %indvars.iv77 = phi i64 [ 0, %for.cond1.preheader ], [ %indvars.iv.next78, %for.cond5.for.cond.cleanup7_crit_edge ]
+ %2 = mul nuw nsw i64 %indvars.iv77, %0
+ %arrayidx = getelementptr inbounds nuw [4 x i8], ptr %path, i64 %2
+ %arrayidx14 = getelementptr inbounds nuw [4 x i8], ptr %arrayidx, i64 %indvars.iv82
+ br label %for.body8
+
+for.cond1.for.cond.cleanup3_crit_edge:
+ %indvars.iv.next83 = add nuw nsw i64 %indvars.iv82, 1
+ %exitcond86.not = icmp eq i64 %indvars.iv.next83, %wide.trip.count85
+ br i1 %exitcond86.not, label %for.cond.cleanup, label %for.cond1.preheader
+
+for.cond5.for.cond.cleanup7_crit_edge:
+ %indvars.iv.next78 = add nuw nsw i64 %indvars.iv77, 1
+ %exitcond81.not = icmp eq i64 %indvars.iv.next78, %wide.trip.count80
+ br i1 %exitcond81.not, label %for.cond1.for.cond.cleanup3_crit_edge, label %for.cond5.preheader
+
+for.body8:
+ %indvars.iv = phi i64 [ 0, %for.cond5.preheader ], [ %indvars.iv.next, %for.body8 ]
+ %arrayidx10 = getelementptr inbounds nuw [4 x i8], ptr %arrayidx, i64 %indvars.iv
+ %3 = load i32, ptr %arrayidx10, align 4
+ %4 = load i32, ptr %arrayidx14, align 4
+ %arrayidx18 = getelementptr inbounds nuw [4 x i8], ptr %arrayidx16, i64 %indvars.iv
+ %5 = load i32, ptr %arrayidx18, align 4
+ %add = add nsw i32 %5, %4
+ %.add = tail call i32 @llvm.smin.i32(i32 %3, i32 %add)
+ store i32 %.add, ptr %arrayidx10, align 4
+ %indvars.iv.next = add nuw nsw i64 %indvars.iv, 1
+ %exitcond.not = icmp eq i64 %indvars.iv.next, %wide.trip.count
+ br i1 %exitcond.not, label %for.cond5.for.cond.cleanup7_crit_edge, label %for.body8
+}
+
+declare i32 @llvm.smin.i32(i32, i32)
More information about the llvm-commits
mailing list