[polly] [Polly] Isolate complete tiles from partial tiles in the tiling path (PR #221087)

Timur Baidusenov via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 22 06:22:25 PDT 2026


https://github.com/bai-tim updated https://github.com/llvm/llvm-project/pull/221087

>From 12dab69c125801673c37e811a54d89b7002843f6 Mon Sep 17 00:00:00 2001
From: Timur Baidusenov <timurbaidusenov at gmail.com>
Date: Fri, 4 Sep 2026 01:58:40 +0300
Subject: [PATCH 1/3] [Polly] Isolate full tiles from partial tiles in the
 tiling path

Tiling an iteration space whose extent is not a multiple of the tile size
gives every point loop a min() upper bound. The loop vectorizer cannot commit
to a vector factor for such a loop without a runtime check, and it has to emit
a scalar epilogue inside the tile loop, so the loops that carry the whole
runtime of a tiled kernel are the ones it handles worst.

Isolating the full tiles from the partial ones gives the point loops of the
full tiles constant bounds. Polly already does this for the loop it strip-mines
in the prevectorization path, added in commit ca7f5bb7674f ("Full/partial tile
separation for vectorization"); this applies the same technique to the loops
produced by the regular tiling. getPartialTilePrefixes is generalized to
several point dimensions as getFullTilePrefixes, and the resulting isolate
option, combined with atomic for the remaining part, is attached to the tile
band.

Two options are added, both leaving the default behaviour unchanged:

 - -polly-isolate-full-tiles enables the separation. It is off by default
   because it grows the generated code considerably.
 - -polly-isolate-complete-tile-dims limits how many innermost dimensions of a
   tile have to be complete. Requiring fewer of them bounds the number of
   copies of the loop nest, at the price of leaving the outer point loops with
   a min() bound.

On PolyBench/C 4.2.1 with -O3 -march=native -mllvm -polly -mllvm
-polly-vectorizer=none on an AMD Ryzen 7 9700X, minimum of 5 to 41 runs per
kernel pinned to a single core, the separation changes the code of 23 of the 30
kernels and improves their geometric mean by 37.5%:

  2mm 0.146, 3mm 0.157, gemm 0.159, covariance 0.241, syrk 0.384,
  jacobi-1d 0.470, mvt 0.668, bicg 0.672, atax 0.673, gemver 0.691,
  jacobi-2d 0.744, symm 0.805, lu 0.844

Kernels whose code the option does not change give the noise floor:
floyd-warshall 0.993, nussinov 1.001, trisolv 0.991.

Two kernels regress: gesummv by 16% and deriche by 13%. In both the vectorizer
keeps the same vector factor and only gets more copies of the loop, so the
added code buys nothing: gesummv streams 27 MB at 24 GB/s against a measured
single-core peak of 29.8 GB/s, and deriche is dominated by eleven loops that
carry a dependence and never vectorize. Restricting the isolation to the
innermost dimension does not help them either (1.178 and 1.125), while it costs
performance elsewhere (geometric mean 0.698 instead of 0.625, with syrk falling
back from 0.384 to 0.964), which is why completeness in every dimension is the
default.

The cost is code size and compile time: .text of the benchmark binaries grows
by 81% to 384%, and compiling a single kernel takes 31% to 96% longer, which
mirrors the compile-time regression reported in ca7f5bb7674f. Requiring only
the innermost dimension to be complete keeps the growth lower, for example
+140% instead of +299% for gemm and +94% instead of +260% for syrk.

Numerical output is unchanged: all 30 kernels produce byte-identical dumps with
and without the option, in both modes, on the MINI and MEDIUM datasets, as do
the eight largest winners on LARGE.

Assisted-by: Claude Opus 5
---
 polly/include/polly/ScheduleTreeTransform.h   |  18 +++
 polly/lib/Transform/ScheduleOptimizer.cpp     |  78 ++++++++++++
 polly/lib/Transform/ScheduleTreeTransform.cpp |  59 ++++++++++
 .../ScheduleOptimizer/isolate-full-tiles.ll   | 111 ++++++++++++++++++
 4 files changed, 266 insertions(+)
 create mode 100644 polly/test/ScheduleOptimizer/isolate-full-tiles.ll

diff --git a/polly/include/polly/ScheduleTreeTransform.h b/polly/include/polly/ScheduleTreeTransform.h
index 6bd5a3abf9ea2..cd2cbbe244203 100644
--- a/polly/include/polly/ScheduleTreeTransform.h
+++ b/polly/include/polly/ScheduleTreeTransform.h
@@ -234,6 +234,24 @@ isl::schedule applyMaxFission(isl::schedule_node BandToFission);
 ///                      relation.
 isl::set getPartialTilePrefixes(isl::set ScheduleRange, int VectorWidth);
 
+/// Compute the prefixes of the complete tiles of a tiled band.
+///
+/// A tile is complete if every point of it belongs to @p ScheduleRange; those
+/// are the tiles whose point loops can be given constant bounds by isolating
+/// them from the partial tiles at the boundary of the iteration space.
+///
+/// @param ScheduleRange   A range of a map, which describes a prefix schedule
+///                        relation whose last @p TileSizes.size() dimensions
+///                        are the point dimensions of the tiling.
+/// @param TileSizes       The tile size of each tiled dimension.
+/// @param NumCompleteDims The number of innermost point dimensions that have to
+///                        be complete. Requiring fewer than all of them
+///                        isolates more tiles, but bounds the number of copies
+///                        of the loop nest that the AST generator creates.
+isl::set getFullTilePrefixes(isl::set ScheduleRange,
+                             llvm::ArrayRef<int> TileSizes,
+                             unsigned NumCompleteDims);
+
 /// Create an isl::union_set, which describes the isolate option based on
 /// IsolateDomain.
 ///
diff --git a/polly/lib/Transform/ScheduleOptimizer.cpp b/polly/lib/Transform/ScheduleOptimizer.cpp
index b9b9abbd85ae4..978855a752b69 100644
--- a/polly/lib/Transform/ScheduleOptimizer.cpp
+++ b/polly/lib/Transform/ScheduleOptimizer.cpp
@@ -171,6 +171,20 @@ static cl::list<int>
                                "with --polly-register-tile-size"),
                       cl::Hidden, cl::CommaSeparated, cl::cat(PollyCategory));
 
+static cl::opt<bool> IsolateFullTiles(
+    "polly-isolate-full-tiles",
+    cl::desc("Separate the full tiles of a tiled band from the partial ones, "
+             "so that the point loops of the full tiles have constant bounds"),
+    cl::Hidden, cl::init(false), cl::cat(PollyCategory));
+
+static cl::opt<int> IsolateCompleteTileDims(
+    "polly-isolate-complete-tile-dims",
+    cl::desc("Number of innermost tile dimensions that have to be complete for "
+             "a tile to be isolated (0: all of them). Requiring fewer of them "
+             "generates less code, but also gives fewer point loops a constant "
+             "bound"),
+    cl::Hidden, cl::init(0), cl::cat(PollyCategory));
+
 static cl::opt<bool> PragmaBasedOpts(
     "polly-pragma-based-opts",
     cl::desc("Apply user-directed transformation from metadata"),
@@ -405,6 +419,56 @@ ScheduleTreeOptimizer::isolateFullPartialTiles(isl::schedule_node Node,
   return Result;
 }
 
+/// Separate the full tiles of a tiled band from the partial ones.
+///
+/// The point loops of a full tile run over the whole tile, so isolating those
+/// tiles gives them constant loop bounds instead of the min() expressions that
+/// a tiling of an iteration space which is not a multiple of the tile size
+/// produces. The partial tiles are left to a single atomic copy of the loop
+/// nest to keep the code growth bounded.
+///
+/// @param Node      The point band of the tiling, as returned by tileNode.
+/// @param TileSizes The tile size of each tiled dimension.
+/// @return          The point band of the modified tree.
+static isl::schedule_node isolateFullTiles(isl::schedule_node Node,
+                                           ArrayRef<int> TileSizes) {
+  assert(isl_schedule_node_get_type(Node.get()) == isl_schedule_node_band);
+
+  // Below the point band, the prefix schedule covers the outer dimensions
+  // followed by the tile and the point dimensions of this tiling.
+  isl::union_set ScheduleRangeUSet =
+      Node.child(0).get_prefix_schedule_relation().range();
+  isl::set ScheduleRange{ScheduleRangeUSet};
+  if (ScheduleRange.is_null())
+    return Node;
+
+  unsigned NumCompleteDims = TileSizes.size();
+  if (IsolateCompleteTileDims > 0)
+    NumCompleteDims =
+        std::min<unsigned>(IsolateCompleteTileDims, TileSizes.size());
+
+  isl::set FullTilePrefixes =
+      getFullTilePrefixes(ScheduleRange, TileSizes, NumCompleteDims);
+  if (FullTilePrefixes.is_null())
+    return Node;
+
+  isl::union_set Options = getIsolateOptions(FullTilePrefixes, TileSizes.size())
+                               .unite(getDimOptions(Node.ctx(), "atomic"));
+
+  // The option describes the tile dimensions, so it belongs to the tile band,
+  // which sits above the marker separating it from the point band.
+  isl::schedule_node TileBand = Node.parent().parent();
+  if (!TileBand.isa<isl::schedule_node_band>())
+    return Node;
+
+  TileBand =
+      TileBand.as<isl::schedule_node_band>().set_ast_build_options(Options);
+  if (TileBand.is_null())
+    return Node;
+
+  return TileBand.child(0).child(0);
+}
+
 struct InsertSimdMarkers final : ScheduleNodeRewriter<InsertSimdMarkers> {
   isl::schedule_node visitBand(isl::schedule_node_band Band) {
     isl::schedule_node Node = visitChildren(Band);
@@ -528,9 +592,23 @@ bool ScheduleTreeOptimizer::isPMOptimizableBandNode(isl::schedule_node Node) {
 __isl_give isl::schedule_node
 ScheduleTreeOptimizer::applyTileBandOpt(isl::schedule_node Node) {
   if (FirstLevelTiling) {
+    // Resolve the tile size of every dimension before tiling splits the band.
+    SmallVector<int, 4> Sizes;
+    if (IsolateFullTiles) {
+      isl::space Space =
+          isl::manage(isl_schedule_node_band_get_space(Node.get()));
+      for (unsigned i : rangeIslSize(0, Space.dim(isl::dim::set)))
+        Sizes.push_back(i < FirstLevelTileSizes.size()
+                            ? FirstLevelTileSizes[i]
+                            : FirstLevelDefaultTileSize.getValue());
+    }
+
     Node = tileNode(Node, "1st level tiling", FirstLevelTileSizes,
                     FirstLevelDefaultTileSize);
     FirstLevelTileOpts++;
+
+    if (IsolateFullTiles)
+      Node = isolateFullTiles(Node, Sizes);
   }
 
   if (SecondLevelTiling) {
diff --git a/polly/lib/Transform/ScheduleTreeTransform.cpp b/polly/lib/Transform/ScheduleTreeTransform.cpp
index c95c55858f038..ebbcf1bb2c050 100644
--- a/polly/lib/Transform/ScheduleTreeTransform.cpp
+++ b/polly/lib/Transform/ScheduleTreeTransform.cpp
@@ -1127,6 +1127,65 @@ isl::set polly::getPartialTilePrefixes(isl::set ScheduleRange,
   return LoopPrefixes.subtract(BadPrefixes);
 }
 
+/// Restrict the last @p Sizes.size() dimensions of @p Set to a complete tile,
+/// that is, to 0 <= dim <= size - 1 for each of them.
+static isl::set addTileExtentConstraints(isl::set Set,
+                                         llvm::ArrayRef<int> Sizes) {
+  unsigned Dims = unsignedFromIslSize(Set.tuple_dim());
+  unsigned NumTileDims = Sizes.size();
+  assert(Dims >= NumTileDims);
+  isl::local_space LocalSpace = isl::local_space(Set.get_space());
+
+  for (unsigned i = 0; i < NumTileDims; ++i) {
+    unsigned Pos = Dims - NumTileDims + i;
+
+    isl::constraint LowerBound = isl::constraint::alloc_inequality(LocalSpace);
+    LowerBound = LowerBound.set_constant_si(0);
+    LowerBound = LowerBound.set_coefficient_si(isl::dim::set, Pos, 1);
+    Set = Set.add_constraint(LowerBound);
+
+    isl::constraint UpperBound = isl::constraint::alloc_inequality(LocalSpace);
+    UpperBound = UpperBound.set_constant_si(Sizes[i] - 1);
+    UpperBound = UpperBound.set_coefficient_si(isl::dim::set, Pos, -1);
+    Set = Set.add_constraint(UpperBound);
+  }
+
+  return Set;
+}
+
+isl::set polly::getFullTilePrefixes(isl::set ScheduleRange,
+                                    llvm::ArrayRef<int> TileSizes,
+                                    unsigned NumCompleteDims) {
+  unsigned Dims = unsignedFromIslSize(ScheduleRange.tuple_dim());
+  unsigned NumTileDims = TileSizes.size();
+  assert(Dims >= NumTileDims);
+  assert(NumCompleteDims >= 1 && NumCompleteDims <= NumTileDims);
+
+  // Same idea as getPartialTilePrefixes, but keeping the prefixes of the
+  // complete tiles instead of the incomplete ones: over-approximate the
+  // schedule range in the dimensions whose completeness is required, add the
+  // constraints of a complete tile there, and drop every prefix for which that
+  // tile leaves the schedule range.
+  //
+  // Requiring completeness in fewer than all point dimensions leaves the outer
+  // tile dimensions unconstrained in the result, so the AST generator has to
+  // separate the isolated part along the innermost dimensions only. That trades
+  // some of the constant loop bounds for a smaller number of copies of the loop
+  // nest.
+  isl::set Relaxed = ScheduleRange.drop_constraints_involving_dims(
+      isl::dim::set, Dims - NumCompleteDims, NumCompleteDims);
+  Relaxed =
+      addTileExtentConstraints(Relaxed, TileSizes.take_back(NumCompleteDims));
+
+  isl::set BadPrefixes = Relaxed.subtract(ScheduleRange);
+  BadPrefixes =
+      BadPrefixes.project_out(isl::dim::set, Dims - NumTileDims, NumTileDims);
+  isl::set LoopPrefixes =
+      ScheduleRange.project_out(isl::dim::set, Dims - NumTileDims, NumTileDims);
+
+  return LoopPrefixes.subtract(BadPrefixes);
+}
+
 isl::union_set polly::getIsolateOptions(isl::set IsolateDomain,
                                         unsigned OutDimsNum) {
   if (IsolateDomain.is_null())
diff --git a/polly/test/ScheduleOptimizer/isolate-full-tiles.ll b/polly/test/ScheduleOptimizer/isolate-full-tiles.ll
new file mode 100644
index 0000000000000..b3f3019827d7f
--- /dev/null
+++ b/polly/test/ScheduleOptimizer/isolate-full-tiles.ll
@@ -0,0 +1,111 @@
+; RUN: opt %loadNPMPolly '-passes=polly-custom<opt-isl;ast>' -polly-print-ast \
+; RUN:     -disable-output < %s | FileCheck %s --check-prefix=DEFAULT
+; RUN: opt %loadNPMPolly -polly-isolate-full-tiles \
+; RUN:     '-passes=polly-custom<opt-isl;ast>' -polly-print-ast \
+; RUN:     -disable-output < %s | FileCheck %s --check-prefix=ALL
+; RUN: opt %loadNPMPolly -polly-isolate-full-tiles \
+; RUN:     -polly-isolate-complete-tile-dims=1 \
+; RUN:     '-passes=polly-custom<opt-isl;ast>' -polly-print-ast \
+; RUN:     -disable-output < %s | FileCheck %s --check-prefix=INNERMOST
+;
+;    void foo(float *A, float *B) {
+;      for (long i = 0; i < 100; i++)
+;        for (long j = 0; j < 100; j++)
+;          A[100 * i + j] = B[100 * i + j] + 1;
+;    }
+;
+; 100 iterations tiled by 32 leave a partial tile of four in each dimension.
+;
+; Without isolation there is a single copy of the loop nest and both point
+; loops carry a min() upper bound.
+;
+; DEFAULT:      // 1st level tiling - Tiles
+; DEFAULT-NEXT: for (int c0 = 0; c0 <= 3; c0 += 1)
+; DEFAULT-NEXT:   for (int c1 = 0; c1 <= 3; c1 += 1) {
+; DEFAULT-NEXT:     // 1st level tiling - Points
+; DEFAULT-NEXT:     for (int c2 = 0; c2 <= min(31, -32 * c0 + 99); c2 += 1)
+; DEFAULT-NEXT:       for (int c3 = 0; c3 <= min(31, -32 * c1 + 99); c3 += 1)
+; DEFAULT-NEXT:         Stmt_for_j(32 * c0 + c2, 32 * c1 + c3);
+; DEFAULT-NOT:  // 1st level tiling - Points
+;
+; With isolation the complete tiles are separated from the two tails. The nine
+; complete tiles have constant bounds in both point loops.
+;
+; ALL:      // 1st level tiling - Tiles
+; ALL:      for (int c0 = 0; c0 <= 2; c0 += 1) {
+; ALL-NEXT:   for (int c1 = 0; c1 <= 2; c1 += 1) {
+; ALL-NEXT:     // 1st level tiling - Points
+; ALL-NEXT:     for (int c2 = 0; c2 <= 31; c2 += 1)
+; ALL-NEXT:       for (int c3 = 0; c3 <= 31; c3 += 1)
+; ALL-NEXT:         Stmt_for_j(32 * c0 + c2, 32 * c1 + c3);
+;
+; The tail of the second dimension, columns 96 to 99, runs once per complete
+; tile row and keeps its constant bound in the first dimension.
+;
+; ALL:      // 1st level tiling - Points
+; ALL-NEXT: for (int c2 = 0; c2 <= 31; c2 += 1)
+; ALL-NEXT:   for (int c3 = 0; c3 <= 3; c3 += 1)
+; ALL-NEXT:     Stmt_for_j(32 * c0 + c2, c3 + 96);
+;
+; The tail of the first dimension, rows 96 to 99, spans the whole width and is
+; the only part left with a min() bound.
+;
+; ALL:      for (int c1 = 0; c1 <= 3; c1 += 1) {
+; ALL-NEXT:   // 1st level tiling - Points
+; ALL-NEXT:   for (int c2 = 0; c2 <= 3; c2 += 1)
+; ALL-NEXT:     for (int c3 = 0; c3 <= min(31, -32 * c1 + 99); c3 += 1)
+; ALL-NEXT:       Stmt_for_j(c2 + 96, 32 * c1 + c3);
+; ALL-NOT:  // 1st level tiling - Points
+;
+; Requiring only the innermost dimension to be complete separates along that
+; dimension alone. The outer tile loop is not split, so the nest is copied
+; twice instead of three times and only the innermost point loop gets a
+; constant bound.
+;
+; INNERMOST:      // 1st level tiling - Tiles
+; INNERMOST-NEXT: for (int c0 = 0; c0 <= 3; c0 += 1) {
+; INNERMOST-NEXT:   for (int c1 = 0; c1 <= 2; c1 += 1) {
+; INNERMOST-NEXT:     // 1st level tiling - Points
+; INNERMOST-NEXT:     for (int c2 = 0; c2 <= min(31, -32 * c0 + 99); c2 += 1)
+; INNERMOST-NEXT:       for (int c3 = 0; c3 <= 31; c3 += 1)
+; INNERMOST-NEXT:         Stmt_for_j(32 * c0 + c2, 32 * c1 + c3);
+;
+; Its tail keeps the min() bound of the dimension that was left alone.
+;
+; INNERMOST:      // 1st level tiling - Points
+; INNERMOST-NEXT: for (int c2 = 0; c2 <= min(31, -32 * c0 + 99); c2 += 1)
+; INNERMOST-NEXT:   for (int c3 = 0; c3 <= 3; c3 += 1)
+; INNERMOST-NEXT:     Stmt_for_j(32 * c0 + c2, c3 + 96);
+; INNERMOST-NOT:  // 1st level tiling - Points
+
+target datalayout = "e-m:e-i64:64-f80:128-n8:16:32:64-S128"
+
+define void @foo(ptr %A, ptr %B) {
+entry:
+  br label %for.i
+
+for.i:
+  %i = phi i64 [ 0, %entry ], [ %i.next, %for.i.inc ]
+  br label %for.j
+
+for.j:
+  %j = phi i64 [ 0, %for.i ], [ %j.next, %for.j ]
+  %mul = mul nuw nsw i64 %i, 100
+  %idx = add nuw nsw i64 %mul, %j
+  %ptrB = getelementptr inbounds float, ptr %B, i64 %idx
+  %valB = load float, ptr %ptrB
+  %add = fadd float %valB, 1.000000e+00
+  %ptrA = getelementptr inbounds float, ptr %A, i64 %idx
+  store float %add, ptr %ptrA
+  %j.next = add nuw nsw i64 %j, 1
+  %j.cmp = icmp eq i64 %j.next, 100
+  br i1 %j.cmp, label %for.i.inc, label %for.j
+
+for.i.inc:
+  %i.next = add nuw nsw i64 %i, 1
+  %i.cmp = icmp eq i64 %i.next, 100
+  br i1 %i.cmp, label %exit, label %for.i
+
+exit:
+  ret void
+}

>From 6c634dfb1754fd94515f49f2005da0f19e091458 Mon Sep 17 00:00:00 2001
From: Timur Baidusenov <timurbaidusenov at gmail.com>
Date: Mon, 21 Sep 2026 17:18:30 +0300
Subject: [PATCH 2/3] [Polly] Address review comments

- Use the term "complete tile" from the OpenMP specification for the tiles
  that lie entirely inside the iteration space: getFullTilePrefixes becomes
  getCompleteTilePrefixes, isolateFullTiles becomes isolateCompleteTiles and
  -polly-isolate-full-tiles becomes -polly-isolate-complete-tiles. The test is
  renamed to match.
- Give every assertion added by the patch a message.

Assisted-by: Claude (Anthropic)
---
 polly/include/polly/ScheduleTreeTransform.h   |  6 +--
 polly/lib/Transform/ScheduleOptimizer.cpp     | 41 ++++++++++---------
 polly/lib/Transform/ScheduleTreeTransform.cpp | 15 ++++---
 ...ull-tiles.ll => isolate-complete-tiles.ll} |  4 +-
 4 files changed, 36 insertions(+), 30 deletions(-)
 rename polly/test/ScheduleOptimizer/{isolate-full-tiles.ll => isolate-complete-tiles.ll} (97%)

diff --git a/polly/include/polly/ScheduleTreeTransform.h b/polly/include/polly/ScheduleTreeTransform.h
index cd2cbbe244203..05f2b06e400ad 100644
--- a/polly/include/polly/ScheduleTreeTransform.h
+++ b/polly/include/polly/ScheduleTreeTransform.h
@@ -248,9 +248,9 @@ isl::set getPartialTilePrefixes(isl::set ScheduleRange, int VectorWidth);
 ///                        be complete. Requiring fewer than all of them
 ///                        isolates more tiles, but bounds the number of copies
 ///                        of the loop nest that the AST generator creates.
-isl::set getFullTilePrefixes(isl::set ScheduleRange,
-                             llvm::ArrayRef<int> TileSizes,
-                             unsigned NumCompleteDims);
+isl::set getCompleteTilePrefixes(isl::set ScheduleRange,
+                                 llvm::ArrayRef<int> TileSizes,
+                                 unsigned NumCompleteDims);
 
 /// Create an isl::union_set, which describes the isolate option based on
 /// IsolateDomain.
diff --git a/polly/lib/Transform/ScheduleOptimizer.cpp b/polly/lib/Transform/ScheduleOptimizer.cpp
index 978855a752b69..325c2d385f753 100644
--- a/polly/lib/Transform/ScheduleOptimizer.cpp
+++ b/polly/lib/Transform/ScheduleOptimizer.cpp
@@ -171,10 +171,11 @@ static cl::list<int>
                                "with --polly-register-tile-size"),
                       cl::Hidden, cl::CommaSeparated, cl::cat(PollyCategory));
 
-static cl::opt<bool> IsolateFullTiles(
-    "polly-isolate-full-tiles",
-    cl::desc("Separate the full tiles of a tiled band from the partial ones, "
-             "so that the point loops of the full tiles have constant bounds"),
+static cl::opt<bool> IsolateCompleteTiles(
+    "polly-isolate-complete-tiles",
+    cl::desc("Separate the complete tiles of a tiled band from the partial "
+             "ones, so that the point loops of the complete tiles have "
+             "constant bounds"),
     cl::Hidden, cl::init(false), cl::cat(PollyCategory));
 
 static cl::opt<int> IsolateCompleteTileDims(
@@ -419,20 +420,21 @@ ScheduleTreeOptimizer::isolateFullPartialTiles(isl::schedule_node Node,
   return Result;
 }
 
-/// Separate the full tiles of a tiled band from the partial ones.
+/// Separate the complete tiles of a tiled band from the partial ones.
 ///
-/// The point loops of a full tile run over the whole tile, so isolating those
-/// tiles gives them constant loop bounds instead of the min() expressions that
-/// a tiling of an iteration space which is not a multiple of the tile size
+/// The point loops of a complete tile run over the whole tile, so isolating
+/// those tiles gives them constant loop bounds instead of the min() expressions
+/// that a tiling of an iteration space which is not a multiple of the tile size
 /// produces. The partial tiles are left to a single atomic copy of the loop
 /// nest to keep the code growth bounded.
 ///
 /// @param Node      The point band of the tiling, as returned by tileNode.
 /// @param TileSizes The tile size of each tiled dimension.
 /// @return          The point band of the modified tree.
-static isl::schedule_node isolateFullTiles(isl::schedule_node Node,
-                                           ArrayRef<int> TileSizes) {
-  assert(isl_schedule_node_get_type(Node.get()) == isl_schedule_node_band);
+static isl::schedule_node isolateCompleteTiles(isl::schedule_node Node,
+                                               ArrayRef<int> TileSizes) {
+  assert(isl_schedule_node_get_type(Node.get()) == isl_schedule_node_band &&
+         "Expecting the point band that tileNode returned");
 
   // Below the point band, the prefix schedule covers the outer dimensions
   // followed by the tile and the point dimensions of this tiling.
@@ -447,13 +449,14 @@ static isl::schedule_node isolateFullTiles(isl::schedule_node Node,
     NumCompleteDims =
         std::min<unsigned>(IsolateCompleteTileDims, TileSizes.size());
 
-  isl::set FullTilePrefixes =
-      getFullTilePrefixes(ScheduleRange, TileSizes, NumCompleteDims);
-  if (FullTilePrefixes.is_null())
+  isl::set CompleteTilePrefixes =
+      getCompleteTilePrefixes(ScheduleRange, TileSizes, NumCompleteDims);
+  if (CompleteTilePrefixes.is_null())
     return Node;
 
-  isl::union_set Options = getIsolateOptions(FullTilePrefixes, TileSizes.size())
-                               .unite(getDimOptions(Node.ctx(), "atomic"));
+  isl::union_set Options =
+      getIsolateOptions(CompleteTilePrefixes, TileSizes.size())
+          .unite(getDimOptions(Node.ctx(), "atomic"));
 
   // The option describes the tile dimensions, so it belongs to the tile band,
   // which sits above the marker separating it from the point band.
@@ -594,7 +597,7 @@ ScheduleTreeOptimizer::applyTileBandOpt(isl::schedule_node Node) {
   if (FirstLevelTiling) {
     // Resolve the tile size of every dimension before tiling splits the band.
     SmallVector<int, 4> Sizes;
-    if (IsolateFullTiles) {
+    if (IsolateCompleteTiles) {
       isl::space Space =
           isl::manage(isl_schedule_node_band_get_space(Node.get()));
       for (unsigned i : rangeIslSize(0, Space.dim(isl::dim::set)))
@@ -607,8 +610,8 @@ ScheduleTreeOptimizer::applyTileBandOpt(isl::schedule_node Node) {
                     FirstLevelDefaultTileSize);
     FirstLevelTileOpts++;
 
-    if (IsolateFullTiles)
-      Node = isolateFullTiles(Node, Sizes);
+    if (IsolateCompleteTiles)
+      Node = isolateCompleteTiles(Node, Sizes);
   }
 
   if (SecondLevelTiling) {
diff --git a/polly/lib/Transform/ScheduleTreeTransform.cpp b/polly/lib/Transform/ScheduleTreeTransform.cpp
index ebbcf1bb2c050..ccd0399bf11a2 100644
--- a/polly/lib/Transform/ScheduleTreeTransform.cpp
+++ b/polly/lib/Transform/ScheduleTreeTransform.cpp
@@ -1133,7 +1133,8 @@ static isl::set addTileExtentConstraints(isl::set Set,
                                          llvm::ArrayRef<int> Sizes) {
   unsigned Dims = unsignedFromIslSize(Set.tuple_dim());
   unsigned NumTileDims = Sizes.size();
-  assert(Dims >= NumTileDims);
+  assert(Dims >= NumTileDims &&
+         "Not enough dimensions for the given tile sizes");
   isl::local_space LocalSpace = isl::local_space(Set.get_space());
 
   for (unsigned i = 0; i < NumTileDims; ++i) {
@@ -1153,13 +1154,15 @@ static isl::set addTileExtentConstraints(isl::set Set,
   return Set;
 }
 
-isl::set polly::getFullTilePrefixes(isl::set ScheduleRange,
-                                    llvm::ArrayRef<int> TileSizes,
-                                    unsigned NumCompleteDims) {
+isl::set polly::getCompleteTilePrefixes(isl::set ScheduleRange,
+                                        llvm::ArrayRef<int> TileSizes,
+                                        unsigned NumCompleteDims) {
   unsigned Dims = unsignedFromIslSize(ScheduleRange.tuple_dim());
   unsigned NumTileDims = TileSizes.size();
-  assert(Dims >= NumTileDims);
-  assert(NumCompleteDims >= 1 && NumCompleteDims <= NumTileDims);
+  assert(Dims >= NumTileDims &&
+         "Not enough dimensions for the given tile sizes");
+  assert(NumCompleteDims >= 1 && NumCompleteDims <= NumTileDims &&
+         "The dimensions that have to be complete must be tiled ones");
 
   // Same idea as getPartialTilePrefixes, but keeping the prefixes of the
   // complete tiles instead of the incomplete ones: over-approximate the
diff --git a/polly/test/ScheduleOptimizer/isolate-full-tiles.ll b/polly/test/ScheduleOptimizer/isolate-complete-tiles.ll
similarity index 97%
rename from polly/test/ScheduleOptimizer/isolate-full-tiles.ll
rename to polly/test/ScheduleOptimizer/isolate-complete-tiles.ll
index b3f3019827d7f..57a0f4afb2e4f 100644
--- a/polly/test/ScheduleOptimizer/isolate-full-tiles.ll
+++ b/polly/test/ScheduleOptimizer/isolate-complete-tiles.ll
@@ -1,9 +1,9 @@
 ; RUN: opt %loadNPMPolly '-passes=polly-custom<opt-isl;ast>' -polly-print-ast \
 ; RUN:     -disable-output < %s | FileCheck %s --check-prefix=DEFAULT
-; RUN: opt %loadNPMPolly -polly-isolate-full-tiles \
+; RUN: opt %loadNPMPolly -polly-isolate-complete-tiles \
 ; RUN:     '-passes=polly-custom<opt-isl;ast>' -polly-print-ast \
 ; RUN:     -disable-output < %s | FileCheck %s --check-prefix=ALL
-; RUN: opt %loadNPMPolly -polly-isolate-full-tiles \
+; RUN: opt %loadNPMPolly -polly-isolate-complete-tiles \
 ; RUN:     -polly-isolate-complete-tile-dims=1 \
 ; RUN:     '-passes=polly-custom<opt-isl;ast>' -polly-print-ast \
 ; RUN:     -disable-output < %s | FileCheck %s --check-prefix=INNERMOST

>From f9b5c389553def22c9ad1e87a4c33fc1f13e90b7 Mon Sep 17 00:00:00 2001
From: Timur Baidusenov <timurbaidusenov at gmail.com>
Date: Tue, 22 Sep 2026 15:27:44 +0300
Subject: [PATCH 3/3] [Polly] Optionally isolate the complete tiles of the
 inner tiling levels

-polly-isolate-complete-tiles separates the complete tiles of the first level
of tiling. The second level of tiling and register tiling tile the point band
of the first level, and a complete first-level tile does not make the inner
tiles complete: they are complete only when the inner tile size divides the
outer one, and not along the diagonal of a triangular or skewed iteration
domain, whatever the sizes.

Add two options that isolate the inner levels in the same way:

 - -polly-isolate-complete-tiles-2nd-level separates the complete tiles of the
   second level of tiling.
 - -polly-isolate-complete-register-tiles separates the complete register
   tiles, so that their unrolled point loops need no guards. The matrix
   multiplication optimization already does this for its micro-kernel, see
   commit 9989088ee99a.

Both are off by default, because whether they pay off depends on whether the
problem sizes are known at compile time. On PolyBench/C 4.2.1 with -O3
-march=native -polly -polly-vectorizer=none on a Ryzen 7 9700X, the geometric
mean of the run time relative to the same flags without the new option, with
first-level isolation and the tiling level itself enabled in both, is:

                                 sizes known    parametric sizes
  second level isolated             0.965            1.056
  register tiles isolated           0.997            0.934

With parametric sizes AST generation also exceeds -polly-astgen-computeout on
7 of the 18 kernels that Polly tiles with the second-level option, and on up
to 13 with the register one, and Polly then silently keeps the original code.
The full measurements are attached to the pull request.

Assisted-by: Claude (Anthropic)
---
 polly/lib/Transform/ScheduleOptimizer.cpp     |  46 +++++-
 .../isolate-complete-tiles-inner-levels.ll    | 138 ++++++++++++++++++
 2 files changed, 176 insertions(+), 8 deletions(-)
 create mode 100644 polly/test/ScheduleOptimizer/isolate-complete-tiles-inner-levels.ll

diff --git a/polly/lib/Transform/ScheduleOptimizer.cpp b/polly/lib/Transform/ScheduleOptimizer.cpp
index 325c2d385f753..4ac9d81cc5114 100644
--- a/polly/lib/Transform/ScheduleOptimizer.cpp
+++ b/polly/lib/Transform/ScheduleOptimizer.cpp
@@ -186,6 +186,18 @@ static cl::opt<int> IsolateCompleteTileDims(
              "bound"),
     cl::Hidden, cl::init(0), cl::cat(PollyCategory));
 
+static cl::opt<bool> IsolateCompleteTiles2ndLevel(
+    "polly-isolate-complete-tiles-2nd-level",
+    cl::desc("Separate the complete tiles of the second level of tiling from "
+             "the partial ones"),
+    cl::Hidden, cl::init(false), cl::cat(PollyCategory));
+
+static cl::opt<bool> IsolateCompleteRegisterTiles(
+    "polly-isolate-complete-register-tiles",
+    cl::desc("Separate the complete register tiles from the partial ones, so "
+             "that their unrolled point loops need no guards"),
+    cl::Hidden, cl::init(false), cl::cat(PollyCategory));
+
 static cl::opt<bool> PragmaBasedOpts(
     "polly-pragma-based-opts",
     cl::desc("Apply user-directed transformation from metadata"),
@@ -592,19 +604,25 @@ bool ScheduleTreeOptimizer::isPMOptimizableBandNode(isl::schedule_node Node) {
   return Node.child(0).isa<isl::schedule_node_leaf>();
 }
 
+/// Resolve the tile size of every dimension of the band @p Node.
+static SmallVector<int, 4> resolveTileSizes(isl::schedule_node Node,
+                                            ArrayRef<int> TileSizes,
+                                            int DefaultTileSize) {
+  SmallVector<int, 4> Sizes;
+  isl::space Space = isl::manage(isl_schedule_node_band_get_space(Node.get()));
+  for (unsigned i : rangeIslSize(0, Space.dim(isl::dim::set)))
+    Sizes.push_back(i < TileSizes.size() ? TileSizes[i] : DefaultTileSize);
+  return Sizes;
+}
+
 __isl_give isl::schedule_node
 ScheduleTreeOptimizer::applyTileBandOpt(isl::schedule_node Node) {
   if (FirstLevelTiling) {
     // Resolve the tile size of every dimension before tiling splits the band.
     SmallVector<int, 4> Sizes;
-    if (IsolateCompleteTiles) {
-      isl::space Space =
-          isl::manage(isl_schedule_node_band_get_space(Node.get()));
-      for (unsigned i : rangeIslSize(0, Space.dim(isl::dim::set)))
-        Sizes.push_back(i < FirstLevelTileSizes.size()
-                            ? FirstLevelTileSizes[i]
-                            : FirstLevelDefaultTileSize.getValue());
-    }
+    if (IsolateCompleteTiles)
+      Sizes = resolveTileSizes(Node, FirstLevelTileSizes,
+                               FirstLevelDefaultTileSize);
 
     Node = tileNode(Node, "1st level tiling", FirstLevelTileSizes,
                     FirstLevelDefaultTileSize);
@@ -615,15 +633,27 @@ ScheduleTreeOptimizer::applyTileBandOpt(isl::schedule_node Node) {
   }
 
   if (SecondLevelTiling) {
+    SmallVector<int, 4> Sizes;
+    if (IsolateCompleteTiles2ndLevel)
+      Sizes = resolveTileSizes(Node, SecondLevelTileSizes,
+                               SecondLevelDefaultTileSize);
     Node = tileNode(Node, "2nd level tiling", SecondLevelTileSizes,
                     SecondLevelDefaultTileSize);
     SecondLevelTileOpts++;
+    if (IsolateCompleteTiles2ndLevel)
+      Node = isolateCompleteTiles(Node, Sizes);
   }
 
   if (RegisterTiling) {
+    SmallVector<int, 4> Sizes;
+    if (IsolateCompleteRegisterTiles)
+      Sizes =
+          resolveTileSizes(Node, RegisterTileSizes, RegisterDefaultTileSize);
     Node =
         applyRegisterTiling(Node, RegisterTileSizes, RegisterDefaultTileSize);
     RegisterTileOpts++;
+    if (IsolateCompleteRegisterTiles)
+      Node = isolateCompleteTiles(Node, Sizes);
   }
 
   return Node;
diff --git a/polly/test/ScheduleOptimizer/isolate-complete-tiles-inner-levels.ll b/polly/test/ScheduleOptimizer/isolate-complete-tiles-inner-levels.ll
new file mode 100644
index 0000000000000..8fe6131a45df5
--- /dev/null
+++ b/polly/test/ScheduleOptimizer/isolate-complete-tiles-inner-levels.ll
@@ -0,0 +1,138 @@
+; RUN: opt %loadNPMPolly -polly-isolate-complete-tiles \
+; RUN:     -polly-2nd-level-tiling -polly-2nd-level-default-tile-size=12 \
+; RUN:     '-passes=polly-custom<opt-isl;ast>' -polly-print-ast \
+; RUN:     -disable-output < %s | FileCheck %s --check-prefix=L2
+; RUN: opt %loadNPMPolly -polly-isolate-complete-tiles \
+; RUN:     -polly-2nd-level-tiling -polly-2nd-level-default-tile-size=12 \
+; RUN:     -polly-isolate-complete-tiles-2nd-level \
+; RUN:     '-passes=polly-custom<opt-isl;ast>' -polly-print-ast \
+; RUN:     -disable-output < %s | FileCheck %s --check-prefix=L2-ISO
+; RUN: opt %loadNPMPolly -polly-isolate-complete-tiles \
+; RUN:     -polly-register-tiling -polly-register-tiling-default-tile-size=3 \
+; RUN:     '-passes=polly-custom<opt-isl;ast>' -polly-print-ast \
+; RUN:     -disable-output < %s | FileCheck %s --check-prefix=REG
+; RUN: opt %loadNPMPolly -polly-isolate-complete-tiles \
+; RUN:     -polly-register-tiling -polly-register-tiling-default-tile-size=3 \
+; RUN:     -polly-isolate-complete-register-tiles \
+; RUN:     '-passes=polly-custom<opt-isl;ast>' -polly-print-ast \
+; RUN:     -disable-output < %s | FileCheck %s --check-prefix=REG-ISO
+;
+;    void foo(float *A, float *B) {
+;      for (long i = 0; i < 100; i++)
+;        for (long j = 0; j < 100; j++)
+;          A[100 * i + j] = B[100 * i + j] + 1;
+;    }
+;
+; A complete first-level tile is 32 wide, which is not a multiple of the inner
+; tile sizes used here, 12 for the second level and 3 for register tiling. The
+; inner tiles of a complete first-level tile are therefore not all complete,
+; and isolating the first level alone does not give their point loops constant
+; bounds.
+;
+; Without the second-level option, the second-level point loops keep a min()
+; bound inside the complete first-level tiles.
+;
+; L2:      // 1st level tiling - Tiles
+; L2:      for (int c0 = 0; c0 <= 2; c0 += 1) {
+; L2-NEXT:   for (int c1 = 0; c1 <= 2; c1 += 1) {
+; L2-NEXT:     // 1st level tiling - Points
+; L2-NEXT:     // 2nd level tiling - Tiles
+; L2-NEXT:     for (int c2 = 0; c2 <= 2; c2 += 1)
+; L2-NEXT:       for (int c3 = 0; c3 <= 2; c3 += 1) {
+; L2-NEXT:         // 2nd level tiling - Points
+; L2-NEXT:         for (int c4 = 0; c4 <= min(11, -12 * c2 + 31); c4 += 1)
+; L2-NEXT:           for (int c5 = 0; c5 <= min(11, -12 * c3 + 31); c5 += 1)
+; L2-NEXT:             Stmt_for_j(32 * c0 + 12 * c2 + c4, 32 * c1 + 12 * c3 + c5);
+;
+; With it, the two by two complete second-level tiles have constant bounds and
+; the remainder of eight is separated from them.
+;
+; L2-ISO:      // 1st level tiling - Tiles
+; L2-ISO:      for (int c0 = 0; c0 <= 2; c0 += 1) {
+; L2-ISO-NEXT:   for (int c1 = 0; c1 <= 2; c1 += 1) {
+; L2-ISO-NEXT:     // 1st level tiling - Points
+; L2-ISO-NEXT:     // 2nd level tiling - Tiles
+; L2-ISO-NEXT:     {
+; L2-ISO-NEXT:       for (int c2 = 0; c2 <= 1; c2 += 1) {
+; L2-ISO-NEXT:         for (int c3 = 0; c3 <= 1; c3 += 1) {
+; L2-ISO-NEXT:           // 2nd level tiling - Points
+; L2-ISO-NEXT:           for (int c4 = 0; c4 <= 11; c4 += 1)
+; L2-ISO-NEXT:             for (int c5 = 0; c5 <= 11; c5 += 1)
+; L2-ISO-NEXT:               Stmt_for_j(32 * c0 + 12 * c2 + c4, 32 * c1 + 12 * c3 + c5);
+; L2-ISO-NEXT:         }
+; L2-ISO-NEXT:         // 2nd level tiling - Points
+; L2-ISO-NEXT:         for (int c4 = 0; c4 <= 11; c4 += 1)
+; L2-ISO-NEXT:           for (int c5 = 0; c5 <= 7; c5 += 1)
+; L2-ISO-NEXT:             Stmt_for_j(32 * c0 + 12 * c2 + c4, 32 * c1 + c5 + 24);
+;
+; Without the register option, the unrolled body of a register tile guards the
+; statements that fall into the remainder of two.
+;
+; REG:      // 1st level tiling - Tiles
+; REG:      for (int c0 = 0; c0 <= 2; c0 += 1) {
+; REG-NEXT:   for (int c1 = 0; c1 <= 2; c1 += 1) {
+; REG-NEXT:     // 1st level tiling - Points
+; REG-NEXT:     // Register tiling - Tiles
+; REG-NEXT:     for (int c2 = 0; c2 <= 10; c2 += 1)
+; REG-NEXT:       for (int c3 = 0; c3 <= 10; c3 += 1) {
+; REG-NEXT:         // Register tiling - Points
+; REG-NEXT:         {
+; REG-NEXT:           Stmt_for_j(32 * c0 + 3 * c2, 32 * c1 + 3 * c3);
+; REG-NEXT:           Stmt_for_j(32 * c0 + 3 * c2, 32 * c1 + 3 * c3 + 1);
+; REG-NEXT:           if (c3 <= 9)
+; REG-NEXT:             Stmt_for_j(32 * c0 + 3 * c2, 32 * c1 + 3 * c3 + 2);
+;
+; With it, the complete register tiles are unrolled without any guard.
+;
+; REG-ISO:      // 1st level tiling - Tiles
+; REG-ISO:      for (int c0 = 0; c0 <= 2; c0 += 1) {
+; REG-ISO-NEXT:   for (int c1 = 0; c1 <= 2; c1 += 1) {
+; REG-ISO-NEXT:     // 1st level tiling - Points
+; REG-ISO-NEXT:     // Register tiling - Tiles
+; REG-ISO-NEXT:     {
+; REG-ISO-NEXT:       for (int c2 = 0; c2 <= 9; c2 += 1) {
+; REG-ISO-NEXT:         for (int c3 = 0; c3 <= 9; c3 += 1) {
+; REG-ISO-NEXT:           // Register tiling - Points
+; REG-ISO-NEXT:           {
+; REG-ISO-NEXT:             Stmt_for_j(32 * c0 + 3 * c2, 32 * c1 + 3 * c3);
+; REG-ISO-NEXT:             Stmt_for_j(32 * c0 + 3 * c2, 32 * c1 + 3 * c3 + 1);
+; REG-ISO-NEXT:             Stmt_for_j(32 * c0 + 3 * c2, 32 * c1 + 3 * c3 + 2);
+; REG-ISO-NEXT:             Stmt_for_j(32 * c0 + 3 * c2 + 1, 32 * c1 + 3 * c3);
+; REG-ISO-NEXT:             Stmt_for_j(32 * c0 + 3 * c2 + 1, 32 * c1 + 3 * c3 + 1);
+; REG-ISO-NEXT:             Stmt_for_j(32 * c0 + 3 * c2 + 1, 32 * c1 + 3 * c3 + 2);
+; REG-ISO-NEXT:             Stmt_for_j(32 * c0 + 3 * c2 + 2, 32 * c1 + 3 * c3);
+; REG-ISO-NEXT:             Stmt_for_j(32 * c0 + 3 * c2 + 2, 32 * c1 + 3 * c3 + 1);
+; REG-ISO-NEXT:             Stmt_for_j(32 * c0 + 3 * c2 + 2, 32 * c1 + 3 * c3 + 2);
+; REG-ISO-NEXT:           }
+
+target datalayout = "e-m:e-i64:64-f80:128-n8:16:32:64-S128"
+
+define void @foo(ptr %A, ptr %B) {
+entry:
+  br label %for.i
+
+for.i:
+  %i = phi i64 [ 0, %entry ], [ %i.next, %for.i.inc ]
+  br label %for.j
+
+for.j:
+  %j = phi i64 [ 0, %for.i ], [ %j.next, %for.j ]
+  %mul = mul nuw nsw i64 %i, 100
+  %idx = add nuw nsw i64 %mul, %j
+  %ptrB = getelementptr inbounds float, ptr %B, i64 %idx
+  %valB = load float, ptr %ptrB
+  %add = fadd float %valB, 1.000000e+00
+  %ptrA = getelementptr inbounds float, ptr %A, i64 %idx
+  store float %add, ptr %ptrA
+  %j.next = add nuw nsw i64 %j, 1
+  %j.cmp = icmp eq i64 %j.next, 100
+  br i1 %j.cmp, label %for.i.inc, label %for.j
+
+for.i.inc:
+  %i.next = add nuw nsw i64 %i, 1
+  %i.cmp = icmp eq i64 %i.next, 100
+  br i1 %i.cmp, label %exit, label %for.i
+
+exit:
+  ret void
+}



More information about the llvm-commits mailing list