[polly] [Polly] Isolate complete tiles from partial tiles in the tiling path (PR #221087)
Timur Baidusenov via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 22 06:23:59 PDT 2026
bai-tim wrote:
> Did you intent to only cover 1st level tiling, and not 2nd level/register tiling?
Yes, the patch only covered the first level. I have now measured the inner levels and added two options for them in the latest commit, both off by default:
- `-polly-isolate-complete-tiles-2nd-level` isolates the complete tiles of `-polly-2nd-level-tiling`.
- `-polly-isolate-complete-register-tiles` isolates the complete register tiles, so that their unrolled point loops need no guards, as `isolateAndUnrollMatMulInnerLoops` already does for the matmul micro-kernel.
First-level isolation does not cover the inner levels by itself: the inner tiles of a complete first-level tile are complete only if the inner tile size divides the outer one, and not along the diagonal of a triangular or skewed domain, whatever the sizes. The new test `isolate-complete-tiles-inner-levels.ll` shows both levels on a 100x100 loop with tile sizes 32/12 and 32/3.
Whether isolating the inner levels pays off depends on whether the problem sizes are known at compile time. PolyBench/C 4.2.1, same setup as in the description; for parametric sizes the kernels are compiled with `-Dstatic= -fno-inline-functions`, so that Polly sees the sizes as parameters instead of the constants it gets after inlining. Geometric mean relative to the same flags without the new option:
| | sizes known | parametric sizes |
|---|---:|---:|
| run time, 2nd-level isolation | 0.965 | 1.056 |
| run time, register-tile isolation | 0.997 | 0.934 |
| compile time, 2nd-level isolation | 1.17 | 2.30 |
| compile time, register-tile isolation | 1.10 | 2.04 |
Run time over 23 and 11 kernels, compile time over 23 and 18.
- With known sizes, 2nd-level isolation helps triangular kernels: syrk and covariance take 16% and 14% less time.
- With parametric sizes, it slows down the stencils: fdtd-2d by 1.47x and jacobi-2d by 1.33x.
- With parametric sizes, AST generation exceeds `-polly-astgen-computeout` on 7 of the 18 kernels Polly tiles (13 with the register option), and Polly silently keeps the original code. The extra compile time goes mostly into schedule optimization, before AST generation, which no limit bounds: ludcmp takes 26 s to compile instead of 0.9 s.
Hence both stay off by default.
A side result that may be worth a follow-up: first-level isolation is what makes `-polly-2nd-level-tiling` pay off. With known sizes, L1+L2 has a geometric mean of 0.579 against 0.627 for first-level isolation alone, and syr2k and syrk run 4.1x and 1.8x faster than with first-level isolation alone. It also slows down gemver by 1.78x and doitgen by 1.30x, though, so enabling it would need a cost model.
All dumps stay byte-identical to Polly without isolation (MINI and MEDIUM, every configuration, both size cases: 1440 comparisons). Per-kernel run time, the AST-limit fallbacks, code size and compile time are in the attached report.
https://github.com/llvm/llvm-project/pull/221087
More information about the llvm-commits
mailing list