[polly] 17065e8 - [Polly] Allocate packed arrays of matrix multiplication on heap (#226163)

via llvm-commits llvm-commits at lists.llvm.org
Wed Sep 30 14:46:29 PDT 2026


Author: Timur Baidusenov
Date: 2026-09-30T23:46:19+02:00
New Revision: 17065e8ad007fa47db2e8541ba00b4fe06cd0205

URL: https://github.com/llvm/llvm-project/commit/17065e8ad007fa47db2e8541ba00b4fe06cd0205
DIFF: https://github.com/llvm/llvm-project/commit/17065e8ad007fa47db2e8541ba00b4fe06cd0205.diff

LOG: [Polly] Allocate packed arrays of matrix multiplication on heap (#226163)

Problem: The matrix multiplication optimization copies blocks of its
operands into the arrays Packed_A and Packed_B. Their sizes are derived
from the cache parameters rather than from the operands; with the
default parameters, Packed_B takes 4 MiB even for a 100x100 product.
They are allocated with alloca in the entry block of the function, so
two optimized multiplications in one function exceed the default 8 MiB
stack: two products of 100x100 int32 matrices declared as VLAs and
inlined into main crash with a segmentation fault.

Solution: Allocate a packed array on the heap, with malloc at the start
of the SCoP and free at its exit, only if it is larger than
-polly-pattern-matching-max-stack-array-size (1 MiB by default), and
keep smaller ones on the stack. -1 keeps all of them on the stack, 0
puts all of them on the heap. With the default parameters, Packed_B goes
to the heap and Packed_A (192 KiB) stays on the stack.

Default size depends only on the cache parameters. The default
parameters are 32 KiB L1 and 256 KiB L2, both 8-way: the X86 TTI reports
these for every CPU, and they are the Polly defaults for targets whose
TTI reports none. With them, for gemm with float and double on x86-64
(generic, skylake-avx512, znver4) and AArch64 (generic, neoverse-v2), I
get:
- Packed_A: 192 KiB in all cases. It is sized to fit into the L2 cache
and takes at most (1 - 2/associativity) of it.
- Packed_B: 3 to 4 MiB. It holds 256
(-polly-pattern-matching-nc-quotient) micro-panels sized for the L1
cache.

Two Packed_B already exceed the default 8 MiB stack, so Packed_B has to
go to the heap, while Packed_A fits on the stack, like any threshold
from 192 KiB to 3 MiB. I chose 1 MiB because it also keeps Packed_A on
the stack if the L2 parameters are set for current CPUs with up to about
1.25 MiB of L2 (768 KiB for 1 MiB 8-way, 1 MiB for 1.25 MiB 10-way). It
also limits each array on the stack to 1/8 of the default 8 MiB stack of
Linux and macOS.

Assisted-by: Claude (Anthropic)

Added: 
    polly/test/ScheduleOptimizer/pattern-matching-based-opts-packed-arrays.ll

Modified: 
    polly/docs/ReleaseNotes.rst
    polly/lib/Transform/MatmulOptimizer.cpp

Removed: 
    


################################################################################
diff  --git a/polly/docs/ReleaseNotes.rst b/polly/docs/ReleaseNotes.rst
index 80890778c3fea..0ddc5a79a48c5 100644
--- a/polly/docs/ReleaseNotes.rst
+++ b/polly/docs/ReleaseNotes.rst
@@ -17,3 +17,7 @@ In Polly |version| the following important changes have been incorporated.
    When Polly is linked into ``opt``, the plain ``-polly-*`` options remain
    accepted for now.
 
+ * The matrix multiplication optimization allocates packed arrays larger than
+   ``-polly-pattern-matching-max-stack-array-size`` (1 MiB by default) on the
+   heap instead of the stack. ``-1`` keeps all of them on the stack.
+

diff  --git a/polly/lib/Transform/MatmulOptimizer.cpp b/polly/lib/Transform/MatmulOptimizer.cpp
index 2acb7b66298da..281fc594f68ff 100644
--- a/polly/lib/Transform/MatmulOptimizer.cpp
+++ b/polly/lib/Transform/MatmulOptimizer.cpp
@@ -128,6 +128,14 @@ static cl::opt<int> PollyPatternMatchingNcQuotient(
              "macro-kernel, by Nr, the parameter of the micro-kernel"),
     cl::Hidden, cl::init(256), cl::cat(PollyCategory));
 
+static cl::opt<int> MaxStackArraySize(
+    "polly-pattern-matching-max-stack-array-size",
+    cl::desc("The maximal size in bytes of a packed array of the matrix "
+             "multiplication optimization that is allocated on the stack; "
+             "larger ones are allocated on the heap (-1: all on the stack, "
+             "0: all on the heap)"),
+    cl::Hidden, cl::init(1024 * 1024), cl::cat(PollyCategory));
+
 static cl::opt<bool>
     PMBasedTCOpts("polly-tc-opt",
                   cl::desc("Perform optimizations of tensor contractions based "
@@ -795,6 +803,19 @@ static isl::schedule_node createExtensionNode(isl::schedule_node Node,
   return Node.graft_before(NewNode);
 }
 
+/// Allocate the packed array @p SAI, whose dimensions have the sizes
+/// @p DimSizes, on the heap if it is larger than
+/// -polly-pattern-matching-max-stack-array-size and that is not negative, and
+/// on the stack otherwise.
+static void setPackedArrayAllocation(ScopArrayInfo *SAI,
+                                     ArrayRef<unsigned> DimSizes) {
+  uint64_t Size = SAI->getElemSizeInBytes();
+  for (unsigned DimSize : DimSizes)
+    Size *= DimSize;
+  SAI->setIsOnHeap(MaxStackArraySize >= 0 &&
+                   Size > uint64_t(MaxStackArraySize));
+}
+
 static isl::schedule_node optimizePackedB(isl::schedule_node Node,
                                           ScopStmt *Stmt, isl::map MapOldIndVar,
                                           MicroKernelParamsTy MicroParams,
@@ -810,6 +831,8 @@ static isl::schedule_node optimizePackedB(isl::schedule_node Node,
   ScopArrayInfo *PackedB =
       S->createScopArrayInfo(MMI.B->getElementType(), "Packed_B",
                              {FirstDimSize, SecondDimSize, ThirdDimSize});
+  setPackedArrayAllocation(PackedB,
+                           {FirstDimSize, SecondDimSize, ThirdDimSize});
 
   // Compute the access relation for copying from B to PackedB.
   isl::map AccRelB = MMI.B->getLatestAccessRelation();
@@ -849,6 +872,8 @@ static isl::schedule_node optimizePackedA(isl::schedule_node Node, ScopStmt *,
   ScopArrayInfo *PackedA = Stmt->getParent()->createScopArrayInfo(
       MMI.A->getElementType(), "Packed_A",
       {FirstDimSize, SecondDimSize, ThirdDimSize});
+  setPackedArrayAllocation(PackedA,
+                           {FirstDimSize, SecondDimSize, ThirdDimSize});
 
   // Compute the access relation for copying from A to PackedA.
   isl::map AccRelA = MMI.A->getLatestAccessRelation();

diff  --git a/polly/test/ScheduleOptimizer/pattern-matching-based-opts-packed-arrays.ll b/polly/test/ScheduleOptimizer/pattern-matching-based-opts-packed-arrays.ll
new file mode 100644
index 0000000000000..9a453302d78ad
--- /dev/null
+++ b/polly/test/ScheduleOptimizer/pattern-matching-based-opts-packed-arrays.ll
@@ -0,0 +1,85 @@
+; RUN: opt %loadNPMPolly -plugin-arg=Polly,-polly-pattern-matching-based-opts=true -plugin-arg=Polly,-polly-target-throughput-vector-fma=1 -plugin-arg=Polly,-polly-target-latency-vector-fma=8 -plugin-arg=Polly,-polly-target-1st-cache-level-associativity=8 -plugin-arg=Polly,-polly-target-2nd-cache-level-associativity=8 -plugin-arg=Polly,-polly-target-1st-cache-level-size=32768 -plugin-arg=Polly,-polly-target-vector-register-bitwidth=256 -plugin-arg=Polly,-polly-target-2nd-cache-level-size=262144 '-passes=polly<no-default-opts;opt-isl>' -S < %s | FileCheck %s
+; RUN: opt %loadNPMPolly -plugin-arg=Polly,-polly-pattern-matching-based-opts=true -plugin-arg=Polly,-polly-target-throughput-vector-fma=1 -plugin-arg=Polly,-polly-target-latency-vector-fma=8 -plugin-arg=Polly,-polly-target-1st-cache-level-associativity=8 -plugin-arg=Polly,-polly-target-2nd-cache-level-associativity=8 -plugin-arg=Polly,-polly-target-1st-cache-level-size=32768 -plugin-arg=Polly,-polly-target-vector-register-bitwidth=256 -plugin-arg=Polly,-polly-target-2nd-cache-level-size=262144 -plugin-arg=Polly,-polly-pattern-matching-max-stack-array-size=-1 '-passes=polly<no-default-opts;opt-isl>' -S < %s | FileCheck %s --check-prefix=STACK
+; RUN: opt %loadNPMPolly -plugin-arg=Polly,-polly-pattern-matching-based-opts=true -plugin-arg=Polly,-polly-target-throughput-vector-fma=1 -plugin-arg=Polly,-polly-target-latency-vector-fma=8 -plugin-arg=Polly,-polly-target-1st-cache-level-associativity=8 -plugin-arg=Polly,-polly-target-2nd-cache-level-associativity=8 -plugin-arg=Polly,-polly-target-1st-cache-level-size=32768 -plugin-arg=Polly,-polly-target-vector-register-bitwidth=256 -plugin-arg=Polly,-polly-target-2nd-cache-level-size=262144 -plugin-arg=Polly,-polly-pattern-matching-max-stack-array-size=0 '-passes=polly<no-default-opts;opt-isl>' -S < %s | FileCheck %s --check-prefix=HEAP
+;
+; The packed arrays of the matrix multiplication optimization are sized by the
+; cache parameters: here Packed_A takes 192 KiB and Packed_B 4 MiB. Arrays
+; larger than -polly-pattern-matching-max-stack-array-size (1 MiB by default)
+; are allocated on the heap, the others on the stack. With -1 all of them are
+; allocated on the stack, with 0 all of them on the heap.
+;
+;    /* C := alpha*A*B + beta*C */
+;    for (i = 0; i < _PB_NI; i++)
+;      for (j = 0; j < _PB_NJ; j++)
+;        {
+;	   C[i][j] *= beta;
+;	   for (k = 0; k < _PB_NK; ++k)
+;	     C[i][j] += alpha * A[i][k] * B[k][j];
+;        }
+;
+; CHECK-LABEL: define internal void @kernel_gemm(
+; CHECK:         %Packed_A = alloca [24 x [256 x [4 x double]]]
+; CHECK:         %Packed_B = tail call ptr @malloc(i64 4194304)
+; CHECK-NOT:     call ptr @malloc
+; CHECK:         tail call void @free(ptr %Packed_B)
+; CHECK-NOT:     call void @free
+;
+; STACK-LABEL: define internal void @kernel_gemm(
+; STACK:         %Packed_B = alloca [256 x [256 x [8 x double]]]
+; STACK-NEXT:    %Packed_A = alloca [24 x [256 x [4 x double]]]
+; STACK-NOT:     call ptr @malloc
+;
+; HEAP-LABEL: define internal void @kernel_gemm(
+; HEAP-NOT:      %Packed_{{[AB]}} = alloca
+; HEAP:          %Packed_B = tail call ptr @malloc(i64 4194304)
+; HEAP-NEXT:     %Packed_A = tail call ptr @malloc(i64 196608)
+; HEAP:          tail call void @free(ptr %Packed_B)
+; HEAP-NEXT:     tail call void @free(ptr %Packed_A)
+
+target datalayout = "e-m:e-i64:64-f80:128-n8:16:32:64-S128"
+target triple = "x86_64-unknown-unknown"
+
+define internal void @kernel_gemm(i32 %arg, i32 %arg1, i32 %arg2, double %arg3, double %arg4, ptr %arg5, ptr %arg6, ptr %arg7) #0 {
+bb:
+  br label %bb8
+
+bb8:                                              ; preds = %bb29, %bb
+  %tmp = phi i64 [ 0, %bb ], [ %tmp30, %bb29 ]
+  br label %bb9
+
+bb9:                                              ; preds = %bb26, %bb8
+  %tmp10 = phi i64 [ 0, %bb8 ], [ %tmp27, %bb26 ]
+  %tmp11 = getelementptr inbounds [1056 x double], ptr %arg5, i64 %tmp, i64 %tmp10
+  %tmp12 = load double, ptr %tmp11, align 8
+  %tmp13 = fmul double %tmp12, %arg4
+  store double %tmp13, ptr %tmp11, align 8
+  br label %Copy_0
+
+Copy_0:                                             ; preds = %Copy_0, %bb9
+  %tmp15 = phi i64 [ 0, %bb9 ], [ %tmp24, %Copy_0 ]
+  %tmp16 = getelementptr inbounds [1024 x double], ptr %arg6, i64 %tmp, i64 %tmp15
+  %tmp17 = load double, ptr %tmp16, align 8
+  %tmp18 = fmul double %tmp17, %arg3
+  %tmp19 = getelementptr inbounds [1056 x double], ptr %arg7, i64 %tmp15, i64 %tmp10
+  %tmp20 = load double, ptr %tmp19, align 8
+  %tmp21 = fmul double %tmp18, %tmp20
+  %tmp22 = load double, ptr %tmp11, align 8
+  %tmp23 = fadd double %tmp22, %tmp21
+  store double %tmp23, ptr %tmp11, align 8
+  %tmp24 = add nuw nsw i64 %tmp15, 1
+  %tmp25 = icmp ne i64 %tmp24, 1024
+  br i1 %tmp25, label %Copy_0, label %bb26
+
+bb26:                                             ; preds = %Copy_0
+  %tmp27 = add nuw nsw i64 %tmp10, 1
+  %tmp28 = icmp ne i64 %tmp27, 1056
+  br i1 %tmp28, label %bb9, label %bb29
+
+bb29:                                             ; preds = %bb26
+  %tmp30 = add nuw nsw i64 %tmp, 1
+  %tmp31 = icmp ne i64 %tmp30, 1056
+  br i1 %tmp31, label %bb8, label %bb32
+
+bb32:                                             ; preds = %bb29
+  ret void
+}


        


More information about the llvm-commits mailing list