[Mlir-commits] [mlir] [mlir][acc] Run acc-specialize-for-host rewrite to convergence (PR #205727)
llvmlistbot at llvm.org
llvmlistbot at llvm.org
Wed Jun 24 22:47:18 PDT 2026
https://github.com/khaki3 created https://github.com/llvm/llvm-project/pull/205727
Example:
```fortran
!$acc routine worker
subroutine transform(p, n)
real*8 p(*); integer n
!$acc loop seq
do i1 = 1, n
!$acc loop seq
do i2 = 1, n
! ... a dozen+ levels of nested acc loops ...
!$acc loop vector
do i = 1, n
p(i) = p(i) + 1.0d0
end do
end do
end do
end subroutine
```
In this code, the routine becomes a `func.func` with deeply nested orphan `acc.loop` ops. `acc-specialize-for-host` lowers them (e.g. `acc.loop` → `scf.for`) via `applyPatternsGreedily`, but leaves `GreedyRewriteConfig::maxIterations` at its default of 10. Since inner loops only become rewritable after their parents convert, a nest deeper than 10 isn't at a fixed point when the cap is hit, so the driver returns `failure()` and the pass calls `signalPassFailure()` — a spurious, diagnostic-less failure even though the conversion was progressing correctly.
Fix:
Run the rewrite to convergence instead of stopping at the default cap:
```cpp
config.setMaxIterations(GreedyRewriteConfig::kNoLimit);
```
The patterns are strictly reductive (ops are lowered/erased, never regenerated), so this is safe. Adds a regression test with a 16-deep orphan `acc.loop` nest.
>From 897609588b303f9d4f25cb5785e858e4cdc862db Mon Sep 17 00:00:00 2001
From: Kazuaki Matsumura <kmatsumura at nvidia.com>
Date: Wed, 24 Jun 2026 22:41:07 -0700
Subject: [PATCH] [mlir][acc] Run acc-specialize-for-host rewrite to
convergence
ACCSpecializeForHost lowers orphan acc operations using
applyPatternsGreedily with the default GreedyRewriteConfig, whose
maxIterations cap is 10. Deeply nested orphan acc.loop nests need
roughly one greedy iteration per level to reach a fixed point, so a
nest deeper than the cap is not yet converged when the limit is hit.
applyPatternsGreedily then returns failure and the pass calls
signalPassFailure(), producing a spurious, diagnostic-less failure even
though the conversion was making correct, monotonic progress.
Set maxIterations to kNoLimit so the strictly-reductive orphan-to-host
patterns run to convergence. Add a regression test with a 16-deep
orphan acc.loop nest.
---
.../Transforms/ACCSpecializeForHost.cpp | 3 +
.../acc-specialize-for-host-deep-nesting.mlir | 65 +++++++++++++++++++
2 files changed, 68 insertions(+)
create mode 100644 mlir/test/Dialect/OpenACC/acc-specialize-for-host-deep-nesting.mlir
diff --git a/mlir/lib/Dialect/OpenACC/Transforms/ACCSpecializeForHost.cpp b/mlir/lib/Dialect/OpenACC/Transforms/ACCSpecializeForHost.cpp
index 633538069c268..aab03b31ae70b 100644
--- a/mlir/lib/Dialect/OpenACC/Transforms/ACCSpecializeForHost.cpp
+++ b/mlir/lib/Dialect/OpenACC/Transforms/ACCSpecializeForHost.cpp
@@ -344,6 +344,9 @@ class ACCSpecializeForHost
populateACCOrphanToHostPatterns(patterns, accSupport);
GreedyRewriteConfig config;
config.setUseTopDownTraversal(true);
+ // Deeply nested orphan acc routines can need more than the default
+ // iteration cap to converge; lift it to avoid spurious pass failure.
+ config.setMaxIterations(GreedyRewriteConfig::kNoLimit);
if (failed(applyPatternsGreedily(funcOp, std::move(patterns), config)))
signalPassFailure();
}
diff --git a/mlir/test/Dialect/OpenACC/acc-specialize-for-host-deep-nesting.mlir b/mlir/test/Dialect/OpenACC/acc-specialize-for-host-deep-nesting.mlir
new file mode 100644
index 0000000000000..4d7940dbda462
--- /dev/null
+++ b/mlir/test/Dialect/OpenACC/acc-specialize-for-host-deep-nesting.mlir
@@ -0,0 +1,65 @@
+// RUN: mlir-opt %s -acc-specialize-for-host | FileCheck %s
+
+// Deeply nested orphan acc.loops in a seq routine require more than the
+// greedy rewriter's default iteration cap to fully lower to scf.for. The
+// pass must run the rewrite to convergence rather than bailing out.
+
+acc.routine @acc_routine_deep func(@deeply_nested_orphan_loops) seq
+// CHECK-LABEL: func.func @deeply_nested_orphan_loops
+// CHECK-NOT: acc.loop
+// CHECK-COUNT-16: scf.for
+func.func @deeply_nested_orphan_loops(%arg0 : memref<i32>) attributes {acc.routine_info = #acc.routine_info<[@acc_routine_deep]>} {
+ %c0 = arith.constant 0 : i32
+ %c10 = arith.constant 10 : i32
+ %c1 = arith.constant 1 : i32
+ acc.loop control(%iv0 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ acc.loop control(%iv1 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ acc.loop control(%iv2 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ acc.loop control(%iv3 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ acc.loop control(%iv4 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ acc.loop control(%iv5 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ acc.loop control(%iv6 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ acc.loop control(%iv7 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ acc.loop control(%iv8 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ acc.loop control(%iv9 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ acc.loop control(%iv10 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ acc.loop control(%iv11 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ acc.loop control(%iv12 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ acc.loop control(%iv13 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ acc.loop control(%iv14 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ acc.loop control(%iv15 : i32) = (%c0 : i32) to (%c10 : i32) step (%c1 : i32) {
+ memref.store %iv15, %arg0[] : memref<i32>
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ acc.yield
+ } attributes {inclusiveUpperbound = array<i1: true>, seq = [#acc.device_type<none>]}
+ return
+}
More information about the Mlir-commits
mailing list