[flang-commits] [flang] [flang] Apply the offload-region stack arrays rule in the allocation policy (PR #228539)
Zhen Wang via flang-commits
flang-commits at lists.llvm.org
Fri Oct 2 12:52:07 PDT 2026
https://github.com/wangzpgi updated https://github.com/llvm/llvm-project/pull/228539
>From 7c777c4142e3527f02318827ec2e9b0e83553116 Mon Sep 17 00:00:00 2001
From: Zhen Wang <zhenw at nvidia.com>
Date: Thu, 1 Oct 2026 14:15:00 -0700
Subject: [PATCH 1/3] Apply the offload-region stack arrays rule in the
allocation policy
---
flang/docs/FlangDriver.md | 2 +
flang/docs/fstack-arrays.md | 33 +++++++++++++++
.../Optimizer/Support/AllocationPolicy.h | 13 +++++-
flang/lib/Lower/OpenACC.cpp | 4 --
flang/lib/Optimizer/Builder/CUFCommon.cpp | 3 +-
.../Optimizer/Support/AllocationPolicy.cpp | 40 ++++++++++++++++---
.../Transforms/AllocationPlacement.cpp | 13 ++----
.../lib/Optimizer/Transforms/StackArrays.cpp | 3 +-
.../OpenACC/acc-routine-allocation-policy.f90 | 39 ------------------
.../acc-routine-bind-clone-signature.f90 | 2 +-
.../acc-routine-bind-devtype-filter.f90 | 10 ++---
.../acc-routine-bind-devtype-undeclared.f90 | 2 +-
.../acc-routine-bind-string-undeclared.f90 | 2 +-
.../OpenACC/acc-routine-bind-undeclared.f90 | 2 +-
.../Lower/OpenACC/acc-routine-multi-name.f90 | 10 ++---
.../test/Lower/OpenACC/acc-routine-named.f90 | 4 +-
.../Lower/OpenACC/acc-routine-use-module.f90 | 2 +-
flang/test/Lower/OpenACC/acc-routine.f90 | 22 +++++-----
flang/test/Lower/OpenACC/acc-routine02.f90 | 2 +-
flang/test/Lower/OpenACC/acc-routine03.f90 | 4 +-
flang/test/Lower/OpenACC/acc-routine04.f90 | 4 +-
21 files changed, 120 insertions(+), 96 deletions(-)
delete mode 100644 flang/test/Lower/OpenACC/acc-routine-allocation-policy.f90
diff --git a/flang/docs/FlangDriver.md b/flang/docs/FlangDriver.md
index c9df05ded7f49..f9701355f80ef 100644
--- a/flang/docs/FlangDriver.md
+++ b/flang/docs/FlangDriver.md
@@ -599,6 +599,8 @@ documentation for more details.
## Ofast and Fast Math
`-Ofast` in Flang means `-O3 -ffast-math -fstack-arrays -fno-protect-parens`.
+`-fstack-arrays` is not applied to code that runs on an accelerator, see the
+"Device code" section of [fstack-arrays.md](fstack-arrays.md).
`-ffast-math` means the following:
- `-fno-honor-infinities`
diff --git a/flang/docs/fstack-arrays.md b/flang/docs/fstack-arrays.md
index 9f3f385b4834a..b21e33bd44456 100644
--- a/flang/docs/fstack-arrays.md
+++ b/flang/docs/fstack-arrays.md
@@ -138,9 +138,42 @@ The attribute will be called `"fir.must_be_heap"` and will have a boolean value:
meaning that stack arrays may move the allocation. Not specifying the attribute
will be equivalent to setting it to `false`.
+### Device code
+`-fstack-arrays` is not honored in code that runs on an accelerator: CUDA
+Fortran `device` and `global` procedures, OpenACC compute constructs
+(`acc.parallel`, `acc.kernels`, `acc.serial`) and `cuf.kernel` loops. The
+device stack is far smaller than the host one
+(1 KB per thread by default on NVIDIA GPUs) and every thread of a launch pays
+for the storage, so a runtime-sized array that fits the host stack easily
+overflows it. In those contexts runtime-sized allocations stay on the heap. The
+size based part of the allocation policy still applies, so small constant size
+temporaries may still be placed on the device stack, and fixed size local arrays
+are unaffected.
+
+The rule is part of the allocation policy (`flang/Optimizer/Support/
+AllocationPolicy.h`), so every stack-or-heap decision applies it:
+- Lowering records a `fir.allocation_policy` attribute with `stack_arrays`
+ disabled on device procedures. `fir::getAllocationPolicy` honors the
+ innermost policy, so the function attribute narrows the module one.
+- `fir::shouldAllocateOnStack` ignores `stack_arrays` when the allocation's
+ context operation is nested in an offload region (`fir::isInOffloadRegion`).
+ The `allocation-placement` pass passes the allocation itself as the context;
+ the `stack-arrays` pass skips such candidates.
+
+An `acc routine` procedure is lowered once for both the host and the device,
+so it keeps the host policy. Keeping its device copy off the device stack is
+left to the device code generation, which sees the runtime-sized allocations
+once the routine has been specialized for the device. Host code, including the
+host copy of an `attributes(host,device)` procedure and of an `acc routine`, is
+unaffected.
+
## Testing Plan
FileCheck tests will be written to check each of the above identified sources of
heap allocated array temporaries are detected and converted by the new pass.
Another test will check that `allocate` statements in source code will not be
moved to the stack.
+
+Tests for device code check that allocations inside offload regions and in
+device procedures are left on the heap while the same allocation in host code
+is moved to the stack.
diff --git a/flang/include/flang/Optimizer/Support/AllocationPolicy.h b/flang/include/flang/Optimizer/Support/AllocationPolicy.h
index 4f9ce3df4c826..1a72e64365edc 100644
--- a/flang/include/flang/Optimizer/Support/AllocationPolicy.h
+++ b/flang/include/flang/Optimizer/Support/AllocationPolicy.h
@@ -84,6 +84,10 @@ struct PendingAllocationInfo {
bool isDynamic = false;
/// The constant size of the allocation in bytes, if it can be determined.
std::optional<std::int64_t> byteSize;
+ /// An operation at the point where the allocation lives or will be inserted,
+ /// used for context dependent decisions such as code that runs on a device.
+ /// Optional: without it the offload region rule is not applied.
+ mlir::Operation *context = nullptr;
};
/// Module-level information needed to compute constant allocation sizes.
@@ -110,10 +114,17 @@ struct AllocationInfo : PendingAllocationInfo {
bool isCurrentlyOnStack = false;
};
+/// Return true if \p op is nested in a region that is offloaded to a device:
+/// an OpenACC compute construct or a CUDA Fortran kernel loop. Code there runs
+/// on the device stack, which is far smaller than the host one.
+bool isInOffloadRegion(mlir::Operation *op);
+
/// Size-based placement policy, usable before the allocation is created.
/// Decides whether an allocation described by \p info should live on the stack,
/// given the \p policy in effect and the per-function stack bytes already
-/// committed to the stack (\p stackBytesUsed).
+/// committed to the stack (\p stackBytesUsed). When \p info carries a context
+/// inside an offload region, -fstack-arrays is not honored there and only the
+/// size based rules apply, as for a device procedure.
bool shouldAllocateOnStack(const PendingAllocationInfo &info,
const AllocationPolicy &policy,
std::size_t stackBytesUsed);
diff --git a/flang/lib/Lower/OpenACC.cpp b/flang/lib/Lower/OpenACC.cpp
index 327edb579fe46..eec08190aa6a1 100644
--- a/flang/lib/Lower/OpenACC.cpp
+++ b/flang/lib/Lower/OpenACC.cpp
@@ -25,7 +25,6 @@
#include "flang/Lower/Support/Utils.h"
#include "flang/Lower/SymbolMap.h"
#include "flang/Optimizer/Builder/BoxValue.h"
-#include "flang/Optimizer/Builder/CUFCommon.h"
#include "flang/Optimizer/Builder/FIRBuilder.h"
#include "flang/Optimizer/Builder/HLFIRTools.h"
#include "flang/Optimizer/Builder/IntrinsicCall.h"
@@ -4808,9 +4807,6 @@ static void attachRoutineInfo(mlir::func::FuncOp func,
func.getOperation()->setAttr(
mlir::acc::getRoutineInfoAttrName(),
mlir::acc::RoutineInfoAttr::get(func.getContext(), routines));
- // The routine is compiled for the device as well, where -fstack-arrays
- // cannot be honored: record the same policy as for a device procedure.
- cuf::setDeviceAllocationPolicy(func.getOperation());
}
static mlir::ArrayAttr
diff --git a/flang/lib/Optimizer/Builder/CUFCommon.cpp b/flang/lib/Optimizer/Builder/CUFCommon.cpp
index de5831a05636a..c9cf48c1eaed9 100644
--- a/flang/lib/Optimizer/Builder/CUFCommon.cpp
+++ b/flang/lib/Optimizer/Builder/CUFCommon.cpp
@@ -74,8 +74,7 @@ bool cuf::isCUDADeviceContext(mlir::Region ®ion,
bool cuf::isExecutingOnDevice(mlir::Operation *op) {
if (!op)
return false;
- if (op->getParentOfType<cuf::KernelOp>() ||
- op->getParentOfType<mlir::acc::OffloadRegionOpInterface>() ||
+ if (fir::isInOffloadRegion(op) ||
op->getParentOfType<mlir::gpu::GPUModuleOp>() ||
op->getParentOfType<mlir::gpu::LaunchOp>() ||
op->getParentOfType<mlir::gpu::GPUFuncOp>())
diff --git a/flang/lib/Optimizer/Support/AllocationPolicy.cpp b/flang/lib/Optimizer/Support/AllocationPolicy.cpp
index 32a64b50b6f51..5e7dfa04455fa 100644
--- a/flang/lib/Optimizer/Support/AllocationPolicy.cpp
+++ b/flang/lib/Optimizer/Support/AllocationPolicy.cpp
@@ -11,8 +11,10 @@
//===----------------------------------------------------------------------===//
#include "flang/Optimizer/Support/AllocationPolicy.h"
+#include "flang/Optimizer/Dialect/CUDAKernelOpInterface.h"
#include "flang/Optimizer/Dialect/FIRAttr.h"
#include "flang/Optimizer/Dialect/FIRType.h"
+#include "mlir/Dialect/OpenACC/OpenACC.h"
#include "mlir/IR/BuiltinOps.h"
#include "llvm/Support/CommandLine.h"
@@ -63,9 +65,29 @@ bool fir::shouldUseStackForCopyin(mlir::Location loc, mlir::Type sequenceType,
return shouldAllocateOnStack(info, copyInPolicy, /*stackBytesUsed=*/0);
}
-bool fir::shouldAllocateOnStack(const PendingAllocationInfo &info,
- const AllocationPolicy &policy,
- std::size_t stackBytesUsed) {
+bool fir::isInOffloadRegion(mlir::Operation *op) {
+ for (mlir::Operation *cur = op ? op->getParentOp() : nullptr; cur;
+ cur = cur->getParentOp())
+ if (mlir::isa<mlir::acc::OffloadRegionOpInterface,
+ fir::CUDAKernelOpInterface>(cur))
+ return true;
+ return false;
+}
+
+/// The policy in effect for \p info. Inside an offload region -fstack-arrays is
+/// not honored, since the device stack is far smaller than the host one.
+static fir::AllocationPolicy policyFor(const fir::PendingAllocationInfo &info,
+ const fir::AllocationPolicy &policy) {
+ fir::AllocationPolicy effective = policy;
+ if (effective.stackArrays && fir::isInOffloadRegion(info.context))
+ effective.stackArrays = false;
+ return effective;
+}
+
+/// shouldAllocateOnStack with the policy already narrowed by policyFor.
+static bool shouldAllocateOnStackImpl(const fir::PendingAllocationInfo &info,
+ const fir::AllocationPolicy &policy,
+ std::size_t stackBytesUsed) {
// -fstack-arrays: put everything on the stack (best effort). For existing
// allocations, the heap-to-stack conversion still only happens where it is
// provably safe.
@@ -93,11 +115,19 @@ bool fir::shouldAllocateOnStack(const PendingAllocationInfo &info,
stackBytesUsed + size <= policy.totalStackLimitBytes;
}
+bool fir::shouldAllocateOnStack(const PendingAllocationInfo &info,
+ const AllocationPolicy &basePolicy,
+ std::size_t stackBytesUsed) {
+ return shouldAllocateOnStackImpl(info, policyFor(info, basePolicy),
+ stackBytesUsed);
+}
+
fir::AllocationPlacement
fir::decideAllocationPlacement(const AllocationInfo &info,
- const AllocationPolicy &policy,
+ const AllocationPolicy &basePolicy,
std::size_t stackBytesUsed) {
using P = fir::AllocationPlacement;
+ AllocationPolicy policy = policyFor(info, basePolicy);
// An allocation that is not known to be dynamic but whose size cannot be
// determined cannot be reasoned about: leave it where it is instead of
@@ -107,7 +137,7 @@ fir::decideAllocationPlacement(const AllocationInfo &info,
// Translate the "should this be on the stack" decision into a placement,
// accounting for where the allocation currently lives.
- bool wantStack = fir::shouldAllocateOnStack(info, policy, stackBytesUsed);
+ bool wantStack = shouldAllocateOnStackImpl(info, policy, stackBytesUsed);
if (wantStack)
return info.isCurrentlyOnStack ? P::Leave : P::Stack;
return info.isCurrentlyOnStack ? P::Heap : P::Leave;
diff --git a/flang/lib/Optimizer/Transforms/AllocationPlacement.cpp b/flang/lib/Optimizer/Transforms/AllocationPlacement.cpp
index de6028b92dae5..89783d7814831 100644
--- a/flang/lib/Optimizer/Transforms/AllocationPlacement.cpp
+++ b/flang/lib/Optimizer/Transforms/AllocationPlacement.cpp
@@ -17,7 +17,6 @@
//===----------------------------------------------------------------------===//
#include "StackArrays.h"
-#include "flang/Optimizer/Builder/CUFCommon.h"
#include "flang/Optimizer/Dialect/FIRAttr.h"
#include "flang/Optimizer/Dialect/FIRDialect.h"
#include "flang/Optimizer/Dialect/FIROps.h"
@@ -201,6 +200,7 @@ void AllocationPlacementPass::runOnOperation() {
fir::AllocationInfo info;
info.op = op;
+ info.context = op;
info.isCurrentlyOnStack = static_cast<bool>(alloca);
info.isTemporary = isTemporaryAllocation(op);
info.isDynamic =
@@ -208,19 +208,12 @@ void AllocationPlacementPass::runOnOperation() {
: (allocmem.hasLenParams() || allocmem.hasShapeOperands());
info.byteSize = getConstantByteSize(op, dl, kindMap);
- // -fstack-arrays cannot be honored in an offload region either: like a
- // device procedure, it runs on the device stack, which is far smaller than
- // the host one. The size based part of the policy still applies.
- fir::AllocationPolicy policy = basePolicy;
- if (policy.stackArrays && cuf::isExecutingOnDevice(op))
- policy.stackArrays = false;
-
// A hook, if provided, fully overrides the default policy; it may delegate
// back to decideAllocationPlacement after adjusting the policy.
fir::AllocationPlacement placement =
placementHook
- ? placementHook(info, policy, stackBytesUsed)
- : fir::decideAllocationPlacement(info, policy, stackBytesUsed);
+ ? placementHook(info, basePolicy, stackBytesUsed)
+ : fir::decideAllocationPlacement(info, basePolicy, stackBytesUsed);
// Account for the decision in the running stack budget.
if (endsUpOnStack(placement, info.isCurrentlyOnStack) && info.byteSize)
diff --git a/flang/lib/Optimizer/Transforms/StackArrays.cpp b/flang/lib/Optimizer/Transforms/StackArrays.cpp
index 575c853b0bc21..863b506a23597 100644
--- a/flang/lib/Optimizer/Transforms/StackArrays.cpp
+++ b/flang/lib/Optimizer/Transforms/StackArrays.cpp
@@ -7,7 +7,6 @@
//===----------------------------------------------------------------------===//
#include "StackArrays.h"
-#include "flang/Optimizer/Builder/CUFCommon.h"
#include "flang/Optimizer/Builder/FIRBuilder.h"
#include "flang/Optimizer/Builder/LowLevelIntrinsics.h"
#include "flang/Optimizer/Dialect/FIRAttr.h"
@@ -797,7 +796,7 @@ void StackArraysPass::runOnOperation() {
llvm::SmallVector<mlir::Operation *> opsToConvert;
opsToConvert.reserve(candidateOps->size());
for (auto [op, _] : *candidateOps)
- if (!cuf::isExecutingOnDevice(op))
+ if (!fir::isInOffloadRegion(op))
opsToConvert.push_back(op);
if (opsToConvert.empty())
diff --git a/flang/test/Lower/OpenACC/acc-routine-allocation-policy.f90 b/flang/test/Lower/OpenACC/acc-routine-allocation-policy.f90
deleted file mode 100644
index aa2510ddda2c1..0000000000000
--- a/flang/test/Lower/OpenACC/acc-routine-allocation-policy.f90
+++ /dev/null
@@ -1,39 +0,0 @@
-! Test that -fstack-arrays is not applied to acc routines. They are compiled for
-! the device as well, where the stack is far smaller than on the host, so
-! lowering records the same policy as for a device procedure. The policy is
-! recorded whether or not -fstack-arrays was requested, so that the opt-out is
-! explicit in the IR.
-
-! RUN: %flang_fc1 -emit-fir -fopenacc %s -o - | FileCheck %s --check-prefixes=CHECK,DEFAULT
-! RUN: %flang_fc1 -emit-fir -fopenacc -fstack-arrays %s -o - | FileCheck %s --check-prefixes=CHECK,STACK
-
-subroutine routine_seq(a, n)
- !$acc routine seq
- integer :: a(*)
- integer :: n
- integer :: auto(n)
- do i = 1, n
- auto(i) = i
- end do
- a(1) = sum(auto(1:n))
-end subroutine
-
-subroutine host_sub(a, n)
- integer :: a(*)
- integer :: n
- integer :: auto(n)
- do i = 1, n
- auto(i) = i
- end do
- a(1) = sum(auto(1:n))
-end subroutine
-
-! The module policy follows -fstack-arrays.
-! DEFAULT: module attributes {{.*}}fir.allocation_policy = #fir.allocation_policy<stack_arrays = false
-! STACK: module attributes {{.*}}fir.allocation_policy = #fir.allocation_policy<stack_arrays = true
-
-! The acc routine opts out of it in both cases.
-! CHECK: func.func @_QProutine_seq({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@acc_routine_0]>, fir.allocation_policy = #fir.allocation_policy<stack_arrays = false
-
-! The host procedure keeps the module policy.
-! CHECK: func.func @_QPhost_sub({{.*}}) {
diff --git a/flang/test/Lower/OpenACC/acc-routine-bind-clone-signature.f90 b/flang/test/Lower/OpenACC/acc-routine-bind-clone-signature.f90
index 63b9cb3cf17d5..181e1883a68d1 100644
--- a/flang/test/Lower/OpenACC/acc-routine-bind-clone-signature.f90
+++ b/flang/test/Lower/OpenACC/acc-routine-bind-clone-signature.f90
@@ -24,4 +24,4 @@ subroutine aclear(y)
! The decorated routine's signature (assumed-shape array descriptor):
! CHECK: func.func private @_QPaclear(!fir.box<!fir.array<?xf32>>
! The bind target is declared with the same type, proving the clone:
-! CHECK: func.func private @_QPaclear_seq(!fir.box<!fir.array<?xf32>>) attributes {acc.routine_info = #acc.routine_info<[@[[ACLEAR_SEQ_ROUTINE]]]>{{.*}}}
+! CHECK: func.func private @_QPaclear_seq(!fir.box<!fir.array<?xf32>>) attributes {acc.routine_info = #acc.routine_info<[@[[ACLEAR_SEQ_ROUTINE]]]>}
diff --git a/flang/test/Lower/OpenACC/acc-routine-bind-devtype-filter.f90 b/flang/test/Lower/OpenACC/acc-routine-bind-devtype-filter.f90
index c17668364b4ff..725dec744b01b 100644
--- a/flang/test/Lower/OpenACC/acc-routine-bind-devtype-filter.f90
+++ b/flang/test/Lower/OpenACC/acc-routine-bind-devtype-filter.f90
@@ -17,8 +17,8 @@ subroutine s_bind_devtype_filter(n, x)
! CHECK-DAG: acc.routine @{{.*}} func(@_QPfoo) bind(@_QPfoo_n [#acc.device_type<nvidia>], @_QPfoo_m [#acc.device_type<multicore>]) worker ([#acc.device_type<multicore>]) vector ([#acc.device_type<nvidia>])
! CHECK-DAG: acc.routine @[[FOO_N_ROUTINE:.*]] func(@_QPfoo_n) vector ([#acc.device_type<nvidia>]){{$}}
! CHECK-DAG: acc.routine @[[FOO_M_ROUTINE:.*]] func(@_QPfoo_m) worker ([#acc.device_type<multicore>]){{$}}
-! CHECK-DAG: func.func private @_QPfoo_n({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[FOO_N_ROUTINE]]]>{{.*}}}
-! CHECK-DAG: func.func private @_QPfoo_m({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[FOO_M_ROUTINE]]]>{{.*}}}
+! CHECK-DAG: func.func private @_QPfoo_n({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[FOO_N_ROUTINE]]]>}
+! CHECK-DAG: func.func private @_QPfoo_m({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[FOO_M_ROUTINE]]]>}
subroutine s_bind_devtype_merged_target(n, x)
integer :: n, i
@@ -33,7 +33,7 @@ subroutine s_bind_devtype_merged_target(n, x)
! CHECK-DAG: acc.routine @{{.*}} func(@_QPfoo_merge) bind(@_QPfoo_dev [#acc.device_type<nvidia>], @_QPfoo_dev [#acc.device_type<multicore>]) worker ([#acc.device_type<multicore>]) vector ([#acc.device_type<nvidia>])
! CHECK-DAG: acc.routine @[[FOO_DEV_ROUTINE:.*]] func(@_QPfoo_dev) worker ([#acc.device_type<multicore>]) vector ([#acc.device_type<nvidia>]){{$}}
-! CHECK-DAG: func.func private @_QPfoo_dev({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[FOO_DEV_ROUTINE]]]>{{.*}}}
+! CHECK-DAG: func.func private @_QPfoo_dev({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[FOO_DEV_ROUTINE]]]>}
subroutine s_bind_before_modality(n, x)
integer :: n, i
@@ -49,5 +49,5 @@ subroutine s_bind_before_modality(n, x)
! CHECK-DAG: acc.routine @{{.*}} func(@_QPbar) bind(@_QPbar_n [#acc.device_type<nvidia>], @_QPbar_m [#acc.device_type<multicore>]) vector ([#acc.device_type<nvidia>]) seq ([#acc.device_type<multicore>])
! CHECK-DAG: acc.routine @[[BAR_N_ROUTINE:.*]] func(@_QPbar_n) vector ([#acc.device_type<nvidia>]){{$}}
! CHECK-DAG: acc.routine @[[BAR_M_ROUTINE:.*]] func(@_QPbar_m) seq ([#acc.device_type<multicore>]){{$}}
-! CHECK-DAG: func.func private @_QPbar_n({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[BAR_N_ROUTINE]]]>{{.*}}}
-! CHECK-DAG: func.func private @_QPbar_m({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[BAR_M_ROUTINE]]]>{{.*}}}
+! CHECK-DAG: func.func private @_QPbar_n({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[BAR_N_ROUTINE]]]>}
+! CHECK-DAG: func.func private @_QPbar_m({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[BAR_M_ROUTINE]]]>}
diff --git a/flang/test/Lower/OpenACC/acc-routine-bind-devtype-undeclared.f90 b/flang/test/Lower/OpenACC/acc-routine-bind-devtype-undeclared.f90
index a2add365202ca..74e13ff3a5bff 100644
--- a/flang/test/Lower/OpenACC/acc-routine-bind-devtype-undeclared.f90
+++ b/flang/test/Lower/OpenACC/acc-routine-bind-devtype-undeclared.f90
@@ -15,4 +15,4 @@ subroutine s_bind_devtype(n, x)
! CHECK: acc.routine @[[ACLEAR_DEV_ROUTINE:.*]] func(@_QPaclear_dev) seq
! CHECK: acc.routine @{{.*}} func(@_QPaclear){{.*}}@_QPaclear_dev
-! CHECK: func.func private @_QPaclear_dev({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[ACLEAR_DEV_ROUTINE]]]>{{.*}}}
+! CHECK: func.func private @_QPaclear_dev({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[ACLEAR_DEV_ROUTINE]]]>}
diff --git a/flang/test/Lower/OpenACC/acc-routine-bind-string-undeclared.f90 b/flang/test/Lower/OpenACC/acc-routine-bind-string-undeclared.f90
index eed98594bb3f4..6647c6986b8b1 100644
--- a/flang/test/Lower/OpenACC/acc-routine-bind-string-undeclared.f90
+++ b/flang/test/Lower/OpenACC/acc-routine-bind-string-undeclared.f90
@@ -5,7 +5,7 @@
! CHECK: acc.routine @[[ACLEAR_DEV_ROUTINE:.*]] func(@aclear_dev) seq
! CHECK: acc.routine @{{.*}} func(@_QPaclear) bind("aclear_dev") seq
-! CHECK: func.func private @aclear_dev({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[ACLEAR_DEV_ROUTINE]]]>{{.*}}}
+! CHECK: func.func private @aclear_dev({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[ACLEAR_DEV_ROUTINE]]]>}
! CHECK-SAME: loc("{{.*}}acc-routine-bind-string-undeclared.f90":{{[0-9]+}}:{{[0-9]+}})
! CHECK-NOT: func.func private @aclear_dev{{.*}}loc(unknown)
diff --git a/flang/test/Lower/OpenACC/acc-routine-bind-undeclared.f90 b/flang/test/Lower/OpenACC/acc-routine-bind-undeclared.f90
index d559aa520e366..c80e062dc8782 100644
--- a/flang/test/Lower/OpenACC/acc-routine-bind-undeclared.f90
+++ b/flang/test/Lower/OpenACC/acc-routine-bind-undeclared.f90
@@ -4,7 +4,7 @@
! CHECK: acc.routine @[[ACLEAR_SEQ_ROUTINE:.*]] func(@_QPaclear_seq) seq
! CHECK: acc.routine @{{.*}} func(@_QPaclear) bind(@_QPaclear_seq) seq
-! CHECK: func.func private @_QPaclear_seq({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[ACLEAR_SEQ_ROUTINE]]]>{{.*}}}
+! CHECK: func.func private @_QPaclear_seq({{.*}}) attributes {acc.routine_info = #acc.routine_info<[@[[ACLEAR_SEQ_ROUTINE]]]>}
! CHECK-SAME: loc("{{.*}}acc-routine-bind-undeclared.f90":{{[0-9]+}}:{{[0-9]+}})
! CHECK-NOT: func.func private @_QPaclear_seq{{.*}}loc(unknown)
diff --git a/flang/test/Lower/OpenACC/acc-routine-multi-name.f90 b/flang/test/Lower/OpenACC/acc-routine-multi-name.f90
index d293eaa5b48b7..6ba46f2da7fbb 100644
--- a/flang/test/Lower/OpenACC/acc-routine-multi-name.f90
+++ b/flang/test/Lower/OpenACC/acc-routine-multi-name.f90
@@ -23,30 +23,30 @@ subroutine seq1()
end subroutine
! CHECK-LABEL: func.func @_QMacc_multi_routinesPseq1()
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r_seq1]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r_seq1]]]>}
subroutine seq2()
end subroutine
! CHECK-LABEL: func.func @_QMacc_multi_routinesPseq2()
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r_seq2]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r_seq2]]]>}
subroutine gang1()
end subroutine
! CHECK-LABEL: func.func @_QMacc_multi_routinesPgang1()
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r_gang1]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r_gang1]]]>}
subroutine gang2()
end subroutine
! CHECK-LABEL: func.func @_QMacc_multi_routinesPgang2()
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r_gang2]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r_gang2]]]>}
subroutine gang3()
end subroutine
! CHECK-LABEL: func.func @_QMacc_multi_routinesPgang3()
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r_gang3]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r_gang3]]]>}
end module
diff --git a/flang/test/Lower/OpenACC/acc-routine-named.f90 b/flang/test/Lower/OpenACC/acc-routine-named.f90
index 3a8a2a21a195a..24d47e58b6e1b 100644
--- a/flang/test/Lower/OpenACC/acc-routine-named.f90
+++ b/flang/test/Lower/OpenACC/acc-routine-named.f90
@@ -15,13 +15,13 @@ subroutine acc1()
end subroutine
! CHECK-LABEL: func.func @_QMacc_routinesPacc1()
-! CHECK-SAME:attributes {acc.routine_info = #acc.routine_info<[@[[r1]]]>{{.*}}}
+! CHECK-SAME:attributes {acc.routine_info = #acc.routine_info<[@[[r1]]]>}
subroutine acc2()
!$acc routine(acc2)
end subroutine
! CHECK-LABEL: func.func @_QMacc_routinesPacc2()
-! CHECK-SAME:attributes {acc.routine_info = #acc.routine_info<[@[[r0]]]>{{.*}}}
+! CHECK-SAME:attributes {acc.routine_info = #acc.routine_info<[@[[r0]]]>}
end module
diff --git a/flang/test/Lower/OpenACC/acc-routine-use-module.f90 b/flang/test/Lower/OpenACC/acc-routine-use-module.f90
index 79316d405f77b..059324230a746 100644
--- a/flang/test/Lower/OpenACC/acc-routine-use-module.f90
+++ b/flang/test/Lower/OpenACC/acc-routine-use-module.f90
@@ -19,5 +19,5 @@ subroutine caller(aa)
!$acc end serial
end subroutine
!CHECK: }
- !CHECK: func.func private @_QMmod1Pcallee(!fir.ref<i32>) attributes {acc.routine_info = #acc.routine_info<[@acc_routine_0]>{{.*}}}
+ !CHECK: func.func private @_QMmod1Pcallee(!fir.ref<i32>) attributes {acc.routine_info = #acc.routine_info<[@acc_routine_0]>}
end module
\ No newline at end of file
diff --git a/flang/test/Lower/OpenACC/acc-routine.f90 b/flang/test/Lower/OpenACC/acc-routine.f90
index 37c2cc17240f5..c281ca5dfc287 100644
--- a/flang/test/Lower/OpenACC/acc-routine.f90
+++ b/flang/test/Lower/OpenACC/acc-routine.f90
@@ -24,56 +24,56 @@ subroutine acc_routine1()
end subroutine
! CHECK-LABEL: func.func @_QPacc_routine1()
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r00]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r00]]]>}
subroutine acc_routine2()
!$acc routine seq
end subroutine
! CHECK-LABEL: func.func @_QPacc_routine2()
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r01]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r01]]]>}
subroutine acc_routine3()
!$acc routine gang
end subroutine
! CHECK-LABEL: func.func @_QPacc_routine3()
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r02]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r02]]]>}
subroutine acc_routine4()
!$acc routine vector
end subroutine
! CHECK-LABEL: func.func @_QPacc_routine4()
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r03]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r03]]]>}
subroutine acc_routine5()
!$acc routine worker
end subroutine
! CHECK-LABEL: func.func @_QPacc_routine5()
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r04]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r04]]]>}
subroutine acc_routine6()
!$acc routine nohost
end subroutine
! CHECK-LABEL: func.func @_QPacc_routine6()
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r05]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r05]]]>}
subroutine acc_routine7()
!$acc routine gang(dim:1)
end subroutine
! CHECK-LABEL: func.func @_QPacc_routine7()
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r06]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r06]]]>}
subroutine acc_routine8()
!$acc routine bind("routine8_")
end subroutine
! CHECK-LABEL: func.func @_QPacc_routine8()
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r07]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r07]]]>}
subroutine acc_routine9a()
end subroutine
@@ -83,14 +83,14 @@ subroutine acc_routine9()
end subroutine
! CHECK-LABEL: func.func @_QPacc_routine9()
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r08]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r08]]]>}
function acc_routine10()
!$acc routine(acc_routine10) seq
end function
! CHECK-LABEL: func.func @_QPacc_routine10() -> f32
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r09]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r09]]]>}
subroutine acc_routine11(a)
real :: a
@@ -98,7 +98,7 @@ subroutine acc_routine11(a)
end subroutine
! CHECK-LABEL: func.func @_QPacc_routine11(%arg0: !fir.ref<f32> {fir.bindc_name = "a"})
-! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r10]]]>{{.*}}}
+! CHECK-SAME: attributes {acc.routine_info = #acc.routine_info<[@[[r10]]]>}
subroutine acc_routine12()
diff --git a/flang/test/Lower/OpenACC/acc-routine02.f90 b/flang/test/Lower/OpenACC/acc-routine02.f90
index 1c350c1ee33c8..dd07cba4b20e3 100644
--- a/flang/test/Lower/OpenACC/acc-routine02.f90
+++ b/flang/test/Lower/OpenACC/acc-routine02.f90
@@ -17,4 +17,4 @@ program test
! CHECK-LABEL: acc.routine @acc_routine_0 func(@_QPsub1)
-! CHECK: func.func @_QPsub1(%ar{{.*}}: !fir.ref<!fir.array<?xf32>> {fir.bindc_name = "a"}, %arg1: !fir.ref<i32> {fir.bindc_name = "n"}) attributes {acc.routine_info = #acc.routine_info<[@acc_routine_0]>{{.*}}}
+! CHECK: func.func @_QPsub1(%ar{{.*}}: !fir.ref<!fir.array<?xf32>> {fir.bindc_name = "a"}, %arg1: !fir.ref<i32> {fir.bindc_name = "n"}) attributes {acc.routine_info = #acc.routine_info<[@acc_routine_0]>}
diff --git a/flang/test/Lower/OpenACC/acc-routine03.f90 b/flang/test/Lower/OpenACC/acc-routine03.f90
index 0c2754ab3fcbf..3fc307746849f 100644
--- a/flang/test/Lower/OpenACC/acc-routine03.f90
+++ b/flang/test/Lower/OpenACC/acc-routine03.f90
@@ -31,5 +31,5 @@ subroutine sub2(a)
! CHECK: acc.routine @acc_routine_1 func(@_QPsub2) worker nohost
! CHECK: acc.routine @acc_routine_0 func(@_QPsub1) bind(@_QPsub2) worker
-! CHECK: func.func @_QPsub1(%arg0: !fir.box<!fir.array<?xf32>> {fir.bindc_name = "a"}) attributes {acc.routine_info = #acc.routine_info<[@acc_routine_0]>{{.*}}}
-! CHECK: func.func @_QPsub2(%arg0: !fir.box<!fir.array<?xf32>> {fir.bindc_name = "a"}) attributes {acc.routine_info = #acc.routine_info<[@acc_routine_1]>{{.*}}}
+! CHECK: func.func @_QPsub1(%arg0: !fir.box<!fir.array<?xf32>> {fir.bindc_name = "a"}) attributes {acc.routine_info = #acc.routine_info<[@acc_routine_0]>}
+! CHECK: func.func @_QPsub2(%arg0: !fir.box<!fir.array<?xf32>> {fir.bindc_name = "a"}) attributes {acc.routine_info = #acc.routine_info<[@acc_routine_1]>}
diff --git a/flang/test/Lower/OpenACC/acc-routine04.f90 b/flang/test/Lower/OpenACC/acc-routine04.f90
index fb8009ebe160c..470440728d2f5 100644
--- a/flang/test/Lower/OpenACC/acc-routine04.f90
+++ b/flang/test/Lower/OpenACC/acc-routine04.f90
@@ -29,6 +29,6 @@ subroutine sub2()
! CHECK: acc.routine @acc_routine_1 func(@_QFPsub2) seq
! CHECK: acc.routine @acc_routine_0 func(@_QMdummy_modPsub1) seq
-! CHECK: func.func @_QMdummy_modPsub1(%arg0: !fir.ref<i32> {fir.bindc_name = "i"}) attributes {acc.routine_info = #acc.routine_info<[@acc_routine_0]>{{.*}}}
+! CHECK: func.func @_QMdummy_modPsub1(%arg0: !fir.ref<i32> {fir.bindc_name = "i"}) attributes {acc.routine_info = #acc.routine_info<[@acc_routine_0]>}
! CHECK: func.func @_QQmain() attributes {fir.bindc_name = "TEST_ACC_ROUTINE"}
-! CHECK: func.func private @_QFPsub2() attributes {acc.routine_info = #acc.routine_info<[@acc_routine_1]>, fir.allocation_policy = {{.*}}, fir.host_symbol = @_QQmain, llvm.linkage = #llvm.linkage<internal>}
+! CHECK: func.func private @_QFPsub2() attributes {acc.routine_info = #acc.routine_info<[@acc_routine_1]>, fir.host_symbol = @_QQmain, llvm.linkage = #llvm.linkage<internal>}
>From fbabb7089c1ff4956ecaaa3a2399078d734b660f Mon Sep 17 00:00:00 2001
From: Zhen Wang <zhenw at nvidia.com>
Date: Fri, 2 Oct 2026 11:48:02 -0700
Subject: [PATCH 2/3] Cover gpu.launch, gpu.func, gpu.module and specialized
routines in isInOffloadRegion
---
.../Optimizer/Support/AllocationPolicy.h | 7 +++---
flang/lib/Optimizer/Builder/CUFCommon.cpp | 7 +-----
.../Optimizer/Support/AllocationPolicy.cpp | 8 +++++--
.../allocation-placement-offload-region.fir | 24 ++++++++++++++++++-
4 files changed, 34 insertions(+), 12 deletions(-)
diff --git a/flang/include/flang/Optimizer/Support/AllocationPolicy.h b/flang/include/flang/Optimizer/Support/AllocationPolicy.h
index 1a72e64365edc..b1b3a2d5ef7f3 100644
--- a/flang/include/flang/Optimizer/Support/AllocationPolicy.h
+++ b/flang/include/flang/Optimizer/Support/AllocationPolicy.h
@@ -114,9 +114,10 @@ struct AllocationInfo : PendingAllocationInfo {
bool isCurrentlyOnStack = false;
};
-/// Return true if \p op is nested in a region that is offloaded to a device:
-/// an OpenACC compute construct or a CUDA Fortran kernel loop. Code there runs
-/// on the device stack, which is far smaller than the host one.
+/// Return true if \p op is nested in code that is offloaded to a device: an
+/// OpenACC compute construct or specialized routine, a CUDA Fortran kernel
+/// loop, a gpu.launch, a gpu.func or a gpu.module. Code there runs on the
+/// device stack, which is far smaller than the host one.
bool isInOffloadRegion(mlir::Operation *op);
/// Size-based placement policy, usable before the allocation is created.
diff --git a/flang/lib/Optimizer/Builder/CUFCommon.cpp b/flang/lib/Optimizer/Builder/CUFCommon.cpp
index c9cf48c1eaed9..2a8dfbc203744 100644
--- a/flang/lib/Optimizer/Builder/CUFCommon.cpp
+++ b/flang/lib/Optimizer/Builder/CUFCommon.cpp
@@ -74,14 +74,9 @@ bool cuf::isCUDADeviceContext(mlir::Region ®ion,
bool cuf::isExecutingOnDevice(mlir::Operation *op) {
if (!op)
return false;
- if (fir::isInOffloadRegion(op) ||
- op->getParentOfType<mlir::gpu::GPUModuleOp>() ||
- op->getParentOfType<mlir::gpu::LaunchOp>() ||
- op->getParentOfType<mlir::gpu::GPUFuncOp>())
+ if (fir::isInOffloadRegion(op))
return true;
if (auto funcOp = op->getParentOfType<mlir::func::FuncOp>()) {
- if (mlir::acc::isSpecializedAccRoutine(funcOp))
- return true;
if (auto cudaProcAttr =
funcOp.getOperation()->getAttrOfType<cuf::ProcAttributeAttr>(
cuf::getProcAttrName())) {
diff --git a/flang/lib/Optimizer/Support/AllocationPolicy.cpp b/flang/lib/Optimizer/Support/AllocationPolicy.cpp
index 5e7dfa04455fa..d75aa89a7041f 100644
--- a/flang/lib/Optimizer/Support/AllocationPolicy.cpp
+++ b/flang/lib/Optimizer/Support/AllocationPolicy.cpp
@@ -14,6 +14,7 @@
#include "flang/Optimizer/Dialect/CUDAKernelOpInterface.h"
#include "flang/Optimizer/Dialect/FIRAttr.h"
#include "flang/Optimizer/Dialect/FIRType.h"
+#include "mlir/Dialect/GPU/IR/GPUDialect.h"
#include "mlir/Dialect/OpenACC/OpenACC.h"
#include "mlir/IR/BuiltinOps.h"
#include "llvm/Support/CommandLine.h"
@@ -67,10 +68,13 @@ bool fir::shouldUseStackForCopyin(mlir::Location loc, mlir::Type sequenceType,
bool fir::isInOffloadRegion(mlir::Operation *op) {
for (mlir::Operation *cur = op ? op->getParentOp() : nullptr; cur;
- cur = cur->getParentOp())
+ cur = cur->getParentOp()) {
if (mlir::isa<mlir::acc::OffloadRegionOpInterface,
- fir::CUDAKernelOpInterface>(cur))
+ fir::CUDAKernelOpInterface, mlir::gpu::LaunchOp,
+ mlir::gpu::GPUFuncOp, mlir::gpu::GPUModuleOp>(cur) ||
+ mlir::acc::isSpecializedAccRoutine(cur))
return true;
+ }
return false;
}
diff --git a/flang/test/Transforms/allocation-placement-offload-region.fir b/flang/test/Transforms/allocation-placement-offload-region.fir
index 50294b4d0fd49..50d8100713756 100644
--- a/flang/test/Transforms/allocation-placement-offload-region.fir
+++ b/flang/test/Transforms/allocation-placement-offload-region.fir
@@ -1,4 +1,5 @@
-// Test that -fstack-arrays is not honored inside an offload region: the region
+// Test that -fstack-arrays is not honored inside an offload region (an OpenACC
+// compute construct, a cuf.kernel loop or a gpu.launch): the region
// runs on the device stack, which is far smaller than the host one, so its
// runtime-sized temporaries stay on the heap while the same allocation in host
// code goes on the stack. The size based part of the policy still applies, so
@@ -105,6 +106,27 @@ func.func @dynamic_temp_in_cuf_kernel(%n: index) {
return
}
+// A runtime-sized alloca inside a gpu.launch is moved to the heap: the launch
+// body runs on the device as well.
+// CHECK-LABEL: func.func @dynamic_alloca_in_gpu_launch
+// CHECK: gpu.launch
+// CHECK: fir.allocmem !fir.array<?xi32>, %{{.*}}
+// CHECK: fir.freemem
+// CHECK-NOT: fir.alloca
+func.func @dynamic_alloca_in_gpu_launch(%n: index) {
+ %c0 = arith.constant 0 : index
+ %c1 = arith.constant 1 : index
+ %v = arith.constant 0 : i32
+ gpu.launch blocks(%bx, %by, %bz) in (%sbx = %c1, %sby = %c1, %sbz = %c1)
+ threads(%tx, %ty, %tz) in (%stx = %c1, %sty = %c1, %stz = %c1) {
+ %0 = fir.alloca !fir.array<?xi32>, %n
+ %e = fir.coordinate_of %0, %c0 : (!fir.ref<!fir.array<?xi32>>, index) -> !fir.ref<i32>
+ fir.store %v to %e : !fir.ref<i32>
+ gpu.terminator
+ }
+ return
+}
+
// <10xi32> is 40 bytes, under the 1024 byte threshold: the size based policy
// still puts it on the (device) stack.
// CHECK-LABEL: func.func @small_temp_in_acc_parallel
>From 4953a81ddbe133b2c0d24eb6be30c53afc97936b Mon Sep 17 00:00:00 2001
From: Zhen Wang <zhenw at nvidia.com>
Date: Fri, 2 Oct 2026 12:51:39 -0700
Subject: [PATCH 3/3] Move isInOffloadRegion to FIROpsSupport
---
.../flang/Optimizer/Dialect/FIROpsSupport.h | 6 ++++++
.../flang/Optimizer/Support/AllocationPolicy.h | 6 ------
flang/lib/Optimizer/Builder/CUFCommon.cpp | 1 +
flang/lib/Optimizer/Dialect/CMakeLists.txt | 1 +
flang/lib/Optimizer/Dialect/FIROps.cpp | 13 +++++++++++++
flang/lib/Optimizer/Support/AllocationPolicy.cpp | 16 +---------------
flang/lib/Optimizer/Transforms/StackArrays.cpp | 1 +
7 files changed, 23 insertions(+), 21 deletions(-)
diff --git a/flang/include/flang/Optimizer/Dialect/FIROpsSupport.h b/flang/include/flang/Optimizer/Dialect/FIROpsSupport.h
index 61864b24f8883..1fa01bdfc9e88 100644
--- a/flang/include/flang/Optimizer/Dialect/FIROpsSupport.h
+++ b/flang/include/flang/Optimizer/Dialect/FIROpsSupport.h
@@ -316,6 +316,12 @@ bool reboxPreservesContinuity(fir::ReboxOp rebox,
/// the checking is done for continuity of the whole result of embox
bool isContiguousEmbox(fir::EmboxOp embox, bool checkWhole = true);
+/// Return true if \p op is nested in code that is offloaded to a device: an
+/// OpenACC compute construct or specialized routine, a CUDA Fortran kernel
+/// loop, a gpu.launch, a gpu.func or a gpu.module. Code there runs on the
+/// device stack, which is far smaller than the host one.
+bool isInOffloadRegion(mlir::Operation *op);
+
} // namespace fir
#endif // FORTRAN_OPTIMIZER_DIALECT_FIROPSSUPPORT_H
diff --git a/flang/include/flang/Optimizer/Support/AllocationPolicy.h b/flang/include/flang/Optimizer/Support/AllocationPolicy.h
index b1b3a2d5ef7f3..0b9a3be169875 100644
--- a/flang/include/flang/Optimizer/Support/AllocationPolicy.h
+++ b/flang/include/flang/Optimizer/Support/AllocationPolicy.h
@@ -114,12 +114,6 @@ struct AllocationInfo : PendingAllocationInfo {
bool isCurrentlyOnStack = false;
};
-/// Return true if \p op is nested in code that is offloaded to a device: an
-/// OpenACC compute construct or specialized routine, a CUDA Fortran kernel
-/// loop, a gpu.launch, a gpu.func or a gpu.module. Code there runs on the
-/// device stack, which is far smaller than the host one.
-bool isInOffloadRegion(mlir::Operation *op);
-
/// Size-based placement policy, usable before the allocation is created.
/// Decides whether an allocation described by \p info should live on the stack,
/// given the \p policy in effect and the per-function stack bytes already
diff --git a/flang/lib/Optimizer/Builder/CUFCommon.cpp b/flang/lib/Optimizer/Builder/CUFCommon.cpp
index 2a8dfbc203744..eb2e64158b10b 100644
--- a/flang/lib/Optimizer/Builder/CUFCommon.cpp
+++ b/flang/lib/Optimizer/Builder/CUFCommon.cpp
@@ -10,6 +10,7 @@
#include "flang/Optimizer/Builder/FIRBuilder.h"
#include "flang/Optimizer/Builder/Todo.h"
#include "flang/Optimizer/Dialect/CUF/CUFOps.h"
+#include "flang/Optimizer/Dialect/FIROpsSupport.h"
#include "flang/Optimizer/Dialect/Support/KindMapping.h"
#include "flang/Optimizer/HLFIR/HLFIROps.h"
#include "flang/Optimizer/Support/AllocationPolicy.h"
diff --git a/flang/lib/Optimizer/Dialect/CMakeLists.txt b/flang/lib/Optimizer/Dialect/CMakeLists.txt
index a50736d17f987..08fd679c74c91 100644
--- a/flang/lib/Optimizer/Dialect/CMakeLists.txt
+++ b/flang/lib/Optimizer/Dialect/CMakeLists.txt
@@ -39,6 +39,7 @@ add_flang_library(FIRDialect
MLIR_LIBS
MLIRArithDialect
+ MLIRGPUDialect
MLIROpenACCDialect
MLIROpenACCUtils
MLIRBuiltinToLLVMIRTranslation
diff --git a/flang/lib/Optimizer/Dialect/FIROps.cpp b/flang/lib/Optimizer/Dialect/FIROps.cpp
index 2480996cb411c..f0e96324f9ae8 100644
--- a/flang/lib/Optimizer/Dialect/FIROps.cpp
+++ b/flang/lib/Optimizer/Dialect/FIROps.cpp
@@ -21,6 +21,7 @@
#include "flang/Optimizer/Dialect/Support/KindMapping.h"
#include "flang/Optimizer/Support/Utils.h"
#include "mlir/Dialect/Func/IR/FuncOps.h"
+#include "mlir/Dialect/GPU/IR/GPUDialect.h"
#include "mlir/Dialect/OpenACC/OpenACC.h"
#include "mlir/Dialect/OpenACC/OpenACCUtils.h"
#include "mlir/Dialect/OpenMP/OpenMPDialect.h"
@@ -7048,3 +7049,15 @@ void fir::FIROpsDialect::registerOpExternalInterfaces() {
#define GET_OP_CLASSES
#include "flang/Optimizer/Dialect/FIROps.cpp.inc"
+
+bool fir::isInOffloadRegion(mlir::Operation *op) {
+ for (mlir::Operation *cur = op ? op->getParentOp() : nullptr; cur;
+ cur = cur->getParentOp()) {
+ if (mlir::isa<mlir::acc::OffloadRegionOpInterface,
+ fir::CUDAKernelOpInterface, mlir::gpu::LaunchOp,
+ mlir::gpu::GPUFuncOp, mlir::gpu::GPUModuleOp>(cur) ||
+ mlir::acc::isSpecializedAccRoutine(cur))
+ return true;
+ }
+ return false;
+}
diff --git a/flang/lib/Optimizer/Support/AllocationPolicy.cpp b/flang/lib/Optimizer/Support/AllocationPolicy.cpp
index d75aa89a7041f..1e460d18096d9 100644
--- a/flang/lib/Optimizer/Support/AllocationPolicy.cpp
+++ b/flang/lib/Optimizer/Support/AllocationPolicy.cpp
@@ -11,11 +11,9 @@
//===----------------------------------------------------------------------===//
#include "flang/Optimizer/Support/AllocationPolicy.h"
-#include "flang/Optimizer/Dialect/CUDAKernelOpInterface.h"
#include "flang/Optimizer/Dialect/FIRAttr.h"
+#include "flang/Optimizer/Dialect/FIROpsSupport.h"
#include "flang/Optimizer/Dialect/FIRType.h"
-#include "mlir/Dialect/GPU/IR/GPUDialect.h"
-#include "mlir/Dialect/OpenACC/OpenACC.h"
#include "mlir/IR/BuiltinOps.h"
#include "llvm/Support/CommandLine.h"
@@ -66,18 +64,6 @@ bool fir::shouldUseStackForCopyin(mlir::Location loc, mlir::Type sequenceType,
return shouldAllocateOnStack(info, copyInPolicy, /*stackBytesUsed=*/0);
}
-bool fir::isInOffloadRegion(mlir::Operation *op) {
- for (mlir::Operation *cur = op ? op->getParentOp() : nullptr; cur;
- cur = cur->getParentOp()) {
- if (mlir::isa<mlir::acc::OffloadRegionOpInterface,
- fir::CUDAKernelOpInterface, mlir::gpu::LaunchOp,
- mlir::gpu::GPUFuncOp, mlir::gpu::GPUModuleOp>(cur) ||
- mlir::acc::isSpecializedAccRoutine(cur))
- return true;
- }
- return false;
-}
-
/// The policy in effect for \p info. Inside an offload region -fstack-arrays is
/// not honored, since the device stack is far smaller than the host one.
static fir::AllocationPolicy policyFor(const fir::PendingAllocationInfo &info,
diff --git a/flang/lib/Optimizer/Transforms/StackArrays.cpp b/flang/lib/Optimizer/Transforms/StackArrays.cpp
index 863b506a23597..0674cc2431960 100644
--- a/flang/lib/Optimizer/Transforms/StackArrays.cpp
+++ b/flang/lib/Optimizer/Transforms/StackArrays.cpp
@@ -12,6 +12,7 @@
#include "flang/Optimizer/Dialect/FIRAttr.h"
#include "flang/Optimizer/Dialect/FIRDialect.h"
#include "flang/Optimizer/Dialect/FIROps.h"
+#include "flang/Optimizer/Dialect/FIROpsSupport.h"
#include "flang/Optimizer/Dialect/FIRType.h"
#include "flang/Optimizer/Dialect/Support/FIRContext.h"
#include "flang/Optimizer/Support/AllocationPolicy.h"
More information about the flang-commits
mailing list