[flang-commits] [flang] [flang][CodeGen] Fold array/struct global initializers to dense constants (PR #226520)
John Otken via flang-commits
flang-commits at lists.llvm.org
Fri Sep 25 08:55:47 PDT 2026
https://github.com/jotken created https://github.com/llvm/llvm-project/pull/226520
Initializing a large COMMON block array with a partial DATA statement produced a long chain of fir.insert_value/fir.insert_on_range ops that GlobalOpConversion lowered one-by-one, making compilation of programs with big COMMON blocks extremely slow (issue #209393).
Add folding helpers to GlobalOpConversion that walk the initializer insert chain a single time and build an equivalent llvm.mlir constant:
* foldScalarArrayConstant: collects all elements of a scalar array (handling fir.insert_value, fir.insert_on_range, and logical fir.convert normalization) into a DenseElementsAttr, emitting a bodyless global when an explicit initializer is present.
* foldStructMemberConstant / foldGlobalStructInitializer: fold tuple (COMMON block) globals whose members are scalars or scalar arrays, emitting a small region that builds the struct with insertvalue.
Globals with only zero/undef initializers, non-scalar element types, or derived-type/named struct fields fall back to the original lowering.
This reduces compile time for the affected COMMON block cases from many seconds to well under a second while producing identical initializers.
Fixes #209393
Co-authored-by: John Otken john.otken at hpe.com
Assisted-by: Copilot and Claude Opus 4.8.
>From 8282f946068e7c884219f2cfb1f04323997a29e5 Mon Sep 17 00:00:00 2001
From: John Otken <john at otken.com>
Date: Fri, 25 Sep 2026 10:39:56 -0500
Subject: [PATCH] [flang][CodeGen] Fold array/struct global initializers to
dense constants
Initializing a large COMMON block array with a partial DATA statement
produced a long chain of fir.insert_value/fir.insert_on_range ops that
GlobalOpConversion lowered one-by-one, making compilation of programs
with big COMMON blocks extremely slow (issue #209393).
Add folding helpers to GlobalOpConversion that walk the initializer
insert chain a single time and build an equivalent llvm.mlir constant:
- foldScalarArrayConstant: collects all elements of a scalar array
(handling fir.insert_value, fir.insert_on_range, and logical
fir.convert normalization) into a DenseElementsAttr, emitting a
bodyless global when an explicit initializer is present.
- foldStructMemberConstant / foldGlobalStructInitializer: fold tuple
(COMMON block) globals whose members are scalars or scalar arrays,
emitting a small region that builds the struct with insertvalue.
Globals with only zero/undef initializers, non-scalar element types,
or derived-type/named struct fields fall back to the original lowering.
This reduces compile time for the affected COMMON block cases from many
seconds to well under a second while producing identical initializers.
Fixes #209393
Co-authored-by: John Otken john.otken at hpe.com
Assisted-by: Copilot and Claude Opus 4.8.
---
flang/lib/Optimizer/CodeGen/CodeGen.cpp | 440 ++++++++++++++++--
flang/test/Fir/convert-to-llvm.fir | 7 +-
flang/test/Fir/global-constant-fold.fir | 153 ++++++
flang/test/Fir/global-initialization.fir | 27 +-
flang/test/Fir/omp-declare-target-data.fir | 2 +-
.../Integration/common-block-large-init.f90 | 15 +
6 files changed, 576 insertions(+), 68 deletions(-)
create mode 100644 flang/test/Fir/global-constant-fold.fir
create mode 100644 flang/test/Integration/common-block-large-init.f90
diff --git a/flang/lib/Optimizer/CodeGen/CodeGen.cpp b/flang/lib/Optimizer/CodeGen/CodeGen.cpp
index 1962beb70174ba..39faee66117430 100644
--- a/flang/lib/Optimizer/CodeGen/CodeGen.cpp
+++ b/flang/lib/Optimizer/CodeGen/CodeGen.cpp
@@ -3790,6 +3790,326 @@ static inline bool attributeTypeIsCompatible(mlir::MLIRContext *ctx,
}
#endif
+/// Fold the initializer body of a scalar array `fir.global` into a single
+/// dense constant attribute, walking the region exactly once.
+///
+/// The body of a `fir.global` for a constant array is a chain of
+/// `fir.insert_value` (single elements) and `fir.insert_on_range` (contiguous
+/// runs of equal elements) operations, rooted at a `fir.undefined` or
+/// `fir.zero_bits`, and terminated by `fir.has_value`. Lowering each
+/// `fir.insert_on_range` to an `llvm.insertvalue` chain (as the fallback does)
+/// is O(N) operations that are then O(N^2) to fold back into a constant when
+/// generating LLVM IR. This is pathologically slow for large arrays that are
+/// only partially initialized (e.g. a big COMMON block array set by a DATA
+/// statement on a handful of leading elements).
+///
+/// Instead of folding each `fir.insert_on_range` individually from its `seq`
+/// operand, this walks the whole region once, from `fir.has_value` back to the
+/// root `fir.undefined`/`fir.zero_bits`, filling a single flat element vector.
+/// It returns a `DenseElementsAttr` that can be used directly as the
+/// initializer of a bodyless `llvm.mlir.global`, or failure if the region is
+/// not a foldable scalar array constant (in which case the caller falls back
+/// to the regular per-operation lowering).
+///
+/// Only arrays whose element lowers to a scalar integer or floating point type
+/// are handled here: these are the ones representable with a
+/// `DenseElementsAttr`. Initializers involving pointers, descriptors
+/// (`fir.box`) or derived-type element values are intentionally left to the
+/// fallback path, as flang has no attribute representation for descriptor
+/// values and cannot always emit constant attributes for them.
+///
+/// `arrayVal` is the array value whose defining chain is walked (typically the
+/// operand of `fir.has_value`, or the value inserted into a struct member).
+/// `llvmArrType` is its lowered `!llvm.array<...>` type. When
+/// `requireExplicitInit` is set, an array that has no explicit element
+/// initialization at all (a bare `fir.zero_bits`/`fir.undefined`) is left to
+/// the regular lowering: there is no long insertvalue chain to collapse in that
+/// case, and a zero/undef global is better emitted as such.
+static llvm::FailureOr<mlir::DenseElementsAttr>
+foldScalarArrayConstant(mlir::Value arrayVal, mlir::Type llvmArrType,
+ bool requireExplicitInit = true) {
+ auto seqTy = mlir::dyn_cast<fir::SequenceType>(arrayVal.getType());
+ if (!seqTy || seqTy.hasDynamicExtents())
+ return llvm::failure();
+ llvm::ArrayRef<int64_t> firShape = seqTy.getShape();
+ if (firShape.empty())
+ return llvm::failure();
+
+ // The element must lower to a scalar integer or floating point type for the
+ // whole array to be representable with a DenseElementsAttr. Peel exactly one
+ // llvm.array level per Fortran dimension so aggregate elements (character,
+ // complex, derived types, descriptors) are rejected.
+ mlir::Type llvmEleTy = llvmArrType;
+ for (size_t i = 0, rank = firShape.size(); i < rank; ++i) {
+ auto arrTy = mlir::dyn_cast<mlir::LLVM::LLVMArrayType>(llvmEleTy);
+ if (!arrTy)
+ return llvm::failure();
+ llvmEleTy = arrTy.getElementType();
+ }
+ if (!mlir::isa<mlir::IntegerType, mlir::FloatType>(llvmEleTy))
+ return llvm::failure();
+
+ // Total number of scalar elements and the per-dimension strides for a
+ // column-major flattening (first Fortran dimension varies fastest). This
+ // matches the coordinate order of fir.insert_value/fir.insert_on_range and
+ // makes an insert_on_range map to a contiguous [start, end] flat range.
+ int64_t total = 1;
+ llvm::SmallVector<int64_t> strides(firShape.size());
+ for (size_t i = 0; i < firShape.size(); ++i) {
+ if (firShape[i] < 0)
+ return llvm::failure();
+ strides[i] = total;
+ total *= firShape[i];
+ }
+ if (total == 0)
+ return llvm::failure();
+
+ auto flatIndex = [&](llvm::ArrayRef<int64_t> coord) -> int64_t {
+ int64_t idx = 0;
+ for (size_t i = 0; i < coord.size(); ++i)
+ idx += coord[i] * strides[i];
+ return idx;
+ };
+
+ // Fold a scalar element value being inserted into a typed attribute of the
+ // LLVM element type. Handles the fir.convert introduced for logical values,
+ // normalizing them to a canonical 0/1 exactly like ConvertOpConversion does.
+ // Returns a null attribute if the value is not a foldable scalar constant.
+ auto foldElement = [&](mlir::Value val) -> mlir::TypedAttr {
+ auto convertOp = val.getDefiningOp<fir::ConvertOp>();
+ mlir::Value src = convertOp ? convertOp.getValue() : val;
+ auto cst = src.getDefiningOp<mlir::arith::ConstantOp>();
+ if (!cst)
+ return {};
+ mlir::TypedAttr valueAttr = cst.getValue();
+ if (valueAttr.getType() == llvmEleTy)
+ return valueAttr;
+ // Only an integer<->logical fir.convert is folded here: it normalizes the
+ // operand to a canonical 0/1 (see ConvertOpConversion).
+ auto intAttr = mlir::dyn_cast<mlir::IntegerAttr>(valueAttr);
+ auto intType = mlir::dyn_cast<mlir::IntegerType>(llvmEleTy);
+ if (!intAttr || !intType || !convertOp ||
+ (!mlir::isa<fir::LogicalType>(convertOp.getType()) &&
+ !mlir::isa<fir::LogicalType>(convertOp.getValue().getType())))
+ return {};
+ return mlir::IntegerAttr::get(intType, intAttr.getValue().isZero() ? 0 : 1);
+ };
+
+ // Element attribute for each flat index. A null entry means "not yet set".
+ std::vector<mlir::Attribute> elements(total, mlir::Attribute{});
+ mlir::Attribute defaultValue; // value for elements never explicitly set
+
+ // Walk the initializer chain once, from the array value back to the root.
+ bool sawExplicitInit = false;
+ mlir::Value cur = arrayVal;
+ while (true) {
+ mlir::Operation *op = cur.getDefiningOp();
+ if (!op)
+ return llvm::failure();
+ if (mlir::isa<fir::UndefOp>(op))
+ break; // undefined root: every element must be set explicitly.
+ if (mlir::isa<fir::ZeroOp>(op)) {
+ if (auto intType = mlir::dyn_cast<mlir::IntegerType>(llvmEleTy))
+ defaultValue = mlir::IntegerAttr::get(intType, 0);
+ else
+ defaultValue = mlir::FloatAttr::get(llvmEleTy, 0.0);
+ break;
+ }
+ if (auto insert = mlir::dyn_cast<fir::InsertValueOp>(op)) {
+ llvm::SmallVector<int64_t> coord;
+ for (mlir::Attribute a : insert.getCoor()) {
+ auto intAttr = mlir::dyn_cast<mlir::IntegerAttr>(a);
+ if (!intAttr) // a named field: not a plain scalar array element.
+ return llvm::failure();
+ coord.push_back(intAttr.getInt());
+ }
+ if (coord.size() != firShape.size())
+ return llvm::failure();
+ mlir::TypedAttr eleAttr = foldElement(insert.getVal());
+ if (!eleAttr)
+ return llvm::failure();
+ int64_t idx = flatIndex(coord);
+ if (idx < 0 || idx >= total)
+ return llvm::failure();
+ // Walking backward: the first value seen for an index wins (it is the
+ // last insert in program order).
+ if (!elements[idx])
+ elements[idx] = eleAttr;
+ sawExplicitInit = true;
+ cur = insert.getAdt();
+ continue;
+ }
+ if (auto range = mlir::dyn_cast<fir::InsertOnRangeOp>(op)) {
+ mlir::TypedAttr eleAttr = foldElement(range.getVal());
+ if (!eleAttr)
+ return llvm::failure();
+ // coor holds [lb0, ub0, lb1, ub1, ...] in Fortran dimension order.
+ auto bounds = range.getCoor().getValues<int64_t>();
+ if (bounds.size() != 2 * firShape.size())
+ return llvm::failure();
+ llvm::SmallVector<int64_t> lb, ub;
+ for (size_t i = 0; i < firShape.size(); ++i) {
+ lb.push_back(bounds[2 * i]);
+ ub.push_back(bounds[2 * i + 1]);
+ }
+ int64_t start = flatIndex(lb);
+ int64_t end = flatIndex(ub);
+ if (start < 0 || end >= total || start > end)
+ return llvm::failure();
+ for (int64_t i = start; i <= end; ++i)
+ if (!elements[i])
+ elements[i] = eleAttr;
+ sawExplicitInit = true;
+ cur = range.getSeq();
+ continue;
+ }
+ // Any other producer (embox, address_of, aggregate insert, ...) cannot be
+ // folded to a scalar array constant.
+ return llvm::failure();
+ }
+
+ // A bare zero/undef array with no explicit element initialization has no
+ // insertvalue chain to collapse: leave it to the regular lowering.
+ if (requireExplicitInit && !sawExplicitInit)
+ return llvm::failure();
+
+ // Fill any element that was not explicitly initialized.
+ for (mlir::Attribute &a : elements)
+ if (!a) {
+ if (!defaultValue) // undefined element with no zero default.
+ return llvm::failure();
+ a = defaultValue;
+ }
+
+ // Build a DenseElementsAttr over a tensor shape matching the row-major LLVM
+ // array layout (the reverse of the Fortran dimension order). The element
+ // vector is already in that row-major order (first Fortran dimension
+ // fastest).
+ llvm::SmallVector<int64_t> tensorShape(firShape.rbegin(), firShape.rend());
+ auto tensorTy = mlir::RankedTensorType::get(tensorShape, llvmEleTy);
+ return mlir::DenseElementsAttr::get(tensorTy, elements);
+}
+
+/// A member of a struct/tuple `fir.global` initializer that has been folded to
+/// a single constant attribute, ready to be materialized as one
+/// `llvm.mlir.constant` and inserted into the aggregate.
+struct FoldedGlobalMember {
+ int64_t index; // position of the member in the struct
+ mlir::Type llvmType; // lowered LLVM type of the member
+ mlir::TypedAttr constant; // the folded constant (dense array or scalar)
+};
+
+/// Fold a scalar array value or a scalar value being inserted into a struct
+/// member into a single typed constant attribute. Returns a null attribute if
+/// the value is not foldable.
+static mlir::TypedAttr foldStructMemberConstant(mlir::Value val,
+ mlir::Type llvmMemberType) {
+ // Scalar array member: fold to a DenseElementsAttr. A struct member is
+ // folded even when it is entirely zero/undef, so that a data-bearing sibling
+ // member still takes the fast path (the whole struct is folded or none of
+ // it).
+ if (mlir::isa<mlir::LLVM::LLVMArrayType>(llvmMemberType)) {
+ llvm::FailureOr<mlir::DenseElementsAttr> folded =
+ foldScalarArrayConstant(val, llvmMemberType,
+ /*requireExplicitInit=*/false);
+ if (llvm::succeeded(folded))
+ return *folded;
+ return {};
+ }
+ // Scalar member: fold to a scalar constant, looking through the fir.convert
+ // introduced for logical values (normalizing them to a canonical 0/1).
+ if (mlir::isa<mlir::IntegerType, mlir::FloatType>(llvmMemberType)) {
+ auto convertOp = val.getDefiningOp<fir::ConvertOp>();
+ mlir::Value src = convertOp ? convertOp.getValue() : val;
+ auto cst = src.getDefiningOp<mlir::arith::ConstantOp>();
+ if (!cst)
+ return {};
+ mlir::TypedAttr valueAttr = cst.getValue();
+ if (valueAttr.getType() == llvmMemberType)
+ return valueAttr;
+ auto intAttr = mlir::dyn_cast<mlir::IntegerAttr>(valueAttr);
+ auto intType = mlir::dyn_cast<mlir::IntegerType>(llvmMemberType);
+ if (!intAttr || !intType || !convertOp ||
+ (!mlir::isa<fir::LogicalType>(convertOp.getType()) &&
+ !mlir::isa<fir::LogicalType>(convertOp.getValue().getType())))
+ return {};
+ return mlir::IntegerAttr::get(intType, intAttr.getValue().isZero() ? 0 : 1);
+ }
+ return {};
+}
+
+/// Fold the initializer body of a struct/tuple `fir.global` whose members are
+/// scalar arrays or scalars (as generated for COMMON blocks) into a list of
+/// per-member constant attributes. The caller emits a small region that
+/// materializes each member with one `llvm.mlir.constant` and inserts it into
+/// the aggregate: a single struct-level `llvm.mlir.constant` cannot represent a
+/// struct containing an array.
+///
+/// Returns failure if the global is not such a struct, or if any member is not
+/// a foldable scalar array/scalar (e.g. it involves a descriptor, pointer or
+/// nested aggregate), in which case the caller falls back to the regular
+/// per-operation lowering.
+static llvm::FailureOr<llvm::SmallVector<FoldedGlobalMember>>
+foldGlobalStructInitializer(fir::GlobalOp global, mlir::Type llvmType) {
+ auto structTy = mlir::dyn_cast<mlir::LLVM::LLVMStructType>(llvmType);
+ if (!structTy || structTy.isOpaque())
+ return llvm::failure();
+ llvm::ArrayRef<mlir::Type> body = structTy.getBody();
+ if (body.empty())
+ return llvm::failure();
+
+ mlir::Region ®ion = global.getRegion();
+ if (region.empty())
+ return llvm::failure();
+ auto hasValue =
+ mlir::dyn_cast<fir::HasValueOp>(region.front().getTerminator());
+ if (!hasValue)
+ return llvm::failure();
+
+ // Walk the struct-level fir.insert_value chain once, back to the root
+ // fir.undefined. Each insert sets one member by its integer field index.
+ llvm::SmallVector<mlir::TypedAttr> members(body.size(), mlir::TypedAttr{});
+ mlir::Value cur = hasValue.getResval();
+ while (true) {
+ mlir::Operation *op = cur.getDefiningOp();
+ if (!op)
+ return llvm::failure();
+ if (mlir::isa<fir::UndefOp, fir::ZeroOp>(op))
+ break;
+ auto insert = mlir::dyn_cast<fir::InsertValueOp>(op);
+ if (!insert)
+ return llvm::failure();
+ // A struct field is addressed by a single integer index.
+ if (insert.getCoor().size() != 1)
+ return llvm::failure();
+ auto idxAttr = mlir::dyn_cast<mlir::IntegerAttr>(insert.getCoor()[0]);
+ if (!idxAttr)
+ return llvm::failure();
+ int64_t field = idxAttr.getInt();
+ if (field < 0 || field >= static_cast<int64_t>(body.size()))
+ return llvm::failure();
+ // Walking backward: the first value seen for a field wins.
+ if (!members[field]) {
+ mlir::TypedAttr memberAttr =
+ foldStructMemberConstant(insert.getVal(), body[field]);
+ if (!memberAttr)
+ return llvm::failure();
+ members[field] = memberAttr;
+ }
+ cur = insert.getAdt();
+ }
+
+ // Every member must have been explicitly set: there is no zero default here
+ // because a struct member may itself need a specific per-element value.
+ llvm::SmallVector<FoldedGlobalMember> folded;
+ for (auto [i, memberAttr] : llvm::enumerate(members)) {
+ if (!memberAttr)
+ return llvm::failure();
+ folded.push_back({static_cast<int64_t>(i), body[i], memberAttr});
+ }
+ return folded;
+}
+
/// Lower `fir.global` operation to `llvm.global` operation.
/// `fir.insert_on_range` operations are replaced with constant dense attribute
/// if they are applied on the full range.
@@ -3818,6 +4138,31 @@ struct GlobalOpConversion : public fir::FIROpConversion<fir::GlobalOp> {
tyAttr = this->lowerTy().convertBoxTypeAsStruct(boxType);
auto loc = global.getLoc();
mlir::Attribute initAttr = global.getInitVal().value_or(mlir::Attribute());
+ // Try to fold the initializer body into constant attributes so a compact
+ // llvm.mlir.global can be emitted, avoiding the O(N^2) materialization and
+ // folding of long insertvalue chains for large, partially initialized
+ // arrays. Two forms are handled:
+ // - a top-level scalar array: a bodyless global with a dense init;
+ // - a struct/tuple of scalar arrays/scalars (COMMON blocks): a region
+ // that materializes each member with a single llvm.mlir.constant.
+ bool foldedInit = false;
+ llvm::SmallVector<FoldedGlobalMember> foldedMembers;
+ if (!initAttr && !global.getRegion().empty()) {
+ if (auto hasValue = mlir::dyn_cast<fir::HasValueOp>(
+ global.getRegion().front().getTerminator())) {
+ llvm::FailureOr<mlir::DenseElementsAttr> foldedArray =
+ foldScalarArrayConstant(hasValue.getResval(), tyAttr);
+ if (llvm::succeeded(foldedArray)) {
+ initAttr = *foldedArray;
+ foldedInit = true;
+ } else {
+ llvm::FailureOr<llvm::SmallVector<FoldedGlobalMember>> foldedStruct =
+ foldGlobalStructInitializer(global, tyAttr);
+ if (llvm::succeeded(foldedStruct))
+ foldedMembers = std::move(*foldedStruct);
+ }
+ }
+ }
assert(attributeTypeIsCompatible(global.getContext(), initAttr, tyAttr));
auto linkage = convertLinkage(global.getLinkage());
auto isConst = global.getConstant().has_value();
@@ -3867,47 +4212,64 @@ struct GlobalOpConversion : public fir::FIROpConversion<fir::GlobalOp> {
g->setDiscardableAttrs(global->getDiscardableAttrDictionary());
auto &gr = g.getInitializerRegion();
- rewriter.inlineRegionBefore(global.getRegion(), gr, gr.end());
- if (!gr.empty()) {
- // Replace insert_on_range with a constant dense attribute if the
- // initialization is on the full range.
- auto insertOnRangeOps = gr.front().getOps<fir::InsertOnRangeOp>();
- for (auto insertOp : insertOnRangeOps) {
- if (!insertOp.isFullRange())
- continue;
- // The dense attribute must use the converted element type of the
- // array, not the type of whatever constant feeds the insertion.
- mlir::Type elementType = convertType(insertOp.getType().getEleTy());
- mlir::Value val = insertOp.getVal();
- // Logical constants reach the insertion through a `fir.convert`.
- auto convertOp = val.getDefiningOp<fir::ConvertOp>();
- if (convertOp)
- val = convertOp.getValue();
- auto constant = val.getDefiningOp<mlir::arith::ConstantOp>();
- if (!constant)
- continue;
- mlir::TypedAttr valueAttr = constant.getValue();
- if (valueAttr.getType() != elementType) {
- // Looking through the `fir.convert` leaves the constant with the
- // source type. Only an integer<->logical conversion is folded here:
- // it normalizes any integer operand to a canonical 0/1, see
- // ConvertOpConversion. Any other mismatching conversion is left to
- // the regular lowering.
- auto intAttr = mlir::dyn_cast<mlir::IntegerAttr>(valueAttr);
- auto intType = mlir::dyn_cast<mlir::IntegerType>(elementType);
- if (!intAttr || !intType || !convertOp ||
- (!mlir::isa<fir::LogicalType>(convertOp.getType()) &&
- !mlir::isa<fir::LogicalType>(convertOp.getValue().getType())))
+ if (!foldedMembers.empty()) {
+ // Materialize the folded struct/tuple members: one llvm.mlir.constant per
+ // member inserted into an undef aggregate, then returned. A single
+ // struct-level llvm.mlir.constant cannot represent a struct containing an
+ // array, so the aggregate is assembled with llvm.insertvalue.
+ mlir::Block *block = rewriter.createBlock(&gr);
+ rewriter.setInsertionPointToStart(block);
+ mlir::Value agg = mlir::LLVM::UndefOp::create(rewriter, loc, tyAttr);
+ for (const FoldedGlobalMember &member : foldedMembers) {
+ mlir::Value cst = mlir::LLVM::ConstantOp::create(
+ rewriter, loc, member.llvmType, member.constant);
+ agg = mlir::LLVM::InsertValueOp::create(rewriter, loc, agg, cst,
+ member.index);
+ }
+ mlir::LLVM::ReturnOp::create(rewriter, loc, agg);
+ } else if (!foldedInit) {
+ rewriter.inlineRegionBefore(global.getRegion(), gr, gr.end());
+ if (!gr.empty()) {
+ // Replace insert_on_range with a constant dense attribute if the
+ // initialization is on the full range.
+ auto insertOnRangeOps = gr.front().getOps<fir::InsertOnRangeOp>();
+ for (auto insertOp : insertOnRangeOps) {
+ if (!insertOp.isFullRange())
+ continue;
+ // The dense attribute must use the converted element type of the
+ // array, not the type of whatever constant feeds the insertion.
+ mlir::Type elementType = convertType(insertOp.getType().getEleTy());
+ mlir::Value val = insertOp.getVal();
+ // Logical constants reach the insertion through a `fir.convert`.
+ auto convertOp = val.getDefiningOp<fir::ConvertOp>();
+ if (convertOp)
+ val = convertOp.getValue();
+ auto constant = val.getDefiningOp<mlir::arith::ConstantOp>();
+ if (!constant)
continue;
- valueAttr = mlir::IntegerAttr::get(
- intType, intAttr.getValue().isZero() ? 0 : 1);
+ mlir::TypedAttr valueAttr = constant.getValue();
+ if (valueAttr.getType() != elementType) {
+ // Looking through the `fir.convert` leaves the constant with the
+ // source type. Only an integer<->logical conversion is folded here:
+ // it normalizes any integer operand to a canonical 0/1, see
+ // ConvertOpConversion. Any other mismatching conversion is left to
+ // the regular lowering.
+ auto intAttr = mlir::dyn_cast<mlir::IntegerAttr>(valueAttr);
+ auto intType = mlir::dyn_cast<mlir::IntegerType>(elementType);
+ if (!intAttr || !intType || !convertOp ||
+ (!mlir::isa<fir::LogicalType>(convertOp.getType()) &&
+ !mlir::isa<fir::LogicalType>(convertOp.getValue().getType())))
+ continue;
+ valueAttr = mlir::IntegerAttr::get(
+ intType, intAttr.getValue().isZero() ? 0 : 1);
+ }
+ auto vecType =
+ mlir::VectorType::get(insertOp.getType().getShape(), elementType);
+ auto denseAttr = mlir::DenseElementsAttr::get(vecType, valueAttr);
+ rewriter.setInsertionPointAfter(insertOp);
+ rewriter.replaceOpWithNewOp<mlir::arith::ConstantOp>(
+ insertOp, convertType(insertOp.getType()), denseAttr);
}
- auto vecType =
- mlir::VectorType::get(insertOp.getType().getShape(), elementType);
- auto denseAttr = mlir::DenseElementsAttr::get(vecType, valueAttr);
- rewriter.setInsertionPointAfter(insertOp);
- rewriter.replaceOpWithNewOp<mlir::arith::ConstantOp>(
- insertOp, convertType(insertOp.getType()), denseAttr);
}
}
diff --git a/flang/test/Fir/convert-to-llvm.fir b/flang/test/Fir/convert-to-llvm.fir
index d00aae198bf138..7261b4df690b23 100644
--- a/flang/test/Fir/convert-to-llvm.fir
+++ b/flang/test/Fir/convert-to-llvm.fir
@@ -107,13 +107,10 @@ fir.global internal @_QEmultiarray : !fir.array<32x32xi32> {
fir.has_value %2 : !fir.array<32x32xi32>
}
-// CHECK: llvm.mlir.global internal @_QEmultiarray()
+// CHECK: llvm.mlir.global internal @_QEmultiarray(dense<1> : tensor<32x32xi32>)
// GENERIC-SAME: {addr_space = 0 : i32}
// AMDGPU-SAME: {addr_space = 1 : i32}
-// CHECK-SAME: : !llvm.array<32 x array<32 x i32>> {
-// CHECK: %[[CST:.*]] = llvm.mlir.constant(dense<1> : vector<32x32xi32>) : !llvm.array<32 x array<32 x i32>>
-// CHECK: llvm.return %[[CST]] : !llvm.array<32 x array<32 x i32>>
-// CHECK: }
+// CHECK-SAME: : !llvm.array<32 x array<32 x i32>>
// -----
diff --git a/flang/test/Fir/global-constant-fold.fir b/flang/test/Fir/global-constant-fold.fir
new file mode 100644
index 00000000000000..17b28b31f9a22b
--- /dev/null
+++ b/flang/test/Fir/global-constant-fold.fir
@@ -0,0 +1,153 @@
+// Test the single-pass folding of constant array `fir.global` initializer
+// bodies into dense constants. A partially initialized large array must be
+// lowered to a compact constant instead of a long `llvm.insertvalue` chain
+// (which is pathologically slow to materialize and fold back, see
+// https://github.com/llvm/llvm-project/issues/209393).
+
+// RUN: fir-opt --split-input-file --fir-to-llvm-ir %s | FileCheck %s
+
+// A 1D array initialized with a couple of leading elements and a trailing
+// range folds to a bodyless dense global.
+fir.global @oned : !fir.array<8xi32> {
+ %c7 = arith.constant 7 : i32
+ %c77 = arith.constant 77 : i32
+ %c0 = arith.constant 0 : i32
+ %0 = fir.undefined !fir.array<8xi32>
+ %1 = fir.insert_value %0, %c7, [0 : index] : (!fir.array<8xi32>, i32) -> !fir.array<8xi32>
+ %2 = fir.insert_value %1, %c77, [1 : index] : (!fir.array<8xi32>, i32) -> !fir.array<8xi32>
+ %3 = fir.insert_on_range %2, %c0 from (2) to (7) : (!fir.array<8xi32>, i32) -> !fir.array<8xi32>
+ fir.has_value %3 : !fir.array<8xi32>
+}
+
+// CHECK: llvm.mlir.global external @oned(dense<[7, 77, 0, 0, 0, 0, 0, 0]> : tensor<8xi32>) {{.*}} : !llvm.array<8 x i32>
+// CHECK-NOT: llvm.insertvalue
+
+// -----
+
+// A 2D array with a mix of a single-element insert and a range insert folds to
+// a dense global. The tensor shape is the reverse of the Fortran dimensions,
+// and the element order is column-major (first Fortran dimension fastest).
+fir.global @twod : !fir.array<3x2xi32> {
+ %c9 = arith.constant 9 : i32
+ %c0 = arith.constant 0 : i32
+ %0 = fir.undefined !fir.array<3x2xi32>
+ %1 = fir.insert_value %0, %c9, [0 : index, 0 : index] : (!fir.array<3x2xi32>, i32) -> !fir.array<3x2xi32>
+ %2 = fir.insert_on_range %1, %c0 from (1,0) to (2,1) : (!fir.array<3x2xi32>, i32) -> !fir.array<3x2xi32>
+ fir.has_value %2 : !fir.array<3x2xi32>
+}
+
+// CHECK: llvm.mlir.global external @twod(dense<{{\[}}[9, 0, 0], [0, 0, 0]]> : tensor<2x3xi32>) {{.*}} : !llvm.array<2 x array<3 x i32>>
+// CHECK-NOT: llvm.insertvalue
+
+// -----
+
+// A 3D array fully initialized by a single range folds to a dense splat.
+fir.global @threed : !fir.array<2x2x2xi32> {
+ %c5 = arith.constant 5 : i32
+ %0 = fir.undefined !fir.array<2x2x2xi32>
+ %1 = fir.insert_on_range %0, %c5 from (0,0,0) to (1,1,1) : (!fir.array<2x2x2xi32>, i32) -> !fir.array<2x2x2xi32>
+ fir.has_value %1 : !fir.array<2x2x2xi32>
+}
+
+// CHECK: llvm.mlir.global external @threed(dense<5> : tensor<2x2x2xi32>) {{.*}} : !llvm.array<2 x array<2 x array<2 x i32>>>
+// CHECK-NOT: llvm.insertvalue
+
+// -----
+
+// A real (floating point) array folds as well.
+fir.global @reals : !fir.array<4xf32> {
+ %c1 = arith.constant 1.5 : f32
+ %c0 = arith.constant 0.0 : f32
+ %0 = fir.undefined !fir.array<4xf32>
+ %1 = fir.insert_value %0, %c1, [0 : index] : (!fir.array<4xf32>, f32) -> !fir.array<4xf32>
+ %2 = fir.insert_on_range %1, %c0 from (1) to (3) : (!fir.array<4xf32>, f32) -> !fir.array<4xf32>
+ fir.has_value %2 : !fir.array<4xf32>
+}
+
+// CHECK: llvm.mlir.global external @reals(dense<[1.500000e+00, 0.000000e+00, 0.000000e+00, 0.000000e+00]> : tensor<4xf32>) {{.*}} : !llvm.array<4 x f32>
+// CHECK-NOT: llvm.insertvalue
+
+// -----
+
+// A logical array reaches the insertion through a fir.convert; the fold
+// normalizes the value to a canonical 0/1 of the lowered element type.
+fir.global @logicals : !fir.array<4x!fir.logical<4>> {
+ %true = arith.constant true
+ %0 = fir.undefined !fir.array<4x!fir.logical<4>>
+ %1 = fir.convert %true : (i1) -> !fir.logical<4>
+ %2 = fir.insert_on_range %0, %1 from (0) to (3) : (!fir.array<4x!fir.logical<4>>, !fir.logical<4>) -> !fir.array<4x!fir.logical<4>>
+ fir.has_value %2 : !fir.array<4x!fir.logical<4>>
+}
+
+// CHECK: llvm.mlir.global external @logicals(dense<1> : tensor<4xi32>) {{.*}} : !llvm.array<4 x i32>
+// CHECK-NOT: llvm.insertvalue
+
+// -----
+
+// A COMMON block lowers to a single-element tuple wrapping the array. A single
+// struct-level llvm.mlir.constant cannot represent a struct containing an
+// array, so the array member is materialized with one llvm.mlir.constant and
+// inserted into the aggregate.
+fir.global @common1 {alignment = 4 : i64} : tuple<!fir.array<8xi32>> {
+ %c7 = arith.constant 7 : i32
+ %c77 = arith.constant 77 : i32
+ %c0 = arith.constant 0 : i32
+ %0 = fir.zero_bits tuple<!fir.array<8xi32>>
+ %1 = fir.undefined !fir.array<8xi32>
+ %2 = fir.insert_value %1, %c7, [0 : index] : (!fir.array<8xi32>, i32) -> !fir.array<8xi32>
+ %3 = fir.insert_value %2, %c77, [1 : index] : (!fir.array<8xi32>, i32) -> !fir.array<8xi32>
+ %4 = fir.insert_on_range %3, %c0 from (2) to (7) : (!fir.array<8xi32>, i32) -> !fir.array<8xi32>
+ %5 = fir.insert_value %0, %4, [0 : index] : (tuple<!fir.array<8xi32>>, !fir.array<8xi32>) -> tuple<!fir.array<8xi32>>
+ fir.has_value %5 : tuple<!fir.array<8xi32>>
+}
+
+// CHECK: llvm.mlir.global external @common1() {{.*}} : !llvm.struct<(array<8 x i32>)> {
+// CHECK: %[[UNDEF:.*]] = llvm.mlir.undef : !llvm.struct<(array<8 x i32>)>
+// CHECK: %[[CST:.*]] = llvm.mlir.constant(dense<[7, 77, 0, 0, 0, 0, 0, 0]> : tensor<8xi32>) : !llvm.array<8 x i32>
+// CHECK: %[[INS:.*]] = llvm.insertvalue %[[CST]], %[[UNDEF]][0] : !llvm.struct<(array<8 x i32>)>
+// CHECK: llvm.return %[[INS]] : !llvm.struct<(array<8 x i32>)>
+// CHECK: }
+
+// -----
+
+// A COMMON block with several members: each member folds to one constant and
+// is inserted at its field index.
+fir.global @common2 : tuple<!fir.array<4xi32>, !fir.array<2xi32>> {
+ %c1 = arith.constant 1 : i32
+ %c0 = arith.constant 0 : i32
+ %c3 = arith.constant 3 : i32
+ %0 = fir.zero_bits tuple<!fir.array<4xi32>, !fir.array<2xi32>>
+ %1 = fir.undefined !fir.array<4xi32>
+ %2 = fir.insert_on_range %1, %c1 from (0) to (3) : (!fir.array<4xi32>, i32) -> !fir.array<4xi32>
+ %3 = fir.insert_value %0, %2, [0 : index] : (tuple<!fir.array<4xi32>, !fir.array<2xi32>>, !fir.array<4xi32>) -> tuple<!fir.array<4xi32>, !fir.array<2xi32>>
+ %4 = fir.undefined !fir.array<2xi32>
+ %5 = fir.insert_value %4, %c3, [0 : index] : (!fir.array<2xi32>, i32) -> !fir.array<2xi32>
+ %6 = fir.insert_value %5, %c0, [1 : index] : (!fir.array<2xi32>, i32) -> !fir.array<2xi32>
+ %7 = fir.insert_value %3, %6, [1 : index] : (tuple<!fir.array<4xi32>, !fir.array<2xi32>>, !fir.array<2xi32>) -> tuple<!fir.array<4xi32>, !fir.array<2xi32>>
+ fir.has_value %7 : tuple<!fir.array<4xi32>, !fir.array<2xi32>>
+}
+
+// CHECK: llvm.mlir.global external @common2() {{.*}} : !llvm.struct<(array<4 x i32>, array<2 x i32>)> {
+// CHECK: %[[UNDEF:.*]] = llvm.mlir.undef : !llvm.struct<(array<4 x i32>, array<2 x i32>)>
+// CHECK: %[[CST0:.*]] = llvm.mlir.constant(dense<1> : tensor<4xi32>) : !llvm.array<4 x i32>
+// CHECK: %[[INS0:.*]] = llvm.insertvalue %[[CST0]], %[[UNDEF]][0] : !llvm.struct<(array<4 x i32>, array<2 x i32>)>
+// CHECK: %[[CST1:.*]] = llvm.mlir.constant(dense<[3, 0]> : tensor<2xi32>) : !llvm.array<2 x i32>
+// CHECK: %[[INS1:.*]] = llvm.insertvalue %[[CST1]], %[[INS0]][1] : !llvm.struct<(array<4 x i32>, array<2 x i32>)>
+// CHECK: llvm.return %[[INS1]] : !llvm.struct<(array<4 x i32>, array<2 x i32>)>
+// CHECK: }
+
+// -----
+
+// An array whose leading elements are left undefined (no zero base) has no
+// constant to fold to for those elements: it must stay on the regular
+// per-operation lowering rather than being force-folded.
+fir.global @partialundef : !fir.array<8xi32> {
+ %c1 = arith.constant 1 : i32
+ %0 = fir.undefined !fir.array<8xi32>
+ %1 = fir.insert_on_range %0, %c1 from (5) to (7) : (!fir.array<8xi32>, i32) -> !fir.array<8xi32>
+ fir.has_value %1 : !fir.array<8xi32>
+}
+
+// CHECK: llvm.mlir.global external @partialundef() {{.*}} : !llvm.array<8 x i32> {
+// CHECK: llvm.insertvalue
+// CHECK: }
diff --git a/flang/test/Fir/global-initialization.fir b/flang/test/Fir/global-initialization.fir
index a683b46de2541a..73cba152d0c202 100644
--- a/flang/test/Fir/global-initialization.fir
+++ b/flang/test/Fir/global-initialization.fir
@@ -8,12 +8,7 @@ fir.global internal @_QEmask : !fir.array<32xi32> {
fir.has_value %2 : !fir.array<32xi32>
}
-// CHECK: llvm.mlir.global internal @_QEmask() {addr_space = 0 : i32} : !llvm.array<32 x i32> {
-// CHECK: [[VAL0:%.*]] = llvm.mlir.constant(1 : i32) : i32
-// CHECK: [[VAL1:%.*]] = llvm.mlir.undef : !llvm.array<32 x i32>
-// CHECK: [[VAL2:%.*]] = llvm.mlir.constant(dense<1> : vector<32xi32>) : !llvm.array<32 x i32>
-// CHECK: llvm.return [[VAL2]] : !llvm.array<32 x i32>
-// CHECK: }
+// CHECK: llvm.mlir.global internal @_QEmask(dense<1> : tensor<32xi32>) {addr_space = 0 : i32} : !llvm.array<32 x i32>
fir.global internal @_QEmultiarray : !fir.array<32x32xi32> {
%c0_i32 = arith.constant 1 : i32
@@ -22,12 +17,7 @@ fir.global internal @_QEmultiarray : !fir.array<32x32xi32> {
fir.has_value %2 : !fir.array<32x32xi32>
}
-// CHECK: llvm.mlir.global internal @_QEmultiarray() {addr_space = 0 : i32} : !llvm.array<32 x array<32 x i32>> {
-// CHECK: [[VAL0:%.*]] = llvm.mlir.constant(1 : i32) : i32
-// CHECK: [[VAL1:%.*]] = llvm.mlir.undef : !llvm.array<32 x array<32 x i32>>
-// CHECK: [[VAL2:%.*]] = llvm.mlir.constant(dense<1> : vector<32x32xi32>) : !llvm.array<32 x array<32 x i32>>
-// CHECK: llvm.return [[VAL2]] : !llvm.array<32 x array<32 x i32>>
-// CHECK: }
+// CHECK: llvm.mlir.global internal @_QEmultiarray(dense<1> : tensor<32x32xi32>) {addr_space = 0 : i32} : !llvm.array<32 x array<32 x i32>>
fir.global internal @_QEmasklogical : !fir.array<32768x!fir.logical<4>> {
%true = arith.constant true
@@ -37,13 +27,7 @@ fir.global internal @_QEmasklogical : !fir.array<32768x!fir.logical<4>> {
fir.has_value %2 : !fir.array<32768x!fir.logical<4>>
}
-// CHECK: llvm.mlir.global internal @_QEmasklogical() {addr_space = 0 : i32} : !llvm.array<32768 x i32> {
-// CHECK: [[VAL0:%.*]] = llvm.mlir.constant(true) : i1
-// CHECK: [[VAL1:%.*]] = llvm.mlir.undef : !llvm.array<32768 x i32>
-// CHECK: [[VAL2:%.*]] = llvm.mlir.constant(1 : i32) : i32
-// CHECK: [[VAL3:%.*]] = llvm.mlir.constant(dense<1> : vector<32768xi32>) : !llvm.array<32768 x i32>
-// CHECK: llvm.return [[VAL3]] : !llvm.array<32768 x i32>
-// CHECK: }
+// CHECK: llvm.mlir.global internal @_QEmasklogical(dense<1> : tensor<32768xi32>) {addr_space = 0 : i32} : !llvm.array<32768 x i32>
// A logical conversion normalizes any integer operand to a canonical 0/1, not
// just an `i1` one, so the full-range fold must apply here as well.
@@ -55,10 +39,7 @@ fir.global internal @_QEmasklogicalkind : !fir.array<8x!fir.logical<8>> {
fir.has_value %2 : !fir.array<8x!fir.logical<8>>
}
-// CHECK: llvm.mlir.global internal @_QEmasklogicalkind() {addr_space = 0 : i32} : !llvm.array<8 x i64> {
-// CHECK: [[VAL2:%.*]] = llvm.mlir.constant(dense<1> : vector<8xi64>) : !llvm.array<8 x i64>
-// CHECK: llvm.return [[VAL2]] : !llvm.array<8 x i64>
-// CHECK: }
+// CHECK: llvm.mlir.global internal @_QEmasklogicalkind(dense<1> : tensor<8xi64>) {addr_space = 0 : i32} : !llvm.array<8 x i64>
fir.global internal @_QElookforme : !fir.type<_QTt{i:!fir.array<500xi32>,j:!fir.array<500xi32>}> {
%c2_i32 = arith.constant 2 : i32
diff --git a/flang/test/Fir/omp-declare-target-data.fir b/flang/test/Fir/omp-declare-target-data.fir
index 539a7acf267f2d..9ad41f08e8b4a1 100644
--- a/flang/test/Fir/omp-declare-target-data.fir
+++ b/flang/test/Fir/omp-declare-target-data.fir
@@ -5,7 +5,7 @@ module attributes {omp.is_target_device = false} {
// CHECK: llvm.mlir.global external @_QMtest_0Earray_1d(dense<[1, 2, 3]> : tensor<3xi32>) {{{.*}}omp.declare_target = #omp.declaretarget<device_type = any, capture_clause = link>{{.*}}} : !llvm.array<3 x i32>
fir.global @_QMtest_0Earray_1d(dense<[1, 2, 3]> : tensor<3xi32>) {omp.declare_target = #omp.declaretarget<device_type = any, capture_clause = link>} : !fir.array<3xi32>
- // CHECK: llvm.mlir.global external @_QMtest_0Earray_2d() {{{.*}}omp.declare_target = #omp.declaretarget<device_type = any, capture_clause = link>{{.*}}} : !llvm.array<2 x array<2 x i32>>
+ // CHECK: llvm.mlir.global external @_QMtest_0Earray_2d(dense<{{\[}}[1, 2], [3, 4]]> : tensor<2x2xi32>) {{{.*}}omp.declare_target = #omp.declaretarget<device_type = any, capture_clause = link>{{.*}}} : !llvm.array<2 x array<2 x i32>>
fir.global @_QMtest_0Earray_2d {omp.declare_target = #omp.declaretarget<device_type = any, capture_clause = link>} : !fir.array<2x2xi32> {
%0 = fir.undefined !fir.array<2x2xi32>
%c1_i32 = arith.constant 1 : i32
diff --git a/flang/test/Integration/common-block-large-init.f90 b/flang/test/Integration/common-block-large-init.f90
new file mode 100644
index 00000000000000..4fd563707aff10
--- /dev/null
+++ b/flang/test/Integration/common-block-large-init.f90
@@ -0,0 +1,15 @@
+! Verify that a large COMMON block array initialized on only a few elements by a
+! DATA statement lowers to a compact constant global instead of an enormous
+! `insertvalue` chain (which is pathologically slow to generate, see
+! https://github.com/llvm/llvm-project/issues/209393).
+
+! RUN: %flang_fc1 -emit-llvm %s -o - | FileCheck %s
+
+block data
+ integer i(55000)
+ common /b/ i
+ data (i(j), j = 1, 2) / 7, 77 /
+end block data
+
+! CHECK: @b_ = {{.*}}global { [55000 x i32] } { [55000 x i32] [i32 7, i32 77, i32 0
+! CHECK-NOT: insertvalue
More information about the flang-commits
mailing list