[flang-commits] [flang] [flang][CodeGen] Fold array/struct global initializers to dense constants (PR #226520)

John Otken via flang-commits flang-commits at lists.llvm.org
Fri Sep 25 08:55:47 PDT 2026


https://github.com/jotken created https://github.com/llvm/llvm-project/pull/226520

Initializing a large COMMON block array with a partial DATA statement produced a long chain of fir.insert_value/fir.insert_on_range ops that GlobalOpConversion lowered one-by-one, making compilation of programs with big COMMON blocks extremely slow (issue #209393).

Add folding helpers to GlobalOpConversion that walk the initializer insert chain a single time and build an equivalent llvm.mlir constant:

* foldScalarArrayConstant: collects all elements of a scalar array (handling fir.insert_value, fir.insert_on_range, and logical fir.convert normalization) into a DenseElementsAttr, emitting a bodyless global when an explicit initializer is present.

* foldStructMemberConstant / foldGlobalStructInitializer: fold tuple (COMMON block) globals whose members are scalars or scalar arrays, emitting a small region that builds the struct with insertvalue.

Globals with only zero/undef initializers, non-scalar element types, or derived-type/named struct fields fall back to the original lowering.

This reduces compile time for the affected COMMON block cases from many seconds to well under a second while producing identical initializers.

Fixes #209393

Co-authored-by: John Otken john.otken at hpe.com
Assisted-by: Copilot and Claude Opus 4.8.

>From 8282f946068e7c884219f2cfb1f04323997a29e5 Mon Sep 17 00:00:00 2001
From: John Otken <john at otken.com>
Date: Fri, 25 Sep 2026 10:39:56 -0500
Subject: [PATCH] [flang][CodeGen] Fold array/struct global initializers to
 dense constants

Initializing a large COMMON block array with a partial DATA statement
produced a long chain of fir.insert_value/fir.insert_on_range ops that
GlobalOpConversion lowered one-by-one, making compilation of programs
with big COMMON blocks extremely slow (issue #209393).

Add folding helpers to GlobalOpConversion that walk the initializer
insert chain a single time and build an equivalent llvm.mlir constant:

- foldScalarArrayConstant: collects all elements of a scalar array
  (handling fir.insert_value, fir.insert_on_range, and logical
  fir.convert normalization) into a DenseElementsAttr, emitting a
  bodyless global when an explicit initializer is present.

- foldStructMemberConstant / foldGlobalStructInitializer: fold tuple
  (COMMON block) globals whose members are scalars or scalar arrays,
  emitting a small region that builds the struct with insertvalue.

Globals with only zero/undef initializers, non-scalar element types,
or derived-type/named struct fields fall back to the original lowering.

This reduces compile time for the affected COMMON block cases from many
seconds to well under a second while producing identical initializers.

Fixes #209393

Co-authored-by: John Otken john.otken at hpe.com
Assisted-by: Copilot and Claude Opus 4.8.
---
 flang/lib/Optimizer/CodeGen/CodeGen.cpp       | 440 ++++++++++++++++--
 flang/test/Fir/convert-to-llvm.fir            |   7 +-
 flang/test/Fir/global-constant-fold.fir       | 153 ++++++
 flang/test/Fir/global-initialization.fir      |  27 +-
 flang/test/Fir/omp-declare-target-data.fir    |   2 +-
 .../Integration/common-block-large-init.f90   |  15 +
 6 files changed, 576 insertions(+), 68 deletions(-)
 create mode 100644 flang/test/Fir/global-constant-fold.fir
 create mode 100644 flang/test/Integration/common-block-large-init.f90

diff --git a/flang/lib/Optimizer/CodeGen/CodeGen.cpp b/flang/lib/Optimizer/CodeGen/CodeGen.cpp
index 1962beb70174ba..39faee66117430 100644
--- a/flang/lib/Optimizer/CodeGen/CodeGen.cpp
+++ b/flang/lib/Optimizer/CodeGen/CodeGen.cpp
@@ -3790,6 +3790,326 @@ static inline bool attributeTypeIsCompatible(mlir::MLIRContext *ctx,
 }
 #endif
 
+/// Fold the initializer body of a scalar array `fir.global` into a single
+/// dense constant attribute, walking the region exactly once.
+///
+/// The body of a `fir.global` for a constant array is a chain of
+/// `fir.insert_value` (single elements) and `fir.insert_on_range` (contiguous
+/// runs of equal elements) operations, rooted at a `fir.undefined` or
+/// `fir.zero_bits`, and terminated by `fir.has_value`. Lowering each
+/// `fir.insert_on_range` to an `llvm.insertvalue` chain (as the fallback does)
+/// is O(N) operations that are then O(N^2) to fold back into a constant when
+/// generating LLVM IR. This is pathologically slow for large arrays that are
+/// only partially initialized (e.g. a big COMMON block array set by a DATA
+/// statement on a handful of leading elements).
+///
+/// Instead of folding each `fir.insert_on_range` individually from its `seq`
+/// operand, this walks the whole region once, from `fir.has_value` back to the
+/// root `fir.undefined`/`fir.zero_bits`, filling a single flat element vector.
+/// It returns a `DenseElementsAttr` that can be used directly as the
+/// initializer of a bodyless `llvm.mlir.global`, or failure if the region is
+/// not a foldable scalar array constant (in which case the caller falls back
+/// to the regular per-operation lowering).
+///
+/// Only arrays whose element lowers to a scalar integer or floating point type
+/// are handled here: these are the ones representable with a
+/// `DenseElementsAttr`. Initializers involving pointers, descriptors
+/// (`fir.box`) or derived-type element values are intentionally left to the
+/// fallback path, as flang has no attribute representation for descriptor
+/// values and cannot always emit constant attributes for them.
+///
+/// `arrayVal` is the array value whose defining chain is walked (typically the
+/// operand of `fir.has_value`, or the value inserted into a struct member).
+/// `llvmArrType` is its lowered `!llvm.array<...>` type. When
+/// `requireExplicitInit` is set, an array that has no explicit element
+/// initialization at all (a bare `fir.zero_bits`/`fir.undefined`) is left to
+/// the regular lowering: there is no long insertvalue chain to collapse in that
+/// case, and a zero/undef global is better emitted as such.
+static llvm::FailureOr<mlir::DenseElementsAttr>
+foldScalarArrayConstant(mlir::Value arrayVal, mlir::Type llvmArrType,
+                        bool requireExplicitInit = true) {
+  auto seqTy = mlir::dyn_cast<fir::SequenceType>(arrayVal.getType());
+  if (!seqTy || seqTy.hasDynamicExtents())
+    return llvm::failure();
+  llvm::ArrayRef<int64_t> firShape = seqTy.getShape();
+  if (firShape.empty())
+    return llvm::failure();
+
+  // The element must lower to a scalar integer or floating point type for the
+  // whole array to be representable with a DenseElementsAttr. Peel exactly one
+  // llvm.array level per Fortran dimension so aggregate elements (character,
+  // complex, derived types, descriptors) are rejected.
+  mlir::Type llvmEleTy = llvmArrType;
+  for (size_t i = 0, rank = firShape.size(); i < rank; ++i) {
+    auto arrTy = mlir::dyn_cast<mlir::LLVM::LLVMArrayType>(llvmEleTy);
+    if (!arrTy)
+      return llvm::failure();
+    llvmEleTy = arrTy.getElementType();
+  }
+  if (!mlir::isa<mlir::IntegerType, mlir::FloatType>(llvmEleTy))
+    return llvm::failure();
+
+  // Total number of scalar elements and the per-dimension strides for a
+  // column-major flattening (first Fortran dimension varies fastest). This
+  // matches the coordinate order of fir.insert_value/fir.insert_on_range and
+  // makes an insert_on_range map to a contiguous [start, end] flat range.
+  int64_t total = 1;
+  llvm::SmallVector<int64_t> strides(firShape.size());
+  for (size_t i = 0; i < firShape.size(); ++i) {
+    if (firShape[i] < 0)
+      return llvm::failure();
+    strides[i] = total;
+    total *= firShape[i];
+  }
+  if (total == 0)
+    return llvm::failure();
+
+  auto flatIndex = [&](llvm::ArrayRef<int64_t> coord) -> int64_t {
+    int64_t idx = 0;
+    for (size_t i = 0; i < coord.size(); ++i)
+      idx += coord[i] * strides[i];
+    return idx;
+  };
+
+  // Fold a scalar element value being inserted into a typed attribute of the
+  // LLVM element type. Handles the fir.convert introduced for logical values,
+  // normalizing them to a canonical 0/1 exactly like ConvertOpConversion does.
+  // Returns a null attribute if the value is not a foldable scalar constant.
+  auto foldElement = [&](mlir::Value val) -> mlir::TypedAttr {
+    auto convertOp = val.getDefiningOp<fir::ConvertOp>();
+    mlir::Value src = convertOp ? convertOp.getValue() : val;
+    auto cst = src.getDefiningOp<mlir::arith::ConstantOp>();
+    if (!cst)
+      return {};
+    mlir::TypedAttr valueAttr = cst.getValue();
+    if (valueAttr.getType() == llvmEleTy)
+      return valueAttr;
+    // Only an integer<->logical fir.convert is folded here: it normalizes the
+    // operand to a canonical 0/1 (see ConvertOpConversion).
+    auto intAttr = mlir::dyn_cast<mlir::IntegerAttr>(valueAttr);
+    auto intType = mlir::dyn_cast<mlir::IntegerType>(llvmEleTy);
+    if (!intAttr || !intType || !convertOp ||
+        (!mlir::isa<fir::LogicalType>(convertOp.getType()) &&
+         !mlir::isa<fir::LogicalType>(convertOp.getValue().getType())))
+      return {};
+    return mlir::IntegerAttr::get(intType, intAttr.getValue().isZero() ? 0 : 1);
+  };
+
+  // Element attribute for each flat index. A null entry means "not yet set".
+  std::vector<mlir::Attribute> elements(total, mlir::Attribute{});
+  mlir::Attribute defaultValue; // value for elements never explicitly set
+
+  // Walk the initializer chain once, from the array value back to the root.
+  bool sawExplicitInit = false;
+  mlir::Value cur = arrayVal;
+  while (true) {
+    mlir::Operation *op = cur.getDefiningOp();
+    if (!op)
+      return llvm::failure();
+    if (mlir::isa<fir::UndefOp>(op))
+      break; // undefined root: every element must be set explicitly.
+    if (mlir::isa<fir::ZeroOp>(op)) {
+      if (auto intType = mlir::dyn_cast<mlir::IntegerType>(llvmEleTy))
+        defaultValue = mlir::IntegerAttr::get(intType, 0);
+      else
+        defaultValue = mlir::FloatAttr::get(llvmEleTy, 0.0);
+      break;
+    }
+    if (auto insert = mlir::dyn_cast<fir::InsertValueOp>(op)) {
+      llvm::SmallVector<int64_t> coord;
+      for (mlir::Attribute a : insert.getCoor()) {
+        auto intAttr = mlir::dyn_cast<mlir::IntegerAttr>(a);
+        if (!intAttr) // a named field: not a plain scalar array element.
+          return llvm::failure();
+        coord.push_back(intAttr.getInt());
+      }
+      if (coord.size() != firShape.size())
+        return llvm::failure();
+      mlir::TypedAttr eleAttr = foldElement(insert.getVal());
+      if (!eleAttr)
+        return llvm::failure();
+      int64_t idx = flatIndex(coord);
+      if (idx < 0 || idx >= total)
+        return llvm::failure();
+      // Walking backward: the first value seen for an index wins (it is the
+      // last insert in program order).
+      if (!elements[idx])
+        elements[idx] = eleAttr;
+      sawExplicitInit = true;
+      cur = insert.getAdt();
+      continue;
+    }
+    if (auto range = mlir::dyn_cast<fir::InsertOnRangeOp>(op)) {
+      mlir::TypedAttr eleAttr = foldElement(range.getVal());
+      if (!eleAttr)
+        return llvm::failure();
+      // coor holds [lb0, ub0, lb1, ub1, ...] in Fortran dimension order.
+      auto bounds = range.getCoor().getValues<int64_t>();
+      if (bounds.size() != 2 * firShape.size())
+        return llvm::failure();
+      llvm::SmallVector<int64_t> lb, ub;
+      for (size_t i = 0; i < firShape.size(); ++i) {
+        lb.push_back(bounds[2 * i]);
+        ub.push_back(bounds[2 * i + 1]);
+      }
+      int64_t start = flatIndex(lb);
+      int64_t end = flatIndex(ub);
+      if (start < 0 || end >= total || start > end)
+        return llvm::failure();
+      for (int64_t i = start; i <= end; ++i)
+        if (!elements[i])
+          elements[i] = eleAttr;
+      sawExplicitInit = true;
+      cur = range.getSeq();
+      continue;
+    }
+    // Any other producer (embox, address_of, aggregate insert, ...) cannot be
+    // folded to a scalar array constant.
+    return llvm::failure();
+  }
+
+  // A bare zero/undef array with no explicit element initialization has no
+  // insertvalue chain to collapse: leave it to the regular lowering.
+  if (requireExplicitInit && !sawExplicitInit)
+    return llvm::failure();
+
+  // Fill any element that was not explicitly initialized.
+  for (mlir::Attribute &a : elements)
+    if (!a) {
+      if (!defaultValue) // undefined element with no zero default.
+        return llvm::failure();
+      a = defaultValue;
+    }
+
+  // Build a DenseElementsAttr over a tensor shape matching the row-major LLVM
+  // array layout (the reverse of the Fortran dimension order). The element
+  // vector is already in that row-major order (first Fortran dimension
+  // fastest).
+  llvm::SmallVector<int64_t> tensorShape(firShape.rbegin(), firShape.rend());
+  auto tensorTy = mlir::RankedTensorType::get(tensorShape, llvmEleTy);
+  return mlir::DenseElementsAttr::get(tensorTy, elements);
+}
+
+/// A member of a struct/tuple `fir.global` initializer that has been folded to
+/// a single constant attribute, ready to be materialized as one
+/// `llvm.mlir.constant` and inserted into the aggregate.
+struct FoldedGlobalMember {
+  int64_t index;            // position of the member in the struct
+  mlir::Type llvmType;      // lowered LLVM type of the member
+  mlir::TypedAttr constant; // the folded constant (dense array or scalar)
+};
+
+/// Fold a scalar array value or a scalar value being inserted into a struct
+/// member into a single typed constant attribute. Returns a null attribute if
+/// the value is not foldable.
+static mlir::TypedAttr foldStructMemberConstant(mlir::Value val,
+                                                mlir::Type llvmMemberType) {
+  // Scalar array member: fold to a DenseElementsAttr. A struct member is
+  // folded even when it is entirely zero/undef, so that a data-bearing sibling
+  // member still takes the fast path (the whole struct is folded or none of
+  // it).
+  if (mlir::isa<mlir::LLVM::LLVMArrayType>(llvmMemberType)) {
+    llvm::FailureOr<mlir::DenseElementsAttr> folded =
+        foldScalarArrayConstant(val, llvmMemberType,
+                                /*requireExplicitInit=*/false);
+    if (llvm::succeeded(folded))
+      return *folded;
+    return {};
+  }
+  // Scalar member: fold to a scalar constant, looking through the fir.convert
+  // introduced for logical values (normalizing them to a canonical 0/1).
+  if (mlir::isa<mlir::IntegerType, mlir::FloatType>(llvmMemberType)) {
+    auto convertOp = val.getDefiningOp<fir::ConvertOp>();
+    mlir::Value src = convertOp ? convertOp.getValue() : val;
+    auto cst = src.getDefiningOp<mlir::arith::ConstantOp>();
+    if (!cst)
+      return {};
+    mlir::TypedAttr valueAttr = cst.getValue();
+    if (valueAttr.getType() == llvmMemberType)
+      return valueAttr;
+    auto intAttr = mlir::dyn_cast<mlir::IntegerAttr>(valueAttr);
+    auto intType = mlir::dyn_cast<mlir::IntegerType>(llvmMemberType);
+    if (!intAttr || !intType || !convertOp ||
+        (!mlir::isa<fir::LogicalType>(convertOp.getType()) &&
+         !mlir::isa<fir::LogicalType>(convertOp.getValue().getType())))
+      return {};
+    return mlir::IntegerAttr::get(intType, intAttr.getValue().isZero() ? 0 : 1);
+  }
+  return {};
+}
+
+/// Fold the initializer body of a struct/tuple `fir.global` whose members are
+/// scalar arrays or scalars (as generated for COMMON blocks) into a list of
+/// per-member constant attributes. The caller emits a small region that
+/// materializes each member with one `llvm.mlir.constant` and inserts it into
+/// the aggregate: a single struct-level `llvm.mlir.constant` cannot represent a
+/// struct containing an array.
+///
+/// Returns failure if the global is not such a struct, or if any member is not
+/// a foldable scalar array/scalar (e.g. it involves a descriptor, pointer or
+/// nested aggregate), in which case the caller falls back to the regular
+/// per-operation lowering.
+static llvm::FailureOr<llvm::SmallVector<FoldedGlobalMember>>
+foldGlobalStructInitializer(fir::GlobalOp global, mlir::Type llvmType) {
+  auto structTy = mlir::dyn_cast<mlir::LLVM::LLVMStructType>(llvmType);
+  if (!structTy || structTy.isOpaque())
+    return llvm::failure();
+  llvm::ArrayRef<mlir::Type> body = structTy.getBody();
+  if (body.empty())
+    return llvm::failure();
+
+  mlir::Region &region = global.getRegion();
+  if (region.empty())
+    return llvm::failure();
+  auto hasValue =
+      mlir::dyn_cast<fir::HasValueOp>(region.front().getTerminator());
+  if (!hasValue)
+    return llvm::failure();
+
+  // Walk the struct-level fir.insert_value chain once, back to the root
+  // fir.undefined. Each insert sets one member by its integer field index.
+  llvm::SmallVector<mlir::TypedAttr> members(body.size(), mlir::TypedAttr{});
+  mlir::Value cur = hasValue.getResval();
+  while (true) {
+    mlir::Operation *op = cur.getDefiningOp();
+    if (!op)
+      return llvm::failure();
+    if (mlir::isa<fir::UndefOp, fir::ZeroOp>(op))
+      break;
+    auto insert = mlir::dyn_cast<fir::InsertValueOp>(op);
+    if (!insert)
+      return llvm::failure();
+    // A struct field is addressed by a single integer index.
+    if (insert.getCoor().size() != 1)
+      return llvm::failure();
+    auto idxAttr = mlir::dyn_cast<mlir::IntegerAttr>(insert.getCoor()[0]);
+    if (!idxAttr)
+      return llvm::failure();
+    int64_t field = idxAttr.getInt();
+    if (field < 0 || field >= static_cast<int64_t>(body.size()))
+      return llvm::failure();
+    // Walking backward: the first value seen for a field wins.
+    if (!members[field]) {
+      mlir::TypedAttr memberAttr =
+          foldStructMemberConstant(insert.getVal(), body[field]);
+      if (!memberAttr)
+        return llvm::failure();
+      members[field] = memberAttr;
+    }
+    cur = insert.getAdt();
+  }
+
+  // Every member must have been explicitly set: there is no zero default here
+  // because a struct member may itself need a specific per-element value.
+  llvm::SmallVector<FoldedGlobalMember> folded;
+  for (auto [i, memberAttr] : llvm::enumerate(members)) {
+    if (!memberAttr)
+      return llvm::failure();
+    folded.push_back({static_cast<int64_t>(i), body[i], memberAttr});
+  }
+  return folded;
+}
+
 /// Lower `fir.global` operation to `llvm.global` operation.
 /// `fir.insert_on_range` operations are replaced with constant dense attribute
 /// if they are applied on the full range.
@@ -3818,6 +4138,31 @@ struct GlobalOpConversion : public fir::FIROpConversion<fir::GlobalOp> {
       tyAttr = this->lowerTy().convertBoxTypeAsStruct(boxType);
     auto loc = global.getLoc();
     mlir::Attribute initAttr = global.getInitVal().value_or(mlir::Attribute());
+    // Try to fold the initializer body into constant attributes so a compact
+    // llvm.mlir.global can be emitted, avoiding the O(N^2) materialization and
+    // folding of long insertvalue chains for large, partially initialized
+    // arrays. Two forms are handled:
+    //  - a top-level scalar array: a bodyless global with a dense init;
+    //  - a struct/tuple of scalar arrays/scalars (COMMON blocks): a region
+    //    that materializes each member with a single llvm.mlir.constant.
+    bool foldedInit = false;
+    llvm::SmallVector<FoldedGlobalMember> foldedMembers;
+    if (!initAttr && !global.getRegion().empty()) {
+      if (auto hasValue = mlir::dyn_cast<fir::HasValueOp>(
+              global.getRegion().front().getTerminator())) {
+        llvm::FailureOr<mlir::DenseElementsAttr> foldedArray =
+            foldScalarArrayConstant(hasValue.getResval(), tyAttr);
+        if (llvm::succeeded(foldedArray)) {
+          initAttr = *foldedArray;
+          foldedInit = true;
+        } else {
+          llvm::FailureOr<llvm::SmallVector<FoldedGlobalMember>> foldedStruct =
+              foldGlobalStructInitializer(global, tyAttr);
+          if (llvm::succeeded(foldedStruct))
+            foldedMembers = std::move(*foldedStruct);
+        }
+      }
+    }
     assert(attributeTypeIsCompatible(global.getContext(), initAttr, tyAttr));
     auto linkage = convertLinkage(global.getLinkage());
     auto isConst = global.getConstant().has_value();
@@ -3867,47 +4212,64 @@ struct GlobalOpConversion : public fir::FIROpConversion<fir::GlobalOp> {
     g->setDiscardableAttrs(global->getDiscardableAttrDictionary());
 
     auto &gr = g.getInitializerRegion();
-    rewriter.inlineRegionBefore(global.getRegion(), gr, gr.end());
-    if (!gr.empty()) {
-      // Replace insert_on_range with a constant dense attribute if the
-      // initialization is on the full range.
-      auto insertOnRangeOps = gr.front().getOps<fir::InsertOnRangeOp>();
-      for (auto insertOp : insertOnRangeOps) {
-        if (!insertOp.isFullRange())
-          continue;
-        // The dense attribute must use the converted element type of the
-        // array, not the type of whatever constant feeds the insertion.
-        mlir::Type elementType = convertType(insertOp.getType().getEleTy());
-        mlir::Value val = insertOp.getVal();
-        // Logical constants reach the insertion through a `fir.convert`.
-        auto convertOp = val.getDefiningOp<fir::ConvertOp>();
-        if (convertOp)
-          val = convertOp.getValue();
-        auto constant = val.getDefiningOp<mlir::arith::ConstantOp>();
-        if (!constant)
-          continue;
-        mlir::TypedAttr valueAttr = constant.getValue();
-        if (valueAttr.getType() != elementType) {
-          // Looking through the `fir.convert` leaves the constant with the
-          // source type. Only an integer<->logical conversion is folded here:
-          // it normalizes any integer operand to a canonical 0/1, see
-          // ConvertOpConversion. Any other mismatching conversion is left to
-          // the regular lowering.
-          auto intAttr = mlir::dyn_cast<mlir::IntegerAttr>(valueAttr);
-          auto intType = mlir::dyn_cast<mlir::IntegerType>(elementType);
-          if (!intAttr || !intType || !convertOp ||
-              (!mlir::isa<fir::LogicalType>(convertOp.getType()) &&
-               !mlir::isa<fir::LogicalType>(convertOp.getValue().getType())))
+    if (!foldedMembers.empty()) {
+      // Materialize the folded struct/tuple members: one llvm.mlir.constant per
+      // member inserted into an undef aggregate, then returned. A single
+      // struct-level llvm.mlir.constant cannot represent a struct containing an
+      // array, so the aggregate is assembled with llvm.insertvalue.
+      mlir::Block *block = rewriter.createBlock(&gr);
+      rewriter.setInsertionPointToStart(block);
+      mlir::Value agg = mlir::LLVM::UndefOp::create(rewriter, loc, tyAttr);
+      for (const FoldedGlobalMember &member : foldedMembers) {
+        mlir::Value cst = mlir::LLVM::ConstantOp::create(
+            rewriter, loc, member.llvmType, member.constant);
+        agg = mlir::LLVM::InsertValueOp::create(rewriter, loc, agg, cst,
+                                                member.index);
+      }
+      mlir::LLVM::ReturnOp::create(rewriter, loc, agg);
+    } else if (!foldedInit) {
+      rewriter.inlineRegionBefore(global.getRegion(), gr, gr.end());
+      if (!gr.empty()) {
+        // Replace insert_on_range with a constant dense attribute if the
+        // initialization is on the full range.
+        auto insertOnRangeOps = gr.front().getOps<fir::InsertOnRangeOp>();
+        for (auto insertOp : insertOnRangeOps) {
+          if (!insertOp.isFullRange())
+            continue;
+          // The dense attribute must use the converted element type of the
+          // array, not the type of whatever constant feeds the insertion.
+          mlir::Type elementType = convertType(insertOp.getType().getEleTy());
+          mlir::Value val = insertOp.getVal();
+          // Logical constants reach the insertion through a `fir.convert`.
+          auto convertOp = val.getDefiningOp<fir::ConvertOp>();
+          if (convertOp)
+            val = convertOp.getValue();
+          auto constant = val.getDefiningOp<mlir::arith::ConstantOp>();
+          if (!constant)
             continue;
-          valueAttr = mlir::IntegerAttr::get(
-              intType, intAttr.getValue().isZero() ? 0 : 1);
+          mlir::TypedAttr valueAttr = constant.getValue();
+          if (valueAttr.getType() != elementType) {
+            // Looking through the `fir.convert` leaves the constant with the
+            // source type. Only an integer<->logical conversion is folded here:
+            // it normalizes any integer operand to a canonical 0/1, see
+            // ConvertOpConversion. Any other mismatching conversion is left to
+            // the regular lowering.
+            auto intAttr = mlir::dyn_cast<mlir::IntegerAttr>(valueAttr);
+            auto intType = mlir::dyn_cast<mlir::IntegerType>(elementType);
+            if (!intAttr || !intType || !convertOp ||
+                (!mlir::isa<fir::LogicalType>(convertOp.getType()) &&
+                 !mlir::isa<fir::LogicalType>(convertOp.getValue().getType())))
+              continue;
+            valueAttr = mlir::IntegerAttr::get(
+                intType, intAttr.getValue().isZero() ? 0 : 1);
+          }
+          auto vecType =
+              mlir::VectorType::get(insertOp.getType().getShape(), elementType);
+          auto denseAttr = mlir::DenseElementsAttr::get(vecType, valueAttr);
+          rewriter.setInsertionPointAfter(insertOp);
+          rewriter.replaceOpWithNewOp<mlir::arith::ConstantOp>(
+              insertOp, convertType(insertOp.getType()), denseAttr);
         }
-        auto vecType =
-            mlir::VectorType::get(insertOp.getType().getShape(), elementType);
-        auto denseAttr = mlir::DenseElementsAttr::get(vecType, valueAttr);
-        rewriter.setInsertionPointAfter(insertOp);
-        rewriter.replaceOpWithNewOp<mlir::arith::ConstantOp>(
-            insertOp, convertType(insertOp.getType()), denseAttr);
       }
     }
 
diff --git a/flang/test/Fir/convert-to-llvm.fir b/flang/test/Fir/convert-to-llvm.fir
index d00aae198bf138..7261b4df690b23 100644
--- a/flang/test/Fir/convert-to-llvm.fir
+++ b/flang/test/Fir/convert-to-llvm.fir
@@ -107,13 +107,10 @@ fir.global internal @_QEmultiarray : !fir.array<32x32xi32> {
   fir.has_value %2 : !fir.array<32x32xi32>
 }
 
-// CHECK: llvm.mlir.global internal @_QEmultiarray() 
+// CHECK: llvm.mlir.global internal @_QEmultiarray(dense<1> : tensor<32x32xi32>) 
 // GENERIC-SAME: {addr_space = 0 : i32}
 // AMDGPU-SAME: {addr_space = 1 : i32}
-// CHECK-SAME: : !llvm.array<32 x array<32 x i32>> {
-// CHECK:   %[[CST:.*]] = llvm.mlir.constant(dense<1> : vector<32x32xi32>) : !llvm.array<32 x array<32 x i32>>
-// CHECK:   llvm.return %[[CST]] : !llvm.array<32 x array<32 x i32>>
-// CHECK: }
+// CHECK-SAME: : !llvm.array<32 x array<32 x i32>>
 
 // -----
 
diff --git a/flang/test/Fir/global-constant-fold.fir b/flang/test/Fir/global-constant-fold.fir
new file mode 100644
index 00000000000000..17b28b31f9a22b
--- /dev/null
+++ b/flang/test/Fir/global-constant-fold.fir
@@ -0,0 +1,153 @@
+// Test the single-pass folding of constant array `fir.global` initializer
+// bodies into dense constants. A partially initialized large array must be
+// lowered to a compact constant instead of a long `llvm.insertvalue` chain
+// (which is pathologically slow to materialize and fold back, see
+// https://github.com/llvm/llvm-project/issues/209393).
+
+// RUN: fir-opt --split-input-file --fir-to-llvm-ir %s | FileCheck %s
+
+// A 1D array initialized with a couple of leading elements and a trailing
+// range folds to a bodyless dense global.
+fir.global @oned : !fir.array<8xi32> {
+  %c7 = arith.constant 7 : i32
+  %c77 = arith.constant 77 : i32
+  %c0 = arith.constant 0 : i32
+  %0 = fir.undefined !fir.array<8xi32>
+  %1 = fir.insert_value %0, %c7, [0 : index] : (!fir.array<8xi32>, i32) -> !fir.array<8xi32>
+  %2 = fir.insert_value %1, %c77, [1 : index] : (!fir.array<8xi32>, i32) -> !fir.array<8xi32>
+  %3 = fir.insert_on_range %2, %c0 from (2) to (7) : (!fir.array<8xi32>, i32) -> !fir.array<8xi32>
+  fir.has_value %3 : !fir.array<8xi32>
+}
+
+// CHECK: llvm.mlir.global external @oned(dense<[7, 77, 0, 0, 0, 0, 0, 0]> : tensor<8xi32>) {{.*}} : !llvm.array<8 x i32>
+// CHECK-NOT: llvm.insertvalue
+
+// -----
+
+// A 2D array with a mix of a single-element insert and a range insert folds to
+// a dense global. The tensor shape is the reverse of the Fortran dimensions,
+// and the element order is column-major (first Fortran dimension fastest).
+fir.global @twod : !fir.array<3x2xi32> {
+  %c9 = arith.constant 9 : i32
+  %c0 = arith.constant 0 : i32
+  %0 = fir.undefined !fir.array<3x2xi32>
+  %1 = fir.insert_value %0, %c9, [0 : index, 0 : index] : (!fir.array<3x2xi32>, i32) -> !fir.array<3x2xi32>
+  %2 = fir.insert_on_range %1, %c0 from (1,0) to (2,1) : (!fir.array<3x2xi32>, i32) -> !fir.array<3x2xi32>
+  fir.has_value %2 : !fir.array<3x2xi32>
+}
+
+// CHECK: llvm.mlir.global external @twod(dense<{{\[}}[9, 0, 0], [0, 0, 0]]> : tensor<2x3xi32>) {{.*}} : !llvm.array<2 x array<3 x i32>>
+// CHECK-NOT: llvm.insertvalue
+
+// -----
+
+// A 3D array fully initialized by a single range folds to a dense splat.
+fir.global @threed : !fir.array<2x2x2xi32> {
+  %c5 = arith.constant 5 : i32
+  %0 = fir.undefined !fir.array<2x2x2xi32>
+  %1 = fir.insert_on_range %0, %c5 from (0,0,0) to (1,1,1) : (!fir.array<2x2x2xi32>, i32) -> !fir.array<2x2x2xi32>
+  fir.has_value %1 : !fir.array<2x2x2xi32>
+}
+
+// CHECK: llvm.mlir.global external @threed(dense<5> : tensor<2x2x2xi32>) {{.*}} : !llvm.array<2 x array<2 x array<2 x i32>>>
+// CHECK-NOT: llvm.insertvalue
+
+// -----
+
+// A real (floating point) array folds as well.
+fir.global @reals : !fir.array<4xf32> {
+  %c1 = arith.constant 1.5 : f32
+  %c0 = arith.constant 0.0 : f32
+  %0 = fir.undefined !fir.array<4xf32>
+  %1 = fir.insert_value %0, %c1, [0 : index] : (!fir.array<4xf32>, f32) -> !fir.array<4xf32>
+  %2 = fir.insert_on_range %1, %c0 from (1) to (3) : (!fir.array<4xf32>, f32) -> !fir.array<4xf32>
+  fir.has_value %2 : !fir.array<4xf32>
+}
+
+// CHECK: llvm.mlir.global external @reals(dense<[1.500000e+00, 0.000000e+00, 0.000000e+00, 0.000000e+00]> : tensor<4xf32>) {{.*}} : !llvm.array<4 x f32>
+// CHECK-NOT: llvm.insertvalue
+
+// -----
+
+// A logical array reaches the insertion through a fir.convert; the fold
+// normalizes the value to a canonical 0/1 of the lowered element type.
+fir.global @logicals : !fir.array<4x!fir.logical<4>> {
+  %true = arith.constant true
+  %0 = fir.undefined !fir.array<4x!fir.logical<4>>
+  %1 = fir.convert %true : (i1) -> !fir.logical<4>
+  %2 = fir.insert_on_range %0, %1 from (0) to (3) : (!fir.array<4x!fir.logical<4>>, !fir.logical<4>) -> !fir.array<4x!fir.logical<4>>
+  fir.has_value %2 : !fir.array<4x!fir.logical<4>>
+}
+
+// CHECK: llvm.mlir.global external @logicals(dense<1> : tensor<4xi32>) {{.*}} : !llvm.array<4 x i32>
+// CHECK-NOT: llvm.insertvalue
+
+// -----
+
+// A COMMON block lowers to a single-element tuple wrapping the array. A single
+// struct-level llvm.mlir.constant cannot represent a struct containing an
+// array, so the array member is materialized with one llvm.mlir.constant and
+// inserted into the aggregate.
+fir.global @common1 {alignment = 4 : i64} : tuple<!fir.array<8xi32>> {
+  %c7 = arith.constant 7 : i32
+  %c77 = arith.constant 77 : i32
+  %c0 = arith.constant 0 : i32
+  %0 = fir.zero_bits tuple<!fir.array<8xi32>>
+  %1 = fir.undefined !fir.array<8xi32>
+  %2 = fir.insert_value %1, %c7, [0 : index] : (!fir.array<8xi32>, i32) -> !fir.array<8xi32>
+  %3 = fir.insert_value %2, %c77, [1 : index] : (!fir.array<8xi32>, i32) -> !fir.array<8xi32>
+  %4 = fir.insert_on_range %3, %c0 from (2) to (7) : (!fir.array<8xi32>, i32) -> !fir.array<8xi32>
+  %5 = fir.insert_value %0, %4, [0 : index] : (tuple<!fir.array<8xi32>>, !fir.array<8xi32>) -> tuple<!fir.array<8xi32>>
+  fir.has_value %5 : tuple<!fir.array<8xi32>>
+}
+
+// CHECK: llvm.mlir.global external @common1() {{.*}} : !llvm.struct<(array<8 x i32>)> {
+// CHECK:   %[[UNDEF:.*]] = llvm.mlir.undef : !llvm.struct<(array<8 x i32>)>
+// CHECK:   %[[CST:.*]] = llvm.mlir.constant(dense<[7, 77, 0, 0, 0, 0, 0, 0]> : tensor<8xi32>) : !llvm.array<8 x i32>
+// CHECK:   %[[INS:.*]] = llvm.insertvalue %[[CST]], %[[UNDEF]][0] : !llvm.struct<(array<8 x i32>)>
+// CHECK:   llvm.return %[[INS]] : !llvm.struct<(array<8 x i32>)>
+// CHECK: }
+
+// -----
+
+// A COMMON block with several members: each member folds to one constant and
+// is inserted at its field index.
+fir.global @common2 : tuple<!fir.array<4xi32>, !fir.array<2xi32>> {
+  %c1 = arith.constant 1 : i32
+  %c0 = arith.constant 0 : i32
+  %c3 = arith.constant 3 : i32
+  %0 = fir.zero_bits tuple<!fir.array<4xi32>, !fir.array<2xi32>>
+  %1 = fir.undefined !fir.array<4xi32>
+  %2 = fir.insert_on_range %1, %c1 from (0) to (3) : (!fir.array<4xi32>, i32) -> !fir.array<4xi32>
+  %3 = fir.insert_value %0, %2, [0 : index] : (tuple<!fir.array<4xi32>, !fir.array<2xi32>>, !fir.array<4xi32>) -> tuple<!fir.array<4xi32>, !fir.array<2xi32>>
+  %4 = fir.undefined !fir.array<2xi32>
+  %5 = fir.insert_value %4, %c3, [0 : index] : (!fir.array<2xi32>, i32) -> !fir.array<2xi32>
+  %6 = fir.insert_value %5, %c0, [1 : index] : (!fir.array<2xi32>, i32) -> !fir.array<2xi32>
+  %7 = fir.insert_value %3, %6, [1 : index] : (tuple<!fir.array<4xi32>, !fir.array<2xi32>>, !fir.array<2xi32>) -> tuple<!fir.array<4xi32>, !fir.array<2xi32>>
+  fir.has_value %7 : tuple<!fir.array<4xi32>, !fir.array<2xi32>>
+}
+
+// CHECK: llvm.mlir.global external @common2() {{.*}} : !llvm.struct<(array<4 x i32>, array<2 x i32>)> {
+// CHECK:   %[[UNDEF:.*]] = llvm.mlir.undef : !llvm.struct<(array<4 x i32>, array<2 x i32>)>
+// CHECK:   %[[CST0:.*]] = llvm.mlir.constant(dense<1> : tensor<4xi32>) : !llvm.array<4 x i32>
+// CHECK:   %[[INS0:.*]] = llvm.insertvalue %[[CST0]], %[[UNDEF]][0] : !llvm.struct<(array<4 x i32>, array<2 x i32>)>
+// CHECK:   %[[CST1:.*]] = llvm.mlir.constant(dense<[3, 0]> : tensor<2xi32>) : !llvm.array<2 x i32>
+// CHECK:   %[[INS1:.*]] = llvm.insertvalue %[[CST1]], %[[INS0]][1] : !llvm.struct<(array<4 x i32>, array<2 x i32>)>
+// CHECK:   llvm.return %[[INS1]] : !llvm.struct<(array<4 x i32>, array<2 x i32>)>
+// CHECK: }
+
+// -----
+
+// An array whose leading elements are left undefined (no zero base) has no
+// constant to fold to for those elements: it must stay on the regular
+// per-operation lowering rather than being force-folded.
+fir.global @partialundef : !fir.array<8xi32> {
+  %c1 = arith.constant 1 : i32
+  %0 = fir.undefined !fir.array<8xi32>
+  %1 = fir.insert_on_range %0, %c1 from (5) to (7) : (!fir.array<8xi32>, i32) -> !fir.array<8xi32>
+  fir.has_value %1 : !fir.array<8xi32>
+}
+
+// CHECK: llvm.mlir.global external @partialundef() {{.*}} : !llvm.array<8 x i32> {
+// CHECK:   llvm.insertvalue
+// CHECK: }
diff --git a/flang/test/Fir/global-initialization.fir b/flang/test/Fir/global-initialization.fir
index a683b46de2541a..73cba152d0c202 100644
--- a/flang/test/Fir/global-initialization.fir
+++ b/flang/test/Fir/global-initialization.fir
@@ -8,12 +8,7 @@ fir.global internal @_QEmask : !fir.array<32xi32> {
   fir.has_value %2 : !fir.array<32xi32>
 }
 
-// CHECK: llvm.mlir.global internal @_QEmask() {addr_space = 0 : i32} : !llvm.array<32 x i32> {
-// CHECK:   [[VAL0:%.*]] = llvm.mlir.constant(1 : i32) : i32
-// CHECK:   [[VAL1:%.*]] = llvm.mlir.undef : !llvm.array<32 x i32>
-// CHECK:   [[VAL2:%.*]] = llvm.mlir.constant(dense<1> : vector<32xi32>) : !llvm.array<32 x i32>
-// CHECK:   llvm.return [[VAL2]] : !llvm.array<32 x i32>
-// CHECK: }
+// CHECK: llvm.mlir.global internal @_QEmask(dense<1> : tensor<32xi32>) {addr_space = 0 : i32} : !llvm.array<32 x i32>
 
 fir.global internal @_QEmultiarray : !fir.array<32x32xi32> {
   %c0_i32 = arith.constant 1 : i32
@@ -22,12 +17,7 @@ fir.global internal @_QEmultiarray : !fir.array<32x32xi32> {
   fir.has_value %2 : !fir.array<32x32xi32>
 }
 
-// CHECK: llvm.mlir.global internal @_QEmultiarray() {addr_space = 0 : i32} : !llvm.array<32 x array<32 x i32>> {
-// CHECK:   [[VAL0:%.*]] = llvm.mlir.constant(1 : i32) : i32
-// CHECK:   [[VAL1:%.*]] = llvm.mlir.undef : !llvm.array<32 x array<32 x i32>>
-// CHECK:   [[VAL2:%.*]] = llvm.mlir.constant(dense<1> : vector<32x32xi32>) : !llvm.array<32 x array<32 x i32>>
-// CHECK:   llvm.return [[VAL2]] : !llvm.array<32 x array<32 x i32>>
-// CHECK: }
+// CHECK: llvm.mlir.global internal @_QEmultiarray(dense<1> : tensor<32x32xi32>) {addr_space = 0 : i32} : !llvm.array<32 x array<32 x i32>>
 
 fir.global internal @_QEmasklogical : !fir.array<32768x!fir.logical<4>> {
   %true = arith.constant true
@@ -37,13 +27,7 @@ fir.global internal @_QEmasklogical : !fir.array<32768x!fir.logical<4>> {
   fir.has_value %2 : !fir.array<32768x!fir.logical<4>>
 }
 
-// CHECK: llvm.mlir.global internal @_QEmasklogical() {addr_space = 0 : i32} : !llvm.array<32768 x i32> {
-// CHECK:   [[VAL0:%.*]] = llvm.mlir.constant(true) : i1
-// CHECK:   [[VAL1:%.*]] = llvm.mlir.undef : !llvm.array<32768 x i32>
-// CHECK:   [[VAL2:%.*]] = llvm.mlir.constant(1 : i32) : i32
-// CHECK:   [[VAL3:%.*]] = llvm.mlir.constant(dense<1> : vector<32768xi32>) : !llvm.array<32768 x i32>
-// CHECK:   llvm.return [[VAL3]] : !llvm.array<32768 x i32>
-// CHECK: }
+// CHECK: llvm.mlir.global internal @_QEmasklogical(dense<1> : tensor<32768xi32>) {addr_space = 0 : i32} : !llvm.array<32768 x i32>
 
 // A logical conversion normalizes any integer operand to a canonical 0/1, not
 // just an `i1` one, so the full-range fold must apply here as well.
@@ -55,10 +39,7 @@ fir.global internal @_QEmasklogicalkind : !fir.array<8x!fir.logical<8>> {
   fir.has_value %2 : !fir.array<8x!fir.logical<8>>
 }
 
-// CHECK: llvm.mlir.global internal @_QEmasklogicalkind() {addr_space = 0 : i32} : !llvm.array<8 x i64> {
-// CHECK:   [[VAL2:%.*]] = llvm.mlir.constant(dense<1> : vector<8xi64>) : !llvm.array<8 x i64>
-// CHECK:   llvm.return [[VAL2]] : !llvm.array<8 x i64>
-// CHECK: }
+// CHECK: llvm.mlir.global internal @_QEmasklogicalkind(dense<1> : tensor<8xi64>) {addr_space = 0 : i32} : !llvm.array<8 x i64>
 
 fir.global internal @_QElookforme : !fir.type<_QTt{i:!fir.array<500xi32>,j:!fir.array<500xi32>}> {
   %c2_i32 = arith.constant 2 : i32
diff --git a/flang/test/Fir/omp-declare-target-data.fir b/flang/test/Fir/omp-declare-target-data.fir
index 539a7acf267f2d..9ad41f08e8b4a1 100644
--- a/flang/test/Fir/omp-declare-target-data.fir
+++ b/flang/test/Fir/omp-declare-target-data.fir
@@ -5,7 +5,7 @@ module attributes {omp.is_target_device = false} {
   // CHECK: llvm.mlir.global external @_QMtest_0Earray_1d(dense<[1, 2, 3]> : tensor<3xi32>) {{{.*}}omp.declare_target = #omp.declaretarget<device_type = any, capture_clause = link>{{.*}}} : !llvm.array<3 x i32>
   fir.global @_QMtest_0Earray_1d(dense<[1, 2, 3]> : tensor<3xi32>) {omp.declare_target = #omp.declaretarget<device_type = any, capture_clause = link>} : !fir.array<3xi32>
 
-  // CHECK: llvm.mlir.global external @_QMtest_0Earray_2d() {{{.*}}omp.declare_target = #omp.declaretarget<device_type = any, capture_clause = link>{{.*}}} : !llvm.array<2 x array<2 x i32>>
+  // CHECK: llvm.mlir.global external @_QMtest_0Earray_2d(dense<{{\[}}[1, 2], [3, 4]]> : tensor<2x2xi32>) {{{.*}}omp.declare_target = #omp.declaretarget<device_type = any, capture_clause = link>{{.*}}} : !llvm.array<2 x array<2 x i32>>
   fir.global @_QMtest_0Earray_2d {omp.declare_target = #omp.declaretarget<device_type = any, capture_clause = link>} : !fir.array<2x2xi32> {
     %0 = fir.undefined !fir.array<2x2xi32>
     %c1_i32 = arith.constant 1 : i32
diff --git a/flang/test/Integration/common-block-large-init.f90 b/flang/test/Integration/common-block-large-init.f90
new file mode 100644
index 00000000000000..4fd563707aff10
--- /dev/null
+++ b/flang/test/Integration/common-block-large-init.f90
@@ -0,0 +1,15 @@
+! Verify that a large COMMON block array initialized on only a few elements by a
+! DATA statement lowers to a compact constant global instead of an enormous
+! `insertvalue` chain (which is pathologically slow to generate, see
+! https://github.com/llvm/llvm-project/issues/209393).
+
+! RUN: %flang_fc1 -emit-llvm %s -o - | FileCheck %s
+
+block data
+  integer i(55000)
+  common /b/ i
+  data (i(j), j = 1, 2) / 7, 77 /
+end block data
+
+! CHECK: @b_ = {{.*}}global { [55000 x i32] } { [55000 x i32] [i32 7, i32 77, i32 0
+! CHECK-NOT: insertvalue



More information about the flang-commits mailing list