[llvm] [AArch64] Avoid expanding arrays to classify consecutive argument registers (PR #226887)

via llvm-commits llvm-commits at lists.llvm.org
Sun Sep 27 21:39:23 PDT 2026


https://github.com/zhouguangyuan0718 created https://github.com/llvm/llvm-project/pull/226887

AArch64's `functionArgumentNeedsConsecutiveRegisters` expands every array element with `ComputeValueVTs` just to check whether all member EVTs are equal. An unused large `byval` array therefore consumes memory proportional to its element count during `SelectionDAGISel::LowerArguments`.

For example, this hand-written IR uses the ordinary C calling convention:

```llvm
define i32 @large_array(ptr byval([1073741808 x i8]) align 1 %arg) {
  ret i32 1
}
```

```sh
llc -mtriple=aarch64-linux-gnu -global-isel=0 -O2 -verify-machineinstrs repro.ll -o out.s
```

On unmodified main `b1f669055159`, this exceeded a 1 GiB RSS watchdog threshold in about 0.82 seconds. The watchdog deliberately sent SIGABRT rather than allowing an unbounded allocation; the resulting stack contains `ComputeValueTypes -> ComputeValueVTs -> functionArgumentNeedsConsecutiveRegisters -> LowerArguments`. After this change, the large and nested-array regression compiles at O0/O2 with roughly 25 MiB peak RSS.

Walk distinct aggregate types and compare their lowered EVTs instead. Nonempty arrays only need their element type inspected once; empty arrays/structs contribute no members. This retains pointer/integer EVT equivalence and leaves the existing non-array/scalable-vector rule unchanged. No calling-convention register assignment or optimization settings are changed.

This is **not a C-source reproducer**. Unmodified Clang 24.0.0git successfully compiled C array parameters and large struct/nested struct/mixed struct parameters at O0/O2 for AArch64 Linux, big-endian Linux, Darwin, Windows, and FreeBSD. Those large parameters lower to ordinary pointers without `byval`, so the tested C forms do not reach this case. The regression exercises the generic LLVM IR interface directly.

Validation on the rebased patch (base `a98dcad4fed0`, Release with assertions, macOS arm64):

- All 102 AArch64 unit tests pass.
- Complete `CodeGen/AArch64` suite: 4,261 passed, 4 expected failures, no unexpected failures (4,265 tests).
- The new large/nested-array regression passes at O0/O2 with machine verification; final measurements are 21.95/24.39 MiB peak RSS.
- 32 assembly comparisons are byte-identical before/after, covering Clang-generated IR, ordinary/mixed/empty/nested aggregates, pointer/integer members, fixed vectors, and SVE calling conventions.
- Added generic O0/O2 SelectionDAG coverage and predicate unit coverage; no Go-specific tests or metadata.

The implementation is adapted from the downstream fix in https://github.com/goallc/llvm-project/pull/170. This patch only bounds the classification query; it does not establish that arbitrary gigabyte-sized calls can execute successfully.

AI assistance: OpenAI Codex was used to adapt the implementation, write tests and this description, and run validation, under the contributor's explicit authorization.


>From cf00c53f2bdea040acbf30a2d602aa447315cc56 Mon Sep 17 00:00:00 2001
From: ZhouGuangyuan <zhouguangyuan.xian at gmail.com>
Date: Mon, 28 Sep 2026 12:27:27 +0800
Subject: [PATCH] [AArch64] Avoid expanding arrays to classify consecutive
 argument registers

Walk distinct aggregate types and compare their lowered EVTs instead of materializing an entry for every array element. This bounds the argument-classification query for large byval arrays while retaining the existing empty-aggregate and non-array behavior.

Assisted-by: OpenAI Codex
---
 .../Target/AArch64/AArch64ISelLowering.cpp    | 28 ++++++++++++--
 .../test/CodeGen/AArch64/large-byval-array.ll | 18 +++++++++
 .../AArch64/AArch64SelectionDAGTest.cpp       | 38 +++++++++++++++++++
 3 files changed, 80 insertions(+), 4 deletions(-)
 create mode 100644 llvm/test/CodeGen/AArch64/large-byval-array.ll

diff --git a/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp b/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
index c98d607c4a22c..6b0439198df9b 100644
--- a/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
+++ b/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
@@ -25,6 +25,7 @@
 #include "llvm/ADT/APInt.h"
 #include "llvm/ADT/ArrayRef.h"
 #include "llvm/ADT/STLExtras.h"
+#include "llvm/ADT/SmallPtrSet.h"
 #include "llvm/ADT/SmallSet.h"
 #include "llvm/ADT/SmallVector.h"
 #include "llvm/ADT/SmallVectorExtras.h"
@@ -33548,10 +33549,29 @@ bool AArch64TargetLowering::functionArgumentNeedsConsecutiveRegisters(
     return TySize.isScalable() && TySize.getKnownMinValue() > 128;
   }
 
-  // All non aggregate members of the type must have the same type
-  SmallVector<EVT> ValueVTs;
-  ComputeValueVTs(*this, DL, Ty, ValueVTs);
-  return all_equal(ValueVTs);
+  // All non-aggregate members must have the same value type. Array elements
+  // repeat the same type, so inspect it once instead of expanding potentially
+  // huge byval arrays into one EVT per element.
+  SmallVector<Type *, 8> Worklist{Ty};
+  SmallPtrSet<Type *, 8> Visited;
+  std::optional<EVT> MemberVT;
+  while (!Worklist.empty()) {
+    Type *Member = Worklist.pop_back_val();
+    if (!Visited.insert(Member).second)
+      continue;
+    if (auto *AT = dyn_cast<ArrayType>(Member)) {
+      if (AT->getNumElements())
+        Worklist.push_back(AT->getElementType());
+    } else if (auto *ST = dyn_cast<StructType>(Member)) {
+      append_range(Worklist, ST->elements());
+    } else if (!Member->isVoidTy()) {
+      EVT VT = getValueType(DL, Member);
+      if (MemberVT && *MemberVT != VT)
+        return false;
+      MemberVT = VT;
+    }
+  }
+  return true;
 }
 
 bool AArch64TargetLowering::shouldNormalizeToSelectSequence(LLVMContext &, EVT,
diff --git a/llvm/test/CodeGen/AArch64/large-byval-array.ll b/llvm/test/CodeGen/AArch64/large-byval-array.ll
new file mode 100644
index 0000000000000..4c57eea611f8d
--- /dev/null
+++ b/llvm/test/CodeGen/AArch64/large-byval-array.ll
@@ -0,0 +1,18 @@
+; RUN: llc -mtriple=aarch64-unknown-linux-gnu -global-isel=0 -O0 -verify-machineinstrs < %s | FileCheck %s
+; RUN: llc -mtriple=aarch64-unknown-linux-gnu -global-isel=0 -O2 -verify-machineinstrs < %s | FileCheck %s
+
+; Checking whether array members need consecutive registers must not allocate
+; one type entry per byte of a large byval argument, even when it is unused.
+define i32 @large_array(ptr byval([1073741808 x i8]) align 1 %arg) {
+; CHECK-LABEL: large_array:
+; CHECK: mov w0, #1
+; CHECK-NEXT: ret
+  ret i32 1
+}
+
+define i32 @nested_array(ptr byval([2 x [536870904 x i8]]) align 1 %arg) {
+; CHECK-LABEL: nested_array:
+; CHECK: mov w0, #2
+; CHECK-NEXT: ret
+  ret i32 2
+}
diff --git a/llvm/unittests/Target/AArch64/AArch64SelectionDAGTest.cpp b/llvm/unittests/Target/AArch64/AArch64SelectionDAGTest.cpp
index ab5cf3ef145ee..957a1d967989e 100644
--- a/llvm/unittests/Target/AArch64/AArch64SelectionDAGTest.cpp
+++ b/llvm/unittests/Target/AArch64/AArch64SelectionDAGTest.cpp
@@ -82,6 +82,44 @@ class AArch64SelectionDAGTest : public testing::Test {
   std::unique_ptr<SelectionDAG> DAG;
 };
 
+TEST_F(AArch64SelectionDAGTest, ConsecutiveArgumentRegisters) {
+  const auto &TLI = DAG->getTargetLoweringInfo();
+  auto NeedsConsecutiveRegisters = [&](Type *Ty) {
+    return TLI.functionArgumentNeedsConsecutiveRegisters(
+        Ty, CallingConv::C, false, M->getDataLayout());
+  };
+  Type *I8 = Type::getInt8Ty(Context);
+  Type *I64 = Type::getInt64Ty(Context);
+  Type *F32 = Type::getFloatTy(Context);
+  Type *Empty = StructType::get(Context);
+  Type *Mixed = StructType::get(Context, {I8, F32});
+  Type *Huge = ArrayType::get(I8, 1ULL << 30);
+
+  EXPECT_TRUE(NeedsConsecutiveRegisters(Huge));
+  EXPECT_TRUE(NeedsConsecutiveRegisters(ArrayType::get(Huge, 2)));
+  EXPECT_TRUE(NeedsConsecutiveRegisters(
+      ArrayType::get(StructType::get(Context, {Huge, I8}), 2)));
+  EXPECT_FALSE(NeedsConsecutiveRegisters(
+      ArrayType::get(StructType::get(Context, {Huge, F32}), 2)));
+  EXPECT_FALSE(NeedsConsecutiveRegisters(ArrayType::get(Mixed, 2)));
+  EXPECT_TRUE(NeedsConsecutiveRegisters(ArrayType::get(Mixed, 0)));
+  EXPECT_TRUE(NeedsConsecutiveRegisters(ArrayType::get(Empty, 2)));
+  EXPECT_TRUE(NeedsConsecutiveRegisters(ArrayType::get(
+      StructType::get(Context, {I8, ArrayType::get(F32, 0), Empty}), 2)));
+  // Compare lowered value types, not IR types: pointers and i64 both use i64.
+  EXPECT_TRUE(NeedsConsecutiveRegisters(ArrayType::get(
+      StructType::get(Context, {PointerType::getUnqual(Context), I64}), 2)));
+  EXPECT_TRUE(NeedsConsecutiveRegisters(
+      ArrayType::get(FixedVectorType::get(F32, 4), 2)));
+  EXPECT_FALSE(NeedsConsecutiveRegisters(ArrayType::get(
+      StructType::get(Context, {FixedVectorType::get(F32, 4), F32}), 2)));
+  EXPECT_FALSE(NeedsConsecutiveRegisters(Mixed));
+  EXPECT_FALSE(NeedsConsecutiveRegisters(I64));
+  EXPECT_FALSE(NeedsConsecutiveRegisters(FixedVectorType::get(F32, 4)));
+  EXPECT_FALSE(NeedsConsecutiveRegisters(ScalableVectorType::get(F32, 4)));
+  EXPECT_TRUE(NeedsConsecutiveRegisters(ScalableVectorType::get(F32, 8)));
+}
+
 TEST_F(AArch64SelectionDAGTest, computeKnownBits_ZERO_EXTEND_VECTOR_INREG) {
   SDLoc Loc;
   auto Int8VT = EVT::getIntegerVT(Context, 8);



More information about the llvm-commits mailing list