[llvm] [AArch64] Avoid expanding arrays to classify consecutive argument registers (PR #226887)
via llvm-commits
llvm-commits at lists.llvm.org
Sun Sep 27 21:39:23 PDT 2026
https://github.com/zhouguangyuan0718 created https://github.com/llvm/llvm-project/pull/226887
AArch64's `functionArgumentNeedsConsecutiveRegisters` expands every array element with `ComputeValueVTs` just to check whether all member EVTs are equal. An unused large `byval` array therefore consumes memory proportional to its element count during `SelectionDAGISel::LowerArguments`.
For example, this hand-written IR uses the ordinary C calling convention:
```llvm
define i32 @large_array(ptr byval([1073741808 x i8]) align 1 %arg) {
ret i32 1
}
```
```sh
llc -mtriple=aarch64-linux-gnu -global-isel=0 -O2 -verify-machineinstrs repro.ll -o out.s
```
On unmodified main `b1f669055159`, this exceeded a 1 GiB RSS watchdog threshold in about 0.82 seconds. The watchdog deliberately sent SIGABRT rather than allowing an unbounded allocation; the resulting stack contains `ComputeValueTypes -> ComputeValueVTs -> functionArgumentNeedsConsecutiveRegisters -> LowerArguments`. After this change, the large and nested-array regression compiles at O0/O2 with roughly 25 MiB peak RSS.
Walk distinct aggregate types and compare their lowered EVTs instead. Nonempty arrays only need their element type inspected once; empty arrays/structs contribute no members. This retains pointer/integer EVT equivalence and leaves the existing non-array/scalable-vector rule unchanged. No calling-convention register assignment or optimization settings are changed.
This is **not a C-source reproducer**. Unmodified Clang 24.0.0git successfully compiled C array parameters and large struct/nested struct/mixed struct parameters at O0/O2 for AArch64 Linux, big-endian Linux, Darwin, Windows, and FreeBSD. Those large parameters lower to ordinary pointers without `byval`, so the tested C forms do not reach this case. The regression exercises the generic LLVM IR interface directly.
Validation on the rebased patch (base `a98dcad4fed0`, Release with assertions, macOS arm64):
- All 102 AArch64 unit tests pass.
- Complete `CodeGen/AArch64` suite: 4,261 passed, 4 expected failures, no unexpected failures (4,265 tests).
- The new large/nested-array regression passes at O0/O2 with machine verification; final measurements are 21.95/24.39 MiB peak RSS.
- 32 assembly comparisons are byte-identical before/after, covering Clang-generated IR, ordinary/mixed/empty/nested aggregates, pointer/integer members, fixed vectors, and SVE calling conventions.
- Added generic O0/O2 SelectionDAG coverage and predicate unit coverage; no Go-specific tests or metadata.
The implementation is adapted from the downstream fix in https://github.com/goallc/llvm-project/pull/170. This patch only bounds the classification query; it does not establish that arbitrary gigabyte-sized calls can execute successfully.
AI assistance: OpenAI Codex was used to adapt the implementation, write tests and this description, and run validation, under the contributor's explicit authorization.
>From cf00c53f2bdea040acbf30a2d602aa447315cc56 Mon Sep 17 00:00:00 2001
From: ZhouGuangyuan <zhouguangyuan.xian at gmail.com>
Date: Mon, 28 Sep 2026 12:27:27 +0800
Subject: [PATCH] [AArch64] Avoid expanding arrays to classify consecutive
argument registers
Walk distinct aggregate types and compare their lowered EVTs instead of materializing an entry for every array element. This bounds the argument-classification query for large byval arrays while retaining the existing empty-aggregate and non-array behavior.
Assisted-by: OpenAI Codex
---
.../Target/AArch64/AArch64ISelLowering.cpp | 28 ++++++++++++--
.../test/CodeGen/AArch64/large-byval-array.ll | 18 +++++++++
.../AArch64/AArch64SelectionDAGTest.cpp | 38 +++++++++++++++++++
3 files changed, 80 insertions(+), 4 deletions(-)
create mode 100644 llvm/test/CodeGen/AArch64/large-byval-array.ll
diff --git a/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp b/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
index c98d607c4a22c..6b0439198df9b 100644
--- a/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
+++ b/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
@@ -25,6 +25,7 @@
#include "llvm/ADT/APInt.h"
#include "llvm/ADT/ArrayRef.h"
#include "llvm/ADT/STLExtras.h"
+#include "llvm/ADT/SmallPtrSet.h"
#include "llvm/ADT/SmallSet.h"
#include "llvm/ADT/SmallVector.h"
#include "llvm/ADT/SmallVectorExtras.h"
@@ -33548,10 +33549,29 @@ bool AArch64TargetLowering::functionArgumentNeedsConsecutiveRegisters(
return TySize.isScalable() && TySize.getKnownMinValue() > 128;
}
- // All non aggregate members of the type must have the same type
- SmallVector<EVT> ValueVTs;
- ComputeValueVTs(*this, DL, Ty, ValueVTs);
- return all_equal(ValueVTs);
+ // All non-aggregate members must have the same value type. Array elements
+ // repeat the same type, so inspect it once instead of expanding potentially
+ // huge byval arrays into one EVT per element.
+ SmallVector<Type *, 8> Worklist{Ty};
+ SmallPtrSet<Type *, 8> Visited;
+ std::optional<EVT> MemberVT;
+ while (!Worklist.empty()) {
+ Type *Member = Worklist.pop_back_val();
+ if (!Visited.insert(Member).second)
+ continue;
+ if (auto *AT = dyn_cast<ArrayType>(Member)) {
+ if (AT->getNumElements())
+ Worklist.push_back(AT->getElementType());
+ } else if (auto *ST = dyn_cast<StructType>(Member)) {
+ append_range(Worklist, ST->elements());
+ } else if (!Member->isVoidTy()) {
+ EVT VT = getValueType(DL, Member);
+ if (MemberVT && *MemberVT != VT)
+ return false;
+ MemberVT = VT;
+ }
+ }
+ return true;
}
bool AArch64TargetLowering::shouldNormalizeToSelectSequence(LLVMContext &, EVT,
diff --git a/llvm/test/CodeGen/AArch64/large-byval-array.ll b/llvm/test/CodeGen/AArch64/large-byval-array.ll
new file mode 100644
index 0000000000000..4c57eea611f8d
--- /dev/null
+++ b/llvm/test/CodeGen/AArch64/large-byval-array.ll
@@ -0,0 +1,18 @@
+; RUN: llc -mtriple=aarch64-unknown-linux-gnu -global-isel=0 -O0 -verify-machineinstrs < %s | FileCheck %s
+; RUN: llc -mtriple=aarch64-unknown-linux-gnu -global-isel=0 -O2 -verify-machineinstrs < %s | FileCheck %s
+
+; Checking whether array members need consecutive registers must not allocate
+; one type entry per byte of a large byval argument, even when it is unused.
+define i32 @large_array(ptr byval([1073741808 x i8]) align 1 %arg) {
+; CHECK-LABEL: large_array:
+; CHECK: mov w0, #1
+; CHECK-NEXT: ret
+ ret i32 1
+}
+
+define i32 @nested_array(ptr byval([2 x [536870904 x i8]]) align 1 %arg) {
+; CHECK-LABEL: nested_array:
+; CHECK: mov w0, #2
+; CHECK-NEXT: ret
+ ret i32 2
+}
diff --git a/llvm/unittests/Target/AArch64/AArch64SelectionDAGTest.cpp b/llvm/unittests/Target/AArch64/AArch64SelectionDAGTest.cpp
index ab5cf3ef145ee..957a1d967989e 100644
--- a/llvm/unittests/Target/AArch64/AArch64SelectionDAGTest.cpp
+++ b/llvm/unittests/Target/AArch64/AArch64SelectionDAGTest.cpp
@@ -82,6 +82,44 @@ class AArch64SelectionDAGTest : public testing::Test {
std::unique_ptr<SelectionDAG> DAG;
};
+TEST_F(AArch64SelectionDAGTest, ConsecutiveArgumentRegisters) {
+ const auto &TLI = DAG->getTargetLoweringInfo();
+ auto NeedsConsecutiveRegisters = [&](Type *Ty) {
+ return TLI.functionArgumentNeedsConsecutiveRegisters(
+ Ty, CallingConv::C, false, M->getDataLayout());
+ };
+ Type *I8 = Type::getInt8Ty(Context);
+ Type *I64 = Type::getInt64Ty(Context);
+ Type *F32 = Type::getFloatTy(Context);
+ Type *Empty = StructType::get(Context);
+ Type *Mixed = StructType::get(Context, {I8, F32});
+ Type *Huge = ArrayType::get(I8, 1ULL << 30);
+
+ EXPECT_TRUE(NeedsConsecutiveRegisters(Huge));
+ EXPECT_TRUE(NeedsConsecutiveRegisters(ArrayType::get(Huge, 2)));
+ EXPECT_TRUE(NeedsConsecutiveRegisters(
+ ArrayType::get(StructType::get(Context, {Huge, I8}), 2)));
+ EXPECT_FALSE(NeedsConsecutiveRegisters(
+ ArrayType::get(StructType::get(Context, {Huge, F32}), 2)));
+ EXPECT_FALSE(NeedsConsecutiveRegisters(ArrayType::get(Mixed, 2)));
+ EXPECT_TRUE(NeedsConsecutiveRegisters(ArrayType::get(Mixed, 0)));
+ EXPECT_TRUE(NeedsConsecutiveRegisters(ArrayType::get(Empty, 2)));
+ EXPECT_TRUE(NeedsConsecutiveRegisters(ArrayType::get(
+ StructType::get(Context, {I8, ArrayType::get(F32, 0), Empty}), 2)));
+ // Compare lowered value types, not IR types: pointers and i64 both use i64.
+ EXPECT_TRUE(NeedsConsecutiveRegisters(ArrayType::get(
+ StructType::get(Context, {PointerType::getUnqual(Context), I64}), 2)));
+ EXPECT_TRUE(NeedsConsecutiveRegisters(
+ ArrayType::get(FixedVectorType::get(F32, 4), 2)));
+ EXPECT_FALSE(NeedsConsecutiveRegisters(ArrayType::get(
+ StructType::get(Context, {FixedVectorType::get(F32, 4), F32}), 2)));
+ EXPECT_FALSE(NeedsConsecutiveRegisters(Mixed));
+ EXPECT_FALSE(NeedsConsecutiveRegisters(I64));
+ EXPECT_FALSE(NeedsConsecutiveRegisters(FixedVectorType::get(F32, 4)));
+ EXPECT_FALSE(NeedsConsecutiveRegisters(ScalableVectorType::get(F32, 4)));
+ EXPECT_TRUE(NeedsConsecutiveRegisters(ScalableVectorType::get(F32, 8)));
+}
+
TEST_F(AArch64SelectionDAGTest, computeKnownBits_ZERO_EXTEND_VECTOR_INREG) {
SDLoc Loc;
auto Int8VT = EVT::getIntegerVT(Context, 8);
More information about the llvm-commits
mailing list