[llvm] [AArch64] Avoid expanding arrays to classify consecutive argument registers (PR #226887)
via llvm-commits
llvm-commits at lists.llvm.org
Sun Sep 27 21:39:57 PDT 2026
llvmorg-github-actions[bot] wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-backend-aarch64
Author: zhouguangyuan0718
<details>
<summary>Changes</summary>
AArch64's `functionArgumentNeedsConsecutiveRegisters` expands every array element with `ComputeValueVTs` just to check whether all member EVTs are equal. An unused large `byval` array therefore consumes memory proportional to its element count during `SelectionDAGISel::LowerArguments`.
For example, this hand-written IR uses the ordinary C calling convention:
```llvm
define i32 @<!-- -->large_array(ptr byval([1073741808 x i8]) align 1 %arg) {
ret i32 1
}
```
```sh
llc -mtriple=aarch64-linux-gnu -global-isel=0 -O2 -verify-machineinstrs repro.ll -o out.s
```
On unmodified main `b1f669055159`, this exceeded a 1 GiB RSS watchdog threshold in about 0.82 seconds. The watchdog deliberately sent SIGABRT rather than allowing an unbounded allocation; the resulting stack contains `ComputeValueTypes -> ComputeValueVTs -> functionArgumentNeedsConsecutiveRegisters -> LowerArguments`. After this change, the large and nested-array regression compiles at O0/O2 with roughly 25 MiB peak RSS.
Walk distinct aggregate types and compare their lowered EVTs instead. Nonempty arrays only need their element type inspected once; empty arrays/structs contribute no members. This retains pointer/integer EVT equivalence and leaves the existing non-array/scalable-vector rule unchanged. No calling-convention register assignment or optimization settings are changed.
This is **not a C-source reproducer**. Unmodified Clang 24.0.0git successfully compiled C array parameters and large struct/nested struct/mixed struct parameters at O0/O2 for AArch64 Linux, big-endian Linux, Darwin, Windows, and FreeBSD. Those large parameters lower to ordinary pointers without `byval`, so the tested C forms do not reach this case. The regression exercises the generic LLVM IR interface directly.
Validation on the rebased patch (base `a98dcad4fed0`, Release with assertions, macOS arm64):
- All 102 AArch64 unit tests pass.
- Complete `CodeGen/AArch64` suite: 4,261 passed, 4 expected failures, no unexpected failures (4,265 tests).
- The new large/nested-array regression passes at O0/O2 with machine verification; final measurements are 21.95/24.39 MiB peak RSS.
- 32 assembly comparisons are byte-identical before/after, covering Clang-generated IR, ordinary/mixed/empty/nested aggregates, pointer/integer members, fixed vectors, and SVE calling conventions.
- Added generic O0/O2 SelectionDAG coverage and predicate unit coverage; no Go-specific tests or metadata.
The implementation is adapted from the downstream fix in https://github.com/goallc/llvm-project/pull/170. This patch only bounds the classification query; it does not establish that arbitrary gigabyte-sized calls can execute successfully.
AI assistance: OpenAI Codex was used to adapt the implementation, write tests and this description, and run validation, under the contributor's explicit authorization.
---
Full diff: https://github.com/llvm/llvm-project/pull/226887.diff
3 Files Affected:
- (modified) llvm/lib/Target/AArch64/AArch64ISelLowering.cpp (+24-4)
- (added) llvm/test/CodeGen/AArch64/large-byval-array.ll (+18)
- (modified) llvm/unittests/Target/AArch64/AArch64SelectionDAGTest.cpp (+38)
``````````diff
The server is unavailable at this time. Please wait a few minutes before you try again.
``````````
</details>
https://github.com/llvm/llvm-project/pull/226887
More information about the llvm-commits
mailing list