[llvm] [AArch64] Avoid expanding arrays to classify consecutive argument registers (PR #226887)

via llvm-commits llvm-commits at lists.llvm.org
Sun Sep 27 21:39:57 PDT 2026


llvmorg-github-actions[bot] wrote:


<!--LLVM PR SUMMARY COMMENT-->

@llvm/pr-subscribers-backend-aarch64

Author: zhouguangyuan0718

<details>
<summary>Changes</summary>

AArch64's `functionArgumentNeedsConsecutiveRegisters` expands every array element with `ComputeValueVTs` just to check whether all member EVTs are equal. An unused large `byval` array therefore consumes memory proportional to its element count during `SelectionDAGISel::LowerArguments`.

For example, this hand-written IR uses the ordinary C calling convention:

```llvm
define i32 @<!-- -->large_array(ptr byval([1073741808 x i8]) align 1 %arg) {
  ret i32 1
}
```

```sh
llc -mtriple=aarch64-linux-gnu -global-isel=0 -O2 -verify-machineinstrs repro.ll -o out.s
```

On unmodified main `b1f669055159`, this exceeded a 1 GiB RSS watchdog threshold in about 0.82 seconds. The watchdog deliberately sent SIGABRT rather than allowing an unbounded allocation; the resulting stack contains `ComputeValueTypes -> ComputeValueVTs -> functionArgumentNeedsConsecutiveRegisters -> LowerArguments`. After this change, the large and nested-array regression compiles at O0/O2 with roughly 25 MiB peak RSS.

Walk distinct aggregate types and compare their lowered EVTs instead. Nonempty arrays only need their element type inspected once; empty arrays/structs contribute no members. This retains pointer/integer EVT equivalence and leaves the existing non-array/scalable-vector rule unchanged. No calling-convention register assignment or optimization settings are changed.

This is **not a C-source reproducer**. Unmodified Clang 24.0.0git successfully compiled C array parameters and large struct/nested struct/mixed struct parameters at O0/O2 for AArch64 Linux, big-endian Linux, Darwin, Windows, and FreeBSD. Those large parameters lower to ordinary pointers without `byval`, so the tested C forms do not reach this case. The regression exercises the generic LLVM IR interface directly.

Validation on the rebased patch (base `a98dcad4fed0`, Release with assertions, macOS arm64):

- All 102 AArch64 unit tests pass.
- Complete `CodeGen/AArch64` suite: 4,261 passed, 4 expected failures, no unexpected failures (4,265 tests).
- The new large/nested-array regression passes at O0/O2 with machine verification; final measurements are 21.95/24.39 MiB peak RSS.
- 32 assembly comparisons are byte-identical before/after, covering Clang-generated IR, ordinary/mixed/empty/nested aggregates, pointer/integer members, fixed vectors, and SVE calling conventions.
- Added generic O0/O2 SelectionDAG coverage and predicate unit coverage; no Go-specific tests or metadata.

The implementation is adapted from the downstream fix in https://github.com/goallc/llvm-project/pull/170. This patch only bounds the classification query; it does not establish that arbitrary gigabyte-sized calls can execute successfully.

AI assistance: OpenAI Codex was used to adapt the implementation, write tests and this description, and run validation, under the contributor's explicit authorization.


---
Full diff: https://github.com/llvm/llvm-project/pull/226887.diff


3 Files Affected:

- (modified) llvm/lib/Target/AArch64/AArch64ISelLowering.cpp (+24-4) 
- (added) llvm/test/CodeGen/AArch64/large-byval-array.ll (+18) 
- (modified) llvm/unittests/Target/AArch64/AArch64SelectionDAGTest.cpp (+38) 


``````````diff
The server is unavailable at this time. Please wait a few minutes before you try again.
``````````

</details>


https://github.com/llvm/llvm-project/pull/226887


More information about the llvm-commits mailing list