[llvm] [AArch64] Add a preserve_none return convention (PR #223922)

via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 15 23:12:36 PDT 2026


llvmorg-github-actions[bot] wrote:


<!--LLVM PR SUMMARY COMMENT-->

@llvm/pr-subscribers-backend-aarch64

Author: vincent163 (vincent-163)

<details>
<summary>Changes</summary>

### Summary

Add `RetCC_AArch64_Preserve_None` so that the integer return values of a
`preserve_none` function are returned in the same registers that
`CC_AArch64_Preserve_None` uses for arguments, mirroring the behaviour of the
`preserve_none` convention on x86-64 and LoongArch.

### Motivation

Without a dedicated return convention, `preserve_none` falls back to the AAPCS
return convention, so a function that returns more than two values -- or an
aggregate larger than 16 bytes -- returns through a hidden pointer and memory.
That defeats the purpose of the convention for its intended users (runtimes
that hand a whole register file back to their caller, e.g. syscall/sysret style
entry points, where the return values should be left in the registers that the
runtime already has loaded).

### Details

* Register order mirrors the argument convention exactly: `W20`-`W28`,
  `W0`-`W7`, `W10`-`W14`, `W9` and, last, `W15` (or the `X` equivalents), so a
  `preserve_none` caller and callee agree on the return registers with no
  rematerialization, and forwarding arguments to the return values usually
  becomes a no-op.
* As in the argument convention, `X15` is only used on targets where it is not
  reserved by the stack probing prologue, that is everywhere except Windows;
  a Windows function can therefore return at most 23 values in registers, and a
  larger aggregate is returned indirectly.
* Pointers are bit-converted to `i64` first.  All other types, including
  floating point and vector types, keep the standard AAPCS return convention.
* `CCAssignFnForReturn()` returns the new convention for
  `CallingConv::PreserveNone`.

### Testing

* New test `llvm/test/CodeGen/AArch64/preserve_nonecc_return.ll`:
  * 24 `i64` values returned in registers (23 + indirect return on Windows),
  * a function forwarding 24 incoming arguments to 24 return values,
  * single `i64` and `double` returns,
  * a caller that consumes the returned registers.
* The existing tests whose expected return register changed are updated.
* `llvm/test/CodeGen/AArch64`: 10/10 targeted tests pass; the full suite runs
  2280 + 1952 tests with only pre-existing failures in this build environment
  (missing external tools such as `lli`, `llvm-objcopy`).


---

Patch is 24.46 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/223922.diff


9 Files Affected:

- (modified) llvm/lib/Target/AArch64/AArch64CallingConvention.h (+4) 
- (modified) llvm/lib/Target/AArch64/AArch64CallingConvention.td (+29) 
- (modified) llvm/lib/Target/AArch64/AArch64ISelLowering.cpp (+2) 
- (modified) llvm/test/CodeGen/AArch64/dynamic-regmask-preserve-none.ll (+5-4) 
- (modified) llvm/test/CodeGen/AArch64/preserve_nonecc_call.ll (+4-1) 
- (added) llvm/test/CodeGen/AArch64/preserve_nonecc_return.ll (+261) 
- (modified) llvm/test/CodeGen/AArch64/preserve_nonecc_varargs_aapcs.ll (+3-2) 
- (modified) llvm/test/CodeGen/AArch64/preserve_nonecc_varargs_darwin.ll (+2-1) 
- (modified) llvm/test/CodeGen/AArch64/preserve_nonecc_varargs_win64.ll (+3-3) 


``````````diff
diff --git a/llvm/lib/Target/AArch64/AArch64CallingConvention.h b/llvm/lib/Target/AArch64/AArch64CallingConvention.h
index 7105fa695334b..e8df44b88cc69 100644
--- a/llvm/lib/Target/AArch64/AArch64CallingConvention.h
+++ b/llvm/lib/Target/AArch64/AArch64CallingConvention.h
@@ -68,6 +68,10 @@ bool CC_AArch64_Preserve_None(unsigned ValNo, MVT ValVT, MVT LocVT,
 bool RetCC_AArch64_AAPCS(unsigned ValNo, MVT ValVT, MVT LocVT,
                          CCValAssign::LocInfo LocInfo, ISD::ArgFlagsTy ArgFlags,
                          Type *OrigTy, CCState &State);
+bool RetCC_AArch64_Preserve_None(unsigned ValNo, MVT ValVT, MVT LocVT,
+                                 CCValAssign::LocInfo LocInfo,
+                                 ISD::ArgFlagsTy ArgFlags, Type *OrigTy,
+                                 CCState &State);
 bool RetCC_AArch64_Arm64EC_Thunk(unsigned ValNo, MVT ValVT, MVT LocVT,
                                  CCValAssign::LocInfo LocInfo,
                                  ISD::ArgFlagsTy ArgFlags, Type *OrigTy,
diff --git a/llvm/lib/Target/AArch64/AArch64CallingConvention.td b/llvm/lib/Target/AArch64/AArch64CallingConvention.td
index 4683f166386e7..8b414b8ba5d48 100644
--- a/llvm/lib/Target/AArch64/AArch64CallingConvention.td
+++ b/llvm/lib/Target/AArch64/AArch64CallingConvention.td
@@ -161,6 +161,35 @@ def RetCC_AArch64_AAPCS : CallingConv<[
            CCAssignToReg<[P0, P1, P2, P3]>>
 ]>;
 
+let Entry = 1 in
+def RetCC_AArch64_Preserve_None : CallingConv<[
+  // Return values are assigned the same registers, in the same order, that
+  // CC_AArch64_Preserve_None assigns arguments to, so that a preserve_none
+  // caller and callee agree on the return registers without any
+  // rematerialization.
+  CCIfType<[iPTR], CCBitConvertToType<i64>>,
+  CCIfType<[i32], CCAssignToReg<[W20, W21, W22, W23,
+                                 W24, W25, W26, W27, W28,
+                                 W0, W1, W2, W3, W4, W5,
+                                 W6, W7, W10, W11,
+                                 W12, W13, W14, W9]>>,
+  CCIfType<[i64], CCAssignToReg<[X20, X21, X22, X23,
+                                 X24, X25, X26, X27, X28,
+                                 X0, X1, X2, X3, X4, X5,
+                                 X6, X7, X10, X11,
+                                 X12, X13, X14, X9]>>,
+
+  // Windows uses X15 for stack allocation
+  CCIf<"!State.getMachineFunction().getSubtarget<AArch64Subtarget>().isTargetWindows()",
+    CCIfType<[i32], CCAssignToReg<[W15]>>>,
+  CCIf<"!State.getMachineFunction().getSubtarget<AArch64Subtarget>().isTargetWindows()",
+    CCIfType<[i64], CCAssignToReg<[X15]>>>,
+
+  // All other types (including floating point and vector types) use the
+  // standard AAPCS return convention.
+  CCDelegateTo<RetCC_AArch64_AAPCS>
+]>;
+
 let Entry = 1 in
 def CC_AArch64_Win64PCS : CallingConv<!listconcat(
   [
diff --git a/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp b/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
index 2a50c0c474ab9..2e6a2cc930c07 100644
--- a/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
+++ b/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
@@ -9355,6 +9355,8 @@ AArch64TargetLowering::CCAssignFnForReturn(CallingConv::ID CC) const {
   switch (CC) {
   default:
     return RetCC_AArch64_AAPCS;
+  case CallingConv::PreserveNone:
+    return RetCC_AArch64_Preserve_None;
   case CallingConv::ARM64EC_Thunk_X64:
     return RetCC_AArch64_Arm64EC_Thunk;
   case CallingConv::CFGuard_Check:
diff --git a/llvm/test/CodeGen/AArch64/dynamic-regmask-preserve-none.ll b/llvm/test/CodeGen/AArch64/dynamic-regmask-preserve-none.ll
index 2d4fefe82b991..10cb127ebc77b 100644
--- a/llvm/test/CodeGen/AArch64/dynamic-regmask-preserve-none.ll
+++ b/llvm/test/CodeGen/AArch64/dynamic-regmask-preserve-none.ll
@@ -1,6 +1,7 @@
 ; RUN: llc -mtriple=aarch64-apple-darwin -stop-after finalize-isel <%s | FileCheck %s
 
-; Check that the callee doesn't have calleeSavedRegisters.
+; Check that the callee doesn't have calleeSavedRegisters, and that return
+; values use the argument registers (x20 onwards), as in LoongArch.
 define preserve_nonecc i64 @callee1(i64 %a0, i64 %b0, i64 %c0, i64 %d0, i64 %e0) nounwind {
   %a1 = mul i64 %a0, %b0
   %a2 = mul i64 %a1, %c0
@@ -10,7 +11,7 @@ define preserve_nonecc i64 @callee1(i64 %a0, i64 %b0, i64 %c0, i64 %d0, i64 %e0)
 }
 ; CHECK:     name: callee1
 ; CHECK-NOT: calleeSavedRegisters:
-; CHECK:     RET_ReallyLR implicit $x0
+; CHECK:     RET_ReallyLR implicit $x20
 
 ; Check that RegMask is csr_aarch64_noneregs.
 define i64 @caller1(i64 %a0) nounwind {
@@ -35,7 +36,7 @@ define preserve_nonecc {i64, i64} @callee2(i64 %a0, i64 %b0, i64 %c0, i64 %d0, i
 }
 ; CHECK:     name: callee2
 ; CHECK-NOT: calleeSavedRegisters:
-; CHECK:     RET_ReallyLR implicit $x0
+; CHECK:     RET_ReallyLR implicit $x20, implicit $x21
 
 
 ; Check that RegMask is csr_aarch64_noneregs.
@@ -76,7 +77,7 @@ define preserve_nonecc {i64, double} @callee4(i64 %a0, i64 %b0, i64 %c0, i64 %d0
 }
 ; CHECK:     name: callee4
 ; CHECK-NOT: calleeSavedRegisters:
-; CHECK:     RET_ReallyLR implicit $x0, implicit $d0
+; CHECK:     RET_ReallyLR implicit $x20, implicit $d0
 
 ; Check that RegMask is csr_aarch64_noneregs.
 define {i64, double} @caller4(i64 %a0) nounwind {
diff --git a/llvm/test/CodeGen/AArch64/preserve_nonecc_call.ll b/llvm/test/CodeGen/AArch64/preserve_nonecc_call.ll
index ca0139f5382dc..61e0a1293daaf 100644
--- a/llvm/test/CodeGen/AArch64/preserve_nonecc_call.ll
+++ b/llvm/test/CodeGen/AArch64/preserve_nonecc_call.ll
@@ -357,9 +357,10 @@ define i64 @caller3() {
 ; CHECK-NEXT:    mov w9, #23 // =0x17
 ; CHECK-NEXT:    mov w15, #24 // =0x18
 ; CHECK-NEXT:    bl callee_with_many_param
+; CHECK-NEXT:    mov x0, x20
 ; CHECK-NEXT:    ldp x20, x19, [sp, #144] // 16-byte Folded Reload
-; CHECK-NEXT:    ldr x30, [sp, #64] // 8-byte Reload
 ; CHECK-NEXT:    ldp x22, x21, [sp, #128] // 16-byte Folded Reload
+; CHECK-NEXT:    ldr x30, [sp, #64] // 8-byte Reload
 ; CHECK-NEXT:    ldp x24, x23, [sp, #112] // 16-byte Folded Reload
 ; CHECK-NEXT:    ldp x26, x25, [sp, #96] // 16-byte Folded Reload
 ; CHECK-NEXT:    ldp x28, x27, [sp, #80] // 16-byte Folded Reload
@@ -427,6 +428,7 @@ define i64 @caller3() {
 ; DARWIN-NEXT:    mov w9, #23 ; =0x17
 ; DARWIN-NEXT:    mov w15, #24 ; =0x18
 ; DARWIN-NEXT:    bl _callee_with_many_param
+; DARWIN-NEXT:    mov x0, x20
 ; DARWIN-NEXT:    ldp x29, x30, [sp, #144] ; 16-byte Folded Reload
 ; DARWIN-NEXT:    ldp x20, x19, [sp, #128] ; 16-byte Folded Reload
 ; DARWIN-NEXT:    ldp x22, x21, [sp, #112] ; 16-byte Folded Reload
@@ -491,6 +493,7 @@ define i64 @caller3() {
 ; WIN-NEXT:    mov w9, #23 // =0x17
 ; WIN-NEXT:    str x8, [sp]
 ; WIN-NEXT:    bl callee_with_many_param
+; WIN-NEXT:    mov x0, x20
 ; WIN-NEXT:    .seh_startepilogue
 ; WIN-NEXT:    ldp d14, d15, [sp, #152] // 16-byte Folded Reload
 ; WIN-NEXT:    .seh_save_fregp d14, 152
diff --git a/llvm/test/CodeGen/AArch64/preserve_nonecc_return.ll b/llvm/test/CodeGen/AArch64/preserve_nonecc_return.ll
new file mode 100644
index 0000000000000..2d528c00f44ac
--- /dev/null
+++ b/llvm/test/CodeGen/AArch64/preserve_nonecc_return.ll
@@ -0,0 +1,261 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py
+; RUN: llc -mtriple=aarch64 < %s | FileCheck %s
+; RUN: llc -mtriple=aarch64-windows < %s | FileCheck %s --check-prefix=WIN
+
+;; A preserve_none return value is returned in the same register that an
+;; argument of the same type would use, in the same order: x20-x28, x0-x7,
+;; x10-x14, x9 and, on non-Windows targets, x15.  x15 is reserved by the stack
+;; probing prologue on Windows, so a Windows function can return at most 23
+;; values in registers and a larger aggregate is returned indirectly.
+
+; All 24 registers of the pool are used for return values.
+define preserve_nonecc {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} @ret_24() {
+; CHECK-LABEL: ret_24:
+; CHECK:       // %bb.0:
+; CHECK-NEXT:    mov w20, #1 // =0x1
+; CHECK-NEXT:    mov w21, #2 // =0x2
+; CHECK-NEXT:    mov w22, #3 // =0x3
+; CHECK-NEXT:    mov w23, #4 // =0x4
+; CHECK-NEXT:    mov w24, #5 // =0x5
+; CHECK-NEXT:    mov w25, #6 // =0x6
+; CHECK-NEXT:    mov w26, #7 // =0x7
+; CHECK-NEXT:    mov w27, #8 // =0x8
+; CHECK-NEXT:    mov w28, #9 // =0x9
+; CHECK-NEXT:    mov w0, #10 // =0xa
+; CHECK-NEXT:    mov w1, #11 // =0xb
+; CHECK-NEXT:    mov w2, #12 // =0xc
+; CHECK-NEXT:    mov w3, #13 // =0xd
+; CHECK-NEXT:    mov w4, #14 // =0xe
+; CHECK-NEXT:    mov w5, #15 // =0xf
+; CHECK-NEXT:    mov w6, #16 // =0x10
+; CHECK-NEXT:    mov w7, #17 // =0x11
+; CHECK-NEXT:    mov w10, #18 // =0x12
+; CHECK-NEXT:    mov w11, #19 // =0x13
+; CHECK-NEXT:    mov w12, #20 // =0x14
+; CHECK-NEXT:    mov w13, #21 // =0x15
+; CHECK-NEXT:    mov w14, #22 // =0x16
+; CHECK-NEXT:    mov w9, #23 // =0x17
+; CHECK-NEXT:    mov w15, #24 // =0x18
+; CHECK-NEXT:    ret
+;
+; WIN-LABEL: ret_24:
+; WIN:       // %bb.0:
+; WIN-NEXT:    mov w8, #24 // =0x18
+; WIN-NEXT:    mov w9, #23 // =0x17
+; WIN-NEXT:    stp x9, x8, [x20, #176]
+; WIN-NEXT:    mov w8, #22 // =0x16
+; WIN-NEXT:    mov w9, #21 // =0x15
+; WIN-NEXT:    stp x9, x8, [x20, #160]
+; WIN-NEXT:    mov w8, #20 // =0x14
+; WIN-NEXT:    mov w9, #19 // =0x13
+; WIN-NEXT:    stp x9, x8, [x20, #144]
+; WIN-NEXT:    mov w8, #18 // =0x12
+; WIN-NEXT:    mov w9, #17 // =0x11
+; WIN-NEXT:    stp x9, x8, [x20, #128]
+; WIN-NEXT:    mov w8, #16 // =0x10
+; WIN-NEXT:    mov w9, #15 // =0xf
+; WIN-NEXT:    stp x9, x8, [x20, #112]
+; WIN-NEXT:    mov w8, #14 // =0xe
+; WIN-NEXT:    mov w9, #13 // =0xd
+; WIN-NEXT:    stp x9, x8, [x20, #96]
+; WIN-NEXT:    mov w8, #12 // =0xc
+; WIN-NEXT:    mov w9, #11 // =0xb
+; WIN-NEXT:    stp x9, x8, [x20, #80]
+; WIN-NEXT:    mov w8, #10 // =0xa
+; WIN-NEXT:    mov w9, #9 // =0x9
+; WIN-NEXT:    stp x9, x8, [x20, #64]
+; WIN-NEXT:    mov w8, #8 // =0x8
+; WIN-NEXT:    mov w9, #7 // =0x7
+; WIN-NEXT:    stp x9, x8, [x20, #48]
+; WIN-NEXT:    mov w8, #6 // =0x6
+; WIN-NEXT:    mov w9, #5 // =0x5
+; WIN-NEXT:    stp x9, x8, [x20, #32]
+; WIN-NEXT:    mov w8, #4 // =0x4
+; WIN-NEXT:    mov w9, #3 // =0x3
+; WIN-NEXT:    stp x9, x8, [x20, #16]
+; WIN-NEXT:    mov w8, #2 // =0x2
+; WIN-NEXT:    mov w9, #1 // =0x1
+; WIN-NEXT:    stp x9, x8, [x20]
+; WIN-NEXT:    ret
+  ret {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} {i64 1, i64 2, i64 3, i64 4, i64 5, i64 6, i64 7, i64 8, i64 9, i64 10, i64 11, i64 12, i64 13, i64 14, i64 15, i64 16, i64 17, i64 18, i64 19, i64 20, i64 21, i64 22, i64 23, i64 24}
+}
+
+; The incoming argument registers are already the outgoing return registers.
+define preserve_nonecc {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} @forward_24(i64 %a0, i64 %a1, i64 %a2, i64 %a3, i64 %a4, i64 %a5, i64 %a6, i64 %a7, i64 %a8, i64 %a9, i64 %a10, i64 %a11, i64 %a12, i64 %a13, i64 %a14, i64 %a15, i64 %a16, i64 %a17, i64 %a18, i64 %a19, i64 %a20, i64 %a21, i64 %a22, i64 %a23) {
+; CHECK-LABEL: forward_24:
+; CHECK:       // %bb.0:
+; CHECK-NEXT:    ret
+;
+; WIN-LABEL: forward_24:
+; WIN:       // %bb.0:
+; WIN-NEXT:    ldp x15, x8, [sp]
+; WIN-NEXT:    stp x14, x9, [x20, #160]
+; WIN-NEXT:    stp x12, x13, [x20, #144]
+; WIN-NEXT:    stp x10, x11, [x20, #128]
+; WIN-NEXT:    stp x15, x8, [x20, #176]
+; WIN-NEXT:    stp x6, x7, [x20, #112]
+; WIN-NEXT:    stp x4, x5, [x20, #96]
+; WIN-NEXT:    stp x2, x3, [x20, #80]
+; WIN-NEXT:    stp x0, x1, [x20, #64]
+; WIN-NEXT:    stp x27, x28, [x20, #48]
+; WIN-NEXT:    stp x25, x26, [x20, #32]
+; WIN-NEXT:    stp x23, x24, [x20, #16]
+; WIN-NEXT:    stp x21, x22, [x20]
+; WIN-NEXT:    ret
+  %v0 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} poison, i64 %a0, 0
+  %v1 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v0, i64 %a1, 1
+  %v2 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v1, i64 %a2, 2
+  %v3 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v2, i64 %a3, 3
+  %v4 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v3, i64 %a4, 4
+  %v5 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v4, i64 %a5, 5
+  %v6 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v5, i64 %a6, 6
+  %v7 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v6, i64 %a7, 7
+  %v8 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v7, i64 %a8, 8
+  %v9 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v8, i64 %a9, 9
+  %v10 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v9, i64 %a10, 10
+  %v11 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v10, i64 %a11, 11
+  %v12 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v11, i64 %a12, 12
+  %v13 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v12, i64 %a13, 13
+  %v14 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v13, i64 %a14, 14
+  %v15 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v14, i64 %a15, 15
+  %v16 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v15, i64 %a16, 16
+  %v17 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v16, i64 %a17, 17
+  %v18 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v17, i64 %a18, 18
+  %v19 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v18, i64 %a19, 19
+  %v20 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v19, i64 %a20, 20
+  %v21 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v20, i64 %a21, 21
+  %v22 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v21, i64 %a22, 22
+  %v23 = insertvalue {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v22, i64 %a23, 23
+  ret {i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64} %v23
+}
+
+; A single value is returned in x20 rather than x0.
+define preserve_nonecc i64 @ret_i64() {
+; CHECK-LABEL: ret_i64:
+; CHECK:       // %bb.0:
+; CHECK-NEXT:    mov w20, #1 // =0x1
+; CHECK-NEXT:    ret
+;
+; WIN-LABEL: ret_i64:
+; WIN:       // %bb.0:
+; WIN-NEXT:    mov w20, #1 // =0x1
+; WIN-NEXT:    ret
+  ret i64 1
+}
+
+; Floating point and vector return values keep the AAPCS convention.
+define preserve_nonecc double @ret_double() {
+; CHECK-LABEL: ret_double:
+; CHECK:       // %bb.0:
+; CHECK-NEXT:    fmov d0, #1.00000000
+; CHECK-NEXT:    ret
+;
+; WIN-LABEL: ret_double:
+; WIN:       // %bb.0:
+; WIN-NEXT:    fmov d0, #1.00000000
+; WIN-NEXT:    ret
+  ret double 1.0
+}
+
+; The caller reads the return value from the preserve_none return register.
+define i64 @caller() {
+; CHECK-LABEL: caller:
+; CHECK:       // %bb.0:
+; CHECK-NEXT:    stp d15, d14, [sp, #-160]! // 16-byte Folded Spill
+; CHECK-NEXT:    stp d13, d12, [sp, #16] // 16-byte Folded Spill
+; CHECK-NEXT:    stp d11, d10, [sp, #32] // 16-byte Folded Spill
+; CHECK-NEXT:    stp d9, d8, [sp, #48] // 16-byte Folded Spill
+; CHECK-NEXT:    str x30, [sp, #64] // 8-byte Spill
+; CHECK-NEXT:    stp x28, x27, [sp, #80] // 16-byte Folded Spill
+; CHECK-NEXT:    stp x26, x25, [sp, #96] // 16-byte Folded Spill
+; CHECK-NEXT:    stp x24, x23, [sp, #112] // 16-byte Folded Spill
+; CHECK-NEXT:    stp x22, x21, [sp, #128] // 16-byte Folded Spill
+; CHECK-NEXT:    stp x20, x19, [sp, #144] // 16-byte Folded Spill
+; CHECK-NEXT:    .cfi_def_cfa_offset 160
+; CHECK-NEXT:    .cfi_offset w19, -8
+; CHECK-NEXT:    .cfi_offset w20, -16
+; CHECK-NEXT:    .cfi_offset w21, -24
+; CHECK-NEXT:    .cfi_offset w22, -32
+; CHECK-NEXT:    .cfi_offset w23, -40
+; CHECK-NEXT:    .cfi_offset w24, -48
+; CHECK-NEXT:    .cfi_offset w25, -56
+; CHECK-NEXT:    .cfi_offset w26, -64
+; CHECK-NEXT:    .cfi_offset w27, -72
+; CHECK-NEXT:    .cfi_offset w28, -80
+; CHECK-NEXT:    .cfi_offset w30, -96
+; CHECK-NEXT:    .cfi_offset b8, -104
+; CHECK-NEXT:    .cfi_offset b9, -112
+; CHECK-NEXT:    .cfi_offset b10, -120
+; CHECK-NEXT:    .cfi_offset b11, -128
+; CHECK-NEXT:    .cfi_offset b12, -136
+; CHECK-NEXT:    .cfi_offset b13, -144
+; CHECK-NEXT:    .cfi_offset b14, -152
+; CHECK-NEXT:    .cfi_offset b15, -160
+; CHECK-NEXT:    bl ret_i64
+; CHECK-NEXT:    mov x0, x20
+; CHECK-NEXT:    ldp x20, x19, [sp, #144] // 16-byte Folded Reload
+; CHECK-NEXT:    ldp x22, x21, [sp, #128] // 16-byte Folded Reload
+; CHECK-NEXT:    ldr x30, [sp, #64] // 8-byte Reload
+; CHECK-NEXT:    ldp x24, x23, [sp, #112] // 16-byte Folded Reload
+; CHECK-NEXT:    ldp x26, x25, [sp, #96] // 16-byte Folded Reload
+; CHECK-NEXT:    ldp x28, x27, [sp, #80] // 16-byte Folded Reload
+; CHECK-NEXT:    ldp d9, d8, [sp, #48] // 16-byte Folded Reload
+; CHECK-NEXT:    ldp d11, d10, [sp, #32] // 16-byte Folded Reload
+; CHECK-NEXT:    ldp d13, d12, [sp, #16] // 16-byte Folded Reload
+; CHECK-NEXT:    ldp d15, d14, [sp], #160 // 16-byte Folded Reload
+; CHECK-NEXT:    ret
+;
+; WIN-LABEL: caller:
+; WIN:       .seh_proc caller
+; WIN-NEXT:  // %bb.0:
+; WIN-NEXT:    stp x19, x20, [sp, #-160]! // 16-byte Folded Spill
+; WIN-NEXT:    .seh_save_regp_x x19, 160
+; WIN-NEXT:    stp x21, x22, [sp, #16] // 16-byte Folded Spill
+; WIN-NEXT:    .seh_save_regp x21, 16
+; WIN-NEXT:    stp x23, x24, [sp, #32] // 16-byte Folded Spill
+; WIN-NEXT:    .seh_save_regp x23, 32
+; WIN-NEXT:    stp x25, x26, [sp, #48] // 16-byte Folded Spill
+; WIN-NEXT:    .seh_save_regp x25, 48
+; WIN-NEXT:    stp x27, x28, [sp, #64] // 16-byte Folded Spill
+; WIN-NEXT:    .seh_save_regp x27, 64
+; WIN-NEXT:    str x30, [sp, #80] // 8-byte Spill
+; WIN-NEXT:    .seh_save_reg x30, 80
+; WIN-NEXT:    stp d8, d9, [sp, #88] // 16-byte Folded Spill
+; WIN-NEXT:    .seh_save_fregp d8, 88
+; WIN-NEXT:    stp d10, d11, [sp, #104] // 16-byte Folded Spill
+; WIN-NEXT:    .seh_save_fregp d10, 104
+; WIN-NEXT:    stp d12, d13, [sp, #120] // 16-byte Folded Spill
+; WIN-NEXT:    .seh_save_fregp d12, 120
+; WIN-NEXT:    stp d14, d15, [sp, #136] // 16-byte Folded Spill
+; WIN-NEXT:    .seh_save_fregp d14, 136
+; WIN-NEXT:    .seh_endprologue
+; WIN-NEXT:    bl ret_i64
+; WIN-NEXT:    mov x0, x20
+; WIN-NEXT:    .seh_startepilogue
+; WIN-NEXT:    ldp d14, d15, [sp, #136] // 16-byte Folded Reload
+; WIN-NEXT:    .seh_save_fregp d14, 136
+; WIN-NEXT:    ldp d12, d13, [sp, #120] // 16-byte Folded Reload
+; WIN-NEXT:    .seh_save_fregp d12, 120
+; WIN-NEXT:    ldp d10, d11, [sp, #104] // 16-byte Folded Reload
+; WIN-NEXT:    .seh_save_fregp d10, 104
+; WIN-NEXT:    ldp d8, d9, [sp, #88] // 16-byte Folded Reload
+; WIN-NEXT...
[truncated]

``````````

</details>


https://github.com/llvm/llvm-project/pull/223922


More information about the llvm-commits mailing list