[llvm] [Xtensa] Fix addressing in dynamically realigned windowed-ABI frames (PR #208947)

via llvm-commits llvm-commits at lists.llvm.org
Sat Jul 11 13:07:32 PDT 2026


llvmorg-github-actions[bot] wrote:


<!--LLVM PR SUMMARY COMMENT-->

@llvm/pr-subscribers-backend-xtensa

Author: Randall Nortman (rnortman)

<details>
<summary>Changes</summary>

Fixes #<!-- -->208946 

When a frame is dynamically realigned (an over-aligned stack object with MaxAlignment > 32), the prologue moves SP in place by a runtime-variable pad. Two addressing bugs follow from that:

1. Incoming stack arguments (fixed frame objects) were still resolved as "SP + StackSize + offset", which is only correct against the SP value at entry. Every stack-passed argument was read displaced by the pad, i.e. through a garbage pointer. Fix: when a realigned windowed-ABI function has fixed stack objects, reserve a slot, store the entry SP into it right after realignment, and reload it in eliminateFrameIndex as the base register for fixed objects. Functions that don't hit the realign+stack-args conjunction generate unchanged code.

2. The realignment pad itself was computed from getFrameRegister(), i.e. from A7 when the function also has dynamic allocas -- before A7 is set up as the frame pointer, while it still holds the sixth argument. The frame was realigned by a garbage amount, breaking the alignment guarantee. Fix: compute the pad from SP, which getFrameRegister() aliased in the only case that worked.

The entry-SP slot is a local object: it is stored after FP setup and reloaded via the standard frame register (FP when the function has one, which stays fixed across MOVSP for dynamic allocas; SP otherwise). If the slot's offset ever exceeds the L32I immediate range the reload materializes the offset instead; that branch is defensive -- frame layout places this late-created slot near SP in all constructible cases.

---
Full diff: https://github.com/llvm/llvm-project/pull/208947.diff


4 Files Affected:

- (modified) llvm/lib/Target/Xtensa/XtensaFrameLowering.cpp (+38-2) 
- (modified) llvm/lib/Target/Xtensa/XtensaMachineFunctionInfo.h (+7) 
- (modified) llvm/lib/Target/Xtensa/XtensaRegisterInfo.cpp (+43) 
- (added) llvm/test/CodeGen/Xtensa/frame-realign-stack-args.ll (+297) 


``````````diff
diff --git a/llvm/lib/Target/Xtensa/XtensaFrameLowering.cpp b/llvm/lib/Target/Xtensa/XtensaFrameLowering.cpp
index 1c0dc66a46144..d97a4f1174b89 100644
--- a/llvm/lib/Target/Xtensa/XtensaFrameLowering.cpp
+++ b/llvm/lib/Target/Xtensa/XtensaFrameLowering.cpp
@@ -89,19 +89,22 @@ void XtensaFrameLowering::emitPrologue(MachineFunction &MF,
     // new_offset = SP + diff_to_128_aligned_address
     // This is safe to do because we increased the stack size by MaxAlignment.
     MCRegister Reg, RegMisAlign;
+    int SP0FI = XtensaFI->getRealignSP0FrameIndex();
     if (MaxAlignment > 32) {
       TII.loadImmediate(MBB, MBBI, &RegMisAlign, MaxAlignment - 1);
       TII.loadImmediate(MBB, MBBI, &Reg, MaxAlignment);
       BuildMI(MBB, MBBI, DL, TII.get(Xtensa::AND))
           .addReg(RegMisAlign, RegState::Define)
-          .addReg(FP)
+          .addReg(SP)
           .addReg(RegMisAlign);
       BuildMI(MBB, MBBI, DL, TII.get(Xtensa::SUB), RegMisAlign)
           .addReg(Reg)
           .addReg(RegMisAlign);
+      // pad == RegMisAlign; realign upward: SP = SP + pad. Keep pad live if it
+      // is still needed below to reconstruct the original SP.
       BuildMI(MBB, MBBI, DL, TII.get(Xtensa::ADD), SP)
           .addReg(SP)
-          .addReg(RegMisAlign, RegState::Kill);
+          .addReg(RegMisAlign, getKillRegState(SP0FI == -1));
     }
 
     // Store FP register in A8, because FP may be used to pass function
@@ -132,6 +135,24 @@ void XtensaFrameLowering::emitPrologue(MachineFunction &MF,
       BuildMI(MBB, MBBI, DL, TII.get(TargetOpcode::CFI_INSTRUCTION))
           .addCFIIndex(CFIIndex);
     }
+
+    // Incoming stack arguments are addressed relative to the caller's stack
+    // pointer, which equals the value of SP right after ENTRY -- i.e. the SP
+    // we just realigned away from. Recover it (original SP = realigned SP -
+    // pad) and stash it in its slot; eliminateFrameIndex reloads it to resolve
+    // those (fixed) frame objects. Emitted after the frame pointer setup so
+    // that the slot's frame index resolves correctly when hasFP.
+    if (SP0FI != -1) {
+      assert(MaxAlignment > 32 &&
+             "entry-SP slot reserved for a frame that is not realigned");
+      BuildMI(MBB, MBBI, DL, TII.get(Xtensa::SUB), RegMisAlign)
+          .addReg(SP)
+          .addReg(RegMisAlign, RegState::Kill);
+      BuildMI(MBB, MBBI, DL, TII.get(Xtensa::S32I))
+          .addReg(RegMisAlign, RegState::Kill)
+          .addFrameIndex(SP0FI)
+          .addImm(0);
+    }
   } else {
     // No need to allocate space on the stack.
     if (StackSize == 0 && !MFI.adjustsStack())
@@ -376,6 +397,21 @@ void XtensaFrameLowering::processFunctionBeforeFrameFinalized(
   if (IsLargeFunction)
     ScavSlotsNum = std::max(ScavSlotsNum, 1u);
 
+  // A dynamically realigned frame (windowed ABI) moves SP away from the value
+  // it had on entry, which is the base incoming stack arguments are addressed
+  // against. If the function has any incoming stack (fixed) arguments, reserve
+  // a slot to preserve the entry SP; eliminateFrameIndex reloads it through a
+  // scavenged register. Two scavenging slots: the entry-SP reload register can
+  // be live at the same point as the register the generic path scavenges to
+  // materialize an out-of-range offset.
+  if (STI.isWindowedABI() && MFI.getMaxAlign().value() > 32 &&
+      MFI.getObjectIndexBegin() < 0) {
+    const TargetRegisterClass &RC = Xtensa::ARRegClass;
+    XtensaFI->setRealignSP0FrameIndex(MFI.CreateStackObject(
+        TRI->getSpillSize(RC), TRI->getSpillAlign(RC), false));
+    ScavSlotsNum = std::max(ScavSlotsNum, 2u);
+  }
+
   const TargetRegisterClass &RC = Xtensa::ARRegClass;
   unsigned Size = TRI->getSpillSize(RC);
   Align Alignment = TRI->getSpillAlign(RC);
diff --git a/llvm/lib/Target/Xtensa/XtensaMachineFunctionInfo.h b/llvm/lib/Target/Xtensa/XtensaMachineFunctionInfo.h
index ff3bba0985c27..55e317b17236a 100644
--- a/llvm/lib/Target/Xtensa/XtensaMachineFunctionInfo.h
+++ b/llvm/lib/Target/Xtensa/XtensaMachineFunctionInfo.h
@@ -29,6 +29,10 @@ class XtensaMachineFunctionInfo : public MachineFunctionInfo {
   int VarArgsInRegsFrameIndex;
   bool SaveFrameRegister = false;
   unsigned CPLabelId = 0;
+  /// FrameIndex of the slot that holds the original (pre-realignment) stack
+  /// pointer. Used to reference incoming (caller-relative) stack arguments in
+  /// functions whose frame is dynamically realigned. -1 when not needed.
+  int RealignSP0FrameIndex = -1;
 
 public:
   explicit XtensaMachineFunctionInfo(const Function &F,
@@ -56,6 +60,9 @@ class XtensaMachineFunctionInfo : public MachineFunctionInfo {
   bool isSaveFrameRegister() const { return SaveFrameRegister; }
   void setSaveFrameRegister() { SaveFrameRegister = true; }
 
+  int getRealignSP0FrameIndex() const { return RealignSP0FrameIndex; }
+  void setRealignSP0FrameIndex(int FI) { RealignSP0FrameIndex = FI; }
+
   unsigned createCPLabelId() { return CPLabelId++; }
 };
 
diff --git a/llvm/lib/Target/Xtensa/XtensaRegisterInfo.cpp b/llvm/lib/Target/Xtensa/XtensaRegisterInfo.cpp
index ed76431bf5495..29c9efe496c3a 100644
--- a/llvm/lib/Target/Xtensa/XtensaRegisterInfo.cpp
+++ b/llvm/lib/Target/Xtensa/XtensaRegisterInfo.cpp
@@ -13,11 +13,13 @@
 #include "XtensaRegisterInfo.h"
 #include "MCTargetDesc/XtensaMCTargetDesc.h"
 #include "XtensaInstrInfo.h"
+#include "XtensaMachineFunctionInfo.h"
 #include "XtensaSubtarget.h"
 #include "llvm/CodeGen/MachineFrameInfo.h"
 #include "llvm/CodeGen/MachineFunction.h"
 #include "llvm/CodeGen/MachineInstrBuilder.h"
 #include "llvm/CodeGen/MachineRegisterInfo.h"
+#include "llvm/CodeGen/RegisterScavenging.h"
 #include "llvm/Support/Debug.h"
 #include "llvm/Support/ErrorHandling.h"
 #include "llvm/Support/raw_ostream.h"
@@ -92,6 +94,47 @@ bool XtensaRegisterInfo::eliminateFrameIndex(MachineBasicBlock::iterator II,
   else
     FrameReg = getFrameRegister(MF);
 
+  // In a dynamically realigned WindowedABI frame, SP no longer holds the value
+  // it had on entry (realignment moved it by a runtime-variable pad). Incoming
+  // stack arguments are fixed objects addressed relative to that entry SP, so
+  // resolving them against the realigned SP would be wrong by the pad. Reload
+  // the preserved original SP from its slot and use it as the base register.
+  {
+    auto *XtensaFI = MF.getInfo<XtensaMachineFunctionInfo>();
+    if (Subtarget.isWindowedABI() && MFI.getMaxAlign().value() > 32 &&
+        MFI.isFixedObjectIndex(FrameIndex) &&
+        XtensaFI->getRealignSP0FrameIndex() != -1) {
+      int SP0FI = XtensaFI->getRealignSP0FrameIndex();
+      int64_t SP0Off = MFI.getObjectOffset(SP0FI) + (int64_t)StackSize;
+      MachineBasicBlock &MBB = *MI.getParent();
+      DebugLoc DL = II->getDebugLoc();
+      const XtensaInstrInfo &TII = *static_cast<const XtensaInstrInfo *>(
+          MBB.getParent()->getSubtarget().getInstrInfo());
+      Register SP0Reg =
+          MF.getRegInfo().createVirtualRegister(&Xtensa::ARRegClass);
+      // The slot is a local object: address it the same way the generic path
+      // below addresses locals -- off FrameReg (FP when the function has one,
+      // which stays fixed while SP moves for dynamic allocas; SP otherwise).
+      if (Xtensa::isValidAddrOffsetForOpcode(Xtensa::L32I, SP0Off)) {
+        BuildMI(MBB, II, DL, TII.get(Xtensa::L32I), SP0Reg)
+            .addReg(FrameReg)
+            .addImm(SP0Off);
+      } else {
+        // The slot offset does not fit the L32I immediate; materialize it and
+        // index off FrameReg through a temporary register.
+        MCRegister OffReg;
+        TII.loadImmediate(MBB, II, &OffReg, SP0Off);
+        BuildMI(MBB, II, DL, TII.get(Xtensa::ADD), OffReg)
+            .addReg(FrameReg)
+            .addReg(OffReg, RegState::Kill);
+        BuildMI(MBB, II, DL, TII.get(Xtensa::L32I), SP0Reg)
+            .addReg(OffReg, RegState::Kill)
+            .addImm(0);
+      }
+      FrameReg = SP0Reg;
+    }
+  }
+
   // Calculate final offset.
   // - There is no need to change the offset if the frame object is one of the
   //   following: an outgoing argument, pointer to a dynamically allocated
diff --git a/llvm/test/CodeGen/Xtensa/frame-realign-stack-args.ll b/llvm/test/CodeGen/Xtensa/frame-realign-stack-args.ll
new file mode 100644
index 0000000000000..a85c9aa4ed3f9
--- /dev/null
+++ b/llvm/test/CodeGen/Xtensa/frame-realign-stack-args.ll
@@ -0,0 +1,297 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 5
+; RUN: llc -mtriple=xtensa -O2 -mattr=+windowed -verify-machineinstrs < %s \
+; RUN:   | FileCheck %s -check-prefix=XTENSA
+
+; A windowed-ABI function whose frame is dynamically realigned (an over-aligned
+; local forces SP to be realigned in place after ENTRY) must still address its
+; incoming stack-passed arguments relative to the stack pointer as it was on
+; entry. The realignment moves SP by a runtime-variable pad, so resolving the
+; incoming (fixed) argument slots against the moved SP is wrong by that pad.
+;
+; The backend preserves the original entry SP in a reserved slot right after the
+; realignment and reloads it to address incoming stack arguments.
+
+declare void @sink(ptr)
+
+; (a) realigned frame AND (b) reads incoming stack-passed args %g/%h/%i.
+; The stores to %g/%h/%i must be based on the reloaded entry SP, not the
+; realigned SP.
+define i32 @victim(i32 %a, i32 %b, i32 %c, i32 %d, i32 %e, i32 %f,
+; XTENSA-LABEL: victim:
+; XTENSA:         .cfi_startproc
+; XTENSA-NEXT:  # %bb.0:
+; XTENSA-NEXT:    entry a1, 224
+; XTENSA-NEXT:    movi a8, 63
+; XTENSA-NEXT:    movi a9, 64
+; XTENSA-NEXT:    and a8, a1, a8
+; XTENSA-NEXT:    sub a8, a9, a8
+; XTENSA-NEXT:    add a1, a1, a8
+; XTENSA-NEXT:    .cfi_def_cfa_offset 224
+; XTENSA-NEXT:    sub a8, a1, a8
+; XTENSA-NEXT:    s32i a8, a1, 60
+; XTENSA-NEXT:    addi a7, a1, 64
+; XTENSA-NEXT:    l32r a8, .LCPI0_0
+; XTENSA-NEXT:    or a10, a7, a7
+; XTENSA-NEXT:    callx8 a8
+; XTENSA-NEXT:    l32i a8, a1, 60
+; XTENSA-NEXT:    l32i a8, a8, 224
+; XTENSA-NEXT:    movi a9, 0
+; XTENSA-NEXT:    s8i a9, a8, 0
+; XTENSA-NEXT:    l32i a8, a1, 60
+; XTENSA-NEXT:    l32i a8, a8, 228
+; XTENSA-NEXT:    s8i a9, a8, 0
+; XTENSA-NEXT:    l32i a8, a1, 60
+; XTENSA-NEXT:    l32i a8, a8, 232
+; XTENSA-NEXT:    s8i a9, a8, 0
+; XTENSA-NEXT:    l32r a8, .LCPI0_1
+; XTENSA-NEXT:    or a10, a7, a7
+; XTENSA-NEXT:    callx8 a8
+; XTENSA-NEXT:    retw
+                   ptr %g, ptr %h, ptr %i) {
+  %buf = alloca [64 x i8], align 64
+  call void @sink(ptr %buf)
+  store i8 0, ptr %g
+  store i8 0, ptr %h
+  store i8 0, ptr %i
+  call void @sink(ptr %buf)
+  ret i32 %a
+}
+
+; (a) only: realigned frame, single arg stays in a register -> no incoming stack
+; object, so no original-SP slot is reserved and no recovery is emitted.
+define i32 @ctrl_align_only(ptr %g) {
+; XTENSA-LABEL: ctrl_align_only:
+; XTENSA:         .cfi_startproc
+; XTENSA-NEXT:  # %bb.0:
+; XTENSA-NEXT:    entry a1, 160
+; XTENSA-NEXT:    movi a8, 63
+; XTENSA-NEXT:    movi a9, 64
+; XTENSA-NEXT:    and a8, a1, a8
+; XTENSA-NEXT:    sub a8, a9, a8
+; XTENSA-NEXT:    add a1, a1, a8
+; XTENSA-NEXT:    .cfi_def_cfa_offset 160
+; XTENSA-NEXT:    addi a7, a1, 0
+; XTENSA-NEXT:    l32r a8, .LCPI1_0
+; XTENSA-NEXT:    or a10, a7, a7
+; XTENSA-NEXT:    callx8 a8
+; XTENSA-NEXT:    movi a6, 0
+; XTENSA-NEXT:    s8i a6, a2, 0
+; XTENSA-NEXT:    l32r a8, .LCPI1_1
+; XTENSA-NEXT:    or a10, a7, a7
+; XTENSA-NEXT:    callx8 a8
+; XTENSA-NEXT:    or a2, a6, a6
+; XTENSA-NEXT:    retw
+  %buf = alloca [64 x i8], align 64
+  call void @sink(ptr %buf)
+  store i8 0, ptr %g
+  call void @sink(ptr %buf)
+  ret i32 0
+}
+
+; (b) only: incoming stack args but ordinary alignment -> SP is not realigned,
+; so the incoming args are addressed off the unmoved SP directly.
+define i32 @ctrl_args_only(i32 %a, i32 %b, i32 %c, i32 %d, i32 %e, i32 %f,
+; XTENSA-LABEL: ctrl_args_only:
+; XTENSA:         .cfi_startproc
+; XTENSA-NEXT:  # %bb.0:
+; XTENSA-NEXT:    entry a1, 96
+; XTENSA-NEXT:    .cfi_def_cfa_offset 96
+; XTENSA-NEXT:    addi a7, a1, 0
+; XTENSA-NEXT:    l32r a8, .LCPI2_0
+; XTENSA-NEXT:    or a10, a7, a7
+; XTENSA-NEXT:    callx8 a8
+; XTENSA-NEXT:    l32i a8, a1, 96
+; XTENSA-NEXT:    movi a9, 0
+; XTENSA-NEXT:    s8i a9, a8, 0
+; XTENSA-NEXT:    l32i a8, a1, 100
+; XTENSA-NEXT:    s8i a9, a8, 0
+; XTENSA-NEXT:    l32i a8, a1, 104
+; XTENSA-NEXT:    s8i a9, a8, 0
+; XTENSA-NEXT:    l32r a8, .LCPI2_1
+; XTENSA-NEXT:    or a10, a7, a7
+; XTENSA-NEXT:    callx8 a8
+; XTENSA-NEXT:    retw
+                           ptr %g, ptr %h, ptr %i) {
+  %buf = alloca [64 x i8], align 4
+  call void @sink(ptr %buf)
+  store i8 0, ptr %g
+  store i8 0, ptr %h
+  store i8 0, ptr %i
+  call void @sink(ptr %buf)
+  ret i32 %a
+}
+
+; Same conjunction with a frame large enough that the incoming-arg offsets
+; exceed the L32I immediate range: the offset is materialized and added to the
+; reloaded entry SP (not the realigned SP).
+define i32 @victim_big(i32 %a, i32 %b, i32 %c, i32 %d, i32 %e, i32 %f,
+; XTENSA-LABEL: victim_big:
+; XTENSA:         .cfi_startproc
+; XTENSA-NEXT:  # %bb.0:
+; XTENSA-NEXT:    entry a1, 2272
+; XTENSA-NEXT:    movi a8, 63
+; XTENSA-NEXT:    movi a9, 64
+; XTENSA-NEXT:    and a8, a1, a8
+; XTENSA-NEXT:    sub a8, a9, a8
+; XTENSA-NEXT:    add a1, a1, a8
+; XTENSA-NEXT:    .cfi_def_cfa_offset 2272
+; XTENSA-NEXT:    sub a8, a1, a8
+; XTENSA-NEXT:    s32i a8, a1, 60
+; XTENSA-NEXT:    movi a8, 64
+; XTENSA-NEXT:    addmi a8, a8, 2048
+; XTENSA-NEXT:    add a8, a1, a8
+; XTENSA-NEXT:    addi a7, a8, 0
+; XTENSA-NEXT:    l32r a8, .LCPI3_0
+; XTENSA-NEXT:    or a10, a7, a7
+; XTENSA-NEXT:    callx8 a8
+; XTENSA-NEXT:    addi a10, a1, 64
+; XTENSA-NEXT:    l32r a8, .LCPI3_1
+; XTENSA-NEXT:    callx8 a8
+; XTENSA-NEXT:    l32i a9, a1, 60
+; XTENSA-NEXT:    movi a8, 224
+; XTENSA-NEXT:    addmi a8, a8, 2048
+; XTENSA-NEXT:    add a8, a9, a8
+; XTENSA-NEXT:    l32i a8, a8, 0
+; XTENSA-NEXT:    movi a9, 0
+; XTENSA-NEXT:    s8i a9, a8, 0
+; XTENSA-NEXT:    l32i a10, a1, 60
+; XTENSA-NEXT:    movi a8, 228
+; XTENSA-NEXT:    addmi a8, a8, 2048
+; XTENSA-NEXT:    add a8, a10, a8
+; XTENSA-NEXT:    l32i a8, a8, 0
+; XTENSA-NEXT:    s8i a9, a8, 0
+; XTENSA-NEXT:    l32i a10, a1, 60
+; XTENSA-NEXT:    movi a8, 232
+; XTENSA-NEXT:    addmi a8, a8, 2048
+; XTENSA-NEXT:    add a8, a10, a8
+; XTENSA-NEXT:    l32i a8, a8, 0
+; XTENSA-NEXT:    s8i a9, a8, 0
+; XTENSA-NEXT:    l32r a8, .LCPI3_2
+; XTENSA-NEXT:    or a10, a7, a7
+; XTENSA-NEXT:    callx8 a8
+; XTENSA-NEXT:    retw
+                       ptr %g, ptr %h, ptr %i) {
+  %buf = alloca [64 x i8], align 64
+  %big = alloca [2048 x i8], align 4
+  call void @sink(ptr %buf)
+  call void @sink(ptr %big)
+  store i8 0, ptr %g
+  store i8 0, ptr %h
+  store i8 0, ptr %i
+  call void @sink(ptr %buf)
+  ret i32 %a
+}
+
+declare void @llvm.va_start(ptr)
+
+; Realigned frame + stack args + dynamic alloca (hasFP): the pad must be
+; computed from SP (not the still-unset frame pointer), the entry-SP slot is
+; stored after FP setup, and reloads go through FP, which is immune to the
+; MOVSP the alloca performs.
+define i32 @realign_alloca(i32 %a, i32 %b, i32 %c, i32 %d, i32 %e, i32 %f,
+; XTENSA-LABEL: realign_alloca:
+; XTENSA:         .cfi_startproc
+; XTENSA-NEXT:  # %bb.0:
+; XTENSA-NEXT:    entry a1, 224
+; XTENSA-NEXT:    movi a9, 63
+; XTENSA-NEXT:    movi a8, 64
+; XTENSA-NEXT:    and a9, a1, a9
+; XTENSA-NEXT:    sub a9, a8, a9
+; XTENSA-NEXT:    add a1, a1, a9
+; XTENSA-NEXT:    or a8, a7, a7
+; XTENSA-NEXT:    or a7, a1, a1
+; XTENSA-NEXT:    .cfi_def_cfa a7, 224
+; XTENSA-NEXT:    sub a9, a1, a9
+; XTENSA-NEXT:    s32i a9, a7, 60
+; XTENSA-NEXT:    l32i a8, a7, 60
+; XTENSA-NEXT:    l32i a8, a8, 232
+; XTENSA-NEXT:    addi a8, a8, 3
+; XTENSA-NEXT:    movi a9, -4
+; XTENSA-NEXT:    and a8, a8, a9
+; XTENSA-NEXT:    addi a8, a8, 31
+; XTENSA-NEXT:    movi a9, -32
+; XTENSA-NEXT:    and a8, a8, a9
+; XTENSA-NEXT:    sub a8, a1, a8
+; XTENSA-NEXT:    movsp a1, a8
+; XTENSA-NEXT:    or a6, a1, a1
+; XTENSA-NEXT:    addi a10, a7, 64
+; XTENSA-NEXT:    l32r a8, .LCPI4_0
+; XTENSA-NEXT:    callx8 a8
+; XTENSA-NEXT:    l32r a8, .LCPI4_1
+; XTENSA-NEXT:    or a10, a6, a6
+; XTENSA-NEXT:    callx8 a8
+; XTENSA-NEXT:    l32i a8, a7, 60
+; XTENSA-NEXT:    l32i a8, a8, 224
+; XTENSA-NEXT:    movi a9, 0
+; XTENSA-NEXT:    s8i a9, a8, 0
+; XTENSA-NEXT:    l32i a8, a7, 60
+; XTENSA-NEXT:    l32i a8, a8, 228
+; XTENSA-NEXT:    s8i a9, a8, 0
+; XTENSA-NEXT:    retw
+                           ptr %g, ptr %h, i32 %n) {
+  %buf = alloca [64 x i8], align 64
+  %dyn = alloca i8, i32 %n
+  call void @sink(ptr %buf)
+  call void @sink(ptr %dyn)
+  store i8 0, ptr %g
+  store i8 0, ptr %h
+  ret i32 %a
+}
+
+; Realigned frame + varargs: the vararg register-save area and the overflow
+; area pointer are fixed objects; both must be addressed via the entry SP so
+; that va_arg reads what the callee spilled and the caller passed.
+define i32 @realign_varargs(ptr %fmt, ...) {
+; XTENSA-LABEL: realign_varargs:
+; XTENSA:         .cfi_startproc
+; XTENSA-NEXT:  # %bb.0:
+; XTENSA-NEXT:    entry a1, 288
+; XTENSA-NEXT:    movi a8, 63
+; XTENSA-NEXT:    movi a9, 64
+; XTENSA-NEXT:    and a8, a1, a8
+; XTENSA-NEXT:    sub a8, a9, a8
+; XTENSA-NEXT:    add a1, a1, a8
+; XTENSA-NEXT:    .cfi_def_cfa_offset 288
+; XTENSA-NEXT:    sub a8, a1, a8
+; XTENSA-NEXT:    s32i a8, a1, 48
+; XTENSA-NEXT:    l32i a8, a1, 48
+; XTENSA-NEXT:    s32i a7, a8, 188
+; XTENSA-NEXT:    l32i a8, a1, 48
+; XTENSA-NEXT:    s32i a6, a8, 184
+; XTENSA-NEXT:    l32i a8, a1, 48
+; XTENSA-NEXT:    s32i a5, a8, 180
+; XTENSA-NEXT:    l32i a8, a1, 48
+; XTENSA-NEXT:    s32i a4, a8, 176
+; XTENSA-NEXT:    l32i a8, a1, 48
+; XTENSA-NEXT:    s32i a3, a8, 172
+; XTENSA-NEXT:    movi a8, 4
+; XTENSA-NEXT:    s32i a8, a1, 60
+; XTENSA-NEXT:    l32i a9, a1, 48
+; XTENSA-NEXT:    movi a8, 172
+; XTENSA-NEXT:    add a8, a9, a8
+; XTENSA-NEXT:    addi a8, a8, 0
+; XTENSA-NEXT:    s32i a8, a1, 56
+; XTENSA-NEXT:    l32i a9, a1, 48
+; XTENSA-NEXT:    movi a8, 288
+; XTENSA-NEXT:    add a8, a9, a8
+; XTENSA-NEXT:    addi a8, a8, 0
+; XTENSA-NEXT:    addi a8, a8, -32
+; XTENSA-NEXT:    s32i a8, a1, 52
+; XTENSA-NEXT:    addi a10, a1, 64
+; XTENSA-NEXT:    l32r a8, .LCPI5_0
+; XTENSA-NEXT:    callx8 a8
+; XTENSA-NEXT:    addi a10, a1, 52
+; XTENSA-NEXT:    l32r a8, .LCPI5_1
+; XTENSA-NEXT:    callx8 a8
+; XTENSA-NEXT:    movi a8, 0
+; XTENSA-NEXT:    s8i a8, a2, 0
+; XTENSA-NEXT:    or a2, a8, a8
+; XTENSA-NEXT:    retw
+  %buf = alloca [64 x i8], align 64
+  %ap = alloca [12 x i8], align 4
+  call void @llvm.va_start(ptr %ap)
+  call void @sink(ptr %buf)
+  call void @sink(ptr %ap)
+  store i8 0, ptr %fmt
+  ret i32 0
+}

``````````

</details>


https://github.com/llvm/llvm-project/pull/208947


More information about the llvm-commits mailing list