[llvm] [X86][CodeGen] Preserve XMM liveness across VZEROUPPER and check reg clobbers in RemoveLoadsIntoFakeUses (#214068) (PR #218873)
Patrick Ribbsaeter via llvm-commits
llvm-commits at lists.llvm.org
Wed Aug 26 03:01:55 PDT 2026
https://github.com/patrickswedish created https://github.com/llvm/llvm-project/pull/218873
Fixes #214068.
### Problem
Under `llc -O2 -mattr=+avx2`, a function returning a floating-point scalar (`double`) via `%xmm0` had its final reload of `%xmm0` incorrectly deleted during `RemoveLoadsIntoFakeUses` when followed by `VZEROUPPER` and `RET`.
### Root Cause Analysis & Proof
1. **`VZEROUPPER` Subregister Def vs Liveness Invariant**:
In `X86InsertVZeroUpper`, `VZEROUPPER` instructions are inserted before calls/returns with `let Defs = [YMM0..YMM15]`.
When `LivePhysRegs.stepBackward` processes `VZEROUPPER`, `LiveRegUnits` treats `implicit-def $ymm0` as clearing liveness for all subregisters of `$ymm0` (including `$xmm0`). Because `VZEROUPPER` lacked implicit uses of the live return value `%xmm0`, `LivePhysRegs.available($xmm0)` evaluated to `true` before `VZEROUPPER`.
When `RemoveLoadsIntoFakeUses` encountered `renamable $xmm0 = VMOVSDrm_alt ...` (the return value reload), it checked `LivePhysRegs.available($xmm0)`. Because `%xmm0` appeared "available" and a prior `FAKE_USE $xmm0` existed in the block, the pass erroneously deleted the real return value reload.
*Fix*: In `X86InsertVZeroUpper`, when inserting `VZEROUPPER` before an instruction `I` (return or call), any physical XMM register read by `I` is attached as `implicit $xmmN` to `VZEROUPPER`. This accurately models hardware behavior (bits 255:128 are zeroed, bits 127:0 are preserved) and keeps `LivePhysRegs` exact across `VZEROUPPER`.
2. **Register Mask Clobbers & Sub/Super Register Aliases in `RemoveLoadsIntoFakeUses`**:
During the backward walk in `RemoveLoadsIntoFakeUses.cpp`, `RegFakeUses` was previously invalidated only by iterating `MO.isReg()`. It did not check `MO.isRegMask()`, meaning call clobbers did not invalidate tracked fake uses.
*Fix*: In `RemoveLoadsIntoFakeUses.cpp`, use `MI.modifiesRegister(MO.getReg(), TRI)` over tracked `RegFakeUses`. This canonical LLVM API cleanly handles explicit defs, sub/super-register aliases, and call `RegMask` clobbers. Additionally, non-deleted reloads fall through so their register liveness updates correctly in `LivePhysRegs`.
### Verification
- Added MIR regression tests in `llvm/test/CodeGen/X86/fake-use-remove-loads.mir`:
- `clobber_regmask`: Verifies reloads before calls clobbering registers are preserved across calls with subsequent fake uses.
- `vzeroupper_return`: Verifies return reloads of `$xmm0` are preserved across `VZEROUPPER` with implicit XMM use.
- Added end-to-end LLVM IR regression test in `llvm/test/CodeGen/X86/pr214068-vzeroupper-fakeuse-reload.ll`.
- Validated end-to-end binary execution of the original reproducer (compiled with `-O2 -mattr=+avx2`), confirming the return value matches expected computation (`0.3125`).
- Validated with `llvm-lit` across the fake-use and X86 CodeGen test suite.
>From cea226e256d61fe3bf9458e7cd718adeead732ee Mon Sep 17 00:00:00 2001
From: Patrick Ribbsaeter <patrick.ribbsaeter at gmail.com>
Date: Wed, 26 Aug 2026 12:00:26 +0200
Subject: [PATCH] [X86][CodeGen] Preserve XMM liveness across VZEROUPPER and
check reg clobbers in RemoveLoadsIntoFakeUses (#214068)
Fix miscompilation where a floating-point return value reload into %xmm0
was incorrectly removed by RemoveLoadsIntoFakeUses when followed by
VZEROUPPER and RET.
1. In X86InsertVZeroUpper, when inserting VZEROUPPER before a return or call,
preserve liveness of active physical XMM registers read by the terminator
by marking them as implicit uses on VZEROUPPER. This accurately reflects
that VZEROUPPER zeroes bits 255:128 (YMM upper halves) while leaving bits
127:0 (XMM registers) intact.
2. In RemoveLoadsIntoFakeUses, check MI.modifiesRegister(MO.getReg(), TRI)
when clearing tracked fake uses during the backward walk. This canonically
handles explicit defs, sub/super register aliases, and call RegMask clobbers.
Additionally, ensure non-deleted reloads fall through to update register
liveness in LivePhysRegs.
Fixes #214068.
---
llvm/lib/CodeGen/RemoveLoadsIntoFakeUses.cpp | 83 +++++-----
llvm/lib/Target/X86/X86InsertVZeroUpper.cpp | 30 +++-
.../CodeGen/X86/fake-use-remove-loads.mir | 61 ++++++++
.../X86/pr214068-vzeroupper-fakeuse-reload.ll | 145 ++++++++++++++++++
4 files changed, 271 insertions(+), 48 deletions(-)
create mode 100644 llvm/test/CodeGen/X86/pr214068-vzeroupper-fakeuse-reload.ll
diff --git a/llvm/lib/CodeGen/RemoveLoadsIntoFakeUses.cpp b/llvm/lib/CodeGen/RemoveLoadsIntoFakeUses.cpp
index 8888bf792b279..5472ae14f2bc3 100644
--- a/llvm/lib/CodeGen/RemoveLoadsIntoFakeUses.cpp
+++ b/llvm/lib/CodeGen/RemoveLoadsIntoFakeUses.cpp
@@ -137,54 +137,55 @@ bool RemoveLoadsIntoFakeUses::run(MachineFunction &MF) {
if (MI.getRestoreSize(TII)) {
Register Reg = MI.getOperand(0).getReg();
// Don't delete live physreg defs, or any reserved register defs.
- if (!LivePhysRegs.available(Reg) || MRI->isReserved(Reg))
- continue;
- // There should typically be an exact match between the loaded register
- // and the FAKE_USE, but sometimes regalloc will choose to load a larger
- // value than is needed. Therefore, as long as the load isn't used by
- // anything except at least one FAKE_USE, we will delete it. If it isn't
- // used by any fake uses, it should still be safe to delete but we
- // choose to ignore it so that this pass has no side effects unrelated
- // to fake uses.
- SmallDenseSet<MachineInstr *> FakeUsesToDelete;
- for (MachineInstr *&FakeUse : reverse(RegFakeUses)) {
- if (FakeUse->readsRegister(Reg, TRI)) {
- FakeUsesToDelete.insert(FakeUse);
- RegFakeUses.erase(&FakeUse);
+ if (LivePhysRegs.available(Reg) && !MRI->isReserved(Reg)) {
+ // There should typically be an exact match between the loaded
+ // register and the FAKE_USE, but sometimes regalloc will choose to
+ // load a larger value than is needed. Therefore, as long as the load
+ // isn't used by anything except at least one FAKE_USE, we will
+ // delete it. If it isn't used by any fake uses, it should still be
+ // safe to delete but we choose to ignore it so that this pass has no
+ // side effects unrelated to fake uses.
+ SmallDenseSet<MachineInstr *> FakeUsesToDelete;
+ for (MachineInstr *&FakeUse : reverse(RegFakeUses)) {
+ if (FakeUse->readsRegister(Reg, TRI)) {
+ FakeUsesToDelete.insert(FakeUse);
+ RegFakeUses.erase(&FakeUse);
+ }
}
- }
- if (!FakeUsesToDelete.empty()) {
- LLVM_DEBUG(dbgs() << "RemoveLoadsIntoFakeUses: DELETING: " << MI);
- // Since this load only exists to restore a spilled register and we
- // haven't, run LiveDebugValues yet, there shouldn't be any DBG_VALUEs
- // for this load; otherwise, deleting this would be incorrect.
- MI.eraseFromParent();
- AnyChanges = true;
- ++NumLoadsDeleted;
- for (MachineInstr *FakeUse : FakeUsesToDelete) {
- LLVM_DEBUG(dbgs()
- << "RemoveLoadsIntoFakeUses: DELETING: " << *FakeUse);
- FakeUse->eraseFromParent();
+ if (!FakeUsesToDelete.empty()) {
+ LLVM_DEBUG(dbgs() << "RemoveLoadsIntoFakeUses: DELETING: " << MI);
+ // Since this load only exists to restore a spilled register and
+ // we haven't run LiveDebugValues yet, there shouldn't be any
+ // DBG_VALUEs for this load; otherwise, deleting this would be
+ // incorrect.
+ MI.eraseFromParent();
+ AnyChanges = true;
+ ++NumLoadsDeleted;
+ for (MachineInstr *FakeUse : FakeUsesToDelete) {
+ LLVM_DEBUG(dbgs()
+ << "RemoveLoadsIntoFakeUses: DELETING: " << *FakeUse);
+ FakeUse->eraseFromParent();
+ }
+ NumFakeUsesDeleted += FakeUsesToDelete.size();
+ continue;
}
- NumFakeUsesDeleted += FakeUsesToDelete.size();
}
- continue;
}
// In addition to tracking LivePhysRegs, we need to clear RegFakeUses each
- // time a register is defined, as existing FAKE_USEs no longer apply to
- // that register.
+ // time a register is defined or clobbered, as existing FAKE_USEs no
+ // longer apply to that register.
if (!RegFakeUses.empty()) {
- for (const MachineOperand &MO : MI.operands()) {
- if (!MO.isReg())
- continue;
- Register Reg = MO.getReg();
- // We clear RegFakeUses for this register and all subregisters,
- // because any such FAKE_USE encountered prior is no longer relevant
- // for later encountered loads.
- for (MachineInstr *&FakeUse : reverse(RegFakeUses))
- if (FakeUse->readsRegister(Reg, TRI))
- RegFakeUses.erase(&FakeUse);
+ for (MachineInstr *&FakeUse : reverse(RegFakeUses)) {
+ bool Modified = false;
+ for (const MachineOperand &MO : FakeUse->operands()) {
+ if (MO.isReg() && MI.modifiesRegister(MO.getReg(), TRI)) {
+ Modified = true;
+ break;
+ }
+ }
+ if (Modified)
+ RegFakeUses.erase(&FakeUse);
}
}
if (!MI.isDebugInstr())
diff --git a/llvm/lib/Target/X86/X86InsertVZeroUpper.cpp b/llvm/lib/Target/X86/X86InsertVZeroUpper.cpp
index 4cbb911a5880f..87496f8675f80 100644
--- a/llvm/lib/Target/X86/X86InsertVZeroUpper.cpp
+++ b/llvm/lib/Target/X86/X86InsertVZeroUpper.cpp
@@ -171,8 +171,21 @@ static bool callHasRegMask(MachineInstr &MI) {
/// Insert a vzeroupper instruction before I.
static bool insertVZeroUpper(MachineBasicBlock::iterator I,
MachineBasicBlock &MBB,
- const TargetInstrInfo *TII) {
- BuildMI(MBB, I, I->getDebugLoc(), TII->get(X86::VZEROUPPER));
+ const TargetInstrInfo *TII,
+ const TargetRegisterInfo *TRI) {
+ MachineInstrBuilder MIB =
+ BuildMI(MBB, I, I->getDebugLoc(), TII->get(X86::VZEROUPPER));
+ if (I != MBB.end()) {
+ for (const MachineOperand &MO : I->operands()) {
+ if (MO.isReg() && MO.readsReg() && MO.getReg().isPhysical()) {
+ Register Reg = MO.getReg();
+ for (MCPhysReg XMM : X86::FR32RegClass) {
+ if (TRI->regsOverlap(Reg, XMM))
+ MIB.addReg(XMM, RegState::Implicit);
+ }
+ }
+ }
+ }
++NumVZU;
return true;
}
@@ -192,7 +205,8 @@ static void addDirtySuccessor(MachineBasicBlock &MBB,
static bool processBasicBlock(MachineBasicBlock &MBB,
BlockStateMap &BlockStates,
DirtySuccessorsWorkList &DirtySuccessors,
- bool IsX86INTR, const TargetInstrInfo *TII) {
+ bool IsX86INTR, const TargetInstrInfo *TII,
+ const TargetRegisterInfo *TRI) {
// Start by assuming that the block is PASS_THROUGH which implies no unguarded
// calls.
BlockExitState CurState = PASS_THROUGH;
@@ -250,7 +264,7 @@ static bool processBasicBlock(MachineBasicBlock &MBB,
// After the inserted VZEROUPPER the state becomes clean again, but
// other YMM/ZMM may appear before other subsequent calls or even before
// the end of the BB.
- MadeChange |= insertVZeroUpper(MI, MBB, TII);
+ MadeChange |= insertVZeroUpper(MI, MBB, TII, TRI);
CurState = EXITS_CLEAN;
} else if (CurState == PASS_THROUGH) {
// If this block is currently in pass-through state and we encounter a
@@ -306,6 +320,7 @@ static bool insertVZeroUpper(MachineFunction &MF) {
return false;
const TargetInstrInfo *TII = ST.getInstrInfo();
+ const TargetRegisterInfo *TRI = ST.getRegisterInfo();
bool IsX86INTR = MF.getFunction().getCallingConv() == CallingConv::X86_INTR;
bool EverMadeChange = false;
BlockStateMap BlockStates(MF.getNumBlockIDs());
@@ -318,8 +333,8 @@ static bool insertVZeroUpper(MachineFunction &MF) {
// unguarded call in each block, and add successors of dirty blocks to the
// DirtySuccessors list.
for (MachineBasicBlock &MBB : MF)
- EverMadeChange |=
- processBasicBlock(MBB, BlockStates, DirtySuccessors, IsX86INTR, TII);
+ EverMadeChange |= processBasicBlock(MBB, BlockStates, DirtySuccessors,
+ IsX86INTR, TII, TRI);
// If any YMM/ZMM regs are live-in to this function, add the entry block to
// the DirtySuccessors list
@@ -337,7 +352,8 @@ static bool insertVZeroUpper(MachineFunction &MF) {
// MBB is a successor of a dirty block, so its first call needs to be
// guarded.
if (BBState.FirstUnguardedCall != MBB.end())
- EverMadeChange |= insertVZeroUpper(BBState.FirstUnguardedCall, MBB, TII);
+ EverMadeChange |=
+ insertVZeroUpper(BBState.FirstUnguardedCall, MBB, TII, TRI);
// If this successor was a pass-through block, then it is now dirty. Its
// successors need to be added to the worklist (if they haven't been
diff --git a/llvm/test/CodeGen/X86/fake-use-remove-loads.mir b/llvm/test/CodeGen/X86/fake-use-remove-loads.mir
index aa9839d2700af..fea07f674b9ae 100644
--- a/llvm/test/CodeGen/X86/fake-use-remove-loads.mir
+++ b/llvm/test/CodeGen/X86/fake-use-remove-loads.mir
@@ -171,3 +171,64 @@ body: |
RET64
...
+---
+name: clobber_regmask
+tracksRegLiveness: true
+noPhis: true
+noVRegs: true
+hasFakeUses: true
+tracksDebugUserValues: true
+debugInstrRef: true
+stack:
+ - { id: 0, name: '', type: spill-slot, offset: -8, size: 8, alignment: 8,
+ stack-id: default, callee-saved-register: '', callee-saved-restored: true }
+ - { id: 1, name: '', type: spill-slot, offset: -16, size: 8, alignment: 8,
+ stack-id: default, callee-saved-register: '', callee-saved-restored: true }
+body: |
+ bb.0:
+ liveins: $rdi, $r12
+
+ ; CHECK-LABEL: name: clobber_regmask
+ ; CHECK: liveins: $rdi, $r12
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: renamable $rax = MOV64rm $rsp, 1, $noreg, -8, $noreg :: (load (s64) from %stack.0)
+ ; CHECK-NEXT: $rdi = MOV64rr $rax
+ ; CHECK-NEXT: CALL64r renamable $r12, csr_64, implicit $rsp, implicit $ssp, implicit $rdi, implicit-def $rsp, implicit-def $ssp
+ ; CHECK-NEXT: RET64
+
+ ;; Verify that a reload before a call whose regmask clobbers the register is
+ ;; not removed due to a fake use after the call.
+ renamable $rax = MOV64rm $rsp, 1, $noreg, -8, $noreg :: (load (s64) from %stack.0)
+ $rdi = MOV64rr $rax
+ CALL64r renamable $r12, csr_64, implicit $rsp, implicit $ssp, implicit $rdi, implicit-def $rsp, implicit-def $ssp
+ renamable $rax = MOV64rm $rsp, 1, $noreg, -16, $noreg :: (load (s64) from %stack.1)
+ FAKE_USE killed renamable $rax
+ RET64
+
+...
+---
+name: vzeroupper_return
+tracksRegLiveness: true
+noPhis: true
+noVRegs: true
+hasFakeUses: true
+tracksDebugUserValues: true
+debugInstrRef: true
+stack:
+ - { id: 0, name: '', type: spill-slot, offset: -8, size: 8, alignment: 8,
+ stack-id: default, callee-saved-register: '', callee-saved-restored: true }
+body: |
+ bb.0:
+ ; CHECK-LABEL: name: vzeroupper_return
+ ; CHECK: renamable $xmm0 = VMOVSDrm_alt $rsp, 1, $noreg, -8, $noreg :: (load (s64) from %stack.0)
+ ; CHECK-NEXT: FAKE_USE renamable $xmm0
+ ; CHECK-NEXT: VZEROUPPER implicit-def $ymm0, implicit-def $ymm1, implicit-def $ymm2, implicit-def $ymm3, implicit-def $ymm4, implicit-def $ymm5, implicit-def $ymm6, implicit-def $ymm7, implicit-def $ymm8, implicit-def $ymm9, implicit-def $ymm10, implicit-def $ymm11, implicit-def $ymm12, implicit-def $ymm13, implicit-def $ymm14, implicit-def $ymm15, implicit $xmm0
+ ; CHECK-NEXT: RET64 $xmm0
+
+ ;; Verify that a reload of a return value is preserved across VZEROUPPER with implicit XMM use.
+ renamable $xmm0 = VMOVSDrm_alt $rsp, 1, $noreg, -8, $noreg :: (load (s64) from %stack.0)
+ FAKE_USE renamable $xmm0
+ VZEROUPPER implicit-def $ymm0, implicit-def $ymm1, implicit-def $ymm2, implicit-def $ymm3, implicit-def $ymm4, implicit-def $ymm5, implicit-def $ymm6, implicit-def $ymm7, implicit-def $ymm8, implicit-def $ymm9, implicit-def $ymm10, implicit-def $ymm11, implicit-def $ymm12, implicit-def $ymm13, implicit-def $ymm14, implicit-def $ymm15, implicit $xmm0
+ RET64 $xmm0
+
+...
diff --git a/llvm/test/CodeGen/X86/pr214068-vzeroupper-fakeuse-reload.ll b/llvm/test/CodeGen/X86/pr214068-vzeroupper-fakeuse-reload.ll
new file mode 100644
index 0000000000000..d448e97cb4582
--- /dev/null
+++ b/llvm/test/CodeGen/X86/pr214068-vzeroupper-fakeuse-reload.ll
@@ -0,0 +1,145 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py
+; RUN: llc < %s -mtriple=x86_64-unknown-linux-gnu -mattr=+avx2 -O2 | FileCheck %s
+
+; PR214068: Verify that Remove Loads Into Fake Uses does not delete the return
+; value reload into %xmm0 when vzeroupper is inserted before ret.
+
+; ModuleID = 'reduced.ll'
+source_filename = "reduced.cpp"
+target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i64:64-i128:128-f80:128-n8:16:32:64-S128"
+target triple = "x86_64-redhat-linux-gnu"
+
+define double @_ZN12_GLOBAL__N_111Accumulator25calculateErrorAndJacobianENS_6vectorEPNS_4LossE(ptr noundef nonnull align 8 dereferenceable(32) %0, i32 %1, ptr %2, ptr noundef nonnull %3) #2 align 2 {
+; CHECK-LABEL: _ZN12_GLOBAL__N_111Accumulator25calculateErrorAndJacobianENS_6vectorEPNS_4LossE:
+; CHECK: # %bb.0:
+; CHECK-NEXT: pushq %rbp
+; CHECK-NEXT: .cfi_def_cfa_offset 16
+; CHECK-NEXT: pushq %r15
+; CHECK-NEXT: .cfi_def_cfa_offset 24
+; CHECK-NEXT: pushq %r14
+; CHECK-NEXT: .cfi_def_cfa_offset 32
+; CHECK-NEXT: pushq %r13
+; CHECK-NEXT: .cfi_def_cfa_offset 40
+; CHECK-NEXT: pushq %r12
+; CHECK-NEXT: .cfi_def_cfa_offset 48
+; CHECK-NEXT: pushq %rbx
+; CHECK-NEXT: .cfi_def_cfa_offset 56
+; CHECK-NEXT: subq $40, %rsp
+; CHECK-NEXT: .cfi_def_cfa_offset 96
+; CHECK-NEXT: .cfi_offset %rbx, -56
+; CHECK-NEXT: .cfi_offset %r12, -48
+; CHECK-NEXT: .cfi_offset %r13, -40
+; CHECK-NEXT: .cfi_offset %r14, -32
+; CHECK-NEXT: .cfi_offset %r15, -24
+; CHECK-NEXT: .cfi_offset %rbp, -16
+; CHECK-NEXT: movq %rcx, %r14
+; CHECK-NEXT: movq %rdx, %rbx
+; CHECK-NEXT: movl %esi, %ebp
+; CHECK-NEXT: vxorpd %xmm0, %xmm0, %xmm0
+; CHECK-NEXT: movq %rdi, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; CHECK-NEXT: vmovupd %ymm0, (%rdi)
+; CHECK-NEXT: testl %esi, %esi
+; CHECK-NEXT: je .LBB0_1
+; CHECK-NEXT: # %bb.3: # %.preheader
+; CHECK-NEXT: movslq %ebp, %r13
+; CHECK-NEXT: shlq $3, %r13
+; CHECK-NEXT: vxorpd %xmm0, %xmm0, %xmm0
+; CHECK-NEXT: vmovsd %xmm0, (%rsp) # 8-byte Spill
+; CHECK-NEXT: xorl %r15d, %r15d
+; CHECK-NEXT: .p2align 4
+; CHECK-NEXT: .LBB0_4: # =>This Inner Loop Header: Depth=1
+; CHECK-NEXT: movq (%rbx,%r15), %r12
+; CHECK-NEXT: movq (%r12), %rax
+; CHECK-NEXT: movq %r12, %rdi
+; CHECK-NEXT: vzeroupper
+; CHECK-NEXT: callq *(%rax)
+; CHECK-NEXT: vmovsd %xmm0, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; CHECK-NEXT: vmovsd %xmm1, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; CHECK-NEXT: vmulsd %xmm1, %xmm1, %xmm2
+; CHECK-NEXT: vmulsd %xmm0, %xmm0, %xmm1
+; CHECK-NEXT: vaddsd %xmm2, %xmm1, %xmm1
+; CHECK-NEXT: vmovsd %xmm1, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; CHECK-NEXT: vmovsd (%rsp), %xmm0 # 8-byte Reload
+; CHECK-NEXT: # xmm0 = mem[0],zero
+; CHECK-NEXT: vaddsd %xmm1, %xmm0, %xmm0
+; CHECK-NEXT: vmovsd %xmm0, (%rsp) # 8-byte Spill
+; CHECK-NEXT: movq (%r14), %rax
+; CHECK-NEXT: movq %r14, %rdi
+; CHECK-NEXT: vmovapd %xmm1, %xmm0
+; CHECK-NEXT: callq *(%rax)
+; CHECK-NEXT: # fake_use: $r12
+; CHECK-NEXT: addq $8, %r15
+; CHECK-NEXT: cmpq %r15, %r13
+; CHECK-NEXT: jne .LBB0_4
+; CHECK-NEXT: jmp .LBB0_2
+; CHECK-NEXT: .LBB0_1:
+; CHECK-NEXT: vxorpd %xmm0, %xmm0, %xmm0
+; CHECK-NEXT: vmovsd %xmm0, (%rsp) # 8-byte Spill
+; CHECK-NEXT: .LBB0_2:
+; CHECK-NEXT: vmovsd (%rsp), %xmm0 # 8-byte Reload
+; CHECK-NEXT: # xmm0 = mem[0],zero
+; CHECK-NEXT: # fake_use: $xmm0
+; CHECK-NEXT: # fake_use: $r14
+; CHECK-NEXT: # fake_use: $ebp
+; CHECK-NEXT: # fake_use: $rbx
+; CHECK-NEXT: addq $40, %rsp
+; CHECK-NEXT: .cfi_def_cfa_offset 56
+; CHECK-NEXT: popq %rbx
+; CHECK-NEXT: .cfi_def_cfa_offset 48
+; CHECK-NEXT: popq %r12
+; CHECK-NEXT: .cfi_def_cfa_offset 40
+; CHECK-NEXT: popq %r13
+; CHECK-NEXT: .cfi_def_cfa_offset 32
+; CHECK-NEXT: popq %r14
+; CHECK-NEXT: .cfi_def_cfa_offset 24
+; CHECK-NEXT: popq %r15
+; CHECK-NEXT: .cfi_def_cfa_offset 16
+; CHECK-NEXT: popq %rbp
+; CHECK-NEXT: .cfi_def_cfa_offset 8
+; CHECK-NEXT: vzeroupper
+; CHECK-NEXT: retq
+ tail call void @llvm.memset.p0.i64(ptr noundef nonnull align 8 dereferenceable(32) %0, i8 0, i64 32, i1 false)
+ %5 = sext i32 %1 to i64
+ %6 = shl nsw i64 %5, 3
+ %7 = getelementptr inbounds i8, ptr %2, i64 %6
+ %8 = icmp eq i32 %1, 0
+ br i1 %8, label %9, label %11
+
+9: ; preds = %11, %4
+ %10 = phi double [ 0.000000e+00, %4 ], [ %22, %11 ]
+ notail call void (...) @llvm.fake.use(double %10)
+ notail call void (...) @llvm.fake.use(ptr nonnull %3)
+ tail call void (...) @llvm.fake.use(i32 %1)
+ tail call void (...) @llvm.fake.use(ptr %2)
+ notail call void (...) @llvm.fake.use(ptr nonnull %0)
+ ret double %10
+
+11: ; preds = %4, %11
+ %12 = phi double [ %22, %11 ], [ 0.000000e+00, %4 ]
+ %13 = phi ptr [ %26, %11 ], [ %2, %4 ]
+ %14 = load ptr, ptr %13, align 8
+ %15 = load ptr, ptr %14, align 8
+ %16 = load ptr, ptr %15, align 8
+ %17 = tail call { double, double } %16(ptr noundef nonnull align 8 dereferenceable(8) %14)
+ %18 = extractvalue { double, double } %17, 0
+ %19 = extractvalue { double, double } %17, 1
+ %20 = fmul double %19, %19
+ %21 = tail call noundef double @llvm.fmuladd.f64(double %18, double %18, double %20)
+ %22 = fadd double %12, %21
+ %23 = load ptr, ptr %3, align 8
+ %24 = load ptr, ptr %23, align 8
+ %25 = tail call noundef double %24(ptr noundef nonnull align 8 dereferenceable(8) %3, double noundef %21)
+ notail call void (...) @llvm.fake.use(double %21)
+ tail call void (...) @llvm.fake.use(double %18)
+ tail call void (...) @llvm.fake.use(double %19)
+ notail call void (...) @llvm.fake.use(ptr nonnull %14)
+ %26 = getelementptr inbounds nuw i8, ptr %13, i64 8
+ %27 = icmp eq ptr %26, %7
+ br i1 %27, label %9, label %11
+}
+
+declare void @llvm.fake.use(...)
+declare double @llvm.fmuladd.f64(double, double, double)
+declare void @llvm.memset.p0.i64(ptr writeonly captures(none), i8, i64, i1 immarg)
+
+attributes #2 = { mustprogress noinline norecurse uwtable "min-legal-vector-width"="0" "no-trapping-math"="true" "stack-protector-buffer-size"="8" "target-cpu"="x86-64" "target-features"="+avx,+avx2,+cmov,+crc32,+cx8,+fxsr,+mmx,+popcnt,+sse,+sse2,+sse3,+sse4.1,+sse4.2,+ssse3,+x87,+xsave" "tune-cpu"="generic" }
More information about the llvm-commits
mailing list