[llvm] [X86][CodeGen] Preserve XMM liveness across VZEROUPPER and check reg clobbers in RemoveLoadsIntoFakeUses (#214068) (PR #218873)

Patrick Ribbsaeter via llvm-commits llvm-commits at lists.llvm.org
Wed Aug 26 03:01:55 PDT 2026


https://github.com/patrickswedish created https://github.com/llvm/llvm-project/pull/218873

Fixes #214068.

### Problem
Under `llc -O2 -mattr=+avx2`, a function returning a floating-point scalar (`double`) via `%xmm0` had its final reload of `%xmm0` incorrectly deleted during `RemoveLoadsIntoFakeUses` when followed by `VZEROUPPER` and `RET`.

### Root Cause Analysis & Proof
1. **`VZEROUPPER` Subregister Def vs Liveness Invariant**:
   In `X86InsertVZeroUpper`, `VZEROUPPER` instructions are inserted before calls/returns with `let Defs = [YMM0..YMM15]`.
   When `LivePhysRegs.stepBackward` processes `VZEROUPPER`, `LiveRegUnits` treats `implicit-def $ymm0` as clearing liveness for all subregisters of `$ymm0` (including `$xmm0`). Because `VZEROUPPER` lacked implicit uses of the live return value `%xmm0`, `LivePhysRegs.available($xmm0)` evaluated to `true` before `VZEROUPPER`.
   When `RemoveLoadsIntoFakeUses` encountered `renamable $xmm0 = VMOVSDrm_alt ...` (the return value reload), it checked `LivePhysRegs.available($xmm0)`. Because `%xmm0` appeared "available" and a prior `FAKE_USE $xmm0` existed in the block, the pass erroneously deleted the real return value reload.
   *Fix*: In `X86InsertVZeroUpper`, when inserting `VZEROUPPER` before an instruction `I` (return or call), any physical XMM register read by `I` is attached as `implicit $xmmN` to `VZEROUPPER`. This accurately models hardware behavior (bits 255:128 are zeroed, bits 127:0 are preserved) and keeps `LivePhysRegs` exact across `VZEROUPPER`.

2. **Register Mask Clobbers & Sub/Super Register Aliases in `RemoveLoadsIntoFakeUses`**:
   During the backward walk in `RemoveLoadsIntoFakeUses.cpp`, `RegFakeUses` was previously invalidated only by iterating `MO.isReg()`. It did not check `MO.isRegMask()`, meaning call clobbers did not invalidate tracked fake uses.
   *Fix*: In `RemoveLoadsIntoFakeUses.cpp`, use `MI.modifiesRegister(MO.getReg(), TRI)` over tracked `RegFakeUses`. This canonical LLVM API cleanly handles explicit defs, sub/super-register aliases, and call `RegMask` clobbers. Additionally, non-deleted reloads fall through so their register liveness updates correctly in `LivePhysRegs`.

### Verification
- Added MIR regression tests in `llvm/test/CodeGen/X86/fake-use-remove-loads.mir`:
  - `clobber_regmask`: Verifies reloads before calls clobbering registers are preserved across calls with subsequent fake uses.
  - `vzeroupper_return`: Verifies return reloads of `$xmm0` are preserved across `VZEROUPPER` with implicit XMM use.
- Added end-to-end LLVM IR regression test in `llvm/test/CodeGen/X86/pr214068-vzeroupper-fakeuse-reload.ll`.
- Validated end-to-end binary execution of the original reproducer (compiled with `-O2 -mattr=+avx2`), confirming the return value matches expected computation (`0.3125`).
- Validated with `llvm-lit` across the fake-use and X86 CodeGen test suite.


>From cea226e256d61fe3bf9458e7cd718adeead732ee Mon Sep 17 00:00:00 2001
From: Patrick Ribbsaeter <patrick.ribbsaeter at gmail.com>
Date: Wed, 26 Aug 2026 12:00:26 +0200
Subject: [PATCH] [X86][CodeGen] Preserve XMM liveness across VZEROUPPER and
 check reg clobbers in RemoveLoadsIntoFakeUses (#214068)

Fix miscompilation where a floating-point return value reload into %xmm0
was incorrectly removed by RemoveLoadsIntoFakeUses when followed by
VZEROUPPER and RET.

1. In X86InsertVZeroUpper, when inserting VZEROUPPER before a return or call,
   preserve liveness of active physical XMM registers read by the terminator
   by marking them as implicit uses on VZEROUPPER. This accurately reflects
   that VZEROUPPER zeroes bits 255:128 (YMM upper halves) while leaving bits
   127:0 (XMM registers) intact.
2. In RemoveLoadsIntoFakeUses, check MI.modifiesRegister(MO.getReg(), TRI)
   when clearing tracked fake uses during the backward walk. This canonically
   handles explicit defs, sub/super register aliases, and call RegMask clobbers.
   Additionally, ensure non-deleted reloads fall through to update register
   liveness in LivePhysRegs.

Fixes #214068.
---
 llvm/lib/CodeGen/RemoveLoadsIntoFakeUses.cpp  |  83 +++++-----
 llvm/lib/Target/X86/X86InsertVZeroUpper.cpp   |  30 +++-
 .../CodeGen/X86/fake-use-remove-loads.mir     |  61 ++++++++
 .../X86/pr214068-vzeroupper-fakeuse-reload.ll | 145 ++++++++++++++++++
 4 files changed, 271 insertions(+), 48 deletions(-)
 create mode 100644 llvm/test/CodeGen/X86/pr214068-vzeroupper-fakeuse-reload.ll

diff --git a/llvm/lib/CodeGen/RemoveLoadsIntoFakeUses.cpp b/llvm/lib/CodeGen/RemoveLoadsIntoFakeUses.cpp
index 8888bf792b279..5472ae14f2bc3 100644
--- a/llvm/lib/CodeGen/RemoveLoadsIntoFakeUses.cpp
+++ b/llvm/lib/CodeGen/RemoveLoadsIntoFakeUses.cpp
@@ -137,54 +137,55 @@ bool RemoveLoadsIntoFakeUses::run(MachineFunction &MF) {
       if (MI.getRestoreSize(TII)) {
         Register Reg = MI.getOperand(0).getReg();
         // Don't delete live physreg defs, or any reserved register defs.
-        if (!LivePhysRegs.available(Reg) || MRI->isReserved(Reg))
-          continue;
-        // There should typically be an exact match between the loaded register
-        // and the FAKE_USE, but sometimes regalloc will choose to load a larger
-        // value than is needed. Therefore, as long as the load isn't used by
-        // anything except at least one FAKE_USE, we will delete it. If it isn't
-        // used by any fake uses, it should still be safe to delete but we
-        // choose to ignore it so that this pass has no side effects unrelated
-        // to fake uses.
-        SmallDenseSet<MachineInstr *> FakeUsesToDelete;
-        for (MachineInstr *&FakeUse : reverse(RegFakeUses)) {
-          if (FakeUse->readsRegister(Reg, TRI)) {
-            FakeUsesToDelete.insert(FakeUse);
-            RegFakeUses.erase(&FakeUse);
+        if (LivePhysRegs.available(Reg) && !MRI->isReserved(Reg)) {
+          // There should typically be an exact match between the loaded
+          // register and the FAKE_USE, but sometimes regalloc will choose to
+          // load a larger value than is needed. Therefore, as long as the load
+          // isn't used by anything except at least one FAKE_USE, we will
+          // delete it. If it isn't used by any fake uses, it should still be
+          // safe to delete but we choose to ignore it so that this pass has no
+          // side effects unrelated to fake uses.
+          SmallDenseSet<MachineInstr *> FakeUsesToDelete;
+          for (MachineInstr *&FakeUse : reverse(RegFakeUses)) {
+            if (FakeUse->readsRegister(Reg, TRI)) {
+              FakeUsesToDelete.insert(FakeUse);
+              RegFakeUses.erase(&FakeUse);
+            }
           }
-        }
-        if (!FakeUsesToDelete.empty()) {
-          LLVM_DEBUG(dbgs() << "RemoveLoadsIntoFakeUses: DELETING: " << MI);
-          // Since this load only exists to restore a spilled register and we
-          // haven't, run LiveDebugValues yet, there shouldn't be any DBG_VALUEs
-          // for this load; otherwise, deleting this would be incorrect.
-          MI.eraseFromParent();
-          AnyChanges = true;
-          ++NumLoadsDeleted;
-          for (MachineInstr *FakeUse : FakeUsesToDelete) {
-            LLVM_DEBUG(dbgs()
-                       << "RemoveLoadsIntoFakeUses: DELETING: " << *FakeUse);
-            FakeUse->eraseFromParent();
+          if (!FakeUsesToDelete.empty()) {
+            LLVM_DEBUG(dbgs() << "RemoveLoadsIntoFakeUses: DELETING: " << MI);
+            // Since this load only exists to restore a spilled register and
+            // we haven't run LiveDebugValues yet, there shouldn't be any
+            // DBG_VALUEs for this load; otherwise, deleting this would be
+            // incorrect.
+            MI.eraseFromParent();
+            AnyChanges = true;
+            ++NumLoadsDeleted;
+            for (MachineInstr *FakeUse : FakeUsesToDelete) {
+              LLVM_DEBUG(dbgs()
+                         << "RemoveLoadsIntoFakeUses: DELETING: " << *FakeUse);
+              FakeUse->eraseFromParent();
+            }
+            NumFakeUsesDeleted += FakeUsesToDelete.size();
+            continue;
           }
-          NumFakeUsesDeleted += FakeUsesToDelete.size();
         }
-        continue;
       }
 
       // In addition to tracking LivePhysRegs, we need to clear RegFakeUses each
-      // time a register is defined, as existing FAKE_USEs no longer apply to
-      // that register.
+      // time a register is defined or clobbered, as existing FAKE_USEs no
+      // longer apply to that register.
       if (!RegFakeUses.empty()) {
-        for (const MachineOperand &MO : MI.operands()) {
-          if (!MO.isReg())
-            continue;
-          Register Reg = MO.getReg();
-          // We clear RegFakeUses for this register and all subregisters,
-          // because any such FAKE_USE encountered prior is no longer relevant
-          // for later encountered loads.
-          for (MachineInstr *&FakeUse : reverse(RegFakeUses))
-            if (FakeUse->readsRegister(Reg, TRI))
-              RegFakeUses.erase(&FakeUse);
+        for (MachineInstr *&FakeUse : reverse(RegFakeUses)) {
+          bool Modified = false;
+          for (const MachineOperand &MO : FakeUse->operands()) {
+            if (MO.isReg() && MI.modifiesRegister(MO.getReg(), TRI)) {
+              Modified = true;
+              break;
+            }
+          }
+          if (Modified)
+            RegFakeUses.erase(&FakeUse);
         }
       }
       if (!MI.isDebugInstr())
diff --git a/llvm/lib/Target/X86/X86InsertVZeroUpper.cpp b/llvm/lib/Target/X86/X86InsertVZeroUpper.cpp
index 4cbb911a5880f..87496f8675f80 100644
--- a/llvm/lib/Target/X86/X86InsertVZeroUpper.cpp
+++ b/llvm/lib/Target/X86/X86InsertVZeroUpper.cpp
@@ -171,8 +171,21 @@ static bool callHasRegMask(MachineInstr &MI) {
 /// Insert a vzeroupper instruction before I.
 static bool insertVZeroUpper(MachineBasicBlock::iterator I,
                              MachineBasicBlock &MBB,
-                             const TargetInstrInfo *TII) {
-  BuildMI(MBB, I, I->getDebugLoc(), TII->get(X86::VZEROUPPER));
+                             const TargetInstrInfo *TII,
+                             const TargetRegisterInfo *TRI) {
+  MachineInstrBuilder MIB =
+      BuildMI(MBB, I, I->getDebugLoc(), TII->get(X86::VZEROUPPER));
+  if (I != MBB.end()) {
+    for (const MachineOperand &MO : I->operands()) {
+      if (MO.isReg() && MO.readsReg() && MO.getReg().isPhysical()) {
+        Register Reg = MO.getReg();
+        for (MCPhysReg XMM : X86::FR32RegClass) {
+          if (TRI->regsOverlap(Reg, XMM))
+            MIB.addReg(XMM, RegState::Implicit);
+        }
+      }
+    }
+  }
   ++NumVZU;
   return true;
 }
@@ -192,7 +205,8 @@ static void addDirtySuccessor(MachineBasicBlock &MBB,
 static bool processBasicBlock(MachineBasicBlock &MBB,
                               BlockStateMap &BlockStates,
                               DirtySuccessorsWorkList &DirtySuccessors,
-                              bool IsX86INTR, const TargetInstrInfo *TII) {
+                              bool IsX86INTR, const TargetInstrInfo *TII,
+                              const TargetRegisterInfo *TRI) {
   // Start by assuming that the block is PASS_THROUGH which implies no unguarded
   // calls.
   BlockExitState CurState = PASS_THROUGH;
@@ -250,7 +264,7 @@ static bool processBasicBlock(MachineBasicBlock &MBB,
       // After the inserted VZEROUPPER the state becomes clean again, but
       // other YMM/ZMM may appear before other subsequent calls or even before
       // the end of the BB.
-      MadeChange |= insertVZeroUpper(MI, MBB, TII);
+      MadeChange |= insertVZeroUpper(MI, MBB, TII, TRI);
       CurState = EXITS_CLEAN;
     } else if (CurState == PASS_THROUGH) {
       // If this block is currently in pass-through state and we encounter a
@@ -306,6 +320,7 @@ static bool insertVZeroUpper(MachineFunction &MF) {
     return false;
 
   const TargetInstrInfo *TII = ST.getInstrInfo();
+  const TargetRegisterInfo *TRI = ST.getRegisterInfo();
   bool IsX86INTR = MF.getFunction().getCallingConv() == CallingConv::X86_INTR;
   bool EverMadeChange = false;
   BlockStateMap BlockStates(MF.getNumBlockIDs());
@@ -318,8 +333,8 @@ static bool insertVZeroUpper(MachineFunction &MF) {
   // unguarded call in each block, and add successors of dirty blocks to the
   // DirtySuccessors list.
   for (MachineBasicBlock &MBB : MF)
-    EverMadeChange |=
-        processBasicBlock(MBB, BlockStates, DirtySuccessors, IsX86INTR, TII);
+    EverMadeChange |= processBasicBlock(MBB, BlockStates, DirtySuccessors,
+                                          IsX86INTR, TII, TRI);
 
   // If any YMM/ZMM regs are live-in to this function, add the entry block to
   // the DirtySuccessors list
@@ -337,7 +352,8 @@ static bool insertVZeroUpper(MachineFunction &MF) {
     // MBB is a successor of a dirty block, so its first call needs to be
     // guarded.
     if (BBState.FirstUnguardedCall != MBB.end())
-      EverMadeChange |= insertVZeroUpper(BBState.FirstUnguardedCall, MBB, TII);
+      EverMadeChange |=
+          insertVZeroUpper(BBState.FirstUnguardedCall, MBB, TII, TRI);
 
     // If this successor was a pass-through block, then it is now dirty. Its
     // successors need to be added to the worklist (if they haven't been
diff --git a/llvm/test/CodeGen/X86/fake-use-remove-loads.mir b/llvm/test/CodeGen/X86/fake-use-remove-loads.mir
index aa9839d2700af..fea07f674b9ae 100644
--- a/llvm/test/CodeGen/X86/fake-use-remove-loads.mir
+++ b/llvm/test/CodeGen/X86/fake-use-remove-loads.mir
@@ -171,3 +171,64 @@ body:             |
     RET64
 
 ...
+---
+name:            clobber_regmask
+tracksRegLiveness: true
+noPhis:          true
+noVRegs:         true
+hasFakeUses:     true
+tracksDebugUserValues: true
+debugInstrRef: true
+stack:
+  - { id: 0, name: '', type: spill-slot, offset: -8, size: 8, alignment: 8,
+      stack-id: default, callee-saved-register: '', callee-saved-restored: true }
+  - { id: 1, name: '', type: spill-slot, offset: -16, size: 8, alignment: 8,
+      stack-id: default, callee-saved-register: '', callee-saved-restored: true }
+body:             |
+  bb.0:
+    liveins: $rdi, $r12
+
+    ; CHECK-LABEL: name: clobber_regmask
+    ; CHECK: liveins: $rdi, $r12
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: renamable $rax = MOV64rm $rsp, 1, $noreg, -8, $noreg :: (load (s64) from %stack.0)
+    ; CHECK-NEXT: $rdi = MOV64rr $rax
+    ; CHECK-NEXT: CALL64r renamable $r12, csr_64, implicit $rsp, implicit $ssp, implicit $rdi, implicit-def $rsp, implicit-def $ssp
+    ; CHECK-NEXT: RET64
+
+    ;; Verify that a reload before a call whose regmask clobbers the register is
+    ;; not removed due to a fake use after the call.
+    renamable $rax = MOV64rm $rsp, 1, $noreg, -8, $noreg :: (load (s64) from %stack.0)
+    $rdi = MOV64rr $rax
+    CALL64r renamable $r12, csr_64, implicit $rsp, implicit $ssp, implicit $rdi, implicit-def $rsp, implicit-def $ssp
+    renamable $rax = MOV64rm $rsp, 1, $noreg, -16, $noreg :: (load (s64) from %stack.1)
+    FAKE_USE killed renamable $rax
+    RET64
+
+...
+---
+name:            vzeroupper_return
+tracksRegLiveness: true
+noPhis:          true
+noVRegs:         true
+hasFakeUses:     true
+tracksDebugUserValues: true
+debugInstrRef: true
+stack:
+  - { id: 0, name: '', type: spill-slot, offset: -8, size: 8, alignment: 8,
+      stack-id: default, callee-saved-register: '', callee-saved-restored: true }
+body:             |
+  bb.0:
+    ; CHECK-LABEL: name: vzeroupper_return
+    ; CHECK: renamable $xmm0 = VMOVSDrm_alt $rsp, 1, $noreg, -8, $noreg :: (load (s64) from %stack.0)
+    ; CHECK-NEXT: FAKE_USE renamable $xmm0
+    ; CHECK-NEXT: VZEROUPPER implicit-def $ymm0, implicit-def $ymm1, implicit-def $ymm2, implicit-def $ymm3, implicit-def $ymm4, implicit-def $ymm5, implicit-def $ymm6, implicit-def $ymm7, implicit-def $ymm8, implicit-def $ymm9, implicit-def $ymm10, implicit-def $ymm11, implicit-def $ymm12, implicit-def $ymm13, implicit-def $ymm14, implicit-def $ymm15, implicit $xmm0
+    ; CHECK-NEXT: RET64 $xmm0
+
+    ;; Verify that a reload of a return value is preserved across VZEROUPPER with implicit XMM use.
+    renamable $xmm0 = VMOVSDrm_alt $rsp, 1, $noreg, -8, $noreg :: (load (s64) from %stack.0)
+    FAKE_USE renamable $xmm0
+    VZEROUPPER implicit-def $ymm0, implicit-def $ymm1, implicit-def $ymm2, implicit-def $ymm3, implicit-def $ymm4, implicit-def $ymm5, implicit-def $ymm6, implicit-def $ymm7, implicit-def $ymm8, implicit-def $ymm9, implicit-def $ymm10, implicit-def $ymm11, implicit-def $ymm12, implicit-def $ymm13, implicit-def $ymm14, implicit-def $ymm15, implicit $xmm0
+    RET64 $xmm0
+
+...
diff --git a/llvm/test/CodeGen/X86/pr214068-vzeroupper-fakeuse-reload.ll b/llvm/test/CodeGen/X86/pr214068-vzeroupper-fakeuse-reload.ll
new file mode 100644
index 0000000000000..d448e97cb4582
--- /dev/null
+++ b/llvm/test/CodeGen/X86/pr214068-vzeroupper-fakeuse-reload.ll
@@ -0,0 +1,145 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py
+; RUN: llc < %s -mtriple=x86_64-unknown-linux-gnu -mattr=+avx2 -O2 | FileCheck %s
+
+; PR214068: Verify that Remove Loads Into Fake Uses does not delete the return
+; value reload into %xmm0 when vzeroupper is inserted before ret.
+
+; ModuleID = 'reduced.ll'
+source_filename = "reduced.cpp"
+target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i64:64-i128:128-f80:128-n8:16:32:64-S128"
+target triple = "x86_64-redhat-linux-gnu"
+
+define double @_ZN12_GLOBAL__N_111Accumulator25calculateErrorAndJacobianENS_6vectorEPNS_4LossE(ptr noundef nonnull align 8 dereferenceable(32) %0, i32 %1, ptr %2, ptr noundef nonnull %3) #2 align 2 {
+; CHECK-LABEL: _ZN12_GLOBAL__N_111Accumulator25calculateErrorAndJacobianENS_6vectorEPNS_4LossE:
+; CHECK:       # %bb.0:
+; CHECK-NEXT:    pushq %rbp
+; CHECK-NEXT:    .cfi_def_cfa_offset 16
+; CHECK-NEXT:    pushq %r15
+; CHECK-NEXT:    .cfi_def_cfa_offset 24
+; CHECK-NEXT:    pushq %r14
+; CHECK-NEXT:    .cfi_def_cfa_offset 32
+; CHECK-NEXT:    pushq %r13
+; CHECK-NEXT:    .cfi_def_cfa_offset 40
+; CHECK-NEXT:    pushq %r12
+; CHECK-NEXT:    .cfi_def_cfa_offset 48
+; CHECK-NEXT:    pushq %rbx
+; CHECK-NEXT:    .cfi_def_cfa_offset 56
+; CHECK-NEXT:    subq $40, %rsp
+; CHECK-NEXT:    .cfi_def_cfa_offset 96
+; CHECK-NEXT:    .cfi_offset %rbx, -56
+; CHECK-NEXT:    .cfi_offset %r12, -48
+; CHECK-NEXT:    .cfi_offset %r13, -40
+; CHECK-NEXT:    .cfi_offset %r14, -32
+; CHECK-NEXT:    .cfi_offset %r15, -24
+; CHECK-NEXT:    .cfi_offset %rbp, -16
+; CHECK-NEXT:    movq %rcx, %r14
+; CHECK-NEXT:    movq %rdx, %rbx
+; CHECK-NEXT:    movl %esi, %ebp
+; CHECK-NEXT:    vxorpd %xmm0, %xmm0, %xmm0
+; CHECK-NEXT:    movq %rdi, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; CHECK-NEXT:    vmovupd %ymm0, (%rdi)
+; CHECK-NEXT:    testl %esi, %esi
+; CHECK-NEXT:    je .LBB0_1
+; CHECK-NEXT:  # %bb.3: # %.preheader
+; CHECK-NEXT:    movslq %ebp, %r13
+; CHECK-NEXT:    shlq $3, %r13
+; CHECK-NEXT:    vxorpd %xmm0, %xmm0, %xmm0
+; CHECK-NEXT:    vmovsd %xmm0, (%rsp) # 8-byte Spill
+; CHECK-NEXT:    xorl %r15d, %r15d
+; CHECK-NEXT:    .p2align 4
+; CHECK-NEXT:  .LBB0_4: # =>This Inner Loop Header: Depth=1
+; CHECK-NEXT:    movq (%rbx,%r15), %r12
+; CHECK-NEXT:    movq (%r12), %rax
+; CHECK-NEXT:    movq %r12, %rdi
+; CHECK-NEXT:    vzeroupper
+; CHECK-NEXT:    callq *(%rax)
+; CHECK-NEXT:    vmovsd %xmm0, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; CHECK-NEXT:    vmovsd %xmm1, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; CHECK-NEXT:    vmulsd %xmm1, %xmm1, %xmm2
+; CHECK-NEXT:    vmulsd %xmm0, %xmm0, %xmm1
+; CHECK-NEXT:    vaddsd %xmm2, %xmm1, %xmm1
+; CHECK-NEXT:    vmovsd %xmm1, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; CHECK-NEXT:    vmovsd (%rsp), %xmm0 # 8-byte Reload
+; CHECK-NEXT:    # xmm0 = mem[0],zero
+; CHECK-NEXT:    vaddsd %xmm1, %xmm0, %xmm0
+; CHECK-NEXT:    vmovsd %xmm0, (%rsp) # 8-byte Spill
+; CHECK-NEXT:    movq (%r14), %rax
+; CHECK-NEXT:    movq %r14, %rdi
+; CHECK-NEXT:    vmovapd %xmm1, %xmm0
+; CHECK-NEXT:    callq *(%rax)
+; CHECK-NEXT:    # fake_use: $r12
+; CHECK-NEXT:    addq $8, %r15
+; CHECK-NEXT:    cmpq %r15, %r13
+; CHECK-NEXT:    jne .LBB0_4
+; CHECK-NEXT:    jmp .LBB0_2
+; CHECK-NEXT:  .LBB0_1:
+; CHECK-NEXT:    vxorpd %xmm0, %xmm0, %xmm0
+; CHECK-NEXT:    vmovsd %xmm0, (%rsp) # 8-byte Spill
+; CHECK-NEXT:  .LBB0_2:
+; CHECK-NEXT:    vmovsd (%rsp), %xmm0 # 8-byte Reload
+; CHECK-NEXT:    # xmm0 = mem[0],zero
+; CHECK-NEXT:    # fake_use: $xmm0
+; CHECK-NEXT:    # fake_use: $r14
+; CHECK-NEXT:    # fake_use: $ebp
+; CHECK-NEXT:    # fake_use: $rbx
+; CHECK-NEXT:    addq $40, %rsp
+; CHECK-NEXT:    .cfi_def_cfa_offset 56
+; CHECK-NEXT:    popq %rbx
+; CHECK-NEXT:    .cfi_def_cfa_offset 48
+; CHECK-NEXT:    popq %r12
+; CHECK-NEXT:    .cfi_def_cfa_offset 40
+; CHECK-NEXT:    popq %r13
+; CHECK-NEXT:    .cfi_def_cfa_offset 32
+; CHECK-NEXT:    popq %r14
+; CHECK-NEXT:    .cfi_def_cfa_offset 24
+; CHECK-NEXT:    popq %r15
+; CHECK-NEXT:    .cfi_def_cfa_offset 16
+; CHECK-NEXT:    popq %rbp
+; CHECK-NEXT:    .cfi_def_cfa_offset 8
+; CHECK-NEXT:    vzeroupper
+; CHECK-NEXT:    retq
+  tail call void @llvm.memset.p0.i64(ptr noundef nonnull align 8 dereferenceable(32) %0, i8 0, i64 32, i1 false)
+  %5 = sext i32 %1 to i64
+  %6 = shl nsw i64 %5, 3
+  %7 = getelementptr inbounds i8, ptr %2, i64 %6
+  %8 = icmp eq i32 %1, 0
+  br i1 %8, label %9, label %11
+
+9:                                                ; preds = %11, %4
+  %10 = phi double [ 0.000000e+00, %4 ], [ %22, %11 ]
+  notail call void (...) @llvm.fake.use(double %10)
+  notail call void (...) @llvm.fake.use(ptr nonnull %3)
+  tail call void (...) @llvm.fake.use(i32 %1)
+  tail call void (...) @llvm.fake.use(ptr %2)
+  notail call void (...) @llvm.fake.use(ptr nonnull %0)
+  ret double %10
+
+11:                                               ; preds = %4, %11
+  %12 = phi double [ %22, %11 ], [ 0.000000e+00, %4 ]
+  %13 = phi ptr [ %26, %11 ], [ %2, %4 ]
+  %14 = load ptr, ptr %13, align 8
+  %15 = load ptr, ptr %14, align 8
+  %16 = load ptr, ptr %15, align 8
+  %17 = tail call { double, double } %16(ptr noundef nonnull align 8 dereferenceable(8) %14)
+  %18 = extractvalue { double, double } %17, 0
+  %19 = extractvalue { double, double } %17, 1
+  %20 = fmul double %19, %19
+  %21 = tail call noundef double @llvm.fmuladd.f64(double %18, double %18, double %20)
+  %22 = fadd double %12, %21
+  %23 = load ptr, ptr %3, align 8
+  %24 = load ptr, ptr %23, align 8
+  %25 = tail call noundef double %24(ptr noundef nonnull align 8 dereferenceable(8) %3, double noundef %21)
+  notail call void (...) @llvm.fake.use(double %21)
+  tail call void (...) @llvm.fake.use(double %18)
+  tail call void (...) @llvm.fake.use(double %19)
+  notail call void (...) @llvm.fake.use(ptr nonnull %14)
+  %26 = getelementptr inbounds nuw i8, ptr %13, i64 8
+  %27 = icmp eq ptr %26, %7
+  br i1 %27, label %9, label %11
+}
+
+declare void @llvm.fake.use(...)
+declare double @llvm.fmuladd.f64(double, double, double)
+declare void @llvm.memset.p0.i64(ptr writeonly captures(none), i8, i64, i1 immarg)
+
+attributes #2 = { mustprogress noinline norecurse uwtable "min-legal-vector-width"="0" "no-trapping-math"="true" "stack-protector-buffer-size"="8" "target-cpu"="x86-64" "target-features"="+avx,+avx2,+cmov,+crc32,+cx8,+fxsr,+mmx,+popcnt,+sse,+sse2,+sse3,+sse4.1,+sse4.2,+ssse3,+x87,+xsave" "tune-cpu"="generic" }



More information about the llvm-commits mailing list