[llvm] RegisterPressure: Detect dead physreg defs from LiveIntervals (PR #225079)
via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 21 06:03:44 PDT 2026
llvmorg-github-actions[bot] wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-llvm-regalloc
@llvm/pr-subscribers-backend-x86
Author: Matt Arsenault (arsenm)
<details>
<summary>Changes</summary>
This reverts the remainder of #<!-- -->222627, which was partially reverted by
on dead flags. This is a prerequisite to deleting LiveVariables.
When constructing the PressureDiff for an instruction during scheduling
DAG construction, dead defs were only recognized from the dead flag on the
operand. This implicitly relied on preprocessing done by LiveVariables to
fixup inconsistent dead flags with overlapping registers in other operands.
Dead flags have no verifier-enforced rules and are thus unreliable.
Before LiveVariables, consider this example:
dead $eax = MOV32r0 implicit-def dead $eflags, implicit-def $rax
; $rax is never used
$rax is never used, but only the $eax def is dead-flagged and the overlapping
implicit-def $rax is not. The shared $eax register units are covered by the
non-dead $rax def and so are counted as live defs. That shared unit is then
decremented on recede without a matching increment, tripping the "PSet
overflow/underflow" assertion in getUpwardPressureDelta.
Operand dead flags only describe physreg liveness reliably when
LiveVariables produced them. For a wide def that is only partly used,
HandlePhysRegKill marks the superregister def dead and adds a separate,
explicitly non-dead def of the used sub-register. In the example, LiveVariables
fixes up the flags to liven $eax, and make $rax dead:
$eax = MOV32r0 implicit-def dead $eflags, implicit-def dead $rax
The rule is a regunit unit is dead iff every def covering it is dead. When
LiveVariables does not run, the superregister dead flag is simply missing and
the flags no longer describe liveness.
Detect dead defs from LiveIntervals instead, matching what the pressure
tracker already does on the recede path. We need to not take a conservative no
answer if getCachedRegUnit doesn't already have the computed interval, so add
a parameter to not use the cache. This is the source of most of the churn, which
needs to de-constify the LiveIntervals passed around.
I'm somewhat dissatisfied with relying on LiveIntervals and ignoring the flags. I'm
separately working on adding some verifier rules for dead flags, though I'm not
sure that will be sufficient to avoid this.
Co-authored-by: Claude claude-opus-4.8 <noreply@<!-- -->anthropic.com>
---
Patch is 83.96 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/225079.diff
18 Files Affected:
- (modified) llvm/include/llvm/CodeGen/RegisterPressure.h (+6-6)
- (modified) llvm/lib/CodeGen/RegisterPressure.cpp (+27-21)
- (modified) llvm/lib/CodeGen/ScheduleDAGInstrs.cpp (+4)
- (modified) llvm/lib/Target/AMDGPU/AMDGPUNextUseAnalysis.cpp (+4-5)
- (modified) llvm/lib/Target/AMDGPU/GCNRegPressure.cpp (+1-1)
- (modified) llvm/lib/Target/AMDGPU/GCNRegPressure.h (+4-4)
- (modified) llvm/test/CodeGen/X86/masked-udiv.ll (+208-225)
- (modified) llvm/test/CodeGen/X86/min-legal-vector-width.ll (+32-32)
- (added) llvm/test/CodeGen/X86/misched-pressure-dead-physreg-superreg.mir (+60)
- (modified) llvm/test/CodeGen/X86/ssub_sat_plus.ll (+2-2)
- (modified) llvm/test/CodeGen/X86/statepoint-ra.ll (+4-1)
- (modified) llvm/test/CodeGen/X86/statepoint-vreg-unlimited-tied-opnds.ll (+21-20)
- (modified) llvm/test/CodeGen/X86/udiv_fix.ll (+16-18)
- (modified) llvm/test/CodeGen/X86/udiv_fix_sat.ll (+40-42)
- (modified) llvm/test/CodeGen/X86/usub_sat_plus.ll (+2-2)
- (modified) llvm/test/CodeGen/X86/vector-idiv-strictfp.ll (+87-104)
- (modified) llvm/test/CodeGen/X86/vector-idiv-udiv-512.ll (+60-72)
- (modified) llvm/test/CodeGen/X86/xmulo.ll (+6-6)
``````````diff
diff --git a/llvm/include/llvm/CodeGen/RegisterPressure.h b/llvm/include/llvm/CodeGen/RegisterPressure.h
index 47c2a31d4317a..bd6b89df9c049 100644
--- a/llvm/include/llvm/CodeGen/RegisterPressure.h
+++ b/llvm/include/llvm/CodeGen/RegisterPressure.h
@@ -187,19 +187,19 @@ class RegisterOperands {
/// Use liveness information to find dead defs at \p MI's dead slot not marked
/// with a dead flag and move them to the DeadDefs vector. This only considers
/// the merged live interval for defs, not the per-lane sub-ranges.
- LLVM_ABI void detectDeadDefs(const MachineInstr &MI, const LiveIntervals &LIS,
+ LLVM_ABI void detectDeadDefs(const MachineInstr &MI, LiveIntervals &LIS,
const MachineRegisterInfo &MRI);
/// Use liveness information to find out which uses/defs are partially
/// undefined/dead at \p Pos and adjust the VRegMaskOrUnits accordingly.
- LLVM_ABI void adjustLaneLiveness(const LiveIntervals &LIS,
+ LLVM_ABI void adjustLaneLiveness(LiveIntervals &LIS,
const MachineRegisterInfo &MRI,
SlotIndex Pos);
/// Use liveness information to find out which uses/defs are partially
/// undefined/dead at the \p MI's position and adjust the VRegMaskOrUnits
/// accordingly. Missing read-undef and dead flags are added to \p MI.
- LLVM_ABI void adjustLaneLiveness(const LiveIntervals &LIS,
+ LLVM_ABI void adjustLaneLiveness(LiveIntervals &LIS,
const MachineRegisterInfo &MRI,
MachineInstr &MI);
@@ -211,7 +211,7 @@ class RegisterOperands {
VRegMaskOrUnit *adjustDef(VRegMaskOrUnit &Def, LaneBitmask LiveAfterDef);
/// Use liveness information at \p Pos to adjust the lanemask of all uses.
- void adjustUses(const LiveIntervals &LIS, const MachineRegisterInfo &MRI,
+ void adjustUses(LiveIntervals &LIS, const MachineRegisterInfo &MRI,
SlotIndex Pos);
};
@@ -382,7 +382,7 @@ class RegPressureTracker {
const TargetRegisterInfo *TRI = nullptr;
const RegisterClassInfo *RCI = nullptr;
const MachineRegisterInfo *MRI = nullptr;
- const LiveIntervals *LIS = nullptr;
+ LiveIntervals *LIS = nullptr;
/// We currently only allow pressure tracking within a block.
const MachineBasicBlock *MBB = nullptr;
@@ -423,7 +423,7 @@ class RegPressureTracker {
LLVM_ABI void reset();
LLVM_ABI void init(const MachineFunction *mf, const RegisterClassInfo *rci,
- const LiveIntervals *lis, const MachineBasicBlock *mbb,
+ LiveIntervals *lis, const MachineBasicBlock *mbb,
MachineBasicBlock::const_iterator pos, bool TrackLaneMasks,
bool TrackUntiedDefs);
diff --git a/llvm/lib/CodeGen/RegisterPressure.cpp b/llvm/lib/CodeGen/RegisterPressure.cpp
index 10a0d9b02d9f4..40e3c003e10fb 100644
--- a/llvm/lib/CodeGen/RegisterPressure.cpp
+++ b/llvm/lib/CodeGen/RegisterPressure.cpp
@@ -254,7 +254,7 @@ void RegPressureTracker::reset() {
/// TODO: Add support for pressure without LiveIntervals.
void RegPressureTracker::init(const MachineFunction *mf,
const RegisterClassInfo *rci,
- const LiveIntervals *lis,
+ LiveIntervals *lis,
const MachineBasicBlock *mbb,
MachineBasicBlock::const_iterator pos,
bool TrackLaneMasks, bool TrackUntiedDefs) {
@@ -411,10 +411,11 @@ static void removeRegLanes(SmallVectorImpl<VRegMaskOrUnit> &RegUnits,
}
static LaneBitmask
-getLanesWithProperty(const LiveIntervals &LIS, const MachineRegisterInfo &MRI,
+getLanesWithProperty(LiveIntervals &LIS, const MachineRegisterInfo &MRI,
bool TrackLaneMasks, VirtRegOrUnit VRegOrUnit,
SlotIndex Pos, LaneBitmask SafeDefault,
- bool (*Property)(const LiveRange &LR, SlotIndex Pos)) {
+ bool (*Property)(const LiveRange &LR, SlotIndex Pos),
+ bool ComputePhysRegs = false) {
if (VRegOrUnit.isVirtualReg()) {
const LiveInterval &LI = LIS.getInterval(VRegOrUnit.asVirtualReg());
LaneBitmask Result;
@@ -431,22 +432,27 @@ getLanesWithProperty(const LiveIntervals &LIS, const MachineRegisterInfo &MRI,
return Result;
} else {
- const LiveRange *LR = LIS.getCachedRegUnit(VRegOrUnit.asMCRegUnit());
- // Be prepared for missing liveranges: We usually do not compute liveranges
- // for physical registers on targets with many registers (GPUs).
+ MCRegUnit Unit = VRegOrUnit.asMCRegUnit();
+ // We usually do not compute liveranges for physical registers on targets
+ // with many registers (GPUs), so the cached range may be absent. Callers
+ // that require an authoritative answer pass ComputePhysRegs to force the
+ // range to be computed on demand.
+ const LiveRange *LR =
+ ComputePhysRegs ? &LIS.getRegUnit(Unit) : LIS.getCachedRegUnit(Unit);
if (LR == nullptr)
return SafeDefault;
return Property(*LR, Pos) ? LaneBitmask::getAll() : LaneBitmask::getNone();
}
}
-static LaneBitmask getLiveLanesAt(const LiveIntervals &LIS,
+static LaneBitmask getLiveLanesAt(LiveIntervals &LIS,
const MachineRegisterInfo &MRI,
bool TrackLaneMasks, VirtRegOrUnit VRegOrUnit,
- SlotIndex Pos) {
+ SlotIndex Pos, bool ComputePhysRegs = false) {
return getLanesWithProperty(
LIS, MRI, TrackLaneMasks, VRegOrUnit, Pos, LaneBitmask::getAll(),
- [](const LiveRange &LR, SlotIndex Pos) { return LR.liveAt(Pos); });
+ [](const LiveRange &LR, SlotIndex Pos) { return LR.liveAt(Pos); },
+ ComputePhysRegs);
}
namespace {
@@ -472,12 +478,9 @@ class RegisterOperandsCollector {
for (ConstMIBundleOperands OperI(MI); OperI.isValid(); ++OperI)
collectOperand(*OperI);
- // An instruction can have overlapping defs where only some carry the dead
- // flag, for example a dead super-register def alongside a live sub-register
- // def. A register unit is dead if any def covering it is dead, so subtract
- // the dead defs from the live defs.
- for (const VRegMaskOrUnit &P : RegOpers.DeadDefs)
- removeRegLanes(RegOpers.Defs, P);
+ // Remove redundant physreg dead defs.
+ for (const VRegMaskOrUnit &P : RegOpers.Defs)
+ removeRegLanes(RegOpers.DeadDefs, P);
}
void collectInstrLanes(const MachineInstr &MI) const {
@@ -573,17 +576,20 @@ void RegisterOperands::collect(const MachineInstr &MI,
}
void RegisterOperands::detectDeadDefs(const MachineInstr &MI,
- const LiveIntervals &LIS,
+ LiveIntervals &LIS,
const MachineRegisterInfo &MRI) {
SlotIndex DeadSlotIdx = LIS.getInstructionIndex(MI).getDeadSlot();
for (auto *I = Defs.begin(); I != Defs.end(); /*empty*/) {
- LaneBitmask LiveAfter = getLiveLanesAt(LIS, MRI, /*TrackLaneMasks=*/false,
- I->VRegOrUnit, DeadSlotIdx);
+ // Force physreg unit ranges to be computed, we need to accurately know if a
+ // physreg is dead.
+ LaneBitmask LiveAfter =
+ getLiveLanesAt(LIS, MRI, /*TrackLaneMasks=*/false, I->VRegOrUnit,
+ DeadSlotIdx, /*ComputePhysRegs=*/true);
I = adjustDef(*I, LiveAfter);
}
}
-void RegisterOperands::adjustLaneLiveness(const LiveIntervals &LIS,
+void RegisterOperands::adjustLaneLiveness(LiveIntervals &LIS,
const MachineRegisterInfo &MRI,
SlotIndex Pos) {
for (auto *I = Defs.begin(); I != Defs.end(); /*empty*/) {
@@ -594,7 +600,7 @@ void RegisterOperands::adjustLaneLiveness(const LiveIntervals &LIS,
adjustUses(LIS, MRI, Pos.getBaseIndex());
}
-void RegisterOperands::adjustLaneLiveness(const LiveIntervals &LIS,
+void RegisterOperands::adjustLaneLiveness(LiveIntervals &LIS,
const MachineRegisterInfo &MRI,
MachineInstr &MI) {
SlotIndex Pos = LIS.getInstructionIndex(MI);
@@ -644,7 +650,7 @@ VRegMaskOrUnit *RegisterOperands::adjustDef(VRegMaskOrUnit &Def,
return &Def + 1;
}
-void RegisterOperands::adjustUses(const LiveIntervals &LIS,
+void RegisterOperands::adjustUses(LiveIntervals &LIS,
const MachineRegisterInfo &MRI,
SlotIndex Pos) {
for (auto &[VRegOrUnit, LaneMask] : Uses) {
diff --git a/llvm/lib/CodeGen/ScheduleDAGInstrs.cpp b/llvm/lib/CodeGen/ScheduleDAGInstrs.cpp
index c929276b219f7..b59898e4cd4e7 100644
--- a/llvm/lib/CodeGen/ScheduleDAGInstrs.cpp
+++ b/llvm/lib/CodeGen/ScheduleDAGInstrs.cpp
@@ -796,6 +796,10 @@ void ScheduleDAGInstrs::buildSchedGraph(AAResults *AA,
if (TrackLaneMasks) {
SlotIndex SlotIdx = LIS->getInstructionIndex(MI);
RegOpers.adjustLaneLiveness(*LIS, MRI, SlotIdx);
+ } else if (LIS) {
+ // Detect dead defs from LiveIntervals instead of trusting operand dead
+ // flags.
+ RegOpers.detectDeadDefs(MI, *LIS, MRI);
}
if (PDiffs != nullptr)
PDiffs->addInstruction(SU->NodeNum, RegOpers, MRI);
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUNextUseAnalysis.cpp b/llvm/lib/Target/AMDGPU/AMDGPUNextUseAnalysis.cpp
index a667b96c0d089..2a19add12b8e4 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUNextUseAnalysis.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUNextUseAnalysis.cpp
@@ -2414,7 +2414,7 @@ void printDistanceFromDefToUse(json::OStream &J, const MachineFunction &MF,
void printNextUseDistancesAsJson(json::OStream &J, const MachineFunction &MF,
const AMDGPUNextUseAnalysis &NUA,
const AMDGPUNextUseAnalysisImpl &NUAImpl,
- const LiveIntervals &LIS) {
+ LiveIntervals &LIS) {
using UseDistancePair = AMDGPUNextUseAnalysis::UseDistancePair;
const Function &F = MF.getFunction();
const Module *M = F.getParent();
@@ -2501,8 +2501,7 @@ void printNextUseDistancesAsJson(json::OStream &J, const MachineFunction &MF,
void printAsJson(raw_ostream &FallbackOS, TimerGroup &JsonTimerGroup,
Timer &JsonTimer, const MachineFunction &MF,
const AMDGPUNextUseAnalysis &NUA,
- const AMDGPUNextUseAnalysisImpl &NUAImpl,
- const LiveIntervals &LIS) {
+ const AMDGPUNextUseAnalysisImpl &NUAImpl, LiveIntervals &LIS) {
std::string FN = DumpNextUseDistanceAsJson;
auto dump = [&](raw_ostream &OS) {
@@ -2551,7 +2550,7 @@ bool AMDGPUNextUseAnalysisPrinterLegacyPass::runOnMachineFunction(
Timer JsonTimer("json", "Total time spent generating json", JsonTimerGroup);
JsonTimer.startTimer();
- const LiveIntervals &LIS = getAnalysis<LiveIntervalsWrapperPass>().getLIS();
+ LiveIntervals &LIS = getAnalysis<LiveIntervalsWrapperPass>().getLIS();
const AMDGPUNextUseAnalysis &NUA =
getAnalysis<AMDGPUNextUseAnalysisLegacyPass>().getNextUseAnalysis();
@@ -2600,7 +2599,7 @@ AMDGPUNextUseAnalysisPrinterPass::run(MachineFunction &MF,
Timer JsonTimer("json", "Total time spent generating json", JsonTimerGroup);
JsonTimer.startTimer();
- const LiveIntervals &LIS = MFAM.getResult<LiveIntervalsAnalysis>(MF);
+ LiveIntervals &LIS = MFAM.getResult<LiveIntervalsAnalysis>(MF);
const AMDGPUNextUseAnalysis &NUA =
MFAM.getResult<AMDGPUNextUseAnalysisPass>(MF);
diff --git a/llvm/lib/Target/AMDGPU/GCNRegPressure.cpp b/llvm/lib/Target/AMDGPU/GCNRegPressure.cpp
index 53617e89af757..312fa4cf852ae 100644
--- a/llvm/lib/Target/AMDGPU/GCNRegPressure.cpp
+++ b/llvm/lib/Target/AMDGPU/GCNRegPressure.cpp
@@ -971,7 +971,7 @@ getRegLiveThroughMask(const MachineRegisterInfo &MRI, const LiveIntervals &LIS,
bool GCNRegPressurePrinter::runOnMachineFunction(MachineFunction &MF) {
const MachineRegisterInfo &MRI = MF.getRegInfo();
const TargetRegisterInfo *TRI = MRI.getTargetRegisterInfo();
- const LiveIntervals &LIS = getAnalysis<LiveIntervalsWrapperPass>().getLIS();
+ LiveIntervals &LIS = getAnalysis<LiveIntervalsWrapperPass>().getLIS();
auto &OS = dbgs();
diff --git a/llvm/lib/Target/AMDGPU/GCNRegPressure.h b/llvm/lib/Target/AMDGPU/GCNRegPressure.h
index 021d879732832..5c8bd5a2cd76a 100644
--- a/llvm/lib/Target/AMDGPU/GCNRegPressure.h
+++ b/llvm/lib/Target/AMDGPU/GCNRegPressure.h
@@ -323,13 +323,13 @@ class GCNRPTracker {
using LiveRegSet = DenseMap<unsigned, LaneBitmask>;
protected:
- const LiveIntervals &LIS;
+ LiveIntervals &LIS;
LiveRegSet LiveRegs;
GCNRegPressure CurPressure, MaxPressure;
const MachineInstr *LastTrackedMI = nullptr;
mutable const MachineRegisterInfo *MRI = nullptr;
- GCNRPTracker(const LiveIntervals &LIS_) : LIS(LIS_) {}
+ GCNRPTracker(LiveIntervals &LIS_) : LIS(LIS_) {}
/// Resets tracker before or \p After the provided \p MI, which can be a debug
/// instruction.
@@ -370,7 +370,7 @@ getLiveRegs(SlotIndex SI, const LiveIntervals &LIS,
class GCNUpwardRPTracker : public GCNRPTracker {
public:
- GCNUpwardRPTracker(const LiveIntervals &LIS_) : GCNRPTracker(LIS_) {}
+ GCNUpwardRPTracker(LiveIntervals &LIS_) : GCNRPTracker(LIS_) {}
using GCNRPTracker::reset;
@@ -408,7 +408,7 @@ class GCNDownwardRPTracker : public GCNRPTracker {
MachineBasicBlock::const_iterator MBBEnd;
public:
- GCNDownwardRPTracker(const LiveIntervals &LIS_) : GCNRPTracker(LIS_) {}
+ GCNDownwardRPTracker(LiveIntervals &LIS_) : GCNRPTracker(LIS_) {}
using GCNRPTracker::reset;
diff --git a/llvm/test/CodeGen/X86/masked-udiv.ll b/llvm/test/CodeGen/X86/masked-udiv.ll
index 51114074232cb..7fbddc22cc38a 100644
--- a/llvm/test/CodeGen/X86/masked-udiv.ll
+++ b/llvm/test/CodeGen/X86/masked-udiv.ll
@@ -204,40 +204,39 @@ define <2 x i64> @udiv_v2i64(<2 x i64> %x, <2 x i64> %y, <2 x i1> %m) {
define <4 x i64> @udiv_v4i64(<4 x i64> %x, <4 x i64> %y, <4 x i1> %m) {
; SSE2-LABEL: udiv_v4i64:
; SSE2: # %bb.0:
-; SSE2-NEXT: pshufd {{.*#+}} xmm6 = xmm4[0,0,1,1]
+; SSE2-NEXT: movdqa %xmm0, %xmm5
+; SSE2-NEXT: pshufd {{.*#+}} xmm6 = xmm4[2,2,3,3]
; SSE2-NEXT: pslld $31, %xmm6
; SSE2-NEXT: psrad $31, %xmm6
-; SSE2-NEXT: movdqa {{.*#+}} xmm5 = [1,1]
-; SSE2-NEXT: pand %xmm6, %xmm2
-; SSE2-NEXT: pandn %xmm5, %xmm6
-; SSE2-NEXT: por %xmm2, %xmm6
-; SSE2-NEXT: movq %xmm6, %rcx
+; SSE2-NEXT: pshufd {{.*#+}} xmm4 = xmm4[0,0,1,1]
+; SSE2-NEXT: pslld $31, %xmm4
+; SSE2-NEXT: psrad $31, %xmm4
+; SSE2-NEXT: movdqa {{.*#+}} xmm7 = [1,1]
+; SSE2-NEXT: pand %xmm4, %xmm2
+; SSE2-NEXT: pandn %xmm7, %xmm4
+; SSE2-NEXT: por %xmm2, %xmm4
+; SSE2-NEXT: movq %xmm4, %rcx
; SSE2-NEXT: movq %xmm0, %rax
; SSE2-NEXT: xorl %edx, %edx
; SSE2-NEXT: divq %rcx
-; SSE2-NEXT: movq %rax, %rcx
-; SSE2-NEXT: pshufd {{.*#+}} xmm2 = xmm6[2,3,2,3]
-; SSE2-NEXT: movq %xmm2, %rsi
-; SSE2-NEXT: pshufd {{.*#+}} xmm0 = xmm0[2,3,2,3]
-; SSE2-NEXT: movq %xmm0, %rax
+; SSE2-NEXT: movq %rax, %xmm0
+; SSE2-NEXT: pshufd {{.*#+}} xmm2 = xmm4[2,3,2,3]
+; SSE2-NEXT: movq %xmm2, %rcx
+; SSE2-NEXT: pshufd {{.*#+}} xmm2 = xmm5[2,3,2,3]
+; SSE2-NEXT: movq %xmm2, %rax
; SSE2-NEXT: xorl %edx, %edx
-; SSE2-NEXT: divq %rsi
-; SSE2-NEXT: movq %rax, %rsi
-; SSE2-NEXT: pshufd {{.*#+}} xmm4 = xmm4[2,2,3,3]
-; SSE2-NEXT: pslld $31, %xmm4
-; SSE2-NEXT: psrad $31, %xmm4
-; SSE2-NEXT: pand %xmm4, %xmm3
-; SSE2-NEXT: pandn %xmm5, %xmm4
-; SSE2-NEXT: por %xmm3, %xmm4
-; SSE2-NEXT: movq %xmm4, %rdi
+; SSE2-NEXT: divq %rcx
+; SSE2-NEXT: movq %rax, %xmm2
+; SSE2-NEXT: punpcklqdq {{.*#+}} xmm0 = xmm0[0],xmm2[0]
+; SSE2-NEXT: pand %xmm6, %xmm3
+; SSE2-NEXT: pandn %xmm7, %xmm6
+; SSE2-NEXT: por %xmm3, %xmm6
+; SSE2-NEXT: movq %xmm6, %rcx
; SSE2-NEXT: movq %xmm1, %rax
; SSE2-NEXT: xorl %edx, %edx
-; SSE2-NEXT: divq %rdi
-; SSE2-NEXT: movq %rcx, %xmm0
-; SSE2-NEXT: movq %rsi, %xmm2
-; SSE2-NEXT: punpcklqdq {{.*#+}} xmm0 = xmm0[0],xmm2[0]
+; SSE2-NEXT: divq %rcx
; SSE2-NEXT: movq %rax, %xmm2
-; SSE2-NEXT: pshufd {{.*#+}} xmm3 = xmm4[2,3,2,3]
+; SSE2-NEXT: pshufd {{.*#+}} xmm3 = xmm6[2,3,2,3]
; SSE2-NEXT: movq %xmm3, %rcx
; SSE2-NEXT: pshufd {{.*#+}} xmm1 = xmm1[2,3,2,3]
; SSE2-NEXT: movq %xmm1, %rax
@@ -251,39 +250,38 @@ define <4 x i64> @udiv_v4i64(<4 x i64> %x, <4 x i64> %y, <4 x i1> %m) {
; SSE42-LABEL: udiv_v4i64:
; SSE42: # %bb.0:
; SSE42-NEXT: movdqa %xmm0, %xmm5
+; SSE42-NEXT: pshufd {{.*#+}} xmm6 = xmm4[2,2,3,3]
; SSE42-NEXT: pmovzxdq {{.*#+}} xmm0 = xmm4[0],zero,xmm4[1],zero
; SSE42-NEXT: psllq $63, %xmm0
-; SSE42-NEXT: movapd {{.*#+}} xmm6 = [1,1]
-; SSE42-NEXT: movapd %xmm6, %xmm7
+; SSE42-NEXT: movapd {{.*#+}} xmm4 = [1,1]
+; SSE42-NEXT: movapd %xmm4, %xmm7
; SSE42-NEXT: blendvpd %xmm0, %xmm2, %xmm7
; SSE42-NEXT: pextrq $1, %xmm7, %rcx
; SSE42-NEXT: pextrq $1, %xmm5, %rax
+; SSE42-NEXT: psllq $63, %xmm6
; SSE42-NEXT: xorl %edx, %edx
; SSE42-NEXT: divq %rcx
-; SSE42-NEXT: movq %rax, %rcx
-; SSE42-NEXT: pshufd {{.*#+}} xmm0 = xmm4[2,2,3,3]
-; SSE42-NEXT: psllq $63, %xmm0
-; SSE42-NEXT: movq %xmm7, %rdi
-; SSE42-NEXT: blendvpd %xmm0, %xmm3, %xmm6
-; SSE42-NEXT: pextrq $1, %xmm6, %r8
-; SSE42-NEXT: pextrq $1, %xmm1, %rsi
+; SSE42-NEXT: movq %rax, %xmm8
+; SSE42-NEXT: movq %xmm7, %rcx
; SSE42-NEXT: movq %xmm5, %rax
; SSE42-NEXT: xorl %edx, %edx
-; SSE42-NEXT: divq %rdi
-; SSE42-NEXT: movq %rax, %rdi
-; SSE42-NEXT: movq %rsi, %rax
-; SSE42-NEXT: xorl %edx, %edx
-; SSE42-NEXT: divq %r8
-; SSE42-NEXT: movq %rcx, %xmm2
-; SSE42-NEXT: movq %rdi, %xmm0
-; SSE42-NEXT: punpcklqdq {{.*#+}} xmm0 = xmm0[0],xmm2[0]
+; SSE42-NEXT: divq %rcx
; SSE42-NEXT: movq %rax, %xmm2
-; SSE42-NEXT: movq %xmm6, %rcx
+; SSE42-NEXT: movdqa %xmm6, %xmm0
+; SSE42-NEXT: blendvpd %xmm0, %xmm3, %xmm4
+; SSE42-NEXT: pextrq $1, %xmm4, %rcx
+; SSE42-NEXT: pextrq $1, %xmm1, %rax
+; SSE42-NEXT: punpcklqdq {{.*#+}} xmm2 = xmm2[0],xmm8[0]
+; SSE42-NEXT: xorl %edx, %edx
+; SSE42-NEXT: divq %rcx
+; SSE42-NEXT: movq %rax, %xmm0
+; SSE42-NEXT: movq %xmm4, %rcx
; SSE42-NEXT: movq %xmm1, %rax
; SSE42-NEXT: xorl %edx, %edx
; SSE42-NEXT: divq %rcx
; SSE42-NEXT: movq %rax, %xmm1
-; SSE42-NEXT: punpcklqdq {{.*#+}} xmm1 = xmm1[0],xmm2[0]
+; SSE42-NEXT: punpcklqdq {{.*#+}} xmm1 = xmm1[0],xmm0[0]
+; SSE42-NEXT: movdqa %xmm2, %xmm0
; SSE42-NEXT: retq
;
; AVX2-LABEL: udiv_v4i64:
@@ -303,20 +301,19 @@ define <4 x i64> @udiv_v4i64(<4 x i64> %x, <4 x i64> %y, <4 x i1> %m) {
; AVX2-NEXT: vmovq %xmm3, %rax
; AVX2-NEXT: xorl %edx, %edx
; AVX2-NEXT: divq %rsi
-; AVX2-NEXT: movq %rax, %rsi
-; AVX2-NEXT: vpextrq $1, %xmm1, %rdi
+; AVX2-NEXT: vmovq %rcx, %xmm2
+; AVX2-NEXT: vpextrq $1, %xmm1, %rcx
+; AVX2-NEXT: vmovq %rax, %xmm3
; AVX2-NEXT: vpextrq $1, %xmm0, %rax
; AVX2-NEXT: xorl %edx, %edx
-; AVX2-NEXT: divq %rdi
-; AVX2-NEXT: movq %rax, %rdi
-; AVX2-NEXT: vmovq %rcx, %xmm2
-; AVX2-NEXT: vmovq %rsi, %xmm3
+; AVX2-NEXT: divq %rcx
+; AVX2-NEXT: movq %rax, %rcx
; AVX2-NEXT: vpunpcklqdq {{.*#+}} xmm2 = xmm3[0],xmm2[0]
-; AVX2-NEXT: vmovq %xmm1, %rcx
+; AVX2-NEXT: vmovq %xmm1, %rsi
; AVX2-NEXT: vmovq %xmm0, %rax
; AVX2-NEXT: xorl %edx, %edx
-; AVX2-NEXT: divq %rcx
-; AVX2-NEXT: vmovq %rdi, %xmm0
+; AVX2-NEXT: divq %rsi
+; AVX2-NEXT: vmovq %rcx, %xmm0
; AVX2-NEXT: vmovq %rax, %xmm1
; AVX2-NEXT: vpunpcklqdq {{.*#+}} xmm0 = xmm1[0],xmm0[0]
; AVX2-NEXT: vinserti128 $1, %xmm2, %ymm0, %ymm0
@@ -819,181 +816,169 @@ define <3 x i10> @udiv_v3i10(<3 x i10> %x, <3 x i10> %y, <3 x i1> %m) {
define <8 x i64> @udiv_v8i64(<8 x i64> %x, <8 x i64> %y, <8 x i1> %m) {
; SSE2-LABEL: udiv_v8i64:
; SSE2: # %bb.0:
-; SSE2-NEXT: movdqa {{[0-9]+}}(%rsp), %xmm8
-; SSE2-NEXT: pshufd {{.*#+}} xmm9 = xmm8[0,0,0,0]
+; SSE2-NEXT: movdqa %xmm0, %xmm8
+; SSE2-NEXT: movdqa {{[0-9]+}}(%rsp), %xm...
[truncated]
``````````
</details>
https://github.com/llvm/llvm-project/pull/225079
More information about the llvm-commits
mailing list