[llvm] [CodeGen] Add target-independent conditional-compare formation pass (PR #219359)
Feng Zou via llvm-commits
llvm-commits at lists.llvm.org
Thu Aug 27 21:24:09 PDT 2026
https://github.com/fzou1 updated https://github.com/llvm/llvm-project/pull/219359
>From 633f50d2754e0de3ac02218991a81e4e7a6895fa Mon Sep 17 00:00:00 2001
From: "Zou, Feng" <feng.zou at intel.com>
Date: Fri, 28 Aug 2026 04:34:54 +0200
Subject: [PATCH] [CodeGen] Add target-independent conditional-compare
formation pass
Unify the near-duplicate X86ConditionalCompares and
AArch64ConditionalCompares passes into a single target-independent
MachineConditionalCompares pass in lib/CodeGen, modeled on
EarlyIfConversion. The pass folds a CFG triangle of chained
short-circuit conditions into a conditional-compare chain
(CCMP/CTEST on X86; CCMP/CCMN/FCCMP on AArch64), eliminating a branch.
All target-specific behavior is delegated through new hooks:
- TargetInstrInfo::getConditionalCompareFlagReg / canConvertToCCMP /
convertToCCMP / getCCMPCodeSizeDelta
- TargetSubtargetInfo::enableCCMPFormation and a CCmpConvHeuristics
tuning struct (a code-size-delta path when minimizing size).
The engine (SSACCmpConv) is heuristics-free and stores condition codes
opaquely, round-tripping them through the target hooks. The driver
carries the shared MachineTraceMetrics cost model: a 3/4-misprediction-
penalty branch-delay budget plus a resource-depth check, applied
identically on both targets. The only per-target difference is the
CCmpConvHeuristics tuning struct: AArch64 enables the MinSize
code-size-delta path (getCCMPCodeSizeDelta), while X86 leaves it off.
This reproduces each target's historical cost model.
Both targets are migrated onto the new pass (legacy and new-PM,
registered as "machine-ccmp") and their private passes removed. X86
gains new-PM support for the first time; AArch64 keeps its dual
registration via the generic pass.
Assisted-by: Claude Opus 4.8 (1M context) <noreply at anthropic.com>
---
.../llvm/CodeGen/MachineConditionalCompares.h | 25 +
llvm/include/llvm/CodeGen/Passes.h | 4 +
llvm/include/llvm/CodeGen/TargetInstrInfo.h | 70 ++
.../llvm/CodeGen/TargetSubtargetInfo.h | 15 +
llvm/include/llvm/InitializePasses.h | 1 +
.../llvm/Passes/MachinePassRegistry.def | 1 +
llvm/lib/CodeGen/CMakeLists.txt | 1 +
llvm/lib/CodeGen/CodeGen.cpp | 1 +
.../MachineConditionalCompares.cpp} | 664 ++++++------------
llvm/lib/Passes/PassBuilder.cpp | 1 +
llvm/lib/Target/AArch64/AArch64.h | 9 -
llvm/lib/Target/AArch64/AArch64InstrInfo.cpp | 316 +++++++++
llvm/lib/Target/AArch64/AArch64InstrInfo.h | 14 +
.../Target/AArch64/AArch64PassRegistry.def | 1 -
llvm/lib/Target/AArch64/AArch64Subtarget.cpp | 14 +
llvm/lib/Target/AArch64/AArch64Subtarget.h | 4 +
.../Target/AArch64/AArch64TargetMachine.cpp | 7 +-
llvm/lib/Target/AArch64/CMakeLists.txt | 1 -
llvm/lib/Target/X86/X86CodeGenPassBuilder.cpp | 2 +
llvm/lib/Target/X86/X86InstrInfo.cpp | 258 +++++++
llvm/lib/Target/X86/X86InstrInfo.h | 12 +
llvm/lib/Target/X86/X86Subtarget.cpp | 8 +
llvm/lib/Target/X86/X86Subtarget.h | 2 +
llvm/lib/Target/X86/X86TargetMachine.cpp | 1 +
llvm/test/CodeGen/AArch64/O3-pipeline.ll | 4 +-
llvm/test/CodeGen/AArch64/arm64-ccmp.ll | 4 +-
.../AArch64/ccmp-look-through-copy.mir | 4 +-
.../CodeGen/AArch64/ccmp-successor-probs.mir | 4 +-
llvm/test/CodeGen/X86/apx/ccmp-cost-model.mir | 164 +++++
.../X86/apx/ccmp-look-through-copy.mir | 41 ++
llvm/test/CodeGen/X86/apx/ccmp-reject.mir | 495 +++++++++++++
llvm/test/CodeGen/X86/apx/ccmp.ll | 29 +-
llvm/test/CodeGen/X86/llc-pipeline-npm.ll | 2 +
llvm/test/CodeGen/X86/opt-pipeline.ll | 3 +
34 files changed, 1698 insertions(+), 484 deletions(-)
create mode 100644 llvm/include/llvm/CodeGen/MachineConditionalCompares.h
rename llvm/lib/{Target/AArch64/AArch64ConditionalCompares.cpp => CodeGen/MachineConditionalCompares.cpp} (55%)
create mode 100644 llvm/test/CodeGen/X86/apx/ccmp-cost-model.mir
create mode 100644 llvm/test/CodeGen/X86/apx/ccmp-look-through-copy.mir
create mode 100644 llvm/test/CodeGen/X86/apx/ccmp-reject.mir
diff --git a/llvm/include/llvm/CodeGen/MachineConditionalCompares.h b/llvm/include/llvm/CodeGen/MachineConditionalCompares.h
new file mode 100644
index 0000000000000..63ddf90617710
--- /dev/null
+++ b/llvm/include/llvm/CodeGen/MachineConditionalCompares.h
@@ -0,0 +1,25 @@
+//===- llvm/CodeGen/MachineConditionalCompares.h ---------------*- C++ -*-===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+#ifndef LLVM_CODEGEN_MACHINECONDITIONALCOMPARES_H
+#define LLVM_CODEGEN_MACHINECONDITIONALCOMPARES_H
+
+#include "llvm/CodeGen/MachinePassManager.h"
+
+namespace llvm {
+
+class MachineConditionalComparesPass
+ : public OptionalPassInfoMixin<MachineConditionalComparesPass> {
+public:
+ LLVM_ABI PreservedAnalyses run(MachineFunction &MF,
+ MachineFunctionAnalysisManager &MFAM);
+};
+
+} // namespace llvm
+
+#endif // LLVM_CODEGEN_MACHINECONDITIONALCOMPARES_H
diff --git a/llvm/include/llvm/CodeGen/Passes.h b/llvm/include/llvm/CodeGen/Passes.h
index 8c2b923aeb807..a912caebcd853 100644
--- a/llvm/include/llvm/CodeGen/Passes.h
+++ b/llvm/include/llvm/CodeGen/Passes.h
@@ -321,6 +321,10 @@ LLVM_ABI extern char &EarlyIfPredicatorID;
/// critical-path and resource depth.
LLVM_ABI extern char &MachineCombinerID;
+/// MachineConditionalCompares - This pass performs target-independent
+/// conditional-compare formation on SSA form, reducing branching.
+LLVM_ABI extern char &MachineConditionalComparesLegacyID;
+
/// StackSlotColoring - This pass performs stack coloring and merging.
/// It merges disjoint allocas to reduce the stack size.
LLVM_ABI extern char &StackColoringLegacyID;
diff --git a/llvm/include/llvm/CodeGen/TargetInstrInfo.h b/llvm/include/llvm/CodeGen/TargetInstrInfo.h
index 4749d06501cb2..65b5c1d2ee1ac 100644
--- a/llvm/include/llvm/CodeGen/TargetInstrInfo.h
+++ b/llvm/include/llvm/CodeGen/TargetInstrInfo.h
@@ -1050,6 +1050,76 @@ class LLVM_ABI TargetInstrInfo : public MCInstrInfo {
llvm_unreachable("Target didn't implement TargetInstrInfo::insertSelect!");
}
+ /// Carries the state that canConvertToCCMP() computes and convertToCCMP()
+ /// consumes when the target-independent conditional-compare formation pass
+ /// (see MachineConditionalCompares) folds a CFG triangle into a conditional
+ /// compare. All fields except CmpMI are target-owned scratch; the generic
+ /// pass never interprets them.
+ struct CCmpConvInfo {
+ /// The convertible compare instruction found in CmpBB.
+ MachineInstr *CmpMI = nullptr;
+ /// Opaque target scratch (parsed condition codes, chosen opcode, etc.);
+ /// sized to cover any target's needs.
+ int64_t TargetData[5] = {};
+ };
+
+ /// Return the physical flag/status register clobbered by conditional-compare
+ /// candidates (EFLAGS on X86, NZCV on AArch64), or a null register if the
+ /// target does not support conditional-compare formation. Used by the generic
+ /// MachineConditionalCompares pass to scan for flag reads/defs.
+ virtual MCRegister getConditionalCompareFlagReg() const {
+ return MCRegister();
+ }
+
+ /// Analyze whether the compare controlling CmpBB's terminator can be folded
+ /// into a conditional compare merged into Head (a triangle Head->CmpBB->Tail).
+ ///
+ /// HeadCond/CmpBBCond are analyzeBranch's opaque condition arrays. The
+ /// HeadTBBIsCmpBB and CmpBBTBBIsTail flags report whether each block's
+ /// analyzeBranch TBB is the wanted successor, so the target can invert its
+ /// condition codes as needed.
+ ///
+ /// On success, fills Info (including Info.CmpMI) and returns true.
+ virtual bool canConvertToCCMP(MachineBasicBlock &Head,
+ MachineBasicBlock &CmpBB,
+ ArrayRef<MachineOperand> HeadCond,
+ bool HeadTBBIsCmpBB,
+ ArrayRef<MachineOperand> CmpBBCond,
+ bool CmpBBTBBIsTail,
+ const MachineRegisterInfo &MRI,
+ CCmpConvInfo &Info) const {
+ return false;
+ }
+
+ /// Emit the conditional compare that replaces Info.CmpMI, using the target
+ /// data stashed by canConvertToCCMP(). This is invoked after the generic pass
+ /// has spliced CmpBB into Head and removed Head's branch, so the target only
+ /// performs the instruction rewrite (and any Head-terminator fixup, e.g.
+ /// synthesizing a compare for a cbz/cbnz head, or re-inserting a conditional
+ /// branch when the folded compare was itself a terminator).
+ ///
+ /// SpliceLoc is the insertion point for a synthesized Head compare (the first
+ /// instruction spliced in from CmpBB); HeadTermDL is the removed Head
+ /// terminator's debug location; HeadCond is its condition array. The generic
+ /// caller erases Info.CmpMI and fixes up Head's terminator afterwards. Returns
+ /// the new conditional-compare instruction.
+ virtual MachineInstr *convertToCCMP(MachineBasicBlock &Head,
+ MachineBasicBlock::iterator SpliceLoc,
+ const DebugLoc &HeadTermDL,
+ ArrayRef<MachineOperand> HeadCond,
+ const CCmpConvInfo &Info,
+ MachineRegisterInfo &MRI) const {
+ llvm_unreachable("Target didn't implement TargetInstrInfo::convertToCCMP!");
+ }
+
+ /// Return the expected code-size delta (in number of instructions) of
+ /// converting the candidate described by Info into a conditional compare.
+ /// Used by the MinSize heuristic. Defaults to 0 (no size change assumed).
+ virtual int getCCMPCodeSizeDelta(const CCmpConvInfo &Info,
+ ArrayRef<MachineOperand> HeadCond) const {
+ return 0;
+ }
+
/// Given an instruction marked as `isSelect = true`, attempt to optimize MI
/// by merging it with one of its operands. Returns nullptr on failure.
///
diff --git a/llvm/include/llvm/CodeGen/TargetSubtargetInfo.h b/llvm/include/llvm/CodeGen/TargetSubtargetInfo.h
index 0dc8b339e8ebb..cdeb8a5a80244 100644
--- a/llvm/include/llvm/CodeGen/TargetSubtargetInfo.h
+++ b/llvm/include/llvm/CodeGen/TargetSubtargetInfo.h
@@ -343,6 +343,21 @@ class LLVM_ABI TargetSubtargetInfo : public MCSubtargetInfo {
/// Enable the use of the early if conversion pass.
virtual bool enableEarlyIfConversion() const { return false; }
+ /// Per-target tuning for the conditional-compare formation pass
+ /// (MachineConditionalCompares). The default leaves the code-size path off
+ /// (X86 behavior); AArch64 overrides it.
+ struct CCmpConvHeuristics {
+ /// In a MinSize function, convert whenever it does not grow code.
+ bool UseCodeSizeDeltaOnMinSize = false;
+ };
+
+ /// Enable the target-independent conditional-compare formation pass
+ /// (MachineConditionalCompares) for this subtarget.
+ virtual bool enableCCMPFormation() const { return false; }
+
+ /// Return the tunable heuristics for conditional-compare formation.
+ virtual CCmpConvHeuristics getCCmpConvHeuristics() const { return {}; }
+
/// Return PBQPConstraint(s) for the target.
///
/// Override to provide custom PBQP constraints.
diff --git a/llvm/include/llvm/InitializePasses.h b/llvm/include/llvm/InitializePasses.h
index dec7a5ff0845d..09aa8caf5aaf6 100644
--- a/llvm/include/llvm/InitializePasses.h
+++ b/llvm/include/llvm/InitializePasses.h
@@ -199,6 +199,7 @@ initializeMachineBranchProbabilityInfoWrapperPassPass(PassRegistry &);
LLVM_ABI void initializeMachineCFGPrinterLegacyPass(PassRegistry &);
LLVM_ABI void initializeMachineCSELegacyPass(PassRegistry &);
LLVM_ABI void initializeMachineCombinerLegacyPass(PassRegistry &);
+LLVM_ABI void initializeMachineConditionalComparesLegacyPass(PassRegistry &);
LLVM_ABI void initializeMachineCopyPropagationLegacyPass(PassRegistry &);
LLVM_ABI void initializeMachineCycleInfoPrinterLegacyPass(PassRegistry &);
LLVM_ABI void initializeMachineCycleInfoWrapperPassPass(PassRegistry &);
diff --git a/llvm/include/llvm/Passes/MachinePassRegistry.def b/llvm/include/llvm/Passes/MachinePassRegistry.def
index 1bdfd72c9d736..88828667c036f 100644
--- a/llvm/include/llvm/Passes/MachinePassRegistry.def
+++ b/llvm/include/llvm/Passes/MachinePassRegistry.def
@@ -92,6 +92,7 @@ MACHINE_FUNCTION_PASS("legalizer", LegalizerPass())
MACHINE_FUNCTION_PASS("load-store-opt", LoadStoreOptPass())
MACHINE_FUNCTION_PASS("localizer", LocalizerPass())
MACHINE_FUNCTION_PASS("localstackalloc", LocalStackSlotAllocationPass())
+MACHINE_FUNCTION_PASS("machine-ccmp", MachineConditionalComparesPass())
MACHINE_FUNCTION_PASS("machine-combiner", MachineCombinerPass())
MACHINE_FUNCTION_PASS("machine-cp", MachineCopyPropagationPass())
MACHINE_FUNCTION_PASS("machine-cse", MachineCSEPass())
diff --git a/llvm/lib/CodeGen/CMakeLists.txt b/llvm/lib/CodeGen/CMakeLists.txt
index 99dfb4bb09df7..b39d9a1a8b71c 100644
--- a/llvm/lib/CodeGen/CMakeLists.txt
+++ b/llvm/lib/CodeGen/CMakeLists.txt
@@ -125,6 +125,7 @@ add_llvm_component_library(LLVMCodeGen
MachineBranchProbabilityInfo.cpp
MachineCFGPrinter.cpp
MachineCombiner.cpp
+ MachineConditionalCompares.cpp
MachineConvergenceVerifier.cpp
MachineCopyPropagation.cpp
MachineCSE.cpp
diff --git a/llvm/lib/CodeGen/CodeGen.cpp b/llvm/lib/CodeGen/CodeGen.cpp
index 7fb2fca6b6494..a7ebd7ae6e0ed 100644
--- a/llvm/lib/CodeGen/CodeGen.cpp
+++ b/llvm/lib/CodeGen/CodeGen.cpp
@@ -89,6 +89,7 @@ void llvm::initializeCodeGen(PassRegistry &Registry) {
initializeMachineCFGPrinterLegacyPass(Registry);
initializeMachineCSELegacyPass(Registry);
initializeMachineCombinerLegacyPass(Registry);
+ initializeMachineConditionalComparesLegacyPass(Registry);
initializeMachineDominanceFrontierWrapperPassPass(Registry);
initializeMachineCopyPropagationLegacyPass(Registry);
initializeMachineCycleInfoPrinterLegacyPass(Registry);
diff --git a/llvm/lib/Target/AArch64/AArch64ConditionalCompares.cpp b/llvm/lib/CodeGen/MachineConditionalCompares.cpp
similarity index 55%
rename from llvm/lib/Target/AArch64/AArch64ConditionalCompares.cpp
rename to llvm/lib/CodeGen/MachineConditionalCompares.cpp
index f490b5f94fe0d..31a258b011ed8 100644
--- a/llvm/lib/Target/AArch64/AArch64ConditionalCompares.cpp
+++ b/llvm/lib/CodeGen/MachineConditionalCompares.cpp
@@ -1,4 +1,4 @@
-//===-- AArch64ConditionalCompares.cpp --- CCMP formation for AArch64 -----===//
+//===-- MachineConditionalCompares.cpp --- CCMP formation -----------------===//
//
// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
// See https://llvm.org/LICENSE.txt for license information.
@@ -6,32 +6,41 @@
//
//===----------------------------------------------------------------------===//
//
-// This file implements the AArch64ConditionalCompares pass which reduces
-// branching and code size by using the conditional compare instructions CCMP,
-// CCMN, and FCMP.
+// This file implements the target-independent MachineConditionalCompares pass
+// which reduces branching by using the conditional-compare instructions
+// (CCMP/CTEST on X86; CCMP/CCMN/FCCMP on AArch64).
//
// The CFG transformations for forming conditional compares are very similar to
// if-conversion, and this pass should run immediately before the early
-// if-conversion pass.
+// if-conversion pass. The transform itself is target-independent; the
+// recognition of a convertible compare and the emission of the conditional
+// compare are delegated to the target via TargetInstrInfo hooks
+// (getConditionalCompareFlagReg / canConvertToCCMP / convertToCCMP /
+// getCCMPCodeSizeDelta) and the pass is gated and tuned per target via
+// TargetSubtargetInfo (enableCCMPFormation / getCCmpConvHeuristics).
//
//===----------------------------------------------------------------------===//
-#include "AArch64.h"
+#include "llvm/CodeGen/MachineConditionalCompares.h"
#include "llvm/ADT/DepthFirstIterator.h"
+#include "llvm/ADT/SmallVector.h"
#include "llvm/ADT/Statistic.h"
+#include "llvm/Analysis/OptimizationRemarkEmitter.h"
#include "llvm/CodeGen/MachineBranchProbabilityInfo.h"
#include "llvm/CodeGen/MachineDominators.h"
#include "llvm/CodeGen/MachineFunction.h"
#include "llvm/CodeGen/MachineFunctionPass.h"
#include "llvm/CodeGen/MachineInstrBuilder.h"
+#include "llvm/CodeGen/MachineInstrBundle.h"
#include "llvm/CodeGen/MachineLoopInfo.h"
-#include "llvm/CodeGen/MachinePassManager.h"
+#include "llvm/CodeGen/MachineOptimizationRemarkEmitter.h"
#include "llvm/CodeGen/MachineRegisterInfo.h"
#include "llvm/CodeGen/MachineTraceMetrics.h"
#include "llvm/CodeGen/Passes.h"
#include "llvm/CodeGen/TargetInstrInfo.h"
#include "llvm/CodeGen/TargetRegisterInfo.h"
#include "llvm/CodeGen/TargetSubtargetInfo.h"
+#include "llvm/IR/Function.h"
#include "llvm/InitializePasses.h"
#include "llvm/Support/CommandLine.h"
#include "llvm/Support/Debug.h"
@@ -39,17 +48,17 @@
using namespace llvm;
-#define DEBUG_TYPE "aarch64-ccmp"
+#define DEBUG_TYPE "machine-ccmp"
// Absolute maximum number of instructions allowed per speculated block.
// This bypasses all other heuristics, so it should be set fairly high.
static cl::opt<unsigned> BlockInstrLimit(
- "aarch64-ccmp-limit", cl::init(30), cl::Hidden,
+ "machine-ccmp-limit", cl::init(30), cl::Hidden,
cl::desc("Maximum number of instructions per speculated block."));
// Stress testing mode - disable heuristics.
-static cl::opt<bool> Stress("aarch64-stress-ccmp", cl::Hidden,
- cl::desc("Turn all knobs to 11"));
+static cl::opt<bool> Stress("stress-machine-ccmp", cl::Hidden,
+ cl::desc("Turn all knobs to 11"), cl::init(false));
STATISTIC(NumConsidered, "Number of ccmps considered");
STATISTIC(NumPhiRejs, "Number of ccmps rejected (PHI)");
@@ -57,16 +66,8 @@ STATISTIC(NumPhysRejs, "Number of ccmps rejected (Physregs)");
STATISTIC(NumPhi2Rejs, "Number of ccmps rejected (PHI2)");
STATISTIC(NumHeadBranchRejs, "Number of ccmps rejected (Head branch)");
STATISTIC(NumCmpBranchRejs, "Number of ccmps rejected (CmpBB branch)");
-STATISTIC(NumCmpTermRejs, "Number of ccmps rejected (CmpBB is cbz...)");
-STATISTIC(NumImmRangeRejs, "Number of ccmps rejected (Imm out of range)");
-STATISTIC(NumLiveDstRejs, "Number of ccmps rejected (Cmp dest live)");
-STATISTIC(NumMultNZCVUses, "Number of ccmps rejected (NZCV used)");
-STATISTIC(NumUnknNZCVDefs, "Number of ccmps rejected (NZCV def unknown)");
-
STATISTIC(NumSpeculateRejs, "Number of ccmps rejected (Can't speculate)");
-
STATISTIC(NumConverted, "Number of ccmp instructions created");
-STATISTIC(NumCompBranches, "Number of cbz/cbnz branches converted");
//===----------------------------------------------------------------------===//
// SSACCmpConv
@@ -132,14 +133,17 @@ STATISTIC(NumCompBranches, "Number of cbz/cbnz branches converted");
// between Head and Tail, just like if-converting a diamond.
//
// FIXME: Handle PHIs in Tail by turning them into selects (if-conversion).
-
+//
namespace {
class SSACCmpConv {
- MachineFunction *MF;
const TargetInstrInfo *TII;
const TargetRegisterInfo *TRI;
MachineRegisterInfo *MRI;
const MachineBranchProbabilityInfo *MBPI;
+ MachineOptimizationRemarkEmitter *ORE;
+
+ /// The physical flag/status register clobbered by ccmp candidates.
+ MCRegister FlagReg;
public:
/// The first block containing a conditional branch, dominating everything
@@ -156,17 +160,16 @@ class SSACCmpConv {
MachineInstr *CmpMI;
private:
- /// The branch condition in Head as determined by analyzeBranch.
+ /// The branch condition in Head as determined by analyzeBranch. Stored
+ /// opaquely; only round-tripped through the target hooks.
SmallVector<MachineOperand, 4> HeadCond;
- /// The condition code that makes Head branch to CmpBB.
- AArch64CC::CondCode HeadCmpBBCC;
-
- /// The branch condition in CmpBB.
+ /// The branch condition in CmpBB as determined by analyzeBranch. Stored
+ /// opaquely; only round-tripped through the target hooks.
SmallVector<MachineOperand, 4> CmpBBCond;
- /// The condition code that makes CmpBB branch to Tail.
- AArch64CC::CondCode CmpBBTailCC;
+ /// Target-owned recognition/emission state produced by canConvertToCCMP().
+ TargetInstrInfo::CCmpConvInfo Info;
/// Check if the Tail PHIs are trivially convertible.
bool trivialTailPHIs();
@@ -174,42 +177,40 @@ class SSACCmpConv {
/// Remove CmpBB from the Tail PHIs.
void updateTailPHIs();
- /// Check if an operand defining DstReg is dead.
- bool isDeadDef(unsigned DstReg);
-
- /// Find the compare instruction in MBB that controls the conditional branch.
- /// Return NULL if a convertible instruction can't be found.
- MachineInstr *findConvertibleCompare(MachineBasicBlock *MBB);
-
/// Return true if all non-terminator instructions in MBB can be safely
/// speculated.
bool canSpeculateInstrs(MachineBasicBlock *MBB, const MachineInstr *CmpMI);
public:
- /// runOnMachineFunction - Initialize per-function data structures.
- void runOnMachineFunction(MachineFunction &MF,
- const MachineBranchProbabilityInfo *MBPI) {
- this->MF = &MF;
+ /// Initialize per-function data structures.
+ void init(MachineFunction &MF, const MachineBranchProbabilityInfo *MBPI,
+ MachineOptimizationRemarkEmitter *ORE) {
this->MBPI = MBPI;
+ this->ORE = ORE;
TII = MF.getSubtarget().getInstrInfo();
TRI = MF.getSubtarget().getRegisterInfo();
MRI = &MF.getRegInfo();
+ FlagReg = TII->getConditionalCompareFlagReg();
}
- /// If the sub-CFG headed by MBB can be cmp-converted, initialize the
- /// internal state, and return true.
+ /// If the sub-CFG headed by MBB can be cmp-converted, initialize the internal
+ /// state, and return true.
bool canConvert(MachineBasicBlock *MBB);
- /// Cmo-convert the last block passed to canConvertCmp(), assuming
- /// it is possible. Add any erased blocks to RemovedBlocks.
+ /// Cmp-convert the last block passed to canConvert(), assuming it is
+ /// possible. Add any erased blocks to RemovedBlocks.
void convert(SmallVectorImpl<MachineBasicBlock *> &RemovedBlocks);
- /// Return the expected code size delta if the conversion into a
- /// conditional compare is performed.
- int expectedCodeSizeDelta() const;
+ /// Return the expected code size delta if the conversion into a conditional
+ /// compare is performed.
+ int expectedCodeSizeDelta() const {
+ return TII->getCCMPCodeSizeDelta(Info, HeadCond);
+ }
};
} // end anonymous namespace
+// Detect a chain of vreg-to-vreg copies feeding Reg, returning the original
+// value. This lets trivialTailPHIs() see copy-equivalent PHI operands as equal.
static Register lookThroughCopies(Register Reg, MachineRegisterInfo *MRI) {
MachineInstr *MI;
while ((MI = MRI->getUniqueVRegDef(Reg)) &&
@@ -224,14 +225,12 @@ static Register lookThroughCopies(Register Reg, MachineRegisterInfo *MRI) {
// Check that all PHIs in Tail are selecting the same value from Head and CmpBB.
// This means that no if-conversion is required when merging CmpBB into Head.
bool SSACCmpConv::trivialTailPHIs() {
- for (auto &I : *Tail) {
- if (!I.isPHI())
- break;
+ for (auto &I : Tail->phis()) {
unsigned HeadReg = 0, CmpBBReg = 0;
// PHI operands come in (VReg, MBB) pairs.
- for (unsigned oi = 1, oe = I.getNumOperands(); oi != oe; oi += 2) {
- MachineBasicBlock *MBB = I.getOperand(oi + 1).getMBB();
- Register Reg = lookThroughCopies(I.getOperand(oi).getReg(), MRI);
+ for (unsigned Idx = 1, End = I.getNumOperands(); Idx != End; Idx += 2) {
+ MachineBasicBlock *MBB = I.getOperand(Idx + 1).getMBB();
+ Register Reg = lookThroughCopies(I.getOperand(Idx).getReg(), MRI);
if (MBB == Head) {
assert((!HeadReg || HeadReg == Reg) && "Inconsistent PHI operands");
HeadReg = Reg;
@@ -250,149 +249,23 @@ bool SSACCmpConv::trivialTailPHIs() {
// Assuming that trivialTailPHIs() is true, update the Tail PHIs by simply
// removing the CmpBB operands. The Head operands will be identical.
void SSACCmpConv::updateTailPHIs() {
- for (auto &I : *Tail) {
- if (!I.isPHI())
- break;
+ for (auto &I : Tail->phis()) {
// I is a PHI. It can have multiple entries for CmpBB.
- for (unsigned oi = I.getNumOperands(); oi > 2; oi -= 2) {
- // PHI operands are (Reg, MBB) at (oi-2, oi-1).
- if (I.getOperand(oi - 1).getMBB() == CmpBB) {
- I.removeOperand(oi - 1);
- I.removeOperand(oi - 2);
+ for (unsigned Idx = I.getNumOperands(); Idx > 2; Idx -= 2) {
+ // PHI operands are (Reg, MBB) at (Idx-2, Idx-1).
+ if (I.getOperand(Idx - 1).getMBB() == CmpBB) {
+ I.removeOperand(Idx - 1);
+ I.removeOperand(Idx - 2);
}
}
}
}
-// This pass runs before the AArch64DeadRegisterDefinitions pass, so compares
-// are still writing virtual registers without any uses.
-bool SSACCmpConv::isDeadDef(unsigned DstReg) {
- // Writes to the zero register are dead.
- if (DstReg == AArch64::WZR || DstReg == AArch64::XZR)
- return true;
- if (!Register::isVirtualRegister(DstReg))
- return false;
- // A virtual register def without any uses will be marked dead later, and
- // eventually replaced by the zero register.
- return MRI->use_nodbg_empty(DstReg);
-}
-
-// Parse a condition code returned by analyzeBranch, and compute the CondCode
-// corresponding to TBB.
-// Return
-static bool parseCond(ArrayRef<MachineOperand> Cond, AArch64CC::CondCode &CC) {
- // A normal br.cond simply has the condition code.
- if (Cond[0].getImm() != -1) {
- assert(Cond.size() == 1 && "Unknown Cond array format");
- CC = (AArch64CC::CondCode)(int)Cond[0].getImm();
- return true;
- }
- // For tbz and cbz instruction, the opcode is next.
- switch (Cond[1].getImm()) {
- default:
- // This includes tbz / tbnz branches which can't be converted to
- // ccmp + br.cond.
- return false;
- case AArch64::CBZW:
- case AArch64::CBZX:
- assert(Cond.size() == 3 && "Unknown Cond array format");
- CC = AArch64CC::EQ;
- return true;
- case AArch64::CBNZW:
- case AArch64::CBNZX:
- assert(Cond.size() == 3 && "Unknown Cond array format");
- CC = AArch64CC::NE;
- return true;
- }
-}
-
-MachineInstr *SSACCmpConv::findConvertibleCompare(MachineBasicBlock *MBB) {
- MachineBasicBlock::iterator I = MBB->getFirstTerminator();
- if (I == MBB->end())
- return nullptr;
- // The terminator must be controlled by the flags.
- if (!I->readsRegister(AArch64::NZCV, /*TRI=*/nullptr)) {
- switch (I->getOpcode()) {
- case AArch64::CBZW:
- case AArch64::CBZX:
- case AArch64::CBNZW:
- case AArch64::CBNZX:
- // These can be converted into a ccmp against #0.
- return &*I;
- }
- ++NumCmpTermRejs;
- LLVM_DEBUG(dbgs() << "Flags not used by terminator: " << *I);
- return nullptr;
- }
-
- // Now find the instruction controlling the terminator.
- for (MachineBasicBlock::iterator B = MBB->begin(); I != B;) {
- I = prev_nodbg(I, MBB->begin());
- assert(!I->isTerminator() && "Spurious terminator");
- switch (I->getOpcode()) {
- // cmp is an alias for subs with a dead destination register.
- case AArch64::SUBSWri:
- case AArch64::SUBSXri:
- // cmn is an alias for adds with a dead destination register.
- case AArch64::ADDSWri:
- case AArch64::ADDSXri:
- // Check that the immediate operand is within range, ccmp wants a uimm5.
- // Rd = SUBSri Rn, imm, shift
- if (I->getOperand(3).getImm() || !isUInt<5>(I->getOperand(2).getImm())) {
- LLVM_DEBUG(dbgs() << "Immediate out of range for ccmp: " << *I);
- ++NumImmRangeRejs;
- return nullptr;
- }
- [[fallthrough]];
- case AArch64::SUBSWrr:
- case AArch64::SUBSXrr:
- case AArch64::ADDSWrr:
- case AArch64::ADDSXrr:
- if (isDeadDef(I->getOperand(0).getReg()))
- return &*I;
- LLVM_DEBUG(dbgs() << "Can't convert compare with live destination: "
- << *I);
- ++NumLiveDstRejs;
- return nullptr;
- case AArch64::FCMPSrr:
- case AArch64::FCMPDrr:
- case AArch64::FCMPESrr:
- case AArch64::FCMPEDrr:
- return &*I;
- }
-
- // Check for flag reads and clobbers.
- PhysRegInfo PRI = AnalyzePhysRegInBundle(*I, AArch64::NZCV, TRI);
-
- if (PRI.Read) {
- // The ccmp doesn't produce exactly the same flags as the original
- // compare, so reject the transform if there are uses of the flags
- // besides the terminators.
- LLVM_DEBUG(dbgs() << "Can't create ccmp with multiple uses: " << *I);
- ++NumMultNZCVUses;
- return nullptr;
- }
-
- if (PRI.Defined || PRI.Clobbered) {
- LLVM_DEBUG(dbgs() << "Not convertible compare: " << *I);
- ++NumUnknNZCVDefs;
- return nullptr;
- }
- }
- LLVM_DEBUG(dbgs() << "Flags not defined in " << printMBBReference(*MBB)
- << '\n');
- return nullptr;
-}
-
-/// Determine if all the instructions in MBB can safely
-/// be speculated. The terminators are not considered.
-///
-/// Only CmpMI is allowed to clobber the flags.
-///
+/// Determine if all the instructions in MBB can safely be speculated. The
+/// terminators are not considered. Only CmpMI is allowed to clobber the flags.
bool SSACCmpConv::canSpeculateInstrs(MachineBasicBlock *MBB,
const MachineInstr *CmpMI) {
- // Reject any live-in physregs. It's probably NZCV/EFLAGS, and very hard to
- // get right.
+ // Reject any live-in physregs. It's very hard to get right.
if (!MBB->livein_empty()) {
LLVM_DEBUG(dbgs() << printMBBReference(*MBB) << " has live-ins.\n");
return false;
@@ -434,7 +307,7 @@ bool SSACCmpConv::canSpeculateInstrs(MachineBasicBlock *MBB,
}
// Only CmpMI is allowed to clobber the flags.
- if (&I != CmpMI && I.modifiesRegister(AArch64::NZCV, TRI)) {
+ if (&I != CmpMI && I.modifiesRegister(FlagReg, TRI)) {
LLVM_DEBUG(dbgs() << "Clobbers flags: " << I);
return false;
}
@@ -444,10 +317,11 @@ bool SSACCmpConv::canSpeculateInstrs(MachineBasicBlock *MBB,
/// Analyze the sub-cfg rooted in MBB, and return true if it is a potential
/// candidate for cmp-conversion. Fill out the internal state.
-///
bool SSACCmpConv::canConvert(MachineBasicBlock *MBB) {
Head = MBB;
Tail = CmpBB = nullptr;
+ CmpMI = nullptr;
+ Info = TargetInstrInfo::CCmpConvInfo();
if (Head->succ_size() != 2)
return false;
@@ -508,8 +382,8 @@ bool SSACCmpConv::canConvert(MachineBasicBlock *MBB) {
// The branch we're looking to eliminate must be analyzable.
HeadCond.clear();
- MachineBasicBlock *TBB = nullptr, *FBB = nullptr;
- if (TII->analyzeBranch(*Head, TBB, FBB, HeadCond)) {
+ MachineBasicBlock *HeadTBB = nullptr, *HeadFBB = nullptr;
+ if (TII->analyzeBranch(*Head, HeadTBB, HeadFBB, HeadCond)) {
LLVM_DEBUG(dbgs() << "Head branch not analyzable.\n");
++NumHeadBranchRejs;
return false;
@@ -517,62 +391,44 @@ bool SSACCmpConv::canConvert(MachineBasicBlock *MBB) {
// This is weird, probably some sort of degenerate CFG, or an edge to a
// landing pad.
- if (!TBB || HeadCond.empty()) {
+ if (!HeadTBB || HeadCond.empty()) {
LLVM_DEBUG(
dbgs() << "analyzeBranch didn't find conditional branch in Head.\n");
++NumHeadBranchRejs;
return false;
}
- if (!parseCond(HeadCond, HeadCmpBBCC)) {
- LLVM_DEBUG(dbgs() << "Unsupported branch type on Head\n");
- ++NumHeadBranchRejs;
- return false;
- }
-
- // Make sure the branch direction is right.
- if (TBB != CmpBB) {
- assert(TBB == Tail && "Unexpected TBB");
- HeadCmpBBCC = AArch64CC::getInvertedCondCode(HeadCmpBBCC);
- }
-
+ // Analyze the branch in CmpBB.
CmpBBCond.clear();
- TBB = FBB = nullptr;
- if (TII->analyzeBranch(*CmpBB, TBB, FBB, CmpBBCond)) {
+ MachineBasicBlock *CmpBBTBB = nullptr, *CmpBBFBB = nullptr;
+ if (TII->analyzeBranch(*CmpBB, CmpBBTBB, CmpBBFBB, CmpBBCond)) {
LLVM_DEBUG(dbgs() << "CmpBB branch not analyzable.\n");
++NumCmpBranchRejs;
return false;
}
- if (!TBB || CmpBBCond.empty()) {
+ if (!CmpBBTBB || CmpBBCond.empty()) {
LLVM_DEBUG(
dbgs() << "analyzeBranch didn't find conditional branch in CmpBB.\n");
++NumCmpBranchRejs;
return false;
}
- if (!parseCond(CmpBBCond, CmpBBTailCC)) {
- LLVM_DEBUG(dbgs() << "Unsupported branch type on CmpBB\n");
- ++NumCmpBranchRejs;
- return false;
- }
-
- if (TBB != Tail)
- CmpBBTailCC = AArch64CC::getInvertedCondCode(CmpBBTailCC);
-
- LLVM_DEBUG(dbgs() << "Head->CmpBB on "
- << AArch64CC::getCondCodeName(HeadCmpBBCC)
- << ", CmpBB->Tail on "
- << AArch64CC::getCondCodeName(CmpBBTailCC) << '\n');
-
- CmpMI = findConvertibleCompare(CmpBB);
- if (!CmpMI)
+ // Delegate target-specific recognition: condition-code parsing/reject-lists,
+ // any target-specific terminator constraints, and finding the convertible
+ // compare. The condition arrays are passed opaquely; the booleans tell the
+ // target whether analyzeBranch's TBB is the desired successor so it can apply
+ // its own condition-code inversion.
+ if (!TII->canConvertToCCMP(*Head, *CmpBB, HeadCond, HeadTBB == CmpBB,
+ CmpBBCond, CmpBBTBB == Tail, *MRI, Info))
return false;
+ CmpMI = Info.CmpMI;
if (!canSpeculateInstrs(CmpBB, CmpMI)) {
++NumSpeculateRejs;
return false;
}
+
return true;
}
@@ -619,104 +475,30 @@ void SSACCmpConv::convert(SmallVectorImpl<MachineBasicBlock *> &RemovedBlocks) {
}
Head->transferSuccessorsAndUpdatePHIs(CmpBB);
- DebugLoc TermDL = Head->getFirstTerminator()->getDebugLoc();
+ DebugLoc HeadTermDL = Head->getFirstTerminator()->getDebugLoc();
TII->removeBranch(*Head);
- // If the Head terminator was one of the cbz / tbz branches with built-in
- // compare, we need to insert an explicit compare instruction in its place.
- if (HeadCond[0].getImm() == -1) {
- ++NumCompBranches;
- unsigned Opc = 0;
- switch (HeadCond[1].getImm()) {
- case AArch64::CBZW:
- case AArch64::CBNZW:
- Opc = AArch64::SUBSWri;
- break;
- case AArch64::CBZX:
- case AArch64::CBNZX:
- Opc = AArch64::SUBSXri;
- break;
- default:
- llvm_unreachable("Cannot convert Head branch");
- }
- const MCInstrDesc &MCID = TII->get(Opc);
- // Create a dummy virtual register for the SUBS def.
- Register DestReg = MRI->createVirtualRegister(TII->getRegClass(MCID, 0));
- // Insert a SUBS Rn, #0 instruction instead of the cbz / cbnz.
- BuildMI(*Head, Head->end(), TermDL, MCID)
- .addReg(DestReg, RegState::Define | RegState::Dead)
- .add(HeadCond[2])
- .addImm(0)
- .addImm(0);
- // SUBS uses the GPR*sp register classes.
- MRI->constrainRegClass(HeadCond[2].getReg(), TII->getRegClass(MCID, 1));
- }
+ // Remember the splice boundary: the first CmpBB instruction, which after the
+ // splice below becomes the insertion point in Head for any synthesized Head
+ // compare (e.g. a cbz/cbnz head).
+ assert(!CmpBB->empty() && "CmpBB unexpectedly empty");
+ MachineInstr *SpliceStart = &CmpBB->front();
Head->splice(Head->end(), CmpBB, CmpBB->begin(), CmpBB->end());
- // Now replace CmpMI with a ccmp instruction that also considers the incoming
- // flags.
- unsigned Opc = 0;
- unsigned FirstOp = 1; // First CmpMI operand to copy.
- bool isZBranch = false; // CmpMI is a cbz/cbnz instruction.
- switch (CmpMI->getOpcode()) {
- default:
- llvm_unreachable("Unknown compare opcode");
- case AArch64::SUBSWri: Opc = AArch64::CCMPWi; break;
- case AArch64::SUBSWrr: Opc = AArch64::CCMPWr; break;
- case AArch64::SUBSXri: Opc = AArch64::CCMPXi; break;
- case AArch64::SUBSXrr: Opc = AArch64::CCMPXr; break;
- case AArch64::ADDSWri: Opc = AArch64::CCMNWi; break;
- case AArch64::ADDSWrr: Opc = AArch64::CCMNWr; break;
- case AArch64::ADDSXri: Opc = AArch64::CCMNXi; break;
- case AArch64::ADDSXrr: Opc = AArch64::CCMNXr; break;
- case AArch64::FCMPSrr: Opc = AArch64::FCCMPSrr; FirstOp = 0; break;
- case AArch64::FCMPDrr: Opc = AArch64::FCCMPDrr; FirstOp = 0; break;
- case AArch64::FCMPESrr: Opc = AArch64::FCCMPESrr; FirstOp = 0; break;
- case AArch64::FCMPEDrr: Opc = AArch64::FCCMPEDrr; FirstOp = 0; break;
- case AArch64::CBZW:
- case AArch64::CBNZW:
- Opc = AArch64::CCMPWi;
- FirstOp = 0;
- isZBranch = true;
- break;
- case AArch64::CBZX:
- case AArch64::CBNZX:
- Opc = AArch64::CCMPXi;
- FirstOp = 0;
- isZBranch = true;
- break;
- }
+ // Let the target emit the conditional compare that replaces CmpMI (and any
+ // Head-terminator fixup). It builds before CmpMI, which now lives in Head.
+ TII->convertToCCMP(*Head, SpliceStart->getIterator(), HeadTermDL, HeadCond,
+ Info, *MRI);
+
+ if (ORE)
+ ORE->emit([&]() {
+ MachineOptimizationRemark R(DEBUG_TYPE, "ConvertedCMP",
+ CmpMI->getDebugLoc(), CmpBB);
+ R << "convert CMP into conditional CMP";
+ return R;
+ });
- // The ccmp instruction should set the flags according to the comparison when
- // Head would have branched to CmpBB.
- // The NZCV immediate operand should provide flags for the case where Head
- // would have branched to Tail. These flags should cause the new Head
- // terminator to branch to tail.
- unsigned NZCV = AArch64CC::getNZCVToSatisfyCondCode(CmpBBTailCC);
- const MCInstrDesc &MCID = TII->get(Opc);
- MRI->constrainRegClass(CmpMI->getOperand(FirstOp).getReg(),
- TII->getRegClass(MCID, 0));
- if (CmpMI->getOperand(FirstOp + 1).isReg())
- MRI->constrainRegClass(CmpMI->getOperand(FirstOp + 1).getReg(),
- TII->getRegClass(MCID, 1));
- MachineInstrBuilder MIB = BuildMI(*Head, CmpMI, CmpMI->getDebugLoc(), MCID)
- .add(CmpMI->getOperand(FirstOp)); // Register Rn
- if (isZBranch)
- MIB.addImm(0); // cbz/cbnz Rn -> ccmp Rn, #0
- else
- MIB.add(CmpMI->getOperand(FirstOp + 1)); // Register Rm / Immediate
- MIB.addImm(NZCV).addImm(HeadCmpBBCC);
-
- // If CmpMI was a terminator, we need a new conditional branch to replace it.
- // This now becomes a Head terminator.
- if (isZBranch) {
- bool isNZ = CmpMI->getOpcode() == AArch64::CBNZW ||
- CmpMI->getOpcode() == AArch64::CBNZX;
- BuildMI(*Head, CmpMI, CmpMI->getDebugLoc(), TII->get(AArch64::Bcc))
- .addImm(isNZ ? AArch64CC::NE : AArch64CC::EQ)
- .add(CmpMI->getOperand(1)); // Branch target.
- }
CmpMI->eraseFromParent();
Head->updateTerminator(CmpBB->getNextNode());
@@ -725,120 +507,43 @@ void SSACCmpConv::convert(SmallVectorImpl<MachineBasicBlock *> &RemovedBlocks) {
++NumConverted;
}
-int SSACCmpConv::expectedCodeSizeDelta() const {
- int delta = 0;
- // If the Head terminator was one of the cbz / tbz branches with built-in
- // compare, we need to insert an explicit compare instruction in its place
- // plus a branch instruction.
- if (HeadCond[0].getImm() == -1) {
- switch (HeadCond[1].getImm()) {
- case AArch64::CBZW:
- case AArch64::CBNZW:
- case AArch64::CBZX:
- case AArch64::CBNZX:
- // Therefore delta += 1
- delta = 1;
- break;
- default:
- llvm_unreachable("Cannot convert Head branch");
- }
- }
- // If the Cmp terminator was one of the cbz / tbz branches with
- // built-in compare, it will be turned into a compare instruction
- // into Head, but we do not save any instruction.
- // Otherwise, we save the branch instruction.
- switch (CmpMI->getOpcode()) {
- default:
- --delta;
- break;
- case AArch64::CBZW:
- case AArch64::CBNZW:
- case AArch64::CBZX:
- case AArch64::CBNZX:
- break;
- }
- return delta;
-}
-
//===----------------------------------------------------------------------===//
-// AArch64ConditionalCompares Pass
+// MachineConditionalCompares Pass
//===----------------------------------------------------------------------===//
namespace {
-class AArch64ConditionalComparesImpl {
- const MachineBranchProbabilityInfo *MBPI;
- const TargetInstrInfo *TII;
- const TargetRegisterInfo *TRI;
- const TargetSubtargetInfo *STI;
- // Does the proceeded function has Oz attribute.
- bool MinSize;
- MachineRegisterInfo *MRI;
- MachineDominatorTree *DomTree;
- MachineLoopInfo *Loops;
- MachineTraceMetrics *Traces;
- MachineTraceMetrics::Ensemble *MinInstr;
+class MachineConditionalCompares {
+ const TargetSubtargetInfo *STI = nullptr;
+ const MachineBranchProbabilityInfo *MBPI = nullptr;
+ MachineDominatorTree *DomTree = nullptr;
+ MachineLoopInfo *Loops = nullptr;
+ MachineTraceMetrics *Traces = nullptr;
+ MachineTraceMetrics::Ensemble *MinInstr = nullptr;
+ MachineOptimizationRemarkEmitter *ORE = nullptr;
+ TargetSubtargetInfo::CCmpConvHeuristics Heur;
+ bool MinSize = false;
SSACCmpConv CmpConv;
public:
- AArch64ConditionalComparesImpl(const MachineBranchProbabilityInfo *MBPI,
- MachineDominatorTree *DomTree,
- MachineLoopInfo *Loops,
- MachineTraceMetrics *Traces)
- : MBPI(MBPI), DomTree(DomTree), Loops(Loops), Traces(Traces) {}
+ MachineConditionalCompares(const MachineBranchProbabilityInfo *MBPI,
+ MachineDominatorTree *DomTree,
+ MachineLoopInfo *Loops, MachineTraceMetrics *Traces,
+ MachineOptimizationRemarkEmitter *ORE)
+ : MBPI(MBPI), DomTree(DomTree), Loops(Loops), Traces(Traces), ORE(ORE) {}
bool run(MachineFunction &MF);
private:
- bool tryConvert(MachineBasicBlock *);
+ bool tryConvert(MachineBasicBlock *MBB);
void updateDomTree(ArrayRef<MachineBasicBlock *> Removed);
void updateLoops(ArrayRef<MachineBasicBlock *> Removed);
void invalidateTraces();
bool shouldConvert();
};
-
-class AArch64ConditionalComparesLegacy : public MachineFunctionPass {
-public:
- static char ID;
- AArch64ConditionalComparesLegacy() : MachineFunctionPass(ID) {
- initializeAArch64ConditionalComparesLegacyPass(
- *PassRegistry::getPassRegistry());
- }
- void getAnalysisUsage(AnalysisUsage &AU) const override;
- bool runOnMachineFunction(MachineFunction &MF) override;
- StringRef getPassName() const override {
- return "AArch64 Conditional Compares";
- }
-};
} // end anonymous namespace
-char AArch64ConditionalComparesLegacy::ID = 0;
-
-INITIALIZE_PASS_BEGIN(AArch64ConditionalComparesLegacy, "aarch64-ccmp",
- "AArch64 CCMP Pass", false, false)
-INITIALIZE_PASS_DEPENDENCY(MachineBranchProbabilityInfoWrapperPass)
-INITIALIZE_PASS_DEPENDENCY(MachineDominatorTreeWrapperPass)
-INITIALIZE_PASS_DEPENDENCY(MachineTraceMetricsWrapperPass)
-INITIALIZE_PASS_END(AArch64ConditionalComparesLegacy, "aarch64-ccmp",
- "AArch64 CCMP Pass", false, false)
-
-FunctionPass *llvm::createAArch64ConditionalCompares() {
- return new AArch64ConditionalComparesLegacy();
-}
-
-void AArch64ConditionalComparesLegacy::getAnalysisUsage(
- AnalysisUsage &AU) const {
- AU.addRequired<MachineBranchProbabilityInfoWrapperPass>();
- AU.addRequired<MachineDominatorTreeWrapperPass>();
- AU.addPreserved<MachineDominatorTreeWrapperPass>();
- AU.addRequired<MachineLoopInfoWrapperPass>();
- AU.addPreserved<MachineLoopInfoWrapperPass>();
- AU.addRequired<MachineTraceMetricsWrapperPass>();
- AU.addPreserved<MachineTraceMetricsWrapperPass>();
- MachineFunctionPass::getAnalysisUsage(AU);
-}
-
/// Update the dominator tree after if-conversion erased some blocks.
-void AArch64ConditionalComparesImpl::updateDomTree(
+void MachineConditionalCompares::updateDomTree(
ArrayRef<MachineBasicBlock *> Removed) {
// convert() removes CmpBB which was previously dominated by Head.
// CmpBB children should be transferred to Head.
@@ -854,7 +559,7 @@ void AArch64ConditionalComparesImpl::updateDomTree(
}
/// Update LoopInfo after if-conversion.
-void AArch64ConditionalComparesImpl::updateLoops(
+void MachineConditionalCompares::updateLoops(
ArrayRef<MachineBasicBlock *> Removed) {
if (!Loops)
return;
@@ -863,76 +568,72 @@ void AArch64ConditionalComparesImpl::updateLoops(
}
/// Invalidate MachineTraceMetrics before if-conversion.
-void AArch64ConditionalComparesImpl::invalidateTraces() {
+void MachineConditionalCompares::invalidateTraces() {
Traces->invalidate(CmpConv.Head);
Traces->invalidate(CmpConv.CmpBB);
}
-/// Apply cost model and heuristics to the if-conversion in IfConv.
-/// Return true if the conversion is a good idea.
-///
-bool AArch64ConditionalComparesImpl::shouldConvert() {
+/// Apply the cost model to the candidate conversion in CmpConv. Return true if
+/// the conversion is a good idea.
+bool MachineConditionalCompares::shouldConvert() {
// Stress testing mode disables all cost considerations.
if (Stress)
return true;
+
if (!MinInstr)
MinInstr = Traces->getEnsemble(MachineTraceStrategy::TS_MinInstrCount);
// Head dominates CmpBB, so it is always included in its trace.
MachineTraceMetrics::Trace Trace = MinInstr->getTrace(CmpConv.CmpBB);
- // If code size is the main concern
- if (MinSize) {
+ // If code size is the main concern.
+ if (Heur.UseCodeSizeDeltaOnMinSize && MinSize) {
int CodeSizeDelta = CmpConv.expectedCodeSizeDelta();
LLVM_DEBUG(dbgs() << "Code size delta: " << CodeSizeDelta << '\n');
- // If we are minimizing the code size, do the conversion whatever
- // the cost is.
+ // If we are minimizing the code size, do the conversion whatever the cost
+ // is.
if (CodeSizeDelta < 0)
return true;
if (CodeSizeDelta > 0) {
LLVM_DEBUG(dbgs() << "Code size is increasing, give up on this one.\n");
return false;
}
- // CodeSizeDelta == 0, continue with the regular heuristics
+ // CodeSizeDelta == 0, continue with the regular heuristics.
}
- // Heuristic: The compare conversion delays the execution of the branch
- // instruction because we must wait for the inputs to the second compare as
- // well. The branch has no dependent instructions, but delaying it increases
- // the cost of a misprediction.
- //
- // Set a limit on the delay we will accept.
- unsigned DelayLimit = STI->getMispredictionPenalty() * 3 / 4;
-
// Instruction depths can be computed for all trace instructions above CmpBB.
unsigned HeadDepth =
Trace.getInstrCycles(*CmpConv.Head->getFirstTerminator()).Depth;
+
+ // The conversion delays the branch because it must also wait for the inputs
+ // to the second compare. The branch has no dependent instructions, but
+ // delaying it increases the cost of a misprediction, so cap the delay at 3/4
+ // of the misprediction penalty.
unsigned CmpBBDepth =
Trace.getInstrCycles(*CmpConv.CmpBB->getFirstTerminator()).Depth;
+ unsigned DelayLimit = STI->getMispredictionPenalty() * 3 / 4;
LLVM_DEBUG(dbgs() << "Head depth: " << HeadDepth
- << "\nCmpBB depth: " << CmpBBDepth << '\n');
+ << "\nCmpBB depth: " << CmpBBDepth
+ << "\nDelay limit: " << DelayLimit << '\n');
if (CmpBBDepth > HeadDepth + DelayLimit) {
LLVM_DEBUG(dbgs() << "Branch delay would be larger than " << DelayLimit
<< " cycles.\n");
return false;
}
- // Check the resource depth at the bottom of CmpBB - these instructions will
- // be speculated.
+ // The speculated instructions at the bottom of CmpBB must fit under the Head
+ // critical path, i.e. their resource depth must not exceed it.
unsigned ResDepth = Trace.getResourceDepth(true);
LLVM_DEBUG(dbgs() << "Resources: " << ResDepth << '\n');
-
- // Heuristic: The speculatively executed instructions must all be able to
- // merge into the Head block. The Head critical path should dominate the
- // resource cost of the speculated instructions.
if (ResDepth > HeadDepth) {
LLVM_DEBUG(dbgs() << "Too many instructions to speculate.\n");
return false;
}
+
return true;
}
-bool AArch64ConditionalComparesImpl::tryConvert(MachineBasicBlock *MBB) {
+bool MachineConditionalCompares::tryConvert(MachineBasicBlock *MBB) {
bool Changed = false;
while (CmpConv.canConvert(MBB) && shouldConvert()) {
invalidateTraces();
@@ -941,25 +642,23 @@ bool AArch64ConditionalComparesImpl::tryConvert(MachineBasicBlock *MBB) {
Changed = true;
updateDomTree(RemovedBlocks);
updateLoops(RemovedBlocks);
- for (MachineBasicBlock *MBB : RemovedBlocks)
- MBB->eraseFromParent();
+ for (MachineBasicBlock *RemovedMBB : RemovedBlocks)
+ RemovedMBB->eraseFromParent();
}
return Changed;
}
-bool AArch64ConditionalComparesImpl::run(MachineFunction &MF) {
- LLVM_DEBUG(dbgs() << "********** AArch64 Conditional Compares **********\n"
+bool MachineConditionalCompares::run(MachineFunction &MF) {
+ LLVM_DEBUG(dbgs() << "********** Machine Conditional Compares **********\n"
<< "********** Function: " << MF.getName() << '\n');
- TII = MF.getSubtarget().getInstrInfo();
- TRI = MF.getSubtarget().getRegisterInfo();
STI = &MF.getSubtarget();
- MRI = &MF.getRegInfo();
MinInstr = nullptr;
MinSize = MF.getFunction().hasMinSize();
+ Heur = STI->getCCmpConvHeuristics();
bool Changed = false;
- CmpConv.runOnMachineFunction(MF, MBPI);
+ CmpConv.init(MF, MBPI, ORE);
// Visit blocks in dominator tree pre-order. The pre-order enables multiple
// cmp-conversions from the same head block.
@@ -970,13 +669,73 @@ bool AArch64ConditionalComparesImpl::run(MachineFunction &MF) {
if (tryConvert(I->getBlock()))
Changed = true;
+ if (Changed && ORE) {
+ ORE->emit([&]() {
+ MachineOptimizationRemarkAnalysis R(DEBUG_TYPE, "NumOfCCMP",
+ MF.getFunction().getSubprogram(),
+ &MF.front());
+ R << "converted compare(s) to CCMP in function "
+ << ore::NV("Function", MF.getName());
+ return R;
+ });
+ }
+
return Changed;
}
-bool AArch64ConditionalComparesLegacy::runOnMachineFunction(
+//===----------------------------------------------------------------------===//
+// Pass wrappers
+//===----------------------------------------------------------------------===//
+
+namespace {
+class MachineConditionalComparesLegacy : public MachineFunctionPass {
+public:
+ static char ID;
+ MachineConditionalComparesLegacy() : MachineFunctionPass(ID) {
+ initializeMachineConditionalComparesLegacyPass(
+ *PassRegistry::getPassRegistry());
+ }
+ void getAnalysisUsage(AnalysisUsage &AU) const override;
+ bool runOnMachineFunction(MachineFunction &MF) override;
+ StringRef getPassName() const override {
+ return "Machine Conditional Compares";
+ }
+};
+} // end anonymous namespace
+
+char MachineConditionalComparesLegacy::ID = 0;
+char &llvm::MachineConditionalComparesLegacyID =
+ MachineConditionalComparesLegacy::ID;
+
+INITIALIZE_PASS_BEGIN(MachineConditionalComparesLegacy, DEBUG_TYPE,
+ "Machine Conditional Compares", false, false)
+INITIALIZE_PASS_DEPENDENCY(MachineBranchProbabilityInfoWrapperPass)
+INITIALIZE_PASS_DEPENDENCY(MachineDominatorTreeWrapperPass)
+INITIALIZE_PASS_DEPENDENCY(MachineLoopInfoWrapperPass)
+INITIALIZE_PASS_DEPENDENCY(MachineTraceMetricsWrapperPass)
+INITIALIZE_PASS_DEPENDENCY(MachineOptimizationRemarkEmitterPass)
+INITIALIZE_PASS_END(MachineConditionalComparesLegacy, DEBUG_TYPE,
+ "Machine Conditional Compares", false, false)
+
+void MachineConditionalComparesLegacy::getAnalysisUsage(
+ AnalysisUsage &AU) const {
+ AU.addRequired<MachineBranchProbabilityInfoWrapperPass>();
+ AU.addRequired<MachineDominatorTreeWrapperPass>();
+ AU.addPreserved<MachineDominatorTreeWrapperPass>();
+ AU.addRequired<MachineLoopInfoWrapperPass>();
+ AU.addPreserved<MachineLoopInfoWrapperPass>();
+ AU.addRequired<MachineTraceMetricsWrapperPass>();
+ AU.addPreserved<MachineTraceMetricsWrapperPass>();
+ AU.addRequired<MachineOptimizationRemarkEmitterPass>();
+ MachineFunctionPass::getAnalysisUsage(AU);
+}
+
+bool MachineConditionalComparesLegacy::runOnMachineFunction(
MachineFunction &MF) {
if (skipFunction(MF.getFunction()))
return false;
+ if (!MF.getSubtarget().enableCCMPFormation())
+ return false;
const MachineBranchProbabilityInfo *MBPI =
&getAnalysis<MachineBranchProbabilityInfoWrapperPass>().getMBPI();
@@ -985,14 +744,19 @@ bool AArch64ConditionalComparesLegacy::runOnMachineFunction(
MachineLoopInfo *Loops = &getAnalysis<MachineLoopInfoWrapperPass>().getLI();
MachineTraceMetrics *Traces =
&getAnalysis<MachineTraceMetricsWrapperPass>().getMTM();
+ MachineOptimizationRemarkEmitter *ORE =
+ &getAnalysis<MachineOptimizationRemarkEmitterPass>().getORE();
- AArch64ConditionalComparesImpl Impl(MBPI, DomTree, Loops, Traces);
+ MachineConditionalCompares Impl(MBPI, DomTree, Loops, Traces, ORE);
return Impl.run(MF);
}
PreservedAnalyses
-AArch64ConditionalComparesPass::run(MachineFunction &MF,
+MachineConditionalComparesPass::run(MachineFunction &MF,
MachineFunctionAnalysisManager &MFAM) {
+ if (!MF.getSubtarget().enableCCMPFormation())
+ return PreservedAnalyses::all();
+
const MachineBranchProbabilityInfo *MBPI =
&MFAM.getResult<MachineBranchProbabilityAnalysis>(MF);
MachineDominatorTree *DomTree =
@@ -1000,8 +764,10 @@ AArch64ConditionalComparesPass::run(MachineFunction &MF,
MachineLoopInfo *Loops = &MFAM.getResult<MachineLoopAnalysis>(MF);
MachineTraceMetrics *Traces =
&MFAM.getResult<MachineTraceMetricsAnalysis>(MF);
+ MachineOptimizationRemarkEmitter *ORE =
+ &MFAM.getResult<MachineOptimizationRemarkEmitterAnalysis>(MF);
- AArch64ConditionalComparesImpl Impl(MBPI, DomTree, Loops, Traces);
+ MachineConditionalCompares Impl(MBPI, DomTree, Loops, Traces, ORE);
bool Changed = Impl.run(MF);
if (!Changed)
return PreservedAnalyses::all();
diff --git a/llvm/lib/Passes/PassBuilder.cpp b/llvm/lib/Passes/PassBuilder.cpp
index 545b2905bdaaa..7782b80186feb 100644
--- a/llvm/lib/Passes/PassBuilder.cpp
+++ b/llvm/lib/Passes/PassBuilder.cpp
@@ -142,6 +142,7 @@
#include "llvm/CodeGen/MachineCSE.h"
#include "llvm/CodeGen/MachineCheckDebugify.h"
#include "llvm/CodeGen/MachineCombiner.h"
+#include "llvm/CodeGen/MachineConditionalCompares.h"
#include "llvm/CodeGen/MachineCopyPropagation.h"
#include "llvm/CodeGen/MachineDebugify.h"
#include "llvm/CodeGen/MachineDominanceFrontier.h"
diff --git a/llvm/lib/Target/AArch64/AArch64.h b/llvm/lib/Target/AArch64/AArch64.h
index bcf40a02e3cf5..406cffc38bdd4 100644
--- a/llvm/lib/Target/AArch64/AArch64.h
+++ b/llvm/lib/Target/AArch64/AArch64.h
@@ -50,7 +50,6 @@ FunctionPass *createAArch64RedundantCopyEliminationPass();
FunctionPass *createAArch64RedundantCondBranchPass();
FunctionPass *createAArch64CondBrTuning();
FunctionPass *createAArch64CompressJumpTablesPass();
-FunctionPass *createAArch64ConditionalCompares();
FunctionPass *createAArch64AdvSIMDScalar();
FunctionPass *createAArch64ISelDag(AArch64TargetMachine &TM,
CodeGenOptLevel OptLevel);
@@ -173,7 +172,6 @@ void initializeAArch64CollectLOHLegacyPass(PassRegistry &);
void initializeAArch64CompressJumpTablesLegacyPass(PassRegistry &);
void initializeAArch64CondBrTuningPass(PassRegistry &);
void initializeAArch64ConditionOptimizerLegacyPass(PassRegistry &);
-void initializeAArch64ConditionalComparesLegacyPass(PassRegistry &);
void initializeAArch64DAGToDAGISelLegacyPass(PassRegistry &);
void initializeAArch64DeadRegisterDefinitionsLegacyPass(PassRegistry &);
void initializeAArch64ExpandPseudoLegacyPass(PassRegistry &);
@@ -359,13 +357,6 @@ class AArch64RedundantCopyEliminationPass
MachineFunctionAnalysisManager &MFAM);
};
-class AArch64ConditionalComparesPass
- : public OptionalPassInfoMixin<AArch64ConditionalComparesPass> {
-public:
- PreservedAnalyses run(MachineFunction &MF,
- MachineFunctionAnalysisManager &MFAM);
-};
-
class AArch64SRLTDefineSuperRegsPass
: public OptionalPassInfoMixin<AArch64SRLTDefineSuperRegsPass> {
public:
diff --git a/llvm/lib/Target/AArch64/AArch64InstrInfo.cpp b/llvm/lib/Target/AArch64/AArch64InstrInfo.cpp
index 43aab1400e81f..7734ff15a2f35 100644
--- a/llvm/lib/Target/AArch64/AArch64InstrInfo.cpp
+++ b/llvm/lib/Target/AArch64/AArch64InstrInfo.cpp
@@ -34,6 +34,7 @@
#include "llvm/CodeGen/MachineFunction.h"
#include "llvm/CodeGen/MachineInstr.h"
#include "llvm/CodeGen/MachineInstrBuilder.h"
+#include "llvm/CodeGen/MachineInstrBundle.h"
#include "llvm/CodeGen/MachineMemOperand.h"
#include "llvm/CodeGen/MachineModuleInfo.h"
#include "llvm/CodeGen/MachineOperand.h"
@@ -79,6 +80,15 @@ STATISTIC(NumZCZeroingInstrsGPR, "Number of zero-cycle GPR zeroing "
"instructions expanded from canonical COPY");
// NumZCZeroingInstrsFPR is counted at AArch64AsmPrinter
+// Conditional-compare formation rejection reasons (see findConvertibleCompare)
+// and the number of cbz/cbnz branches converted (see convertToCCMP).
+STATISTIC(NumCmpTermRejs, "Number of ccmps rejected (CmpBB is cbz...)");
+STATISTIC(NumImmRangeRejs, "Number of ccmps rejected (Imm out of range)");
+STATISTIC(NumLiveDstRejs, "Number of ccmps rejected (Cmp dest live)");
+STATISTIC(NumMultNZCVUses, "Number of ccmps rejected (NZCV used)");
+STATISTIC(NumUnknNZCVDefs, "Number of ccmps rejected (NZCV def unknown)");
+STATISTIC(NumCompBranches, "Number of cbz/cbnz branches converted");
+
static cl::opt<unsigned>
CBDisplacementBits("aarch64-cb-offset-bits", cl::Hidden, cl::init(9),
cl::desc("Restrict range of CB instructions (DEBUG)"));
@@ -1334,6 +1344,312 @@ void AArch64InstrInfo::insertSelect(MachineBasicBlock &MBB,
.addImm(CC);
}
+//===----------------------------------------------------------------------===//
+// Conditional-compare formation hooks (see MachineConditionalCompares).
+//===----------------------------------------------------------------------===//
+
+// Parse a condition code returned by analyzeBranch, and compute the CondCode
+// corresponding to TBB. Returns false for branch shapes that cannot be turned
+// into a ccmp + br.cond (e.g. tbz/tbnz and the newer compare-and-branch forms).
+static bool parseCCMPCond(ArrayRef<MachineOperand> Cond,
+ AArch64CC::CondCode &CC) {
+ // A normal br.cond simply has the condition code.
+ if (Cond[0].getImm() != -1) {
+ if (Cond.size() != 1)
+ return false;
+ CC = (AArch64CC::CondCode)(int)Cond[0].getImm();
+ return true;
+ }
+ // For tbz and cbz instructions, the opcode is next.
+ switch (Cond[1].getImm()) {
+ default:
+ // This includes tbz / tbnz branches which can't be converted to
+ // ccmp + br.cond.
+ return false;
+ case AArch64::CBZW:
+ case AArch64::CBZX:
+ if (Cond.size() != 3)
+ return false;
+ CC = AArch64CC::EQ;
+ return true;
+ case AArch64::CBNZW:
+ case AArch64::CBNZX:
+ if (Cond.size() != 3)
+ return false;
+ CC = AArch64CC::NE;
+ return true;
+ }
+}
+
+static bool isDeadDef(const MachineRegisterInfo &MRI, unsigned DstReg) {
+ // Writes to the zero register are dead.
+ if (DstReg == AArch64::WZR || DstReg == AArch64::XZR)
+ return true;
+ if (!Register::isVirtualRegister(DstReg))
+ return false;
+ // A virtual register def without any uses will be marked dead later, and
+ // eventually replaced by the zero register.
+ return MRI.use_nodbg_empty(DstReg);
+}
+
+// Find the compare instruction in MBB that controls the conditional branch and
+// can be converted to a ccmp/ccmn/fccmp. Return nullptr if none is found.
+static MachineInstr *findConvertibleCompare(MachineBasicBlock *MBB,
+ const MachineRegisterInfo &MRI,
+ const TargetRegisterInfo *TRI) {
+ MachineBasicBlock::iterator I = MBB->getFirstTerminator();
+ if (I == MBB->end())
+ return nullptr;
+ // The terminator must be controlled by the flags.
+ if (!I->readsRegister(AArch64::NZCV, /*TRI=*/nullptr)) {
+ switch (I->getOpcode()) {
+ case AArch64::CBZW:
+ case AArch64::CBZX:
+ case AArch64::CBNZW:
+ case AArch64::CBNZX:
+ // These can be converted into a ccmp against #0.
+ return &*I;
+ }
+ ++NumCmpTermRejs;
+ return nullptr;
+ }
+
+ // Now find the instruction controlling the terminator.
+ for (MachineBasicBlock::iterator B = MBB->begin(); I != B;) {
+ I = prev_nodbg(I, MBB->begin());
+ assert(!I->isTerminator() && "Spurious terminator");
+ switch (I->getOpcode()) {
+ // cmp is an alias for subs with a dead destination register.
+ case AArch64::SUBSWri:
+ case AArch64::SUBSXri:
+ // cmn is an alias for adds with a dead destination register.
+ case AArch64::ADDSWri:
+ case AArch64::ADDSXri:
+ // Check that the immediate operand is within range, ccmp wants a uimm5.
+ // Rd = SUBSri Rn, imm, shift
+ if (I->getOperand(3).getImm() || !isUInt<5>(I->getOperand(2).getImm())) {
+ ++NumImmRangeRejs;
+ return nullptr;
+ }
+ [[fallthrough]];
+ case AArch64::SUBSWrr:
+ case AArch64::SUBSXrr:
+ case AArch64::ADDSWrr:
+ case AArch64::ADDSXrr:
+ if (isDeadDef(MRI, I->getOperand(0).getReg()))
+ return &*I;
+ ++NumLiveDstRejs;
+ return nullptr;
+ case AArch64::FCMPSrr:
+ case AArch64::FCMPDrr:
+ case AArch64::FCMPESrr:
+ case AArch64::FCMPEDrr:
+ return &*I;
+ }
+
+ // Check for flag reads and clobbers.
+ PhysRegInfo PRI = AnalyzePhysRegInBundle(*I, AArch64::NZCV, TRI);
+
+ if (PRI.Read) {
+ // The ccmp doesn't produce exactly the same flags as the original
+ // compare, so reject the transform if there are uses of the flags
+ // besides the terminators.
+ ++NumMultNZCVUses;
+ return nullptr;
+ }
+
+ if (PRI.Defined || PRI.Clobbered) {
+ ++NumUnknNZCVDefs;
+ return nullptr;
+ }
+ }
+ return nullptr;
+}
+
+// Layout of Info.TargetData for AArch64.
+enum { CCMP_HeadCC = 0, CCMP_TailCC = 1 };
+
+MCRegister AArch64InstrInfo::getConditionalCompareFlagReg() const {
+ return AArch64::NZCV;
+}
+
+bool AArch64InstrInfo::canConvertToCCMP(MachineBasicBlock &Head,
+ MachineBasicBlock &CmpBB,
+ ArrayRef<MachineOperand> HeadCond,
+ bool HeadTBBIsCmpBB,
+ ArrayRef<MachineOperand> CmpBBCond,
+ bool CmpBBTBBIsTail,
+ const MachineRegisterInfo &MRI,
+ CCmpConvInfo &Info) const {
+ AArch64CC::CondCode HeadCmpBBCC;
+ if (!parseCCMPCond(HeadCond, HeadCmpBBCC))
+ return false;
+ // The condition code should make Head branch to CmpBB.
+ if (!HeadTBBIsCmpBB)
+ HeadCmpBBCC = AArch64CC::getInvertedCondCode(HeadCmpBBCC);
+
+ AArch64CC::CondCode CmpBBTailCC;
+ if (!parseCCMPCond(CmpBBCond, CmpBBTailCC))
+ return false;
+ // The condition code should make CmpBB branch to Tail.
+ if (!CmpBBTBBIsTail)
+ CmpBBTailCC = AArch64CC::getInvertedCondCode(CmpBBTailCC);
+
+ MachineInstr *CmpMI = findConvertibleCompare(&CmpBB, MRI, &getRegisterInfo());
+ if (!CmpMI)
+ return false;
+
+ Info.CmpMI = CmpMI;
+ Info.TargetData[CCMP_HeadCC] = HeadCmpBBCC;
+ Info.TargetData[CCMP_TailCC] = CmpBBTailCC;
+ return true;
+}
+
+MachineInstr *
+AArch64InstrInfo::convertToCCMP(MachineBasicBlock &Head,
+ MachineBasicBlock::iterator SpliceLoc,
+ const DebugLoc &HeadTermDL,
+ ArrayRef<MachineOperand> HeadCond,
+ const CCmpConvInfo &Info,
+ MachineRegisterInfo &MRI) const {
+ MachineInstr *CmpMI = Info.CmpMI;
+ auto HeadCmpBBCC =
+ static_cast<AArch64CC::CondCode>(Info.TargetData[CCMP_HeadCC]);
+ auto CmpBBTailCC =
+ static_cast<AArch64CC::CondCode>(Info.TargetData[CCMP_TailCC]);
+
+ // If the Head terminator was one of the cbz / cbnz branches with built-in
+ // compare, we need to insert an explicit compare instruction in its place.
+ if (HeadCond[0].getImm() == -1) {
+ ++NumCompBranches;
+ unsigned Opc = 0;
+ switch (HeadCond[1].getImm()) {
+ case AArch64::CBZW:
+ case AArch64::CBNZW:
+ Opc = AArch64::SUBSWri;
+ break;
+ case AArch64::CBZX:
+ case AArch64::CBNZX:
+ Opc = AArch64::SUBSXri;
+ break;
+ default:
+ llvm_unreachable("Cannot convert Head branch");
+ }
+ const MCInstrDesc &MCID = get(Opc);
+ // Create a dummy virtual register for the SUBS def.
+ Register DestReg = MRI.createVirtualRegister(getRegClass(MCID, 0));
+ // Insert a SUBS Rn, #0 instruction instead of the cbz / cbnz.
+ BuildMI(Head, SpliceLoc, HeadTermDL, MCID)
+ .addReg(DestReg, RegState::Define | RegState::Dead)
+ .add(HeadCond[2])
+ .addImm(0)
+ .addImm(0);
+ // SUBS uses the GPR*sp register classes.
+ MRI.constrainRegClass(HeadCond[2].getReg(), getRegClass(MCID, 1));
+ }
+
+ // Now replace CmpMI with a ccmp instruction that also considers the incoming
+ // flags.
+ unsigned Opc = 0;
+ unsigned FirstOp = 1; // First CmpMI operand to copy.
+ bool isZBranch = false; // CmpMI is a cbz/cbnz instruction.
+ switch (CmpMI->getOpcode()) {
+ default:
+ llvm_unreachable("Unknown compare opcode");
+ case AArch64::SUBSWri: Opc = AArch64::CCMPWi; break;
+ case AArch64::SUBSWrr: Opc = AArch64::CCMPWr; break;
+ case AArch64::SUBSXri: Opc = AArch64::CCMPXi; break;
+ case AArch64::SUBSXrr: Opc = AArch64::CCMPXr; break;
+ case AArch64::ADDSWri: Opc = AArch64::CCMNWi; break;
+ case AArch64::ADDSWrr: Opc = AArch64::CCMNWr; break;
+ case AArch64::ADDSXri: Opc = AArch64::CCMNXi; break;
+ case AArch64::ADDSXrr: Opc = AArch64::CCMNXr; break;
+ case AArch64::FCMPSrr: Opc = AArch64::FCCMPSrr; FirstOp = 0; break;
+ case AArch64::FCMPDrr: Opc = AArch64::FCCMPDrr; FirstOp = 0; break;
+ case AArch64::FCMPESrr: Opc = AArch64::FCCMPESrr; FirstOp = 0; break;
+ case AArch64::FCMPEDrr: Opc = AArch64::FCCMPEDrr; FirstOp = 0; break;
+ case AArch64::CBZW:
+ case AArch64::CBNZW:
+ Opc = AArch64::CCMPWi;
+ FirstOp = 0;
+ isZBranch = true;
+ break;
+ case AArch64::CBZX:
+ case AArch64::CBNZX:
+ Opc = AArch64::CCMPXi;
+ FirstOp = 0;
+ isZBranch = true;
+ break;
+ }
+
+ // The ccmp instruction should set the flags according to the comparison when
+ // Head would have branched to CmpBB.
+ // The NZCV immediate operand should provide flags for the case where Head
+ // would have branched to Tail. These flags should cause the new Head
+ // terminator to branch to tail.
+ unsigned NZCV = AArch64CC::getNZCVToSatisfyCondCode(CmpBBTailCC);
+ const MCInstrDesc &MCID = get(Opc);
+ MRI.constrainRegClass(CmpMI->getOperand(FirstOp).getReg(),
+ getRegClass(MCID, 0));
+ if (CmpMI->getOperand(FirstOp + 1).isReg())
+ MRI.constrainRegClass(CmpMI->getOperand(FirstOp + 1).getReg(),
+ getRegClass(MCID, 1));
+ MachineInstrBuilder MIB = BuildMI(Head, CmpMI, CmpMI->getDebugLoc(), MCID)
+ .add(CmpMI->getOperand(FirstOp)); // Register Rn
+ if (isZBranch)
+ MIB.addImm(0); // cbz/cbnz Rn -> ccmp Rn, #0
+ else
+ MIB.add(CmpMI->getOperand(FirstOp + 1)); // Register Rm / Immediate
+ MIB.addImm(NZCV).addImm(HeadCmpBBCC);
+
+ // If CmpMI was a terminator, we need a new conditional branch to replace it.
+ // This now becomes a Head terminator.
+ if (isZBranch) {
+ bool isNZ = CmpMI->getOpcode() == AArch64::CBNZW ||
+ CmpMI->getOpcode() == AArch64::CBNZX;
+ BuildMI(Head, CmpMI, CmpMI->getDebugLoc(), get(AArch64::Bcc))
+ .addImm(isNZ ? AArch64CC::NE : AArch64CC::EQ)
+ .add(CmpMI->getOperand(1)); // Branch target.
+ }
+ return MIB;
+}
+
+int AArch64InstrInfo::getCCMPCodeSizeDelta(
+ const CCmpConvInfo &Info, ArrayRef<MachineOperand> HeadCond) const {
+ int delta = 0;
+ // If the Head terminator was one of the cbz / tbz branches with built-in
+ // compare, we need to insert an explicit compare instruction in its place
+ // plus a branch instruction.
+ if (HeadCond[0].getImm() == -1) {
+ switch (HeadCond[1].getImm()) {
+ case AArch64::CBZW:
+ case AArch64::CBNZW:
+ case AArch64::CBZX:
+ case AArch64::CBNZX:
+ // Therefore delta += 1
+ delta = 1;
+ break;
+ default:
+ llvm_unreachable("Cannot convert Head branch");
+ }
+ }
+ // If the Cmp terminator was one of the cbz / tbz branches with
+ // built-in compare, it will be turned into a compare instruction
+ // into Head, but we do not save any instruction.
+ // Otherwise, we save the branch instruction.
+ switch (Info.CmpMI->getOpcode()) {
+ default:
+ --delta;
+ break;
+ case AArch64::CBZW:
+ case AArch64::CBNZW:
+ case AArch64::CBZX:
+ case AArch64::CBNZX:
+ break;
+ }
+ return delta;
+}
+
// Return true if Imm can be loaded into a register by a "cheap" sequence of
// instructions. For now, "cheap" means at most two instructions.
static bool isCheapImmediate(const MachineInstr &MI, unsigned BitSize) {
diff --git a/llvm/lib/Target/AArch64/AArch64InstrInfo.h b/llvm/lib/Target/AArch64/AArch64InstrInfo.h
index d9e34365479f1..29e5801ebb460 100644
--- a/llvm/lib/Target/AArch64/AArch64InstrInfo.h
+++ b/llvm/lib/Target/AArch64/AArch64InstrInfo.h
@@ -430,6 +430,20 @@ class AArch64InstrInfo final : public AArch64GenInstrInfo {
const DebugLoc &DL, Register DstReg,
ArrayRef<MachineOperand> Cond, Register TrueReg,
Register FalseReg) const override;
+ MCRegister getConditionalCompareFlagReg() const override;
+ bool canConvertToCCMP(MachineBasicBlock &Head, MachineBasicBlock &CmpBB,
+ ArrayRef<MachineOperand> HeadCond, bool HeadTBBIsCmpBB,
+ ArrayRef<MachineOperand> CmpBBCond, bool CmpBBTBBIsTail,
+ const MachineRegisterInfo &MRI,
+ CCmpConvInfo &Info) const override;
+ MachineInstr *convertToCCMP(MachineBasicBlock &Head,
+ MachineBasicBlock::iterator SpliceLoc,
+ const DebugLoc &HeadTermDL,
+ ArrayRef<MachineOperand> HeadCond,
+ const CCmpConvInfo &Info,
+ MachineRegisterInfo &MRI) const override;
+ int getCCMPCodeSizeDelta(const CCmpConvInfo &Info,
+ ArrayRef<MachineOperand> HeadCond) const override;
void insertNoop(MachineBasicBlock &MBB,
MachineBasicBlock::iterator MI) const override;
diff --git a/llvm/lib/Target/AArch64/AArch64PassRegistry.def b/llvm/lib/Target/AArch64/AArch64PassRegistry.def
index 7f1fcaf638bf4..7830be422a286 100644
--- a/llvm/lib/Target/AArch64/AArch64PassRegistry.def
+++ b/llvm/lib/Target/AArch64/AArch64PassRegistry.def
@@ -40,7 +40,6 @@ LOOP_PASS(
MACHINE_FUNCTION_PASS("aarch64-a57-fp-load-balancing", AArch64A57FPLoadBalancingPass())
MACHINE_FUNCTION_PASS("aarch64-asm-printer", AArch64AsmPrinterPass())
MACHINE_FUNCTION_PASS("aarch64-branch-targets", AArch64BranchTargetsPass())
-MACHINE_FUNCTION_PASS("aarch64-ccmp", AArch64ConditionalComparesPass())
MACHINE_FUNCTION_PASS("aarch64-collect-loh", AArch64CollectLOHPass())
MACHINE_FUNCTION_PASS("aarch64-condopt", AArch64ConditionOptimizerPass())
MACHINE_FUNCTION_PASS("aarch64-copyelim", AArch64RedundantCopyEliminationPass())
diff --git a/llvm/lib/Target/AArch64/AArch64Subtarget.cpp b/llvm/lib/Target/AArch64/AArch64Subtarget.cpp
index 4fef12ab37ad8..5b693e04e8d06 100644
--- a/llvm/lib/Target/AArch64/AArch64Subtarget.cpp
+++ b/llvm/lib/Target/AArch64/AArch64Subtarget.cpp
@@ -39,6 +39,10 @@ static cl::opt<bool>
EnableEarlyIfConvert("aarch64-early-ifcvt", cl::desc("Enable the early if "
"converter pass"), cl::init(true), cl::Hidden);
+static cl::opt<bool> EnableCCMP("aarch64-enable-ccmp",
+ cl::desc("Enable the CCMP formation pass"),
+ cl::init(true), cl::Hidden);
+
// If OS supports TBI, use this flag to enable it.
static cl::opt<bool>
UseAddressTopByteIgnored("aarch64-use-tbi", cl::desc("Assume that top byte of "
@@ -565,6 +569,16 @@ bool AArch64Subtarget::enableEarlyIfConversion() const {
return EnableEarlyIfConvert;
}
+bool AArch64Subtarget::enableCCMPFormation() const { return EnableCCMP; }
+
+TargetSubtargetInfo::CCmpConvHeuristics
+AArch64Subtarget::getCCmpConvHeuristics() const {
+ CCmpConvHeuristics H;
+ // AArch64 converts freely when minimizing code size.
+ H.UseCodeSizeDeltaOnMinSize = true;
+ return H;
+}
+
bool AArch64Subtarget::supportsAddressTopByteIgnored() const {
if (!UseAddressTopByteIgnored)
return false;
diff --git a/llvm/lib/Target/AArch64/AArch64Subtarget.h b/llvm/lib/Target/AArch64/AArch64Subtarget.h
index 98cc97ae0a695..80911bffab05c 100644
--- a/llvm/lib/Target/AArch64/AArch64Subtarget.h
+++ b/llvm/lib/Target/AArch64/AArch64Subtarget.h
@@ -390,6 +390,10 @@ class AArch64Subtarget final : public AArch64GenSubtargetInfo {
bool enableEarlyIfConversion() const override;
+ bool enableCCMPFormation() const override;
+
+ CCmpConvHeuristics getCCmpConvHeuristics() const override;
+
std::unique_ptr<PBQPRAConstraint> getCustomPBQPConstraints() const override;
bool isCallingConvWin64(CallingConv::ID CC, bool IsVarArg) const {
diff --git a/llvm/lib/Target/AArch64/AArch64TargetMachine.cpp b/llvm/lib/Target/AArch64/AArch64TargetMachine.cpp
index 8195ad04c3556..4a0e5a55eb87d 100644
--- a/llvm/lib/Target/AArch64/AArch64TargetMachine.cpp
+++ b/llvm/lib/Target/AArch64/AArch64TargetMachine.cpp
@@ -58,9 +58,6 @@
using namespace llvm;
-static cl::opt<bool> EnableCCMP("aarch64-enable-ccmp",
- cl::desc("Enable the CCMP formation pass"),
- cl::init(true), cl::Hidden);
static cl::opt<bool>
EnableCondBrTuning("aarch64-enable-cond-br-tune",
@@ -249,7 +246,6 @@ LLVMInitializeAArch64Target() {
initializeAArch64BranchTargetsLegacyPass(PR);
initializeAArch64CollectLOHLegacyPass(PR);
initializeAArch64CompressJumpTablesLegacyPass(PR);
- initializeAArch64ConditionalComparesLegacyPass(PR);
initializeAArch64ConditionOptimizerLegacyPass(PR);
initializeAArch64DeadRegisterDefinitionsLegacyPass(PR);
initializeAArch64ExpandPseudoLegacyPass(PR);
@@ -833,8 +829,7 @@ void AArch64PassConfig::addMachineSSAOptimization() {
bool AArch64PassConfig::addILPOpts() {
if (EnableCondOpt)
addPass(createAArch64ConditionOptimizerLegacyPass());
- if (EnableCCMP)
- addPass(createAArch64ConditionalCompares());
+ addPass(&MachineConditionalComparesLegacyID);
if (EnableMCR)
addPass(&MachineCombinerID);
if (EnableCondBrTuning)
diff --git a/llvm/lib/Target/AArch64/CMakeLists.txt b/llvm/lib/Target/AArch64/CMakeLists.txt
index 12a2214f8e58e..d9762c135c33d 100644
--- a/llvm/lib/Target/AArch64/CMakeLists.txt
+++ b/llvm/lib/Target/AArch64/CMakeLists.txt
@@ -51,7 +51,6 @@ add_llvm_target(AArch64CodeGen
AArch64CleanupLocalDynamicTLSPass.cpp
AArch64CollectLOH.cpp
AArch64CondBrTuning.cpp
- AArch64ConditionalCompares.cpp
AArch64DeadRegisterDefinitionsPass.cpp
AArch64ExpandImm.cpp
AArch64ExpandPseudoInsts.cpp
diff --git a/llvm/lib/Target/X86/X86CodeGenPassBuilder.cpp b/llvm/lib/Target/X86/X86CodeGenPassBuilder.cpp
index c05bd2ca7f3be..761364a9f0bd5 100644
--- a/llvm/lib/Target/X86/X86CodeGenPassBuilder.cpp
+++ b/llvm/lib/Target/X86/X86CodeGenPassBuilder.cpp
@@ -24,6 +24,7 @@
#include "llvm/CodeGen/JMCInstrumenter.h"
#include "llvm/CodeGen/KCFI.h"
#include "llvm/CodeGen/MachineCombiner.h"
+#include "llvm/CodeGen/MachineConditionalCompares.h"
#include "llvm/MC/MCAsmInfo.h"
#include "llvm/MC/MCStreamer.h"
#include "llvm/Passes/CodeGenPassBuilder.h"
@@ -130,6 +131,7 @@ void X86CodeGenPassBuilder::addPreLegalizeMachineIR(PassManagerWrapper &PMW) {
}
void X86CodeGenPassBuilder::addILPOpts(PassManagerWrapper &PMW) {
+ addMachineFunctionPass(MachineConditionalComparesPass(), PMW);
addMachineFunctionPass(EarlyIfConverterPass(), PMW);
if (X86EnableMachineCombinerPass)
addMachineFunctionPass(MachineCombinerPass(), PMW);
diff --git a/llvm/lib/Target/X86/X86InstrInfo.cpp b/llvm/lib/Target/X86/X86InstrInfo.cpp
index c2beff2ec524c..f16c27063239f 100644
--- a/llvm/lib/Target/X86/X86InstrInfo.cpp
+++ b/llvm/lib/Target/X86/X86InstrInfo.cpp
@@ -19,6 +19,7 @@
#include "X86TargetMachine.h"
#include "llvm/ADT/STLExtras.h"
#include "llvm/ADT/Sequence.h"
+#include "llvm/ADT/Statistic.h"
#include "llvm/CodeGen/LiveIntervals.h"
#include "llvm/CodeGen/LivePhysRegs.h"
#include "llvm/CodeGen/LiveVariables.h"
@@ -26,6 +27,7 @@
#include "llvm/CodeGen/MachineFrameInfo.h"
#include "llvm/CodeGen/MachineInstr.h"
#include "llvm/CodeGen/MachineInstrBuilder.h"
+#include "llvm/CodeGen/MachineInstrBundle.h"
#include "llvm/CodeGen/MachineModuleInfo.h"
#include "llvm/CodeGen/MachineOperand.h"
#include "llvm/CodeGen/MachineRegisterInfo.h"
@@ -51,6 +53,10 @@ using namespace llvm;
#define DEBUG_TYPE "x86-instr-info"
+// Conditional-compare formation rejection reasons (see findConvertibleCompare).
+STATISTIC(NumMultEFLAGSUses, "Number of ccmps rejected (EFLAGS used)");
+STATISTIC(NumUnknEFLAGSDefs, "Number of ccmps rejected (EFLAGS def unknown)");
+
#define GET_INSTRINFO_CTOR_DTOR
#include "X86GenInstrInfo.inc"
@@ -3270,6 +3276,258 @@ int X86::getCCMPCondFlagsFromCondCode(X86::CondCode CC) {
}
}
+//===----------------------------------------------------------------------===//
+// Conditional-compare formation hooks (see MachineConditionalCompares).
+//===----------------------------------------------------------------------===//
+
+// Parse a condition code returned by analyzeBranch. Reject the compound and any
+// explicitly unsupported condition codes.
+static bool parseCCMPCond(ArrayRef<MachineOperand> Cond, X86::CondCode &CC,
+ ArrayRef<X86::CondCode> UnsupportedCCs = {}) {
+ if (Cond.size() != 1)
+ return false;
+
+ CC = static_cast<X86::CondCode>(Cond[0].getImm());
+
+ for (const auto &UnsupportedCC : UnsupportedCCs)
+ if (CC == UnsupportedCC)
+ return false;
+
+ return CC != X86::COND_INVALID;
+}
+
+// Count the number of conditional branches in the terminator sequence of MBB.
+static unsigned getNumOfJcc(const MachineBasicBlock *MBB) {
+ unsigned NumOfJcc = 0;
+ for (auto It = MBB->rbegin(); It != MBB->rend(); ++It) {
+ if (!It->isTerminator())
+ return NumOfJcc;
+ if (It->getOpcode() == X86::JCC_1)
+ ++NumOfJcc;
+ }
+ return NumOfJcc;
+}
+
+// Return true if DstReg is a virtual-register def with no uses; it will be
+// marked dead later. CCMP/CTEST produce no general-purpose result, so only a
+// dead compare def can be folded.
+static bool isDeadDef(const MachineRegisterInfo &MRI, Register DstReg) {
+ if (!DstReg.isVirtual())
+ return false;
+ return MRI.use_nodbg_empty(DstReg);
+}
+
+// Find the compare instruction in MBB that controls the conditional branch and
+// can be converted to a ccmp/ctest. Return nullptr if none is found.
+static MachineInstr *findConvertibleCompare(MachineBasicBlock *MBB,
+ const MachineRegisterInfo &MRI,
+ const TargetRegisterInfo *TRI) {
+ MachineBasicBlock::iterator I = MBB->getFirstTerminator();
+ if (I == MBB->end())
+ return nullptr;
+ // The caller only reaches this hook once analyzeBranch has accepted CmpBB's
+ // conditional branch, so its terminator is a JCC_1 that reads EFLAGS.
+
+ // Now find the instruction controlling the terminator.
+ for (MachineBasicBlock::iterator B = MBB->begin(); I != B;) {
+ I = prev_nodbg(I, MBB->begin());
+ assert(!I->isTerminator() && "Spurious terminator");
+
+ switch (I->getOpcode()) {
+ // This pass runs before peephole optimization, so the SUB has not been
+ // optimized to CMP yet.
+ case X86::SUB8rr:
+ case X86::SUB16rr:
+ case X86::SUB32rr:
+ case X86::SUB64rr:
+ case X86::SUB8ri:
+ case X86::SUB16ri:
+ case X86::SUB32ri:
+ case X86::SUB64ri32:
+ case X86::SUB8rr_ND:
+ case X86::SUB16rr_ND:
+ case X86::SUB32rr_ND:
+ case X86::SUB64rr_ND:
+ case X86::SUB8ri_ND:
+ case X86::SUB16ri_ND:
+ case X86::SUB32ri_ND:
+ case X86::SUB64ri32_ND: {
+ if (!isDeadDef(MRI, I->getOperand(0).getReg()))
+ return nullptr;
+ return &*I;
+ }
+ case X86::CMP8rr:
+ case X86::CMP16rr:
+ case X86::CMP32rr:
+ case X86::CMP64rr:
+ case X86::CMP8ri:
+ case X86::CMP16ri:
+ case X86::CMP32ri:
+ case X86::CMP64ri32:
+ case X86::TEST8rr:
+ case X86::TEST16rr:
+ case X86::TEST32rr:
+ case X86::TEST64rr:
+ case X86::TEST8ri:
+ case X86::TEST16ri:
+ case X86::TEST32ri:
+ case X86::TEST64ri32:
+ return &*I;
+ default:
+ break;
+ }
+
+ // Check for flag reads and clobbers.
+ PhysRegInfo PRI = AnalyzePhysRegInBundle(*I, X86::EFLAGS, TRI);
+
+ if (PRI.Read) {
+ // The ccmp doesn't produce exactly the same flags as the original
+ // compare, so reject the transform if there are uses of the flags
+ // besides the terminators.
+ ++NumMultEFLAGSUses;
+ return nullptr;
+ }
+
+ if (PRI.Defined || PRI.Clobbered) {
+ ++NumUnknEFLAGSDefs;
+ return nullptr;
+ }
+ }
+ return nullptr;
+}
+
+// Return the conditional-compare opcode for a flag-setting compare opcode.
+static unsigned getCCMPOpcode(unsigned CmpOpc) {
+ switch (CmpOpc) {
+ default:
+ llvm_unreachable("Unknown compare opcode");
+ case X86::SUB8rr:
+ case X86::SUB8rr_ND:
+ return X86::CCMP8rr;
+ case X86::SUB16rr:
+ case X86::SUB16rr_ND:
+ return X86::CCMP16rr;
+ case X86::SUB32rr:
+ case X86::SUB32rr_ND:
+ return X86::CCMP32rr;
+ case X86::SUB64rr:
+ case X86::SUB64rr_ND:
+ return X86::CCMP64rr;
+ case X86::SUB8ri:
+ case X86::SUB8ri_ND:
+ return X86::CCMP8ri;
+ case X86::SUB16ri:
+ case X86::SUB16ri_ND:
+ return X86::CCMP16ri;
+ case X86::SUB32ri:
+ case X86::SUB32ri_ND:
+ return X86::CCMP32ri;
+ case X86::SUB64ri32:
+ case X86::SUB64ri32_ND:
+ return X86::CCMP64ri32;
+ case X86::CMP8rr:
+ return X86::CCMP8rr;
+ case X86::CMP16rr:
+ return X86::CCMP16rr;
+ case X86::CMP32rr:
+ return X86::CCMP32rr;
+ case X86::CMP64rr:
+ return X86::CCMP64rr;
+ case X86::CMP8ri:
+ return X86::CCMP8ri;
+ case X86::CMP16ri:
+ return X86::CCMP16ri;
+ case X86::CMP32ri:
+ return X86::CCMP32ri;
+ case X86::CMP64ri32:
+ return X86::CCMP64ri32;
+ case X86::TEST8rr:
+ return X86::CTEST8rr;
+ case X86::TEST16rr:
+ return X86::CTEST16rr;
+ case X86::TEST32rr:
+ return X86::CTEST32rr;
+ case X86::TEST64rr:
+ return X86::CTEST64rr;
+ case X86::TEST8ri:
+ return X86::CTEST8ri;
+ case X86::TEST16ri:
+ return X86::CTEST16ri;
+ case X86::TEST32ri:
+ return X86::CTEST32ri;
+ case X86::TEST64ri32:
+ return X86::CTEST64ri32;
+ }
+}
+
+// Layout of Info.TargetData for X86.
+enum { TDHeadCC = 0, TDTailCC = 1 };
+
+MCRegister X86InstrInfo::getConditionalCompareFlagReg() const {
+ return X86::EFLAGS;
+}
+
+bool X86InstrInfo::canConvertToCCMP(MachineBasicBlock &Head,
+ MachineBasicBlock &CmpBB,
+ ArrayRef<MachineOperand> HeadCond,
+ bool HeadTBBIsCmpBB,
+ ArrayRef<MachineOperand> CmpBBCond,
+ bool CmpBBTBBIsTail,
+ const MachineRegisterInfo &MRI,
+ CCmpConvInfo &Info) const {
+ // CCMP/CTEST resets all the bits of EFLAGS, so Head must contain only a
+ // single conditional branch.
+ if (getNumOfJcc(&Head) > 1)
+ return false;
+
+ X86::CondCode HeadCmpBBCC;
+ if (!parseCCMPCond(HeadCond, HeadCmpBBCC, {X86::COND_P, X86::COND_NP}))
+ return false;
+ // The condition code should make Head branch to CmpBB.
+ if (!HeadTBBIsCmpBB)
+ HeadCmpBBCC = X86::GetOppositeBranchCondition(HeadCmpBBCC);
+
+ X86::CondCode CmpBBTailCC;
+ if (!parseCCMPCond(CmpBBCond, CmpBBTailCC,
+ {X86::COND_NE_OR_P, X86::COND_E_AND_NP}))
+ return false;
+ // The condition code should make CmpBB branch to Tail.
+ if (!CmpBBTBBIsTail)
+ CmpBBTailCC = X86::GetOppositeBranchCondition(CmpBBTailCC);
+
+ MachineInstr *CmpMI = findConvertibleCompare(&CmpBB, MRI, &getRegisterInfo());
+ if (!CmpMI)
+ return false;
+
+ Info.CmpMI = CmpMI;
+ Info.TargetData[TDHeadCC] = HeadCmpBBCC;
+ Info.TargetData[TDTailCC] = CmpBBTailCC;
+ return true;
+}
+
+MachineInstr *X86InstrInfo::convertToCCMP(MachineBasicBlock &Head,
+ MachineBasicBlock::iterator SpliceLoc,
+ const DebugLoc &HeadTermDL,
+ ArrayRef<MachineOperand> HeadCond,
+ const CCmpConvInfo &Info,
+ MachineRegisterInfo &MRI) const {
+ MachineInstr *CmpMI = Info.CmpMI;
+ auto HeadCmpBBCC = static_cast<X86::CondCode>(Info.TargetData[TDHeadCC]);
+ auto CmpBBTailCC = static_cast<X86::CondCode>(Info.TargetData[TDTailCC]);
+
+ unsigned Opc = getCCMPOpcode(CmpMI->getOpcode());
+ const MCInstrDesc &MCID = get(Opc);
+ unsigned NumDefs = CmpMI->getDesc().getNumDefs();
+ MachineOperand Op0 = CmpMI->getOperand(NumDefs);
+ MachineOperand Op1 = CmpMI->getOperand(NumDefs + 1);
+ DebugLoc DL = CmpMI->getDebugLoc();
+ return BuildMI(Head, CmpMI, DL, MCID)
+ .add(Op0)
+ .add(Op1)
+ .addImm(X86::getCCMPCondFlagsFromCondCode(CmpBBTailCC))
+ .addImm(HeadCmpBBCC);
+}
+
#define GET_X86_NF_TRANSFORM_TABLE
#define GET_X86_ND2NONND_TABLE
#include "X86GenInstrMapping.inc"
diff --git a/llvm/lib/Target/X86/X86InstrInfo.h b/llvm/lib/Target/X86/X86InstrInfo.h
index b6fac55373d05..d65c1c110a478 100644
--- a/llvm/lib/Target/X86/X86InstrInfo.h
+++ b/llvm/lib/Target/X86/X86InstrInfo.h
@@ -474,6 +474,18 @@ class X86InstrInfo final : public X86GenInstrInfo {
const DebugLoc &DL, Register DstReg,
ArrayRef<MachineOperand> Cond, Register TrueReg,
Register FalseReg) const override;
+ MCRegister getConditionalCompareFlagReg() const override;
+ bool canConvertToCCMP(MachineBasicBlock &Head, MachineBasicBlock &CmpBB,
+ ArrayRef<MachineOperand> HeadCond, bool HeadTBBIsCmpBB,
+ ArrayRef<MachineOperand> CmpBBCond, bool CmpBBTBBIsTail,
+ const MachineRegisterInfo &MRI,
+ CCmpConvInfo &Info) const override;
+ MachineInstr *convertToCCMP(MachineBasicBlock &Head,
+ MachineBasicBlock::iterator SpliceLoc,
+ const DebugLoc &HeadTermDL,
+ ArrayRef<MachineOperand> HeadCond,
+ const CCmpConvInfo &Info,
+ MachineRegisterInfo &MRI) const override;
void copyPhysReg(MachineBasicBlock &MBB, MachineBasicBlock::iterator MI,
const DebugLoc &DL, Register DestReg, Register SrcReg,
bool KillSrc, bool RenamableDest = false,
diff --git a/llvm/lib/Target/X86/X86Subtarget.cpp b/llvm/lib/Target/X86/X86Subtarget.cpp
index ed2da3128b44a..dd658da3fdd57 100644
--- a/llvm/lib/Target/X86/X86Subtarget.cpp
+++ b/llvm/lib/Target/X86/X86Subtarget.cpp
@@ -53,6 +53,10 @@ static cl::opt<bool>
X86EarlyIfConv("x86-early-ifcvt", cl::Hidden,
cl::desc("Enable early if-conversion on X86"));
+// Enable the conditional-compare formation pass for X86.
+static cl::opt<bool> X86EnableCCMPOpt("x86-enable-ccmp-opt", cl::Hidden,
+ cl::desc("Enable CCMP formation on X86"),
+ cl::init(true));
/// Classify a blockaddress reference for the current subtarget according to how
/// we should reference it in a non-pcrel context.
@@ -371,6 +375,10 @@ bool X86Subtarget::enableEarlyIfConversion() const {
return canUseCMOV() && X86EarlyIfConv;
}
+bool X86Subtarget::enableCCMPFormation() const {
+ return X86EnableCCMPOpt && hasCCMP();
+}
+
void X86Subtarget::getPostRAMutations(
std::vector<std::unique_ptr<ScheduleDAGMutation>> &Mutations) const {
Mutations.push_back(createX86MacroFusionDAGMutation());
diff --git a/llvm/lib/Target/X86/X86Subtarget.h b/llvm/lib/Target/X86/X86Subtarget.h
index 6cfea56457910..c4a5f08da09aa 100644
--- a/llvm/lib/Target/X86/X86Subtarget.h
+++ b/llvm/lib/Target/X86/X86Subtarget.h
@@ -437,6 +437,8 @@ class X86Subtarget final : public X86GenSubtargetInfo {
bool enableEarlyIfConversion() const override;
+ bool enableCCMPFormation() const override;
+
void getPostRAMutations(std::vector<std::unique_ptr<ScheduleDAGMutation>>
&Mutations) const override;
diff --git a/llvm/lib/Target/X86/X86TargetMachine.cpp b/llvm/lib/Target/X86/X86TargetMachine.cpp
index 886405a0c7bae..f256155940d7f 100644
--- a/llvm/lib/Target/X86/X86TargetMachine.cpp
+++ b/llvm/lib/Target/X86/X86TargetMachine.cpp
@@ -497,6 +497,7 @@ void X86PassConfig::addPreLegalizeMachineIR() {
}
bool X86PassConfig::addILPOpts() {
+ addPass(&MachineConditionalComparesLegacyID);
addPass(&EarlyIfConverterLegacyID);
if (X86EnableMachineCombinerPass)
addPass(&MachineCombinerID);
diff --git a/llvm/test/CodeGen/AArch64/O3-pipeline.ll b/llvm/test/CodeGen/AArch64/O3-pipeline.ll
index 9b36359c339c8..355e8232b4f70 100644
--- a/llvm/test/CodeGen/AArch64/O3-pipeline.ll
+++ b/llvm/test/CodeGen/AArch64/O3-pipeline.ll
@@ -149,7 +149,9 @@
; CHECK-NEXT: AArch64 Condition Optimizer
; CHECK-NEXT: Machine Natural Loop Construction
; CHECK-NEXT: Machine Trace Metrics
-; CHECK-NEXT: AArch64 Conditional Compares
+; CHECK-NEXT: Lazy Machine Block Frequency Analysis
+; CHECK-NEXT: Machine Optimization Remark Emitter
+; CHECK-NEXT: Machine Conditional Compares
; CHECK-NEXT: Machine Register Class Info Analysis
; CHECK-NEXT: Lazy Machine Block Frequency Analysis
; CHECK-NEXT: Machine InstCombiner
diff --git a/llvm/test/CodeGen/AArch64/arm64-ccmp.ll b/llvm/test/CodeGen/AArch64/arm64-ccmp.ll
index 54d05c581bf2c..12341695b9ec4 100644
--- a/llvm/test/CodeGen/AArch64/arm64-ccmp.ll
+++ b/llvm/test/CodeGen/AArch64/arm64-ccmp.ll
@@ -1,6 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py
-; RUN: llc < %s -debugify-and-strip-all-safe -mcpu=cyclone -verify-machineinstrs -aarch64-enable-ccmp -aarch64-stress-ccmp | FileCheck %s --check-prefixes=CHECK,CHECK-SD
-; RUN: llc < %s -debugify-and-strip-all-safe -mcpu=cyclone -verify-machineinstrs -aarch64-enable-ccmp -aarch64-stress-ccmp -global-isel | FileCheck %s --check-prefixes=CHECK,CHECK-GI
+; RUN: llc < %s -debugify-and-strip-all-safe -mcpu=cyclone -verify-machineinstrs -aarch64-enable-ccmp -stress-machine-ccmp | FileCheck %s --check-prefixes=CHECK,CHECK-SD
+; RUN: llc < %s -debugify-and-strip-all-safe -mcpu=cyclone -verify-machineinstrs -aarch64-enable-ccmp -stress-machine-ccmp -global-isel | FileCheck %s --check-prefixes=CHECK,CHECK-GI
target triple = "arm64-apple-ios"
define i32 @single_same(i32 %a, i32 %b) nounwind ssp {
diff --git a/llvm/test/CodeGen/AArch64/ccmp-look-through-copy.mir b/llvm/test/CodeGen/AArch64/ccmp-look-through-copy.mir
index 0893c69009396..3a64fa0a4186f 100644
--- a/llvm/test/CodeGen/AArch64/ccmp-look-through-copy.mir
+++ b/llvm/test/CodeGen/AArch64/ccmp-look-through-copy.mir
@@ -1,5 +1,5 @@
-# RUN: llc -o - %s -mtriple=aarch64 -run-pass=aarch64-ccmp -aarch64-stress-ccmp | FileCheck %s
-# RUN: llc -o - %s -mtriple=aarch64 -passes=aarch64-ccmp -aarch64-stress-ccmp | FileCheck %s
+# RUN: llc -o - %s -mtriple=aarch64 -run-pass=machine-ccmp -stress-machine-ccmp | FileCheck %s
+# RUN: llc -o - %s -mtriple=aarch64 -passes=machine-ccmp -stress-machine-ccmp | FileCheck %s
---
name: ccmp-look-through-copy
tracksRegLiveness: true
diff --git a/llvm/test/CodeGen/AArch64/ccmp-successor-probs.mir b/llvm/test/CodeGen/AArch64/ccmp-successor-probs.mir
index b99b49fd86f56..95ea9b4c517fa 100644
--- a/llvm/test/CodeGen/AArch64/ccmp-successor-probs.mir
+++ b/llvm/test/CodeGen/AArch64/ccmp-successor-probs.mir
@@ -1,5 +1,5 @@
-# RUN: llc -o - %s -mtriple=aarch64--linux-gnu -mcpu=falkor -run-pass=aarch64-ccmp | FileCheck %s
-# RUN: llc -o - %s -mtriple=aarch64--linux-gnu -mcpu=falkor -passes=aarch64-ccmp | FileCheck %s
+# RUN: llc -o - %s -mtriple=aarch64--linux-gnu -mcpu=falkor -run-pass=machine-ccmp | FileCheck %s
+# RUN: llc -o - %s -mtriple=aarch64--linux-gnu -mcpu=falkor -passes=machine-ccmp | FileCheck %s
---
# This test checks that successor probabilties are properly updated after a
# ccmp-conversion.
diff --git a/llvm/test/CodeGen/X86/apx/ccmp-cost-model.mir b/llvm/test/CodeGen/X86/apx/ccmp-cost-model.mir
new file mode 100644
index 0000000000000..1492ecb2037c4
--- /dev/null
+++ b/llvm/test/CodeGen/X86/apx/ccmp-cost-model.mir
@@ -0,0 +1,164 @@
+# RUN: llc -o - %s -mtriple=x86_64 -mattr=+ccmp -mcpu=diamondrapids -x86-enable-ccmp-opt -run-pass=machine-ccmp | FileCheck %s
+# RUN: llc -o - %s -mtriple=x86_64 -mattr=+ccmp -mcpu=diamondrapids -x86-enable-ccmp-opt -passes=machine-ccmp | FileCheck %s
+
+# This test exercises the profitability cost model of the CCMP formation pass
+# (machine-ccmp), i.e. the paths that -stress-machine-ccmp bypasses. A fixed
+# -mcpu is required so MachineTraceMetrics produces deterministic depths.
+# positive_control confirms the pass is active by folding to a CTEST; the other
+# two functions are structurally convertible but rejected by shouldConvert().
+
+--- |
+ define void @positive_control(i32 %a, i32 %b) { ret void }
+ define void @reject_branch_delay(i32 %a, i32 %b) { ret void }
+ define void @reject_resource_depth(i32 %a, i32 %b) { ret void }
+...
+
+# Structurally convertible and cheap: the cost model accepts and a CTEST forms.
+---
+name: positive_control
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: positive_control
+ ; CHECK: bb.0:
+ ; CHECK: CTEST32rr
+ bb.0:
+ successors: %bb.1, %bb.3
+ liveins: $edi, $esi
+
+ %1:gr32 = COPY $esi
+ %0:gr32 = COPY $edi
+ TEST32rr %0, %0, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.1
+
+ bb.1:
+ successors: %bb.2, %bb.3
+
+ TEST32rr %1, %1, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.2
+
+ bb.2:
+ %3:gr32 = MOV32ri 1
+ $eax = COPY %3
+ RET 0, $eax
+
+ bb.3:
+ %2:gr32 = MOV32r0 implicit-def dead $eflags
+ $eax = COPY %2
+ RET 0, $eax
+...
+
+# Branch-delay budget: a long (flag-free) dependency chain feeds the CmpBB
+# compare, so folding it would delay the branch by more than 3/4 of the
+# misprediction penalty. shouldConvert() rejects; the two branches are kept.
+---
+name: reject_branch_delay
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: reject_branch_delay
+ ; CHECK: bb.0:
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK: bb.1:
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK-NOT: CCMP
+ ; CHECK-NOT: CTEST
+ bb.0:
+ successors: %bb.1, %bb.3
+ liveins: $edi, $esi
+
+ %1:gr32 = COPY $esi
+ %0:gr32 = COPY $edi
+ TEST32rr %0, %0, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.1
+
+ bb.1:
+ successors: %bb.2, %bb.3
+
+ %10:gr32 = NOT32r %1
+ %11:gr32 = NOT32r %10
+ %12:gr32 = NOT32r %11
+ %13:gr32 = NOT32r %12
+ %14:gr32 = NOT32r %13
+ %15:gr32 = NOT32r %14
+ %16:gr32 = NOT32r %15
+ %17:gr32 = NOT32r %16
+ %18:gr32 = NOT32r %17
+ %19:gr32 = NOT32r %18
+ %20:gr32 = NOT32r %19
+ %21:gr32 = NOT32r %20
+ %22:gr32 = NOT32r %21
+ %23:gr32 = NOT32r %22
+ TEST32rr %23, %23, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.2
+
+ bb.2:
+ %3:gr32 = MOV32ri 1
+ $eax = COPY %3
+ RET 0, $eax
+
+ bb.3:
+ %2:gr32 = MOV32r0 implicit-def dead $eflags
+ $eax = COPY %2
+ RET 0, $eax
+...
+
+# Resource depth: many independent speculatable instructions in CmpBB push the
+# resource depth above the Head critical path, so shouldConvert() rejects even
+# though the branch delay is within budget. The two branches are kept.
+---
+name: reject_resource_depth
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: reject_resource_depth
+ ; CHECK: bb.0:
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK: bb.1:
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK-NOT: CCMP
+ ; CHECK-NOT: CTEST
+ bb.0:
+ successors: %bb.1, %bb.3
+ liveins: $edi, $esi
+
+ %1:gr32 = COPY $esi
+ %0:gr32 = COPY $edi
+ TEST32rr %0, %0, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.1
+
+ bb.1:
+ successors: %bb.2, %bb.3
+
+ %10:gr32 = MOV32ri 1
+ %11:gr32 = MOV32ri 2
+ %12:gr32 = MOV32ri 3
+ %13:gr32 = MOV32ri 4
+ %14:gr32 = MOV32ri 5
+ %15:gr32 = MOV32ri 6
+ %16:gr32 = MOV32ri 7
+ %17:gr32 = MOV32ri 8
+ %18:gr32 = MOV32ri 9
+ %19:gr32 = MOV32ri 10
+ %20:gr32 = MOV32ri 11
+ %21:gr32 = MOV32ri 12
+ TEST32rr %1, %1, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.2
+
+ bb.2:
+ %3:gr32 = MOV32ri 1
+ $eax = COPY %3
+ RET 0, $eax
+
+ bb.3:
+ %2:gr32 = MOV32r0 implicit-def dead $eflags
+ $eax = COPY %2
+ RET 0, $eax
+...
diff --git a/llvm/test/CodeGen/X86/apx/ccmp-look-through-copy.mir b/llvm/test/CodeGen/X86/apx/ccmp-look-through-copy.mir
new file mode 100644
index 0000000000000..1e4f1869389ff
--- /dev/null
+++ b/llvm/test/CodeGen/X86/apx/ccmp-look-through-copy.mir
@@ -0,0 +1,41 @@
+# RUN: llc -o - %s -mtriple=x86_64-- -mattr=+ccmp -x86-enable-ccmp-opt -run-pass=machine-ccmp -stress-machine-ccmp | FileCheck %s
+# RUN: llc -o - %s -mtriple=x86_64-- -mattr=+ccmp -x86-enable-ccmp-opt -passes=machine-ccmp -stress-machine-ccmp | FileCheck %s
+
+# The Tail PHI selects copy-equivalent values from Head (%4) and CmpBB (%7),
+# both tracing to %0. lookThroughCopies must see them as equal so the triangle
+# is still recognized as trivially convertible and folds into a CCMP.
+---
+name: ccmp_look_through_copy
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: ccmp_look_through_copy
+ ; CHECK: bb.0:
+ ; CHECK: CCMP32rr
+ bb.0:
+ successors: %bb.1, %bb.2
+ liveins: $edi, $esi
+
+ %0:gr32 = COPY $edi
+ %1:gr32 = COPY $esi
+ %4:gr32 = COPY %0
+ %7:gr32 = COPY %4
+ %6:gr32 = SUB32rr %0, %1, implicit-def $eflags
+ JCC_1 %bb.2, 12, implicit $eflags
+ JMP_1 %bb.1
+
+ bb.1:
+ successors: %bb.2, %bb.3
+
+ %8:gr32 = SUB32rr %1, %0, implicit-def $eflags
+ JCC_1 %bb.2, 13, implicit $eflags
+ JMP_1 %bb.3
+
+ bb.2:
+ %9:gr32 = PHI %4, %bb.0, %7, %bb.1
+ $eax = COPY %9
+ RET 0, $eax
+
+ bb.3:
+ $eax = COPY %0
+ RET 0, $eax
+...
diff --git a/llvm/test/CodeGen/X86/apx/ccmp-reject.mir b/llvm/test/CodeGen/X86/apx/ccmp-reject.mir
new file mode 100644
index 0000000000000..7c25c547ae71c
--- /dev/null
+++ b/llvm/test/CodeGen/X86/apx/ccmp-reject.mir
@@ -0,0 +1,495 @@
+# RUN: llc -o - %s -mtriple=x86_64 -mattr=+ccmp -x86-enable-ccmp-opt -run-pass=machine-ccmp -stress-machine-ccmp | FileCheck %s
+# RUN: llc -o - %s -mtriple=x86_64 -mattr=+ccmp -x86-enable-ccmp-opt -passes=machine-ccmp -stress-machine-ccmp | FileCheck %s
+
+# This test exercises the structural/recognition reject paths of the
+# target-independent CCMP formation pass (machine-ccmp) and its X86 hook.
+# -stress-machine-ccmp bypasses the profitability cost model, so only the
+# canConvert() and X86 canConvertToCCMP()/findConvertibleCompare() checks
+# decide whether a CCMP/CTEST forms. Each reject function keeps its original
+# two compares/branches; the positive controls fold to a CTEST (from TEST) and
+# a CCMP (from a dead-def SUB).
+
+--- |
+ define void @nontrivial_tail_phi(i32 %a, i32 %b) { ret void }
+ define void @phi_in_cmpbb(i32 %a, i32 %b) { ret void }
+ define void @load_in_cmpbb(i32 %b, ptr %p) { ret void }
+ define void @eflags_read_in_cmpbb(i32 %a, i32 %b) { ret void }
+ define void @eflags_def_in_cmpbb(i32 %a, i32 %b) { ret void }
+ define void @head_two_jcc(i32 %a, i32 %b) { ret void }
+ define void @head_cond_parity(i32 %a, i32 %b) { ret void }
+ define void @cmpbb_cond_compound(i32 %a, i32 %b) { ret void }
+ define void @positive_control(i32 %a, i32 %b) { ret void }
+ define void @sub_live_def(i32 %a, i32 %b) { ret void }
+ define void @sub_dead_def(i32 %a, i32 %b) { ret void }
+...
+
+# 1. Non-trivial Tail PHI: the Tail PHI selects different values from Head vs
+# CmpBB, so trivialTailPHIs() rejects (stat NumPhiRejs).
+---
+name: nontrivial_tail_phi
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: nontrivial_tail_phi
+ ; CHECK: bb.0:
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK: bb.1:
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK-NOT: CCMP
+ ; CHECK-NOT: CTEST
+ bb.0:
+ successors: %bb.1, %bb.3
+ liveins: $edi, $esi
+
+ %0:gr32 = COPY $edi
+ %1:gr32 = COPY $esi
+ %2:gr32 = MOV32ri 7
+ TEST32rr %0, %0, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.1
+
+ bb.1:
+ successors: %bb.2, %bb.3
+
+ %3:gr32 = MOV32ri 9
+ TEST32rr %1, %1, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.2
+
+ bb.2:
+ %4:gr32 = MOV32ri 1
+ $eax = COPY %4
+ RET 0, $eax
+
+ bb.3:
+ %5:gr32 = PHI %2, %bb.0, %3, %bb.1
+ $eax = COPY %5
+ RET 0, $eax
+...
+
+# 2. PHI in CmpBB: a PHI at the front of CmpBB triggers the generic reject
+# (stat NumPhi2Rejs).
+---
+name: phi_in_cmpbb
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: phi_in_cmpbb
+ ; CHECK: bb.0:
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK: bb.1:
+ ; CHECK: PHI
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK-NOT: CCMP
+ ; CHECK-NOT: CTEST
+ bb.0:
+ successors: %bb.1, %bb.3
+ liveins: $edi, $esi
+
+ %0:gr32 = COPY $edi
+ %1:gr32 = COPY $esi
+ TEST32rr %0, %0, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.1
+
+ bb.1:
+ successors: %bb.2, %bb.3
+
+ %2:gr32 = PHI %0, %bb.0
+ TEST32rr %1, %1, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.2
+
+ bb.2:
+ $eax = COPY %2
+ RET 0, $eax
+
+ bb.3:
+ %3:gr32 = MOV32r0 implicit-def dead $eflags
+ $eax = COPY %3
+ RET 0, $eax
+...
+
+# 3. Load in CmpBB: a load before the compare cannot be speculated
+# (canSpeculateInstrs, stat NumSpeculateRejs).
+---
+name: load_in_cmpbb
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: load_in_cmpbb
+ ; CHECK: bb.0:
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK: bb.1:
+ ; CHECK: MOV32rm
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK-NOT: CCMP
+ ; CHECK-NOT: CTEST
+ bb.0:
+ successors: %bb.1, %bb.3
+ liveins: $edi, $rsi
+
+ %0:gr32 = COPY $edi
+ %1:gr64 = COPY $rsi
+ TEST32rr %0, %0, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.1
+
+ bb.1:
+ successors: %bb.2, %bb.3
+
+ %2:gr32 = MOV32rm %1, 1, $noreg, 0, $noreg
+ TEST32rr %2, %2, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.2
+
+ bb.2:
+ $eax = COPY %2
+ RET 0, $eax
+
+ bb.3:
+ %3:gr32 = MOV32r0 implicit-def dead $eflags
+ $eax = COPY %3
+ RET 0, $eax
+...
+
+# 4. An instruction that READS EFLAGS between the compare and the terminator
+# (X86 findConvertibleCompare PRI.Read, stat NumMultEFLAGSUses).
+---
+name: eflags_read_in_cmpbb
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: eflags_read_in_cmpbb
+ ; CHECK: bb.0:
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK: bb.1:
+ ; CHECK: TEST32rr
+ ; CHECK: SETCCr
+ ; CHECK: JCC_1
+ ; CHECK-NOT: CCMP
+ ; CHECK-NOT: CTEST
+ bb.0:
+ successors: %bb.1, %bb.3
+ liveins: $edi, $esi
+
+ %0:gr32 = COPY $edi
+ %1:gr32 = COPY $esi
+ TEST32rr %0, %0, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.1
+
+ bb.1:
+ successors: %bb.2, %bb.3
+
+ TEST32rr %1, %1, implicit-def $eflags
+ %2:gr8 = SETCCr 4, implicit $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.2
+
+ bb.2:
+ $al = COPY %2
+ RET 0, $al
+
+ bb.3:
+ %3:gr32 = MOV32r0 implicit-def dead $eflags
+ $eax = COPY %3
+ RET 0, $eax
+...
+
+# 5. An instruction that DEFINES/CLOBBERS EFLAGS between the flag-setting
+# compare and the terminator (X86 findConvertibleCompare PRI.Defined,
+# stat NumUnknEFLAGSDefs).
+---
+name: eflags_def_in_cmpbb
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: eflags_def_in_cmpbb
+ ; CHECK: bb.0:
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK: bb.1:
+ ; CHECK: TEST32rr
+ ; CHECK: ADD32ri
+ ; CHECK: JCC_1
+ ; CHECK-NOT: CCMP
+ ; CHECK-NOT: CTEST
+ bb.0:
+ successors: %bb.1, %bb.3
+ liveins: $edi, $esi
+
+ %0:gr32 = COPY $edi
+ %1:gr32 = COPY $esi
+ TEST32rr %0, %0, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.1
+
+ bb.1:
+ successors: %bb.2, %bb.3
+
+ TEST32rr %1, %1, implicit-def dead $eflags
+ %2:gr32 = ADD32ri %1, 1, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.2
+
+ bb.2:
+ $eax = COPY %2
+ RET 0, $eax
+
+ bb.3:
+ %3:gr32 = MOV32r0 implicit-def dead $eflags
+ $eax = COPY %3
+ RET 0, $eax
+...
+
+# 6. Head has more than one JCC_1 terminator (a compound COND_NE_OR_P head):
+# X86 canConvertToCCMP getNumOfJcc(Head) > 1 reject.
+---
+name: head_two_jcc
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: head_two_jcc
+ ; CHECK: bb.0:
+ ; CHECK: JCC_1 %bb.3, 10
+ ; CHECK: JCC_1 %bb.3, 5
+ ; CHECK: bb.1:
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK-NOT: CCMP
+ ; CHECK-NOT: CTEST
+ bb.0:
+ successors: %bb.3, %bb.1
+ liveins: $edi, $esi
+
+ %0:gr32 = COPY $edi
+ %1:gr32 = COPY $esi
+ TEST32rr %0, %0, implicit-def $eflags
+ JCC_1 %bb.3, 10, implicit $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.1
+
+ bb.1:
+ successors: %bb.2, %bb.3
+
+ TEST32rr %1, %1, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.2
+
+ bb.2:
+ %2:gr32 = MOV32ri 1
+ $eax = COPY %2
+ RET 0, $eax
+
+ bb.3:
+ %3:gr32 = MOV32r0 implicit-def dead $eflags
+ $eax = COPY %3
+ RET 0, $eax
+...
+
+# 7. Head branch condition is COND_P (parity): X86 parseCCMPCond reject-list
+# on Head ({COND_P, COND_NP}).
+---
+name: head_cond_parity
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: head_cond_parity
+ ; CHECK: bb.0:
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1 %bb.3, 10
+ ; CHECK: bb.1:
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK-NOT: CCMP
+ ; CHECK-NOT: CTEST
+ bb.0:
+ successors: %bb.1, %bb.3
+ liveins: $edi, $esi
+
+ %0:gr32 = COPY $edi
+ %1:gr32 = COPY $esi
+ TEST32rr %0, %0, implicit-def $eflags
+ JCC_1 %bb.3, 10, implicit $eflags
+ JMP_1 %bb.1
+
+ bb.1:
+ successors: %bb.2, %bb.3
+
+ TEST32rr %1, %1, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.2
+
+ bb.2:
+ %2:gr32 = MOV32ri 1
+ $eax = COPY %2
+ RET 0, $eax
+
+ bb.3:
+ %3:gr32 = MOV32r0 implicit-def dead $eflags
+ $eax = COPY %3
+ RET 0, $eax
+...
+
+# 8. CmpBB branch condition is a compound COND_NE_OR_P: X86 parseCCMPCond
+# reject-list on CmpBB ({COND_NE_OR_P, COND_E_AND_NP}).
+---
+name: cmpbb_cond_compound
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: cmpbb_cond_compound
+ ; CHECK: bb.0:
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK: bb.1:
+ ; CHECK: JCC_1 %bb.3, 10
+ ; CHECK: JCC_1 %bb.3, 5
+ ; CHECK-NOT: CCMP
+ ; CHECK-NOT: CTEST
+ bb.0:
+ successors: %bb.1, %bb.3
+ liveins: $edi, $esi
+
+ %0:gr32 = COPY $edi
+ %1:gr32 = COPY $esi
+ TEST32rr %0, %0, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.1
+
+ bb.1:
+ successors: %bb.3, %bb.2
+
+ TEST32rr %1, %1, implicit-def $eflags
+ JCC_1 %bb.3, 10, implicit $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.2
+
+ bb.2:
+ %2:gr32 = MOV32ri 1
+ $eax = COPY %2
+ RET 0, $eax
+
+ bb.3:
+ %3:gr32 = MOV32r0 implicit-def dead $eflags
+ $eax = COPY %3
+ RET 0, $eax
+...
+
+# 9. Positive control: a plain cmp/jcc -> cmp/jcc triangle DOES fold to a
+# conditional test (CTEST).
+---
+name: positive_control
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: positive_control
+ ; CHECK: bb.0:
+ ; CHECK: CTEST32rr
+ bb.0:
+ successors: %bb.1, %bb.3
+ liveins: $edi, $esi
+
+ %1:gr32 = COPY $esi
+ %0:gr32 = COPY $edi
+ TEST32rr %0, %0, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.1
+
+ bb.1:
+ successors: %bb.2, %bb.3
+
+ TEST32rr %1, %1, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.2
+
+ bb.2:
+ %3:gr32 = MOV32ri 1
+ $eax = COPY %3
+ RET 0, $eax
+
+ bb.3:
+ %2:gr32 = MOV32r0 implicit-def dead $eflags
+ $eax = COPY %2
+ RET 0, $eax
+...
+
+# 9. SUB with a live destination: the compare feeding CmpBB's branch is a SUB
+# whose result is used, so findConvertibleCompare()'s isDeadDef() check fails
+# (CCMP/CTEST produce no GP result). Rejected; both compares are kept.
+---
+name: sub_live_def
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: sub_live_def
+ ; CHECK: bb.0:
+ ; CHECK: TEST32rr
+ ; CHECK: JCC_1
+ ; CHECK: bb.1:
+ ; CHECK: SUB32rr
+ ; CHECK: JCC_1
+ ; CHECK-NOT: CCMP
+ ; CHECK-NOT: CTEST
+ bb.0:
+ successors: %bb.1, %bb.3
+ liveins: $edi, $esi
+
+ %1:gr32 = COPY $esi
+ %0:gr32 = COPY $edi
+ TEST32rr %0, %0, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.1
+
+ bb.1:
+ successors: %bb.2, %bb.3
+
+ %10:gr32 = SUB32rr %1, %0, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.2
+
+ bb.2:
+ $eax = COPY %10
+ RET 0, $eax
+
+ bb.3:
+ %2:gr32 = MOV32r0 implicit-def dead $eflags
+ $eax = COPY %2
+ RET 0, $eax
+...
+
+# 10. Positive control (SUB): this pass runs before the SUB is shrunk to CMP, so
+# a SUB with a dead (unused) destination is recognized like CMP and folds to
+# a CCMP.
+---
+name: sub_dead_def
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: sub_dead_def
+ ; CHECK: bb.0:
+ ; CHECK: CCMP32rr
+ bb.0:
+ successors: %bb.1, %bb.3
+ liveins: $edi, $esi
+
+ %1:gr32 = COPY $esi
+ %0:gr32 = COPY $edi
+ TEST32rr %0, %0, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.1
+
+ bb.1:
+ successors: %bb.2, %bb.3
+
+ %10:gr32 = SUB32rr %1, %0, implicit-def $eflags
+ JCC_1 %bb.3, 5, implicit $eflags
+ JMP_1 %bb.2
+
+ bb.2:
+ %3:gr32 = MOV32ri 1
+ $eax = COPY %3
+ RET 0, $eax
+
+ bb.3:
+ %2:gr32 = MOV32r0 implicit-def dead $eflags
+ $eax = COPY %2
+ RET 0, $eax
+...
diff --git a/llvm/test/CodeGen/X86/apx/ccmp.ll b/llvm/test/CodeGen/X86/apx/ccmp.ll
index 8629b179a1797..28d502d6c82ca 100644
--- a/llvm/test/CodeGen/X86/apx/ccmp.ll
+++ b/llvm/test/CodeGen/X86/apx/ccmp.ll
@@ -1,6 +1,7 @@
; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py
-; RUN: llc < %s -mtriple=x86_64-unknown -mattr=+ccmp -show-mc-encoding -verify-machineinstrs | FileCheck %s
-; RUN: llc < %s -mtriple=x86_64-unknown -mattr=+ccmp,+ndd -show-mc-encoding -verify-machineinstrs | FileCheck %s --check-prefix=NDD
+; RUN: llc < %s -mtriple=x86_64-unknown -mattr=+ccmp -x86-enable-ccmp-opt=false -show-mc-encoding -verify-machineinstrs | FileCheck %s --check-prefixes=CHECK,CCMP
+; RUN: llc < %s -mtriple=x86_64-unknown -mattr=+ccmp -x86-enable-ccmp-opt=true -show-mc-encoding -verify-machineinstrs | FileCheck %s --check-prefixes=CHECK,CCMP_OPT
+; RUN: llc < %s -mtriple=x86_64-unknown -mattr=+ccmp,+ndd -x86-enable-ccmp-opt=false -show-mc-encoding -verify-machineinstrs | FileCheck %s --check-prefix=NDD
; RUN: llc < %s -mtriple=x86_64-unknown -mattr=+zu -show-mc-encoding -verify-machineinstrs | FileCheck %s --check-prefixes=ZU_COMMON,PREFER_NO_LEGACY_SETCC
; RUN: llc < %s -mtriple=x86_64-unknown -mattr=+zu,+prefer-legacy-setcc -show-mc-encoding -verify-machineinstrs | FileCheck %s --check-prefixes=ZU_COMMON,PREFER_LEGACY_SETCC
@@ -509,15 +510,21 @@ if.end: ; preds = %entry, %if.then
}
define void @ccmp64rr_of_crossbb(i64 %a, i64 %b) {
-; CHECK-LABEL: ccmp64rr_of_crossbb:
-; CHECK: # %bb.0: # %bb
-; CHECK-NEXT: testq %rdi, %rdi # encoding: [0x48,0x85,0xff]
-; CHECK-NEXT: je .LBB7_2 # encoding: [0x74,A]
-; CHECK-NEXT: # fixup A - offset: 1, value: .LBB7_2, kind: FK_PCRel_1
-; CHECK-NEXT: # %bb.1: # %bb1
-; CHECK-NEXT: cmpq %rsi, %rdi # encoding: [0x48,0x39,0xf7]
-; CHECK-NEXT: .LBB7_2: # %bb3
-; CHECK-NEXT: retq # encoding: [0xc3]
+; CCMP-LABEL: ccmp64rr_of_crossbb:
+; CCMP: # %bb.0: # %bb
+; CCMP-NEXT: testq %rdi, %rdi # encoding: [0x48,0x85,0xff]
+; CCMP-NEXT: je .LBB7_2 # encoding: [0x74,A]
+; CCMP-NEXT: # fixup A - offset: 1, value: .LBB7_2, kind: FK_PCRel_1
+; CCMP-NEXT: # %bb.1: # %bb1
+; CCMP-NEXT: cmpq %rsi, %rdi # encoding: [0x48,0x39,0xf7]
+; CCMP-NEXT: .LBB7_2: # %bb3
+; CCMP-NEXT: retq # encoding: [0xc3]
+;
+; CCMP_OPT-LABEL: ccmp64rr_of_crossbb:
+; CCMP_OPT: # %bb.0: # %bb
+; CCMP_OPT-NEXT: testq %rdi, %rdi # encoding: [0x48,0x85,0xff]
+; CCMP_OPT-NEXT: ccmpneq {dfv=of} %rsi, %rdi # encoding: [0x62,0xf4,0xc4,0x05,0x39,0xf7]
+; CCMP_OPT-NEXT: retq # encoding: [0xc3]
;
; NDD-LABEL: ccmp64rr_of_crossbb:
; NDD: # %bb.0: # %bb
diff --git a/llvm/test/CodeGen/X86/llc-pipeline-npm.ll b/llvm/test/CodeGen/X86/llc-pipeline-npm.ll
index 52363d9844c00..f192a33b4d53c 100644
--- a/llvm/test/CodeGen/X86/llc-pipeline-npm.ll
+++ b/llvm/test/CodeGen/X86/llc-pipeline-npm.ll
@@ -131,6 +131,7 @@
; O2-NEXT: stack-coloring
; O2-NEXT: localstackalloc
; O2-NEXT: dead-mi-elimination
+; O2-NEXT: machine-ccmp
; O2-NEXT: early-ifcvt
; O2-NEXT: machine-combiner
; O2-NEXT: x86-cmov-conversion
@@ -334,6 +335,7 @@
; O3-WINDOWS-NEXT: stack-coloring
; O3-WINDOWS-NEXT: localstackalloc
; O3-WINDOWS-NEXT: dead-mi-elimination
+; O3-WINDOWS-NEXT: machine-ccmp
; O3-WINDOWS-NEXT: early-ifcvt
; O3-WINDOWS-NEXT: machine-combiner
; O3-WINDOWS-NEXT: x86-cmov-conversion
diff --git a/llvm/test/CodeGen/X86/opt-pipeline.ll b/llvm/test/CodeGen/X86/opt-pipeline.ll
index e0256b66fff89..aeac727384244 100644
--- a/llvm/test/CodeGen/X86/opt-pipeline.ll
+++ b/llvm/test/CodeGen/X86/opt-pipeline.ll
@@ -104,6 +104,9 @@
; CHECK-NEXT: MachineDominator Tree Construction
; CHECK-NEXT: Machine Natural Loop Construction
; CHECK-NEXT: Machine Trace Metrics
+; CHECK-NEXT: Lazy Machine Block Frequency Analysis
+; CHECK-NEXT: Machine Optimization Remark Emitter
+; CHECK-NEXT: Machine Conditional Compares
; CHECK-NEXT: Early If-Conversion
; CHECK-NEXT: Machine Register Class Info Analysis
; CHECK-NEXT: Lazy Machine Block Frequency Analysis
More information about the llvm-commits
mailing list