[llvm] [RegAllocFast] Lower tied operands, absorbing TwoAddressInstructionPass (PR #225316)

Fangrui Song via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 22 00:26:40 PDT 2026


https://github.com/MaskRay created https://github.com/llvm/llvm-project/pull/225316

TwoAddressInstructionPass inserts a copy for every tied operand it
cannot rewrite, which register allocation then tries to fold away. Teach
RegAllocFast to lower tied operands itself so that the pipeline can skip
the pass:

* A tied use that dies at the instruction takes over the tied def's
  register when its class contains it, and is otherwise copied into it,
  with the instruction's other reads of the value following the copy.
* Copy hints follow ties as chain links, so argument copies feeding
  two-address chains still fold.
* REG_SEQUENCE and INSERT_SUBREG expand to subregister COPYs; an undef
  REG_SEQUENCE source needs one only where a use reads its lane.

PHIElimination still runs, so the allocator's input is not SSA. The
lowering keys on the TiedOpsRewritten property, which the allocator now
sets itself, so partial pipelines (-run-pass, -start-before) follow the
MIR they are given.

X86 will opt in via TargetMachine::setEnableTiedFastRegAlloc(), with
only 5 llvm/test/CodeGen/X86 tests updated (better assembly), followed
by all targets.

On clang -O0 IR the output is identical to the standard pipeline's
except for spill slot and register numbering, and llc retires 2.3% fewer
instructions; on -O2 IR .text changes are within 0.01%.

RFC: https://discourse.llvm.org/t/make-regallocfast-consume-machineir-in-ssa-form/91607

>From 105d8a6324b129f5de038381ea5efa11cdeca87f Mon Sep 17 00:00:00 2001
From: Fangrui Song <i at maskray.me>
Date: Fri, 21 Aug 2026 02:38:31 -0700
Subject: [PATCH] [RegAllocFast] Lower tied operands, absorbing
 TwoAddressInstructionPass

TwoAddressInstructionPass inserts a copy for every tied operand it
cannot rewrite, which register allocation then tries to fold away. Teach
RegAllocFast to lower tied operands itself so that the pipeline can skip
the pass:

* A tied use that dies at the instruction takes over the tied def's
  register when its class contains it, and is otherwise copied into it,
  with the instruction's other reads of the value following the copy.
* Copy hints follow ties as chain links, so argument copies feeding
  two-address chains still fold.
* REG_SEQUENCE and INSERT_SUBREG expand to subregister COPYs; an undef
  REG_SEQUENCE source needs one only where a use reads its lane.

PHIElimination still runs, so the allocator's input is not SSA. The
lowering keys on the TiedOpsRewritten property, which the allocator now
sets itself, so partial pipelines (-run-pass, -start-before) follow the
MIR they are given.

X86 will opt in via TargetMachine::setEnableTiedFastRegAlloc(), with
only 5 llvm/test/CodeGen/X86 tests updated (better assembly), followed
by all targets.

On clang -O0 IR the output is identical to the standard pipeline's
except for spill slot and register numbering, and llc retires 2.3% fewer
instructions; on -O2 IR .text changes are within 0.01%.

RFC: https://discourse.llvm.org/t/make-regallocfast-consume-machineir-in-ssa-form/91607
---
 llvm/include/llvm/CodeGen/RegAllocFast.h      |   9 +-
 .../include/llvm/Target/CGPassBuilderOption.h |   1 +
 llvm/include/llvm/Target/TargetMachine.h      |   7 +
 llvm/lib/CodeGen/RegAllocFast.cpp             | 228 +++++++++++--
 llvm/lib/CodeGen/TargetPassConfig.cpp         |  15 +-
 llvm/lib/Passes/CodeGenPassBuilder.cpp        |   7 +-
 llvm/lib/Target/TargetMachine.cpp             |   2 +-
 .../AArch64/regallocfast-tied-reg-sequence.ll |  20 ++
 .../regallocfast-tied-reg-sequence.mir        |  67 ++++
 llvm/test/CodeGen/X86/O0-pipeline.ll          |   8 +
 llvm/test/CodeGen/X86/clobber_base_ptr.ll     |   1 +
 llvm/test/CodeGen/X86/norex-subreg.ll         |   1 +
 llvm/test/CodeGen/X86/pr11415.ll              |  12 +
 llvm/test/CodeGen/X86/regallocfast-tied.ll    |  58 ++++
 llvm/test/CodeGen/X86/regallocfast-tied.mir   | 312 ++++++++++++++++++
 15 files changed, 715 insertions(+), 33 deletions(-)
 create mode 100644 llvm/test/CodeGen/AArch64/regallocfast-tied-reg-sequence.ll
 create mode 100644 llvm/test/CodeGen/AArch64/regallocfast-tied-reg-sequence.mir
 create mode 100644 llvm/test/CodeGen/X86/regallocfast-tied.ll
 create mode 100644 llvm/test/CodeGen/X86/regallocfast-tied.mir

diff --git a/llvm/include/llvm/CodeGen/RegAllocFast.h b/llvm/include/llvm/CodeGen/RegAllocFast.h
index 82d7f5ba914a1..102f0f655a7dd 100644
--- a/llvm/include/llvm/CodeGen/RegAllocFast.h
+++ b/llvm/include/llvm/CodeGen/RegAllocFast.h
@@ -32,11 +32,10 @@ class RegAllocFastPass : public RequiredPassInfoMixin<RegAllocFastPass> {
   }
 
   MachineFunctionProperties getSetProperties() const {
-    if (Opts.ClearVRegs) {
-      return MachineFunctionProperties().setNoVRegs();
-    }
-
-    return MachineFunctionProperties();
+    MachineFunctionProperties P;
+    if (Opts.ClearVRegs)
+      P.setNoVRegs().setTiedOpsRewritten();
+    return P;
   }
 
   MachineFunctionProperties getClearedProperties() const {
diff --git a/llvm/include/llvm/Target/CGPassBuilderOption.h b/llvm/include/llvm/Target/CGPassBuilderOption.h
index f825cbdca439b..731961a27eea9 100644
--- a/llvm/include/llvm/Target/CGPassBuilderOption.h
+++ b/llvm/include/llvm/Target/CGPassBuilderOption.h
@@ -86,6 +86,7 @@ struct CGPassBuilderOption {
 
   cl::boolOrDefault VerifyMachineCode = cl::boolOrDefault::BOU_UNSET;
   cl::boolOrDefault EnableFastISelOption = cl::boolOrDefault::BOU_UNSET;
+  cl::boolOrDefault EnableRegAllocFastTied = cl::boolOrDefault::BOU_UNSET;
   cl::boolOrDefault EnableGlobalISelOption = cl::boolOrDefault::BOU_UNSET;
   cl::boolOrDefault DebugifyAndStripAll = cl::boolOrDefault::BOU_UNSET;
   cl::boolOrDefault DebugifyCheckAndStripAll = cl::boolOrDefault::BOU_UNSET;
diff --git a/llvm/include/llvm/Target/TargetMachine.h b/llvm/include/llvm/Target/TargetMachine.h
index 5b1dede9d7893..87cc90d015134 100644
--- a/llvm/include/llvm/Target/TargetMachine.h
+++ b/llvm/include/llvm/Target/TargetMachine.h
@@ -120,6 +120,7 @@ class LLVM_ABI TargetMachine {
 
   unsigned RequireStructuredCFG : 1;
   unsigned O0WantsFastISel : 1;
+  unsigned EnableTiedFastRegAlloc : 1;
 
   // PGO related tunables.
   std::optional<PGOOptions> PGOOption;
@@ -269,6 +270,12 @@ class LLVM_ABI TargetMachine {
   bool requiresStructuredCFG() const { return RequireStructuredCFG; }
   void setRequiresStructuredCFG(bool Value) { RequireStructuredCFG = Value; }
 
+  /// Whether the fast register allocator lowers tied operands itself instead
+  /// of running TwoAddressInstructionPass. AMDGPU anchors passes on
+  /// TwoAddressInstructionPassID and cannot enable it.
+  bool enableTiedFastRegAlloc() const { return EnableTiedFastRegAlloc; }
+  void setEnableTiedFastRegAlloc(bool Value) { EnableTiedFastRegAlloc = Value; }
+
   /// Returns the code generation relocation model. The choices are static, PIC,
   /// and dynamic-no-pic, and target default.
   Reloc::Model getRelocationModel() const;
diff --git a/llvm/lib/CodeGen/RegAllocFast.cpp b/llvm/lib/CodeGen/RegAllocFast.cpp
index 524ed5e069941..b5c8fff4108d4 100644
--- a/llvm/lib/CodeGen/RegAllocFast.cpp
+++ b/llvm/lib/CodeGen/RegAllocFast.cpp
@@ -16,6 +16,10 @@
 ///
 /// Each block is walked backwards: a use is the first reference reached and
 /// acquires a register, a def is the last and releases one.
+///
+/// Where the target enables it, TwoAddressInstructionPass is left out of the
+/// pipeline: this pass lowers tied operands and expands REG_SEQUENCE and
+/// INSERT_SUBREG itself.
 //
 //===----------------------------------------------------------------------===//
 
@@ -49,6 +53,7 @@
 #include "llvm/Support/Debug.h"
 #include "llvm/Support/ErrorHandling.h"
 #include "llvm/Support/raw_ostream.h"
+#include "llvm/Target/TargetMachine.h"
 #include <cassert>
 #include <tuple>
 #include <vector>
@@ -197,6 +202,10 @@ class RegAllocFastImpl {
   RegisterClassInfo RegClassInfo;
   const RegAllocFilterFunc ShouldAllocateRegisterImpl;
 
+  /// Tied operands reach this pass unrewritten (TwoAddressInstructionPass was
+  /// left out of the pipeline): lower them here.
+  bool LowerTiedOps = false;
+
   /// Basic block currently being allocated.
   MachineBasicBlock *MBB = nullptr;
 
@@ -342,6 +351,7 @@ class RegAllocFastImpl {
 
 private:
   void allocateBasicBlock(MachineBasicBlock &MBB);
+  void expandSubregPseudo(MachineInstr &MI);
 
   void addRegClassDefCounts(MutableArrayRef<unsigned> RegClassDefCounts,
                             Register Reg) const;
@@ -378,6 +388,7 @@ class RegAllocFastImpl {
   bool defineVirtReg(MachineInstr &MI, unsigned OpNum, Register VirtReg,
                      bool LookAtPhysRegUses = false);
   bool useVirtReg(MachineInstr &MI, MachineOperand &MO, Register VirtReg);
+  bool lowerTiedUse(MachineInstr &MI, MachineOperand &MO, LiveReg &LR);
 
   MCPhysReg getErrorAssignment(const LiveReg &LR, MachineInstr &MI,
                                const TargetRegisterClass &RC);
@@ -433,11 +444,10 @@ class RegAllocFast : public MachineFunctionPass {
   }
 
   MachineFunctionProperties getSetProperties() const override {
-    if (Impl.ClearVirtRegs) {
-      return MachineFunctionProperties().setNoVRegs();
-    }
-
-    return MachineFunctionProperties();
+    MachineFunctionProperties P;
+    if (Impl.ClearVirtRegs)
+      P.setNoVRegs().setTiedOpsRewritten();
+    return P;
   }
 
   MachineFunctionProperties getClearedProperties() const override {
@@ -760,8 +770,9 @@ bool RegAllocFastImpl::usePhysReg(MachineInstr &MI, MCRegister Reg) {
 
 /// Displace whatever holds \p Reg and reserve it, so a virtual register def
 /// cannot land on a register this instruction already writes. Released in the
-/// free-def-operands step, or after the uses for an early clobber; if the
-/// instruction also reads \p Reg it ends up reserved for the code above.
+/// free-def-operands step, after the uses for an early clobber, or by
+/// lowerTiedUse(); if the instruction also reads \p Reg it ends up reserved
+/// for the code above.
 bool RegAllocFastImpl::definePhysReg(MachineInstr &MI, MCRegister Reg) {
   bool displacedAny = displacePhysReg(MI, Reg);
   setPhysRegState(Reg, regPreAssigned);
@@ -862,11 +873,14 @@ void RegAllocFastImpl::assignDanglingDebugValues(MachineInstr &Definition,
       continue;
 
     // Test whether the physreg survives from the definition to the DBG_VALUE.
+    // A tied use that took over its def's register is assigned at an
+    // instruction that overwrites it, so start the scan there.
     MCRegister SetToReg = Reg;
     unsigned Limit = 20;
-    for (MachineBasicBlock::iterator I = std::next(Definition.getIterator()),
-                                     E = DbgValue->getIterator();
-         I != E; ++I) {
+    MachineBasicBlock::iterator I = Definition.getIterator();
+    if (!Definition.definesRegister(Reg, TRI))
+      ++I;
+    for (MachineBasicBlock::iterator E = DbgValue->getIterator(); I != E; ++I) {
       if (I->modifiesRegister(Reg, TRI) || --Limit == 0) {
         LLVM_DEBUG(dbgs() << "Register did not survive for " << *DbgValue
                           << '\n');
@@ -901,6 +915,22 @@ void RegAllocFastImpl::assignVirtToPhysReg(MachineInstr &AtMI, LiveReg &LR,
 
 static bool isCoalescable(const MachineInstr &MI) { return MI.isFullCopy(); }
 
+/// The operand \p MO is tied to.
+static const MachineOperand &getTiedOperand(const MachineInstr &MI,
+                                            const MachineOperand &MO) {
+  return MI.getOperand(MI.findTiedOperandIdx(MI.getOperandNo(&MO)));
+}
+
+/// The register \p DefMO's tied use reads, when the two end up in the same
+/// register: a subregister index on either side makes them differ.
+static Register getTiedUseReg(const MachineInstr &MI,
+                              const MachineOperand &DefMO) {
+  if (!DefMO.isTied() || DefMO.getSubReg())
+    return Register();
+  const MachineOperand &UseMO = getTiedOperand(MI, DefMO);
+  return UseMO.getSubReg() ? Register() : UseMO.getReg();
+}
+
 Register RegAllocFastImpl::traceCopyChain(Register Reg) const {
   static const unsigned ChainLengthLimit = 3;
   for (unsigned C = 0; C <= ChainLengthLimit; ++C) {
@@ -912,22 +942,33 @@ Register RegAllocFastImpl::traceCopyChain(Register Reg) const {
     if (!DefMO)
       return Register();
     const MachineInstr *Def = DefMO->getParent();
-    if (!isCoalescable(*Def))
+    if (isCoalescable(*Def)) {
+      Reg = Def->getOperand(1).getReg();
+      continue;
+    }
+    // A two-address instruction's def and tied use end up in the same
+    // register, so the tie continues the chain.
+    Reg = LowerTiedOps ? getTiedUseReg(*Def, *DefMO) : Register();
+    if (!Reg)
       return Register();
-    Reg = Def->getOperand(1).getReg();
   }
   return Register();
 }
 
-/// Check if any of \p VirtReg's definitions is a copy. If it is follow the
-/// chain of copies to check whether we reach a physical register we can
-/// coalesce with.
+/// Check if any of \p VirtReg's definitions is a copy or a tied def. If it is
+/// follow the chain of copies to check whether we reach a physical register we
+/// can coalesce with.
 Register RegAllocFastImpl::traceCopies(Register VirtReg) const {
   static const unsigned DefLimit = 3;
   unsigned C = 0;
-  for (const MachineInstr &MI : MRI->def_instructions(VirtReg)) {
-    if (isCoalescable(MI)) {
-      Register Reg = MI.getOperand(1).getReg();
+  for (const MachineOperand &DefMO : MRI->def_operands(VirtReg)) {
+    const MachineInstr &MI = *DefMO.getParent();
+    Register Reg;
+    if (isCoalescable(MI))
+      Reg = MI.getOperand(1).getReg();
+    else if (LowerTiedOps)
+      Reg = getTiedUseReg(MI, DefMO);
+    if (Reg) {
       Reg = traceCopyChain(Reg);
       if (Reg.isValid())
         return Reg;
@@ -1037,10 +1078,7 @@ void RegAllocFastImpl::allocVirtRegUndef(MachineOperand &MO) {
   for (const MachineOperand &Tied : MI.all_uses()) {
     if (!Tied.isTied() || Tied.getReg() != VirtReg)
       continue;
-    MCRegister DefReg =
-        MI.getOperand(MI.findTiedOperandIdx(MI.getOperandNo(&Tied)))
-            .getReg()
-            .asMCReg();
+    MCRegister DefReg = getTiedOperand(MI, Tied).getReg().asMCReg();
     for (MachineOperand &O : MI.all_uses()) {
       if (O.getReg() != VirtReg)
         continue;
@@ -1192,6 +1230,70 @@ bool RegAllocFastImpl::defineVirtReg(MachineInstr &MI, unsigned OpNum,
   return setPhysReg(MI, MO, *LRI);
 }
 
+/// Satisfy \p MO's tie: the value must be in the register the def phase already
+/// assigned to the tied def. Take that register over when nothing below \p MI
+/// holds the value and it may live there, otherwise copy into it ahead of
+/// \p MI.
+/// \return true if \p MO is fully lowered; false when it took the register
+/// over and useVirtReg() finishes it.
+bool RegAllocFastImpl::lowerTiedUse(MachineInstr &MI, MachineOperand &MO,
+                                    LiveReg &LR) {
+  const MachineOperand &DefMO = getTiedOperand(MI, MO);
+  assert(DefMO.getReg().isPhysical() && "tied def allocated before its use");
+  MCRegister DefReg = DefMO.getReg().asMCReg();
+  unsigned SubReg = MO.getSubReg();
+  if (!LR.PhysReg) {
+    // A subregister read is a truncation, and DefReg may be unallocatable
+    // (e.g. the base pointer, tied to an inline asm operand). An early-clobber
+    // def is written before the uses are read, so it must not share a register
+    // with another operand reading the value.
+    bool MustCopy = SubReg || !MRI->isAllocatable(DefReg) ||
+                    !MRI->getRegClass(LR.VirtReg)->contains(DefReg) ||
+                    (DefMO.isEarlyClobber() &&
+                     any_of(MI.all_uses(), [&](const MachineOperand &O) {
+                       return &O != &MO && O.getReg() == LR.VirtReg;
+                     }));
+    if (!MustCopy) {
+      freePhysReg(DefReg);
+      assignVirtToPhysReg(MI, LR, DefReg);
+    } else {
+      allocVirtReg(MI, LR, Register(), false);
+    }
+  }
+  MCRegister SrcReg = LR.PhysReg;
+  if (SubReg)
+    SrcReg = TRI->getSubReg(SrcReg, SubReg);
+  if (SrcReg == DefReg)
+    return false;
+
+  BuildMI(*MBB, MI, MI.getDebugLoc(), TII->get(TargetOpcode::COPY), DefReg)
+      .addReg(SrcReg);
+  LR.LastUse = &MI;
+  // The copy above MI reads SrcReg, so no other operand may take it.
+  markRegUsedInInstr(LR.PhysReg);
+  bool Renamable = !MRI->isReserved(DefReg);
+  MO.setReg(DefReg);
+  MO.setSubReg(0);
+  MO.setIsRenamable(Renamable);
+  // The copy makes DefReg hold the value, so the instruction's other reads
+  // follow it and SrcReg dies above MI. A read tied to another def owes that
+  // def's register instead.
+  if (!DefMO.isEarlyClobber()) {
+    for (MachineOperand &O : MI.all_uses()) {
+      if (O.isTied() || O.getReg() != LR.VirtReg || O.getSubReg() != SubReg)
+        continue;
+      O.setReg(DefReg);
+      O.setSubReg(0);
+      O.setIsKill(false);
+      O.setIsRenamable(Renamable);
+    }
+  }
+  // Def processing skips tied operands when freeing, so DefReg is still
+  // assigned here; the copy ends its live range.
+  freePhysReg(DefReg);
+  return true;
+}
+
 /// Allocates a register for a VirtReg use.
 /// \return true if MI's MachineOperands were re-arranged/invalidated.
 bool RegAllocFastImpl::useVirtReg(MachineInstr &MI, MachineOperand &MO,
@@ -1215,6 +1317,9 @@ bool RegAllocFastImpl::useVirtReg(MachineInstr &MI, MachineOperand &MO,
     assert((!MO.isKill() || LRI->LastUse == &MI) && "Invalid kill flag");
   }
 
+  if (LowerTiedOps && MO.isTied() && lowerTiedUse(MI, MO, *LRI))
+    return false;
+
   // If necessary allocate a register.
   if (!LRI->PhysReg) {
     assert(!MO.isTied() && "tied op should be allocated");
@@ -1507,7 +1612,7 @@ void RegAllocFastImpl::allocateInstruction(MachineInstr &MI) {
   // * free the def operands' registers
   // * displace registers clobbered by regmasks
   // * pre-assigned physreg uses
-  // * virtual register uses, inserting reloads
+  // * virtual register uses, inserting reloads and tied-operand copies
   // * undef uses
   // * free early-clobber defs
   //
@@ -1844,8 +1949,10 @@ void RegAllocFastImpl::allocateBasicBlock(MachineBasicBlock &MBB) {
 
   Coalesced.clear();
 
-  // Traverse block in reverse order allocating instructions one by one.
-  for (MachineInstr &MI : reverse(MBB)) {
+  // Lowering a tied operand inserts a copy ahead of MI. Its registers are
+  // already assigned, so visiting it would evict what still lives in the
+  // source.
+  for (MachineInstr &MI : make_early_inc_range(reverse(MBB))) {
     LLVM_DEBUG(dbgs() << "\n>> " << MI << "Regs:"; dumpState());
 
     // Special handling for debug values. Note that they are not allowed to
@@ -1892,6 +1999,65 @@ void RegAllocFastImpl::allocateBasicBlock(MachineBasicBlock &MBB) {
   LLVM_DEBUG(MBB.dump());
 }
 
+/// Expand REG_SEQUENCE and INSERT_SUBREG into subregister COPYs: the lowering
+/// TwoAddressInstructionPass performs when it runs before allocation, plus the
+/// base-value copy its tie processing provides.
+void RegAllocFastImpl::expandSubregPseudo(MachineInstr &MI) {
+  MachineBasicBlock &MBB = *MI.getParent();
+  const DebugLoc &DL = MI.getDebugLoc();
+  if (MI.isInsertSubreg()) {
+    // %d = INSERT_SUBREG %base, %sub, idx  ->  %d = COPY %base
+    //                                          %d.idx = COPY %sub
+    const MachineOperand &BaseMO = MI.getOperand(1);
+    if (!BaseMO.isUndef())
+      BuildMI(MBB, MI, DL, TII->get(TargetOpcode::COPY),
+              MI.getOperand(0).getReg())
+          .addReg(BaseMO.getReg(), RegState::NoFlags, BaseMO.getSubReg());
+    unsigned SubIdx = MI.getOperand(3).getImm();
+    MI.removeOperand(3);
+    assert(MI.getOperand(0).getSubReg() == 0 && "Unexpected subreg idx");
+    MI.getOperand(0).setSubReg(SubIdx);
+    MI.getOperand(0).setIsUndef(MI.getOperand(1).isUndef());
+    MI.removeOperand(1);
+    MI.setDesc(TII->get(TargetOpcode::COPY));
+    return;
+  }
+
+  // %d = REG_SEQUENCE %s1, idx1, ...  ->  undef %d.idx1 = COPY %s1
+  //                                       %d.idx2 = COPY %s2 ...
+  assert(MI.isRegSequence());
+  Register Dst = MI.getOperand(0).getReg();
+  // An undef source needs no copy: the read-undef flag on the first copy
+  // defines the whole register. One is still needed where a use reads that
+  // lane on its own, which would otherwise read an undefined subregister.
+  LaneBitmask ReadLanes = LaneBitmask::getNone();
+  for (const MachineOperand &Use : MRI->use_nodbg_operands(Dst))
+    if (unsigned UseSubIdx = Use.getSubReg())
+      ReadLanes |= TRI->getSubRegIndexLaneMask(UseSubIdx);
+
+  bool DefEmitted = false;
+  for (unsigned I = 1, E = MI.getNumOperands(); I + 1 < E; I += 2) {
+    const MachineOperand &SrcMO = MI.getOperand(I);
+    unsigned SubIdx = MI.getOperand(I + 1).getImm();
+    if (SrcMO.isUndef() &&
+        (ReadLanes & TRI->getSubRegIndexLaneMask(SubIdx)).none())
+      continue;
+    BuildMI(MBB, MI, DL, TII->get(TargetOpcode::COPY))
+        .addReg(Dst, RegState::Define | getUndefRegState(!DefEmitted), SubIdx)
+        .addReg(SrcMO.getReg(), getUndefRegState(SrcMO.isUndef()),
+                SrcMO.getSubReg());
+    DefEmitted = true;
+  }
+  // Every source was undef: uses of Dst still need a definition.
+  if (!DefEmitted) {
+    MI.setDesc(TII->get(TargetOpcode::IMPLICIT_DEF));
+    while (MI.getNumOperands() > 1)
+      MI.removeOperand(MI.getNumOperands() - 1);
+    return;
+  }
+  MI.eraseFromParent();
+}
+
 bool RegAllocFastImpl::runOnMachineFunction(MachineFunction &MF) {
   LLVM_DEBUG(dbgs() << "********** FAST REGISTER ALLOCATION **********\n"
                     << "********** Function: " << MF.getName() << '\n');
@@ -1906,6 +2072,18 @@ bool RegAllocFastImpl::runOnMachineFunction(MachineFunction &MF) {
   InstrGen = 0;
   UsedInInstr.assign(NumRegUnits, 0);
 
+  // MIR that already went through TwoAddressInstructionPass carries
+  // TiedOpsRewritten, so partial pipelines (-run-pass, -start-before) follow
+  // the input they are given.
+  LowerTiedOps = MF.getTarget().enableTiedFastRegAlloc() &&
+                 !MF.getProperties().hasTiedOpsRewritten();
+  if (LowerTiedOps) {
+    for (MachineBasicBlock &MBB : MF)
+      for (MachineInstr &MI : make_early_inc_range(MBB))
+        if (MI.isRegSequence() || MI.isInsertSubreg())
+          expandSubregPseudo(MI);
+  }
+
   // initialize the virtual->physical register map to have a 'null'
   // mapping for all virtual registers
   unsigned NumVirtRegs = MRI->getNumVirtRegs();
diff --git a/llvm/lib/CodeGen/TargetPassConfig.cpp b/llvm/lib/CodeGen/TargetPassConfig.cpp
index 9e73048264b0a..5e726da7cad46 100644
--- a/llvm/lib/CodeGen/TargetPassConfig.cpp
+++ b/llvm/lib/CodeGen/TargetPassConfig.cpp
@@ -56,6 +56,13 @@
 
 using namespace llvm;
 
+static cl::opt<cl::boolOrDefault> EnableRegAllocFastTied(
+    "regalloc-fast-tied",
+    cl::desc("Have the fast register allocator lower tied operands itself "
+             "instead of running TwoAddressInstructionPass (overrides the "
+             "target's default)"),
+    cl::Hidden);
+
 static cl::opt<bool>
     EnableIPRA("enable-ipra", cl::init(false), cl::Hidden,
                cl::desc("Enable interprocedural register allocation "
@@ -520,6 +527,7 @@ CGPassBuilderOption llvm::getCGPassBuilderOption() {
 #define SET_OPTION(Option) Opt.Option = Option;
 
   SET_OPTION(EnableFastISelOption)
+  SET_OPTION(EnableRegAllocFastTied)
   SET_OPTION(EnableGlobalISelOption)
   SET_OPTION(VerifyMachineCode)
   SET_OPTION(DisableAtExitBasedGlobalDtorLowering)
@@ -638,6 +646,10 @@ TargetPassConfig::TargetPassConfig(TargetMachine &TM, PassManagerBase &PM)
   if (TM.Options.EnableIPRA)
     setRequiresCodeGenSCCOrder();
 
+  if (EnableRegAllocFastTied != cl::boolOrDefault::BOU_UNSET)
+    TM.setEnableTiedFastRegAlloc(EnableRegAllocFastTied ==
+                                 cl::boolOrDefault::BOU_TRUE);
+
   if (EnableGlobalISelAbort.getNumOccurrences())
     TM.Options.GlobalISelAbort = EnableGlobalISelAbort;
 
@@ -1485,7 +1497,8 @@ bool TargetPassConfig::usingDefaultRegAlloc() const {
 /// register allocation. No coalescing or scheduling.
 void TargetPassConfig::addFastRegAlloc() {
   addPass(&PHIEliminationID);
-  addPass(&TwoAddressInstructionPassID);
+  if (!TM->enableTiedFastRegAlloc())
+    addPass(&TwoAddressInstructionPassID);
 
   addRegAssignAndRewriteFast();
 }
diff --git a/llvm/lib/Passes/CodeGenPassBuilder.cpp b/llvm/lib/Passes/CodeGenPassBuilder.cpp
index a24323d816caa..a040f11bbc2a4 100644
--- a/llvm/lib/Passes/CodeGenPassBuilder.cpp
+++ b/llvm/lib/Passes/CodeGenPassBuilder.cpp
@@ -150,6 +150,10 @@ CodeGenPassBuilder::CodeGenPassBuilder(TargetMachine &TM,
   if (Opt.EnableGlobalISelAbort)
     TM.Options.GlobalISelAbort = *Opt.EnableGlobalISelAbort;
 
+  if (Opt.EnableRegAllocFastTied != cl::boolOrDefault::BOU_UNSET)
+    TM.setEnableTiedFastRegAlloc(Opt.EnableRegAllocFastTied ==
+                                 cl::boolOrDefault::BOU_TRUE);
+
   // An explicit RegAlloc choice implies its pipeline: only the fast
   // allocator uses the unoptimized one.
   if (Opt.OptimizeRegAlloc == cl::boolOrDefault::BOU_UNSET) {
@@ -842,7 +846,8 @@ CodeGenPassBuilder::addRegAssignAndRewriteOptimized(PassManagerWrapper &PMW) {
 /// register allocation. No coalescing or scheduling.
 Error CodeGenPassBuilder::addFastRegAlloc(PassManagerWrapper &PMW) {
   addMachineFunctionPass(PHIEliminationPass(), PMW);
-  addMachineFunctionPass(TwoAddressInstructionPass(), PMW);
+  if (!TM.enableTiedFastRegAlloc())
+    addMachineFunctionPass(TwoAddressInstructionPass(), PMW);
   return addRegAssignAndRewriteFast(PMW);
 }
 
diff --git a/llvm/lib/Target/TargetMachine.cpp b/llvm/lib/Target/TargetMachine.cpp
index abf3afcc6969b..626dc6f88c4e3 100644
--- a/llvm/lib/Target/TargetMachine.cpp
+++ b/llvm/lib/Target/TargetMachine.cpp
@@ -43,7 +43,7 @@ TargetMachine::TargetMachine(const Target &T, StringRef DataLayoutString,
     : TheTarget(T), DL(DataLayoutString), TargetTriple(TT),
       TargetCPU(std::string(CPU)), TargetFS(std::string(FS)), AsmInfo(nullptr),
       MRI(nullptr), MII(nullptr), STI(nullptr), RequireStructuredCFG(false),
-      O0WantsFastISel(false), Options(Options) {}
+      O0WantsFastISel(false), EnableTiedFastRegAlloc(false), Options(Options) {}
 
 TargetMachine::~TargetMachine() = default;
 
diff --git a/llvm/test/CodeGen/AArch64/regallocfast-tied-reg-sequence.ll b/llvm/test/CodeGen/AArch64/regallocfast-tied-reg-sequence.ll
new file mode 100644
index 0000000000000..c2541ed060e27
--- /dev/null
+++ b/llvm/test/CodeGen/AArch64/regallocfast-tied-reg-sequence.ll
@@ -0,0 +1,20 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py
+; RUN: llc -mtriple=aarch64 -O0 -regalloc-fast-tied -verify-machineinstrs < %s | FileCheck %s
+
+;; REG_SEQUENCE reaches the allocator only when TwoAddressInstructionPass is
+;; skipped. Left in place it survives to the AsmPrinter, which cannot print it.
+;; NEON tuple intrinsics are a reachable source at -O0.
+
+declare void @llvm.aarch64.neon.st2.v4i32.p0(<4 x i32>, <4 x i32>, ptr)
+
+define void @reg_sequence_pair(<4 x i32> %a, <4 x i32> %b, ptr %p) nounwind {
+; CHECK-LABEL: reg_sequence_pair:
+; CHECK:       // %bb.0:
+; CHECK-NEXT:    mov v2.16b, v1.16b
+; CHECK-NEXT:    // kill: def $q0 killed $q0 def $q0_q1
+; CHECK-NEXT:    mov v1.16b, v2.16b
+; CHECK-NEXT:    st2 { v0.4s, v1.4s }, [x0]
+; CHECK-NEXT:    ret
+  call void @llvm.aarch64.neon.st2.v4i32.p0(<4 x i32> %a, <4 x i32> %b, ptr %p)
+  ret void
+}
diff --git a/llvm/test/CodeGen/AArch64/regallocfast-tied-reg-sequence.mir b/llvm/test/CodeGen/AArch64/regallocfast-tied-reg-sequence.mir
new file mode 100644
index 0000000000000..8f14d957232fc
--- /dev/null
+++ b/llvm/test/CodeGen/AArch64/regallocfast-tied-reg-sequence.mir
@@ -0,0 +1,67 @@
+# NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
+# RUN: llc -mtriple=aarch64 -run-pass=regallocfast -regalloc-fast-tied -verify-machineinstrs %s -o - | FileCheck %s
+
+# REG_SEQUENCE sources that instruction selection does not mark undef on its
+# own.
+
+--- |
+  define void @undef_source(ptr %p) { ret void }
+  define void @undef_source_lane_read(ptr %p) { ret void }
+  define void @all_sources_undef(ptr %p) { ret void }
+...
+---
+# The first copy defines the whole register, so the undef source needs none.
+name: undef_source
+tracksRegLiveness: true
+body: |
+  bb.0:
+    liveins: $d0, $x0
+    ; CHECK-LABEL: name: undef_source
+    ; CHECK: liveins: $d0, $x0
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: undef renamable $d0 = COPY killed renamable $d0, implicit-def $d0_d1
+    ; CHECK-NEXT: ST1Twov1d killed renamable $d0_d1, killed $x0
+    ; CHECK-NEXT: RET_ReallyLR
+    %0:fpr64 = COPY $d0
+    %2:dd = REG_SEQUENCE %0, %subreg.dsub0, undef %1:fpr64, %subreg.dsub1
+    ST1Twov1d killed %2, $x0
+    RET_ReallyLR
+...
+---
+# dsub1 is read on its own, so dropping the copy would read an undefined
+# subregister.
+name: undef_source_lane_read
+tracksRegLiveness: true
+body: |
+  bb.0:
+    liveins: $d0, $x0
+    ; CHECK-LABEL: name: undef_source_lane_read
+    ; CHECK: liveins: $d0, $x0
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: undef renamable $d0 = COPY killed renamable $d0, implicit-def $d0_d1
+    ; CHECK-NEXT: renamable $d1 = COPY undef renamable $d0
+    ; CHECK-NEXT: ST1Twov1d renamable $d0_d1, killed $x0
+    ; CHECK-NEXT: $d2 = COPY renamable $d1, implicit killed $d0_d1
+    ; CHECK-NEXT: RET_ReallyLR implicit killed $d2
+    %0:fpr64 = COPY $d0
+    %2:dd = REG_SEQUENCE %0, %subreg.dsub0, undef %1:fpr64, %subreg.dsub1
+    ST1Twov1d %2, $x0
+    $d2 = COPY %2.dsub1
+    RET_ReallyLR implicit $d2
+...
+---
+name: all_sources_undef
+tracksRegLiveness: true
+body: |
+  bb.0:
+    liveins: $x0
+    ; CHECK-LABEL: name: all_sources_undef
+    ; CHECK: liveins: $x0
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: renamable $d0_d1 = IMPLICIT_DEF
+    ; CHECK-NEXT: ST1Twov1d killed renamable $d0_d1, killed $x0
+    ; CHECK-NEXT: RET_ReallyLR
+    %2:dd = REG_SEQUENCE undef %0:fpr64, %subreg.dsub0, undef %1:fpr64, %subreg.dsub1
+    ST1Twov1d killed %2, $x0
+    RET_ReallyLR
+...
diff --git a/llvm/test/CodeGen/X86/O0-pipeline.ll b/llvm/test/CodeGen/X86/O0-pipeline.ll
index e8a3084563573..f722c37dbc198 100644
--- a/llvm/test/CodeGen/X86/O0-pipeline.ll
+++ b/llvm/test/CodeGen/X86/O0-pipeline.ll
@@ -2,9 +2,17 @@
 ; pass. Ignore it with 'grep -v'.
 ; RUN: llc -mtriple=x86_64-- -O0 -debug-pass=Structure < %s -o /dev/null 2>&1 \
 ; RUN:   | grep -v 'Verify generated machine code' | FileCheck %s
+; RUN: llc -mtriple=x86_64-- -O0 -debug-pass=Structure -regalloc-fast-tied < %s -o /dev/null 2>&1 \
+; RUN:   | FileCheck %s --check-prefix=TIED
 
 ; REQUIRES: asserts
 
+; The fast register allocator lowers tied operands itself, so
+; TwoAddressInstructionPass drops out while PHIElimination stays.
+; TIED: Eliminate PHI nodes for register allocation
+; TIED-NOT: Two-Address instruction pass
+; TIED: Fast Register Allocator
+
 ; CHECK-LABEL: Pass Arguments:
 ; CHECK-NEXT: Target Library Information
 ; CHECK-NEXT: Runtime Library Function Analysis
diff --git a/llvm/test/CodeGen/X86/clobber_base_ptr.ll b/llvm/test/CodeGen/X86/clobber_base_ptr.ll
index 2c39560f02d16..7257755d1591a 100644
--- a/llvm/test/CodeGen/X86/clobber_base_ptr.ll
+++ b/llvm/test/CodeGen/X86/clobber_base_ptr.ll
@@ -1,5 +1,6 @@
 ; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 4
 ; RUN: llc < %s | FileCheck %s
+; RUN: llc -regalloc=fast -regalloc-fast-tied -verify-machineinstrs < %s -o /dev/null
 
 target datalayout = "e-m:x-p:32:32-p270:32:32-p271:32:32-p272:64:64-i64:64-i128:128-f80:32-n8:16:32-a:0:32-S32"
 target triple = "i386-pc-windows-gnu"
diff --git a/llvm/test/CodeGen/X86/norex-subreg.ll b/llvm/test/CodeGen/X86/norex-subreg.ll
index 616194abd26a0..22201796444d9 100644
--- a/llvm/test/CodeGen/X86/norex-subreg.ll
+++ b/llvm/test/CodeGen/X86/norex-subreg.ll
@@ -1,4 +1,5 @@
 ; RUN: llc -O0 < %s -verify-machineinstrs
+; RUN: llc -O0 -regalloc-fast-tied < %s -verify-machineinstrs
 ; RUN: llc < %s -verify-machineinstrs
 target triple = "x86_64-apple-macosx10.7"
 
diff --git a/llvm/test/CodeGen/X86/pr11415.ll b/llvm/test/CodeGen/X86/pr11415.ll
index c8c902c4e05bd..65f3cdd9578e8 100644
--- a/llvm/test/CodeGen/X86/pr11415.ll
+++ b/llvm/test/CodeGen/X86/pr11415.ll
@@ -1,4 +1,5 @@
 ; RUN: llc -mtriple=x86_64-pc-linux %s -o - -regalloc=fast | FileCheck %s
+; RUN: llc -mtriple=x86_64-pc-linux -O0 -regalloc-fast-tied %s -o - | FileCheck --check-prefix=O0 %s
 
 ; We used to consider the early clobber in the second asm statement as
 ; defining %0 before it was read. This caused us to omit the
@@ -13,6 +14,17 @@
 ; CHECK-NEXT:	movq	-8(%rsp), %rax
 ; CHECK-NEXT:	ret
 
+; The asm reads the value twice, so the early-clobber tie needs a copy.
+; O0: 	#APP
+; O0-NEXT:	#NO_APP
+; O0-NEXT:	movq	%rcx, %rdx
+; O0-NEXT:	movq	%rdx, %rcx
+; O0-NEXT:	#APP
+; O0-NEXT:	#NO_APP
+; O0-NEXT:	movq	%rcx, -8(%rsp)
+; O0-NEXT:	movq	-8(%rsp), %rax
+; O0-NEXT:	retq
+
 define i64 @foo() {
 entry:
   %0 = tail call i64 asm "", "={cx}"() nounwind
diff --git a/llvm/test/CodeGen/X86/regallocfast-tied.ll b/llvm/test/CodeGen/X86/regallocfast-tied.ll
new file mode 100644
index 0000000000000..66c32c287a32b
--- /dev/null
+++ b/llvm/test/CodeGen/X86/regallocfast-tied.ll
@@ -0,0 +1,58 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 6
+; RUN: llc -mtriple=x86_64-linux -O0 -regalloc-fast-tied -verify-machineinstrs < %s | FileCheck %s
+
+;; Tied operands, which the allocator rewrites itself instead of leaving to
+;; TwoAddressInstructionPass.
+
+;; The tied use dies at the instruction, so it takes over the tied def's
+;; register and no copy is needed.
+define i32 @tied_use_dies(i32 %a, i32 %b) nounwind {
+; CHECK-LABEL: tied_use_dies:
+; CHECK:       # %bb.0:
+; CHECK-NEXT:    movl %edi, %eax
+; CHECK-NEXT:    addl %esi, %eax
+; CHECK-NEXT:    retq
+  %c = add i32 %a, %b
+  ret i32 %c
+}
+
+;; The tied use is read again afterwards, so it has to be copied into the def's
+;; register first, as TwoAddressInstructionPass would have done.
+define i32 @tied_use_lives_on(i32 %a, i32 %b) nounwind {
+; CHECK-LABEL: tied_use_lives_on:
+; CHECK:       # %bb.0:
+; CHECK-NEXT:    movl %edi, %eax
+; CHECK-NEXT:    addl %esi, %eax
+; CHECK-NEXT:    addl %edi, %eax
+; CHECK-NEXT:    retq
+  %c = add i32 %a, %b
+  %d = add i32 %c, %a
+  ret i32 %d
+}
+
+;; The result is allocated first and has no copy to take a hint from; tracing
+;; the hint through the tied def to the tied use reaches the argument's
+;; register, so the argument copy coalesces away and the add runs in %edi.
+define void @tied_hint_follows_tie(i32 %a, i32 %b, ptr %p) nounwind {
+; CHECK-LABEL: tied_hint_follows_tie:
+; CHECK:       # %bb.0:
+; CHECK-NEXT:    addl %esi, %edi
+; CHECK-NEXT:    movl %edi, (%rdx)
+; CHECK-NEXT:    retq
+  %r = add i32 %a, %b
+  store volatile i32 %r, ptr %p
+  ret void
+}
+
+;; The early-clobber scan finds no other read: the use takes the register over.
+define i32 @tied_early_clobber(i32 %a) nounwind {
+; CHECK-LABEL: tied_early_clobber:
+; CHECK:       # %bb.0:
+; CHECK-NEXT:    movl %edi, %eax
+; CHECK-NEXT:    #APP
+; CHECK-NEXT:    # %eax %eax %edi
+; CHECK-NEXT:    #NO_APP
+; CHECK-NEXT:    retq
+  %r = call i32 asm "// $0 $1 $2", "=&r,0,r"(i32 %a, i32 %a)
+  ret i32 %r
+}
diff --git a/llvm/test/CodeGen/X86/regallocfast-tied.mir b/llvm/test/CodeGen/X86/regallocfast-tied.mir
new file mode 100644
index 0000000000000..4a48abb066204
--- /dev/null
+++ b/llvm/test/CodeGen/X86/regallocfast-tied.mir
@@ -0,0 +1,312 @@
+# NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
+# RUN: llc -mtriple=x86_64-linux -run-pass=regallocfast -regalloc-fast-tied -verify-machineinstrs %s -o - | FileCheck %s
+
+# Tied-operand shapes that instruction selection does not produce reliably.
+
+--- |
+  define i32 @tied_use_live_out(i32 %a, i32 %b) { ret i32 0 }
+  define i32 @tied_use_live_in(i32 %a, i32 %b) { ret i32 0 }
+  define i32 @tied_use_dbg_value(i32 %a, i32 %b) !dbg !4 { ret i32 0 }
+  define i32 @tied_early_clobber_shared_use(i32 %a) { ret i32 0 }
+  define i16 @tied_use_subreg(i32 %a, i16 %b) { ret i16 0 }
+  define void @tied_use_class_mismatch(i32 %a) { ret void }
+  define i64 @tied_def_reserved(i64 %a) { ret i64 0 }
+  define i64 @insert_subreg_undef_base(i32 %a) { ret i64 0 }
+  define i64 @insert_subreg_live_base(i64 %a, i32 %b) { ret i64 0 }
+  define i32 @tied_use_read_twice(i32 %a) { ret i32 0 }
+  define i32 @tied_use_undef() { ret i32 0 }
+  define void @two_tied_uses_one_value(i32 %a) { ret void }
+
+  !llvm.dbg.cu = !{!0}
+  !llvm.module.flags = !{!2}
+  !0 = distinct !DICompileUnit(language: DW_LANG_C, file: !1, emissionKind: FullDebug)
+  !1 = !DIFile(filename: "t.c", directory: "/")
+  !2 = !{i32 2, !"Debug Info Version", i32 3}
+  !4 = distinct !DISubprogram(name: "tied_use_dbg_value", scope: !1, file: !1, line: 1, type: !5, spFlags: DISPFlagDefinition, unit: !0)
+  !5 = !DISubroutineType(types: !6)
+  !6 = !{}
+  !7 = !DILocalVariable(name: "a", scope: !4, file: !1, line: 1, type: !8)
+  !8 = !DIBasicType(name: "int", size: 32, encoding: DW_ATE_signed)
+  !9 = !DILocation(line: 1, scope: !4)
+...
+---
+# %0 is read again in a successor: it still takes the def's register over
+# here, since the spill at its def serves the later read.
+name: tied_use_live_out
+tracksRegLiveness: true
+body: |
+  ; CHECK-LABEL: name: tied_use_live_out
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT:   liveins: $edi, $esi
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   renamable $eax = COPY killed $edi
+  ; CHECK-NEXT:   MOV32mr %stack.0, 1, $noreg, 0, $noreg, $eax :: (store (s32) into %stack.0)
+  ; CHECK-NEXT:   renamable $eax = ADD32rr renamable $eax, killed renamable $esi, implicit-def dead $eflags
+  ; CHECK-NEXT:   JMP_1 %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   liveins: $eax
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   $ecx = MOV32rm %stack.0, 1, $noreg, 0, $noreg :: (load (s32) from %stack.0)
+  ; CHECK-NEXT:   RET64 implicit killed $eax, implicit killed $ecx
+  bb.0:
+    liveins: $edi, $esi
+    %0:gr32 = COPY $edi
+    %1:gr32 = COPY $esi
+    %2:gr32 = ADD32rr %0, %1, implicit-def dead $eflags
+    $eax = COPY %2
+    JMP_1 %bb.1
+
+  bb.1:
+    liveins: $eax
+    $ecx = COPY %0
+    RET64 implicit $eax, implicit $ecx
+...
+---
+# %0 is defined in a predecessor: the reload at block begin lands in the def's
+# register.
+name: tied_use_live_in
+tracksRegLiveness: true
+body: |
+  ; CHECK-LABEL: name: tied_use_live_in
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT:   liveins: $edi, $esi
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   MOV32mr %stack.1, 1, $noreg, 0, $noreg, killed $edi :: (store (s32) into %stack.1)
+  ; CHECK-NEXT:   MOV32mr %stack.0, 1, $noreg, 0, $noreg, killed $esi :: (store (s32) into %stack.0)
+  ; CHECK-NEXT:   JMP_1 %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   $eax = MOV32rm %stack.1, 1, $noreg, 0, $noreg :: (load (s32) from %stack.1)
+  ; CHECK-NEXT:   $ecx = MOV32rm %stack.0, 1, $noreg, 0, $noreg :: (load (s32) from %stack.0)
+  ; CHECK-NEXT:   renamable $eax = ADD32rr killed renamable $eax, killed renamable $ecx, implicit-def dead $eflags
+  ; CHECK-NEXT:   RET64 implicit killed $eax
+  bb.0:
+    liveins: $edi, $esi
+    %0:gr32 = COPY $edi
+    %1:gr32 = COPY $esi
+    JMP_1 %bb.1
+
+  bb.1:
+    %2:gr32 = ADD32rr %0, %1, implicit-def dead $eflags
+    $eax = COPY %2
+    RET64 implicit $eax
+...
+---
+# The tied use takes over $eax, which the ADD then overwrites, so a DBG_VALUE
+# below cannot find %0 there.
+name: tied_use_dbg_value
+tracksRegLiveness: true
+body: |
+  bb.0:
+    liveins: $edi, $esi
+    ; CHECK-LABEL: name: tied_use_dbg_value
+    ; CHECK: liveins: $edi, $esi
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: renamable $eax = COPY killed $edi
+    ; CHECK-NEXT: renamable $eax = ADD32rr killed renamable $eax, killed renamable $esi, implicit-def dead $eflags
+    ; CHECK-NEXT: DBG_VALUE $noreg, $noreg, !6, !DIExpression(), debug-location !DILocation(line: 1, scope: !3)
+    ; CHECK-NEXT: RET64 implicit killed $eax
+    %0:gr32 = COPY $edi
+    %1:gr32 = COPY $esi
+    %2:gr32 = ADD32rr %0, %1, implicit-def dead $eflags
+    DBG_VALUE %0, $noreg, !7, !DIExpression(), debug-location !9
+    $eax = COPY %2
+    RET64 implicit $eax
+...
+---
+name: tied_early_clobber_shared_use
+tracksRegLiveness: true
+body: |
+  bb.0:
+    liveins: $edi
+    ; CHECK-LABEL: name: tied_early_clobber_shared_use
+    ; CHECK: liveins: $edi
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: $eax = COPY $edi
+    ; CHECK-NEXT: INLINEASM &"// $0 $1 $2", attdialect, regdef-ec:GR32, def early-clobber renamable $eax, reguse tiedto:$0, killed renamable $eax(tied-def 3), reguse:GR32, renamable $edi
+    ; CHECK-NEXT: RET64 implicit killed $eax
+    %0:gr32 = COPY $edi
+    INLINEASM &"// $0 $1 $2", attdialect, regdef-ec:GR32, def early-clobber %1:gr32, reguse tiedto:$0, %0(tied-def 3), reguse:GR32, %0
+    $eax = COPY %1
+    RET64 implicit $eax
+...
+---
+# A tied use reading a subregister is a truncation, so it cannot take over the
+# def's register: the copy performs the truncation and the operand reads the
+# def's register whole, as TwoAddressInstructionPass does. %0 is read again
+# afterwards, so it keeps its own register.
+name: tied_use_subreg
+tracksRegLiveness: true
+body: |
+  bb.0:
+    liveins: $edi, $esi
+    ; CHECK-LABEL: name: tied_use_subreg
+    ; CHECK: liveins: $edi, $esi
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: renamable $ecx = COPY killed $edi
+    ; CHECK-NEXT: $ax = COPY $cx
+    ; CHECK-NEXT: renamable $ax = ADD16rr renamable $ax, killed renamable $si, implicit-def dead $eflags
+    ; CHECK-NEXT: RET64 implicit killed $ax, implicit killed $ecx
+    %0:gr32 = COPY $edi
+    %1:gr16 = COPY $si
+    %2:gr16 = ADD16rr %0.sub_16bit, %1, implicit-def dead $eflags
+    $ecx = COPY %0
+    $ax = COPY %2
+    RET64 implicit $ax, implicit $ecx
+...
+---
+# The high-byte read forces GR32_ABCD on %0, which does not contain the $esi
+# the tied def was assigned, so %0 cannot take that register over and is copied
+# into it instead.
+name: tied_use_class_mismatch
+tracksRegLiveness: true
+body: |
+  bb.0:
+    liveins: $edi
+    ; CHECK-LABEL: name: tied_use_class_mismatch
+    ; CHECK: liveins: $edi
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: renamable $eax = COPY killed $edi
+    ; CHECK-NEXT: renamable $cl = COPY renamable $ah
+    ; CHECK-NEXT: $esi = COPY $eax
+    ; CHECK-NEXT: renamable $esi = SHR32ri killed renamable $esi, 1, implicit-def dead $eflags
+    ; CHECK-NEXT: RET64 implicit killed $esi, implicit killed $cl
+    %0:gr32_abcd = COPY $edi
+    %1:gr8_norex = COPY %0.sub_8bit_hi
+    %2:gr32 = SHR32ri %0, 1, implicit-def dead $eflags
+    $esi = COPY %2
+    $cl = COPY %1
+    RET64 implicit $esi, implicit $cl
+...
+---
+# Stack realignment with a variable-sized object reserves $rbx as the base
+# pointer, so the tied use is copied into it rather than taking it over.
+name: tied_def_reserved
+tracksRegLiveness: true
+frameInfo:
+  maxAlignment: 32
+stack:
+  - { id: 0, type: variable-sized, alignment: 1 }
+body: |
+  bb.0:
+    liveins: $rdi
+    ; CHECK-LABEL: name: tied_def_reserved
+    ; CHECK: liveins: $rdi
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: renamable $rdi = ADD64ri32 killed renamable $rdi, 1, implicit-def dead $eflags
+    ; CHECK-NEXT: $rbx = COPY $rdi
+    ; CHECK-NEXT: INLINEASM &"// $0", attdialect, regdef, implicit-def $rbx, reguse tiedto:$0, killed $rbx(tied-def 3), clobber, implicit-def dead early-clobber $eflags
+    ; CHECK-NEXT: $rax = COPY $rbx
+    ; CHECK-NEXT: RET64 implicit killed $rax
+    %0:gr64 = COPY $rdi
+    %1:gr64 = ADD64ri32 %0, 1, implicit-def dead $eflags
+    INLINEASM &"// $0", attdialect, regdef, implicit-def $rbx, reguse tiedto:$0, %1(tied-def 3), clobber, implicit-def early-clobber $eflags
+    $rax = COPY $rbx
+    RET64 implicit $rax
+...
+---
+name: insert_subreg_undef_base
+tracksRegLiveness: true
+body: |
+  bb.0:
+    liveins: $edi
+    ; CHECK-LABEL: name: insert_subreg_undef_base
+    ; CHECK: liveins: $edi
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: renamable $eax = COPY killed $edi
+    ; CHECK-NEXT: undef renamable $eax = COPY killed renamable $eax, implicit-def $rax
+    ; CHECK-NEXT: RET64 implicit killed $rax
+    %0:gr32 = COPY $edi
+    %1:gr64 = INSERT_SUBREG undef %2:gr64, %0, %subreg.sub_32bit
+    $rax = COPY %1
+    RET64 implicit $rax
+...
+---
+# The base is live afterwards, so the expansion copies it into the destination
+# before writing the subregister.
+name: insert_subreg_live_base
+tracksRegLiveness: true
+body: |
+  bb.0:
+    liveins: $rdi, $esi
+    ; CHECK-LABEL: name: insert_subreg_live_base
+    ; CHECK: liveins: $rdi, $esi
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: renamable $rcx = COPY killed $rdi
+    ; CHECK-NEXT: renamable $rax = COPY renamable $rcx
+    ; CHECK-NEXT: renamable $eax = COPY killed renamable $esi
+    ; CHECK-NEXT: RET64 implicit killed $rax, implicit killed $rcx
+    %0:gr64 = COPY $rdi
+    %1:gr32 = COPY $esi
+    %2:gr64 = INSERT_SUBREG %0, %1, %subreg.sub_32bit
+    $rax = COPY %2
+    $rcx = COPY %0
+    RET64 implicit $rax, implicit $rcx
+...
+---
+# %0 is read again below, so the tied use is copied into the def's register.
+# The instruction's other read of %0 takes that register too.
+name: tied_use_read_twice
+tracksRegLiveness: true
+body: |
+  bb.0:
+    liveins: $edi
+    ; CHECK-LABEL: name: tied_use_read_twice
+    ; CHECK: liveins: $edi
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: renamable $ecx = COPY killed $edi
+    ; CHECK-NEXT: $eax = COPY $ecx
+    ; CHECK-NEXT: renamable $eax = ADD32rr renamable $eax, renamable $eax, implicit-def dead $eflags
+    ; CHECK-NEXT: RET64 implicit killed $eax, implicit killed $ecx
+    %0:gr32 = COPY $edi
+    %1:gr32 = ADD32rr %0, %0, implicit-def dead $eflags
+    $eax = COPY %1
+    $ecx = COPY %0
+    RET64 implicit $eax, implicit $ecx
+...
+---
+# An undef tied use has no value to place; it just names the tied def's
+# register.
+name: tied_use_undef
+tracksRegLiveness: true
+body: |
+  bb.0:
+    liveins: $edi
+    ; CHECK-LABEL: name: tied_use_undef
+    ; CHECK: liveins: $edi
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: renamable $eax = COPY killed $edi
+    ; CHECK-NEXT: renamable $ecx = ADD32rr undef renamable $ecx, renamable $eax, implicit-def dead $eflags
+    ; CHECK-NEXT: RET64 implicit killed $ecx, implicit killed $eax
+    %0:gr32 = COPY $edi
+    %1:gr32 = ADD32rr undef %2:gr32, %0, implicit-def dead $eflags
+    $ecx = COPY %1
+    $eax = COPY %0
+    RET64 implicit $ecx, implicit $eax
+...
+---
+# Both tied uses read %0, which is live below, so each is copied into its own
+# tied def's register and neither may follow the other's copy.
+name: two_tied_uses_one_value
+tracksRegLiveness: true
+body: |
+  bb.0:
+    liveins: $edi
+    ; CHECK-LABEL: name: two_tied_uses_one_value
+    ; CHECK: liveins: $edi
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: renamable $edx = COPY killed $edi
+    ; CHECK-NEXT: $eax = COPY $edx
+    ; CHECK-NEXT: $ecx = COPY $edx
+    ; CHECK-NEXT: INLINEASM &"// $0 $1 $2 $3", attdialect, regdef:GR32, def renamable $eax, regdef:GR32, def renamable $ecx, reguse tiedto:$0, renamable $eax(tied-def 3), reguse tiedto:$1, renamable $ecx(tied-def 5)
+    ; CHECK-NEXT: RET64 implicit killed $eax, implicit killed $ecx, implicit killed $edx
+    %0:gr32 = COPY $edi
+    INLINEASM &"// $0 $1 $2 $3", attdialect, regdef:GR32, def %1:gr32, regdef:GR32, def %2:gr32, reguse tiedto:$0, %0(tied-def 3), reguse tiedto:$1, %0(tied-def 5)
+    $eax = COPY %1
+    $ecx = COPY %2
+    $edx = COPY %0
+    RET64 implicit $eax, implicit $ecx, implicit $edx
+...



More information about the llvm-commits mailing list