[llvm-branch-commits] [llvm] [RegAllocFast] Fold foldable inline asm operands under register pressure (PR #229630)

Bill Wendling via llvm-branch-commits llvm-branch-commits at lists.llvm.org
Tue Oct 6 19:12:58 PDT 2026


https://github.com/isanbard updated https://github.com/llvm/llvm-project/pull/229630

>From eecaf708a9ac5d0ee6cb095b061dcf539688f7b2 Mon Sep 17 00:00:00 2001
From: Bill Wendling <isanbard at gmail.com>
Date: Tue, 6 Oct 2026 06:12:08 -0700
Subject: [PATCH 1/2] [CodeGen] Report an error for a direct inline asm output
 in memory

An inline asm output returned by value has no memory to write to, yet a
constraint such as "=rm" picks memory, the most general constraint, as
does "=m". SelectionDAG asserted on that ("Can only indirectify direct
input operands!"), and GlobalISel dereferenced a null pointer. Clang
never emits such an output, since it passes the address of a memory
output, but other IR can. Report "cannot handle direct memory outputs
yet for constraint 'm'" instead, like the other inline asm errors there.

Assisted-by: Claude Opus 5.5

>From 64677f9bd83580906e64640abd4d5c7bde3af0ac Mon Sep 17 00:00:00 2001
From: Bill Wendling <isanbard at gmail.com>
Date: Tue, 6 Oct 2026 06:14:26 -0700
Subject: [PATCH 2/2] [RegAllocFast] Fold foldable inline asm operands under
 register pressure

An inline asm register operand marked foldable (from an "rm" constraint)
may be replaced with a stack slot when the register allocator runs out
of registers. The greedy allocator does that when it spills the value.
The fast allocator can't: it assigns the operands of an instruction one
at a time, and folding replaces the instruction. So it reported
"inline assembly requires more registers than available" instead.

Before allocating a block, estimate from each inline asm's own operands
whether they fit in registers, and fold as many foldable registers as
needed to make them fit, so the common case without pressure still gets
a register. Values that only live across the asm don't count, since the
allocator spills them when it needs their registers. The estimate
follows allocateInstruction():

- First all defs get distinct registers, avoiding physreg defs; then all
  uses, along with the defs still occupied while the uses are read
  (early-clobber and tied defs, see isLiveThroughDef()), get distinct
  registers, avoiding physreg uses and early-clobber physreg defs.
- A register counts against another register class as many times as one
  of its registers overlaps that class's registers, so e.g. 32-bit "rm"
  operands compete with 64-bit operands for the same registers.
- A value used by both an "rm" and an "r" operand isn't folded, since
  that frees no register.
- An input tied to a def keeps a register of its own when the allocator
  lowers the tie itself and the input lives on.

Folding a register folds all of its operands at once, with one store
before the asm and one reload after it. It uses the register's own spill
slot, so that a block that reloads the register, such as an
INLINEASM_BR's indirect target, which the reload after the asm doesn't
reach, finds the value the asm wrote.

The whole thing is skipped for functions without inline asm.

Nothing in tree marks an inline asm register operand foldable yet, so
the tests set the "foldable" flag in MIR.

This follows the shape of the unmerged #74344, "[RegAllocFast] fold
foldable inline asm", with a pressure estimate instead of folding
unconditionally, per the review there.

Assisted-by: Claude Opus 5.5
---
 llvm/lib/CodeGen/RegAllocFast.cpp             | 280 ++++++++++
 .../regallocfast-inline-asm-fold-pressure.mir | 479 ++++++++++++++++++
 .../X86/regallocfast-inline-asm-fold.mir      |  58 +++
 3 files changed, 817 insertions(+)
 create mode 100644 llvm/test/CodeGen/X86/regallocfast-inline-asm-fold-pressure.mir
 create mode 100644 llvm/test/CodeGen/X86/regallocfast-inline-asm-fold.mir

diff --git a/llvm/lib/CodeGen/RegAllocFast.cpp b/llvm/lib/CodeGen/RegAllocFast.cpp
index dfdbe7d4d2121..112e336a99b7f 100644
--- a/llvm/lib/CodeGen/RegAllocFast.cpp
+++ b/llvm/lib/CodeGen/RegAllocFast.cpp
@@ -20,6 +20,10 @@
 /// Where the target enables it, TwoAddressInstructionPass is left out of the
 /// pipeline: this pass lowers tied operands and expands REG_SEQUENCE and
 /// INSERT_SUBREG itself.
+///
+/// An inline asm register operand that may be folded to memory (from an "rm"
+/// constraint) is folded to a stack slot before its block is allocated, if
+/// the asm's register operands wouldn't fit in registers otherwise.
 //
 //===----------------------------------------------------------------------===//
 
@@ -64,6 +68,8 @@ using namespace llvm;
 STATISTIC(NumStores, "Number of stores added");
 STATISTIC(NumLoads, "Number of loads added");
 STATISTIC(NumCoalesced, "Number of copies coalesced");
+STATISTIC(NumInlineAsmFolds,
+          "Number of inline asm registers folded to stack slots");
 
 static RegisterRegAlloc fastRegAlloc("fast", "fast register allocator",
                                      createFastRegisterAllocator);
@@ -431,6 +437,11 @@ class RegAllocFastImpl {
 
   bool mayBeSpillFromInlineAsmBr(const MachineInstr &MI) const;
 
+  void foldInlineAsmOperands(MachineBasicBlock &MBB);
+  void selectInlineAsmRegsToFold(const MachineInstr &MI,
+                                 SmallVectorImpl<Register> &ToFold) const;
+  void foldInlineAsmReg(MachineInstr *&MI, Register Reg);
+
   void dumpState() const;
 };
 
@@ -1956,6 +1967,270 @@ void RegAllocFastImpl::handleBundle(MachineInstr &MI) {
   }
 }
 
+/// Choose which foldable ("rm") registers of the inline asm \p MI to fold to
+/// stack slots so that the rest of its register operands fit, and add them to
+/// \p ToFold.
+///
+/// The greedy allocator folds such an operand when it runs out of registers.
+/// This allocator can't: it assigns the operands of an instruction one at a
+/// time, and folding replaces the instruction. So estimate up front, from the
+/// asm's own operands, whether they fit. Values that only live across the asm
+/// don't count, since this allocator spills them when it needs their
+/// registers. The estimate follows allocateInstruction(): first all defs get
+/// distinct registers, avoiding physreg defs; then all uses, along with the
+/// defs still occupied while the uses are read (early-clobber and tied defs),
+/// get distinct registers, avoiding physreg uses and early-clobber physreg
+/// defs. Folding a register removes it from both.
+///
+/// A register counts against another register class as many times as one of
+/// its registers overlaps registers of that class, e.g. once for GR64 against
+/// GR32. That, and keeping physreg uses out of the defs' registers, errs
+/// toward folding more than necessary rather than running out of registers.
+void RegAllocFastImpl::selectInlineAsmRegsToFold(
+    const MachineInstr &MI, SmallVectorImpl<Register> &ToFold) const {
+  // The virtual registers of MI's register operands.
+  struct AsmReg {
+    Register Reg;
+    const TargetRegisterClass *RC;
+    // Needs a register while the defs are assigned.
+    bool InDefs = false;
+    // Needs a register while the uses are read.
+    bool InUses = false;
+    // Inputs tied to this def that keep registers of their own, to be copied
+    // from (see lowerTiedUse()). Folding the def folds them too.
+    unsigned TiedInputs = 0;
+    // Read by an operand that isn't tied to its def.
+    bool HasUntiedUse = false;
+    // Every untied operand of the register can be folded.
+    bool Foldable = true;
+    bool Folded = false;
+  };
+  SmallVector<AsmReg, 8> Regs;
+  auto getAsmReg = [&](Register Reg) -> AsmReg & {
+    for (AsmReg &R : Regs)
+      if (R.Reg == Reg)
+        return R;
+
+    return Regs.emplace_back(AsmReg{Reg, MRI->getRegClass(Reg)});
+  };
+
+  // Physical registers taken while the defs, and while the uses, are assigned.
+  SmallVector<MCRegister, 8> DefTaken;
+  SmallVector<MCRegister, 8> UseTaken;
+
+  for (unsigned I = InlineAsm::MIOp_FirstOperand, E = MI.getNumOperands();
+       I != E; ++I) {
+    const MachineOperand &MO = MI.getOperand(I);
+    if (!MO.isReg() || !MO.getReg() || (MO.isUse() && MO.isUndef()))
+      continue;
+
+    Register Reg = MO.getReg();
+    if (Reg.isPhysical()) {
+      DefTaken.push_back(Reg);
+      if (MO.isUse() || isLiveThroughDef(MI, MO))
+        UseTaken.push_back(Reg);
+
+      continue;
+    }
+
+    if (!shouldAllocateRegister(Reg))
+      continue;
+
+    // A tied use is read through its def's register. Unless that is its own
+    // register (once tied operands have been rewritten), it can also need a
+    // register of its own to copy from: if it doesn't die here, and whenever
+    // the def is a fixed register, which might not be allocatable.
+    if (MO.isUse() && MO.isTied()) {
+      Register DefReg = MI.getOperand(MI.findTiedOperandIdx(I)).getReg();
+      if (Reg == DefReg)
+        continue;
+
+      if (DefReg.isVirtual()) {
+        if (!MO.isKill())
+          ++getAsmReg(DefReg).TiedInputs;
+      } else {
+        AsmReg &R = getAsmReg(Reg);
+        R.InUses = true;
+        R.Foldable = false;
+      }
+
+      continue;
+    }
+
+    AsmReg &R = getAsmReg(Reg);
+    R.Foldable &= MI.mayFoldInlineAsmRegOp(I);
+    if (MO.isUse()) {
+      R.InUses = true;
+      R.HasUntiedUse = true;
+    } else {
+      R.InDefs = true;
+      R.InUses |= isLiveThroughDef(MI, MO);
+    }
+  }
+
+  // The asm's own read of a register it also writes can't share the def's
+  // stack slot unless that's how the two are tied, so only fold a register
+  // the asm either reads or writes.
+  for (AsmReg &R : Regs)
+    R.Foldable &= !(R.InDefs && R.HasUntiedUse);
+
+  // Fold within one register class at a time, in operand order. Registers
+  // folded for one class no longer count against the next.
+  SmallVector<const TargetRegisterClass *, 4> Done;
+  for (const AsmReg &Candidate : Regs) {
+    const TargetRegisterClass *RC = Candidate.RC;
+    if (!Candidate.Foldable || is_contained(Done, RC))
+      continue;
+
+    Done.push_back(RC);
+
+    ArrayRef<MCPhysReg> Order = RegClassInfo.getOrder(RC);
+    auto countFree = [&](ArrayRef<MCRegister> Taken) -> unsigned {
+      return count_if(Order, [&](MCPhysReg PhysReg) {
+        return none_of(
+            Taken, [&](MCRegister T) { return TRI->regsOverlap(PhysReg, T); });
+      });
+    };
+    unsigned DefsFree = countFree(DefTaken);
+    unsigned UsesFree = countFree(UseTaken);
+
+    // How many registers of RC one register of each AsmReg can take.
+    SmallVector<unsigned, 8> Weight;
+    for (const AsmReg &R : Regs) {
+      if (RC->hasSubClassEq(R.RC) || R.RC->hasSubClassEq(RC)) {
+        Weight.push_back(1);
+        continue;
+      }
+
+      unsigned Max = 0;
+      for (MCPhysReg PhysReg : RegClassInfo.getOrder(R.RC))
+        Max = std::max<unsigned>(Max, count_if(Order, [&](MCPhysReg Other) {
+                                   return TRI->regsOverlap(PhysReg, Other);
+                                 }));
+
+      Weight.push_back(Max);
+    }
+
+    while (true) {
+      unsigned DefsNeeded = 0;
+      unsigned UsesNeeded = 0;
+      for (auto [R, W] : zip(Regs, Weight)) {
+        if (R.Folded)
+          continue;
+
+        DefsNeeded += R.InDefs * W;
+        UsesNeeded += (R.InUses + R.TiedInputs) * W;
+      }
+
+      bool DefsShort = DefsNeeded > DefsFree;
+      bool UsesShort = UsesNeeded > UsesFree;
+      if (!DefsShort && !UsesShort)
+        break;
+
+      // Fold the first register that relieves the most phases that are short.
+      AsmReg *Best = nullptr;
+      unsigned BestRelief = 0;
+      for (auto [R, W] : zip(Regs, Weight)) {
+        if (R.Folded || !R.Foldable || !W)
+          continue;
+
+        unsigned Relief =
+            (DefsShort && R.InDefs) + (UsesShort && (R.InUses || R.TiedInputs));
+        if (Relief > BestRelief) {
+          Best = &R;
+          BestRelief = Relief;
+        }
+      }
+
+      // Without anything left to fold, allocation reports the error.
+      if (!Best)
+        break;
+
+      Best->Folded = true;
+      ToFold.push_back(Best->Reg);
+    }
+  }
+}
+
+/// Fold the operands of \p Reg in the inline asm \p MI to Reg's stack slot,
+/// storing the value the asm reads there before it, and reloading Reg from it
+/// afterward if the asm writes Reg. \p MI is replaced by the new instruction.
+void RegAllocFastImpl::foldInlineAsmReg(MachineInstr *&MI, Register Reg) {
+  // Fold every operand of Reg. A tied use goes along with its def, and reads
+  // the def's slot, which therefore needs the tied input's value, whichever
+  // register that is.
+  SmallVector<unsigned, 2> Ops;
+  Register ReadReg;
+  bool Writes = false;
+  for (unsigned I = InlineAsm::MIOp_FirstOperand, E = MI->getNumOperands();
+       I != E; ++I) {
+    const MachineOperand &MO = MI->getOperand(I);
+    if (!MO.isReg() || MO.getReg() != Reg)
+      continue;
+
+    if (MO.isUse()) {
+      if (!MO.isTied())
+        Ops.push_back(I);
+
+      if (MO.readsReg())
+        ReadReg = Reg;
+
+      continue;
+    }
+
+    Ops.push_back(I);
+    Writes |= !MO.isDead();
+    if (MO.isTied()) {
+      const MachineOperand &Use = MI->getOperand(MI->findTiedOperandIdx(I));
+      if (Use.readsReg())
+        ReadReg = Use.getReg();
+    }
+  }
+
+  // Use Reg's own spill slot: a block that reloads Reg, such as an
+  // INLINEASM_BR's indirect target, which the reload after the asm doesn't
+  // reach, then finds the value the asm wrote.
+  MachineBasicBlock &MBB = *MI->getParent();
+  const TargetRegisterClass *RC = MRI->getRegClass(Reg);
+  int FI = getStackSpaceFor(Reg);
+  MachineInstr *CopyMI = nullptr;
+  MachineInstr *NewMI = TII->foldMemoryOperand(*MI, Ops, FI, CopyMI);
+  assert(NewMI && !CopyMI && "inline asm register should fold");
+
+  // NewMI is inserted before MI. The allocator works out the kill flags.
+  if (ReadReg) {
+    TII->storeRegToStackSlot(MBB, NewMI->getIterator(), ReadReg,
+                             /*isKill=*/false, FI, RC, ReadReg);
+    ++NumStores;
+  }
+
+  if (Writes) {
+    TII->loadRegFromStackSlot(MBB, std::next(NewMI->getIterator()), Reg, FI, RC,
+                              Reg);
+    ++NumLoads;
+  }
+
+  ++NumInlineAsmFolds;
+
+  MI->eraseFromParent();
+  MI = NewMI;
+}
+
+/// Fold the operands of each inline asm in \p MBB that won't fit in registers
+/// (see selectInlineAsmRegsToFold()).
+void RegAllocFastImpl::foldInlineAsmOperands(MachineBasicBlock &MBB) {
+  for (MachineInstr &MI : make_early_inc_range(MBB)) {
+    if (!MI.isInlineAsm())
+      continue;
+
+    SmallVector<Register, 4> ToFold;
+    selectInlineAsmRegsToFold(MI, ToFold);
+    MachineInstr *AsmMI = &MI;
+    for (Register Reg : ToFold)
+      foldInlineAsmReg(AsmMI, Reg);
+  }
+}
+
 void RegAllocFastImpl::allocateBasicBlock(MachineBasicBlock &MBB) {
   this->MBB = &MBB;
   LLVM_DEBUG(dbgs() << "\nAllocating " << MBB);
@@ -1969,6 +2244,11 @@ void RegAllocFastImpl::allocateBasicBlock(MachineBasicBlock &MBB) {
 
   Coalesced.clear();
 
+  // Folding replaces an inline asm, which allocateInstruction() can't do
+  // partway through, so fold the operands that won't fit in registers first.
+  if (MBB.getParent()->hasInlineAsm())
+    foldInlineAsmOperands(MBB);
+
   // Lowering a tied operand inserts a copy ahead of MI. Its registers are
   // already assigned, so visiting it would evict what still lives in the
   // source.
diff --git a/llvm/test/CodeGen/X86/regallocfast-inline-asm-fold-pressure.mir b/llvm/test/CodeGen/X86/regallocfast-inline-asm-fold-pressure.mir
new file mode 100644
index 0000000000000..87f7239f56bdb
--- /dev/null
+++ b/llvm/test/CodeGen/X86/regallocfast-inline-asm-fold-pressure.mir
@@ -0,0 +1,479 @@
+# NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
+# RUN: llc -mtriple=x86_64-unknown-linux-gnu -run-pass=regallocfast \
+# RUN:   -verify-machineinstrs -verify-regalloc %s -o - | FileCheck %s
+
+# The fast register allocator folds foldable ("rm") inline asm operands to
+# stack slots when the asm's register operands wouldn't fit in registers
+# otherwise, and only then.
+
+---
+# The asm clobbers rax/rbx/rbp/r10-r15, leaving six of the 15 GPRs for its
+# register operands, and its six "r" inputs take all of them. The "=&rm"
+# output can't share a register with an input, so it's folded to a stack slot.
+name:            test_rm_output_pressure
+tracksRegLiveness: true
+registers:
+  - { id: 0, class: gr64 }
+  - { id: 1, class: gr64 }
+  - { id: 2, class: gr64 }
+  - { id: 3, class: gr64 }
+  - { id: 4, class: gr64 }
+  - { id: 5, class: gr64 }
+  - { id: 6, class: gr64_norex2 }
+  - { id: 7, class: gr64_norex2 }
+  - { id: 8, class: gr64_norex2 }
+  - { id: 9, class: gr64_norex2 }
+  - { id: 10, class: gr64_norex2 }
+  - { id: 11, class: gr64_norex2 }
+  - { id: 12, class: gr64_norex2 }
+body:             |
+  bb.0:
+    liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+
+    ; CHECK-LABEL: name: test_rm_output_pressure
+    ; CHECK: liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: INLINEASM &"# $0 $1 $2 $3 $4 $5 $6", sideeffect maystore attdialect, mem:m, %stack.0, 1, $noreg, 0, $noreg, reguse:GR64_NOREX2, killed renamable $rdi, reguse:GR64_NOREX2, killed renamable $rsi, reguse:GR64_NOREX2, killed renamable $rdx, reguse:GR64_NOREX2, killed renamable $rcx, reguse:GR64_NOREX2, killed renamable $r8, reguse:GR64_NOREX2, killed renamable $r9, clobber, implicit-def dead early-clobber $rax, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r10, clobber, implicit-def dead early-clobber $r11, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14, clobber, implicit-def dead early-clobber $r15 :: (store (s64) into %stack.0)
+    ; CHECK-NEXT: renamable $rax = MOV64rm %stack.0, 1, $noreg, 0, $noreg :: (load (s64) from %stack.0)
+    ; CHECK-NEXT: RET 0, killed $rax
+    %5:gr64 = COPY $r9
+    %4:gr64 = COPY $r8
+    %3:gr64 = COPY $rcx
+    %2:gr64 = COPY $rdx
+    %1:gr64 = COPY $rsi
+    %0:gr64 = COPY $rdi
+    %7:gr64_norex2 = COPY %0
+    %8:gr64_norex2 = COPY %1
+    %9:gr64_norex2 = COPY %2
+    %10:gr64_norex2 = COPY %3
+    %11:gr64_norex2 = COPY %4
+    %12:gr64_norex2 = COPY %5
+    INLINEASM &"# $0 $1 $2 $3 $4 $5 $6", sideeffect attdialect, regdef-ec:GR64_NOREX2 foldable, def early-clobber %6, reguse:GR64_NOREX2, %7, reguse:GR64_NOREX2, %8, reguse:GR64_NOREX2, %9, reguse:GR64_NOREX2, %10, reguse:GR64_NOREX2, %11, reguse:GR64_NOREX2, %12, clobber, implicit-def dead early-clobber $rax, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r10, clobber, implicit-def dead early-clobber $r11, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14, clobber, implicit-def dead early-clobber $r15
+    $rax = COPY %6
+    RET 0, $rax
+...
+---
+# One value passed to two "rm" operands: spilling it means folding both
+# operands into its stack slot at once.
+name:            test_rm_same_value_twice_pressure
+tracksRegLiveness: true
+registers:
+  - { id: 0, class: gr64 }
+  - { id: 1, class: gr64 }
+  - { id: 2, class: gr64 }
+  - { id: 3, class: gr64 }
+  - { id: 4, class: gr64 }
+  - { id: 5, class: gr64 }
+  - { id: 6, class: gr64_norex2 }
+  - { id: 7, class: gr64_norex2 }
+  - { id: 8, class: gr64_norex2 }
+  - { id: 9, class: gr64_norex2 }
+  - { id: 10, class: gr64_norex2 }
+  - { id: 11, class: gr64_norex2 }
+  - { id: 12, class: gr64_norex2 }
+  - { id: 13, class: gr64_norex2 }
+  - { id: 14, class: gr64 }
+fixedStack:
+  - { id: 0, size: 8, alignment: 16, isImmutable: true }
+body:             |
+  bb.0:
+    liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+
+    ; CHECK-LABEL: name: test_rm_same_value_twice_pressure
+    ; CHECK: liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: renamable $rax = COPY killed $rdi
+    ; CHECK-NEXT: renamable $rdi = MOV64rm %fixed-stack.0, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.0, align 16)
+    ; CHECK-NEXT: MOV64mr %stack.0, 1, $noreg, 0, $noreg, killed renamable $rax :: (store (s64) into %stack.0)
+    ; CHECK-NEXT: INLINEASM &"# $0 $1 $2 $3 $4 $5 $6 $7", sideeffect mayload attdialect, mem:m, %stack.0, 1, $noreg, 0, $noreg, mem:m, %stack.0, 1, $noreg, 0, $noreg, reguse:GR64_NOREX2, killed renamable $rsi, reguse:GR64_NOREX2, killed renamable $rdx, reguse:GR64_NOREX2, killed renamable $rcx, reguse:GR64_NOREX2, killed renamable $r8, reguse:GR64_NOREX2, killed renamable $r9, reguse:GR64_NOREX2, killed renamable $rdi, clobber, implicit-def dead early-clobber $rax, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r10, clobber, implicit-def dead early-clobber $r11, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14, clobber, implicit-def dead early-clobber $r15 :: (load (s64) from %stack.0)
+    ; CHECK-NEXT: RET 0
+    %5:gr64 = COPY $r9
+    %4:gr64 = COPY $r8
+    %3:gr64 = COPY $rcx
+    %2:gr64 = COPY $rdx
+    %1:gr64 = COPY $rsi
+    %0:gr64 = COPY $rdi
+    %14:gr64 = MOV64rm %fixed-stack.0, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.0, align 16)
+    %6:gr64_norex2 = COPY %0
+    %8:gr64_norex2 = COPY %1
+    %9:gr64_norex2 = COPY %2
+    %10:gr64_norex2 = COPY %3
+    %11:gr64_norex2 = COPY %4
+    %12:gr64_norex2 = COPY %5
+    %13:gr64_norex2 = COPY %14
+    INLINEASM &"# $0 $1 $2 $3 $4 $5 $6 $7", sideeffect attdialect, reguse:GR64_NOREX2 foldable, %6, reguse:GR64_NOREX2 foldable, %6, reguse:GR64_NOREX2, %8, reguse:GR64_NOREX2, %9, reguse:GR64_NOREX2, %10, reguse:GR64_NOREX2, %11, reguse:GR64_NOREX2, %12, reguse:GR64_NOREX2, %13, clobber, implicit-def dead early-clobber $rax, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r10, clobber, implicit-def dead early-clobber $r11, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14, clobber, implicit-def dead early-clobber $r15
+    RET 0
+...
+---
+# Folding the "rm" input moves the operands after it, including the input
+# tied to the "=r" output, whose tie must survive.
+name:            test_rm_before_tied_pressure
+tracksRegLiveness: true
+registers:
+  - { id: 0, class: gr64 }
+  - { id: 1, class: gr64 }
+  - { id: 2, class: gr64 }
+  - { id: 3, class: gr64 }
+  - { id: 4, class: gr64 }
+  - { id: 5, class: gr64 }
+  - { id: 6, class: gr64_norex2 }
+  - { id: 7, class: gr64_norex2 }
+  - { id: 8, class: gr64_norex2 }
+  - { id: 9, class: gr64_norex2 }
+  - { id: 10, class: gr64_norex2 }
+  - { id: 11, class: gr64_norex2 }
+  - { id: 12, class: gr64_norex2 }
+  - { id: 13, class: gr64_norex2 }
+  - { id: 14, class: gr64 }
+fixedStack:
+  - { id: 0, size: 8, alignment: 16, isImmutable: true }
+body:             |
+  bb.0:
+    liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+
+    ; CHECK-LABEL: name: test_rm_before_tied_pressure
+    ; CHECK: liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: MOV64mr %stack.2, 1, $noreg, 0, $noreg, $rcx :: (store (s64) into %stack.2)
+    ; CHECK-NEXT: renamable $rax = COPY $rdi
+    ; CHECK-NEXT: $rdi = MOV64rm %stack.2, 1, $noreg, 0, $noreg :: (load (s64) from %stack.2)
+    ; CHECK-NEXT: renamable $rcx = MOV64rm %fixed-stack.0, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.0, align 16)
+    ; CHECK-NEXT: MOV64mr %stack.0, 1, $noreg, 0, $noreg, killed renamable $rax :: (store (s64) into %stack.0)
+    ; CHECK-NEXT: INLINEASM &"# $0 $1 $2 $3 $4 $5 $6 $7", sideeffect mayload attdialect, regdef:GR64_NOREX2, def renamable $rcx, mem:m, %stack.0, 1, $noreg, 0, $noreg, reguse:GR64_NOREX2, killed renamable $rsi, reguse:GR64_NOREX2, killed renamable $rdx, reguse:GR64_NOREX2, killed renamable $rdi, reguse:GR64_NOREX2, killed renamable $r8, reguse:GR64_NOREX2, killed renamable $r9, reguse tiedto:$0, renamable $rcx(tied-def 3), clobber, implicit-def dead early-clobber $rax, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r10, clobber, implicit-def dead early-clobber $r11, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14, clobber, implicit-def dead early-clobber $r15 :: (load (s64) from %stack.0)
+    ; CHECK-NEXT: MOV64mr %stack.1, 1, $noreg, 0, $noreg, $rcx :: (store (s64) into %stack.1)
+    ; CHECK-NEXT: $rax = MOV64rm %stack.1, 1, $noreg, 0, $noreg :: (load (s64) from %stack.1)
+    ; CHECK-NEXT: RET 0, killed $rax
+    %5:gr64 = COPY $r9
+    %4:gr64 = COPY $r8
+    %3:gr64 = COPY $rcx
+    %2:gr64 = COPY $rdx
+    %1:gr64 = COPY $rsi
+    %0:gr64 = COPY $rdi
+    %14:gr64 = MOV64rm %fixed-stack.0, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.0, align 16)
+    %7:gr64_norex2 = COPY %0
+    %8:gr64_norex2 = COPY %1
+    %9:gr64_norex2 = COPY %2
+    %10:gr64_norex2 = COPY %3
+    %11:gr64_norex2 = COPY %4
+    %12:gr64_norex2 = COPY %5
+    %13:gr64_norex2 = COPY %14
+    %6:gr64_norex2 = COPY %13
+    INLINEASM &"# $0 $1 $2 $3 $4 $5 $6 $7", sideeffect attdialect, regdef:GR64_NOREX2, def %6, reguse:GR64_NOREX2 foldable, %7, reguse:GR64_NOREX2, %8, reguse:GR64_NOREX2, %9, reguse:GR64_NOREX2, %10, reguse:GR64_NOREX2, %11, reguse:GR64_NOREX2, %12, reguse tiedto:$0, %6(tied-def 3), clobber, implicit-def dead early-clobber $rax, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r10, clobber, implicit-def dead early-clobber $r11, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14, clobber, implicit-def dead early-clobber $r15
+    $rax = COPY %6
+    RET 0, $rax
+...
+---
+# Two read-write "+rm" operands: folding one tied pair must keep the other
+# pair tied.
+name:            test_two_tied_rm_pressure
+tracksRegLiveness: true
+registers:
+  - { id: 0, class: gr64 }
+  - { id: 1, class: gr64 }
+  - { id: 2, class: gr64 }
+  - { id: 3, class: gr64 }
+  - { id: 4, class: gr64 }
+  - { id: 5, class: gr64 }
+  - { id: 6, class: gr64_norex2 }
+  - { id: 7, class: gr64_norex2 }
+  - { id: 8, class: gr64_norex2 }
+  - { id: 9, class: gr64_norex2 }
+  - { id: 10, class: gr64_norex2 }
+  - { id: 11, class: gr64_norex2 }
+  - { id: 12, class: gr64_norex2 }
+  - { id: 13, class: gr64_norex2 }
+  - { id: 14, class: gr64_norex2 }
+  - { id: 15, class: gr64 }
+fixedStack:
+  - { id: 0, size: 8, alignment: 16, isImmutable: true }
+body:             |
+  bb.0:
+    liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+
+    ; CHECK-LABEL: name: test_two_tied_rm_pressure
+    ; CHECK: liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: MOV64mr %stack.1, 1, $noreg, 0, $noreg, $rcx :: (store (s64) into %stack.1)
+    ; CHECK-NEXT: renamable $rcx = COPY killed $rdx
+    ; CHECK-NEXT: renamable $rdx = COPY $rsi
+    ; CHECK-NEXT: $rsi = MOV64rm %stack.1, 1, $noreg, 0, $noreg :: (load (s64) from %stack.1)
+    ; CHECK-NEXT: renamable $rax = COPY killed $rdi
+    ; CHECK-NEXT: renamable $rdi = MOV64rm %fixed-stack.0, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.0, align 16)
+    ; CHECK-NEXT: MOV64mr %stack.0, 1, $noreg, 0, $noreg, renamable $rax :: (store (s64) into %stack.0)
+    ; CHECK-NEXT: INLINEASM &"# $0 $1 $2 $3 $4 $5 $6", sideeffect mayload maystore attdialect, mem:m, %stack.0, 1, $noreg, 0, $noreg, regdef:GR64_NOREX2 foldable, def renamable $rdx, reguse:GR64_NOREX2, killed renamable $rcx, reguse:GR64_NOREX2, killed renamable $rsi, reguse:GR64_NOREX2, killed renamable $r8, reguse:GR64_NOREX2, killed renamable $r9, reguse:GR64_NOREX2, killed renamable $rdi, mem:m, %stack.0, 1, $noreg, 0, $noreg, reguse tiedto:$1, renamable $rdx(tied-def 9), clobber, implicit-def dead early-clobber $rax, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r10, clobber, implicit-def dead early-clobber $r11, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14, clobber, implicit-def dead early-clobber $r15 :: (load store (s64) on %stack.0)
+    ; CHECK-NEXT: renamable $rax = MOV64rm %stack.0, 1, $noreg, 0, $noreg :: (load (s64) from %stack.0)
+    ; CHECK-NEXT: RET 0, killed $rax, killed $rdx
+    %5:gr64 = COPY $r9
+    %4:gr64 = COPY $r8
+    %3:gr64 = COPY $rcx
+    %2:gr64 = COPY $rdx
+    %1:gr64 = COPY $rsi
+    %0:gr64 = COPY $rdi
+    %15:gr64 = MOV64rm %fixed-stack.0, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.0, align 16)
+    %8:gr64_norex2 = COPY %2
+    %9:gr64_norex2 = COPY %3
+    %10:gr64_norex2 = COPY %4
+    %11:gr64_norex2 = COPY %5
+    %12:gr64_norex2 = COPY %15
+    %13:gr64_norex2 = COPY %0
+    %14:gr64_norex2 = COPY %1
+    %6:gr64_norex2 = COPY %13
+    %7:gr64_norex2 = COPY %14
+    INLINEASM &"# $0 $1 $2 $3 $4 $5 $6", sideeffect attdialect, regdef:GR64_NOREX2 foldable, def %6, regdef:GR64_NOREX2 foldable, def %7, reguse:GR64_NOREX2, %8, reguse:GR64_NOREX2, %9, reguse:GR64_NOREX2, %10, reguse:GR64_NOREX2, %11, reguse:GR64_NOREX2, %12, reguse tiedto:$0, %6(tied-def 3), reguse tiedto:$1, %7(tied-def 5), clobber, implicit-def dead early-clobber $rax, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r10, clobber, implicit-def dead early-clobber $r11, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14, clobber, implicit-def dead early-clobber $r15
+    $rax = COPY %6
+    $rdx = COPY %7
+    RET 0, $rax, $rdx
+...
+---
+# Like test_rm_output_pressure, but the output isn't early-clobber, so it can
+# share a register with an input and nothing needs to be folded.
+name:            test_rm_output_shares_input_register
+tracksRegLiveness: true
+registers:
+  - { id: 0, class: gr64 }
+  - { id: 1, class: gr64 }
+  - { id: 2, class: gr64 }
+  - { id: 3, class: gr64 }
+  - { id: 4, class: gr64 }
+  - { id: 5, class: gr64 }
+  - { id: 6, class: gr64_norex2 }
+  - { id: 7, class: gr64_norex2 }
+  - { id: 8, class: gr64_norex2 }
+  - { id: 9, class: gr64_norex2 }
+  - { id: 10, class: gr64_norex2 }
+  - { id: 11, class: gr64_norex2 }
+  - { id: 12, class: gr64_norex2 }
+body:             |
+  bb.0:
+    liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+
+    ; CHECK-LABEL: name: test_rm_output_shares_input_register
+    ; CHECK: liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: INLINEASM &"# $0 $1 $2 $3 $4 $5 $6", sideeffect attdialect, regdef:GR64_NOREX2 foldable, def renamable $rcx, reguse:GR64_NOREX2, killed renamable $rdi, reguse:GR64_NOREX2, killed renamable $rsi, reguse:GR64_NOREX2, killed renamable $rdx, reguse:GR64_NOREX2, killed renamable $rcx, reguse:GR64_NOREX2, killed renamable $r8, reguse:GR64_NOREX2, killed renamable $r9, clobber, implicit-def dead early-clobber $rax, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r10, clobber, implicit-def dead early-clobber $r11, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14, clobber, implicit-def dead early-clobber $r15
+    ; CHECK-NEXT: MOV64mr %stack.0, 1, $noreg, 0, $noreg, $rcx :: (store (s64) into %stack.0)
+    ; CHECK-NEXT: $rax = MOV64rm %stack.0, 1, $noreg, 0, $noreg :: (load (s64) from %stack.0)
+    ; CHECK-NEXT: RET 0, killed $rax
+    %5:gr64 = COPY $r9
+    %4:gr64 = COPY $r8
+    %3:gr64 = COPY $rcx
+    %2:gr64 = COPY $rdx
+    %1:gr64 = COPY $rsi
+    %0:gr64 = COPY $rdi
+    %7:gr64_norex2 = COPY %0
+    %8:gr64_norex2 = COPY %1
+    %9:gr64_norex2 = COPY %2
+    %10:gr64_norex2 = COPY %3
+    %11:gr64_norex2 = COPY %4
+    %12:gr64_norex2 = COPY %5
+    INLINEASM &"# $0 $1 $2 $3 $4 $5 $6", sideeffect attdialect, regdef:GR64_NOREX2 foldable, def %6, reguse:GR64_NOREX2, %7, reguse:GR64_NOREX2, %8, reguse:GR64_NOREX2, %9, reguse:GR64_NOREX2, %10, reguse:GR64_NOREX2, %11, reguse:GR64_NOREX2, %12, clobber, implicit-def dead early-clobber $rax, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r10, clobber, implicit-def dead early-clobber $r11, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14, clobber, implicit-def dead early-clobber $r15
+    $rax = COPY %6
+    RET 0, $rax
+...
+---
+# Eleven inputs for the ten GPRs left after the clobbers: one "rm" input has
+# to be folded even though the 32-bit inputs alone would fit, because each
+# 64-bit input takes a 32-bit register too.
+name:            test_rm_mixed_widths_pressure
+tracksRegLiveness: true
+registers:
+  - { id: 0, class: gr32 }
+  - { id: 1, class: gr32 }
+  - { id: 2, class: gr32 }
+  - { id: 3, class: gr32 }
+  - { id: 4, class: gr32 }
+  - { id: 5, class: gr64 }
+  - { id: 6, class: gr32_norex2 }
+  - { id: 7, class: gr32_norex2 }
+  - { id: 8, class: gr32_norex2 }
+  - { id: 9, class: gr32_norex2 }
+  - { id: 10, class: gr32_norex2 }
+  - { id: 11, class: gr64_norex2 }
+  - { id: 12, class: gr64_norex2 }
+  - { id: 13, class: gr64_norex2 }
+  - { id: 14, class: gr64_norex2 }
+  - { id: 15, class: gr64_norex2 }
+  - { id: 16, class: gr64_norex2 }
+  - { id: 17, class: gr64 }
+  - { id: 18, class: gr64 }
+  - { id: 19, class: gr64 }
+  - { id: 20, class: gr64 }
+  - { id: 21, class: gr64 }
+fixedStack:
+  - { id: 0, offset: 32, size: 8, alignment: 16, isImmutable: true }
+  - { id: 1, offset: 24, size: 8, alignment: 8, isImmutable: true }
+  - { id: 2, offset: 16, size: 8, alignment: 16, isImmutable: true }
+  - { id: 3, offset: 8, size: 8, alignment: 8, isImmutable: true }
+  - { id: 4, size: 8, alignment: 16, isImmutable: true }
+body:             |
+  bb.0:
+    liveins: $edi, $esi, $edx, $ecx, $r8d, $r9
+
+    ; CHECK-LABEL: name: test_rm_mixed_widths_pressure
+    ; CHECK: liveins: $edi, $esi, $edx, $ecx, $r8d, $r9
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: renamable $ebx = COPY killed $edi
+    ; CHECK-NEXT: renamable $rax = MOV64rm %fixed-stack.0, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.0, align 16)
+    ; CHECK-NEXT: renamable $rdi = MOV64rm %fixed-stack.1, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.1)
+    ; CHECK-NEXT: renamable $r10 = MOV64rm %fixed-stack.2, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.2, align 16)
+    ; CHECK-NEXT: renamable $r11 = MOV64rm %fixed-stack.3, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.3)
+    ; CHECK-NEXT: renamable $r15 = MOV64rm %fixed-stack.4, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.4, align 16)
+    ; CHECK-NEXT: MOV32mr %stack.0, 1, $noreg, 0, $noreg, killed renamable $ebx :: (store (s32) into %stack.0)
+    ; CHECK-NEXT: INLINEASM &"# $0 $1 $2 $3 $4 $5 $6 $7 $8 $9 $10", sideeffect mayload attdialect, mem:m, %stack.0, 1, $noreg, 0, $noreg, reguse:GR32_NOREX2 foldable, killed renamable $esi, reguse:GR32_NOREX2 foldable, killed renamable $edx, reguse:GR32_NOREX2 foldable, killed renamable $ecx, reguse:GR32_NOREX2 foldable, killed renamable $r8d, reguse:GR64_NOREX2, killed renamable $r9, reguse:GR64_NOREX2, killed renamable $rax, reguse:GR64_NOREX2, killed renamable $rdi, reguse:GR64_NOREX2, killed renamable $r10, reguse:GR64_NOREX2, killed renamable $r11, reguse:GR64_NOREX2, killed renamable $r15, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14 :: (load (s32) from %stack.0)
+    ; CHECK-NEXT: RET 0
+    %5:gr64 = COPY $r9
+    %4:gr32 = COPY $r8d
+    %3:gr32 = COPY $ecx
+    %2:gr32 = COPY $edx
+    %1:gr32 = COPY $esi
+    %0:gr32 = COPY $edi
+    %17:gr64 = MOV64rm %fixed-stack.4, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.4, align 16)
+    %18:gr64 = MOV64rm %fixed-stack.3, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.3)
+    %19:gr64 = MOV64rm %fixed-stack.2, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.2, align 16)
+    %20:gr64 = MOV64rm %fixed-stack.1, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.1)
+    %21:gr64 = MOV64rm %fixed-stack.0, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.0, align 16)
+    %6:gr32_norex2 = COPY %0
+    %7:gr32_norex2 = COPY %1
+    %8:gr32_norex2 = COPY %2
+    %9:gr32_norex2 = COPY %3
+    %10:gr32_norex2 = COPY %4
+    %11:gr64_norex2 = COPY %5
+    %12:gr64_norex2 = COPY %17
+    %13:gr64_norex2 = COPY %18
+    %14:gr64_norex2 = COPY %19
+    %15:gr64_norex2 = COPY %20
+    %16:gr64_norex2 = COPY %21
+    INLINEASM &"# $0 $1 $2 $3 $4 $5 $6 $7 $8 $9 $10", sideeffect attdialect, reguse:GR32_NOREX2 foldable, %6, reguse:GR32_NOREX2 foldable, %7, reguse:GR32_NOREX2 foldable, %8, reguse:GR32_NOREX2 foldable, %9, reguse:GR32_NOREX2 foldable, %10, reguse:GR64_NOREX2, %11, reguse:GR64_NOREX2, %12, reguse:GR64_NOREX2, %13, reguse:GR64_NOREX2, %14, reguse:GR64_NOREX2, %15, reguse:GR64_NOREX2, %16, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14
+    RET 0
+...
+---
+# Seven values for six GPRs. Folding the "rm" operand of %6 wouldn't free a
+# register, since the "r" operand needs %6 in one anyway, so fold %8's.
+name:            test_rm_and_r_same_value_pressure
+tracksRegLiveness: true
+registers:
+  - { id: 0, class: gr64 }
+  - { id: 1, class: gr64 }
+  - { id: 2, class: gr64 }
+  - { id: 3, class: gr64 }
+  - { id: 4, class: gr64 }
+  - { id: 5, class: gr64 }
+  - { id: 6, class: gr64_norex2 }
+  - { id: 7, class: gr64_norex2 }
+  - { id: 8, class: gr64_norex2 }
+  - { id: 9, class: gr64_norex2 }
+  - { id: 10, class: gr64_norex2 }
+  - { id: 11, class: gr64_norex2 }
+  - { id: 12, class: gr64_norex2 }
+  - { id: 13, class: gr64_norex2 }
+  - { id: 14, class: gr64 }
+fixedStack:
+  - { id: 0, size: 8, alignment: 16, isImmutable: true }
+body:             |
+  bb.0:
+    liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+
+    ; CHECK-LABEL: name: test_rm_and_r_same_value_pressure
+    ; CHECK: liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: renamable $rax = COPY killed $rsi
+    ; CHECK-NEXT: renamable $rsi = MOV64rm %fixed-stack.0, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.0, align 16)
+    ; CHECK-NEXT: MOV64mr %stack.0, 1, $noreg, 0, $noreg, killed renamable $rax :: (store (s64) into %stack.0)
+    ; CHECK-NEXT: INLINEASM &"# $0 $1 $2 $3 $4 $5 $6 $7", sideeffect mayload attdialect, reguse:GR64_NOREX2 foldable, killed renamable $rdi, reguse:GR64_NOREX2, renamable $rdi, mem:m, %stack.0, 1, $noreg, 0, $noreg, reguse:GR64_NOREX2, killed renamable $rdx, reguse:GR64_NOREX2, killed renamable $rcx, reguse:GR64_NOREX2, killed renamable $r8, reguse:GR64_NOREX2, killed renamable $r9, reguse:GR64_NOREX2, killed renamable $rsi, clobber, implicit-def dead early-clobber $rax, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r10, clobber, implicit-def dead early-clobber $r11, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14, clobber, implicit-def dead early-clobber $r15 :: (load (s64) from %stack.0)
+    ; CHECK-NEXT: RET 0
+    %5:gr64 = COPY $r9
+    %4:gr64 = COPY $r8
+    %3:gr64 = COPY $rcx
+    %2:gr64 = COPY $rdx
+    %1:gr64 = COPY $rsi
+    %0:gr64 = COPY $rdi
+    %14:gr64 = MOV64rm %fixed-stack.0, 1, $noreg, 0, $noreg :: (load (s64) from %fixed-stack.0, align 16)
+    %6:gr64_norex2 = COPY %0
+    %8:gr64_norex2 = COPY %1
+    %9:gr64_norex2 = COPY %2
+    %10:gr64_norex2 = COPY %3
+    %11:gr64_norex2 = COPY %4
+    %12:gr64_norex2 = COPY %5
+    %13:gr64_norex2 = COPY %14
+    INLINEASM &"# $0 $1 $2 $3 $4 $5 $6 $7", sideeffect attdialect, reguse:GR64_NOREX2 foldable, %6, reguse:GR64_NOREX2, %6, reguse:GR64_NOREX2 foldable, %8, reguse:GR64_NOREX2, %9, reguse:GR64_NOREX2, %10, reguse:GR64_NOREX2, %11, reguse:GR64_NOREX2, %12, reguse:GR64_NOREX2, %13, clobber, implicit-def dead early-clobber $rax, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r10, clobber, implicit-def dead early-clobber $r11, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14, clobber, implicit-def dead early-clobber $r15
+    RET 0
+...
+---
+# Like test_rm_output_pressure, with an INLINEASM_BR. The reload of %7 after
+# the asm only covers the fallthrough path: on the indirect path, %7 must come
+# from the slot the asm wrote, which has to be %7's own spill slot, so that
+# the indirect block's ordinary reload of %7 finds it.
+name:            test_callbr_rm_indirect_use
+tracksRegLiveness: true
+registers:
+  - { id: 0, class: gr64 }
+  - { id: 1, class: gr64_nosp }
+  - { id: 2, class: gr64 }
+  - { id: 3, class: gr64 }
+  - { id: 4, class: gr64 }
+  - { id: 5, class: gr64 }
+  - { id: 6, class: gr64 }
+  - { id: 7, class: gr64_norex2 }
+  - { id: 8, class: gr64_norex2 }
+  - { id: 9, class: gr64_norex2 }
+  - { id: 10, class: gr64_norex2 }
+  - { id: 11, class: gr64_norex2 }
+  - { id: 12, class: gr64_norex2 }
+  - { id: 13, class: gr64_norex2 }
+  - { id: 14, class: gr64_norex2 }
+  - { id: 15, class: gr64 }
+  - { id: 16, class: gr32 }
+  - { id: 17, class: gr64 }
+body:             |
+  ; CHECK-LABEL: name: test_callbr_rm_indirect_use
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000), %bb.2(0x00000000)
+  ; CHECK-NEXT:   liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   MOV64mr %stack.1, 1, $noreg, 0, $noreg, $rdi :: (store (s64) into %stack.1)
+  ; CHECK-NEXT:   INLINEASM_BR &"# $0 $1 $2 $3 $4 $5 $6 ${7:l}", sideeffect maystore attdialect, mem:m, %stack.0, 1, $noreg, 0, $noreg, reguse:GR64_NOREX2, killed renamable $rdi, reguse:GR64_NOREX2, killed renamable $rsi, reguse:GR64_NOREX2, killed renamable $rdx, reguse:GR64_NOREX2, killed renamable $rcx, reguse:GR64_NOREX2, killed renamable $r8, reguse:GR64_NOREX2, killed renamable $r9, imm, %bb.2, clobber, implicit-def dead early-clobber $rax, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r10, clobber, implicit-def dead early-clobber $r11, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14, clobber, implicit-def dead early-clobber $r15 :: (store (s64) into %stack.0)
+  ; CHECK-NEXT:   renamable $rax = MOV64rm %stack.0, 1, $noreg, 0, $noreg :: (load (s64) from %stack.0)
+  ; CHECK-NEXT:   MOV64mr %stack.0, 1, $noreg, 0, $noreg, killed $rax :: (store (s64) into %stack.0)
+  ; CHECK-NEXT:   JMP_1 %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   renamable $eax = MOV32r0 implicit-def dead $eflags
+  ; CHECK-NEXT:   renamable $rax = SUBREG_TO_REG killed renamable $eax, %subreg.sub_32bit
+  ; CHECK-NEXT:   RET 0, killed $rax
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2 (inlineasm-br-indirect-target):
+  ; CHECK-NEXT:   $rax = MOV64rm %stack.0, 1, $noreg, 0, $noreg :: (load (s64) from %stack.0)
+  ; CHECK-NEXT:   $rcx = MOV64rm %stack.1, 1, $noreg, 0, $noreg :: (load (s64) from %stack.1)
+  ; CHECK-NEXT:   renamable $rax = LEA64r killed renamable $rax, 1, killed renamable $rcx, 0, $noreg
+  ; CHECK-NEXT:   RET 0, killed $rax
+  bb.0:
+    successors: %bb.1(0x80000000), %bb.2(0x00000000)
+    liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+
+    %6:gr64 = COPY $r9
+    %5:gr64 = COPY $r8
+    %4:gr64 = COPY $rcx
+    %3:gr64 = COPY $rdx
+    %2:gr64 = COPY $rsi
+    %1:gr64_nosp = COPY $rdi
+    %8:gr64_norex2 = COPY %1
+    %9:gr64_norex2 = COPY %2
+    %10:gr64_norex2 = COPY %3
+    %11:gr64_norex2 = COPY %4
+    %12:gr64_norex2 = COPY %5
+    %13:gr64_norex2 = COPY %6
+    INLINEASM_BR &"# $0 $1 $2 $3 $4 $5 $6 ${7:l}", sideeffect attdialect, regdef-ec:GR64_NOREX2 foldable, def early-clobber %7, reguse:GR64_NOREX2, %8, reguse:GR64_NOREX2, %9, reguse:GR64_NOREX2, %10, reguse:GR64_NOREX2, %11, reguse:GR64_NOREX2, %12, reguse:GR64_NOREX2, %13, imm, %bb.2, clobber, implicit-def dead early-clobber $rax, clobber, implicit-def dead early-clobber $rbx, clobber, implicit-def dead early-clobber $rbp, clobber, implicit-def dead early-clobber $r10, clobber, implicit-def dead early-clobber $r11, clobber, implicit-def dead early-clobber $r12, clobber, implicit-def dead early-clobber $r13, clobber, implicit-def dead early-clobber $r14, clobber, implicit-def dead early-clobber $r15
+    JMP_1 %bb.1
+
+  bb.1:
+    %16:gr32 = MOV32r0 implicit-def dead $eflags
+    %17:gr64 = SUBREG_TO_REG killed %16, %subreg.sub_32bit
+    $rax = COPY %17
+    RET 0, $rax
+
+  bb.2 (inlineasm-br-indirect-target):
+    %15:gr64 = LEA64r %7, 1, %1, 0, $noreg
+    $rax = COPY %15
+    RET 0, $rax
+...
+
diff --git a/llvm/test/CodeGen/X86/regallocfast-inline-asm-fold.mir b/llvm/test/CodeGen/X86/regallocfast-inline-asm-fold.mir
new file mode 100644
index 0000000000000..9100b5ee320b3
--- /dev/null
+++ b/llvm/test/CodeGen/X86/regallocfast-inline-asm-fold.mir
@@ -0,0 +1,58 @@
+# RUN: llc -mtriple=x86_64-unknown-linux-gnu -run-pass=regallocfast \
+# RUN:   -verify-machineinstrs %s -o - | FileCheck %s
+
+# The fast register allocator folds a foldable ("rm") inline asm operand to a
+# stack slot when the asm's register operands wouldn't fit otherwise. The
+# clobbers below leave six GPRs. The tied operands haven't been rewritten, so
+# a tied def and its input are different virtual registers: the allocator
+# lowers the tie itself, copying the input into the def's register if the
+# input stays live.
+
+---
+# Five inputs plus the tied pair fit: the input dies, so the def takes over its
+# register.
+name: tied_input_dies
+tracksRegLiveness: true
+body: |
+  bb.0:
+    liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+
+    ; CHECK-LABEL: name: tied_input_dies
+    ; CHECK-NOT: MOV64mr
+    ; CHECK: INLINEASM &"", sideeffect attdialect, regdef:GR64 foldable, def renamable [[REG:\$[a-z0-9]+]], {{.*}}, reguse tiedto:$0, killed renamable [[REG]](tied-def 3)
+    %0:gr64 = COPY $rdi
+    %1:gr64 = COPY $rsi
+    %2:gr64 = COPY $rdx
+    %3:gr64 = COPY $rcx
+    %4:gr64 = COPY $r8
+    %5:gr64 = COPY $r9
+    INLINEASM &"", sideeffect attdialect, regdef:GR64 foldable, def %6:gr64, reguse:GR64, killed %1, reguse:GR64, killed %2, reguse:GR64, killed %3, reguse:GR64, killed %4, reguse:GR64, killed %5, reguse tiedto:$0, killed %0(tied-def 3), clobber, implicit-def early-clobber $rax, clobber, implicit-def early-clobber $rbx, clobber, implicit-def early-clobber $rbp, clobber, implicit-def early-clobber $r10, clobber, implicit-def early-clobber $r11, clobber, implicit-def early-clobber $r12, clobber, implicit-def early-clobber $r13, clobber, implicit-def early-clobber $r14, clobber, implicit-def early-clobber $r15
+    $rax = COPY %6
+    RET 0, $rax
+...
+---
+# The same, except that the input is used after the asm, so it keeps its own
+# register while it's copied into the def's: seven registers. Fold the tied
+# pair instead, storing the input's value to the def's stack slot first.
+name: tied_input_lives
+tracksRegLiveness: true
+body: |
+  bb.0:
+    liveins: $rdi, $rsi, $rdx, $rcx, $r8, $r9
+
+    ; CHECK-LABEL: name: tied_input_lives
+    ; CHECK: MOV64mr %stack.[[SLOT:[0-9]+]], 1, $noreg, 0, $noreg, renamable [[INPUT:\$[a-z0-9]+]]
+    ; CHECK-NEXT: INLINEASM &"", sideeffect mayload maystore attdialect, mem:m, %stack.[[SLOT]], 1, $noreg, 0, $noreg, {{.*}}, mem:m, %stack.[[SLOT]], 1, $noreg, 0, $noreg, clobber,
+    ; CHECK-NEXT: renamable [[DEF:\$[a-z0-9]+]] = MOV64rm %stack.[[SLOT]], 1, $noreg, 0, $noreg
+    ; CHECK-NEXT: ADD64rr killed renamable [[DEF]], killed renamable [[INPUT]]
+    %0:gr64 = COPY $rdi
+    %1:gr64 = COPY $rsi
+    %2:gr64 = COPY $rdx
+    %3:gr64 = COPY $rcx
+    %4:gr64 = COPY $r8
+    %5:gr64 = COPY $r9
+    INLINEASM &"", sideeffect attdialect, regdef:GR64 foldable, def %6:gr64, reguse:GR64, killed %1, reguse:GR64, killed %2, reguse:GR64, killed %3, reguse:GR64, killed %4, reguse:GR64, killed %5, reguse tiedto:$0, %0(tied-def 3), clobber, implicit-def early-clobber $rax, clobber, implicit-def early-clobber $rbx, clobber, implicit-def early-clobber $rbp, clobber, implicit-def early-clobber $r10, clobber, implicit-def early-clobber $r11, clobber, implicit-def early-clobber $r12, clobber, implicit-def early-clobber $r13, clobber, implicit-def early-clobber $r14, clobber, implicit-def early-clobber $r15
+    %7:gr64 = ADD64rr %6, killed %0, implicit-def dead $eflags
+    $rax = COPY %7
+    RET 0, $rax
+...



More information about the llvm-branch-commits mailing list