[llvm-branch-commits] [llvm] [TargetLowering][X86] Prefer 'r' over 'm' for foldable "rm" inline as… (PR #229632)
Bill Wendling via llvm-branch-commits
llvm-branch-commits at lists.llvm.org
Tue Oct 6 18:58:12 PDT 2026
https://github.com/isanbard created https://github.com/llvm/llvm-project/pull/229632
…m operands
An "rm" (register-or-memory) inline asm operand has always resolved to 'm', because getConstraintPreferences() picks the most general constraint present, and 'm' is more general than 'r'. That's safe, since memory can't run out, but it forces a value that could stay in a register through a stack slot even when there's no register pressure (https://github.com/llvm/llvm-project/issues/20571).
Prefer 'r' instead where the register allocator can fold the register back to a stack slot when it runs out of registers, and mark the register operand foldable (InlineAsm::Flag::setRegMayBeFolded()) so it does. Both allocators can: the greedy allocator folds an operand when it spills its value, and the fast allocator folds operands up front when the asm's register operands wouldn't fit.
ParseConstraints() sets AsmOperandInfo::MayFoldRegister for an operand whose constraint codes are exactly {r, m}, above -O0, on a target that opts in through the new supportsRegMemInlineAsmFolding() hook, which needs TargetInstrInfo::getFrameIndexOperands() and is only enabled for X86. It leaves out operands where the fold can't work:
- An indirect operand, which Clang emits for every "=rm" output and for an "rm" input of aggregate type, is already in memory.
- A value no 'r' register can hold, e.g. an x87 or vector value.
- A value that takes two registers, e.g. __int128 on x86-64 or a 64-bit integer on i686, since the allocator folds a single register operand.
The tied input of a "+rm" operand has its own constraint ("0"), so only the output is marked.
SelectionDAG and GlobalISel, which shares getConstraintPreferences(), both mark the operand foldable, though no in-tree target both lowers inline asm with GlobalISel and supports the fold yet.
inlineasm-sched-bug.ll uses an "rm" input, which now stays in a register instead of round-tripping through a stack slot.
Assisted-by: Claude Opus 5.5
>From c9f2a1f19b859ed349bc7376031b5edcb6ea267a Mon Sep 17 00:00:00 2001
From: Bill Wendling <isanbard at gmail.com>
Date: Tue, 6 Oct 2026 06:16:29 -0700
Subject: [PATCH] [TargetLowering][X86] Prefer 'r' over 'm' for foldable "rm"
inline asm operands
An "rm" (register-or-memory) inline asm operand has always resolved to
'm', because getConstraintPreferences() picks the most general
constraint present, and 'm' is more general than 'r'. That's safe, since
memory can't run out, but it forces a value that could stay in a
register through a stack slot even when there's no register pressure
(https://github.com/llvm/llvm-project/issues/20571).
Prefer 'r' instead where the register allocator can fold the register
back to a stack slot when it runs out of registers, and mark the
register operand foldable (InlineAsm::Flag::setRegMayBeFolded()) so it
does. Both allocators can: the greedy allocator folds an operand when it
spills its value, and the fast allocator folds operands up front when
the asm's register operands wouldn't fit.
ParseConstraints() sets AsmOperandInfo::MayFoldRegister for an operand
whose constraint codes are exactly {r, m}, above -O0, on a target that
opts in through the new supportsRegMemInlineAsmFolding() hook, which
needs TargetInstrInfo::getFrameIndexOperands() and is only enabled for
X86. It leaves out operands where the fold can't work:
- An indirect operand, which Clang emits for every "=rm" output and for
an "rm" input of aggregate type, is already in memory.
- A value no 'r' register can hold, e.g. an x87 or vector value.
- A value that takes two registers, e.g. __int128 on x86-64 or a 64-bit
integer on i686, since the allocator folds a single register operand.
The tied input of a "+rm" operand has its own constraint ("0"), so only
the output is marked.
SelectionDAG and GlobalISel, which shares getConstraintPreferences(),
both mark the operand foldable, though no in-tree target both lowers
inline asm with GlobalISel and supports the fold yet.
inlineasm-sched-bug.ll uses an "rm" input, which now stays in a
register instead of round-tripping through a stack slot.
Assisted-by: Claude Opus 5.5
---
llvm/include/llvm/CodeGen/TargetLowering.h | 18 +
.../CodeGen/GlobalISel/InlineAsmLowering.cpp | 7 +
.../SelectionDAG/SelectionDAGBuilder.cpp | 14 +-
.../SelectionDAG/SelectionDAGBuilder.h | 6 +-
.../CodeGen/SelectionDAG/TargetLowering.cpp | 53 +-
llvm/lib/Target/X86/X86ISelLowering.h | 4 +
.../inline-asm-direct-mem-output-error.ll | 3 +-
.../X86/asm-constraints-rm-callbr-fold.ll | 74 ++
.../CodeGen/X86/asm-constraints-rm-isel.ll | 97 +++
.../X86/asm-constraints-rm-pressure.ll | 689 ++++++++++++++++
llvm/test/CodeGen/X86/asm-constraints-rm.ll | 35 +
.../CodeGen/X86/asm-constraints-torture.ll | 767 ++++++++++++++++++
llvm/test/CodeGen/X86/inline-asm-callbase.ll | 227 ++++++
.../X86/inline-asm-direct-mem-output-error.ll | 5 +-
llvm/test/CodeGen/X86/inlineasm-sched-bug.ll | 5 +-
15 files changed, 1988 insertions(+), 16 deletions(-)
create mode 100644 llvm/test/CodeGen/X86/asm-constraints-rm-callbr-fold.ll
create mode 100644 llvm/test/CodeGen/X86/asm-constraints-rm-isel.ll
create mode 100644 llvm/test/CodeGen/X86/asm-constraints-rm-pressure.ll
create mode 100644 llvm/test/CodeGen/X86/asm-constraints-rm.ll
create mode 100644 llvm/test/CodeGen/X86/asm-constraints-torture.ll
create mode 100644 llvm/test/CodeGen/X86/inline-asm-callbase.ll
diff --git a/llvm/include/llvm/CodeGen/TargetLowering.h b/llvm/include/llvm/CodeGen/TargetLowering.h
index ab92757abeb66b..1c0efaeadb52a9 100644
--- a/llvm/include/llvm/CodeGen/TargetLowering.h
+++ b/llvm/include/llvm/CodeGen/TargetLowering.h
@@ -5473,6 +5473,16 @@ class LLVM_ABI TargetLowering : public TargetLoweringBase {
/// The ValueType for the operand value.
MVT ConstraintVT = MVT::Other;
+ /// True if this "rm" operand should prefer a register, leaving the
+ /// register allocator to fold it to a stack slot if it runs out of
+ /// registers. ParseConstraints() sets this for a direct operand whose
+ /// value fits in one register, on a target that can fold it.
+ /// getConstraintPreferences() then picks 'r', and instruction selection
+ /// marks the register operand foldable (see
+ /// InlineAsm::Flag::setRegMayBeFolded()). The tied input of a "+rm"
+ /// operand has its own constraint ("0"), so only the output gets this.
+ bool MayFoldRegister = false;
+
/// Copy constructor for copying from a ConstraintInfo.
AsmOperandInfo(InlineAsm::ConstraintInfo Info)
: InlineAsm::ConstraintInfo(std::move(Info)) {}
@@ -5517,6 +5527,14 @@ class LLVM_ABI TargetLowering : public TargetLoweringBase {
/// Given a constraint, return the type of constraint it is for this target.
virtual ConstraintType getConstraintType(StringRef Constraint) const;
+ /// Return true if the register allocator can fold an inline asm register
+ /// operand to a stack slot on this target, which needs an override of
+ /// TargetInstrInfo::getFrameIndexOperands(). Only then may an "rm" operand
+ /// prefer a register (see AsmOperandInfo::MayFoldRegister): that choice
+ /// relies on the allocator falling back to memory when it runs out of
+ /// registers.
+ virtual bool supportsRegMemInlineAsmFolding() const { return false; }
+
using ConstraintPair = std::pair<StringRef, TargetLowering::ConstraintType>;
using ConstraintGroup = SmallVector<ConstraintPair>;
/// Given an OpInfo with list of constraints codes as strings, return a
diff --git a/llvm/lib/CodeGen/GlobalISel/InlineAsmLowering.cpp b/llvm/lib/CodeGen/GlobalISel/InlineAsmLowering.cpp
index 5d66cb5e332276..9aa2bfdf265d4e 100644
--- a/llvm/lib/CodeGen/GlobalISel/InlineAsmLowering.cpp
+++ b/llvm/lib/CodeGen/GlobalISel/InlineAsmLowering.cpp
@@ -393,6 +393,11 @@ bool InlineAsmLowering::lowerInlineAsm(
// tied operands that can use the regclass information from the def.
const TargetRegisterClass *RC = MRI->getRegClass(OpInfo.Regs.front());
Flag.setRegClass(RC->getID());
+
+ // An "rm" operand that preferred a register may still be folded to
+ // a stack slot by the register allocator.
+ if (OpInfo.MayFoldRegister)
+ Flag.setRegMayBeFolded(true);
}
Inst.addImm(Flag);
@@ -576,6 +581,8 @@ bool InlineAsmLowering::lowerInlineAsm(
// Put the register class of the virtual registers in the flag word.
const TargetRegisterClass *RC = MRI->getRegClass(OpInfo.Regs.front());
Flag.setRegClass(RC->getID());
+ if (OpInfo.MayFoldRegister)
+ Flag.setRegMayBeFolded(true);
}
Inst.addImm(Flag);
if (!buildAnyextOrCopy(OpInfo.Regs[0], SourceRegs[0], MIRBuilder))
diff --git a/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp b/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp
index de009524f97c37..24af1b5d43467a 100644
--- a/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp
@@ -1036,7 +1036,8 @@ void RegsForValue::getCopyToRegs(SDValue Val, SelectionDAG &DAG,
void RegsForValue::AddInlineAsmOperands(InlineAsm::Kind Code, bool HasMatching,
unsigned MatchingIdx, const SDLoc &dl,
SelectionDAG &DAG,
- std::vector<SDValue> &Ops) const {
+ std::vector<SDValue> &Ops,
+ bool MayFoldRegister) const {
const TargetLowering &TLI = DAG.getTargetLoweringInfo();
InlineAsm::Flag Flag(Code, Regs.size());
@@ -1051,6 +1052,10 @@ void RegsForValue::AddInlineAsmOperands(InlineAsm::Kind Code, bool HasMatching,
const MachineRegisterInfo &MRI = DAG.getMachineFunction().getRegInfo();
const TargetRegisterClass *RC = MRI.getRegClass(Regs.front());
Flag.setRegClass(RC->getID());
+ if (MayFoldRegister) {
+ assert(Regs.size() == 1 && "only a single register can be folded");
+ Flag.setRegMayBeFolded(true);
+ }
}
SDValue Res = DAG.getTargetConstant(Flag, dl, MVT::i32);
@@ -10655,7 +10660,7 @@ static bool prepareDAGLevelOperands(ConstraintDecisionInfo &Info,
OpInfo.AssignedRegs.AddInlineAsmOperands(
OpInfo.isEarlyClobber ? InlineAsm::Kind::RegDefEarlyClobber
: InlineAsm::Kind::RegDef,
- false, 0, DL, DAG, Info.AsmNodeOperands);
+ false, 0, DL, DAG, Info.AsmNodeOperands, OpInfo.MayFoldRegister);
}
break;
@@ -10821,8 +10826,9 @@ static bool prepareDAGLevelOperands(ConstraintDecisionInfo &Info,
OpInfo.AssignedRegs.getCopyToRegs(InOperandVal, DAG, DL, Info.Chain,
&Info.Glue, &Call);
- OpInfo.AssignedRegs.AddInlineAsmOperands(
- InlineAsm::Kind::RegUse, false, 0, DL, DAG, Info.AsmNodeOperands);
+ OpInfo.AssignedRegs.AddInlineAsmOperands(InlineAsm::Kind::RegUse, false,
+ 0, DL, DAG, Info.AsmNodeOperands,
+ OpInfo.MayFoldRegister);
break;
}
diff --git a/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.h b/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.h
index c9bb96b86b7bde..7981c24e83865f 100644
--- a/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.h
+++ b/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.h
@@ -826,10 +826,12 @@ struct RegsForValue {
/// Add this value to the specified inlineasm node operand list. This adds the
/// code marker, matching input operand index (if applicable), and includes
- /// the number of values added into it.
+ /// the number of values added into it. If \p MayFoldRegister, the register
+ /// allocator may fold the (single) register to a stack slot.
void AddInlineAsmOperands(InlineAsm::Kind Code, bool HasMatching,
unsigned MatchingIdx, const SDLoc &dl,
- SelectionDAG &DAG, std::vector<SDValue> &Ops) const;
+ SelectionDAG &DAG, std::vector<SDValue> &Ops,
+ bool MayFoldRegister = false) const;
/// Check if the total RegCount is greater than one.
bool occupiesMultipleRegs() const {
diff --git a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
index e89be8b69929f8..0b5a4534d3af9b 100644
--- a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
@@ -6228,6 +6228,42 @@ unsigned TargetLowering::AsmOperandInfo::getMatchedOperand() const {
return atoi(ConstraintCode.c_str());
}
+/// Return true if \p OpInfo is an "rm" operand that should prefer a register,
+/// leaving the register allocator to fold it to a stack slot if it runs out of
+/// registers (see AsmOperandInfo::MayFoldRegister).
+static bool
+shouldPreferFoldableRegister(const TargetLowering &TLI,
+ const TargetRegisterInfo *TRI,
+ const TargetLowering::AsmOperandInfo &OpInfo) {
+ // Codes stays empty for a constraint with alternatives ("r|m") until one is
+ // selected, so this only ever matches a plain "rm".
+ if (OpInfo.Codes.size() != 2 || !is_contained(OpInfo.Codes, "r") ||
+ !is_contained(OpInfo.Codes, "m"))
+ return false;
+
+ // Without a fold to fall back on, 'm' is the only choice that can't run out
+ // of registers. Above -O0 only, which leaves -O0 code as it was.
+ if (!TLI.supportsRegMemInlineAsmFolding() ||
+ TLI.getTargetMachine().getOptLevel() == CodeGenOptLevel::None)
+ return false;
+
+ // An indirect operand is already in memory: a register would only add a
+ // load or store around the asm, and isn't supported for an indirect input.
+ if (OpInfo.isIndirect)
+ return false;
+
+ // The register allocator folds one register operand, so the value must fit
+ // in a single register, and some types (e.g. x87 and vector types) can't go
+ // in an 'r' register at all.
+ MVT VT = OpInfo.ConstraintVT;
+ if (VT == MVT::Other || VT.isScalableVector())
+ return false;
+
+ const TargetRegisterClass *RC =
+ TLI.getRegForInlineAsmConstraint(TRI, "r", VT).second;
+ return RC && VT.getFixedSizeInBits() <= TRI->getRegSizeInBits(*RC);
+}
+
/// Split up the constraint string from the inline assembly value into the
/// specific constraints and their prefixes, and also tie in the associated
/// operand values.
@@ -6326,6 +6362,8 @@ TargetLowering::ParseConstraints(const DataLayout &DL,
OpInfo.ConstraintVT = VT.isSimple() ? VT.getSimpleVT() : MVT::Other;
ArgNo++;
}
+
+ OpInfo.MayFoldRegister = shouldPreferFoldableRegister(*this, TRI, OpInfo);
}
// If we have multiple alternative constraints, select the best alternative.
@@ -6424,7 +6462,8 @@ TargetLowering::ParseConstraints(const DataLayout &DL,
/// preferrable (when they can be emitted). A higher return value means a
/// stronger preference for one constraint type relative to another.
/// FIXME: We should prefer registers over memory but doing so may lead to
-/// unrecoverable register exhaustion later.
+/// unrecoverable register exhaustion later, unless the register allocator can
+/// fold the register to memory (see AsmOperandInfo::MayFoldRegister).
/// https://github.com/llvm/llvm-project/issues/20571
static unsigned getConstraintPiority(TargetLowering::ConstraintType CT) {
switch (CT) {
@@ -6538,7 +6577,10 @@ TargetLowering::ConstraintWeight
/// 1) If there is an 'other' constraint, and if the operand is valid for
/// that constraint, use it. This makes us take advantage of 'i'
/// constraints when available.
-/// 2) Otherwise, pick the most general constraint present. This prefers
+/// 2) For an "rm" operand that the register allocator can fold to a stack
+/// slot (see AsmOperandInfo::MayFoldRegister), pick 'r': the allocator
+/// still falls back to memory if it runs out of registers.
+/// 3) Otherwise, pick the most general constraint present. This prefers
/// 'm' over 'r', for example.
///
TargetLowering::ConstraintGroup TargetLowering::getConstraintPreferences(
@@ -6546,6 +6588,13 @@ TargetLowering::ConstraintGroup TargetLowering::getConstraintPreferences(
ConstraintGroup Ret;
Ret.reserve(OpInfo.Codes.size());
+
+ if (OpInfo.MayFoldRegister) {
+ Ret.emplace_back("r", getConstraintType("r"));
+ Ret.emplace_back("m", getConstraintType("m"));
+ return Ret;
+ }
+
for (StringRef Code : OpInfo.Codes) {
TargetLowering::ConstraintType CType = getConstraintType(Code);
diff --git a/llvm/lib/Target/X86/X86ISelLowering.h b/llvm/lib/Target/X86/X86ISelLowering.h
index ba04080573c421..45897dd0f527eb 100644
--- a/llvm/lib/Target/X86/X86ISelLowering.h
+++ b/llvm/lib/Target/X86/X86ISelLowering.h
@@ -414,6 +414,10 @@ namespace llvm {
ConstraintType getConstraintType(StringRef Constraint) const override;
+ // X86InstrInfo::getFrameIndexOperands() lets the register allocator fold
+ // an inline asm register operand to a stack slot.
+ bool supportsRegMemInlineAsmFolding() const override { return true; }
+
/// Examine constraint string and operand type and determine a weight value.
/// The operand object must already have been set up with the operand type.
ConstraintWeight
diff --git a/llvm/test/CodeGen/AArch64/inline-asm-direct-mem-output-error.ll b/llvm/test/CodeGen/AArch64/inline-asm-direct-mem-output-error.ll
index 7ebe9f006943f9..2165b79681476e 100644
--- a/llvm/test/CodeGen/AArch64/inline-asm-direct-mem-output-error.ll
+++ b/llvm/test/CodeGen/AArch64/inline-asm-direct-mem-output-error.ll
@@ -4,7 +4,8 @@
; A direct output is the asm's result, so there is no memory to write it to.
; Picking a memory constraint for one is an error, not a crash, with either
-; instruction selector.
+; instruction selector. AArch64 can't fold a register operand to memory, so
+; "rm" keeps picking memory.
; CHECK: error: cannot handle direct memory outputs yet for constraint 'm'
define i64 @rm_output() {
diff --git a/llvm/test/CodeGen/X86/asm-constraints-rm-callbr-fold.ll b/llvm/test/CodeGen/X86/asm-constraints-rm-callbr-fold.ll
new file mode 100644
index 00000000000000..f1f7608b8f718e
--- /dev/null
+++ b/llvm/test/CodeGen/X86/asm-constraints-rm-callbr-fold.ll
@@ -0,0 +1,74 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 6
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --regalloc=fast \
+; RUN: -verify-machineinstrs -verify-regalloc < %s | FileCheck %s
+
+; The fast allocator folds this "=&rm" output to a stack slot, since the six
+; "r" inputs and the clobbers leave no register for it. The reload of %r after
+; the callbr only covers the fallthrough path: on the indirect path, %r must
+; come from the slot the asm wrote, which has to be %r's own spill slot, so
+; that the indirect block's ordinary reload of %r finds it.
+define i64 @test_callbr_rm_indirect_use(i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %f) {
+; CHECK-LABEL: test_callbr_rm_indirect_use:
+; CHECK: # %bb.0: # %entry
+; CHECK-NEXT: pushq %rbp
+; CHECK-NEXT: .cfi_def_cfa_offset 16
+; CHECK-NEXT: pushq %r15
+; CHECK-NEXT: .cfi_def_cfa_offset 24
+; CHECK-NEXT: pushq %r14
+; CHECK-NEXT: .cfi_def_cfa_offset 32
+; CHECK-NEXT: pushq %r13
+; CHECK-NEXT: .cfi_def_cfa_offset 40
+; CHECK-NEXT: pushq %r12
+; CHECK-NEXT: .cfi_def_cfa_offset 48
+; CHECK-NEXT: pushq %rbx
+; CHECK-NEXT: .cfi_def_cfa_offset 56
+; CHECK-NEXT: .cfi_offset %rbx, -56
+; CHECK-NEXT: .cfi_offset %r12, -48
+; CHECK-NEXT: .cfi_offset %r13, -40
+; CHECK-NEXT: .cfi_offset %r14, -32
+; CHECK-NEXT: .cfi_offset %r15, -24
+; CHECK-NEXT: .cfi_offset %rbp, -16
+; CHECK-NEXT: movq %rdi, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; CHECK-NEXT: #APP # 8-byte Folded Spill
+; CHECK-NEXT: # -{{[0-9]+}}(%rsp) %rdi %rsi %rdx %rcx %r8 %r9 .LBB0_3
+; CHECK-NEXT: #NO_APP
+; CHECK-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rax # 8-byte Reload
+; CHECK-NEXT: movq %rax, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; CHECK-NEXT: # %bb.1: # %fallthrough
+; CHECK-NEXT: xorl %eax, %eax
+; CHECK-NEXT: # kill: def $rax killed $eax
+; CHECK-NEXT: .LBB0_2: # %fallthrough
+; CHECK-NEXT: popq %rbx
+; CHECK-NEXT: .cfi_def_cfa_offset 48
+; CHECK-NEXT: popq %r12
+; CHECK-NEXT: .cfi_def_cfa_offset 40
+; CHECK-NEXT: popq %r13
+; CHECK-NEXT: .cfi_def_cfa_offset 32
+; CHECK-NEXT: popq %r14
+; CHECK-NEXT: .cfi_def_cfa_offset 24
+; CHECK-NEXT: popq %r15
+; CHECK-NEXT: .cfi_def_cfa_offset 16
+; CHECK-NEXT: popq %rbp
+; CHECK-NEXT: .cfi_def_cfa_offset 8
+; CHECK-NEXT: retq
+; CHECK-NEXT: .LBB0_3: # Inline asm indirect target
+; CHECK-NEXT: # %indirect
+; CHECK-NEXT: # Label of block must be emitted
+; CHECK-NEXT: .cfi_def_cfa_offset 56
+; CHECK-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rax # 8-byte Reload
+; CHECK-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rcx # 8-byte Reload
+; CHECK-NEXT: addq %rcx, %rax
+; CHECK-NEXT: jmp .LBB0_2
+entry:
+ %r = callbr i64 asm sideeffect "# $0 $1 $2 $3 $4 $5 $6 ${7:l}",
+ "=&rm,r,r,r,r,r,r,!i,~{rax},~{rbx},~{rbp},~{r10},~{r11},~{r12},~{r13},~{r14},~{r15}"
+ (i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %f)
+ to label %fallthrough [label %indirect]
+
+fallthrough:
+ ret i64 0
+
+indirect:
+ %sum = add i64 %r, %a
+ ret i64 %sum
+}
diff --git a/llvm/test/CodeGen/X86/asm-constraints-rm-isel.ll b/llvm/test/CodeGen/X86/asm-constraints-rm-isel.ll
new file mode 100644
index 00000000000000..22a492a7baaf56
--- /dev/null
+++ b/llvm/test/CodeGen/X86/asm-constraints-rm-isel.ll
@@ -0,0 +1,97 @@
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu -stop-after=finalize-isel < %s \
+; RUN: | FileCheck --check-prefixes=CHECK,X64 %s
+; RUN: llc -mtriple=i686-unknown-linux-gnu -stop-after=finalize-isel < %s \
+; RUN: | FileCheck --check-prefixes=CHECK,X86 %s
+
+; An "rm" operand picks a register, marked foldable so that the register
+; allocator can still put it in a stack slot, but only when the allocator can
+; fold it: a direct operand whose value fits in one register. Everything else
+; keeps using memory.
+
+%struct.S = type { [3 x i8] }
+
+define void @input_i32(i32 %x) {
+; CHECK-LABEL: name: input_i32
+; CHECK: INLINEASM {{.*}}, reguse:GR32{{[_A-Z0-9]*}} foldable, %{{[0-9]+}}
+ call void asm sideeffect "# $0", "rm"(i32 %x)
+ ret void
+}
+
+; Two registers on i686, which can't be folded as one.
+define void @input_i64(i64 %x) {
+; CHECK-LABEL: name: input_i64
+; X64: INLINEASM {{.*}}, reguse:GR64{{[_A-Z0-9]*}} foldable, %{{[0-9]+}}
+; X86: INLINEASM {{.*}}, mem:m, %stack.0,
+ call void asm sideeffect "# $0", "rm"(i64 %x)
+ ret void
+}
+
+define void @input_double(double %x) {
+; CHECK-LABEL: name: input_double
+; X64: INLINEASM {{.*}}, reguse:GR64{{[_A-Z0-9]*}} foldable, %{{[0-9]+}}
+; X86: INLINEASM {{.*}}, mem:m, %stack.0,
+ call void asm sideeffect "# $0", "rm"(double %x)
+ ret void
+}
+
+define void @input_i128(i128 %x) {
+; CHECK-LABEL: name: input_i128
+; CHECK: INLINEASM {{.*}}, mem:m, %stack.0,
+ call void asm sideeffect "# $0", "rm"(i128 %x)
+ ret void
+}
+
+; No 'r' register can hold these at all.
+define void @input_x86_fp80(x86_fp80 %x) {
+; CHECK-LABEL: name: input_x86_fp80
+; CHECK: INLINEASM {{.*}}, mem:m, %stack.0,
+ call void asm sideeffect "# $0", "rm"(x86_fp80 %x)
+ ret void
+}
+
+define void @input_v4f32(<4 x float> %x) {
+; CHECK-LABEL: name: input_v4f32
+; CHECK: INLINEASM {{.*}}, mem:m, %stack.0,
+ call void asm sideeffect "# $0", "rm"(<4 x float> %x)
+ ret void
+}
+
+; An indirect operand already names memory, as Clang emits for an "rm" input
+; of aggregate type and for every "=rm" output.
+define void @input_indirect(ptr %p) {
+; CHECK-LABEL: name: input_indirect
+; CHECK: INLINEASM {{.*}}, mem:m, {{(killed )?}}%0,
+ call void asm sideeffect "# $0", "*rm"(ptr elementtype(%struct.S) %p)
+ ret void
+}
+
+define void @output_indirect(ptr %p) {
+; CHECK-LABEL: name: output_indirect
+; CHECK: INLINEASM {{.*}}, mem:m, {{(killed )?}}%0,
+ call void asm sideeffect "# $0", "=*rm"(ptr elementtype(i64) %p)
+ ret void
+}
+
+define i32 @output_i32() {
+; CHECK-LABEL: name: output_i32
+; CHECK: INLINEASM {{.*}}, regdef:GR32{{[_A-Z0-9]*}} foldable, def %{{[0-9]+}}
+ %r = call i32 asm sideeffect "# $0", "=rm"()
+ ret i32 %r
+}
+
+; A tied pair is folded through its output.
+define i32 @inout_i32(i32 %x) {
+; CHECK-LABEL: name: inout_i32
+; CHECK: INLINEASM {{.*}}, regdef:GR32{{[_A-Z0-9]*}} foldable, def %{{[0-9]+}}, reguse tiedto:$0, %{{[0-9]+}}(tied-def 3)
+ %r = call i32 asm sideeffect "# $0", "=rm,0"(i32 %x)
+ ret i32 %r
+}
+
+; Clang's form of "+rm": an indirect output can't use memory when tied, so it
+; was already a register, which stays unfoldable.
+define void @inout_indirect_i32(ptr %p, i32 %x) {
+; CHECK-LABEL: name: inout_indirect_i32
+; CHECK: INLINEASM {{.*}}, regdef:GR32{{[_A-Z0-9]*}}, def %{{[0-9]+}}, reguse tiedto:$0, %{{[0-9]+}}(tied-def 3)
+ call void asm sideeffect "# $0", "=*rm,0"(ptr elementtype(i32) %p, i32 %x)
+ ret void
+}
diff --git a/llvm/test/CodeGen/X86/asm-constraints-rm-pressure.ll b/llvm/test/CodeGen/X86/asm-constraints-rm-pressure.ll
new file mode 100644
index 00000000000000..00a6a626147cc4
--- /dev/null
+++ b/llvm/test/CodeGen/X86/asm-constraints-rm-pressure.ll
@@ -0,0 +1,689 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 6
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --regalloc=greedy -verify-machineinstrs < %s | FileCheck --check-prefix=GREEDY_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --regalloc=fast -verify-machineinstrs < %s | FileCheck --check-prefix=FAST_RA %s
+
+; No -O0 RUN line: -O0 picks memory for "rm", which a direct output like the
+; ones below can't use (see inline-asm-direct-mem-output-error.ll).
+
+; The asm clobbers rax/rbx/rbp/r10-r15, leaving six of the 15 GPRs for its
+; register operands, and its six "r" inputs take all of them. The "=&rm"
+; output can't share a register with an input, so both allocators fold it to
+; a stack slot.
+define i64 @test_rm_output_pressure(i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %f) {
+; GREEDY_RA-LABEL: test_rm_output_pressure:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: pushq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: pushq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: pushq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: pushq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: pushq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: pushq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 56
+; GREEDY_RA-NEXT: .cfi_offset %rbx, -56
+; GREEDY_RA-NEXT: .cfi_offset %r12, -48
+; GREEDY_RA-NEXT: .cfi_offset %r13, -40
+; GREEDY_RA-NEXT: .cfi_offset %r14, -32
+; GREEDY_RA-NEXT: .cfi_offset %r15, -24
+; GREEDY_RA-NEXT: .cfi_offset %rbp, -16
+; GREEDY_RA-NEXT: movq %r9, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %r8, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rcx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rdx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rsi, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %r9 # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %r8 # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rcx # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rdx # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rsi # 8-byte Reload
+; GREEDY_RA-NEXT: #APP # 8-byte Folded Spill
+; GREEDY_RA-NEXT: # -{{[0-9]+}}(%rsp) %rdi %rsi %rdx %rcx %r8 %r9
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rax # 8-byte Reload
+; GREEDY_RA-NEXT: popq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: popq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: popq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: popq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: popq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: popq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 8
+; GREEDY_RA-NEXT: retq
+;
+; FAST_RA-LABEL: test_rm_output_pressure:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: pushq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: pushq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: pushq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: pushq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: pushq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: pushq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 56
+; FAST_RA-NEXT: .cfi_offset %rbx, -56
+; FAST_RA-NEXT: .cfi_offset %r12, -48
+; FAST_RA-NEXT: .cfi_offset %r13, -40
+; FAST_RA-NEXT: .cfi_offset %r14, -32
+; FAST_RA-NEXT: .cfi_offset %r15, -24
+; FAST_RA-NEXT: .cfi_offset %rbp, -16
+; FAST_RA-NEXT: #APP # 8-byte Folded Spill
+; FAST_RA-NEXT: # -{{[0-9]+}}(%rsp) %rdi %rsi %rdx %rcx %r8 %r9
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rax # 8-byte Reload
+; FAST_RA-NEXT: popq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: popq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: popq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: popq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: popq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: popq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 8
+; FAST_RA-NEXT: retq
+entry:
+ %0 = call i64 asm sideeffect "# $0 $1 $2 $3 $4 $5 $6",
+ "=&rm,r,r,r,r,r,r,~{rax},~{rbx},~{rbp},~{r10},~{r11},~{r12},~{r13},~{r14},~{r15}"
+ (i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %f)
+ ret i64 %0
+}
+
+; One value passed to two "rm" operands: spilling it means folding both
+; operands into its stack slot at once.
+define void @test_rm_same_value_twice_pressure(i64 %x, i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %f) {
+; GREEDY_RA-LABEL: test_rm_same_value_twice_pressure:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: pushq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: pushq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: pushq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: pushq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: pushq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: pushq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 56
+; GREEDY_RA-NEXT: .cfi_offset %rbx, -56
+; GREEDY_RA-NEXT: .cfi_offset %r12, -48
+; GREEDY_RA-NEXT: .cfi_offset %r13, -40
+; GREEDY_RA-NEXT: .cfi_offset %r14, -32
+; GREEDY_RA-NEXT: .cfi_offset %r15, -24
+; GREEDY_RA-NEXT: .cfi_offset %rbp, -16
+; GREEDY_RA-NEXT: movq %r9, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %r8, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rcx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rdx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rsi, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rdi, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq {{[0-9]+}}(%rsp), %r9
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %r8 # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rcx # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rdx # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rsi # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rdi # 8-byte Reload
+; GREEDY_RA-NEXT: #APP # 8-byte Folded Reload
+; GREEDY_RA-NEXT: # -{{[0-9]+}}(%rsp) -{{[0-9]+}}(%rsp) %rdi %rsi %rdx %rcx %r8 %r9
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: popq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: popq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: popq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: popq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: popq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: popq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 8
+; GREEDY_RA-NEXT: retq
+;
+; FAST_RA-LABEL: test_rm_same_value_twice_pressure:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: pushq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: pushq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: pushq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: pushq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: pushq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: pushq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 56
+; FAST_RA-NEXT: .cfi_offset %rbx, -56
+; FAST_RA-NEXT: .cfi_offset %r12, -48
+; FAST_RA-NEXT: .cfi_offset %r13, -40
+; FAST_RA-NEXT: .cfi_offset %r14, -32
+; FAST_RA-NEXT: .cfi_offset %r15, -24
+; FAST_RA-NEXT: .cfi_offset %rbp, -16
+; FAST_RA-NEXT: movq %rdi, %rax
+; FAST_RA-NEXT: movq {{[0-9]+}}(%rsp), %rdi
+; FAST_RA-NEXT: movq %rax, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; FAST_RA-NEXT: #APP # 8-byte Folded Reload
+; FAST_RA-NEXT: # -{{[0-9]+}}(%rsp) -{{[0-9]+}}(%rsp) %rsi %rdx %rcx %r8 %r9 %rdi
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: popq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: popq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: popq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: popq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: popq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: popq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 8
+; FAST_RA-NEXT: retq
+entry:
+ call void asm sideeffect "# $0 $1 $2 $3 $4 $5 $6 $7",
+ "rm,rm,r,r,r,r,r,r,~{rax},~{rbx},~{rbp},~{r10},~{r11},~{r12},~{r13},~{r14},~{r15}"
+ (i64 %x, i64 %x, i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %f)
+ ret void
+}
+
+; Folding the "rm" input moves the operands after it, including the input
+; tied to the "=r" output, whose tie must survive.
+define i64 @test_rm_before_tied_pressure(i64 %x, i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %y) {
+; GREEDY_RA-LABEL: test_rm_before_tied_pressure:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: pushq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: pushq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: pushq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: pushq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: pushq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: pushq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 56
+; GREEDY_RA-NEXT: .cfi_offset %rbx, -56
+; GREEDY_RA-NEXT: .cfi_offset %r12, -48
+; GREEDY_RA-NEXT: .cfi_offset %r13, -40
+; GREEDY_RA-NEXT: .cfi_offset %r14, -32
+; GREEDY_RA-NEXT: .cfi_offset %r15, -24
+; GREEDY_RA-NEXT: .cfi_offset %rbp, -16
+; GREEDY_RA-NEXT: movq %r9, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %r8, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rcx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rdx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rsi, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rdi, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq {{[0-9]+}}(%rsp), %r9
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %r8 # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rcx # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rdx # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rsi # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rdi # 8-byte Reload
+; GREEDY_RA-NEXT: #APP # 8-byte Folded Reload
+; GREEDY_RA-NEXT: # %r9 -{{[0-9]+}}(%rsp) %rdi %rsi %rdx %rcx %r8 %r9
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: movq %r9, %rax
+; GREEDY_RA-NEXT: popq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: popq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: popq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: popq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: popq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: popq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 8
+; GREEDY_RA-NEXT: retq
+;
+; FAST_RA-LABEL: test_rm_before_tied_pressure:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: pushq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: pushq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: pushq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: pushq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: pushq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: pushq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 56
+; FAST_RA-NEXT: .cfi_offset %rbx, -56
+; FAST_RA-NEXT: .cfi_offset %r12, -48
+; FAST_RA-NEXT: .cfi_offset %r13, -40
+; FAST_RA-NEXT: .cfi_offset %r14, -32
+; FAST_RA-NEXT: .cfi_offset %r15, -24
+; FAST_RA-NEXT: .cfi_offset %rbp, -16
+; FAST_RA-NEXT: movq %rcx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; FAST_RA-NEXT: movq %rdi, %rax
+; FAST_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rdi # 8-byte Reload
+; FAST_RA-NEXT: movq {{[0-9]+}}(%rsp), %rcx
+; FAST_RA-NEXT: movq %rax, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; FAST_RA-NEXT: #APP # 8-byte Folded Reload
+; FAST_RA-NEXT: # %rcx -{{[0-9]+}}(%rsp) %rsi %rdx %rdi %r8 %r9 %rcx
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: movq %rcx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; FAST_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rax # 8-byte Reload
+; FAST_RA-NEXT: popq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: popq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: popq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: popq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: popq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: popq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 8
+; FAST_RA-NEXT: retq
+entry:
+ %r = call i64 asm sideeffect "# $0 $1 $2 $3 $4 $5 $6 $7",
+ "=r,rm,r,r,r,r,r,0,~{rax},~{rbx},~{rbp},~{r10},~{r11},~{r12},~{r13},~{r14},~{r15}"
+ (i64 %x, i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %y)
+ ret i64 %r
+}
+
+; Two read-write "+rm" operands: folding one tied pair must keep the other
+; pair tied.
+define { i64, i64 } @test_two_tied_rm_pressure(i64 %x, i64 %y, i64 %a, i64 %b, i64 %c, i64 %d, i64 %e) {
+; GREEDY_RA-LABEL: test_two_tied_rm_pressure:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: pushq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: pushq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: pushq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: pushq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: pushq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: pushq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 56
+; GREEDY_RA-NEXT: .cfi_offset %rbx, -56
+; GREEDY_RA-NEXT: .cfi_offset %r12, -48
+; GREEDY_RA-NEXT: .cfi_offset %r13, -40
+; GREEDY_RA-NEXT: .cfi_offset %r14, -32
+; GREEDY_RA-NEXT: .cfi_offset %r15, -24
+; GREEDY_RA-NEXT: .cfi_offset %rbp, -16
+; GREEDY_RA-NEXT: movq %r9, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %r8, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rcx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rdx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rsi, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rdi, %rdx
+; GREEDY_RA-NEXT: movq {{[0-9]+}}(%rsp), %rsi
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %r9 # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %r8 # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rcx # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rdi # 8-byte Reload
+; GREEDY_RA-NEXT: #APP # 8-byte Folded Reload
+; GREEDY_RA-NEXT: # %rdx -{{[0-9]+}}(%rsp) %rdi %rcx %r8 %r9 %rsi
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: movq %rdx, %rax
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rdx # 8-byte Reload
+; GREEDY_RA-NEXT: popq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: popq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: popq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: popq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: popq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: popq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 8
+; GREEDY_RA-NEXT: retq
+;
+; FAST_RA-LABEL: test_two_tied_rm_pressure:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: pushq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: pushq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: pushq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: pushq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: pushq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: pushq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 56
+; FAST_RA-NEXT: .cfi_offset %rbx, -56
+; FAST_RA-NEXT: .cfi_offset %r12, -48
+; FAST_RA-NEXT: .cfi_offset %r13, -40
+; FAST_RA-NEXT: .cfi_offset %r14, -32
+; FAST_RA-NEXT: .cfi_offset %r15, -24
+; FAST_RA-NEXT: .cfi_offset %rbp, -16
+; FAST_RA-NEXT: movq %rcx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; FAST_RA-NEXT: movq %rdx, %rcx
+; FAST_RA-NEXT: movq %rsi, %rdx
+; FAST_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rsi # 8-byte Reload
+; FAST_RA-NEXT: movq %rdi, %rax
+; FAST_RA-NEXT: movq {{[0-9]+}}(%rsp), %rdi
+; FAST_RA-NEXT: movq %rax, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; FAST_RA-NEXT: #APP # 8-byte Folded Reload
+; FAST_RA-NEXT: # -{{[0-9]+}}(%rsp) %rdx %rcx %rsi %r8 %r9 %rdi
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rax # 8-byte Reload
+; FAST_RA-NEXT: popq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: popq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: popq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: popq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: popq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: popq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 8
+; FAST_RA-NEXT: retq
+entry:
+ %r = call { i64, i64 } asm sideeffect "# $0 $1 $2 $3 $4 $5 $6",
+ "=rm,=rm,r,r,r,r,r,0,1,~{rax},~{rbx},~{rbp},~{r10},~{r11},~{r12},~{r13},~{r14},~{r15}"
+ (i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %x, i64 %y)
+ ret { i64, i64 } %r
+}
+
+; Like test_rm_output_pressure, but the output isn't early-clobber, so it can
+; share a register with an input and nothing needs to be folded.
+define i64 @test_rm_output_shares_input_register(i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %f) {
+; GREEDY_RA-LABEL: test_rm_output_shares_input_register:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: pushq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: pushq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: pushq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: pushq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: pushq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: pushq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 56
+; GREEDY_RA-NEXT: .cfi_offset %rbx, -56
+; GREEDY_RA-NEXT: .cfi_offset %r12, -48
+; GREEDY_RA-NEXT: .cfi_offset %r13, -40
+; GREEDY_RA-NEXT: .cfi_offset %r14, -32
+; GREEDY_RA-NEXT: .cfi_offset %r15, -24
+; GREEDY_RA-NEXT: .cfi_offset %rbp, -16
+; GREEDY_RA-NEXT: #APP
+; GREEDY_RA-NEXT: # %rcx %rdi %rsi %rdx %rcx %r8 %r9
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: movq %rcx, %rax
+; GREEDY_RA-NEXT: popq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: popq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: popq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: popq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: popq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: popq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 8
+; GREEDY_RA-NEXT: retq
+;
+; FAST_RA-LABEL: test_rm_output_shares_input_register:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: pushq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: pushq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: pushq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: pushq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: pushq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: pushq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 56
+; FAST_RA-NEXT: .cfi_offset %rbx, -56
+; FAST_RA-NEXT: .cfi_offset %r12, -48
+; FAST_RA-NEXT: .cfi_offset %r13, -40
+; FAST_RA-NEXT: .cfi_offset %r14, -32
+; FAST_RA-NEXT: .cfi_offset %r15, -24
+; FAST_RA-NEXT: .cfi_offset %rbp, -16
+; FAST_RA-NEXT: #APP
+; FAST_RA-NEXT: # %rcx %rdi %rsi %rdx %rcx %r8 %r9
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: movq %rcx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; FAST_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rax # 8-byte Reload
+; FAST_RA-NEXT: popq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: popq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: popq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: popq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: popq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: popq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 8
+; FAST_RA-NEXT: retq
+entry:
+ %0 = call i64 asm sideeffect "# $0 $1 $2 $3 $4 $5 $6",
+ "=rm,r,r,r,r,r,r,~{rax},~{rbx},~{rbp},~{r10},~{r11},~{r12},~{r13},~{r14},~{r15}"
+ (i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %f)
+ ret i64 %0
+}
+
+; Eleven inputs for the ten GPRs left after the clobbers: one "rm" input has
+; to be folded even though the 32-bit inputs alone would fit, because each
+; 64-bit input takes a 32-bit register too.
+define void @test_rm_mixed_widths_pressure(i32 %x0, i32 %x1, i32 %x2, i32 %x3, i32 %x4, i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %f) {
+; GREEDY_RA-LABEL: test_rm_mixed_widths_pressure:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: pushq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: pushq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: pushq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: pushq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: pushq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: pushq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 56
+; GREEDY_RA-NEXT: .cfi_offset %rbx, -56
+; GREEDY_RA-NEXT: .cfi_offset %r12, -48
+; GREEDY_RA-NEXT: .cfi_offset %r13, -40
+; GREEDY_RA-NEXT: .cfi_offset %r14, -32
+; GREEDY_RA-NEXT: .cfi_offset %r15, -24
+; GREEDY_RA-NEXT: .cfi_offset %rbp, -16
+; GREEDY_RA-NEXT: movq %r9, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movl %r8d, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; GREEDY_RA-NEXT: movq {{[0-9]+}}(%rsp), %rax
+; GREEDY_RA-NEXT: movq {{[0-9]+}}(%rsp), %r10
+; GREEDY_RA-NEXT: movq {{[0-9]+}}(%rsp), %r11
+; GREEDY_RA-NEXT: movq {{[0-9]+}}(%rsp), %r15
+; GREEDY_RA-NEXT: movq {{[0-9]+}}(%rsp), %r9
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %r8 # 8-byte Reload
+; GREEDY_RA-NEXT: #APP # 4-byte Folded Reload
+; GREEDY_RA-NEXT: # %edi %esi %edx %ecx -{{[0-9]+}}(%rsp) %r8 %r10 %r11 %r15 %r9 %rax
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: popq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: popq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: popq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: popq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: popq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: popq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 8
+; GREEDY_RA-NEXT: retq
+;
+; FAST_RA-LABEL: test_rm_mixed_widths_pressure:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: pushq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: pushq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: pushq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: pushq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: pushq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: pushq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 56
+; FAST_RA-NEXT: .cfi_offset %rbx, -56
+; FAST_RA-NEXT: .cfi_offset %r12, -48
+; FAST_RA-NEXT: .cfi_offset %r13, -40
+; FAST_RA-NEXT: .cfi_offset %r14, -32
+; FAST_RA-NEXT: .cfi_offset %r15, -24
+; FAST_RA-NEXT: .cfi_offset %rbp, -16
+; FAST_RA-NEXT: movl %edi, %ebx
+; FAST_RA-NEXT: movq {{[0-9]+}}(%rsp), %rax
+; FAST_RA-NEXT: movq {{[0-9]+}}(%rsp), %rdi
+; FAST_RA-NEXT: movq {{[0-9]+}}(%rsp), %r10
+; FAST_RA-NEXT: movq {{[0-9]+}}(%rsp), %r11
+; FAST_RA-NEXT: movq {{[0-9]+}}(%rsp), %r15
+; FAST_RA-NEXT: movl %ebx, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_RA-NEXT: #APP # 4-byte Folded Reload
+; FAST_RA-NEXT: # -{{[0-9]+}}(%rsp) %esi %edx %ecx %r8d %r9 %rax %rdi %r10 %r11 %r15
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: popq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: popq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: popq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: popq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: popq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: popq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 8
+; FAST_RA-NEXT: retq
+entry:
+ call void asm sideeffect "# $0 $1 $2 $3 $4 $5 $6 $7 $8 $9 $10",
+ "rm,rm,rm,rm,rm,r,r,r,r,r,r,~{rbx},~{rbp},~{r12},~{r13},~{r14}"
+ (i32 %x0, i32 %x1, i32 %x2, i32 %x3, i32 %x4, i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %f)
+ ret void
+}
+
+; Seven values for six GPRs. Folding the "rm" operand of %x wouldn't free a
+; register, since the "r" operand needs %x in one anyway, so fold %y's.
+define void @test_rm_and_r_same_value_pressure(i64 %x, i64 %y, i64 %a, i64 %b, i64 %c, i64 %d, i64 %e) {
+; GREEDY_RA-LABEL: test_rm_and_r_same_value_pressure:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: pushq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: pushq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: pushq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: pushq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: pushq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: pushq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 56
+; GREEDY_RA-NEXT: .cfi_offset %rbx, -56
+; GREEDY_RA-NEXT: .cfi_offset %r12, -48
+; GREEDY_RA-NEXT: .cfi_offset %r13, -40
+; GREEDY_RA-NEXT: .cfi_offset %r14, -32
+; GREEDY_RA-NEXT: .cfi_offset %r15, -24
+; GREEDY_RA-NEXT: .cfi_offset %rbp, -16
+; GREEDY_RA-NEXT: movq %r9, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %r8, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rcx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rdx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rsi, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq {{[0-9]+}}(%rsp), %r9
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %r8 # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rcx # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rdx # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rsi # 8-byte Reload
+; GREEDY_RA-NEXT: #APP # 8-byte Folded Reload
+; GREEDY_RA-NEXT: # %rdi %rdi -{{[0-9]+}}(%rsp) %rsi %rdx %rcx %r8 %r9
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: popq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: popq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: popq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: popq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: popq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: popq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 8
+; GREEDY_RA-NEXT: retq
+;
+; FAST_RA-LABEL: test_rm_and_r_same_value_pressure:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: pushq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: pushq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: pushq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: pushq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: pushq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: pushq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 56
+; FAST_RA-NEXT: .cfi_offset %rbx, -56
+; FAST_RA-NEXT: .cfi_offset %r12, -48
+; FAST_RA-NEXT: .cfi_offset %r13, -40
+; FAST_RA-NEXT: .cfi_offset %r14, -32
+; FAST_RA-NEXT: .cfi_offset %r15, -24
+; FAST_RA-NEXT: .cfi_offset %rbp, -16
+; FAST_RA-NEXT: movq %rsi, %rax
+; FAST_RA-NEXT: movq {{[0-9]+}}(%rsp), %rsi
+; FAST_RA-NEXT: movq %rax, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; FAST_RA-NEXT: #APP # 8-byte Folded Reload
+; FAST_RA-NEXT: # %rdi %rdi -{{[0-9]+}}(%rsp) %rdx %rcx %r8 %r9 %rsi
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: popq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: popq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: popq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: popq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: popq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: popq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 8
+; FAST_RA-NEXT: retq
+entry:
+ call void asm sideeffect "# $0 $1 $2 $3 $4 $5 $6 $7",
+ "rm,r,rm,r,r,r,r,r,~{rax},~{rbx},~{rbp},~{r10},~{r11},~{r12},~{r13},~{r14},~{r15}"
+ (i64 %x, i64 %x, i64 %y, i64 %a, i64 %b, i64 %c, i64 %d, i64 %e)
+ ret void
+}
diff --git a/llvm/test/CodeGen/X86/asm-constraints-rm.ll b/llvm/test/CodeGen/X86/asm-constraints-rm.ll
new file mode 100644
index 00000000000000..dc1cee8f14c06c
--- /dev/null
+++ b/llvm/test/CodeGen/X86/asm-constraints-rm.ll
@@ -0,0 +1,35 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 6
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --regalloc=greedy < %s | FileCheck --check-prefix=GREEDY_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --regalloc=fast < %s | FileCheck --check-prefix=FAST_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu -O0 < %s | FileCheck --check-prefix=O0 %s
+
+; Above -O0, an "rm" input picks a register that the register allocator may
+; fold to a stack slot (see AsmOperandInfo::MayFoldRegister), instead of
+; always picking memory. With no register pressure, both allocators keep the
+; value in a register. -O0 still picks memory.
+define i64 @test_rm_input_no_pressure(i64 %a) {
+; GREEDY_RA-LABEL: test_rm_input_no_pressure:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: #APP
+; GREEDY_RA-NEXT: bsfq %rdi, %rax
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: retq
+;
+; FAST_RA-LABEL: test_rm_input_no_pressure:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: #APP
+; FAST_RA-NEXT: bsfq %rdi, %rax
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: retq
+;
+; O0-LABEL: test_rm_input_no_pressure:
+; O0: # %bb.0: # %entry
+; O0-NEXT: movq %rdi, -{{[0-9]+}}(%rsp)
+; O0-NEXT: #APP
+; O0-NEXT: bsfq -{{[0-9]+}}(%rsp), %rax
+; O0-NEXT: #NO_APP
+; O0-NEXT: retq
+entry:
+ %0 = call i64 asm "bsfq $1,$0", "=r,rm"(i64 %a)
+ ret i64 %0
+}
diff --git a/llvm/test/CodeGen/X86/asm-constraints-torture.ll b/llvm/test/CodeGen/X86/asm-constraints-torture.ll
new file mode 100644
index 00000000000000..b8b9b0b4c9e362
--- /dev/null
+++ b/llvm/test/CodeGen/X86/asm-constraints-torture.ll
@@ -0,0 +1,767 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --filter "^\t#" --version 4
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=true --regalloc=greedy < %s \
+; RUN: | FileCheck --check-prefix=FAST_ISEL_GREEDY_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=true --regalloc=fast < %s \
+; RUN: | FileCheck --check-prefix=FAST_ISEL_FAST_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=false --regalloc=greedy < %s \
+; RUN: | FileCheck --check-prefix=GREEDY_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=false --regalloc=fast < %s \
+; RUN: | FileCheck --check-prefix=FAST_RA %s
+
+; Combinatorial coverage for "rm" inline asm operands: inputs, outputs, tied
+; read-write operands, and mixes with plain "r"/"m" operands and early-clobber
+; outputs. Each comes without and with pressure from values live across the
+; asm, which the register allocators spill, so every "rm" operand stays in a
+; register. asm-constraints-rm-pressure.ll covers operands that have to be
+; folded to memory.
+
+define dso_local i32 @test_rm_input_no_pressure(ptr noundef readonly captures(none) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_rm_input_no_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm input: no pressure
+; FAST_ISEL_GREEDY_RA: # %eax, %ecx, %edx, %esi, %r8d
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_rm_input_no_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm input: no pressure
+; FAST_ISEL_FAST_RA: # %eax, %ecx, %edx, %esi, %r8d
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_rm_input_no_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm input: no pressure
+; GREEDY_RA: # %eax, %ecx, %edx, %esi, %r8d
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_rm_input_no_pressure:
+; FAST_RA: #APP
+; FAST_RA: # rm input: no pressure
+; FAST_RA: # %eax, %ecx, %edx, %esi, %r8d
+; FAST_RA: #NO_APP
+entry:
+ %0 = load i32, ptr %foo, align 4
+ %b2 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %1 = load i32, ptr %b2, align 4
+ %c3 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %2 = load i32, ptr %c3, align 4
+ %d4 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %3 = load i32, ptr %d4, align 4
+ %e5 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %4 = load i32, ptr %e5, align 4
+ tail call void asm sideeffect "# rm input: no pressure\0A\09# $0, $1, $2, $3, $4", "rm,rm,rm,rm,rm,~{dirflag},~{fpsr},~{flags}"(i32 %0, i32 %1, i32 %2, i32 %3, i32 %4)
+ %5 = load i32, ptr %foo, align 4
+ ret i32 %5
+}
+
+define dso_local i32 @test_rm_pressure(ptr noundef readonly captures(none) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_rm_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_GREEDY_RA: #NO_APP
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm input: pressure
+; FAST_ISEL_GREEDY_RA: # %esi, %edi, %r8d, %r9d, %eax
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_rm_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_FAST_RA: #NO_APP
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm input: pressure
+; FAST_ISEL_FAST_RA: # %eax, %ecx, %edx, %esi, %edi
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_rm_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; GREEDY_RA: #NO_APP
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm input: pressure
+; GREEDY_RA: # %esi, %edi, %r8d, %r9d, %eax
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_rm_pressure:
+; FAST_RA: #APP
+; FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_RA: #NO_APP
+; FAST_RA: #APP
+; FAST_RA: # rm input: pressure
+; FAST_RA: # %eax, %ecx, %edx, %esi, %edi
+; FAST_RA: #NO_APP
+entry:
+ %0 = tail call { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } asm sideeffect "# $0,$1,$2,$3,$4,$5,$6,$7,$8,$9,$10, $11, $12, $13, $14", "={rax},={rcx},={rdx},={rsi},={rdi},={rbx},={rbp},={r8},={r9},={r10},={r11},={r12},={r13},={r14},={r15},~{dirflag},~{fpsr},~{flags}"()
+ %1 = load i32, ptr %foo, align 4
+ %b16 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %2 = load i32, ptr %b16, align 4
+ %c17 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %3 = load i32, ptr %c17, align 4
+ %d18 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %4 = load i32, ptr %d18, align 4
+ %e19 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %5 = load i32, ptr %e19, align 4
+ tail call void asm sideeffect "# rm input: pressure\0A\09# $0, $1, $2, $3, $4", "rm,rm,rm,rm,rm,~{dirflag},~{fpsr},~{flags}"(i32 %1, i32 %2, i32 %3, i32 %4, i32 %5)
+ %asmresult14 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 14
+ %asmresult13 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 13
+ %asmresult12 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 12
+ %asmresult11 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 11
+ %asmresult10 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 10
+ %asmresult9 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 9
+ %asmresult8 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 8
+ %asmresult7 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 7
+ %asmresult6 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 6
+ %asmresult5 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 5
+ %asmresult4 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 4
+ %asmresult3 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 3
+ %asmresult2 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 2
+ %asmresult1 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 1
+ %asmresult = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 0
+ tail call void @g(i64 noundef %asmresult, i64 noundef %asmresult1, i64 noundef %asmresult2, i64 noundef %asmresult3, i64 noundef %asmresult4, i64 noundef %asmresult5, i64 noundef %asmresult6, i64 noundef %asmresult7, i64 noundef %asmresult8, i64 noundef %asmresult9, i64 noundef %asmresult10, i64 noundef %asmresult11, i64 noundef %asmresult12, i64 noundef %asmresult13, i64 noundef %asmresult14)
+ %6 = load i32, ptr %foo, align 4
+ ret i32 %6
+}
+
+define dso_local i32 @test_output_no_pressure(ptr noundef writeonly captures(none) initializes((0, 20)) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_output_no_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm output: no pressure
+; FAST_ISEL_GREEDY_RA: # %eax, %ecx, %edx, %esi, %r8d
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_output_no_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm output: no pressure
+; FAST_ISEL_FAST_RA: # %eax, %r8d, %esi, %edx, %ecx
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_output_no_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm output: no pressure
+; GREEDY_RA: # %eax, %ecx, %edx, %esi, %r8d
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_output_no_pressure:
+; FAST_RA: #APP
+; FAST_RA: # rm output: no pressure
+; FAST_RA: # %eax, %r8d, %esi, %edx, %ecx
+; FAST_RA: #NO_APP
+entry:
+ %b2 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %c3 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %d4 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %e5 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %0 = tail call { i32, i32, i32, i32, i32 } asm sideeffect "# rm output: no pressure\0A\09# $0, $1, $2, $3, $4", "=rm,=rm,=rm,=rm,=rm,~{dirflag},~{fpsr},~{flags}"()
+ %asmresult = extractvalue { i32, i32, i32, i32, i32 } %0, 0
+ %asmresult6 = extractvalue { i32, i32, i32, i32, i32 } %0, 1
+ %asmresult7 = extractvalue { i32, i32, i32, i32, i32 } %0, 2
+ %asmresult8 = extractvalue { i32, i32, i32, i32, i32 } %0, 3
+ %asmresult9 = extractvalue { i32, i32, i32, i32, i32 } %0, 4
+ store i32 %asmresult, ptr %foo, align 4
+ store i32 %asmresult6, ptr %b2, align 4
+ store i32 %asmresult7, ptr %c3, align 4
+ store i32 %asmresult8, ptr %d4, align 4
+ store i32 %asmresult9, ptr %e5, align 4
+ ret i32 %asmresult
+}
+
+define dso_local i32 @test_output_pressure(ptr noundef writeonly captures(none) initializes((0, 20)) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_output_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_GREEDY_RA: #NO_APP
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm output: pressure
+; FAST_ISEL_GREEDY_RA: # %esi, %edi, %r8d, %r9d, %eax
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_output_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_FAST_RA: #NO_APP
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm output: pressure
+; FAST_ISEL_FAST_RA: # %eax, %edi, %ecx, %edx, %esi
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_output_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; GREEDY_RA: #NO_APP
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm output: pressure
+; GREEDY_RA: # %esi, %edi, %r8d, %r9d, %eax
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_output_pressure:
+; FAST_RA: #APP
+; FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_RA: #NO_APP
+; FAST_RA: #APP
+; FAST_RA: # rm output: pressure
+; FAST_RA: # %eax, %edi, %ecx, %edx, %esi
+; FAST_RA: #NO_APP
+entry:
+ %0 = tail call { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } asm sideeffect "# $0,$1,$2,$3,$4,$5,$6,$7,$8,$9,$10, $11, $12, $13, $14", "={rax},={rcx},={rdx},={rsi},={rdi},={rbx},={rbp},={r8},={r9},={r10},={r11},={r12},={r13},={r14},={r15},~{dirflag},~{fpsr},~{flags}"()
+ %b16 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %c17 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %d18 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %e19 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %1 = tail call { i32, i32, i32, i32, i32 } asm sideeffect "# rm output: pressure\0A\09# $0, $1, $2, $3, $4", "=rm,=rm,=rm,=rm,=rm,~{dirflag},~{fpsr},~{flags}"()
+ %asmresult20 = extractvalue { i32, i32, i32, i32, i32 } %1, 0
+ %asmresult21 = extractvalue { i32, i32, i32, i32, i32 } %1, 1
+ %asmresult22 = extractvalue { i32, i32, i32, i32, i32 } %1, 2
+ %asmresult23 = extractvalue { i32, i32, i32, i32, i32 } %1, 3
+ %asmresult24 = extractvalue { i32, i32, i32, i32, i32 } %1, 4
+ store i32 %asmresult20, ptr %foo, align 4
+ store i32 %asmresult21, ptr %b16, align 4
+ store i32 %asmresult22, ptr %c17, align 4
+ store i32 %asmresult23, ptr %d18, align 4
+ store i32 %asmresult24, ptr %e19, align 4
+ %asmresult14 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 14
+ %asmresult13 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 13
+ %asmresult12 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 12
+ %asmresult11 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 11
+ %asmresult10 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 10
+ %asmresult9 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 9
+ %asmresult8 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 8
+ %asmresult7 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 7
+ %asmresult6 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 6
+ %asmresult5 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 5
+ %asmresult4 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 4
+ %asmresult3 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 3
+ %asmresult2 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 2
+ %asmresult1 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 1
+ %asmresult = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 0
+ tail call void @g(i64 noundef %asmresult, i64 noundef %asmresult1, i64 noundef %asmresult2, i64 noundef %asmresult3, i64 noundef %asmresult4, i64 noundef %asmresult5, i64 noundef %asmresult6, i64 noundef %asmresult7, i64 noundef %asmresult8, i64 noundef %asmresult9, i64 noundef %asmresult10, i64 noundef %asmresult11, i64 noundef %asmresult12, i64 noundef %asmresult13, i64 noundef %asmresult14)
+ %2 = load i32, ptr %foo, align 4
+ ret i32 %2
+}
+
+define dso_local i32 @test_tied_output_no_pressure(ptr noundef captures(none) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_tied_output_no_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm tied output: no pressure
+; FAST_ISEL_GREEDY_RA: # %eax, %ecx, %edx, %esi, %r8d
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_tied_output_no_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm tied output: no pressure
+; FAST_ISEL_FAST_RA: # %eax, %r8d, %esi, %edx, %ecx
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_tied_output_no_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm tied output: no pressure
+; GREEDY_RA: # %eax, %ecx, %edx, %esi, %r8d
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_tied_output_no_pressure:
+; FAST_RA: #APP
+; FAST_RA: # rm tied output: no pressure
+; FAST_RA: # %eax, %r8d, %esi, %edx, %ecx
+; FAST_RA: #NO_APP
+entry:
+ %0 = load i32, ptr %foo, align 4
+ %b2 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %1 = load i32, ptr %b2, align 4
+ %c3 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %2 = load i32, ptr %c3, align 4
+ %d4 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %3 = load i32, ptr %d4, align 4
+ %e5 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %4 = load i32, ptr %e5, align 4
+ %5 = tail call { i32, i32, i32, i32, i32 } asm sideeffect "# rm tied output: no pressure\0A\09# $0, $1, $2, $3, $4", "=rm,=rm,=rm,=rm,=rm,0,1,2,3,4,~{dirflag},~{fpsr},~{flags}"(i32 %0, i32 %1, i32 %2, i32 %3, i32 %4)
+ %asmresult = extractvalue { i32, i32, i32, i32, i32 } %5, 0
+ %asmresult6 = extractvalue { i32, i32, i32, i32, i32 } %5, 1
+ %asmresult7 = extractvalue { i32, i32, i32, i32, i32 } %5, 2
+ %asmresult8 = extractvalue { i32, i32, i32, i32, i32 } %5, 3
+ %asmresult9 = extractvalue { i32, i32, i32, i32, i32 } %5, 4
+ store i32 %asmresult, ptr %foo, align 4
+ store i32 %asmresult6, ptr %b2, align 4
+ store i32 %asmresult7, ptr %c3, align 4
+ store i32 %asmresult8, ptr %d4, align 4
+ store i32 %asmresult9, ptr %e5, align 4
+ ret i32 %asmresult
+}
+
+define dso_local i32 @test_tied_output_pressure(ptr noundef captures(none) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_tied_output_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_GREEDY_RA: #NO_APP
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm tied output: pressure
+; FAST_ISEL_GREEDY_RA: # %esi, %edi, %r8d, %r9d, %eax
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_tied_output_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_FAST_RA: #NO_APP
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm tied output: pressure
+; FAST_ISEL_FAST_RA: # %eax, %edi, %ecx, %edx, %esi
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_tied_output_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; GREEDY_RA: #NO_APP
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm tied output: pressure
+; GREEDY_RA: # %esi, %edi, %r8d, %r9d, %eax
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_tied_output_pressure:
+; FAST_RA: #APP
+; FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_RA: #NO_APP
+; FAST_RA: #APP
+; FAST_RA: # rm tied output: pressure
+; FAST_RA: # %eax, %edi, %ecx, %edx, %esi
+; FAST_RA: #NO_APP
+entry:
+ %0 = tail call { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } asm sideeffect "# $0,$1,$2,$3,$4,$5,$6,$7,$8,$9,$10, $11, $12, $13, $14", "={rax},={rcx},={rdx},={rsi},={rdi},={rbx},={rbp},={r8},={r9},={r10},={r11},={r12},={r13},={r14},={r15},~{dirflag},~{fpsr},~{flags}"()
+ %1 = load i32, ptr %foo, align 4
+ %b16 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %2 = load i32, ptr %b16, align 4
+ %c17 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %3 = load i32, ptr %c17, align 4
+ %d18 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %4 = load i32, ptr %d18, align 4
+ %e19 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %5 = load i32, ptr %e19, align 4
+ %6 = tail call { i32, i32, i32, i32, i32 } asm sideeffect "# rm tied output: pressure\0A\09# $0, $1, $2, $3, $4", "=rm,=rm,=rm,=rm,=rm,0,1,2,3,4,~{dirflag},~{fpsr},~{flags}"(i32 %1, i32 %2, i32 %3, i32 %4, i32 %5)
+ %asmresult20 = extractvalue { i32, i32, i32, i32, i32 } %6, 0
+ %asmresult21 = extractvalue { i32, i32, i32, i32, i32 } %6, 1
+ %asmresult22 = extractvalue { i32, i32, i32, i32, i32 } %6, 2
+ %asmresult23 = extractvalue { i32, i32, i32, i32, i32 } %6, 3
+ %asmresult24 = extractvalue { i32, i32, i32, i32, i32 } %6, 4
+ store i32 %asmresult20, ptr %foo, align 4
+ store i32 %asmresult21, ptr %b16, align 4
+ store i32 %asmresult22, ptr %c17, align 4
+ store i32 %asmresult23, ptr %d18, align 4
+ store i32 %asmresult24, ptr %e19, align 4
+ %asmresult14 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 14
+ %asmresult13 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 13
+ %asmresult12 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 12
+ %asmresult11 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 11
+ %asmresult10 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 10
+ %asmresult9 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 9
+ %asmresult8 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 8
+ %asmresult7 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 7
+ %asmresult6 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 6
+ %asmresult5 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 5
+ %asmresult4 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 4
+ %asmresult3 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 3
+ %asmresult2 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 2
+ %asmresult1 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 1
+ %asmresult = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 0
+ tail call void @g(i64 noundef %asmresult, i64 noundef %asmresult1, i64 noundef %asmresult2, i64 noundef %asmresult3, i64 noundef %asmresult4, i64 noundef %asmresult5, i64 noundef %asmresult6, i64 noundef %asmresult7, i64 noundef %asmresult8, i64 noundef %asmresult9, i64 noundef %asmresult10, i64 noundef %asmresult11, i64 noundef %asmresult12, i64 noundef %asmresult13, i64 noundef %asmresult14)
+ %7 = load i32, ptr %foo, align 4
+ ret i32 %7
+}
+
+define dso_local i32 @test_rm_output_r_input_no_pressure(ptr noundef captures(none) initializes((0, 4)) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_rm_output_r_input_no_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm output, r input: no pressure
+; FAST_ISEL_GREEDY_RA: # %eax, %eax
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_rm_output_r_input_no_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm output, r input: no pressure
+; FAST_ISEL_FAST_RA: # %eax, %eax
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_rm_output_r_input_no_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm output, r input: no pressure
+; GREEDY_RA: # %eax, %eax
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_rm_output_r_input_no_pressure:
+; FAST_RA: #APP
+; FAST_RA: # rm output, r input: no pressure
+; FAST_RA: # %eax, %eax
+; FAST_RA: #NO_APP
+entry:
+ %b2 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %0 = load i32, ptr %b2, align 4
+ %1 = tail call i32 asm sideeffect "# rm output, r input: no pressure\0A\09# $0, $1", "=rm,r,~{dirflag},~{fpsr},~{flags}"(i32 %0)
+ store i32 %1, ptr %foo, align 4
+ ret i32 %1
+}
+
+define dso_local i32 @test_rm_output_r_input_pressure(ptr noundef captures(none) initializes((0, 4)) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_rm_output_r_input_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_GREEDY_RA: #NO_APP
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm output, r input: pressure
+; FAST_ISEL_GREEDY_RA: # %esi, %esi
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_rm_output_r_input_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_FAST_RA: #NO_APP
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm output, r input: pressure
+; FAST_ISEL_FAST_RA: # %eax, %eax
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_rm_output_r_input_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; GREEDY_RA: #NO_APP
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm output, r input: pressure
+; GREEDY_RA: # %esi, %esi
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_rm_output_r_input_pressure:
+; FAST_RA: #APP
+; FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_RA: #NO_APP
+; FAST_RA: #APP
+; FAST_RA: # rm output, r input: pressure
+; FAST_RA: # %eax, %eax
+; FAST_RA: #NO_APP
+entry:
+ %0 = tail call { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } asm sideeffect "# $0,$1,$2,$3,$4,$5,$6,$7,$8,$9,$10, $11, $12, $13, $14", "={rax},={rcx},={rdx},={rsi},={rdi},={rbx},={rbp},={r8},={r9},={r10},={r11},={r12},={r13},={r14},={r15},~{dirflag},~{fpsr},~{flags}"()
+ %b16 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %1 = load i32, ptr %b16, align 4
+ %2 = tail call i32 asm sideeffect "# rm output, r input: pressure\0A\09# $0, $1", "=rm,r,~{dirflag},~{fpsr},~{flags}"(i32 %1)
+ store i32 %2, ptr %foo, align 4
+ %asmresult14 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 14
+ %asmresult13 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 13
+ %asmresult12 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 12
+ %asmresult11 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 11
+ %asmresult10 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 10
+ %asmresult9 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 9
+ %asmresult8 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 8
+ %asmresult7 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 7
+ %asmresult6 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 6
+ %asmresult5 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 5
+ %asmresult4 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 4
+ %asmresult3 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 3
+ %asmresult2 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 2
+ %asmresult1 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 1
+ %asmresult = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 0
+ tail call void @g(i64 noundef %asmresult, i64 noundef %asmresult1, i64 noundef %asmresult2, i64 noundef %asmresult3, i64 noundef %asmresult4, i64 noundef %asmresult5, i64 noundef %asmresult6, i64 noundef %asmresult7, i64 noundef %asmresult8, i64 noundef %asmresult9, i64 noundef %asmresult10, i64 noundef %asmresult11, i64 noundef %asmresult12, i64 noundef %asmresult13, i64 noundef %asmresult14)
+ %3 = load i32, ptr %foo, align 4
+ ret i32 %3
+}
+
+define dso_local i32 @test_m_output_rm_input_no_pressure(ptr noundef %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_m_output_rm_input_no_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # m output, rm input: no pressure
+; FAST_ISEL_GREEDY_RA: # (%rdi), %eax
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_m_output_rm_input_no_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # m output, rm input: no pressure
+; FAST_ISEL_FAST_RA: # (%rdi), %eax
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_m_output_rm_input_no_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # m output, rm input: no pressure
+; GREEDY_RA: # (%rdi), %eax
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_m_output_rm_input_no_pressure:
+; FAST_RA: #APP
+; FAST_RA: # m output, rm input: no pressure
+; FAST_RA: # (%rdi), %eax
+; FAST_RA: #NO_APP
+entry:
+ %b2 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %0 = load i32, ptr %b2, align 4
+ tail call void asm sideeffect "# m output, rm input: no pressure\0A\09# $0, $1", "=*m,rm,~{dirflag},~{fpsr},~{flags}"(ptr elementtype(i32) %foo, i32 %0)
+ %1 = load i32, ptr %foo, align 4
+ ret i32 %1
+}
+
+define dso_local i32 @test_m_output_rm_input_pressure(ptr noundef %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_m_output_rm_input_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_GREEDY_RA: #NO_APP
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # m output, rm input: pressure
+; FAST_ISEL_GREEDY_RA: # (%rbp), %esi
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_m_output_rm_input_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_FAST_RA: #NO_APP
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # m output, rm input: pressure
+; FAST_ISEL_FAST_RA: # (%rdi), %eax
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_m_output_rm_input_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; GREEDY_RA: #NO_APP
+; GREEDY_RA: #APP
+; GREEDY_RA: # m output, rm input: pressure
+; GREEDY_RA: # (%rbp), %esi
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_m_output_rm_input_pressure:
+; FAST_RA: #APP
+; FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_RA: #NO_APP
+; FAST_RA: #APP
+; FAST_RA: # m output, rm input: pressure
+; FAST_RA: # (%rdi), %eax
+; FAST_RA: #NO_APP
+entry:
+ %0 = tail call { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } asm sideeffect "# $0,$1,$2,$3,$4,$5,$6,$7,$8,$9,$10, $11, $12, $13, $14", "={rax},={rcx},={rdx},={rsi},={rdi},={rbx},={rbp},={r8},={r9},={r10},={r11},={r12},={r13},={r14},={r15},~{dirflag},~{fpsr},~{flags}"()
+ %b16 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %1 = load i32, ptr %b16, align 4
+ tail call void asm sideeffect "# m output, rm input: pressure\0A\09# $0, $1", "=*m,rm,~{dirflag},~{fpsr},~{flags}"(ptr elementtype(i32) %foo, i32 %1)
+ %asmresult14 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 14
+ %asmresult13 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 13
+ %asmresult12 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 12
+ %asmresult11 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 11
+ %asmresult10 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 10
+ %asmresult9 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 9
+ %asmresult8 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 8
+ %asmresult7 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 7
+ %asmresult6 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 6
+ %asmresult5 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 5
+ %asmresult4 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 4
+ %asmresult3 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 3
+ %asmresult2 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 2
+ %asmresult1 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 1
+ %asmresult = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 0
+ tail call void @g(i64 noundef %asmresult, i64 noundef %asmresult1, i64 noundef %asmresult2, i64 noundef %asmresult3, i64 noundef %asmresult4, i64 noundef %asmresult5, i64 noundef %asmresult6, i64 noundef %asmresult7, i64 noundef %asmresult8, i64 noundef %asmresult9, i64 noundef %asmresult10, i64 noundef %asmresult11, i64 noundef %asmresult12, i64 noundef %asmresult13, i64 noundef %asmresult14)
+ %2 = load i32, ptr %foo, align 4
+ ret i32 %2
+}
+
+define dso_local i32 @test_mult_m_output_rm_input_no_pressure(ptr noundef %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_mult_m_output_rm_input_no_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # multiple m output, rm input: no pressure
+; FAST_ISEL_GREEDY_RA: # (%rdi), (%rax), (%rcx), (%rdx), (%rsi), %r8d, %r9d
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_mult_m_output_rm_input_no_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # multiple m output, rm input: no pressure
+; FAST_ISEL_FAST_RA: # (%rdi), (%rax), (%rcx), (%rdx), (%rsi), %r8d, %r9d
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_mult_m_output_rm_input_no_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # multiple m output, rm input: no pressure
+; GREEDY_RA: # (%rdi), 4(%rdi), 8(%rdi), 12(%rdi), 16(%rdi), %eax, %ecx
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_mult_m_output_rm_input_no_pressure:
+; FAST_RA: #APP
+; FAST_RA: # multiple m output, rm input: no pressure
+; FAST_RA: # (%rdi), 4(%rdi), 8(%rdi), 12(%rdi), 16(%rdi), %eax, %ecx
+; FAST_RA: #NO_APP
+entry:
+ %b4 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %c5 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %d6 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %e7 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %0 = load i32, ptr %foo, align 4
+ %1 = load i32, ptr %b4, align 4
+ tail call void asm sideeffect "# multiple m output, rm input: no pressure\0A\09# $0, $1, $2, $3, $4, $5, $6", "=*m,=*m,=*m,=*m,=*m,rm,rm,~{dirflag},~{fpsr},~{flags}"(ptr nonnull elementtype(i32) %foo, ptr nonnull elementtype(i32) %b4, ptr nonnull elementtype(i32) %c5, ptr nonnull elementtype(i32) %d6, ptr nonnull elementtype(i32) %e7, i32 %0, i32 %1)
+ %2 = load i32, ptr %foo, align 4
+ ret i32 %2
+}
+
+define dso_local i32 @test_mult_m_output_rm_input_pressure(ptr noundef %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_mult_m_output_rm_input_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_GREEDY_RA: #NO_APP
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # multiple m output, rm input: pressure
+; FAST_ISEL_GREEDY_RA: # (%rbp), (%r8), (%r9), (%rbx), (%rsi), %eax, %edi
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_mult_m_output_rm_input_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_FAST_RA: #NO_APP
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # multiple m output, rm input: pressure
+; FAST_ISEL_FAST_RA: # (%rdi), (%rax), (%rcx), (%rdx), (%rsi), %r8d, %r9d
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_mult_m_output_rm_input_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; GREEDY_RA: #NO_APP
+; GREEDY_RA: #APP
+; GREEDY_RA: # multiple m output, rm input: pressure
+; GREEDY_RA: # (%rbp), 4(%rbp), 8(%rbp), 12(%rbp), 16(%rbp), %esi, %edi
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_mult_m_output_rm_input_pressure:
+; FAST_RA: #APP
+; FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_RA: #NO_APP
+; FAST_RA: #APP
+; FAST_RA: # multiple m output, rm input: pressure
+; FAST_RA: # (%rdi), 4(%rdi), 8(%rdi), 12(%rdi), 16(%rdi), %eax, %ecx
+; FAST_RA: #NO_APP
+entry:
+ %0 = tail call { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } asm sideeffect "# $0,$1,$2,$3,$4,$5,$6,$7,$8,$9,$10, $11, $12, $13, $14", "={rax},={rcx},={rdx},={rsi},={rdi},={rbx},={rbp},={r8},={r9},={r10},={r11},={r12},={r13},={r14},={r15},~{dirflag},~{fpsr},~{flags}"()
+ %b18 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %c19 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %d20 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %e21 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %1 = load i32, ptr %foo, align 4
+ %2 = load i32, ptr %b18, align 4
+ tail call void asm sideeffect "# multiple m output, rm input: pressure\0A\09# $0, $1, $2, $3, $4, $5, $6", "=*m,=*m,=*m,=*m,=*m,rm,rm,~{dirflag},~{fpsr},~{flags}"(ptr nonnull elementtype(i32) %foo, ptr nonnull elementtype(i32) %b18, ptr nonnull elementtype(i32) %c19, ptr nonnull elementtype(i32) %d20, ptr nonnull elementtype(i32) %e21, i32 %1, i32 %2)
+ %asmresult14 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 14
+ %asmresult13 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 13
+ %asmresult12 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 12
+ %asmresult11 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 11
+ %asmresult10 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 10
+ %asmresult9 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 9
+ %asmresult8 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 8
+ %asmresult7 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 7
+ %asmresult6 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 6
+ %asmresult5 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 5
+ %asmresult4 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 4
+ %asmresult3 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 3
+ %asmresult2 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 2
+ %asmresult1 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 1
+ %asmresult = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 0
+ tail call void @g(i64 noundef %asmresult, i64 noundef %asmresult1, i64 noundef %asmresult2, i64 noundef %asmresult3, i64 noundef %asmresult4, i64 noundef %asmresult5, i64 noundef %asmresult6, i64 noundef %asmresult7, i64 noundef %asmresult8, i64 noundef %asmresult9, i64 noundef %asmresult10, i64 noundef %asmresult11, i64 noundef %asmresult12, i64 noundef %asmresult13, i64 noundef %asmresult14)
+ %3 = load i32, ptr %foo, align 4
+ ret i32 %3
+}
+
+define dso_local i32 @test_mult_m_early_clobber_output_rm_input_no_pressure(ptr noundef %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_mult_m_early_clobber_output_rm_input_no_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # multiple m output, rm input: no pressure
+; FAST_ISEL_GREEDY_RA: # %eax, %esi, %r8d, %ecx, %edx
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_mult_m_early_clobber_output_rm_input_no_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # multiple m output, rm input: no pressure
+; FAST_ISEL_FAST_RA: # %eax, %edx, %ecx, %esi, %r8d
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_mult_m_early_clobber_output_rm_input_no_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # multiple m output, rm input: no pressure
+; GREEDY_RA: # %eax, %esi, %r8d, %ecx, %edx
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_mult_m_early_clobber_output_rm_input_no_pressure:
+; FAST_RA: #APP
+; FAST_RA: # multiple m output, rm input: no pressure
+; FAST_RA: # %eax, %edx, %ecx, %esi, %r8d
+; FAST_RA: #NO_APP
+entry:
+ %b2 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %c3 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %d4 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %0 = load i32, ptr %d4, align 4
+ %e5 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %1 = load i32, ptr %e5, align 4
+ %2 = tail call { i32, i32, i32 } asm sideeffect "# multiple m output, rm input: no pressure\0A\09# $0, $1, $2, $3, $4", "=&rm,=&rm,=&rm,rm,rm,~{dirflag},~{fpsr},~{flags}"(i32 %0, i32 %1)
+ %asmresult = extractvalue { i32, i32, i32 } %2, 0
+ %asmresult6 = extractvalue { i32, i32, i32 } %2, 1
+ %asmresult7 = extractvalue { i32, i32, i32 } %2, 2
+ store i32 %asmresult, ptr %foo, align 4
+ store i32 %asmresult6, ptr %b2, align 4
+ store i32 %asmresult7, ptr %c3, align 4
+ ret i32 %asmresult
+}
+
+define dso_local i32 @test_mult_m_early_clobber_output_rm_input_pressure(ptr noundef %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_mult_m_early_clobber_output_rm_input_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_GREEDY_RA: #NO_APP
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # multiple m output, rm input: pressure
+; FAST_ISEL_GREEDY_RA: # %r8d, %r9d, %eax, %esi, %edi
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_mult_m_early_clobber_output_rm_input_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_FAST_RA: #NO_APP
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # multiple m output, rm input: pressure
+; FAST_ISEL_FAST_RA: # %eax, %edi, %ecx, %edx, %esi
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_mult_m_early_clobber_output_rm_input_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; GREEDY_RA: #NO_APP
+; GREEDY_RA: #APP
+; GREEDY_RA: # multiple m output, rm input: pressure
+; GREEDY_RA: # %r8d, %r9d, %eax, %esi, %edi
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_mult_m_early_clobber_output_rm_input_pressure:
+; FAST_RA: #APP
+; FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_RA: #NO_APP
+; FAST_RA: #APP
+; FAST_RA: # multiple m output, rm input: pressure
+; FAST_RA: # %eax, %edi, %ecx, %edx, %esi
+; FAST_RA: #NO_APP
+entry:
+ %0 = tail call { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } asm sideeffect "# $0,$1,$2,$3,$4,$5,$6,$7,$8,$9,$10, $11, $12, $13, $14", "={rax},={rcx},={rdx},={rsi},={rdi},={rbx},={rbp},={r8},={r9},={r10},={r11},={r12},={r13},={r14},={r15},~{dirflag},~{fpsr},~{flags}"()
+ %b16 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %c17 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %d18 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %1 = load i32, ptr %d18, align 4
+ %e19 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %2 = load i32, ptr %e19, align 4
+ %3 = tail call { i32, i32, i32 } asm sideeffect "# multiple m output, rm input: pressure\0A\09# $0, $1, $2, $3, $4", "=&rm,=&rm,=&rm,rm,rm,~{dirflag},~{fpsr},~{flags}"(i32 %1, i32 %2)
+ %asmresult20 = extractvalue { i32, i32, i32 } %3, 0
+ %asmresult21 = extractvalue { i32, i32, i32 } %3, 1
+ %asmresult22 = extractvalue { i32, i32, i32 } %3, 2
+ store i32 %asmresult20, ptr %foo, align 4
+ store i32 %asmresult21, ptr %b16, align 4
+ store i32 %asmresult22, ptr %c17, align 4
+ %asmresult14 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 14
+ %asmresult13 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 13
+ %asmresult12 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 12
+ %asmresult11 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 11
+ %asmresult10 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 10
+ %asmresult9 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 9
+ %asmresult8 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 8
+ %asmresult7 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 7
+ %asmresult6 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 6
+ %asmresult5 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 5
+ %asmresult4 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 4
+ %asmresult3 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 3
+ %asmresult2 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 2
+ %asmresult1 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 1
+ %asmresult = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 0
+ tail call void @g(i64 noundef %asmresult, i64 noundef %asmresult1, i64 noundef %asmresult2, i64 noundef %asmresult3, i64 noundef %asmresult4, i64 noundef %asmresult5, i64 noundef %asmresult6, i64 noundef %asmresult7, i64 noundef %asmresult8, i64 noundef %asmresult9, i64 noundef %asmresult10, i64 noundef %asmresult11, i64 noundef %asmresult12, i64 noundef %asmresult13, i64 noundef %asmresult14)
+ %4 = load i32, ptr %foo, align 4
+ ret i32 %4
+}
+
+declare void @g(i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef)
diff --git a/llvm/test/CodeGen/X86/inline-asm-callbase.ll b/llvm/test/CodeGen/X86/inline-asm-callbase.ll
new file mode 100644
index 00000000000000..d9dc0211ec837e
--- /dev/null
+++ b/llvm/test/CodeGen/X86/inline-asm-callbase.ll
@@ -0,0 +1,227 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 6
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=true --regalloc=greedy < %s \
+; RUN: | FileCheck --check-prefix=FAST_ISEL_GREEDY_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=true --regalloc=fast < %s \
+; RUN: | FileCheck --check-prefix=FAST_ISEL_FAST_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=false --regalloc=greedy < %s \
+; RUN: | FileCheck --check-prefix=GREEDY_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=false --regalloc=fast < %s \
+; RUN: | FileCheck --check-prefix=FAST_RA %s
+
+; "rm" constraints on CallBase subclasses other than a plain call: invoke
+; (EH unwind edges) and callbr (indirect-branch asm goto).
+
+declare i32 @__gxx_personality_v0(...)
+
+define i32 @test_invoke_rm(i32 %x) personality ptr @__gxx_personality_v0 {
+; FAST_ISEL_GREEDY_RA-LABEL: test_invoke_rm:
+; FAST_ISEL_GREEDY_RA: # %bb.0: # %entry
+; FAST_ISEL_GREEDY_RA-NEXT: .Ltmp0: # EH_LABEL
+; FAST_ISEL_GREEDY_RA-NEXT: #APP
+; FAST_ISEL_GREEDY_RA-NEXT: # %eax, %edi
+; FAST_ISEL_GREEDY_RA-NEXT: #NO_APP
+; FAST_ISEL_GREEDY_RA-NEXT: .Ltmp1: # EH_LABEL
+; FAST_ISEL_GREEDY_RA-NEXT: # %bb.1: # %normal
+; FAST_ISEL_GREEDY_RA-NEXT: retq
+; FAST_ISEL_GREEDY_RA-NEXT: .LBB0_2: # %unwind
+; FAST_ISEL_GREEDY_RA-NEXT: pushq %rax
+; FAST_ISEL_GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_ISEL_GREEDY_RA-NEXT: .Ltmp2: # EH_LABEL
+; FAST_ISEL_GREEDY_RA-NEXT: movq %rax, %rdi
+; FAST_ISEL_GREEDY_RA-NEXT: callq _Unwind_Resume at PLT
+;
+; FAST_ISEL_FAST_RA-LABEL: test_invoke_rm:
+; FAST_ISEL_FAST_RA: # %bb.0: # %entry
+; FAST_ISEL_FAST_RA-NEXT: pushq %rax
+; FAST_ISEL_FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_ISEL_FAST_RA-NEXT: .Ltmp0: # EH_LABEL
+; FAST_ISEL_FAST_RA-NEXT: #APP
+; FAST_ISEL_FAST_RA-NEXT: # %eax, %edi
+; FAST_ISEL_FAST_RA-NEXT: #NO_APP
+; FAST_ISEL_FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_ISEL_FAST_RA-NEXT: .Ltmp1: # EH_LABEL
+; FAST_ISEL_FAST_RA-NEXT: # %bb.1: # %normal
+; FAST_ISEL_FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_ISEL_FAST_RA-NEXT: popq %rcx
+; FAST_ISEL_FAST_RA-NEXT: .cfi_def_cfa_offset 8
+; FAST_ISEL_FAST_RA-NEXT: retq
+; FAST_ISEL_FAST_RA-NEXT: .LBB0_2: # %unwind
+; FAST_ISEL_FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_ISEL_FAST_RA-NEXT: .Ltmp2: # EH_LABEL
+; FAST_ISEL_FAST_RA-NEXT: movq %rax, %rdi
+; FAST_ISEL_FAST_RA-NEXT: callq _Unwind_Resume at PLT
+;
+; GREEDY_RA-LABEL: test_invoke_rm:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: .Ltmp0: # EH_LABEL
+; GREEDY_RA-NEXT: #APP
+; GREEDY_RA-NEXT: # %eax, %edi
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: .Ltmp1: # EH_LABEL
+; GREEDY_RA-NEXT: # %bb.1: # %normal
+; GREEDY_RA-NEXT: retq
+; GREEDY_RA-NEXT: .LBB0_2: # %unwind
+; GREEDY_RA-NEXT: pushq %rax
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: .Ltmp2: # EH_LABEL
+; GREEDY_RA-NEXT: movq %rax, %rdi
+; GREEDY_RA-NEXT: callq _Unwind_Resume at PLT
+;
+; FAST_RA-LABEL: test_invoke_rm:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: pushq %rax
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: .Ltmp0: # EH_LABEL
+; FAST_RA-NEXT: #APP
+; FAST_RA-NEXT: # %eax, %edi
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_RA-NEXT: .Ltmp1: # EH_LABEL
+; FAST_RA-NEXT: # %bb.1: # %normal
+; FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_RA-NEXT: popq %rcx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 8
+; FAST_RA-NEXT: retq
+; FAST_RA-NEXT: .LBB0_2: # %unwind
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: .Ltmp2: # EH_LABEL
+; FAST_RA-NEXT: movq %rax, %rdi
+; FAST_RA-NEXT: callq _Unwind_Resume at PLT
+entry:
+ %0 = invoke i32 asm "# $0, $1", "=r,rm"(i32 %x)
+ to label %normal unwind label %unwind
+
+normal:
+ ret i32 %0
+
+unwind:
+ %1 = landingpad { ptr, i32 }
+ cleanup
+ resume { ptr, i32 } %1
+}
+
+define i32 @test_callbr_rm(i32 %x) {
+; FAST_ISEL_GREEDY_RA-LABEL: test_callbr_rm:
+; FAST_ISEL_GREEDY_RA: # %bb.0: # %entry
+; FAST_ISEL_GREEDY_RA-NEXT: #APP
+; FAST_ISEL_GREEDY_RA-NEXT: # %eax, %edi
+; FAST_ISEL_GREEDY_RA-NEXT: #NO_APP
+; FAST_ISEL_GREEDY_RA-NEXT: .LBB1_1: # Inline asm indirect target
+; FAST_ISEL_GREEDY_RA-NEXT: # %indirect
+; FAST_ISEL_GREEDY_RA-NEXT: # Label of block must be emitted
+; FAST_ISEL_GREEDY_RA-NEXT: retq
+;
+; FAST_ISEL_FAST_RA-LABEL: test_callbr_rm:
+; FAST_ISEL_FAST_RA: # %bb.0: # %entry
+; FAST_ISEL_FAST_RA-NEXT: #APP
+; FAST_ISEL_FAST_RA-NEXT: # %eax, %edi
+; FAST_ISEL_FAST_RA-NEXT: #NO_APP
+; FAST_ISEL_FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_ISEL_FAST_RA-NEXT: # %bb.1: # %normal
+; FAST_ISEL_FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_ISEL_FAST_RA-NEXT: retq
+; FAST_ISEL_FAST_RA-NEXT: .LBB1_2: # Inline asm indirect target
+; FAST_ISEL_FAST_RA-NEXT: # %indirect
+; FAST_ISEL_FAST_RA-NEXT: # Label of block must be emitted
+; FAST_ISEL_FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_ISEL_FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_ISEL_FAST_RA-NEXT: retq
+;
+; GREEDY_RA-LABEL: test_callbr_rm:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: #APP
+; GREEDY_RA-NEXT: # %eax, %edi
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: .LBB1_1: # Inline asm indirect target
+; GREEDY_RA-NEXT: # %indirect
+; GREEDY_RA-NEXT: # Label of block must be emitted
+; GREEDY_RA-NEXT: retq
+;
+; FAST_RA-LABEL: test_callbr_rm:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: #APP
+; FAST_RA-NEXT: # %eax, %edi
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_RA-NEXT: # %bb.1: # %normal
+; FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_RA-NEXT: retq
+; FAST_RA-NEXT: .LBB1_2: # Inline asm indirect target
+; FAST_RA-NEXT: # %indirect
+; FAST_RA-NEXT: # Label of block must be emitted
+; FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_RA-NEXT: retq
+entry:
+ %0 = callbr i32 asm "# $0, $1", "=r,rm,!i"(i32 %x)
+ to label %normal [label %indirect]
+
+normal:
+ ret i32 %0
+
+indirect:
+ ret i32 %0
+}
+
+define i32 @test_callbr_convert(i32 %x) {
+; FAST_ISEL_GREEDY_RA-LABEL: test_callbr_convert:
+; FAST_ISEL_GREEDY_RA: # %bb.0: # %entry
+; FAST_ISEL_GREEDY_RA-NEXT: #APP
+; FAST_ISEL_GREEDY_RA-NEXT: # %eax, %edi
+; FAST_ISEL_GREEDY_RA-NEXT: #NO_APP
+; FAST_ISEL_GREEDY_RA-NEXT: .LBB2_1: # Inline asm indirect target
+; FAST_ISEL_GREEDY_RA-NEXT: # %indirect
+; FAST_ISEL_GREEDY_RA-NEXT: # Label of block must be emitted
+; FAST_ISEL_GREEDY_RA-NEXT: retq
+;
+; FAST_ISEL_FAST_RA-LABEL: test_callbr_convert:
+; FAST_ISEL_FAST_RA: # %bb.0: # %entry
+; FAST_ISEL_FAST_RA-NEXT: #APP
+; FAST_ISEL_FAST_RA-NEXT: # %eax, %edi
+; FAST_ISEL_FAST_RA-NEXT: #NO_APP
+; FAST_ISEL_FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_ISEL_FAST_RA-NEXT: # %bb.1: # %normal
+; FAST_ISEL_FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_ISEL_FAST_RA-NEXT: retq
+; FAST_ISEL_FAST_RA-NEXT: .LBB2_2: # Inline asm indirect target
+; FAST_ISEL_FAST_RA-NEXT: # %indirect
+; FAST_ISEL_FAST_RA-NEXT: # Label of block must be emitted
+; FAST_ISEL_FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_ISEL_FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_ISEL_FAST_RA-NEXT: retq
+;
+; GREEDY_RA-LABEL: test_callbr_convert:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: #APP
+; GREEDY_RA-NEXT: # %eax, %edi
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: .LBB2_1: # Inline asm indirect target
+; GREEDY_RA-NEXT: # %indirect
+; GREEDY_RA-NEXT: # Label of block must be emitted
+; GREEDY_RA-NEXT: retq
+;
+; FAST_RA-LABEL: test_callbr_convert:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: #APP
+; FAST_RA-NEXT: # %eax, %edi
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_RA-NEXT: # %bb.1: # %normal
+; FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_RA-NEXT: retq
+; FAST_RA-NEXT: .LBB2_2: # Inline asm indirect target
+; FAST_RA-NEXT: # %indirect
+; FAST_RA-NEXT: # Label of block must be emitted
+; FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_RA-NEXT: retq
+entry:
+ %0 = callbr i32 asm "# $0, $1", "=rm,rm,!i"(i32 %x)
+ to label %normal [label %indirect]
+
+normal:
+ ret i32 %0
+
+indirect:
+ ret i32 %0
+}
diff --git a/llvm/test/CodeGen/X86/inline-asm-direct-mem-output-error.ll b/llvm/test/CodeGen/X86/inline-asm-direct-mem-output-error.ll
index 3f3a9168de38f8..248f6538d4f042 100644
--- a/llvm/test/CodeGen/X86/inline-asm-direct-mem-output-error.ll
+++ b/llvm/test/CodeGen/X86/inline-asm-direct-mem-output-error.ll
@@ -6,15 +6,14 @@
; A direct output is the asm's result, so there is no memory to write it to.
; Picking a memory constraint for one is an error, not a crash.
-; "rm" picks memory, the most general constraint.
+; -O0 picks memory for "rm"; above -O0 the register is preferred.
; O0: error: cannot handle direct memory outputs yet for constraint 'm'
-; O2: error: cannot handle direct memory outputs yet for constraint 'm'
define i32 @rm_i32() {
%r = call i32 asm "# $0", "=rm"()
ret i32 %r
}
-; The same for a value no 'r' register can hold.
+; No 'r' register can hold an x87 value, so "rm" picks memory either way.
; O0: error: cannot handle direct memory outputs yet for constraint 'm'
; O2: error: cannot handle direct memory outputs yet for constraint 'm'
define x86_fp80 @rm_x86_fp80() {
diff --git a/llvm/test/CodeGen/X86/inlineasm-sched-bug.ll b/llvm/test/CodeGen/X86/inlineasm-sched-bug.ll
index be4d1c29332f77..a322bd3003a58b 100644
--- a/llvm/test/CodeGen/X86/inlineasm-sched-bug.ll
+++ b/llvm/test/CodeGen/X86/inlineasm-sched-bug.ll
@@ -6,16 +6,13 @@
define i32 @foo(i32 %treemap) nounwind {
; CHECK-LABEL: foo:
; CHECK: # %bb.0: # %entry
-; CHECK-NEXT: pushl %eax
; CHECK-NEXT: movl {{[0-9]+}}(%esp), %eax
; CHECK-NEXT: movl %eax, %ecx
; CHECK-NEXT: negl %ecx
; CHECK-NEXT: andl %eax, %ecx
-; CHECK-NEXT: movl %ecx, (%esp)
; CHECK-NEXT: #APP
-; CHECK-NEXT: bsfl (%esp), %eax
+; CHECK-NEXT: bsfl %ecx, %eax
; CHECK-NEXT: #NO_APP
-; CHECK-NEXT: popl %ecx
; CHECK-NEXT: retl
entry:
%sub = sub i32 0, %treemap
More information about the llvm-branch-commits
mailing list