[clang] [llvm] [WIP][TargetLowering] Prefer 'r' over 'm' for "rm" inline asm constraints (PR #214061)
Bill Wendling via cfe-commits
cfe-commits at lists.llvm.org
Wed Sep 16 19:22:47 PDT 2026
https://github.com/isanbard updated https://github.com/llvm/llvm-project/pull/214061
>From f35ed27683b4ff951467c4c08dd9274c893f896d Mon Sep 17 00:00:00 2001
From: Bill Wendling <isanbard at gmail.com>
Date: Mon, 3 Aug 2026 23:36:30 -0700
Subject: [PATCH 01/12] [TargetLowering] Prefer 'r' over 'm' for "rm" inline
asm constraints
GCC/Clang-generated "rm" (register-or-memory) inline asm constraints
have historically always resolved to 'm', because TargetLowering's
constraint-preference ordering picks the most general constraint
present, and 'm' is more general than 'r'. This is safe but leaves
performance on the table: a value that could live in a register for
the asm's duration gets forced through a stack slot instead, even when
there's no real register pressure.
Add TargetLowering::AsmOperandInfo::MayFoldRegister, set in
ParseConstraints() whenever an operand's constraint codes are exactly
{"r", "m"} (covering both a direct "=rm"/"rm" operand and the tied
output half of a read-write "+rm" operand). getConstraintPreferences()
then opts for 'r' first when MayFoldRegister is set, falling back to
'm' -- but only above -O0: at -O0 the fast register allocator has no
on-demand spill/fold machinery of its own yet (that's what the
following RegAllocFast commit adds), so preferring 'r' there would
just move the "ran out of registers" failure earlier without a way to
recover from it.
Plumb the resulting MayFoldRegister choice through to the actual
INLINEASM MachineInstr by threading it into
RegsForValue::AddInlineAsmOperands(), which sets
InlineAsm::Flag::RegMayBeFolded on the emitted register operand.
Downstream consumers already exist for this bit: the greedy allocator
already honors it via InlineSpiller for on-demand memory folding under
pressure. RegAllocFast does not yet -- see the following commit.
Update inlineasm-sched-bug.ll's expected output: it uses a "=r,rm"
constraint pair, so its "rm" input now resolves to a register instead
of round-tripping through a stack slot, which is the intended effect
of this change and produces strictly better code for that test.
Co-Authored-By: Claude Sonnet 5 <noreply at anthropic.com>
Claude-Session: https://claude.ai/code/session_0143ozRFwhFjWmJjtfZUgmYe
---
llvm/include/llvm/CodeGen/TargetLowering.h | 8 +++++
.../SelectionDAG/SelectionDAGBuilder.cpp | 22 +++++++-----
.../SelectionDAG/SelectionDAGBuilder.h | 5 +--
.../CodeGen/SelectionDAG/TargetLowering.cpp | 33 ++++++++++++++++-
llvm/test/CodeGen/X86/asm-constraints-rm.ll | 36 +++++++++++++++++++
llvm/test/CodeGen/X86/inlineasm-sched-bug.ll | 5 +--
6 files changed, 93 insertions(+), 16 deletions(-)
create mode 100644 llvm/test/CodeGen/X86/asm-constraints-rm.ll
diff --git a/llvm/include/llvm/CodeGen/TargetLowering.h b/llvm/include/llvm/CodeGen/TargetLowering.h
index c663bb8ea65b7..55c8f0599f411 100644
--- a/llvm/include/llvm/CodeGen/TargetLowering.h
+++ b/llvm/include/llvm/CodeGen/TargetLowering.h
@@ -5349,6 +5349,14 @@ class LLVM_ABI TargetLowering : public TargetLoweringBase {
/// The ValueType for the operand value.
MVT ConstraintVT = MVT::Other;
+ /// True if this operand's constraint codes are exactly {"r", "m"} (the
+ /// "rm" constraint, or the tied output half of "+rm"). getConstraintType
+ /// preference selection uses this to opt for 'r' while still allowing
+ /// the register allocator to fall back to 'm' under register pressure,
+ /// instead of picking 'm' unconditionally as it would for a generic
+ /// multi-alternative constraint.
+ bool MayFoldRegister = false;
+
/// Copy constructor for copying from a ConstraintInfo.
AsmOperandInfo(InlineAsm::ConstraintInfo Info)
: InlineAsm::ConstraintInfo(std::move(Info)) {}
diff --git a/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp b/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp
index 385f3729d46c8..e51e02e81b628 100644
--- a/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp
@@ -1033,7 +1033,8 @@ void RegsForValue::getCopyToRegs(SDValue Val, SelectionDAG &DAG,
}
void RegsForValue::AddInlineAsmOperands(InlineAsm::Kind Code, bool HasMatching,
- unsigned MatchingIdx, const SDLoc &dl,
+ unsigned MatchingIdx,
+ bool MayFoldRegister, const SDLoc &dl,
SelectionDAG &DAG,
std::vector<SDValue> &Ops) const {
const TargetLowering &TLI = DAG.getTargetLoweringInfo();
@@ -1050,6 +1051,7 @@ void RegsForValue::AddInlineAsmOperands(InlineAsm::Kind Code, bool HasMatching,
const MachineRegisterInfo &MRI = DAG.getMachineFunction().getRegInfo();
const TargetRegisterClass *RC = MRI.getRegClass(Regs.front());
Flag.setRegClass(RC->getID());
+ Flag.setRegMayBeFolded(MayFoldRegister);
}
SDValue Res = DAG.getTargetConstant(Flag, dl, MVT::i32);
@@ -10551,7 +10553,7 @@ static bool prepareDAGLevelOperands(ConstraintDecisionInfo &Info,
OpInfo.AssignedRegs.AddInlineAsmOperands(
OpInfo.isEarlyClobber ? InlineAsm::Kind::RegDefEarlyClobber
: InlineAsm::Kind::RegDef,
- false, 0, DL, DAG, Info.AsmNodeOperands);
+ false, 0, OpInfo.MayFoldRegister, DL, DAG, Info.AsmNodeOperands);
}
break;
@@ -10592,9 +10594,9 @@ static bool prepareDAGLevelOperands(ConstraintDecisionInfo &Info,
// Use the produced MatchedRegs object to
MatchedRegs.getCopyToRegs(InOperandVal, DAG, DL, Info.Chain,
&Info.Glue, &Call);
- MatchedRegs.AddInlineAsmOperands(InlineAsm::Kind::RegUse, true,
- OpInfo.getMatchedOperand(), DL, DAG,
- Info.AsmNodeOperands);
+ MatchedRegs.AddInlineAsmOperands(
+ InlineAsm::Kind::RegUse, true, OpInfo.getMatchedOperand(),
+ OpInfo.MayFoldRegister, DL, DAG, Info.AsmNodeOperands);
break;
}
@@ -10717,8 +10719,9 @@ static bool prepareDAGLevelOperands(ConstraintDecisionInfo &Info,
OpInfo.AssignedRegs.getCopyToRegs(InOperandVal, DAG, DL, Info.Chain,
&Info.Glue, &Call);
- OpInfo.AssignedRegs.AddInlineAsmOperands(
- InlineAsm::Kind::RegUse, false, 0, DL, DAG, Info.AsmNodeOperands);
+ OpInfo.AssignedRegs.AddInlineAsmOperands(InlineAsm::Kind::RegUse, false,
+ 0, OpInfo.MayFoldRegister, DL,
+ DAG, Info.AsmNodeOperands);
break;
}
@@ -10726,8 +10729,9 @@ static bool prepareDAGLevelOperands(ConstraintDecisionInfo &Info,
// Add the clobbered value to the operand list, so that the register
// allocator is aware that the physreg got clobbered.
if (!OpInfo.AssignedRegs.Regs.empty())
- OpInfo.AssignedRegs.AddInlineAsmOperands(
- InlineAsm::Kind::Clobber, false, 0, DL, DAG, Info.AsmNodeOperands);
+ OpInfo.AssignedRegs.AddInlineAsmOperands(InlineAsm::Kind::Clobber,
+ false, 0, false, DL, DAG,
+ Info.AsmNodeOperands);
break;
}
}
diff --git a/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.h b/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.h
index 6c7711af078f0..9ce96fe7c1c05 100644
--- a/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.h
+++ b/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.h
@@ -829,8 +829,9 @@ struct RegsForValue {
/// code marker, matching input operand index (if applicable), and includes
/// the number of values added into it.
void AddInlineAsmOperands(InlineAsm::Kind Code, bool HasMatching,
- unsigned MatchingIdx, const SDLoc &dl,
- SelectionDAG &DAG, std::vector<SDValue> &Ops) const;
+ unsigned MatchingIdx, bool MayFoldRegister,
+ const SDLoc &dl, SelectionDAG &DAG,
+ std::vector<SDValue> &Ops) const;
/// Check if the total RegCount is greater than one.
bool occupiesMultipleRegs() const {
diff --git a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
index 6f47bf1b81c57..df56b983ce46f 100644
--- a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
@@ -6120,6 +6120,20 @@ TargetLowering::ParseConstraints(const DataLayout &DL,
OpInfo.ConstraintVT = MVT::Other;
+ // Special treatment for all platforms that can fold a register into a
+ // spill. This is used for the "rm" constraint, where we would vastly
+ // prefer to use 'r' over 'm'. The non-fast register allocators are able to
+ // handle the 'r' default by folding. The fast register allocator needs
+ // special handling to convert the instruction to use 'm' instead.
+ //
+ // This also applies to read-write "+rm" constraints (which generate a
+ // direct "=rm" output with a matching tied input). The register allocator
+ // can fold both the output and its tied input to the same memory slot when
+ // under pressure.
+ if (OpInfo.Codes.size() == 2 && llvm::is_contained(OpInfo.Codes, "r") &&
+ llvm::is_contained(OpInfo.Codes, "m"))
+ OpInfo.MayFoldRegister = true;
+
// Compute the value type for each operand.
switch (OpInfo.Type) {
case InlineAsm::isOutput: {
@@ -6400,7 +6414,14 @@ TargetLowering::ConstraintWeight
/// 1) If there is an 'other' constraint, and if the operand is valid for
/// that constraint, use it. This makes us take advantage of 'i'
/// constraints when available.
-/// 2) Otherwise, pick the most general constraint present. This prefers
+/// 2) Special processing is done for the "rm" constraint. If specified, we
+/// opt for the 'r' constraint, but mark the operand as being "foldable."
+/// In the face of register exhaustion, the register allocator is free to
+/// choose to use a stack slot: InlineSpiller does this on demand for the
+/// greedy allocator, and RegAllocFast::foldFoldableInlineAsmOperands()
+/// does the equivalent for the fast allocator, which has no on-demand
+/// spilling machinery of its own to fall back on otherwise.
+/// 3) Otherwise, pick the most general constraint present. This prefers
/// 'm' over 'r', for example.
///
TargetLowering::ConstraintGroup TargetLowering::getConstraintPreferences(
@@ -6408,6 +6429,16 @@ TargetLowering::ConstraintGroup TargetLowering::getConstraintPreferences(
ConstraintGroup Ret;
Ret.reserve(OpInfo.Codes.size());
+
+ // If we can fold the register (i.e. it has an "rm" constraint), opt for the
+ // 'r' constraint, and allow the register allocator to spill if need be.
+ const TargetMachine &TM = getTargetMachine();
+ if (TM.getOptLevel() != CodeGenOptLevel::None && OpInfo.MayFoldRegister) {
+ Ret.emplace_back(ConstraintPair("r", getConstraintType("r")));
+ Ret.emplace_back(ConstraintPair("m", getConstraintType("m")));
+ return Ret;
+ }
+
for (StringRef Code : OpInfo.Codes) {
TargetLowering::ConstraintType CType = getConstraintType(Code);
diff --git a/llvm/test/CodeGen/X86/asm-constraints-rm.ll b/llvm/test/CodeGen/X86/asm-constraints-rm.ll
new file mode 100644
index 0000000000000..1da68a51198c9
--- /dev/null
+++ b/llvm/test/CodeGen/X86/asm-constraints-rm.ll
@@ -0,0 +1,36 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 6
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --regalloc=greedy < %s | FileCheck --check-prefix=GREEDY_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --regalloc=fast < %s | FileCheck --check-prefix=FAST_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu -O0 < %s | FileCheck --check-prefix=O0 %s
+
+; TargetLowering::getConstraintPreferences() now prefers 'r' over 'm' for a
+; bare "rm" constraint (see TargetLowering::ParseConstraints setting
+; MayFoldRegister, and ComputeConstraintToUse's -O0 opt-out). With no real
+; register pressure here, both allocators should just keep the value in a
+; register end to end, unlike the historical "always pick m" default.
+define i64 @test_rm_output_no_pressure(i64 %a) {
+; GREEDY_RA-LABEL: test_rm_output_no_pressure:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: #APP
+; GREEDY_RA-NEXT: bsfq %rdi, %rax
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: retq
+;
+; FAST_RA-LABEL: test_rm_output_no_pressure:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: #APP
+; FAST_RA-NEXT: bsfq %rdi, %rax
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: retq
+;
+; O0-LABEL: test_rm_output_no_pressure:
+; O0: # %bb.0: # %entry
+; O0-NEXT: movq %rdi, -{{[0-9]+}}(%rsp)
+; O0-NEXT: #APP
+; O0-NEXT: bsfq -{{[0-9]+}}(%rsp), %rax
+; O0-NEXT: #NO_APP
+; O0-NEXT: retq
+entry:
+ %0 = call i64 asm "bsfq $1,$0", "=r,rm"(i64 %a)
+ ret i64 %0
+}
diff --git a/llvm/test/CodeGen/X86/inlineasm-sched-bug.ll b/llvm/test/CodeGen/X86/inlineasm-sched-bug.ll
index be4d1c29332f7..a322bd3003a58 100644
--- a/llvm/test/CodeGen/X86/inlineasm-sched-bug.ll
+++ b/llvm/test/CodeGen/X86/inlineasm-sched-bug.ll
@@ -6,16 +6,13 @@
define i32 @foo(i32 %treemap) nounwind {
; CHECK-LABEL: foo:
; CHECK: # %bb.0: # %entry
-; CHECK-NEXT: pushl %eax
; CHECK-NEXT: movl {{[0-9]+}}(%esp), %eax
; CHECK-NEXT: movl %eax, %ecx
; CHECK-NEXT: negl %ecx
; CHECK-NEXT: andl %eax, %ecx
-; CHECK-NEXT: movl %ecx, (%esp)
; CHECK-NEXT: #APP
-; CHECK-NEXT: bsfl (%esp), %eax
+; CHECK-NEXT: bsfl %ecx, %eax
; CHECK-NEXT: #NO_APP
-; CHECK-NEXT: popl %ecx
; CHECK-NEXT: retl
entry:
%sub = sub i32 0, %treemap
>From 31d76f0f7df794fc778330891111640bd59ddf6f Mon Sep 17 00:00:00 2001
From: Bill Wendling <isanbard at gmail.com>
Date: Mon, 3 Aug 2026 23:37:02 -0700
Subject: [PATCH 02/12] [RegAllocFast] Fold foldable "rm" inline asm operands
under register pressure
Now that TargetLowering prefers 'r' over 'm' for "rm" constraints (see
previous commit) and sets InlineAsm::Flag::RegMayBeFolded accordingly,
the fast register allocator needs a way to undo that choice when it
actually runs out of registers -- previously it had none: RegAllocFast
either found a physical register or hard-errored ("inline assembly
requires more registers than available"), with nothing in between. The
greedy allocator already handles this via InlineSpiller's on-demand
folding; RegAllocFast has no equivalent on-demand spill machinery.
This follows the shape of Nick Desaulniers' unmerged upstream PR
"[RegAllocFast] fold foldable inline asm" (llvm/llvm-project#74344,
open since 2023) with two changes responding to that PR's review
thread (MatzeB, qcolombet):
- Per-instruction pressure estimate instead of unconditional folding.
MatzeB's thread ultimately converged on accepting unconditional
folding as good enough, but estimating each inline asm instruction's
own register demand (defs + uses + clobbers) against its register
class and only folding as many "rm" operands as needed to make it
fit is barely more code and avoids regressing the common, no-pressure
case from "uses a register, like greedy" to "always spills."
- A cached MachineFunction::hasInlineAsm() check gates the whole fold
pass per function. MatzeB's main unresolved concern was that scanning
every basic block for foldable operands before the main allocation
loop isn't free, and the vast majority of functions never contain
inline asm at all. This turns that per-block scan into a single
function-wide check for the common case, addressing the concern
without the deeper single-pass restructuring MatzeB also floated
(folding inline asm operands inline within the main reverse-iteration
loop via early_inc_range instead of as a separate pre-pass): that
runs into the exact problem it would need to avoid, since the newly
inserted store/fold/reload instructions still need to be visited by
the allocator like any other instruction, and early-inc iteration
order does not naturally do that.
Adds asm-constraints-rm-pressure.ll (register-exhaustion case for both
greedy and fast) and asm-constraints-rm-callbr-fold.ll, a regression
test for a case the original PR's author flagged with a bare "TODO:
asm goto" but never resolved: a folded "=&rm" callbr def used only
along an indirect (not fallthrough) successor edge. Verified this
already works correctly: the fold's memory write is an unconditional
side effect of executing the asm, so it's visible in the stack slot on
every successor path, and the indirect block's own reload -- inserted
independently by RegAllocFast's normal live-in handling, driven by the
shared StackSlotForVirtReg entry for the folded register -- picks it
up the same way any other cross-block virtual register use would.
Co-Authored-By: Claude Sonnet 5 <noreply at anthropic.com>
Claude-Session: https://claude.ai/code/session_0143ozRFwhFjWmJjtfZUgmYe
---
llvm/lib/CodeGen/RegAllocFast.cpp | 256 ++++++++++++++++++
.../X86/asm-constraints-rm-callbr-fold.ll | 83 ++++++
.../X86/asm-constraints-rm-pressure.ll | 109 ++++++++
3 files changed, 448 insertions(+)
create mode 100644 llvm/test/CodeGen/X86/asm-constraints-rm-callbr-fold.ll
create mode 100644 llvm/test/CodeGen/X86/asm-constraints-rm-pressure.ll
diff --git a/llvm/lib/CodeGen/RegAllocFast.cpp b/llvm/lib/CodeGen/RegAllocFast.cpp
index f10283272f1d9..1109f25f3b214 100644
--- a/llvm/lib/CodeGen/RegAllocFast.cpp
+++ b/llvm/lib/CodeGen/RegAllocFast.cpp
@@ -53,6 +53,10 @@ using namespace llvm;
STATISTIC(NumStores, "Number of stores added");
STATISTIC(NumLoads, "Number of loads added");
STATISTIC(NumCoalesced, "Number of copies coalesced");
+STATISTIC(NumInlineAsmFoldedStores,
+ "Number of stores added by inline asm operand folding");
+STATISTIC(NumInlineAsmFoldedLoads,
+ "Number of loads added by inline asm operand folding");
// FIXME: Remove this switch when all testcases are fixed!
static cl::opt<bool> IgnoreMissingDefs("rafast-ignore-missing-defs",
@@ -398,6 +402,19 @@ class RegAllocFastImpl {
bool mayBeSpillFromInlineAsmBr(const MachineInstr &MI) const;
+ /// Cached copy of MachineFunction::hasInlineAsm(), set once in
+ /// runOnMachineFunction(). Lets allocateBasicBlock() skip scanning each
+ /// block for foldable inline asm operands with a single function-wide
+ /// check, since the overwhelming majority of functions contain no inline
+ /// asm at all.
+ bool MFHasInlineAsm = false;
+
+ void selectInlineAsmOperandsToFold(MachineInstr &MI,
+ SmallSet<Register, 8> &ToFold);
+ void foldFoldableInlineAsmOperands(MachineBasicBlock &MBB);
+ void foldFoldableInlineAsmOperands(MachineInstr *&MI,
+ const SmallSet<Register, 8> &ToFold);
+
void dumpState() const;
};
@@ -1789,6 +1806,233 @@ void RegAllocFastImpl::handleBundle(MachineInstr &MI) {
}
}
+/// Decide which of \p MI's foldable ("rm"-style) register operands actually
+/// need to be converted to memory, and record their registers in \p ToFold.
+///
+/// A foldable operand is one where SelectionDAG chose 'r' (registers give
+/// better code than always spilling to memory) but recorded, via
+/// InlineAsm::Flag::RegMayBeFolded, that 'm' is an available fallback -- see
+/// TargetLowering::ComputeConstraintToUse. The greedy allocator can leave
+/// this decision until it actually runs out of registers: InlineSpiller
+/// folds on demand, informed by real, global register pressure. RegAllocFast
+/// has no such on-demand spilling machinery, and restructuring its
+/// single-pass, iterate-MI's-live-operand-list design to support it safely
+/// is a larger change (folding an operand replaces MI with a new
+/// instruction, which can't happen mid-iteration of the def/use loops in
+/// allocateInstruction()). Folding every foldable operand unconditionally
+/// would sidestep that, but is strictly pessimistic -- verified against
+/// this file's own asm-constraints-torture.ll, it regresses cases with no
+/// real pressure at all from "uses a register, like greedy" to "always
+/// spills," which is worse than doing nothing.
+///
+/// This is the middle ground: estimate, from MI's own operand list alone,
+/// whether its simultaneous register demands -- defs, uses, and clobbers,
+/// all alive only for this one instruction -- fit in the relevant register
+/// class, and only fold as many foldable operands as needed to make them
+/// fit. Non-foldable operands get first claim on the available registers,
+/// since they have no fallback.
+///
+/// Three simplifications, all biased toward folding too little rather than
+/// too much (i.e. toward RegAllocFast's pre-existing "hard error if a
+/// register genuinely isn't available" behavior, never toward silently
+/// wrong codegen):
+/// - It only accounts for pressure local to this one instruction, not
+/// registers already committed to values live across other instructions
+/// in the block. RegAllocFast has no on-demand spilling for those either
+/// way, so this is a pre-existing limitation, not a regression.
+/// - It buckets demand by exact TargetRegisterClass rather than unifying
+/// classes that alias the same physical registers (e.g. GR32 and GR64),
+/// so it can undercount pressure when an instruction mixes classes of
+/// different widths. RegAllocFast's normal out-of-registers error path
+/// remains a backstop for any such case this estimate gets wrong.
+/// - It doesn't model early-clobber's stricter requirement that a def be
+/// disjoint from every input, not just other defs -- it counts an
+/// early-clobber def as ordinary same-class demand, the same as it would
+/// a non-early-clobber one. This is really a specific, easy-to-hit
+/// instance of the previous point (the disjointness early-clobber
+/// requires isn't scoped to one register class either), called out
+/// separately because "=&rm" early-clobber outputs are a common shape
+/// for this constraint in practice. Same backstop applies.
+void RegAllocFastImpl::selectInlineAsmOperandsToFold(
+ MachineInstr &MI, SmallSet<Register, 8> &ToFold) {
+ struct Demand {
+ unsigned NonFoldable = 0;
+ SmallVector<Register, 4> Foldable;
+ };
+ // SmallMapVector, not SmallDenseMap: each register class's bucket is
+ // resolved independently below with no state carried between iterations,
+ // so plain DenseMap's pointer-hash-ordered iteration wouldn't actually be
+ // wrong here -- but a reviewer shouldn't have to re-derive that from
+ // scratch every time this function is read, and it costs nothing to make
+ // the iteration order a non-question by construction (insertion order,
+ // i.e. the order classes are first seen scanning MI's operands).
+ SmallMapVector<const TargetRegisterClass *, Demand, 8> DemandByClass;
+ SmallVector<MCPhysReg, 8> Blocked;
+
+ for (unsigned I = InlineAsm::MIOp_FirstOperand, E = MI.getNumOperands();
+ I != E; ++I) {
+ MachineOperand &MO = MI.getOperand(I);
+ if (!MO.isReg() || !MO.getReg())
+ continue;
+ Register Reg = MO.getReg();
+ if (Reg.isPhysical()) {
+ Blocked.push_back(Reg.asMCReg());
+ continue;
+ }
+ Demand &D = DemandByClass[MRI->getRegClass(Reg)];
+ if (MI.mayFoldInlineAsmRegOp(I))
+ D.Foldable.push_back(Reg);
+ else
+ ++D.NonFoldable;
+ }
+
+ for (auto &KV : DemandByClass) {
+ const TargetRegisterClass *RC = KV.first;
+ Demand &D = KV.second;
+ if (D.Foldable.empty())
+ continue;
+
+ unsigned Available = 0;
+ for (MCPhysReg PhysReg : RegClassInfo.getOrder(RC))
+ if (!llvm::any_of(Blocked, [&](MCPhysReg B) {
+ return TRI->regsOverlap(PhysReg, B);
+ }))
+ ++Available;
+
+ unsigned Spare = Available > D.NonFoldable ? Available - D.NonFoldable : 0;
+ unsigned NumToFold =
+ D.Foldable.size() > Spare ? D.Foldable.size() - Spare : 0;
+ for (unsigned I = 0; I < NumToFold; ++I)
+ ToFold.insert(D.Foldable[I]);
+ }
+}
+
+/// Convert one selected foldable register operand of an inline asm
+/// instruction to its memory ('m') form, in place. \p MI is replaced with
+/// the folded instruction on return, since folding creates a new
+/// instruction rather than mutating MI. Only operands whose register is in
+/// \p ToFold (computed by selectInlineAsmOperandsToFold()) are converted;
+/// the rest are left in register form.
+void RegAllocFastImpl::foldFoldableInlineAsmOperands(
+ MachineInstr *&MI, const SmallSet<Register, 8> &ToFold) {
+ assert(MI->isInlineAsm() && "should only be used on inline asm");
+
+ // Folding a register operand replaces it with a multi-operand frame-index
+ // reference (base/scale/index/disp/segment, on X86), which shifts the
+ // indices of every later operand. Re-read getNumOperands() each iteration
+ // rather than caching it, and never cache MI itself -- foldMemoryOperand()
+ // returns a new instruction rather than mutating in place. ToFold is keyed
+ // by register rather than operand index for exactly this reason: indices
+ // shift as earlier operands in this same loop are folded, but a register
+ // number stays meaningful throughout.
+ for (unsigned I = InlineAsm::MIOp_FirstOperand; I < MI->getNumOperands();
+ ++I) {
+ MachineOperand &MO = MI->getOperand(I);
+ if (!(MO.isReg() && MI->mayFoldInlineAsmRegOp(I) &&
+ ToFold.contains(MO.getReg())))
+ continue;
+
+ const bool IsDef = MO.isDef();
+
+ // A tied "+rm" operand (LLVM IR "=rm,0") is really one memory location
+ // shared between the def and its tied input: TargetInstrInfo already
+ // folds both halves together when given the def's operand index (see
+ // foldInlineAsmMemOperand()'s untie-then-recurse handling), matching how
+ // InlineSpiller.cpp only ever passes the def side to foldMemoryOperand.
+ //
+ // This branch is reached in practice -- TargetLowering::ParseConstraints()
+ // sets MayFoldRegister for a tied def's own "rm" codes the same as for
+ // any other exactly-{r,m} operand (see asm-constraints-torture.ll's
+ // test_tied_output_pressure and asm-reg-mem-constraints.c's
+ // test_tied_rm_output for the pressure and no-pressure cases
+ // respectively). TiedUse below is what makes that correct: without it,
+ // folding only the def's own operand would store back an undefined
+ // value instead of the tied input's.
+ const MachineOperand *TiedUse = nullptr;
+ if (MO.isTied()) {
+ MachineOperand &T = MI->getOperand(MI->findTiedOperandIdx(I));
+ if (T.isUse())
+ TiedUse = &T;
+ }
+
+ Register Reg = MO.getReg();
+ const bool IsVirt = Reg.isVirtual();
+ const TargetRegisterClass *RC =
+ IsVirt ? MRI->getRegClass(Reg) : TRI->getMinimalPhysRegClass(Reg);
+
+ // Reuse the same slot-assignment bookkeeping as ordinary spills for
+ // virtual registers, so a register that's folded here and also spilled
+ // elsewhere in the function shares one stack slot. Physical register
+ // operands (a fixed-register constraint that happens to also allow
+ // "rm") aren't tracked by StackSlotForVirtReg, so give them their own
+ // object.
+ int FrameIndex;
+ if (IsVirt) {
+ FrameIndex = getStackSpaceFor(Reg);
+ } else {
+ unsigned Size = TRI->getSpillSize(*RC);
+ Align Alignment = TRI->getSpillAlign(*RC);
+ FrameIndex = MFI->CreateSpillStackObject(Size, Alignment,
+ TRI->getSpillStackID(*RC));
+ }
+
+ MachineInstr *CopyMI = nullptr;
+ MachineInstr *NewMI = TII->foldMemoryOperand(*MI, {I}, FrameIndex, CopyMI);
+ assert(NewMI && "operand was reported foldable but folding failed");
+ if (!NewMI)
+ continue;
+
+ // foldMemoryOperand() inserts NewMI immediately before MI and leaves MI
+ // in the instruction list. Move NewMI to where MI was, so the rest of
+ // this loop (and the caller's iteration over the block) sees
+ // instructions in program order once MI is erased below.
+ MI->getParent()->splice(std::next(MI->getIterator()), NewMI->getParent(),
+ NewMI->getIterator());
+
+ if (IsDef) {
+ // The asm now writes the stack slot instead of Reg. Reload afterward
+ // so that Reg -- which may still have uses elsewhere -- is available
+ // in a register again immediately after the asm.
+ TII->loadRegFromStackSlot(*MBB, std::next(NewMI->getIterator()), Reg,
+ FrameIndex, RC, Reg);
+ ++NumLoads;
+ ++NumInlineAsmFoldedLoads;
+ }
+
+ if (!IsDef || TiedUse) {
+ // Store this operand's pre-asm value into the slot: either the plain
+ // input's own value, or -- for a tied pair -- the tied input's value,
+ // deliberately reusing the def side's FrameIndex so the asm reads and
+ // writes the one memory location the tie requires.
+ Register StoreReg = TiedUse ? TiedUse->getReg() : Reg;
+ bool IsKill = TiedUse ? TiedUse->isKill() : MO.isKill();
+ TII->storeRegToStackSlot(*MBB, MI->getIterator(), StoreReg, IsKill,
+ FrameIndex, RC, StoreReg);
+ ++NumStores;
+ ++NumInlineAsmFoldedStores;
+ }
+
+ MI->eraseFromParent();
+ MI = NewMI;
+ }
+}
+
+/// Fold whichever of \p MBB's inline asm operands
+/// selectInlineAsmOperandsToFold() determines are actually needed, before the
+/// main allocation loop runs.
+void RegAllocFastImpl::foldFoldableInlineAsmOperands(MachineBasicBlock &MBB) {
+ SmallVector<MachineInstr *, 4> InlineAsms;
+ for (MachineInstr &MI : MBB)
+ if (MI.isInlineAsm())
+ InlineAsms.push_back(&MI);
+ for (MachineInstr *MI : InlineAsms) {
+ SmallSet<Register, 8> ToFold;
+ selectInlineAsmOperandsToFold(*MI, ToFold);
+ if (!ToFold.empty())
+ foldFoldableInlineAsmOperands(MI, ToFold);
+ }
+}
+
void RegAllocFastImpl::allocateBasicBlock(MachineBasicBlock &MBB) {
this->MBB = &MBB;
LLVM_DEBUG(dbgs() << "\nAllocating " << MBB);
@@ -1802,6 +2046,17 @@ void RegAllocFastImpl::allocateBasicBlock(MachineBasicBlock &MBB) {
Coalesced.clear();
+ // Fold operands that ISel flagged as foldable (rm-style constraints) to
+ // their memory form before the main allocation loop runs, so those
+ // operands never compete for a register at all -- see
+ // foldFoldableInlineAsmOperands() for why this can't be done lazily like
+ // greedy's on-demand InlineSpiller folding. Gated on MFHasInlineAsm so
+ // functions with no inline asm at all -- the common case -- pay nothing
+ // beyond the one function-wide check already made in
+ // runOnMachineFunction(), instead of an extra per-block scan.
+ if (MFHasInlineAsm)
+ foldFoldableInlineAsmOperands(MBB);
+
// Traverse block in reverse order allocating instructions one by one.
for (MachineInstr &MI : reverse(MBB)) {
LLVM_DEBUG(dbgs() << "\n>> " << MI << "Regs:"; dumpState());
@@ -1858,6 +2113,7 @@ bool RegAllocFastImpl::runOnMachineFunction(MachineFunction &MF) {
TRI = STI.getRegisterInfo();
TII = STI.getInstrInfo();
MFI = &MF.getFrameInfo();
+ MFHasInlineAsm = MF.hasInlineAsm();
MRI->freezeReservedRegs();
RegClassInfo.runOnMachineFunction(MF);
unsigned NumRegUnits = TRI->getNumRegUnits();
diff --git a/llvm/test/CodeGen/X86/asm-constraints-rm-callbr-fold.ll b/llvm/test/CodeGen/X86/asm-constraints-rm-callbr-fold.ll
new file mode 100644
index 0000000000000..f19e432ca4de2
--- /dev/null
+++ b/llvm/test/CodeGen/X86/asm-constraints-rm-callbr-fold.ll
@@ -0,0 +1,83 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 6
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --regalloc=fast \
+; RUN: -verify-machineinstrs -verify-regalloc < %s | FileCheck %s
+
+; RegAllocFast::foldFoldableInlineAsmOperands() folds this "=&rm" output to
+; memory under the register pressure from the six "r" inputs plus clobbers.
+; Unlike an ordinary instruction, a folded callbr's def is reachable from two
+; different points in the block that follows it: the fallthrough successor
+; (where RegAllocFast inserts an explicit reload right after the asm) and any
+; indirect successor (which has no such reload -- it is a different block,
+; reached by branching out of the asm itself, bypassing the fallthrough
+; instruction stream entirely). This is only correct because the memory
+; write happens as a side effect of executing the folded asm itself, so it is
+; unconditionally visible on every successor path; the indirect block then
+; picks it up via RegAllocFast's ordinary live-in reload (driven by the
+; shared StackSlotForVirtReg entry for %r), the same as any other cross-block
+; virtual register use. Guards against a regression where this only works by
+; accident of shared bookkeeping and silently reads an uninitialized slot on
+; the indirect path.
+define i64 @test_callbr_rm_indirect_use(i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %f) {
+; CHECK-LABEL: test_callbr_rm_indirect_use:
+; CHECK: # %bb.0: # %entry
+; CHECK-NEXT: pushq %rbp
+; CHECK-NEXT: .cfi_def_cfa_offset 16
+; CHECK-NEXT: pushq %r15
+; CHECK-NEXT: .cfi_def_cfa_offset 24
+; CHECK-NEXT: pushq %r14
+; CHECK-NEXT: .cfi_def_cfa_offset 32
+; CHECK-NEXT: pushq %r13
+; CHECK-NEXT: .cfi_def_cfa_offset 40
+; CHECK-NEXT: pushq %r12
+; CHECK-NEXT: .cfi_def_cfa_offset 48
+; CHECK-NEXT: pushq %rbx
+; CHECK-NEXT: .cfi_def_cfa_offset 56
+; CHECK-NEXT: .cfi_offset %rbx, -56
+; CHECK-NEXT: .cfi_offset %r12, -48
+; CHECK-NEXT: .cfi_offset %r13, -40
+; CHECK-NEXT: .cfi_offset %r14, -32
+; CHECK-NEXT: .cfi_offset %r15, -24
+; CHECK-NEXT: .cfi_offset %rbp, -16
+; CHECK-NEXT: movq %rdi, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; CHECK-NEXT: #APP # 8-byte Folded Spill
+; CHECK-NEXT: # -{{[0-9]+}}(%rsp) %rdi %rsi %rdx %rcx %r8 %r9 .LBB0_3
+; CHECK-NEXT: #NO_APP
+; CHECK-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rax # 8-byte Reload
+; CHECK-NEXT: movq %rax, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; CHECK-NEXT: # %bb.1: # %fallthrough
+; CHECK-NEXT: xorl %eax, %eax
+; CHECK-NEXT: .LBB0_2: # %fallthrough
+; CHECK-NEXT: popq %rbx
+; CHECK-NEXT: .cfi_def_cfa_offset 48
+; CHECK-NEXT: popq %r12
+; CHECK-NEXT: .cfi_def_cfa_offset 40
+; CHECK-NEXT: popq %r13
+; CHECK-NEXT: .cfi_def_cfa_offset 32
+; CHECK-NEXT: popq %r14
+; CHECK-NEXT: .cfi_def_cfa_offset 24
+; CHECK-NEXT: popq %r15
+; CHECK-NEXT: .cfi_def_cfa_offset 16
+; CHECK-NEXT: popq %rbp
+; CHECK-NEXT: .cfi_def_cfa_offset 8
+; CHECK-NEXT: retq
+; CHECK-NEXT: .LBB0_3: # Inline asm indirect target
+; CHECK-NEXT: # %indirect
+; CHECK-NEXT: # Label of block must be emitted
+; CHECK-NEXT: .cfi_def_cfa_offset 56
+; CHECK-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rax # 8-byte Reload
+; CHECK-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rcx # 8-byte Reload
+; CHECK-NEXT: addq %rcx, %rax
+; CHECK-NEXT: jmp .LBB0_2
+entry:
+ %r = callbr i64 asm sideeffect "# $0 $1 $2 $3 $4 $5 $6 ${7:l}",
+ "=&rm,r,r,r,r,r,r,!i,~{rax},~{rbx},~{rbp},~{r10},~{r11},~{r12},~{r13},~{r14},~{r15}"
+ (i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %f)
+ to label %fallthrough [label %indirect]
+
+fallthrough:
+ ret i64 0
+
+indirect:
+ %sum = add i64 %r, %a
+ ret i64 %sum
+}
diff --git a/llvm/test/CodeGen/X86/asm-constraints-rm-pressure.ll b/llvm/test/CodeGen/X86/asm-constraints-rm-pressure.ll
new file mode 100644
index 0000000000000..82b897e374a88
--- /dev/null
+++ b/llvm/test/CodeGen/X86/asm-constraints-rm-pressure.ll
@@ -0,0 +1,109 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 6
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --regalloc=greedy < %s | FileCheck --check-prefix=GREEDY_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --regalloc=fast < %s | FileCheck --check-prefix=FAST_RA %s
+
+; No -O0 RUN line here: -O0 codegen for "rm" is untouched by this patch (see
+; getConstraintPreferences()'s CodeGenOptLevel::None opt-out), and this
+; specific combination -- a register-exhausting "=&rm" output, at -O0 --
+; already hits a separate, pre-existing crash in computeConstraintToUse()
+; ("Can only indirectify direct input operands!") on unmodified upstream
+; main, unrelated to MayFoldRegister. Not this patch's to fix.
+
+; Exhaust all 14 usable x86 GPRs so the fast allocator has no register left
+; for the "=&rm" output: 6 inputs consume rdi/rsi/rdx/rcx/r8/r9, and the asm
+; clobbers rax/rbx/rbp/r10-r15. Before RegAllocFast::foldFoldableInlineAsmOperands()
+; this hard-errored ("inline assembly requires more registers than
+; available"); now it folds to a stack slot exactly like the greedy
+; allocator already could via InlineSpiller.
+define i64 @test_rm_output_pressure(i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %f) {
+; GREEDY_RA-LABEL: test_rm_output_pressure:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: pushq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: pushq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: pushq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: pushq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: pushq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: pushq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 56
+; GREEDY_RA-NEXT: .cfi_offset %rbx, -56
+; GREEDY_RA-NEXT: .cfi_offset %r12, -48
+; GREEDY_RA-NEXT: .cfi_offset %r13, -40
+; GREEDY_RA-NEXT: .cfi_offset %r14, -32
+; GREEDY_RA-NEXT: .cfi_offset %r15, -24
+; GREEDY_RA-NEXT: .cfi_offset %rbp, -16
+; GREEDY_RA-NEXT: movq %r9, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %r8, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rcx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rdx, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq %rsi, {{[-0-9]+}}(%r{{[sb]}}p) # 8-byte Spill
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %r9 # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %r8 # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rcx # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rdx # 8-byte Reload
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rsi # 8-byte Reload
+; GREEDY_RA-NEXT: #APP # 8-byte Folded Spill
+; GREEDY_RA-NEXT: # -{{[0-9]+}}(%rsp) %rdi %rsi %rdx %rcx %r8 %r9
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rax # 8-byte Reload
+; GREEDY_RA-NEXT: popq %rbx
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 48
+; GREEDY_RA-NEXT: popq %r12
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 40
+; GREEDY_RA-NEXT: popq %r13
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 32
+; GREEDY_RA-NEXT: popq %r14
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 24
+; GREEDY_RA-NEXT: popq %r15
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: popq %rbp
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 8
+; GREEDY_RA-NEXT: retq
+;
+; FAST_RA-LABEL: test_rm_output_pressure:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: pushq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: pushq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: pushq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: pushq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: pushq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: pushq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 56
+; FAST_RA-NEXT: .cfi_offset %rbx, -56
+; FAST_RA-NEXT: .cfi_offset %r12, -48
+; FAST_RA-NEXT: .cfi_offset %r13, -40
+; FAST_RA-NEXT: .cfi_offset %r14, -32
+; FAST_RA-NEXT: .cfi_offset %r15, -24
+; FAST_RA-NEXT: .cfi_offset %rbp, -16
+; FAST_RA-NEXT: #APP # 8-byte Folded Spill
+; FAST_RA-NEXT: # -{{[0-9]+}}(%rsp) %rdi %rsi %rdx %rcx %r8 %r9
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: movq {{[-0-9]+}}(%r{{[sb]}}p), %rax # 8-byte Reload
+; FAST_RA-NEXT: popq %rbx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 48
+; FAST_RA-NEXT: popq %r12
+; FAST_RA-NEXT: .cfi_def_cfa_offset 40
+; FAST_RA-NEXT: popq %r13
+; FAST_RA-NEXT: .cfi_def_cfa_offset 32
+; FAST_RA-NEXT: popq %r14
+; FAST_RA-NEXT: .cfi_def_cfa_offset 24
+; FAST_RA-NEXT: popq %r15
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: popq %rbp
+; FAST_RA-NEXT: .cfi_def_cfa_offset 8
+; FAST_RA-NEXT: retq
+entry:
+ %0 = call i64 asm sideeffect "# $0 $1 $2 $3 $4 $5 $6",
+ "=&rm,r,r,r,r,r,r,~{rax},~{rbx},~{rbp},~{r10},~{r11},~{r12},~{r13},~{r14},~{r15}"
+ (i64 %a, i64 %b, i64 %c, i64 %d, i64 %e, i64 %f)
+ ret i64 %0
+}
>From aaa8d29990ad708cbc8f5f771c0d636ea89f86ba Mon Sep 17 00:00:00 2001
From: Bill Wendling <isanbard at gmail.com>
Date: Tue, 4 Aug 2026 03:29:48 -0700
Subject: [PATCH 03/12] [RegAllocFast] Assert the CopyMI invariant instead of
silently ignoring it
foldFoldableInlineAsmOperands() declares CopyMI, passes it to
TII->foldMemoryOperand(), and never checked it afterward.
foldMemoryOperand()'s general contract is that CopyMI reports a copy
instruction the fold synthesized, which the caller is responsible for
(see InlineSpiller's own caller, which does handle it) -- but
TargetInstrInfo.cpp's implementation early-returns via
foldInlineAsmMemOperand() for the MI.isInlineAsm() case specifically,
before CopyMI is ever touched, so it's genuinely always null on this
call path today. That guarantee held only by virtue of an
implementation detail three layers away with nothing at this call
site documenting or checking it -- if it ever stopped holding (e.g. a
future change to how inline asm folding is implemented, or a new
target-specific path), this would silently drop whatever instruction
CopyMI pointed to instead of inserting it.
Add an assert(!CopyMI) with a comment explaining why, so a future
violation fails loudly at the point it happens instead of miscompiling
silently. Verified it doesn't fire across the full rm-specific test
suite (asm-constraints-torture.ll, the pressure/callbr-fold tests, the
.mir-based exhaustion test).
Co-Authored-By: Claude Sonnet 5 <noreply at anthropic.com>
Claude-Session: https://claude.ai/code/session_0143ozRFwhFjWmJjtfZUgmYe
---
llvm/lib/CodeGen/RegAllocFast.cpp | 11 +++++++++++
1 file changed, 11 insertions(+)
diff --git a/llvm/lib/CodeGen/RegAllocFast.cpp b/llvm/lib/CodeGen/RegAllocFast.cpp
index 1109f25f3b214..14cfdd5935741 100644
--- a/llvm/lib/CodeGen/RegAllocFast.cpp
+++ b/llvm/lib/CodeGen/RegAllocFast.cpp
@@ -1976,9 +1976,20 @@ void RegAllocFastImpl::foldFoldableInlineAsmOperands(
TRI->getSpillStackID(*RC));
}
+ // CopyMI is an out-parameter foldMemoryOperand() uses to report a copy
+ // instruction it had to synthesize as part of folding (see its own
+ // declaration for that general contract) -- but TargetInstrInfo.cpp's
+ // implementation early-returns via foldInlineAsmMemOperand() for the
+ // MI.isInlineAsm() case specifically, before CopyMI is ever touched, so
+ // it's guaranteed to stay null here. Asserted rather than silently
+ // ignored so that guarantee breaking -- e.g. a future change to how
+ // inline asm folding is implemented -- fails loudly instead of quietly
+ // dropping whatever CopyMI would have pointed to.
MachineInstr *CopyMI = nullptr;
MachineInstr *NewMI = TII->foldMemoryOperand(*MI, {I}, FrameIndex, CopyMI);
assert(NewMI && "operand was reported foldable but folding failed");
+ assert(!CopyMI &&
+ "inline asm folding shouldn't synthesize a copy instruction");
if (!NewMI)
continue;
>From b72a68bf28fb1230cc381cb5b0d936eddd0c4819 Mon Sep 17 00:00:00 2001
From: Bill Wendling <isanbard at gmail.com>
Date: Tue, 4 Aug 2026 01:45:30 -0700
Subject: [PATCH 04/12] [CodeGen] Track inline asm presence per
MachineBasicBlock, not just per function
RegAllocFast::allocateBasicBlock() was gated on a single
function-wide MFHasInlineAsm flag (MachineFunction::hasInlineAsm()),
so a function with inline asm in just one block still paid a full
per-instruction scan (foldFoldableInlineAsmOperands()'s
isInlineAsm() loop) for every other block in that function. For a
large function with inline asm concentrated in a small part of it,
that's needless work on every asm-free block.
Add MachineBasicBlock::HasInlineAsm, mirroring
MachineFunction::HasInlineAsm at block granularity. The key property
that makes this worth doing is that it costs nothing new to compute:
every place that already sets the function-wide flag is already
iterating every instruction of every block for other reasons, so
recording the enclosing block at the same time is free.
- SelectionDAGISel.cpp's post-ISel scan (determines MFI.hasCalls()
and MF->hasInlineAsm()) now also calls MBB.setHasInlineAsm(). This
removes that loop's early-exit optimization (previously breaking
out once both function-wide aggregates were known true) --
per-block accuracy means every block must actually be visited, so
the two are incompatible. In practice this rarely matters: the
break only fired for functions with both a call and inline asm
found early, and most functions contain no inline asm at all, so
the loop was already running to completion for the common case.
- GlobalISel/InstructionSelect.cpp's equivalent (explicitly "ported
from SelectionDAG") gets the identical treatment.
- MIRParser.cpp's computeFunctionProperties(), which independently
recomputes MF.HasInlineAsm by scanning when parsing a .mir file
directly (bypassing ISel), gets the same treatment. This matters
for any test that hand-writes MIR with INLINEASM instructions
(e.g. CodeGen/MIR/X86/inline-asm-rm-exhaustion.mir) -- without
this, such tests would silently get a false block-level flag and
RegAllocFast would skip folding for them.
RegAllocFast::allocateBasicBlock() now checks MBB.hasInlineAsm()
directly instead of the function-wide MFHasInlineAsm, which is
removed as redundant.
Like the function-wide flag this mirrors, this is a snapshot from
ISel/parse time, not incrementally maintained by later passes that
might insert, remove, or move an INLINEASM instruction between
blocks. A stale false-negative fails safe: RegAllocFast's existing
"inline assembly requires more registers than available" hard-error
path remains the backstop, not silent miscompilation.
Verified: -verify-machineinstrs on the .mir-based exhaustion test and
all rm-specific X86/AArch64 tests added in prior commits (identical
codegen, confirming this is a pure cost optimization with no
behavior change). Full CodeGen/ + clang/test/CodeGen/ + MIR/ +
MachineVerifier/ (37929 tests) passes; the handful of failures seen
across repeated runs are non-reproducible (different tests fail each
run, all unrelated to inline asm, all pass standalone) -- pre-existing
ThinLTO/plugin test-harness flakiness under parallel execution on
this machine, not a regression.
Co-Authored-By: Claude Sonnet 5 <noreply at anthropic.com>
Claude-Session: https://claude.ai/code/session_0143ozRFwhFjWmJjtfZUgmYe
---
llvm/include/llvm/CodeGen/MachineBasicBlock.h | 21 +++++++++++++++++++
.../CodeGen/GlobalISel/InstructionSelect.cpp | 15 ++++++-------
llvm/lib/CodeGen/MIRParser/MIRParser.cpp | 6 ++++--
llvm/lib/CodeGen/RegAllocFast.cpp | 20 +++++++-----------
.../CodeGen/SelectionDAG/SelectionDAGISel.cpp | 14 ++++++++-----
5 files changed, 49 insertions(+), 27 deletions(-)
diff --git a/llvm/include/llvm/CodeGen/MachineBasicBlock.h b/llvm/include/llvm/CodeGen/MachineBasicBlock.h
index 47430a100b3cc..152d8205867f4 100644
--- a/llvm/include/llvm/CodeGen/MachineBasicBlock.h
+++ b/llvm/include/llvm/CodeGen/MachineBasicBlock.h
@@ -227,6 +227,18 @@ class MachineBasicBlock
/// Indicate that this basic block is the indirect dest of an INLINEASM_BR.
bool IsInlineAsmBrIndirectTarget = false;
+ /// Indicate that this basic block directly contains an INLINEASM or
+ /// INLINEASM_BR instruction, as of the last time it was computed (see
+ /// MachineFunction::hasInlineAsm(), which this mirrors at block
+ /// granularity). Like that flag, this is set once -- during instruction
+ /// selection, or when parsing MIR directly -- and is not
+ /// incrementally maintained afterward, so a pass that inserts, removes, or
+ /// moves an INLINEASM/INLINEASM_BR between blocks after that point can make
+ /// it stale. Consumers that rely on it (e.g. to skip per-block scanning)
+ /// must tolerate that: a false negative here should fail safe (e.g. by
+ /// falling back to a slower correct path), never silently miscompile.
+ bool HasInlineAsm = false;
+
/// since getSymbol is a relatively heavy-weight operation, the symbol
/// is only computed once and is cached.
mutable MCSymbol *CachedMCSymbol = nullptr;
@@ -742,6 +754,15 @@ class MachineBasicBlock
IsInlineAsmBrIndirectTarget = V;
}
+ /// Returns true if this block directly contains an INLINEASM or
+ /// INLINEASM_BR instruction, as of the last time it was computed. See the
+ /// HasInlineAsm field comment for the staleness caveat.
+ bool hasInlineAsm() const { return HasInlineAsm; }
+
+ /// Indicates if this block directly contains an INLINEASM or INLINEASM_BR
+ /// instruction.
+ void setHasInlineAsm(bool V = true) { HasInlineAsm = V; }
+
/// Returns true if it is legal to hoist instructions into this block.
LLVM_ABI bool isLegalToHoistInto() const;
diff --git a/llvm/lib/CodeGen/GlobalISel/InstructionSelect.cpp b/llvm/lib/CodeGen/GlobalISel/InstructionSelect.cpp
index c750f643e99e2..d042e952ba39a 100644
--- a/llvm/lib/CodeGen/GlobalISel/InstructionSelect.cpp
+++ b/llvm/lib/CodeGen/GlobalISel/InstructionSelect.cpp
@@ -308,18 +308,19 @@ bool InstructionSelect::selectMachineFunction(MachineFunction &MF) {
return false;
}
- // Determine if there are any calls in this machine function. Ported from
- // SelectionDAG.
+ // Determine if there are any calls in this machine function, and which
+ // blocks directly contain inline asm. Ported from SelectionDAG -- see the
+ // comment there on why this can no longer bail out early once
+ // MFI.hasCalls() && MF.hasInlineAsm() are both known true.
MachineFrameInfo &MFI = MF.getFrameInfo();
- for (const auto &MBB : MF) {
- if (MFI.hasCalls() && MF.hasInlineAsm())
- break;
-
+ for (auto &MBB : MF) {
for (const auto &MI : MBB) {
if ((MI.isCall() && !MI.isReturn()) || MI.isStackAligningInlineAsm())
MFI.setHasCalls(true);
- if (MI.isInlineAsm())
+ if (MI.isInlineAsm()) {
MF.setHasInlineAsm(true);
+ MBB.setHasInlineAsm(true);
+ }
}
}
diff --git a/llvm/lib/CodeGen/MIRParser/MIRParser.cpp b/llvm/lib/CodeGen/MIRParser/MIRParser.cpp
index 6f1e7594f34da..5df1fb7a1590b 100644
--- a/llvm/lib/CodeGen/MIRParser/MIRParser.cpp
+++ b/llvm/lib/CodeGen/MIRParser/MIRParser.cpp
@@ -410,12 +410,14 @@ bool MIRParserImpl::computeFunctionProperties(
bool HasInlineAsm = false;
bool HasFakeUses = false;
bool AllTiedOpsRewritten = true, HasTiedOps = false;
- for (const MachineBasicBlock &MBB : MF) {
+ for (MachineBasicBlock &MBB : MF) {
for (const MachineInstr &MI : MBB) {
if (MI.isPHI())
HasPHI = true;
- if (MI.isInlineAsm())
+ if (MI.isInlineAsm()) {
HasInlineAsm = true;
+ MBB.setHasInlineAsm(true);
+ }
if (MI.isFakeUse())
HasFakeUses = true;
for (unsigned I = 0; I < MI.getNumOperands(); ++I) {
diff --git a/llvm/lib/CodeGen/RegAllocFast.cpp b/llvm/lib/CodeGen/RegAllocFast.cpp
index 14cfdd5935741..0a0007eaf14dc 100644
--- a/llvm/lib/CodeGen/RegAllocFast.cpp
+++ b/llvm/lib/CodeGen/RegAllocFast.cpp
@@ -402,13 +402,6 @@ class RegAllocFastImpl {
bool mayBeSpillFromInlineAsmBr(const MachineInstr &MI) const;
- /// Cached copy of MachineFunction::hasInlineAsm(), set once in
- /// runOnMachineFunction(). Lets allocateBasicBlock() skip scanning each
- /// block for foldable inline asm operands with a single function-wide
- /// check, since the overwhelming majority of functions contain no inline
- /// asm at all.
- bool MFHasInlineAsm = false;
-
void selectInlineAsmOperandsToFold(MachineInstr &MI,
SmallSet<Register, 8> &ToFold);
void foldFoldableInlineAsmOperands(MachineBasicBlock &MBB);
@@ -2061,11 +2054,13 @@ void RegAllocFastImpl::allocateBasicBlock(MachineBasicBlock &MBB) {
// their memory form before the main allocation loop runs, so those
// operands never compete for a register at all -- see
// foldFoldableInlineAsmOperands() for why this can't be done lazily like
- // greedy's on-demand InlineSpiller folding. Gated on MFHasInlineAsm so
- // functions with no inline asm at all -- the common case -- pay nothing
- // beyond the one function-wide check already made in
- // runOnMachineFunction(), instead of an extra per-block scan.
- if (MFHasInlineAsm)
+ // greedy's on-demand InlineSpiller folding. Gated on
+ // MachineBasicBlock::hasInlineAsm(), which ISel (or the MIR parser, for
+ // directly-parsed .mir input) already populates for free as a byproduct of
+ // a scan it does anyway -- see its declaration for the staleness caveat.
+ // This means blocks that don't themselves contain inline asm pay only an
+ // O(1) check, even in functions where some other block does.
+ if (MBB.hasInlineAsm())
foldFoldableInlineAsmOperands(MBB);
// Traverse block in reverse order allocating instructions one by one.
@@ -2124,7 +2119,6 @@ bool RegAllocFastImpl::runOnMachineFunction(MachineFunction &MF) {
TRI = STI.getRegisterInfo();
TII = STI.getInstrInfo();
MFI = &MF.getFrameInfo();
- MFHasInlineAsm = MF.hasInlineAsm();
MRI->freezeReservedRegs();
RegClassInfo.runOnMachineFunction(MF);
unsigned NumRegUnits = TRI->getNumRegUnits();
diff --git a/llvm/lib/CodeGen/SelectionDAG/SelectionDAGISel.cpp b/llvm/lib/CodeGen/SelectionDAG/SelectionDAGISel.cpp
index a9140c740eeb4..b8aecdb2e3119 100644
--- a/llvm/lib/CodeGen/SelectionDAG/SelectionDAGISel.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/SelectionDAGISel.cpp
@@ -793,12 +793,15 @@ bool SelectionDAGISel::runOnMachineFunction(MachineFunction &mf) {
if (MF->useDebugInstrRef())
MF->finalizeDebugInstrRefs();
- // Determine if there are any calls in this machine function.
+ // Determine if there are any calls in this machine function, and which
+ // blocks directly contain inline asm (INLINEASM/INLINEASM_BR) -- the
+ // latter lets RegAllocFast skip its own per-block scan later. This can no
+ // longer bail out once MFI.hasCalls() && MF->hasInlineAsm() are both known
+ // true, unlike before: later blocks may still need their own
+ // MachineBasicBlock::HasInlineAsm bit set even after the function-wide
+ // aggregates are already satisfied.
MachineFrameInfo &MFI = MF->getFrameInfo();
- for (const auto &MBB : *MF) {
- if (MFI.hasCalls() && MF->hasInlineAsm())
- break;
-
+ for (auto &MBB : *MF) {
for (const auto &MI : MBB) {
const MCInstrDesc &MCID = TII->get(MI.getOpcode());
if ((MCID.isCall() && !MCID.isReturn()) ||
@@ -807,6 +810,7 @@ bool SelectionDAGISel::runOnMachineFunction(MachineFunction &mf) {
}
if (MI.isInlineAsm()) {
MF->setHasInlineAsm(true);
+ MBB.setHasInlineAsm(true);
}
}
}
>From cc4892bb1161ed73ac9e9df3b4a5ac5138cf6cb3 Mon Sep 17 00:00:00 2001
From: Bill Wendling <isanbard at gmail.com>
Date: Tue, 4 Aug 2026 00:06:17 -0700
Subject: [PATCH 05/12] [test] Port applicable "rm" constraint coverage from
asm-constraint-br
Ports the parts of asm-constraint-br's test suite that exercise the
same outcome this branch's simpler mechanism achieves -- "rm" prefers
'r', falling back to 'm' under pressure -- adapted to no longer need
that branch's llvm.asm.constraint.br intrinsic / dual-path callbr
scaffolding (Clang emitting two IR arms and InlineAsmPrepare picking
one pre-ISel). Since this branch makes the same choice directly in
TargetLowering::getConstraintPreferences() and RegAllocFast, each
ported test keeps only the "prefer register" arm as a plain function
and drops the redundant "prefer memory" arm and the callbr chooser.
asm-constraints-torture.ll: seven scenarios (each with a no-pressure
and pressure variant) not covered by the existing rm test files on
this branch: multiple simultaneous "rm" inputs, multiple simultaneous
"rm" outputs, tied "+rm" read-write outputs, a mixed "=rm" output with
a plain "r" input, a forced-memory "m" output alongside an "rm" input,
multiple forced-memory outputs alongside "rm" inputs, and multiple
early-clobber "=&rm" outputs alongside "rm" inputs. The tied-output and
multiple-simultaneous-early-clobber-output cases are the most
significant of these: they weren't exercised by any test on this
branch before, and they're exactly the shape of case
RegAllocFast::selectInlineAsmOperandsToFold()'s per-register-class
demand counting exists to get right.
inline-asm-callbase.ll: "rm" constraints on CallBase subclasses other
than a plain call (invoke's EH unwind edges, callbr's indirect-branch
asm goto). None of its three functions used the constraint-br
intrinsic to begin with.
AArch64/inline-asm-rm.ll: cross-ISel-framework coverage (GlobalISel,
FastISel, SelectionDAG) x (greedy, fast) that this branch had zero of
before (all prior rm tests are X86-only). Notably, X86's FastISel path
here comes out cleaner than on asm-constraint-br: there, FastISel took
a different path from GlobalISel/SelectionDAG and needlessly spilled
through a stack slot even with no register pressure; here FastISel
shares the same TargetLowering::getConstraintPreferences() call as
every other framework, so it keeps the value in a register directly.
Not ported, and why:
- asm-constraints-torture-exhaust-regs.ll: duplicates the register-
exhaustion scenario already covered by asm-constraints-rm-pressure.ll.
- inline-asm-prepare.ll, inline-asm-prepare-memory.ll: unit tests for
the InlineAsmPrepare pass itself, which this branch doesn't have.
- clang/test/CodeGen/asm-reg-mem-constraints.c,
AArch64/asm-constraint-ro-no-dual-path.c: assert on Clang's dual-path
callbr IR emission, a CGStmt.cpp behavior this branch's Clang doesn't
have -- vanilla Clang here just emits "rm" directly and lets the
backend decide, so there's no dual-path structure to check.
- *-pipeline.ll / *-pipeline-npm.ll across every target: assert on
InlineAsmPrepare's position in the pass pipeline. That pass doesn't
exist on this branch, so there's nothing meaningful to port -- the
existing (untouched) pipeline tests already pass unmodified, which
is the correct state here.
Verified: all six ported/added RUN configurations compile cleanly
(-verify-machineinstrs where applicable) before check-line generation;
full CodeGen/ + clang/test/CodeGen/ (37375 tests) passes with zero
failures beyond the pre-existing unrelated ones already present on
unmodified asm-rm.
Co-Authored-By: Claude Sonnet 5 <noreply at anthropic.com>
Claude-Session: https://claude.ai/code/session_0143ozRFwhFjWmJjtfZUgmYe
---
llvm/test/CodeGen/AArch64/inline-asm-rm.ll | 62 ++
.../CodeGen/X86/asm-constraints-torture.ll | 768 ++++++++++++++++++
llvm/test/CodeGen/X86/inline-asm-callbase.ll | 227 ++++++
3 files changed, 1057 insertions(+)
create mode 100644 llvm/test/CodeGen/AArch64/inline-asm-rm.ll
create mode 100644 llvm/test/CodeGen/X86/asm-constraints-torture.ll
create mode 100644 llvm/test/CodeGen/X86/inline-asm-callbase.ll
diff --git a/llvm/test/CodeGen/AArch64/inline-asm-rm.ll b/llvm/test/CodeGen/AArch64/inline-asm-rm.ll
new file mode 100644
index 0000000000000..817b8fb021864
--- /dev/null
+++ b/llvm/test/CodeGen/AArch64/inline-asm-rm.ll
@@ -0,0 +1,62 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 6
+; RUN: llc -mtriple=aarch64-unknown-linux-gnu --global-isel=true --fast-isel=false --regalloc=greedy < %s \
+; RUN: | FileCheck --check-prefix=GLOBAL_ISEL_GREEDY_RA %s
+; RUN: llc -mtriple=aarch64-unknown-linux-gnu --global-isel=true --fast-isel=false --regalloc=fast < %s \
+; RUN: | FileCheck --check-prefix=GLOBAL_ISEL_FAST_RA %s
+; RUN: llc -mtriple=aarch64-unknown-linux-gnu --global-isel=false --fast-isel=true --regalloc=greedy < %s \
+; RUN: | FileCheck --check-prefix=FAST_ISEL_GREEDY_RA %s
+; RUN: llc -mtriple=aarch64-unknown-linux-gnu --global-isel=false --fast-isel=true --regalloc=fast < %s \
+; RUN: | FileCheck --check-prefix=FAST_ISEL_FAST_RA %s
+; RUN: llc -mtriple=aarch64-unknown-linux-gnu --global-isel=false --fast-isel=false --regalloc=greedy < %s \
+; RUN: | FileCheck --check-prefix=GREEDY_RA %s
+; RUN: llc -mtriple=aarch64-unknown-linux-gnu --global-isel=false --fast-isel=false --regalloc=fast < %s \
+; RUN: | FileCheck --check-prefix=FAST_RA %s
+
+; Test rm constraints on AArch64 under all three ISel frameworks.
+
+define i64 @test_rm_output() {
+; GLOBAL_ISEL_GREEDY_RA-LABEL: test_rm_output:
+; GLOBAL_ISEL_GREEDY_RA: // %bb.0: // %entry
+; GLOBAL_ISEL_GREEDY_RA-NEXT: //APP
+; GLOBAL_ISEL_GREEDY_RA-NEXT: // x0
+; GLOBAL_ISEL_GREEDY_RA-NEXT: //NO_APP
+; GLOBAL_ISEL_GREEDY_RA-NEXT: ret
+;
+; GLOBAL_ISEL_FAST_RA-LABEL: test_rm_output:
+; GLOBAL_ISEL_FAST_RA: // %bb.0: // %entry
+; GLOBAL_ISEL_FAST_RA-NEXT: //APP
+; GLOBAL_ISEL_FAST_RA-NEXT: // x0
+; GLOBAL_ISEL_FAST_RA-NEXT: //NO_APP
+; GLOBAL_ISEL_FAST_RA-NEXT: ret
+;
+; FAST_ISEL_GREEDY_RA-LABEL: test_rm_output:
+; FAST_ISEL_GREEDY_RA: // %bb.0: // %entry
+; FAST_ISEL_GREEDY_RA-NEXT: //APP
+; FAST_ISEL_GREEDY_RA-NEXT: // x0
+; FAST_ISEL_GREEDY_RA-NEXT: //NO_APP
+; FAST_ISEL_GREEDY_RA-NEXT: ret
+;
+; FAST_ISEL_FAST_RA-LABEL: test_rm_output:
+; FAST_ISEL_FAST_RA: // %bb.0: // %entry
+; FAST_ISEL_FAST_RA-NEXT: //APP
+; FAST_ISEL_FAST_RA-NEXT: // x0
+; FAST_ISEL_FAST_RA-NEXT: //NO_APP
+; FAST_ISEL_FAST_RA-NEXT: ret
+;
+; GREEDY_RA-LABEL: test_rm_output:
+; GREEDY_RA: // %bb.0: // %entry
+; GREEDY_RA-NEXT: //APP
+; GREEDY_RA-NEXT: // x0
+; GREEDY_RA-NEXT: //NO_APP
+; GREEDY_RA-NEXT: ret
+;
+; FAST_RA-LABEL: test_rm_output:
+; FAST_RA: // %bb.0: // %entry
+; FAST_RA-NEXT: //APP
+; FAST_RA-NEXT: // x0
+; FAST_RA-NEXT: //NO_APP
+; FAST_RA-NEXT: ret
+entry:
+ %0 = tail call i64 asm "# $0", "=rm"()
+ ret i64 %0
+}
diff --git a/llvm/test/CodeGen/X86/asm-constraints-torture.ll b/llvm/test/CodeGen/X86/asm-constraints-torture.ll
new file mode 100644
index 0000000000000..68f547ef0e1ee
--- /dev/null
+++ b/llvm/test/CodeGen/X86/asm-constraints-torture.ll
@@ -0,0 +1,768 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --filter "^\t#" --version 4
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=true --regalloc=greedy < %s \
+; RUN: | FileCheck --check-prefix=FAST_ISEL_GREEDY_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=true --regalloc=fast < %s \
+; RUN: | FileCheck --check-prefix=FAST_ISEL_FAST_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=false --regalloc=greedy < %s \
+; RUN: | FileCheck --check-prefix=GREEDY_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=false --regalloc=fast < %s \
+; RUN: | FileCheck --check-prefix=FAST_RA %s
+
+; Combinatorial coverage for "rm"-style inline asm constraints: inputs,
+; outputs, tied read-write operands, mixed with plain "r"/"m" operands and
+; early-clobbers, each with and without enough register pressure to force a
+; fold to memory. TargetLowering::getConstraintPreferences() should keep
+; each "rm" operand in a register when there's no pressure, and
+; RegAllocFast::foldFoldableInlineAsmOperands() (fast RA) /
+; InlineSpiller (greedy RA) should fold it to a stack slot when there is.
+
+define dso_local i32 @test_rm_input_no_pressure(ptr noundef readonly captures(none) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_rm_input_no_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm input: no pressure
+; FAST_ISEL_GREEDY_RA: # %eax, %ecx, %edx, %esi, %r8d
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_rm_input_no_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm input: no pressure
+; FAST_ISEL_FAST_RA: # %eax, %ecx, %edx, %esi, %r8d
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_rm_input_no_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm input: no pressure
+; GREEDY_RA: # %eax, %ecx, %edx, %esi, %r8d
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_rm_input_no_pressure:
+; FAST_RA: #APP
+; FAST_RA: # rm input: no pressure
+; FAST_RA: # %eax, %ecx, %edx, %esi, %r8d
+; FAST_RA: #NO_APP
+entry:
+ %0 = load i32, ptr %foo, align 4
+ %b2 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %1 = load i32, ptr %b2, align 4
+ %c3 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %2 = load i32, ptr %c3, align 4
+ %d4 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %3 = load i32, ptr %d4, align 4
+ %e5 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %4 = load i32, ptr %e5, align 4
+ tail call void asm sideeffect "# rm input: no pressure\0A\09# $0, $1, $2, $3, $4", "rm,rm,rm,rm,rm,~{dirflag},~{fpsr},~{flags}"(i32 %0, i32 %1, i32 %2, i32 %3, i32 %4)
+ %5 = load i32, ptr %foo, align 4
+ ret i32 %5
+}
+
+define dso_local i32 @test_rm_pressure(ptr noundef readonly captures(none) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_rm_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_GREEDY_RA: #NO_APP
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm input: pressure
+; FAST_ISEL_GREEDY_RA: # %esi, %edi, %r8d, %r9d, %eax
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_rm_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_FAST_RA: #NO_APP
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm input: pressure
+; FAST_ISEL_FAST_RA: # %eax, %ecx, %edx, %esi, %edi
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_rm_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; GREEDY_RA: #NO_APP
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm input: pressure
+; GREEDY_RA: # %esi, %edi, %r8d, %r9d, %eax
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_rm_pressure:
+; FAST_RA: #APP
+; FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_RA: #NO_APP
+; FAST_RA: #APP
+; FAST_RA: # rm input: pressure
+; FAST_RA: # %eax, %ecx, %edx, %esi, %edi
+; FAST_RA: #NO_APP
+entry:
+ %0 = tail call { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } asm sideeffect "# $0,$1,$2,$3,$4,$5,$6,$7,$8,$9,$10, $11, $12, $13, $14", "={rax},={rcx},={rdx},={rsi},={rdi},={rbx},={rbp},={r8},={r9},={r10},={r11},={r12},={r13},={r14},={r15},~{dirflag},~{fpsr},~{flags}"()
+ %1 = load i32, ptr %foo, align 4
+ %b16 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %2 = load i32, ptr %b16, align 4
+ %c17 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %3 = load i32, ptr %c17, align 4
+ %d18 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %4 = load i32, ptr %d18, align 4
+ %e19 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %5 = load i32, ptr %e19, align 4
+ tail call void asm sideeffect "# rm input: pressure\0A\09# $0, $1, $2, $3, $4", "rm,rm,rm,rm,rm,~{dirflag},~{fpsr},~{flags}"(i32 %1, i32 %2, i32 %3, i32 %4, i32 %5)
+ %asmresult14 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 14
+ %asmresult13 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 13
+ %asmresult12 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 12
+ %asmresult11 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 11
+ %asmresult10 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 10
+ %asmresult9 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 9
+ %asmresult8 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 8
+ %asmresult7 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 7
+ %asmresult6 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 6
+ %asmresult5 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 5
+ %asmresult4 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 4
+ %asmresult3 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 3
+ %asmresult2 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 2
+ %asmresult1 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 1
+ %asmresult = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 0
+ tail call void @g(i64 noundef %asmresult, i64 noundef %asmresult1, i64 noundef %asmresult2, i64 noundef %asmresult3, i64 noundef %asmresult4, i64 noundef %asmresult5, i64 noundef %asmresult6, i64 noundef %asmresult7, i64 noundef %asmresult8, i64 noundef %asmresult9, i64 noundef %asmresult10, i64 noundef %asmresult11, i64 noundef %asmresult12, i64 noundef %asmresult13, i64 noundef %asmresult14)
+ %6 = load i32, ptr %foo, align 4
+ ret i32 %6
+}
+
+define dso_local i32 @test_output_no_pressure(ptr noundef writeonly captures(none) initializes((0, 20)) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_output_no_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm output: no pressure
+; FAST_ISEL_GREEDY_RA: # %eax, %ecx, %edx, %esi, %r8d
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_output_no_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm output: no pressure
+; FAST_ISEL_FAST_RA: # %eax, %r8d, %esi, %edx, %ecx
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_output_no_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm output: no pressure
+; GREEDY_RA: # %eax, %ecx, %edx, %esi, %r8d
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_output_no_pressure:
+; FAST_RA: #APP
+; FAST_RA: # rm output: no pressure
+; FAST_RA: # %eax, %r8d, %esi, %edx, %ecx
+; FAST_RA: #NO_APP
+entry:
+ %b2 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %c3 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %d4 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %e5 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %0 = tail call { i32, i32, i32, i32, i32 } asm sideeffect "# rm output: no pressure\0A\09# $0, $1, $2, $3, $4", "=rm,=rm,=rm,=rm,=rm,~{dirflag},~{fpsr},~{flags}"()
+ %asmresult = extractvalue { i32, i32, i32, i32, i32 } %0, 0
+ %asmresult6 = extractvalue { i32, i32, i32, i32, i32 } %0, 1
+ %asmresult7 = extractvalue { i32, i32, i32, i32, i32 } %0, 2
+ %asmresult8 = extractvalue { i32, i32, i32, i32, i32 } %0, 3
+ %asmresult9 = extractvalue { i32, i32, i32, i32, i32 } %0, 4
+ store i32 %asmresult, ptr %foo, align 4
+ store i32 %asmresult6, ptr %b2, align 4
+ store i32 %asmresult7, ptr %c3, align 4
+ store i32 %asmresult8, ptr %d4, align 4
+ store i32 %asmresult9, ptr %e5, align 4
+ ret i32 %asmresult
+}
+
+define dso_local i32 @test_output_pressure(ptr noundef writeonly captures(none) initializes((0, 20)) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_output_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_GREEDY_RA: #NO_APP
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm output: pressure
+; FAST_ISEL_GREEDY_RA: # %esi, %edi, %r8d, %r9d, %eax
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_output_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_FAST_RA: #NO_APP
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm output: pressure
+; FAST_ISEL_FAST_RA: # %eax, %edi, %ecx, %edx, %esi
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_output_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; GREEDY_RA: #NO_APP
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm output: pressure
+; GREEDY_RA: # %esi, %edi, %r8d, %r9d, %eax
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_output_pressure:
+; FAST_RA: #APP
+; FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_RA: #NO_APP
+; FAST_RA: #APP
+; FAST_RA: # rm output: pressure
+; FAST_RA: # %eax, %edi, %ecx, %edx, %esi
+; FAST_RA: #NO_APP
+entry:
+ %0 = tail call { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } asm sideeffect "# $0,$1,$2,$3,$4,$5,$6,$7,$8,$9,$10, $11, $12, $13, $14", "={rax},={rcx},={rdx},={rsi},={rdi},={rbx},={rbp},={r8},={r9},={r10},={r11},={r12},={r13},={r14},={r15},~{dirflag},~{fpsr},~{flags}"()
+ %b16 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %c17 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %d18 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %e19 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %1 = tail call { i32, i32, i32, i32, i32 } asm sideeffect "# rm output: pressure\0A\09# $0, $1, $2, $3, $4", "=rm,=rm,=rm,=rm,=rm,~{dirflag},~{fpsr},~{flags}"()
+ %asmresult20 = extractvalue { i32, i32, i32, i32, i32 } %1, 0
+ %asmresult21 = extractvalue { i32, i32, i32, i32, i32 } %1, 1
+ %asmresult22 = extractvalue { i32, i32, i32, i32, i32 } %1, 2
+ %asmresult23 = extractvalue { i32, i32, i32, i32, i32 } %1, 3
+ %asmresult24 = extractvalue { i32, i32, i32, i32, i32 } %1, 4
+ store i32 %asmresult20, ptr %foo, align 4
+ store i32 %asmresult21, ptr %b16, align 4
+ store i32 %asmresult22, ptr %c17, align 4
+ store i32 %asmresult23, ptr %d18, align 4
+ store i32 %asmresult24, ptr %e19, align 4
+ %asmresult14 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 14
+ %asmresult13 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 13
+ %asmresult12 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 12
+ %asmresult11 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 11
+ %asmresult10 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 10
+ %asmresult9 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 9
+ %asmresult8 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 8
+ %asmresult7 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 7
+ %asmresult6 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 6
+ %asmresult5 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 5
+ %asmresult4 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 4
+ %asmresult3 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 3
+ %asmresult2 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 2
+ %asmresult1 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 1
+ %asmresult = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 0
+ tail call void @g(i64 noundef %asmresult, i64 noundef %asmresult1, i64 noundef %asmresult2, i64 noundef %asmresult3, i64 noundef %asmresult4, i64 noundef %asmresult5, i64 noundef %asmresult6, i64 noundef %asmresult7, i64 noundef %asmresult8, i64 noundef %asmresult9, i64 noundef %asmresult10, i64 noundef %asmresult11, i64 noundef %asmresult12, i64 noundef %asmresult13, i64 noundef %asmresult14)
+ %2 = load i32, ptr %foo, align 4
+ ret i32 %2
+}
+
+define dso_local i32 @test_tied_output_no_pressure(ptr noundef captures(none) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_tied_output_no_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm tied output: no pressure
+; FAST_ISEL_GREEDY_RA: # %eax, %ecx, %edx, %esi, %r8d
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_tied_output_no_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm tied output: no pressure
+; FAST_ISEL_FAST_RA: # %eax, %r8d, %esi, %edx, %ecx
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_tied_output_no_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm tied output: no pressure
+; GREEDY_RA: # %eax, %ecx, %edx, %esi, %r8d
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_tied_output_no_pressure:
+; FAST_RA: #APP
+; FAST_RA: # rm tied output: no pressure
+; FAST_RA: # %eax, %r8d, %esi, %edx, %ecx
+; FAST_RA: #NO_APP
+entry:
+ %0 = load i32, ptr %foo, align 4
+ %b2 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %1 = load i32, ptr %b2, align 4
+ %c3 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %2 = load i32, ptr %c3, align 4
+ %d4 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %3 = load i32, ptr %d4, align 4
+ %e5 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %4 = load i32, ptr %e5, align 4
+ %5 = tail call { i32, i32, i32, i32, i32 } asm sideeffect "# rm tied output: no pressure\0A\09# $0, $1, $2, $3, $4", "=rm,=rm,=rm,=rm,=rm,0,1,2,3,4,~{dirflag},~{fpsr},~{flags}"(i32 %0, i32 %1, i32 %2, i32 %3, i32 %4)
+ %asmresult = extractvalue { i32, i32, i32, i32, i32 } %5, 0
+ %asmresult6 = extractvalue { i32, i32, i32, i32, i32 } %5, 1
+ %asmresult7 = extractvalue { i32, i32, i32, i32, i32 } %5, 2
+ %asmresult8 = extractvalue { i32, i32, i32, i32, i32 } %5, 3
+ %asmresult9 = extractvalue { i32, i32, i32, i32, i32 } %5, 4
+ store i32 %asmresult, ptr %foo, align 4
+ store i32 %asmresult6, ptr %b2, align 4
+ store i32 %asmresult7, ptr %c3, align 4
+ store i32 %asmresult8, ptr %d4, align 4
+ store i32 %asmresult9, ptr %e5, align 4
+ ret i32 %asmresult
+}
+
+define dso_local i32 @test_tied_output_pressure(ptr noundef captures(none) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_tied_output_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_GREEDY_RA: #NO_APP
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm tied output: pressure
+; FAST_ISEL_GREEDY_RA: # %esi, %edi, %r8d, %r9d, %eax
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_tied_output_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_FAST_RA: #NO_APP
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm tied output: pressure
+; FAST_ISEL_FAST_RA: # %eax, %edi, %ecx, %edx, %esi
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_tied_output_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; GREEDY_RA: #NO_APP
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm tied output: pressure
+; GREEDY_RA: # %esi, %edi, %r8d, %r9d, %eax
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_tied_output_pressure:
+; FAST_RA: #APP
+; FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_RA: #NO_APP
+; FAST_RA: #APP
+; FAST_RA: # rm tied output: pressure
+; FAST_RA: # %eax, %edi, %ecx, %edx, %esi
+; FAST_RA: #NO_APP
+entry:
+ %0 = tail call { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } asm sideeffect "# $0,$1,$2,$3,$4,$5,$6,$7,$8,$9,$10, $11, $12, $13, $14", "={rax},={rcx},={rdx},={rsi},={rdi},={rbx},={rbp},={r8},={r9},={r10},={r11},={r12},={r13},={r14},={r15},~{dirflag},~{fpsr},~{flags}"()
+ %1 = load i32, ptr %foo, align 4
+ %b16 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %2 = load i32, ptr %b16, align 4
+ %c17 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %3 = load i32, ptr %c17, align 4
+ %d18 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %4 = load i32, ptr %d18, align 4
+ %e19 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %5 = load i32, ptr %e19, align 4
+ %6 = tail call { i32, i32, i32, i32, i32 } asm sideeffect "# rm tied output: pressure\0A\09# $0, $1, $2, $3, $4", "=rm,=rm,=rm,=rm,=rm,0,1,2,3,4,~{dirflag},~{fpsr},~{flags}"(i32 %1, i32 %2, i32 %3, i32 %4, i32 %5)
+ %asmresult20 = extractvalue { i32, i32, i32, i32, i32 } %6, 0
+ %asmresult21 = extractvalue { i32, i32, i32, i32, i32 } %6, 1
+ %asmresult22 = extractvalue { i32, i32, i32, i32, i32 } %6, 2
+ %asmresult23 = extractvalue { i32, i32, i32, i32, i32 } %6, 3
+ %asmresult24 = extractvalue { i32, i32, i32, i32, i32 } %6, 4
+ store i32 %asmresult20, ptr %foo, align 4
+ store i32 %asmresult21, ptr %b16, align 4
+ store i32 %asmresult22, ptr %c17, align 4
+ store i32 %asmresult23, ptr %d18, align 4
+ store i32 %asmresult24, ptr %e19, align 4
+ %asmresult14 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 14
+ %asmresult13 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 13
+ %asmresult12 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 12
+ %asmresult11 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 11
+ %asmresult10 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 10
+ %asmresult9 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 9
+ %asmresult8 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 8
+ %asmresult7 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 7
+ %asmresult6 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 6
+ %asmresult5 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 5
+ %asmresult4 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 4
+ %asmresult3 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 3
+ %asmresult2 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 2
+ %asmresult1 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 1
+ %asmresult = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 0
+ tail call void @g(i64 noundef %asmresult, i64 noundef %asmresult1, i64 noundef %asmresult2, i64 noundef %asmresult3, i64 noundef %asmresult4, i64 noundef %asmresult5, i64 noundef %asmresult6, i64 noundef %asmresult7, i64 noundef %asmresult8, i64 noundef %asmresult9, i64 noundef %asmresult10, i64 noundef %asmresult11, i64 noundef %asmresult12, i64 noundef %asmresult13, i64 noundef %asmresult14)
+ %7 = load i32, ptr %foo, align 4
+ ret i32 %7
+}
+
+define dso_local i32 @test_rm_output_r_input_no_pressure(ptr noundef captures(none) initializes((0, 4)) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_rm_output_r_input_no_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm output, r input: no pressure
+; FAST_ISEL_GREEDY_RA: # %eax, %eax
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_rm_output_r_input_no_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm output, r input: no pressure
+; FAST_ISEL_FAST_RA: # %eax, %eax
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_rm_output_r_input_no_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm output, r input: no pressure
+; GREEDY_RA: # %eax, %eax
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_rm_output_r_input_no_pressure:
+; FAST_RA: #APP
+; FAST_RA: # rm output, r input: no pressure
+; FAST_RA: # %eax, %eax
+; FAST_RA: #NO_APP
+entry:
+ %b2 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %0 = load i32, ptr %b2, align 4
+ %1 = tail call i32 asm sideeffect "# rm output, r input: no pressure\0A\09# $0, $1", "=rm,r,~{dirflag},~{fpsr},~{flags}"(i32 %0)
+ store i32 %1, ptr %foo, align 4
+ ret i32 %1
+}
+
+define dso_local i32 @test_rm_output_r_input_pressure(ptr noundef captures(none) initializes((0, 4)) %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_rm_output_r_input_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_GREEDY_RA: #NO_APP
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # rm output, r input: pressure
+; FAST_ISEL_GREEDY_RA: # %esi, %esi
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_rm_output_r_input_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_FAST_RA: #NO_APP
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # rm output, r input: pressure
+; FAST_ISEL_FAST_RA: # %eax, %eax
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_rm_output_r_input_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; GREEDY_RA: #NO_APP
+; GREEDY_RA: #APP
+; GREEDY_RA: # rm output, r input: pressure
+; GREEDY_RA: # %esi, %esi
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_rm_output_r_input_pressure:
+; FAST_RA: #APP
+; FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_RA: #NO_APP
+; FAST_RA: #APP
+; FAST_RA: # rm output, r input: pressure
+; FAST_RA: # %eax, %eax
+; FAST_RA: #NO_APP
+entry:
+ %0 = tail call { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } asm sideeffect "# $0,$1,$2,$3,$4,$5,$6,$7,$8,$9,$10, $11, $12, $13, $14", "={rax},={rcx},={rdx},={rsi},={rdi},={rbx},={rbp},={r8},={r9},={r10},={r11},={r12},={r13},={r14},={r15},~{dirflag},~{fpsr},~{flags}"()
+ %b16 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %1 = load i32, ptr %b16, align 4
+ %2 = tail call i32 asm sideeffect "# rm output, r input: pressure\0A\09# $0, $1", "=rm,r,~{dirflag},~{fpsr},~{flags}"(i32 %1)
+ store i32 %2, ptr %foo, align 4
+ %asmresult14 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 14
+ %asmresult13 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 13
+ %asmresult12 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 12
+ %asmresult11 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 11
+ %asmresult10 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 10
+ %asmresult9 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 9
+ %asmresult8 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 8
+ %asmresult7 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 7
+ %asmresult6 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 6
+ %asmresult5 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 5
+ %asmresult4 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 4
+ %asmresult3 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 3
+ %asmresult2 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 2
+ %asmresult1 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 1
+ %asmresult = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 0
+ tail call void @g(i64 noundef %asmresult, i64 noundef %asmresult1, i64 noundef %asmresult2, i64 noundef %asmresult3, i64 noundef %asmresult4, i64 noundef %asmresult5, i64 noundef %asmresult6, i64 noundef %asmresult7, i64 noundef %asmresult8, i64 noundef %asmresult9, i64 noundef %asmresult10, i64 noundef %asmresult11, i64 noundef %asmresult12, i64 noundef %asmresult13, i64 noundef %asmresult14)
+ %3 = load i32, ptr %foo, align 4
+ ret i32 %3
+}
+
+define dso_local i32 @test_m_output_rm_input_no_pressure(ptr noundef %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_m_output_rm_input_no_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # m output, rm input: no pressure
+; FAST_ISEL_GREEDY_RA: # (%rdi), %eax
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_m_output_rm_input_no_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # m output, rm input: no pressure
+; FAST_ISEL_FAST_RA: # (%rdi), %eax
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_m_output_rm_input_no_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # m output, rm input: no pressure
+; GREEDY_RA: # (%rdi), %eax
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_m_output_rm_input_no_pressure:
+; FAST_RA: #APP
+; FAST_RA: # m output, rm input: no pressure
+; FAST_RA: # (%rdi), %eax
+; FAST_RA: #NO_APP
+entry:
+ %b2 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %0 = load i32, ptr %b2, align 4
+ tail call void asm sideeffect "# m output, rm input: no pressure\0A\09# $0, $1", "=*m,rm,~{dirflag},~{fpsr},~{flags}"(ptr elementtype(i32) %foo, i32 %0)
+ %1 = load i32, ptr %foo, align 4
+ ret i32 %1
+}
+
+define dso_local i32 @test_m_output_rm_input_pressure(ptr noundef %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_m_output_rm_input_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_GREEDY_RA: #NO_APP
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # m output, rm input: pressure
+; FAST_ISEL_GREEDY_RA: # (%rbp), %esi
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_m_output_rm_input_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_FAST_RA: #NO_APP
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # m output, rm input: pressure
+; FAST_ISEL_FAST_RA: # (%rdi), %eax
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_m_output_rm_input_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; GREEDY_RA: #NO_APP
+; GREEDY_RA: #APP
+; GREEDY_RA: # m output, rm input: pressure
+; GREEDY_RA: # (%rbp), %esi
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_m_output_rm_input_pressure:
+; FAST_RA: #APP
+; FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_RA: #NO_APP
+; FAST_RA: #APP
+; FAST_RA: # m output, rm input: pressure
+; FAST_RA: # (%rdi), %eax
+; FAST_RA: #NO_APP
+entry:
+ %0 = tail call { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } asm sideeffect "# $0,$1,$2,$3,$4,$5,$6,$7,$8,$9,$10, $11, $12, $13, $14", "={rax},={rcx},={rdx},={rsi},={rdi},={rbx},={rbp},={r8},={r9},={r10},={r11},={r12},={r13},={r14},={r15},~{dirflag},~{fpsr},~{flags}"()
+ %b16 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %1 = load i32, ptr %b16, align 4
+ tail call void asm sideeffect "# m output, rm input: pressure\0A\09# $0, $1", "=*m,rm,~{dirflag},~{fpsr},~{flags}"(ptr elementtype(i32) %foo, i32 %1)
+ %asmresult14 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 14
+ %asmresult13 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 13
+ %asmresult12 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 12
+ %asmresult11 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 11
+ %asmresult10 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 10
+ %asmresult9 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 9
+ %asmresult8 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 8
+ %asmresult7 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 7
+ %asmresult6 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 6
+ %asmresult5 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 5
+ %asmresult4 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 4
+ %asmresult3 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 3
+ %asmresult2 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 2
+ %asmresult1 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 1
+ %asmresult = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 0
+ tail call void @g(i64 noundef %asmresult, i64 noundef %asmresult1, i64 noundef %asmresult2, i64 noundef %asmresult3, i64 noundef %asmresult4, i64 noundef %asmresult5, i64 noundef %asmresult6, i64 noundef %asmresult7, i64 noundef %asmresult8, i64 noundef %asmresult9, i64 noundef %asmresult10, i64 noundef %asmresult11, i64 noundef %asmresult12, i64 noundef %asmresult13, i64 noundef %asmresult14)
+ %2 = load i32, ptr %foo, align 4
+ ret i32 %2
+}
+
+define dso_local i32 @test_mult_m_output_rm_input_no_pressure(ptr noundef %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_mult_m_output_rm_input_no_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # multiple m output, rm input: no pressure
+; FAST_ISEL_GREEDY_RA: # (%rdi), (%rax), (%rcx), (%rdx), (%rsi), %r8d, %r9d
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_mult_m_output_rm_input_no_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # multiple m output, rm input: no pressure
+; FAST_ISEL_FAST_RA: # (%rdi), (%rax), (%rcx), (%rdx), (%rsi), %r8d, %r9d
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_mult_m_output_rm_input_no_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # multiple m output, rm input: no pressure
+; GREEDY_RA: # (%rdi), 4(%rdi), 8(%rdi), 12(%rdi), 16(%rdi), %eax, %ecx
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_mult_m_output_rm_input_no_pressure:
+; FAST_RA: #APP
+; FAST_RA: # multiple m output, rm input: no pressure
+; FAST_RA: # (%rdi), 4(%rdi), 8(%rdi), 12(%rdi), 16(%rdi), %eax, %ecx
+; FAST_RA: #NO_APP
+entry:
+ %b4 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %c5 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %d6 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %e7 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %0 = load i32, ptr %foo, align 4
+ %1 = load i32, ptr %b4, align 4
+ tail call void asm sideeffect "# multiple m output, rm input: no pressure\0A\09# $0, $1, $2, $3, $4, $5, $6", "=*m,=*m,=*m,=*m,=*m,rm,rm,~{dirflag},~{fpsr},~{flags}"(ptr nonnull elementtype(i32) %foo, ptr nonnull elementtype(i32) %b4, ptr nonnull elementtype(i32) %c5, ptr nonnull elementtype(i32) %d6, ptr nonnull elementtype(i32) %e7, i32 %0, i32 %1)
+ %2 = load i32, ptr %foo, align 4
+ ret i32 %2
+}
+
+define dso_local i32 @test_mult_m_output_rm_input_pressure(ptr noundef %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_mult_m_output_rm_input_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_GREEDY_RA: #NO_APP
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # multiple m output, rm input: pressure
+; FAST_ISEL_GREEDY_RA: # (%rbp), (%r8), (%r9), (%rbx), (%rsi), %eax, %edi
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_mult_m_output_rm_input_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_FAST_RA: #NO_APP
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # multiple m output, rm input: pressure
+; FAST_ISEL_FAST_RA: # (%rdi), (%rax), (%rcx), (%rdx), (%rsi), %r8d, %r9d
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_mult_m_output_rm_input_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; GREEDY_RA: #NO_APP
+; GREEDY_RA: #APP
+; GREEDY_RA: # multiple m output, rm input: pressure
+; GREEDY_RA: # (%rbp), 4(%rbp), 8(%rbp), 12(%rbp), 16(%rbp), %esi, %edi
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_mult_m_output_rm_input_pressure:
+; FAST_RA: #APP
+; FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_RA: #NO_APP
+; FAST_RA: #APP
+; FAST_RA: # multiple m output, rm input: pressure
+; FAST_RA: # (%rdi), 4(%rdi), 8(%rdi), 12(%rdi), 16(%rdi), %eax, %ecx
+; FAST_RA: #NO_APP
+entry:
+ %0 = tail call { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } asm sideeffect "# $0,$1,$2,$3,$4,$5,$6,$7,$8,$9,$10, $11, $12, $13, $14", "={rax},={rcx},={rdx},={rsi},={rdi},={rbx},={rbp},={r8},={r9},={r10},={r11},={r12},={r13},={r14},={r15},~{dirflag},~{fpsr},~{flags}"()
+ %b18 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %c19 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %d20 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %e21 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %1 = load i32, ptr %foo, align 4
+ %2 = load i32, ptr %b18, align 4
+ tail call void asm sideeffect "# multiple m output, rm input: pressure\0A\09# $0, $1, $2, $3, $4, $5, $6", "=*m,=*m,=*m,=*m,=*m,rm,rm,~{dirflag},~{fpsr},~{flags}"(ptr nonnull elementtype(i32) %foo, ptr nonnull elementtype(i32) %b18, ptr nonnull elementtype(i32) %c19, ptr nonnull elementtype(i32) %d20, ptr nonnull elementtype(i32) %e21, i32 %1, i32 %2)
+ %asmresult14 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 14
+ %asmresult13 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 13
+ %asmresult12 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 12
+ %asmresult11 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 11
+ %asmresult10 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 10
+ %asmresult9 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 9
+ %asmresult8 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 8
+ %asmresult7 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 7
+ %asmresult6 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 6
+ %asmresult5 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 5
+ %asmresult4 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 4
+ %asmresult3 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 3
+ %asmresult2 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 2
+ %asmresult1 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 1
+ %asmresult = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 0
+ tail call void @g(i64 noundef %asmresult, i64 noundef %asmresult1, i64 noundef %asmresult2, i64 noundef %asmresult3, i64 noundef %asmresult4, i64 noundef %asmresult5, i64 noundef %asmresult6, i64 noundef %asmresult7, i64 noundef %asmresult8, i64 noundef %asmresult9, i64 noundef %asmresult10, i64 noundef %asmresult11, i64 noundef %asmresult12, i64 noundef %asmresult13, i64 noundef %asmresult14)
+ %3 = load i32, ptr %foo, align 4
+ ret i32 %3
+}
+
+define dso_local i32 @test_mult_m_early_clobber_output_rm_input_no_pressure(ptr noundef %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_mult_m_early_clobber_output_rm_input_no_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # multiple m output, rm input: no pressure
+; FAST_ISEL_GREEDY_RA: # %eax, %esi, %r8d, %ecx, %edx
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_mult_m_early_clobber_output_rm_input_no_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # multiple m output, rm input: no pressure
+; FAST_ISEL_FAST_RA: # %eax, %edx, %ecx, %esi, %r8d
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_mult_m_early_clobber_output_rm_input_no_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # multiple m output, rm input: no pressure
+; GREEDY_RA: # %eax, %esi, %r8d, %ecx, %edx
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_mult_m_early_clobber_output_rm_input_no_pressure:
+; FAST_RA: #APP
+; FAST_RA: # multiple m output, rm input: no pressure
+; FAST_RA: # %eax, %edx, %ecx, %esi, %r8d
+; FAST_RA: #NO_APP
+entry:
+ %b2 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %c3 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %d4 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %0 = load i32, ptr %d4, align 4
+ %e5 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %1 = load i32, ptr %e5, align 4
+ %2 = tail call { i32, i32, i32 } asm sideeffect "# multiple m output, rm input: no pressure\0A\09# $0, $1, $2, $3, $4", "=&rm,=&rm,=&rm,rm,rm,~{dirflag},~{fpsr},~{flags}"(i32 %0, i32 %1)
+ %asmresult = extractvalue { i32, i32, i32 } %2, 0
+ %asmresult6 = extractvalue { i32, i32, i32 } %2, 1
+ %asmresult7 = extractvalue { i32, i32, i32 } %2, 2
+ store i32 %asmresult, ptr %foo, align 4
+ store i32 %asmresult6, ptr %b2, align 4
+ store i32 %asmresult7, ptr %c3, align 4
+ ret i32 %asmresult
+}
+
+define dso_local i32 @test_mult_m_early_clobber_output_rm_input_pressure(ptr noundef %foo) local_unnamed_addr {
+; FAST_ISEL_GREEDY_RA-LABEL: test_mult_m_early_clobber_output_rm_input_pressure:
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_GREEDY_RA: #NO_APP
+; FAST_ISEL_GREEDY_RA: #APP
+; FAST_ISEL_GREEDY_RA: # multiple m output, rm input: pressure
+; FAST_ISEL_GREEDY_RA: # %r8d, %r9d, %eax, %esi, %edi
+; FAST_ISEL_GREEDY_RA: #NO_APP
+;
+; FAST_ISEL_FAST_RA-LABEL: test_mult_m_early_clobber_output_rm_input_pressure:
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_ISEL_FAST_RA: #NO_APP
+; FAST_ISEL_FAST_RA: #APP
+; FAST_ISEL_FAST_RA: # multiple m output, rm input: pressure
+; FAST_ISEL_FAST_RA: # %eax, %edi, %ecx, %edx, %esi
+; FAST_ISEL_FAST_RA: #NO_APP
+;
+; GREEDY_RA-LABEL: test_mult_m_early_clobber_output_rm_input_pressure:
+; GREEDY_RA: #APP
+; GREEDY_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; GREEDY_RA: #NO_APP
+; GREEDY_RA: #APP
+; GREEDY_RA: # multiple m output, rm input: pressure
+; GREEDY_RA: # %r8d, %r9d, %eax, %esi, %edi
+; GREEDY_RA: #NO_APP
+;
+; FAST_RA-LABEL: test_mult_m_early_clobber_output_rm_input_pressure:
+; FAST_RA: #APP
+; FAST_RA: # %rax,%rcx,%rdx,%rsi,%rdi,%rbx,%rbp,%r8,%r9,%r10,%r11, %r12, %r13, %r14, %r15
+; FAST_RA: #NO_APP
+; FAST_RA: #APP
+; FAST_RA: # multiple m output, rm input: pressure
+; FAST_RA: # %eax, %edi, %ecx, %edx, %esi
+; FAST_RA: #NO_APP
+entry:
+ %0 = tail call { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } asm sideeffect "# $0,$1,$2,$3,$4,$5,$6,$7,$8,$9,$10, $11, $12, $13, $14", "={rax},={rcx},={rdx},={rsi},={rdi},={rbx},={rbp},={r8},={r9},={r10},={r11},={r12},={r13},={r14},={r15},~{dirflag},~{fpsr},~{flags}"()
+ %b16 = getelementptr inbounds nuw i8, ptr %foo, i64 4
+ %c17 = getelementptr inbounds nuw i8, ptr %foo, i64 8
+ %d18 = getelementptr inbounds nuw i8, ptr %foo, i64 12
+ %1 = load i32, ptr %d18, align 4
+ %e19 = getelementptr inbounds nuw i8, ptr %foo, i64 16
+ %2 = load i32, ptr %e19, align 4
+ %3 = tail call { i32, i32, i32 } asm sideeffect "# multiple m output, rm input: pressure\0A\09# $0, $1, $2, $3, $4", "=&rm,=&rm,=&rm,rm,rm,~{dirflag},~{fpsr},~{flags}"(i32 %1, i32 %2)
+ %asmresult20 = extractvalue { i32, i32, i32 } %3, 0
+ %asmresult21 = extractvalue { i32, i32, i32 } %3, 1
+ %asmresult22 = extractvalue { i32, i32, i32 } %3, 2
+ store i32 %asmresult20, ptr %foo, align 4
+ store i32 %asmresult21, ptr %b16, align 4
+ store i32 %asmresult22, ptr %c17, align 4
+ %asmresult14 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 14
+ %asmresult13 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 13
+ %asmresult12 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 12
+ %asmresult11 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 11
+ %asmresult10 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 10
+ %asmresult9 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 9
+ %asmresult8 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 8
+ %asmresult7 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 7
+ %asmresult6 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 6
+ %asmresult5 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 5
+ %asmresult4 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 4
+ %asmresult3 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 3
+ %asmresult2 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 2
+ %asmresult1 = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 1
+ %asmresult = extractvalue { i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64, i64 } %0, 0
+ tail call void @g(i64 noundef %asmresult, i64 noundef %asmresult1, i64 noundef %asmresult2, i64 noundef %asmresult3, i64 noundef %asmresult4, i64 noundef %asmresult5, i64 noundef %asmresult6, i64 noundef %asmresult7, i64 noundef %asmresult8, i64 noundef %asmresult9, i64 noundef %asmresult10, i64 noundef %asmresult11, i64 noundef %asmresult12, i64 noundef %asmresult13, i64 noundef %asmresult14)
+ %4 = load i32, ptr %foo, align 4
+ ret i32 %4
+}
+
+declare void @g(i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef, i64 noundef)
diff --git a/llvm/test/CodeGen/X86/inline-asm-callbase.ll b/llvm/test/CodeGen/X86/inline-asm-callbase.ll
new file mode 100644
index 0000000000000..d9dc0211ec837
--- /dev/null
+++ b/llvm/test/CodeGen/X86/inline-asm-callbase.ll
@@ -0,0 +1,227 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 6
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=true --regalloc=greedy < %s \
+; RUN: | FileCheck --check-prefix=FAST_ISEL_GREEDY_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=true --regalloc=fast < %s \
+; RUN: | FileCheck --check-prefix=FAST_ISEL_FAST_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=false --regalloc=greedy < %s \
+; RUN: | FileCheck --check-prefix=GREEDY_RA %s
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu --global-isel=false --fast-isel=false --regalloc=fast < %s \
+; RUN: | FileCheck --check-prefix=FAST_RA %s
+
+; "rm" constraints on CallBase subclasses other than a plain call: invoke
+; (EH unwind edges) and callbr (indirect-branch asm goto).
+
+declare i32 @__gxx_personality_v0(...)
+
+define i32 @test_invoke_rm(i32 %x) personality ptr @__gxx_personality_v0 {
+; FAST_ISEL_GREEDY_RA-LABEL: test_invoke_rm:
+; FAST_ISEL_GREEDY_RA: # %bb.0: # %entry
+; FAST_ISEL_GREEDY_RA-NEXT: .Ltmp0: # EH_LABEL
+; FAST_ISEL_GREEDY_RA-NEXT: #APP
+; FAST_ISEL_GREEDY_RA-NEXT: # %eax, %edi
+; FAST_ISEL_GREEDY_RA-NEXT: #NO_APP
+; FAST_ISEL_GREEDY_RA-NEXT: .Ltmp1: # EH_LABEL
+; FAST_ISEL_GREEDY_RA-NEXT: # %bb.1: # %normal
+; FAST_ISEL_GREEDY_RA-NEXT: retq
+; FAST_ISEL_GREEDY_RA-NEXT: .LBB0_2: # %unwind
+; FAST_ISEL_GREEDY_RA-NEXT: pushq %rax
+; FAST_ISEL_GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_ISEL_GREEDY_RA-NEXT: .Ltmp2: # EH_LABEL
+; FAST_ISEL_GREEDY_RA-NEXT: movq %rax, %rdi
+; FAST_ISEL_GREEDY_RA-NEXT: callq _Unwind_Resume at PLT
+;
+; FAST_ISEL_FAST_RA-LABEL: test_invoke_rm:
+; FAST_ISEL_FAST_RA: # %bb.0: # %entry
+; FAST_ISEL_FAST_RA-NEXT: pushq %rax
+; FAST_ISEL_FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_ISEL_FAST_RA-NEXT: .Ltmp0: # EH_LABEL
+; FAST_ISEL_FAST_RA-NEXT: #APP
+; FAST_ISEL_FAST_RA-NEXT: # %eax, %edi
+; FAST_ISEL_FAST_RA-NEXT: #NO_APP
+; FAST_ISEL_FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_ISEL_FAST_RA-NEXT: .Ltmp1: # EH_LABEL
+; FAST_ISEL_FAST_RA-NEXT: # %bb.1: # %normal
+; FAST_ISEL_FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_ISEL_FAST_RA-NEXT: popq %rcx
+; FAST_ISEL_FAST_RA-NEXT: .cfi_def_cfa_offset 8
+; FAST_ISEL_FAST_RA-NEXT: retq
+; FAST_ISEL_FAST_RA-NEXT: .LBB0_2: # %unwind
+; FAST_ISEL_FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_ISEL_FAST_RA-NEXT: .Ltmp2: # EH_LABEL
+; FAST_ISEL_FAST_RA-NEXT: movq %rax, %rdi
+; FAST_ISEL_FAST_RA-NEXT: callq _Unwind_Resume at PLT
+;
+; GREEDY_RA-LABEL: test_invoke_rm:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: .Ltmp0: # EH_LABEL
+; GREEDY_RA-NEXT: #APP
+; GREEDY_RA-NEXT: # %eax, %edi
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: .Ltmp1: # EH_LABEL
+; GREEDY_RA-NEXT: # %bb.1: # %normal
+; GREEDY_RA-NEXT: retq
+; GREEDY_RA-NEXT: .LBB0_2: # %unwind
+; GREEDY_RA-NEXT: pushq %rax
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: .Ltmp2: # EH_LABEL
+; GREEDY_RA-NEXT: movq %rax, %rdi
+; GREEDY_RA-NEXT: callq _Unwind_Resume at PLT
+;
+; FAST_RA-LABEL: test_invoke_rm:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: pushq %rax
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: .Ltmp0: # EH_LABEL
+; FAST_RA-NEXT: #APP
+; FAST_RA-NEXT: # %eax, %edi
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_RA-NEXT: .Ltmp1: # EH_LABEL
+; FAST_RA-NEXT: # %bb.1: # %normal
+; FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_RA-NEXT: popq %rcx
+; FAST_RA-NEXT: .cfi_def_cfa_offset 8
+; FAST_RA-NEXT: retq
+; FAST_RA-NEXT: .LBB0_2: # %unwind
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: .Ltmp2: # EH_LABEL
+; FAST_RA-NEXT: movq %rax, %rdi
+; FAST_RA-NEXT: callq _Unwind_Resume at PLT
+entry:
+ %0 = invoke i32 asm "# $0, $1", "=r,rm"(i32 %x)
+ to label %normal unwind label %unwind
+
+normal:
+ ret i32 %0
+
+unwind:
+ %1 = landingpad { ptr, i32 }
+ cleanup
+ resume { ptr, i32 } %1
+}
+
+define i32 @test_callbr_rm(i32 %x) {
+; FAST_ISEL_GREEDY_RA-LABEL: test_callbr_rm:
+; FAST_ISEL_GREEDY_RA: # %bb.0: # %entry
+; FAST_ISEL_GREEDY_RA-NEXT: #APP
+; FAST_ISEL_GREEDY_RA-NEXT: # %eax, %edi
+; FAST_ISEL_GREEDY_RA-NEXT: #NO_APP
+; FAST_ISEL_GREEDY_RA-NEXT: .LBB1_1: # Inline asm indirect target
+; FAST_ISEL_GREEDY_RA-NEXT: # %indirect
+; FAST_ISEL_GREEDY_RA-NEXT: # Label of block must be emitted
+; FAST_ISEL_GREEDY_RA-NEXT: retq
+;
+; FAST_ISEL_FAST_RA-LABEL: test_callbr_rm:
+; FAST_ISEL_FAST_RA: # %bb.0: # %entry
+; FAST_ISEL_FAST_RA-NEXT: #APP
+; FAST_ISEL_FAST_RA-NEXT: # %eax, %edi
+; FAST_ISEL_FAST_RA-NEXT: #NO_APP
+; FAST_ISEL_FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_ISEL_FAST_RA-NEXT: # %bb.1: # %normal
+; FAST_ISEL_FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_ISEL_FAST_RA-NEXT: retq
+; FAST_ISEL_FAST_RA-NEXT: .LBB1_2: # Inline asm indirect target
+; FAST_ISEL_FAST_RA-NEXT: # %indirect
+; FAST_ISEL_FAST_RA-NEXT: # Label of block must be emitted
+; FAST_ISEL_FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_ISEL_FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_ISEL_FAST_RA-NEXT: retq
+;
+; GREEDY_RA-LABEL: test_callbr_rm:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: #APP
+; GREEDY_RA-NEXT: # %eax, %edi
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: .LBB1_1: # Inline asm indirect target
+; GREEDY_RA-NEXT: # %indirect
+; GREEDY_RA-NEXT: # Label of block must be emitted
+; GREEDY_RA-NEXT: retq
+;
+; FAST_RA-LABEL: test_callbr_rm:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: #APP
+; FAST_RA-NEXT: # %eax, %edi
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_RA-NEXT: # %bb.1: # %normal
+; FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_RA-NEXT: retq
+; FAST_RA-NEXT: .LBB1_2: # Inline asm indirect target
+; FAST_RA-NEXT: # %indirect
+; FAST_RA-NEXT: # Label of block must be emitted
+; FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_RA-NEXT: retq
+entry:
+ %0 = callbr i32 asm "# $0, $1", "=r,rm,!i"(i32 %x)
+ to label %normal [label %indirect]
+
+normal:
+ ret i32 %0
+
+indirect:
+ ret i32 %0
+}
+
+define i32 @test_callbr_convert(i32 %x) {
+; FAST_ISEL_GREEDY_RA-LABEL: test_callbr_convert:
+; FAST_ISEL_GREEDY_RA: # %bb.0: # %entry
+; FAST_ISEL_GREEDY_RA-NEXT: #APP
+; FAST_ISEL_GREEDY_RA-NEXT: # %eax, %edi
+; FAST_ISEL_GREEDY_RA-NEXT: #NO_APP
+; FAST_ISEL_GREEDY_RA-NEXT: .LBB2_1: # Inline asm indirect target
+; FAST_ISEL_GREEDY_RA-NEXT: # %indirect
+; FAST_ISEL_GREEDY_RA-NEXT: # Label of block must be emitted
+; FAST_ISEL_GREEDY_RA-NEXT: retq
+;
+; FAST_ISEL_FAST_RA-LABEL: test_callbr_convert:
+; FAST_ISEL_FAST_RA: # %bb.0: # %entry
+; FAST_ISEL_FAST_RA-NEXT: #APP
+; FAST_ISEL_FAST_RA-NEXT: # %eax, %edi
+; FAST_ISEL_FAST_RA-NEXT: #NO_APP
+; FAST_ISEL_FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_ISEL_FAST_RA-NEXT: # %bb.1: # %normal
+; FAST_ISEL_FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_ISEL_FAST_RA-NEXT: retq
+; FAST_ISEL_FAST_RA-NEXT: .LBB2_2: # Inline asm indirect target
+; FAST_ISEL_FAST_RA-NEXT: # %indirect
+; FAST_ISEL_FAST_RA-NEXT: # Label of block must be emitted
+; FAST_ISEL_FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_ISEL_FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_ISEL_FAST_RA-NEXT: retq
+;
+; GREEDY_RA-LABEL: test_callbr_convert:
+; GREEDY_RA: # %bb.0: # %entry
+; GREEDY_RA-NEXT: #APP
+; GREEDY_RA-NEXT: # %eax, %edi
+; GREEDY_RA-NEXT: #NO_APP
+; GREEDY_RA-NEXT: .LBB2_1: # Inline asm indirect target
+; GREEDY_RA-NEXT: # %indirect
+; GREEDY_RA-NEXT: # Label of block must be emitted
+; GREEDY_RA-NEXT: retq
+;
+; FAST_RA-LABEL: test_callbr_convert:
+; FAST_RA: # %bb.0: # %entry
+; FAST_RA-NEXT: #APP
+; FAST_RA-NEXT: # %eax, %edi
+; FAST_RA-NEXT: #NO_APP
+; FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_RA-NEXT: # %bb.1: # %normal
+; FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_RA-NEXT: retq
+; FAST_RA-NEXT: .LBB2_2: # Inline asm indirect target
+; FAST_RA-NEXT: # %indirect
+; FAST_RA-NEXT: # Label of block must be emitted
+; FAST_RA-NEXT: movl %eax, {{[-0-9]+}}(%r{{[sb]}}p) # 4-byte Spill
+; FAST_RA-NEXT: movl {{[-0-9]+}}(%r{{[sb]}}p), %eax # 4-byte Reload
+; FAST_RA-NEXT: retq
+entry:
+ %0 = callbr i32 asm "# $0, $1", "=rm,rm,!i"(i32 %x)
+ to label %normal [label %indirect]
+
+normal:
+ ret i32 %0
+
+indirect:
+ ret i32 %0
+}
>From b90e42e5d333e5597d21258bdb0fee34a90b6c76 Mon Sep 17 00:00:00 2001
From: Bill Wendling <isanbard at gmail.com>
Date: Tue, 4 Aug 2026 02:21:09 -0700
Subject: [PATCH 06/12] [Clang][SelectionDAG] Fix "rm" tied-output codegen gap;
fail safe on a pre-existing crash
Two related fixes, landed together because the first widens exposure
to the second.
1. clang/lib/CodeGen/CGStmt.cpp: CGStmt.cpp's output-operand lowering
only takes the direct (by-value, non-indirect) codegen path when a
constraint disallows memory entirely (!Info.allowsMemory()). Any
constraint that permits memory -- including a plain "rm"/"mr", which
is exactly the case the backend now prefers a register for and can
fall back to memory on its own (TargetLowering::MayFoldRegister,
RegAllocFast's inline asm folding) -- fell through to the historical
indirect (alloca + out-parameter) path unconditionally.
For a non-tied "=rm" output this mostly self-corrected: SelectionDAG
has an existing optimization that eliminates the indirection when
the only use is an immediate reload and nothing escapes, recovering
the intended register-preferred behavior anyway. But a tied "+rm"
(read-write) output never got that benefit -- it kept an explicit
store-before/load-after around the asm even with zero register
pressure, verified by direct inspection of real Clang -O2 output.
Fix: take the direct path for a constraint whose alternatives are
exactly {r, m} (mirroring MayFoldRegister's own check -- deliberately
not attempting to replicate the backend's full constraint-code
parser here; anything with an extra modifier this doesn't recognize,
e.g. a commutative '%', simply falls through to the existing,
always-safe indirect path), gated on optimization level > 0 to match
the backend's own opt-level guard. A pointer-escape case (the
output's address later passed to another function) was verified
still correct: the asm's own write still prefers a register, but a
real store to memory is still inserted before the escaping use --
pre-existing, general "keep in a register until something forces it
to memory" logic, unrelated to and unaffected by this change.
2. llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp: this widens
how often a direct (non-indirect) output constraint reaches
computeConstraintToUse() with a memory constraint chosen, which
already had a pre-existing, unrelated latent crash for exactly that
shape: an assertion ("Can only indirectify direct input operands!")
that fires whenever indirectifying is needed for anything but an
input. This was never reachable through prior Clang output (which
always used the indirect form whenever memory was a possible
outcome, by the very logic fix 1 just changed), but is directly
reachable through hand-written IR today -- reproducible on
completely unmodified upstream main, with no register pressure
required, merely by writing a direct-form "=rm"-shaped output and
compiling at -O0 (where the default constraint-priority order
already prefers memory with nothing overriding it).
computeConstraintToUse() now returns bool and reports a clean
"unsupported inline asm" diagnostic through the existing
Info.ErrorMsg/emitInlineAsmError() plumbing instead of asserting.
Implementing actual output-indirectification (synthesizing a stack
slot, threading its address through as an out-parameter, reloading
the result afterward) was considered and rejected as out of
proportion to the benefit: on this branch that shape is only
reachable through hand-written or other-frontend IR, never through
Clang, so failing cleanly is the correct, minimal fix -- matching
RegAllocFast's own existing philosophy of erroring out rather than
silently miscompiling when a constraint genuinely can't be
satisfied.
Verified: the tied-output no-pressure case now compiles with zero
memory traffic; the pressure case still folds correctly (and, as a
side effect of fix 1, now uses one shared stack slot instead of two);
the escape case remains correct; the former crash now produces a
clean diagnostic instead of aborting the compiler. Full CodeGen/ +
clang/test/CodeGen/ + Sema/ + SemaCXX/ + MIR/ + MachineVerifier/
(55000+ tests across this and the prior verification pass) passes,
including every existing rm-specific test added in prior commits,
unchanged.
Co-Authored-By: Claude Sonnet 5 <noreply at anthropic.com>
Claude-Session: https://claude.ai/code/session_0143ozRFwhFjWmJjtfZUgmYe
---
clang/lib/CodeGen/CGStmt.cpp | 26 ++++++++++++-
.../SelectionDAG/SelectionDAGBuilder.cpp | 39 ++++++++++++++++---
2 files changed, 58 insertions(+), 7 deletions(-)
diff --git a/clang/lib/CodeGen/CGStmt.cpp b/clang/lib/CodeGen/CGStmt.cpp
index 27e74d966eca1..892da1cb99855 100644
--- a/clang/lib/CodeGen/CGStmt.cpp
+++ b/clang/lib/CodeGen/CGStmt.cpp
@@ -2963,7 +2963,31 @@ void AsmConstraintsInfo::HandleOutputConstraints() {
CodeGenFunction::hasScalarEvaluationKind(QTy) ||
CodeGenFunction::hasAggregateEvaluationKind(QTy);
- if (!Info.allowsMemory() && IsScalarOrAggregate) {
+ // A plain "rm"/"mr" constraint (register-or-memory, no other
+ // alternatives) also takes the by-value path above -O0, mirroring
+ // TargetLowering::MayFoldRegister on the backend: the backend now
+ // prefers a register for exactly this constraint shape and falls back
+ // to memory under pressure on its own (see RegAllocFast's inline asm
+ // folding), so there's no need to pessimistically commit to memory here
+ // the way a plain "m" (or a broader multi-alternative constraint like
+ // "g") still must. This check is deliberately conservative -- an exact
+ // 2-character match after stripping a leading early-clobber '&' --
+ // rather than trying to replicate the backend's full constraint-code
+ // parser here: anything it doesn't recognize (e.g. a commutative '%'
+ // modifier, which isn't tracked on TargetInfo::ConstraintInfo) simply
+ // falls through to the existing, always-safe indirect path below.
+ bool IsExactRegMem = false;
+ if (CGM.getCodeGenOpts().OptimizationLevel != 0 && Info.allowsRegister() &&
+ Info.allowsMemory()) {
+ StringRef Codes = OutputConstraint;
+ if (Info.earlyClobber() && Codes.starts_with("&"))
+ Codes = Codes.drop_front();
+ IsExactRegMem =
+ Codes.size() == 2 && ((Codes[0] == 'r' && Codes[1] == 'm') ||
+ (Codes[0] == 'm' && Codes[1] == 'r'));
+ }
+
+ if ((!Info.allowsMemory() || IsExactRegMem) && IsScalarOrAggregate) {
Constraints += "=" + OutputConstraint;
ResultRegQualTys.push_back(QTy);
ResultRegDests.push_back(Dest);
diff --git a/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp b/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp
index e51e02e81b628..2291ebe60bacc 100644
--- a/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp
@@ -10397,8 +10397,9 @@ constructOperandInfo(ConstraintDecisionInfo &Info,
return false;
}
-/// Compute which constraint option to use for each operand.
-static void
+/// Compute which constraint option to use for each operand. Returns true (and
+/// sets Info.ErrorMsg) on failure.
+static bool
computeConstraintToUse(ConstraintDecisionInfo &Info, const CallBase &Call,
TargetLowering::AsmOperandInfoVector &TargetConstraints,
SelectionDAGBuilder &Builder, const TargetLowering &TLI,
@@ -10460,9 +10461,31 @@ computeConstraintToUse(ConstraintDecisionInfo &Info, const CallBase &Call,
// need to provide an address for the memory input.
if (OpInfo.ConstraintType == TargetLowering::C_Memory &&
!OpInfo.isIndirect) {
- assert((OpInfo.isMultipleAlternative ||
- (OpInfo.Type == InlineAsm::isInput)) &&
- "Can only indirectify direct input operands!");
+ if (!OpInfo.isMultipleAlternative && OpInfo.Type != InlineAsm::isInput) {
+ // Indirectifying a direct operand this way -- taking the address of
+ // an existing value and switching the operand to reference it in
+ // memory -- only makes sense for an input: there's already a value
+ // to spill and reference by address. An output or clobber has no
+ // value yet (it's about to be produced), so there's nothing to
+ // spill; supporting that would mean synthesizing a stack slot,
+ // threading its address through as an out-parameter, and reloading
+ // the result afterward -- machinery this function doesn't have.
+ //
+ // In practice this is unreachable for Clang-generated IR: Clang only
+ // emits a direct (non-indirect) output for a constraint that
+ // doesn't allow memory at all (or, for an exact register-or-memory
+ // constraint like "rm", above -O0 -- see
+ // TargetLowering::MayFoldRegister and CGStmt.cpp's mirroring
+ // check), so a direct output constraint should never end up
+ // choosing C_Memory here. But hand-written or other-frontend IR can
+ // still construct this shape, so fail with a clean diagnostic
+ // rather than the assertion this used to be.
+ Info.ErrorMsg << "unsupported inline asm: constraint '"
+ << OpInfo.ConstraintCode
+ << "' cannot be satisfied in a register and has no "
+ "memory to fall back to";
+ return true;
+ }
// Memory operands really want the address of the value.
Info.Chain = getAddressForMemoryInput(Info.Chain, Builder.getCurSDLoc(),
@@ -10475,6 +10498,8 @@ computeConstraintToUse(ConstraintDecisionInfo &Info, const CallBase &Call,
OpInfo.isIndirect = true;
}
}
+
+ return false;
}
/// Prepare DAG-level operands. As part of this, assign virtual and physical
@@ -10770,7 +10795,9 @@ determineConstraints(ConstraintDecisionInfo &Info,
Info.Chain = Builder.lowerStartEH(Info.Chain, EHPadBB, Info.BeginLabel);
// Second pass: Compute which constraint option to use.
- computeConstraintToUse(Info, Call, TargetConstraints, Builder, TLI, TM, DAG);
+ if (computeConstraintToUse(Info, Call, TargetConstraints, Builder, TLI, TM,
+ DAG))
+ return true;
// AsmNodeOperands - The operands for the ISD::INLINEASM node.
Info.AsmNodeOperands.push_back(SDValue()); // reserve space for input chain
>From 26812160f46be14cdb88bc7c06f19daf3caf1867 Mon Sep 17 00:00:00 2001
From: Bill Wendling <isanbard at gmail.com>
Date: Tue, 4 Aug 2026 02:45:29 -0700
Subject: [PATCH 07/12] [TargetLowering][Clang][GlobalISel] Gate "rm" register
preference on target support
Critical regression fix, found while adding AArch64 test coverage for
the "rm" constraint work: preferring 'r' for an "rm" constraint is
only safe if the target's backend can actually fall back to memory
when a register genuinely isn't available. That fallback --
TargetInstrInfo::getFrameIndexOperands(), used to materialize a
folded inline asm operand's frame-index addressing -- has a default
implementation of llvm_unreachable(), and until this commit, X86 was
the only target that overrides it. Nothing prior to this fix checked
that before setting MayFoldRegister / emitting the direct-form
output, so on every other target:
- Hand-written or other-frontend-generated direct-form "rm"-shaped
IR under real register pressure crashed outright: reproduced on
AArch64 with both the SelectionDAG (both regalloc=fast and
regalloc=greedy -- this is backend-generic, not RegAllocFast-
specific) and GlobalISel frameworks, hitting
TargetInstrInfo.h:2398's unreachable().
- Far more seriously, real Clang-compiled code broke outright:
ordinary, previously-working inline asm like
`unsigned long f(void) { unsigned long out; asm("..." : "=rm"(out));
return out; }` stopped compiling on AArch64 (and, spot-checked,
ARM and RISCV) with "error: unsupported inline asm: constraint 'm'
cannot be satisfied in a register and has no memory to fall back
to" at any optimization level above -O0. This was fix (1)'s
CGStmt.cpp change (widening which "rm" outputs take the direct,
by-value path) combined with the backend, on these targets, no
longer preferring 'r' for such an operand -- Clang emitted direct
IR the backend then couldn't satisfy.
Fix: llvm::TargetLowering::supportsRegMemInlineAsmFolding() (default
false) gates MayFoldRegister in ParseConstraints(); X86TargetLowering
overrides it true, matching its actual getFrameIndexOperands()
support. clang::TargetCodeGenInfo::supportsRegMemInlineAsmFolding()
(default false) mirrors this on the Clang side, gating CGStmt.cpp's
direct-path decision so it never emits IR the backend can't handle;
X86's TargetCodeGenInfo (both the 32- and 64-bit ABI variants, which
share the one X86InstrInfo implementation) overrides it true.
Also hardened both crash sites -- one already covered for the
SelectionDAG path by the earlier "fail safe" fix in
computeConstraintToUse(), the other newly discovered here in
GlobalISel's InlineAsmLowering::lowerInlineAsm(), which
null-dereferenced OpInfo.CallOperandVal (never set for a genuinely
direct output, which has no argument to dereference) instead of
crashing on an assertion -- so that even a target that hasn't been
capability-gated correctly, or hand-written IR constructing this
shape directly, fails with a clean diagnostic rather than a crash.
This is defense in depth on top of the capability gate, not a
substitute for it: the gate is what keeps real compiled code working
at all; the diagnostic is what keeps the failure mode clean when
something still reaches this path.
CodeGen/AArch64/inline-asm-rm.ll (added in the prior test-port
commit) is rewritten: it previously demonstrated register-preferred
codegen for a direct-form "=rm" output on AArch64, which was only
ever possible because the capability gate didn't exist yet. It now
tests the realistic case -- the indirect form Clang actually emits
for this target, working correctly and identically across all three
ISel frameworks, the same as it always has. A new
inline-asm-rm-unsupported-direct.ll is the regression guard for the
direct-form crash specifically (using `not` + 2>&1, since the clean
diagnostic makes llc's exit code non-zero by design).
Verified: the reported real-world regression (plain "=rm" C code on
AArch64) compiles correctly again; X86 codegen is byte-for-byte
unaffected (confirmed against the existing rm-specific test suite,
unchanged); ARM and RISCV spot-checked to confirm both the
hand-written-IR crash guard and real Clang output are safe. Full
CodeGen/ + clang/test/{CodeGen,Sema,SemaCXX}/ + MIR/ +
MachineVerifier/ + CodeGen/GlobalISel/ (40769 tests) passes, modulo
one non-reproducible failure (passes standalone, unrelated to inline
asm) matching this session's already-established flakiness pattern
under parallel execution on this machine.
Co-Authored-By: Claude Sonnet 5 <noreply at anthropic.com>
Claude-Session: https://claude.ai/code/session_0143ozRFwhFjWmJjtfZUgmYe
---
clang/lib/CodeGen/CGStmt.cpp | 12 ++-
clang/lib/CodeGen/TargetInfo.h | 14 ++++
clang/lib/CodeGen/Targets/X86.cpp | 11 +++
llvm/include/llvm/CodeGen/TargetLowering.h | 16 ++++
.../CodeGen/GlobalISel/InlineAsmLowering.cpp | 24 ++++++
.../CodeGen/SelectionDAG/TargetLowering.cpp | 12 ++-
llvm/lib/Target/X86/X86ISelLowering.h | 5 ++
.../inline-asm-rm-unsupported-direct.ll | 27 +++++++
llvm/test/CodeGen/AArch64/inline-asm-rm.ll | 73 +++++++++++++++----
9 files changed, 176 insertions(+), 18 deletions(-)
create mode 100644 llvm/test/CodeGen/AArch64/inline-asm-rm-unsupported-direct.ll
diff --git a/clang/lib/CodeGen/CGStmt.cpp b/clang/lib/CodeGen/CGStmt.cpp
index 892da1cb99855..baaedd8da75f7 100644
--- a/clang/lib/CodeGen/CGStmt.cpp
+++ b/clang/lib/CodeGen/CGStmt.cpp
@@ -2976,9 +2976,19 @@ void AsmConstraintsInfo::HandleOutputConstraints() {
// parser here: anything it doesn't recognize (e.g. a commutative '%'
// modifier, which isn't tracked on TargetInfo::ConstraintInfo) simply
// falls through to the existing, always-safe indirect path below.
+ //
+ // getTargetHooks().supportsRegMemInlineAsmFolding() must agree with the
+ // backend's own gate
+ // (llvm::TargetLowering::supportsRegMemInlineAsmFolding()): taking the
+ // by-value path commits to the backend being able to fall back to memory if
+ // a register genuinely isn't available, and only one target has that
+ // fallback implemented today. Getting this wrong is not a missed
+ // optimization -- it turns ordinary, previously-working inline asm on every
+ // other target into a hard compile error.
bool IsExactRegMem = false;
if (CGM.getCodeGenOpts().OptimizationLevel != 0 && Info.allowsRegister() &&
- Info.allowsMemory()) {
+ Info.allowsMemory() &&
+ getTargetHooks().supportsRegMemInlineAsmFolding()) {
StringRef Codes = OutputConstraint;
if (Info.earlyClobber() && Codes.starts_with("&"))
Codes = Codes.drop_front();
diff --git a/clang/lib/CodeGen/TargetInfo.h b/clang/lib/CodeGen/TargetInfo.h
index 9405cf0124b29..35023d212ed93 100644
--- a/clang/lib/CodeGen/TargetInfo.h
+++ b/clang/lib/CodeGen/TargetInfo.h
@@ -207,6 +207,20 @@ class TargetCodeGenInfo {
return false;
}
+ /// Returns true if this target's backend can fold a register-preferred
+ /// inline asm operand back to memory under register pressure (mirrors
+ /// llvm::TargetLowering::supportsRegMemInlineAsmFolding() -- see that
+ /// declaration for why this defaults to false everywhere except the one
+ /// target that's actually implemented the fold). CGStmt.cpp's
+ /// direct-vs-indirect output lowering decision must agree with the
+ /// backend's own MayFoldRegister gate: if this target's backend won't
+ /// prefer 'r' for an exact "rm"/"+rm" constraint, emitting the IR as if
+ /// it will (skipping the historical indirect/alloca form) leaves nothing
+ /// to fall back to, and inline asm that has always compiled -- an
+ /// ordinary "=rm" or "+rm" output with no exotic pressure at all -- would
+ /// start failing outright.
+ virtual bool supportsRegMemInlineAsmFolding() const { return false; }
+
/// Adds constraints and types for result registers.
virtual void addReturnRegisterOutputs(
CodeGen::CodeGenFunction &CGF, CodeGen::LValue ReturnValue,
diff --git a/clang/lib/CodeGen/Targets/X86.cpp b/clang/lib/CodeGen/Targets/X86.cpp
index f0d108f3279fd..7ede631d6379c 100644
--- a/clang/lib/CodeGen/Targets/X86.cpp
+++ b/clang/lib/CodeGen/Targets/X86.cpp
@@ -231,6 +231,11 @@ class X86_32TargetCodeGenInfo : public TargetCodeGenInfo {
return X86AdjustInlineAsmType(CGF, Constraint, Ty);
}
+ // See X86_64TargetCodeGenInfo::supportsRegMemInlineAsmFolding(): the same
+ // shared X86InstrInfo::getFrameIndexOperands() implementation backs both
+ // ABIs here.
+ bool supportsRegMemInlineAsmFolding() const override { return true; }
+
void addReturnRegisterOutputs(CodeGenFunction &CGF, LValue ReturnValue,
std::string &Constraints,
std::vector<llvm::Type *> &ResultRegTypes,
@@ -1472,6 +1477,12 @@ class X86_64TargetCodeGenInfo : public TargetCodeGenInfo {
return X86AdjustInlineAsmType(CGF, Constraint, Ty);
}
+ // X86InstrInfo::getFrameIndexOperands() implements the addressing-mode
+ // encoding llvm::TargetLowering::supportsRegMemInlineAsmFolding() needs;
+ // see TargetCodeGenInfo's declaration for why other targets default to
+ // false here.
+ bool supportsRegMemInlineAsmFolding() const override { return true; }
+
bool isNoProtoCallVariadic(const CallArgList &args,
const FunctionNoProtoType *fnType) const override {
// The default CC on x86-64 sets %al to the number of SSA
diff --git a/llvm/include/llvm/CodeGen/TargetLowering.h b/llvm/include/llvm/CodeGen/TargetLowering.h
index 55c8f0599f411..51f3502f88c3a 100644
--- a/llvm/include/llvm/CodeGen/TargetLowering.h
+++ b/llvm/include/llvm/CodeGen/TargetLowering.h
@@ -5401,6 +5401,22 @@ class LLVM_ABI TargetLowering : public TargetLoweringBase {
/// Given a constraint, return the type of constraint it is for this target.
virtual ConstraintType getConstraintType(StringRef Constraint) const;
+ /// Returns true if this target can fold a register operand of an inline
+ /// asm instruction back to a memory operand (see
+ /// TargetInstrInfo::getFrameIndexOperands(), which a target must override
+ /// with its own addressing-mode encoding for this to work -- the base
+ /// TargetInstrInfo implementation is unreachable()). ParseConstraints()
+ /// only sets MayFoldRegister -- and so only ever prefers 'r' over 'm' for
+ /// an exact "rm"/"+rm" constraint -- when this returns true, so that
+ /// register-pressure fallback (RegAllocFast's inline asm folding, or
+ /// InlineSpiller for the greedy allocator) has an actual implementation to
+ /// fall back to instead of crashing. Defaults to false: without this,
+ /// preferring 'r' for "rm" and then genuinely running out of registers
+ /// would attempt to fold to memory and hit that unreachable() instead of
+ /// RegAllocFast's or InlineSpiller's normal "ran out of registers"
+ /// diagnostic.
+ virtual bool supportsRegMemInlineAsmFolding() const { return false; }
+
using ConstraintPair = std::pair<StringRef, TargetLowering::ConstraintType>;
using ConstraintGroup = SmallVector<ConstraintPair>;
/// Given an OpInfo with list of constraints codes as strings, return a
diff --git a/llvm/lib/CodeGen/GlobalISel/InlineAsmLowering.cpp b/llvm/lib/CodeGen/GlobalISel/InlineAsmLowering.cpp
index 14ce305376459..6708e3ea95b33 100644
--- a/llvm/lib/CodeGen/GlobalISel/InlineAsmLowering.cpp
+++ b/llvm/lib/CodeGen/GlobalISel/InlineAsmLowering.cpp
@@ -337,6 +337,30 @@ bool InlineAsmLowering::lowerInlineAsm(
switch (OpInfo.Type) {
case InlineAsm::isOutput:
if (OpInfo.ConstraintType == TargetLowering::C_Memory) {
+ // A direct (non-indirect) output has no CallOperandVal -- it's a
+ // return value, not a call argument, so there's nothing to
+ // reference in memory. This is only reachable for an output
+ // constraint whose alternatives include memory (e.g. "rm") if
+ // TargetLowering::supportsRegMemInlineAsmFolding() is false for
+ // this target: with it true, TargetLowering::ParseConstraints()
+ // never lets such a constraint prefer memory in the first place
+ // (see MayFoldRegister), and with it false, the frontend is
+ // expected to have already lowered this to an indirect (pointer
+ // argument) form itself -- see CGStmt.cpp's mirroring check. Fail
+ // with a clean diagnostic rather than dereferencing the null
+ // CallOperandVal below, for the same reason
+ // computeConstraintToUse() in SelectionDAGBuilder.cpp does for the
+ // equivalent SelectionDAG-path case.
+ if (!OpInfo.isIndirect) {
+ emitInlineAsmError(MIRBuilder, Call,
+ "unsupported inline asm: constraint '" +
+ Twine(OpInfo.ConstraintCode) +
+ "' cannot be satisfied in a register and "
+ "has no memory to fall back to",
+ GetOrCreateVRegs(Call));
+ return true;
+ }
+
const InlineAsm::ConstraintCode ConstraintID =
TLI->getInlineAsmMemConstraint(OpInfo.ConstraintCode);
assert(ConstraintID != InlineAsm::ConstraintCode::Unknown &&
diff --git a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
index df56b983ce46f..81e8052be9309 100644
--- a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
@@ -6130,8 +6130,18 @@ TargetLowering::ParseConstraints(const DataLayout &DL,
// direct "=rm" output with a matching tied input). The register allocator
// can fold both the output and its tied input to the same memory slot when
// under pressure.
+ //
+ // Gated on supportsRegMemInlineAsmFolding(): preferring 'r' here commits
+ // to a register allocator being able to fold back to memory if 'r'
+ // doesn't pan out, and that fold has no generic implementation -- it
+ // needs a target-specific TargetInstrInfo::getFrameIndexOperands()
+ // override, which today only X86 has. Without this guard, any other
+ // target would prefer 'r', then hit an unreachable() the moment a
+ // register genuinely wasn't available, instead of the register
+ // allocator's normal, clean "ran out of registers" diagnostic.
if (OpInfo.Codes.size() == 2 && llvm::is_contained(OpInfo.Codes, "r") &&
- llvm::is_contained(OpInfo.Codes, "m"))
+ llvm::is_contained(OpInfo.Codes, "m") &&
+ supportsRegMemInlineAsmFolding())
OpInfo.MayFoldRegister = true;
// Compute the value type for each operand.
diff --git a/llvm/lib/Target/X86/X86ISelLowering.h b/llvm/lib/Target/X86/X86ISelLowering.h
index 798050028c15a..239140d3f438d 100644
--- a/llvm/lib/Target/X86/X86ISelLowering.h
+++ b/llvm/lib/Target/X86/X86ISelLowering.h
@@ -410,6 +410,11 @@ namespace llvm {
ConstraintType getConstraintType(StringRef Constraint) const override;
+ // X86InstrInfo::getFrameIndexOperands() implements the addressing-mode
+ // encoding this needs; see TargetLowering's declaration for why other
+ // targets default to false here.
+ bool supportsRegMemInlineAsmFolding() const override { return true; }
+
/// Examine constraint string and operand type and determine a weight value.
/// The operand object must already have been set up with the operand type.
ConstraintWeight
diff --git a/llvm/test/CodeGen/AArch64/inline-asm-rm-unsupported-direct.ll b/llvm/test/CodeGen/AArch64/inline-asm-rm-unsupported-direct.ll
new file mode 100644
index 0000000000000..c0cafb175ada0
--- /dev/null
+++ b/llvm/test/CodeGen/AArch64/inline-asm-rm-unsupported-direct.ll
@@ -0,0 +1,27 @@
+; RUN: not llc -mtriple=aarch64-unknown-linux-gnu --global-isel=true --fast-isel=false --regalloc=greedy < %s 2>&1 \
+; RUN: | FileCheck %s
+; RUN: not llc -mtriple=aarch64-unknown-linux-gnu --global-isel=false --fast-isel=true --regalloc=greedy < %s 2>&1 \
+; RUN: | FileCheck %s
+; RUN: not llc -mtriple=aarch64-unknown-linux-gnu --global-isel=false --fast-isel=false --regalloc=greedy < %s 2>&1 \
+; RUN: | FileCheck %s
+; RUN: not llc -mtriple=aarch64-unknown-linux-gnu --global-isel=false --fast-isel=false --regalloc=fast < %s 2>&1 \
+; RUN: | FileCheck %s
+
+; A *direct* (non-indirect) "=rm" output -- something Clang itself would
+; never emit for this target, precisely because
+; TargetLowering::supportsRegMemInlineAsmFolding() is false for it and
+; CGStmt.cpp mirrors that check -- is still directly constructible in IR, by
+; another frontend or by hand, as here. Regression guard for a case that
+; used to hard-crash instead of failing cleanly: an UNREACHABLE in
+; SelectionDAGBuilder's computeConstraintToUse() for the SelectionDAG
+; frameworks, and a null CallOperandVal dereference in
+; InlineAsmLowering::lowerInlineAsm() for GlobalISel. A clean diagnostic is
+; the correct behavior for a target with no register-pressure fallback to
+; offer.
+
+; CHECK: error: unsupported inline asm: constraint 'm' cannot be satisfied in a register and has no memory to fall back to
+define i64 @test_rm_output_direct_unsupported() {
+entry:
+ %0 = call i64 asm "// $0", "=rm"()
+ ret i64 %0
+}
diff --git a/llvm/test/CodeGen/AArch64/inline-asm-rm.ll b/llvm/test/CodeGen/AArch64/inline-asm-rm.ll
index 817b8fb021864..f2288c919969c 100644
--- a/llvm/test/CodeGen/AArch64/inline-asm-rm.ll
+++ b/llvm/test/CodeGen/AArch64/inline-asm-rm.ll
@@ -12,51 +12,92 @@
; RUN: llc -mtriple=aarch64-unknown-linux-gnu --global-isel=false --fast-isel=false --regalloc=fast < %s \
; RUN: | FileCheck --check-prefix=FAST_RA %s
-; Test rm constraints on AArch64 under all three ISel frameworks.
-
-define i64 @test_rm_output() {
-; GLOBAL_ISEL_GREEDY_RA-LABEL: test_rm_output:
+; AArch64 has not implemented TargetInstrInfo::getFrameIndexOperands(), so
+; TargetLowering::supportsRegMemInlineAsmFolding() correctly returns false
+; for it: preferring a register for an "rm" constraint here would have
+; nothing to fall back to if that register genuinely isn't available.
+; CGStmt.cpp's mirroring check means Clang, for this target, keeps emitting
+; the indirect (pointer-argument) form for a plain "rm"/"+rm" output that it
+; always has -- this is the realistic case AArch64 users actually hit, and
+; it's expected to work correctly and identically across all three ISel
+; frameworks, the same as it always has. See
+; inline-asm-rm-unsupported-direct.ll for the direct-form (Clang would never
+; emit this for this target) regression guard.
+define i64 @test_rm_output_indirect() {
+; GLOBAL_ISEL_GREEDY_RA-LABEL: test_rm_output_indirect:
; GLOBAL_ISEL_GREEDY_RA: // %bb.0: // %entry
+; GLOBAL_ISEL_GREEDY_RA-NEXT: sub sp, sp, #16
+; GLOBAL_ISEL_GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GLOBAL_ISEL_GREEDY_RA-NEXT: add x8, sp, #8
; GLOBAL_ISEL_GREEDY_RA-NEXT: //APP
-; GLOBAL_ISEL_GREEDY_RA-NEXT: // x0
+; GLOBAL_ISEL_GREEDY_RA-NEXT: // [x8]
; GLOBAL_ISEL_GREEDY_RA-NEXT: //NO_APP
+; GLOBAL_ISEL_GREEDY_RA-NEXT: ldr x0, [sp, #8]
+; GLOBAL_ISEL_GREEDY_RA-NEXT: add sp, sp, #16
; GLOBAL_ISEL_GREEDY_RA-NEXT: ret
;
-; GLOBAL_ISEL_FAST_RA-LABEL: test_rm_output:
+; GLOBAL_ISEL_FAST_RA-LABEL: test_rm_output_indirect:
; GLOBAL_ISEL_FAST_RA: // %bb.0: // %entry
+; GLOBAL_ISEL_FAST_RA-NEXT: sub sp, sp, #16
+; GLOBAL_ISEL_FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; GLOBAL_ISEL_FAST_RA-NEXT: add x8, sp, #8
; GLOBAL_ISEL_FAST_RA-NEXT: //APP
-; GLOBAL_ISEL_FAST_RA-NEXT: // x0
+; GLOBAL_ISEL_FAST_RA-NEXT: // [x8]
; GLOBAL_ISEL_FAST_RA-NEXT: //NO_APP
+; GLOBAL_ISEL_FAST_RA-NEXT: ldr x0, [sp, #8]
+; GLOBAL_ISEL_FAST_RA-NEXT: add sp, sp, #16
; GLOBAL_ISEL_FAST_RA-NEXT: ret
;
-; FAST_ISEL_GREEDY_RA-LABEL: test_rm_output:
+; FAST_ISEL_GREEDY_RA-LABEL: test_rm_output_indirect:
; FAST_ISEL_GREEDY_RA: // %bb.0: // %entry
+; FAST_ISEL_GREEDY_RA-NEXT: sub sp, sp, #16
+; FAST_ISEL_GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_ISEL_GREEDY_RA-NEXT: add x8, sp, #8
; FAST_ISEL_GREEDY_RA-NEXT: //APP
-; FAST_ISEL_GREEDY_RA-NEXT: // x0
+; FAST_ISEL_GREEDY_RA-NEXT: // [x8]
; FAST_ISEL_GREEDY_RA-NEXT: //NO_APP
+; FAST_ISEL_GREEDY_RA-NEXT: ldr x0, [sp, #8]
+; FAST_ISEL_GREEDY_RA-NEXT: add sp, sp, #16
; FAST_ISEL_GREEDY_RA-NEXT: ret
;
-; FAST_ISEL_FAST_RA-LABEL: test_rm_output:
+; FAST_ISEL_FAST_RA-LABEL: test_rm_output_indirect:
; FAST_ISEL_FAST_RA: // %bb.0: // %entry
+; FAST_ISEL_FAST_RA-NEXT: sub sp, sp, #16
+; FAST_ISEL_FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_ISEL_FAST_RA-NEXT: add x8, sp, #8
; FAST_ISEL_FAST_RA-NEXT: //APP
-; FAST_ISEL_FAST_RA-NEXT: // x0
+; FAST_ISEL_FAST_RA-NEXT: // [x8]
; FAST_ISEL_FAST_RA-NEXT: //NO_APP
+; FAST_ISEL_FAST_RA-NEXT: ldr x0, [sp, #8]
+; FAST_ISEL_FAST_RA-NEXT: add sp, sp, #16
; FAST_ISEL_FAST_RA-NEXT: ret
;
-; GREEDY_RA-LABEL: test_rm_output:
+; GREEDY_RA-LABEL: test_rm_output_indirect:
; GREEDY_RA: // %bb.0: // %entry
+; GREEDY_RA-NEXT: sub sp, sp, #16
+; GREEDY_RA-NEXT: .cfi_def_cfa_offset 16
+; GREEDY_RA-NEXT: add x8, sp, #8
; GREEDY_RA-NEXT: //APP
-; GREEDY_RA-NEXT: // x0
+; GREEDY_RA-NEXT: // [x8]
; GREEDY_RA-NEXT: //NO_APP
+; GREEDY_RA-NEXT: ldr x0, [sp, #8]
+; GREEDY_RA-NEXT: add sp, sp, #16
; GREEDY_RA-NEXT: ret
;
-; FAST_RA-LABEL: test_rm_output:
+; FAST_RA-LABEL: test_rm_output_indirect:
; FAST_RA: // %bb.0: // %entry
+; FAST_RA-NEXT: sub sp, sp, #16
+; FAST_RA-NEXT: .cfi_def_cfa_offset 16
+; FAST_RA-NEXT: add x8, sp, #8
; FAST_RA-NEXT: //APP
-; FAST_RA-NEXT: // x0
+; FAST_RA-NEXT: // [x8]
; FAST_RA-NEXT: //NO_APP
+; FAST_RA-NEXT: ldr x0, [sp, #8]
+; FAST_RA-NEXT: add sp, sp, #16
; FAST_RA-NEXT: ret
entry:
- %0 = tail call i64 asm "# $0", "=rm"()
+ %out = alloca i64, align 8
+ call void asm "// $0", "=*rm"(ptr elementtype(i64) %out)
+ %0 = load i64, ptr %out, align 8
ret i64 %0
}
>From 919f10b1f9e2cf22a27f9a47b3598bcbabc72dc9 Mon Sep 17 00:00:00 2001
From: Bill Wendling <isanbard at gmail.com>
Date: Tue, 4 Aug 2026 02:47:09 -0700
Subject: [PATCH 08/12] [test] Add Clang-level coverage for "rm"/"+rm" output
codegen
Closes a real blind spot: every rm-specific test on this branch so
far operates at the LLVM IR level, using hand-written constraint
strings that bypass Clang's own CGStmt.cpp lowering decision
entirely. Nothing exercised what real users actually compile -- and,
as the preceding commit's regression demonstrates, that's exactly
where a real break can hide undetected.
clang/test/CodeGen/asm-reg-mem-constraints.c checks IR generation
(not final codegen -- the fold itself happens during register
allocation, well after IR generation, and is already covered by the
existing LLVM-IR-level torture tests) for the three cases that matter
at the Clang level specifically:
- A plain "=rm" output takes the direct (by-value, no alloca) path
above -O0, and the historical indirect (alloca + out-parameter)
path at -O0, mirroring the backend's own optimization-level guard.
- A tied "+rm" output does too -- the specific case fixed two commits
ago; this is its regression test.
- An output whose address later escapes to another call still gets a
real store to memory before that call, regardless of the asm's own
operand preferring a register -- pre-existing, general "materialize
to memory only when something forces it" logic this change doesn't
touch, included as a guard against future regressions making the
escaped address diverge from what the asm actually wrote.
Co-Authored-By: Claude Sonnet 5 <noreply at anthropic.com>
Claude-Session: https://claude.ai/code/session_0143ozRFwhFjWmJjtfZUgmYe
---
clang/test/CodeGen/asm-reg-mem-constraints.c | 109 +++++++++++++++++++
1 file changed, 109 insertions(+)
create mode 100644 clang/test/CodeGen/asm-reg-mem-constraints.c
diff --git a/clang/test/CodeGen/asm-reg-mem-constraints.c b/clang/test/CodeGen/asm-reg-mem-constraints.c
new file mode 100644
index 0000000000000..d50173eb5a953
--- /dev/null
+++ b/clang/test/CodeGen/asm-reg-mem-constraints.c
@@ -0,0 +1,109 @@
+// NOTE: Assertions have been autogenerated by utils/update_cc_test_checks.py UTC_ARGS: --version 6
+// RUN: %clang_cc1 -triple x86_64-unknown-linux-gnu -O2 -emit-llvm %s -o - | FileCheck --check-prefix=O2 %s
+// RUN: %clang_cc1 -triple x86_64-unknown-linux-gnu -O0 -emit-llvm %s -o - | FileCheck --check-prefix=O0 %s
+
+// CGStmt.cpp emits an exact register-or-memory ("rm"/"mr") output directly
+// (by value, no alloca) above -O0, the same as a plain "=r" output, and
+// leaves the register-vs-memory decision -- and, if needed, the fallback to
+// memory -- to the backend (TargetLowering::MayFoldRegister,
+// RegAllocFast's inline asm folding). At -O0 it keeps the historical
+// indirect (alloca + out-parameter) form, matching every other memory-
+// capable constraint: MayFoldRegister is disabled at -O0 too (the fast
+// allocator's own on-demand folding is what -O0 relies on instead), so
+// there would be nothing gained by taking the direct path there, and the
+// indirect form is what every codegen path has always been tested against.
+// O2-LABEL: define dso_local i64 @test_rm_output(
+// O2-SAME: ) local_unnamed_addr #[[ATTR0:[0-9]+]] {
+// O2-NEXT: [[ENTRY:.*:]]
+// O2-NEXT: [[TMP0:%.*]] = tail call i64 asm sideeffect "# $0", "=rm,~{dirflag},~{fpsr},~{flags}"() #[[ATTR3:[0-9]+]], !srcloc [[META6:![0-9]+]]
+// O2-NEXT: ret i64 [[TMP0]]
+//
+// O0-LABEL: define dso_local i64 @test_rm_output(
+// O0-SAME: ) #[[ATTR0:[0-9]+]] {
+// O0-NEXT: [[ENTRY:.*:]]
+// O0-NEXT: [[OUT:%.*]] = alloca i64, align 8
+// O0-NEXT: call void asm sideeffect "# $0", "=*rm,~{dirflag},~{fpsr},~{flags}"(ptr elementtype(i64) [[OUT]]) #[[ATTR2:[0-9]+]], !srcloc [[META1:![0-9]+]]
+// O0-NEXT: [[TMP0:%.*]] = load i64, ptr [[OUT]], align 8
+// O0-NEXT: ret i64 [[TMP0]]
+//
+unsigned long test_rm_output(void) {
+ unsigned long out;
+ __asm__ volatile("# %0" : "=rm"(out));
+ return out;
+}
+
+// The tied "+rm" case is the one this file exists to cover: before, a tied
+// register-or-memory output kept an indirect alloca even at -O2 with no
+// register pressure at all, because CGStmt.cpp's direct-vs-indirect
+// decision didn't recognize "rm" as different from a plain "m" -- so a
+// read-modify-write asm operand paid for a real memory round trip it never
+// needed.
+// O2-LABEL: define dso_local i64 @test_tied_rm_output(
+// O2-SAME: i64 noundef [[LEN:%.*]]) local_unnamed_addr #[[ATTR0]] {
+// O2-NEXT: [[ENTRY:.*:]]
+// O2-NEXT: [[TMP0:%.*]] = tail call i64 asm sideeffect "incq $0", "=rm,0,~{dirflag},~{fpsr},~{flags}"(i64 [[LEN]]) #[[ATTR3]], !srcloc [[META7:![0-9]+]]
+// O2-NEXT: ret i64 [[TMP0]]
+//
+// O0-LABEL: define dso_local i64 @test_tied_rm_output(
+// O0-SAME: i64 noundef [[LEN:%.*]]) #[[ATTR0]] {
+// O0-NEXT: [[ENTRY:.*:]]
+// O0-NEXT: [[LEN_ADDR:%.*]] = alloca i64, align 8
+// O0-NEXT: store i64 [[LEN]], ptr [[LEN_ADDR]], align 8
+// O0-NEXT: [[TMP0:%.*]] = load i64, ptr [[LEN_ADDR]], align 8
+// O0-NEXT: call void asm sideeffect "incq $0", "=*rm,0,~{dirflag},~{fpsr},~{flags}"(ptr elementtype(i64) [[LEN_ADDR]], i64 [[TMP0]]) #[[ATTR2]], !srcloc [[META2:![0-9]+]]
+// O0-NEXT: [[TMP1:%.*]] = load i64, ptr [[LEN_ADDR]], align 8
+// O0-NEXT: ret i64 [[TMP1]]
+//
+unsigned long test_tied_rm_output(unsigned long len) {
+ __asm__ volatile("incq %0" : "+rm"(len));
+ return len;
+}
+
+// The output's address escaping to another call must still force a real
+// store to memory before that call, regardless of the register preference
+// for the asm's own operand. This exercises pre-existing, general
+// "materialize to memory only when something forces it" logic, not
+// anything specific to the "rm" constraint -- included here as a guard
+// against a regression that would make the escaped address diverge from
+// what the asm actually wrote.
+void observe(unsigned long *p);
+// O2-LABEL: define dso_local i64 @test_rm_output_escapes(
+// O2-SAME: ) local_unnamed_addr #[[ATTR0]] {
+// O2-NEXT: [[ENTRY:.*:]]
+// O2-NEXT: [[OUT:%.*]] = alloca i64, align 8
+// O2-NEXT: call void @llvm.lifetime.start.p0(ptr nonnull [[OUT]]) #[[ATTR3]]
+// O2-NEXT: [[TMP0:%.*]] = tail call i64 asm sideeffect "# $0", "=rm,~{dirflag},~{fpsr},~{flags}"() #[[ATTR3]], !srcloc [[META8:![0-9]+]]
+// O2-NEXT: store i64 [[TMP0]], ptr [[OUT]], align 8, !tbaa [[LONG_TBAA9:![0-9]+]]
+// O2-NEXT: call void @observe(ptr noundef nonnull [[OUT]]) #[[ATTR3]]
+// O2-NEXT: [[TMP1:%.*]] = load i64, ptr [[OUT]], align 8, !tbaa [[LONG_TBAA9]]
+// O2-NEXT: call void @llvm.lifetime.end.p0(ptr nonnull [[OUT]]) #[[ATTR3]]
+// O2-NEXT: ret i64 [[TMP1]]
+//
+// O0-LABEL: define dso_local i64 @test_rm_output_escapes(
+// O0-SAME: ) #[[ATTR0]] {
+// O0-NEXT: [[ENTRY:.*:]]
+// O0-NEXT: [[OUT:%.*]] = alloca i64, align 8
+// O0-NEXT: call void asm sideeffect "# $0", "=*rm,~{dirflag},~{fpsr},~{flags}"(ptr elementtype(i64) [[OUT]]) #[[ATTR2]], !srcloc [[META3:![0-9]+]]
+// O0-NEXT: call void @observe(ptr noundef [[OUT]])
+// O0-NEXT: [[TMP0:%.*]] = load i64, ptr [[OUT]], align 8
+// O0-NEXT: ret i64 [[TMP0]]
+//
+unsigned long test_rm_output_escapes(void) {
+ unsigned long out;
+ __asm__ volatile("# %0" : "=rm"(out));
+ observe(&out);
+ return out;
+}
+//.
+// O2: [[META4:![0-9]+]] = !{!"omnipotent char", [[META5:![0-9]+]], i64 0}
+// O2: [[META5]] = !{!"Simple C/C++ TBAA"}
+// O2: [[META6]] = !{i64 1849}
+// O2: [[META7]] = !{i64 3272}
+// O2: [[META8]] = !{i64 5115}
+// O2: [[LONG_TBAA9]] = !{[[META10:![0-9]+]], [[META10]], i64 0}
+// O2: [[META10]] = !{!"long", [[META4]], i64 0}
+//.
+// O0: [[META1]] = !{i64 1849}
+// O0: [[META2]] = !{i64 3272}
+// O0: [[META3]] = !{i64 5115}
+//.
>From e7ceb299a4682c1fb7f0f07fc5f8f7d1f9f97e3d Mon Sep 17 00:00:00 2001
From: Bill Wendling <isanbard at gmail.com>
Date: Tue, 4 Aug 2026 03:26:45 -0700
Subject: [PATCH 09/12] [Clang] Add missing supportsRegMemInlineAsmFolding()
override for Win64
WinX86_64TargetCodeGenInfo inherits directly from TargetCodeGenInfo,
not from X86_64TargetCodeGenInfo (unlike WinX86_32, which inherits
from X86_32TargetCodeGenInfo and so already picked up the override
transitively). It was missed when the override was added to the
System V 32- and 64-bit classes, so x86_64-pc-windows-msvc/mingw
silently fell back to TargetCodeGenInfo's default of false --
correct in the sense that it fails safe (the historical indirect
codegen path, not a miscompile or a crash), but an unintended
asymmetry: the underlying LLVM-side capability
(X86TargetLowering::supportsRegMemInlineAsmFolding(), via
X86InstrInfo::getFrameIndexOperands()) is OS-independent, so there's
no actual reason Windows x86-64 should get worse "rm"/"+rm" output
codegen than System V x86-64 gets.
Verified: x86_64-pc-windows-msvc now emits the direct (by-value, no
alloca) form for a plain "=rm" output, matching
x86_64-unknown-linux-gnu. Full clang/test/CodeGen/ (6099 tests)
passes.
Co-Authored-By: Claude Sonnet 5 <noreply at anthropic.com>
Claude-Session: https://claude.ai/code/session_0143ozRFwhFjWmJjtfZUgmYe
---
clang/lib/CodeGen/Targets/X86.cpp | 11 +++++++++++
1 file changed, 11 insertions(+)
diff --git a/clang/lib/CodeGen/Targets/X86.cpp b/clang/lib/CodeGen/Targets/X86.cpp
index 7ede631d6379c..813e4e7e8d474 100644
--- a/clang/lib/CodeGen/Targets/X86.cpp
+++ b/clang/lib/CodeGen/Targets/X86.cpp
@@ -1756,6 +1756,17 @@ class WinX86_64TargetCodeGenInfo : public TargetCodeGenInfo {
std::make_unique<SwiftABIInfo>(CGT, /*SwiftErrorInRegister=*/true);
}
+ // Unlike adjustInlineAsmType() (an ABI-specific type-lowering hook this
+ // class deliberately doesn't override, inheriting TargetCodeGenInfo's
+ // no-op default), register-or-memory inline asm folding support is a
+ // backend/ISA fact, not an ABI one: X86InstrInfo::getFrameIndexOperands()
+ // backs every X86 subtarget the same way regardless of OS. This class
+ // doesn't inherit from X86_64TargetCodeGenInfo (unlike WinX86_32, which
+ // inherits from X86_32TargetCodeGenInfo and so already gets this), so it
+ // needs its own override to avoid silently falling back to the
+ // TargetCodeGenInfo default of false.
+ bool supportsRegMemInlineAsmFolding() const override { return true; }
+
void setTargetAttributes(const Decl *D, llvm::GlobalValue *GV,
CodeGen::CodeGenModule &CGM) const override;
>From d2a49e1cf12769b8731e34b359bf078c3e1a5346 Mon Sep 17 00:00:00 2001
From: Bill Wendling <isanbard at gmail.com>
Date: Tue, 4 Aug 2026 03:49:20 -0700
Subject: [PATCH 10/12] [NFC][TargetLowering] Share the "unsupported inline
asm" diagnostic text
SelectionDAGBuilder.cpp's computeConstraintToUse() and
InlineAsmLowering.cpp's lowerInlineAsm() -- otherwise entirely
independent implementations of the SelectionDAG and GlobalISel
constraint-selection paths -- each had their own hand-written copy of
the same diagnostic string ("unsupported inline asm: constraint '...'
cannot be satisfied in a register and has no memory to fall back
to"). Nothing kept them in sync; a future edit to one with no
corresponding edit to the other would silently produce two different
error messages for what's conceptually the same failure.
Add TargetLowering::getRegMemInlineAsmUnsupportedDiag() as a small
static helper, next to supportsRegMemInlineAsmFolding() since it's
the diagnostic for exactly that gate's failure mode, and have both
call sites build their message through it instead of duplicating the
text.
No functional change -- verified both diagnostics still fire with
identical text (SelectionDAG via a direct-form hand-written IR test,
GlobalISel via inline-asm-rm-unsupported-direct.ll on AArch64), and
the full CodeGen/ + clang/test/CodeGen/ + CodeGen/GlobalISel/ suite
(37377 tests) passes.
Co-Authored-By: Claude Sonnet 5 <noreply at anthropic.com>
Claude-Session: https://claude.ai/code/session_0143ozRFwhFjWmJjtfZUgmYe
---
llvm/include/llvm/CodeGen/TargetLowering.h | 17 +++++++++++++++++
.../CodeGen/GlobalISel/InlineAsmLowering.cpp | 6 ++----
.../SelectionDAG/SelectionDAGBuilder.cpp | 6 ++----
3 files changed, 21 insertions(+), 8 deletions(-)
diff --git a/llvm/include/llvm/CodeGen/TargetLowering.h b/llvm/include/llvm/CodeGen/TargetLowering.h
index 51f3502f88c3a..c8760a69f8135 100644
--- a/llvm/include/llvm/CodeGen/TargetLowering.h
+++ b/llvm/include/llvm/CodeGen/TargetLowering.h
@@ -27,6 +27,7 @@
#include "llvm/ADT/DenseMap.h"
#include "llvm/ADT/SmallVector.h"
#include "llvm/ADT/StringRef.h"
+#include "llvm/ADT/Twine.h"
#include "llvm/CodeGen/DAGCombine.h"
#include "llvm/CodeGen/ISDOpcodes.h"
#include "llvm/CodeGen/LibcallLoweringInfo.h"
@@ -5417,6 +5418,22 @@ class LLVM_ABI TargetLowering : public TargetLoweringBase {
/// diagnostic.
virtual bool supportsRegMemInlineAsmFolding() const { return false; }
+ /// The diagnostic to report when a direct (non-indirect) inline asm
+ /// operand needed a memory constraint but has nowhere to spill to -- see
+ /// supportsRegMemInlineAsmFolding() above for when this can happen.
+ /// Shared between the SelectionDAG (SelectionDAGBuilder.cpp's
+ /// computeConstraintToUse()) and GlobalISel (InlineAsmLowering.cpp's
+ /// lowerInlineAsm()) constraint-selection paths, which are otherwise
+ /// entirely independent implementations, so the two can't drift apart by
+ /// having their own hand-written copies of the same message text.
+ static std::string
+ getRegMemInlineAsmUnsupportedDiag(StringRef ConstraintCode) {
+ return ("unsupported inline asm: constraint '" + ConstraintCode +
+ "' cannot be satisfied in a register and has no memory to fall "
+ "back to")
+ .str();
+ }
+
using ConstraintPair = std::pair<StringRef, TargetLowering::ConstraintType>;
using ConstraintGroup = SmallVector<ConstraintPair>;
/// Given an OpInfo with list of constraints codes as strings, return a
diff --git a/llvm/lib/CodeGen/GlobalISel/InlineAsmLowering.cpp b/llvm/lib/CodeGen/GlobalISel/InlineAsmLowering.cpp
index 6708e3ea95b33..3b78c6245a973 100644
--- a/llvm/lib/CodeGen/GlobalISel/InlineAsmLowering.cpp
+++ b/llvm/lib/CodeGen/GlobalISel/InlineAsmLowering.cpp
@@ -353,10 +353,8 @@ bool InlineAsmLowering::lowerInlineAsm(
// equivalent SelectionDAG-path case.
if (!OpInfo.isIndirect) {
emitInlineAsmError(MIRBuilder, Call,
- "unsupported inline asm: constraint '" +
- Twine(OpInfo.ConstraintCode) +
- "' cannot be satisfied in a register and "
- "has no memory to fall back to",
+ TargetLowering::getRegMemInlineAsmUnsupportedDiag(
+ OpInfo.ConstraintCode),
GetOrCreateVRegs(Call));
return true;
}
diff --git a/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp b/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp
index 2291ebe60bacc..86ce1774c7a88 100644
--- a/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp
@@ -10480,10 +10480,8 @@ computeConstraintToUse(ConstraintDecisionInfo &Info, const CallBase &Call,
// choosing C_Memory here. But hand-written or other-frontend IR can
// still construct this shape, so fail with a clean diagnostic
// rather than the assertion this used to be.
- Info.ErrorMsg << "unsupported inline asm: constraint '"
- << OpInfo.ConstraintCode
- << "' cannot be satisfied in a register and has no "
- "memory to fall back to";
+ Info.ErrorMsg << TargetLowering::getRegMemInlineAsmUnsupportedDiag(
+ OpInfo.ConstraintCode);
return true;
}
>From db67ad5e6cb44bf0c86660f9a05e06e7461bfeb5 Mon Sep 17 00:00:00 2001
From: Bill Wendling <isanbard at gmail.com>
Date: Tue, 4 Aug 2026 04:02:39 -0700
Subject: [PATCH 11/12] [NFC][Clang] Factor supportsRegMemInlineAsmFolding()
into a shared helper
X86_32TargetCodeGenInfo, X86_64TargetCodeGenInfo, and
WinX86_64TargetCodeGenInfo each hand-duplicated the same
"bool supportsRegMemInlineAsmFolding() const override { return true; }"
one-liner, with three near-identical explanations of why the answer is
the same across all of them.
Factor the answer into a single free function,
X86SupportsRegMemInlineAsmFolding(), mirroring the existing
X86AdjustInlineAsmType() precedent in this file for logic shared across
the X86 TargetCodeGenInfo subclasses. Each override now just forwards
to it, and the canonical "why" explanation lives in one place instead
of three.
Verified via a clean ninja build and a full pass of
clang/test/CodeGen/asm-reg-mem-constraints.c and the "rm"/"+rm"
constraint-folding test suites under clang/test/CodeGen/X86 and
llvm/test/CodeGen/{X86,AArch64} (269/269 passing).
---
clang/lib/CodeGen/Targets/X86.cpp | 42 ++++++++++++++++++-------------
1 file changed, 25 insertions(+), 17 deletions(-)
diff --git a/clang/lib/CodeGen/Targets/X86.cpp b/clang/lib/CodeGen/Targets/X86.cpp
index 813e4e7e8d474..f29dd660ea712 100644
--- a/clang/lib/CodeGen/Targets/X86.cpp
+++ b/clang/lib/CodeGen/Targets/X86.cpp
@@ -45,6 +45,16 @@ static llvm::Type *X86AdjustInlineAsmType(CodeGen::CodeGenFunction &CGF,
return Ty;
}
+// Register-or-memory inline asm folding support is a backend/ISA fact, not
+// an ABI one: X86InstrInfo::getFrameIndexOperands() backs every X86
+// subtarget (32-bit, 64-bit, and their Windows variants alike) the same
+// way regardless of OS or calling convention. Every X86TargetCodeGenInfo
+// subclass shares this one answer, so it's factored into a single free
+// function (mirroring X86AdjustInlineAsmType() above) rather than
+// hand-duplicated in each override; see TargetCodeGenInfo's declaration
+// for why other targets default to false instead.
+static bool X86SupportsRegMemInlineAsmFolding() { return true; }
+
/// Returns true if this type can be passed in SSE registers with the
/// X86_VectorCall calling convention. Shared between x86_32 and x86_64.
static bool isX86VectorTypeForVectorCall(ASTContext &Context, QualType Ty) {
@@ -231,10 +241,9 @@ class X86_32TargetCodeGenInfo : public TargetCodeGenInfo {
return X86AdjustInlineAsmType(CGF, Constraint, Ty);
}
- // See X86_64TargetCodeGenInfo::supportsRegMemInlineAsmFolding(): the same
- // shared X86InstrInfo::getFrameIndexOperands() implementation backs both
- // ABIs here.
- bool supportsRegMemInlineAsmFolding() const override { return true; }
+ bool supportsRegMemInlineAsmFolding() const override {
+ return X86SupportsRegMemInlineAsmFolding();
+ }
void addReturnRegisterOutputs(CodeGenFunction &CGF, LValue ReturnValue,
std::string &Constraints,
@@ -1477,11 +1486,9 @@ class X86_64TargetCodeGenInfo : public TargetCodeGenInfo {
return X86AdjustInlineAsmType(CGF, Constraint, Ty);
}
- // X86InstrInfo::getFrameIndexOperands() implements the addressing-mode
- // encoding llvm::TargetLowering::supportsRegMemInlineAsmFolding() needs;
- // see TargetCodeGenInfo's declaration for why other targets default to
- // false here.
- bool supportsRegMemInlineAsmFolding() const override { return true; }
+ bool supportsRegMemInlineAsmFolding() const override {
+ return X86SupportsRegMemInlineAsmFolding();
+ }
bool isNoProtoCallVariadic(const CallArgList &args,
const FunctionNoProtoType *fnType) const override {
@@ -1758,14 +1765,15 @@ class WinX86_64TargetCodeGenInfo : public TargetCodeGenInfo {
// Unlike adjustInlineAsmType() (an ABI-specific type-lowering hook this
// class deliberately doesn't override, inheriting TargetCodeGenInfo's
- // no-op default), register-or-memory inline asm folding support is a
- // backend/ISA fact, not an ABI one: X86InstrInfo::getFrameIndexOperands()
- // backs every X86 subtarget the same way regardless of OS. This class
- // doesn't inherit from X86_64TargetCodeGenInfo (unlike WinX86_32, which
- // inherits from X86_32TargetCodeGenInfo and so already gets this), so it
- // needs its own override to avoid silently falling back to the
- // TargetCodeGenInfo default of false.
- bool supportsRegMemInlineAsmFolding() const override { return true; }
+ // no-op default), this class doesn't inherit from X86_64TargetCodeGenInfo
+ // (unlike WinX86_32, which inherits from X86_32TargetCodeGenInfo and so
+ // already gets this), so it needs its own override to avoid silently
+ // falling back to the TargetCodeGenInfo default of false. See
+ // X86SupportsRegMemInlineAsmFolding() above for why the answer is the
+ // same across every X86 subclass.
+ bool supportsRegMemInlineAsmFolding() const override {
+ return X86SupportsRegMemInlineAsmFolding();
+ }
void setTargetAttributes(const Decl *D, llvm::GlobalValue *GV,
CodeGen::CodeGenModule &CGM) const override;
>From 705f24c082f256fb416b15d2231d7a75f044aaee Mon Sep 17 00:00:00 2001
From: Bill Wendling <isanbard at gmail.com>
Date: Tue, 4 Aug 2026 04:28:10 -0700
Subject: [PATCH 12/12] [TargetLowering][RegAllocFast] Fix tied "+rm" operand
pressure over-counting
TargetLowering::ParseConstraints only set AsmOperandInfo::MayFoldRegister
by inspecting an operand's own constraint codes. For a tied "+rm" pair,
the output's codes are {"r","m"} and get it set correctly, but the
matching input's own codes are just the matching digit (e.g. "0"), so
its MayFoldRegister stayed false -- and with it, the RegMayBeFolded bit
on the corresponding MachineInstr operand.
RegAllocFast::selectInlineAsmOperandsToFold() reads that per-operand bit
to estimate register-class pressure before deciding what to fold. Since
the tied input is a separate virtual register from its def (a fresh vreg
created in SelectionDAGBuilder to carry the input value into the tied
slot), the missing propagation made the estimator treat the pair as two
independent demand units -- one correctly foldable (the def), one
incorrectly "must have a real register" (the tied input) -- when the two
are always assigned the same physical register and so only ever need
one register between them. That inflated the class's non-foldable
demand and could force a fold even when exactly enough registers were
available, e.g.:
%r = tail call i32 asm sideeffect "# tied: $0", "=rm,0,~{eax},~{ecx},
~{edx},~{edi},~{ebx},~{ebp},~{r8d},~{r9d},~{r10d},~{r11d},~{r12d},
~{r13d},~{r14d},~{r15d}"(i32 %x)
With exactly one GR32 register left unclobbered, `llc -O2` (greedy)
correctly keeps the value in that register; `llc -O2 -regalloc=fast`
spilled it to a stack slot instead, unnecessarily.
Fix in two parts:
- TargetLowering.cpp's tied-operand hookup loop in ParseConstraints now
propagates MayFoldRegister from a matched output to its tied input,
so every other current or future consumer of the flag (e.g.
CalcSpillWeights.cpp's canMemFoldInlineAsm) sees a consistent answer
for both halves of the pair.
- That alone isn't enough: with both halves now correctly flagged,
selectInlineAsmOperandsToFold() would count them as two foldable
units instead of one. It now skips the tied-use side of a tied pair
when accumulating per-class demand, counting the pair once via its
def, matching the one physical register the pair actually needs.
Verified via the repro above (now keeps the register at exactly enough
availability, still correctly folds at zero availability) and a full
regression run: llvm/test/CodeGen/{X86,AArch64,MIR,RISCV,ARM}/ and
clang/test/CodeGen/ (20,000+ tests, all passing modulo pre-existing
UNSUPPORTED/XFAIL).
---
llvm/include/llvm/CodeGen/TargetLowering.h | 13 ++++++++-----
llvm/lib/CodeGen/RegAllocFast.cpp | 9 +++++++++
llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp | 11 +++++++++++
3 files changed, 28 insertions(+), 5 deletions(-)
diff --git a/llvm/include/llvm/CodeGen/TargetLowering.h b/llvm/include/llvm/CodeGen/TargetLowering.h
index c8760a69f8135..31cdf0426751b 100644
--- a/llvm/include/llvm/CodeGen/TargetLowering.h
+++ b/llvm/include/llvm/CodeGen/TargetLowering.h
@@ -5351,11 +5351,14 @@ class LLVM_ABI TargetLowering : public TargetLoweringBase {
MVT ConstraintVT = MVT::Other;
/// True if this operand's constraint codes are exactly {"r", "m"} (the
- /// "rm" constraint, or the tied output half of "+rm"). getConstraintType
- /// preference selection uses this to opt for 'r' while still allowing
- /// the register allocator to fall back to 'm' under register pressure,
- /// instead of picking 'm' unconditionally as it would for a generic
- /// multi-alternative constraint.
+ /// "rm" constraint), or -- for the tied input half of a "+rm" pair,
+ /// whose own constraint codes are just the matching digit (e.g. "0") --
+ /// if the output operand it's tied to has this set (see
+ /// ParseConstraints()'s tied-operand hookup loop, which propagates it
+ /// there). getConstraintType preference selection uses this to opt for 'r'
+ /// while still allowing the register allocator to fall back to 'm' under
+ /// register pressure, instead of picking 'm' unconditionally as it would
+ /// for a generic multi-alternative constraint.
bool MayFoldRegister = false;
/// Copy constructor for copying from a ConstraintInfo.
diff --git a/llvm/lib/CodeGen/RegAllocFast.cpp b/llvm/lib/CodeGen/RegAllocFast.cpp
index 0a0007eaf14dc..2a71bf4cee1a1 100644
--- a/llvm/lib/CodeGen/RegAllocFast.cpp
+++ b/llvm/lib/CodeGen/RegAllocFast.cpp
@@ -1872,6 +1872,15 @@ void RegAllocFastImpl::selectInlineAsmOperandsToFold(
Blocked.push_back(Reg.asMCReg());
continue;
}
+ // A tied "+rm" pair (def + matching input) is one memory location
+ // shared between two different virtual registers once folded (see
+ // foldFoldableInlineAsmOperands()'s TiedUse handling below), and the
+ // register allocator must assign them the same physical register
+ // either way -- it's one unit of demand, not two. Count it once, via
+ // the def side; skip the tied use here so it isn't double-counted
+ // against the same register class.
+ if (MO.isTied() && MO.isUse())
+ continue;
Demand &D = DemandByClass[MRI->getRegClass(Reg)];
if (MI.mayFoldInlineAsmRegOp(I))
D.Foldable.push_back(Reg);
diff --git a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
index 81e8052be9309..de80e6da54383 100644
--- a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
@@ -6281,6 +6281,17 @@ TargetLowering::ParseConstraints(const DataLayout &DL,
if (OpInfo.hasMatchingInput()) {
AsmOperandInfo &Input = ConstraintOperands[OpInfo.MatchingInput];
+ // The matching input's own constraint codes are just the matching
+ // digit (e.g. "0"), never {"r","m"}, so the {r,m}-exact-match check
+ // above that sets MayFoldRegister never fires for it directly. A tied
+ // "+rm" pair is one memory location shared between the def and its
+ // input once folded (see RegAllocFast::foldFoldableInlineAsmOperands's
+ // TiedUse handling and InlineSpiller's equivalent untie-then-recurse
+ // fold), so whether the pair may fold is really a property of the
+ // output side; propagate it here rather than leaving every other
+ // MayFoldRegister/RegMayBeFolded consumer to special-case tied inputs.
+ Input.MayFoldRegister = OpInfo.MayFoldRegister;
+
if (OpInfo.ConstraintVT != Input.ConstraintVT) {
std::pair<unsigned, const TargetRegisterClass *> MatchRC =
getRegForInlineAsmConstraint(TRI, OpInfo.ConstraintCode,
More information about the cfe-commits
mailing list