[llvm] [AMDGPU] Add register allocation handoff intrinsic (PR #219842)
via llvm-commits
llvm-commits at lists.llvm.org
Sun Aug 30 12:52:53 PDT 2026
llvmorg-github-actions[bot] wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-llvm-ir
Author: Bangtian Liu (bangtianliu)
<details>
<summary>Changes</summary>
Add llvm.amdgcn.regalloc.handoff to carry register allocation requirements through SelectionDAG and GlobalISel.
Preserve nomerge semantics across IR and machine optimizations, teach the AMDGPU Attributor and register-bank pipeline about the handoff, and add end-to-end regression coverage.
Fixes #<!-- -->219377
---
Patch is 86.28 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/219842.diff
45 Files Affected:
- (modified) llvm/docs/AMDGPUUsage.rst (+33)
- (modified) llvm/include/llvm/IR/IntrinsicsAMDGPU.td (+8)
- (modified) llvm/lib/CodeGen/BranchFolding.cpp (+4)
- (modified) llvm/lib/CodeGen/EarlyIfConversion.cpp (+4)
- (modified) llvm/lib/CodeGen/MIRParser/MILexer.cpp (+1)
- (modified) llvm/lib/CodeGen/MIRParser/MILexer.h (+1)
- (modified) llvm/lib/CodeGen/MIRParser/MIParser.cpp (+3)
- (modified) llvm/lib/CodeGen/MachineCSE.cpp (+1-1)
- (modified) llvm/lib/CodeGen/MachineLICM.cpp (+7-1)
- (modified) llvm/lib/Target/AMDGPU/AMDGPUAttributor.cpp (+6)
- (modified) llvm/lib/Target/AMDGPU/AMDGPUISelDAGToDAG.cpp (+50)
- (modified) llvm/lib/Target/AMDGPU/AMDGPUInstructionSelector.cpp (+55)
- (modified) llvm/lib/Target/AMDGPU/AMDGPURegBankLegalizeRules.cpp (+3)
- (modified) llvm/lib/Target/AMDGPU/AMDGPURegBankSelect.cpp (+96-1)
- (modified) llvm/lib/Target/AMDGPU/AMDGPURegisterBankInfo.cpp (+39)
- (modified) llvm/lib/Target/AMDGPU/SIISelLowering.cpp (+50)
- (modified) llvm/lib/Target/AMDGPU/SIInstrInfo.cpp (+11)
- (modified) llvm/lib/Target/AMDGPU/SIInstructions.td (+13)
- (modified) llvm/lib/Transforms/Scalar/EarlyCSE.cpp (+4-1)
- (modified) llvm/lib/Transforms/Scalar/GVN.cpp (+5)
- (modified) llvm/lib/Transforms/Scalar/GVNSink.cpp (+5)
- (modified) llvm/lib/Transforms/Scalar/NewGVN.cpp (+3)
- (added) llvm/test/CodeGen/AMDGPU/GlobalISel/inst-select-amdgcn.regalloc.handoff.mir (+74)
- (added) llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.regalloc.handoff-hard-vgpr.ll (+19)
- (added) llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.regalloc.handoff-phi-loop.ll (+101)
- (added) llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.regalloc.handoff-phi.ll (+95)
- (added) llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.regalloc.handoff-unsupported.ll (+24)
- (added) llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.regalloc.handoff.ll (+66)
- (added) llvm/test/CodeGen/AMDGPU/GlobalISel/regbankselect-amdgcn.regalloc.handoff.mir (+194)
- (modified) llvm/test/CodeGen/AMDGPU/amdgpu-attributor-min-agpr-alloc.ll (+42)
- (added) llvm/test/CodeGen/AMDGPU/amdgpu-attributor-regalloc-handoff.ll (+21)
- (added) llvm/test/CodeGen/AMDGPU/branch-folder-nomerge.mir (+92)
- (added) llvm/test/CodeGen/AMDGPU/early-ifcvt-nomerge.mir (+99)
- (added) llvm/test/CodeGen/AMDGPU/llvm.amdgcn.regalloc.handoff.ll (+112)
- (added) llvm/test/CodeGen/AMDGPU/machine-cse-nomerge.mir (+20)
- (added) llvm/test/CodeGen/AMDGPU/machine-licm-nomerge.mir (+78)
- (added) llvm/test/CodeGen/AMDGPU/regalloc-handoff-machine-passes.mir (+32)
- (added) llvm/test/CodeGen/AMDGPU/regalloc-handoff-postra.mir (+55)
- (added) llvm/test/CodeGen/AMDGPU/regalloc-handoff-stress-regalloc.mir (+109)
- (added) llvm/test/Transforms/EarlyCSE/nomerge.ll (+31)
- (added) llvm/test/Transforms/GVN/nomerge.ll (+17)
- (added) llvm/test/Transforms/GVNSink/nomerge.ll (+42)
- (added) llvm/test/Transforms/InstCombine/AMDGPU/regalloc-handoff.ll (+23)
- (added) llvm/test/Transforms/NewGVN/nomerge.ll (+17)
- (added) llvm/test/Verifier/AMDGPU/regalloc-handoff.ll (+40)
``````````diff
diff --git a/llvm/docs/AMDGPUUsage.rst b/llvm/docs/AMDGPUUsage.rst
index 40a2820b9e04f..3a3c32dfb1ae0 100644
--- a/llvm/docs/AMDGPUUsage.rst
+++ b/llvm/docs/AMDGPUUsage.rst
@@ -2120,6 +2120,39 @@ The AMDGPU backend implements the following LLVM IR intrinsics.
List AMDGPU intrinsics.
+.. _amdgpu-regalloc-handoff:
+
+'``llvm.amdgcn.regalloc.handoff``' Intrinsic
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+The ``llvm.amdgcn.regalloc.handoff`` intrinsic creates a local register
+allocation boundary around one 32-bit value:
+
+.. code-block:: llvm
+
+ i32 @llvm.amdgcn.regalloc.handoff(i32 %value, i32 immarg %mode)
+
+The result is equal to ``%value``, but the incoming and outgoing values have
+separate register-allocation intervals. The immediate mode controls outgoing
+register allocation:
+
+* ``0`` uses automatic AV outgoing allocation with a source allocation hint,
+ preferring the incoming value's allocation. Scalar and otherwise unhintable
+ values follow the normal AV allocation policy.
+* ``1`` selects VGPR.
+* ``2`` selects AGPR. This mode is unsupported on targets without matrix
+ instruction support.
+
+The intrinsic is pure. An unused call can be removed, and a live call can move
+like an ordinary data-dependent operation, but separate live occurrences are
+not merged or duplicated. It is not a scheduling, memory, inline-assembly, or
+``EXEC`` barrier.
+
+Register allocation may assign compatible incoming and outgoing intervals to
+the same physical register, in which case the handoff emits no instruction.
+Otherwise post-register-allocation expansion uses the normal AMDGPU physical
+copy lowering, including the required VGPR/AGPR transition when banks differ.
+
.. _amdgpu-av-load-store:
'``llvm.amdgcn.av``' Intrinsics
diff --git a/llvm/include/llvm/IR/IntrinsicsAMDGPU.td b/llvm/include/llvm/IR/IntrinsicsAMDGPU.td
index 037630ea86904..936ed86e22647 100644
--- a/llvm/include/llvm/IR/IntrinsicsAMDGPU.td
+++ b/llvm/include/llvm/IR/IntrinsicsAMDGPU.td
@@ -159,6 +159,14 @@ let TargetPrefix = "amdgcn" in {
// ABI Special Intrinsics
//===----------------------------------------------------------------------===//
+// Preserve a local register-allocation boundary around a 32-bit value. The
+// mode is immediate: 0 uses automatic AV allocation with a source allocation
+// preference, 1 selects VGPR, and 2 selects AGPR.
+def int_amdgcn_regalloc_handoff : PureIntrinsic<
+ [llvm_i32_ty], [llvm_i32_ty, llvm_i32_ty],
+ [ImmArg<ArgIndex<1>>, Range<ArgIndex<1>, 0, 3>,
+ IntrNoMerge, IntrNoDuplicate]>;
+
defm int_amdgcn_workitem_id
: AMDGPUReadPreloadRegisterIntrinsic_xyz_named<
"__builtin_amdgcn_workitem_id", [Range<RetIndex, 0, 1024>]>;
diff --git a/llvm/lib/CodeGen/BranchFolding.cpp b/llvm/lib/CodeGen/BranchFolding.cpp
index eb36cad8ab3db..ce5f1578a8a96 100644
--- a/llvm/lib/CodeGen/BranchFolding.cpp
+++ b/llvm/lib/CodeGen/BranchFolding.cpp
@@ -2057,6 +2057,10 @@ bool BranchFolder::HoistCommonCodeInSuccs(MachineBasicBlock *MBB) {
if (TIB == TIE || FIB == FIE)
break;
+ if (TIB->getFlag(MachineInstr::NoMerge) ||
+ FIB->getFlag(MachineInstr::NoMerge))
+ break;
+
if (!TIB->isIdenticalTo(*FIB, MachineInstr::CheckKillDead))
break;
diff --git a/llvm/lib/CodeGen/EarlyIfConversion.cpp b/llvm/lib/CodeGen/EarlyIfConversion.cpp
index bf2664fd4225c..2175ce53f39bd 100644
--- a/llvm/lib/CodeGen/EarlyIfConversion.cpp
+++ b/llvm/lib/CodeGen/EarlyIfConversion.cpp
@@ -610,6 +610,10 @@ static bool hasSameValue(const MachineRegisterInfo &MRI,
if (!TDef || !FDef)
return false;
+ if (TDef->getFlag(MachineInstr::NoMerge) ||
+ FDef->getFlag(MachineInstr::NoMerge))
+ return false;
+
// If there are side-effects, all bets are off.
if (TDef->hasUnmodeledSideEffects())
return false;
diff --git a/llvm/lib/CodeGen/MIRParser/MILexer.cpp b/llvm/lib/CodeGen/MIRParser/MILexer.cpp
index 67fbc00edd9db..d3aa96fe94c24 100644
--- a/llvm/lib/CodeGen/MIRParser/MILexer.cpp
+++ b/llvm/lib/CodeGen/MIRParser/MILexer.cpp
@@ -299,6 +299,7 @@ static MIToken::TokenKind getIdentifierKind(StringRef Identifier) {
.Case("machine-block-address-taken",
MIToken::kw_machine_block_address_taken)
.Case("call-frame-size", MIToken::kw_call_frame_size)
+ .Case("nomerge", MIToken::kw_nomerge)
.Case("noconvergent", MIToken::kw_noconvergent)
.Case("mmra", MIToken::kw_mmra)
.Case("lr-split", MIToken::kw_lr_split)
diff --git a/llvm/lib/CodeGen/MIRParser/MILexer.h b/llvm/lib/CodeGen/MIRParser/MILexer.h
index f5947bfe59b9d..fd50316a662e1 100644
--- a/llvm/lib/CodeGen/MIRParser/MILexer.h
+++ b/llvm/lib/CodeGen/MIRParser/MILexer.h
@@ -151,6 +151,7 @@ struct MIToken {
kw_ir_block_address_taken,
kw_machine_block_address_taken,
kw_call_frame_size,
+ kw_nomerge,
kw_noconvergent,
kw_mmra,
kw_lr_split,
diff --git a/llvm/lib/CodeGen/MIRParser/MIParser.cpp b/llvm/lib/CodeGen/MIRParser/MIParser.cpp
index bb0b87cc042d0..77ddcb6c0ccbc 100644
--- a/llvm/lib/CodeGen/MIRParser/MIParser.cpp
+++ b/llvm/lib/CodeGen/MIRParser/MIParser.cpp
@@ -1530,6 +1530,7 @@ bool MIParser::parseInstruction(unsigned &OpCode, unsigned &Flags) {
Token.is(MIToken::kw_nsw) ||
Token.is(MIToken::kw_exact) ||
Token.is(MIToken::kw_nofpexcept) ||
+ Token.is(MIToken::kw_nomerge) ||
Token.is(MIToken::kw_noconvergent) ||
Token.is(MIToken::kw_unpredictable) ||
Token.is(MIToken::kw_nneg) ||
@@ -1566,6 +1567,8 @@ bool MIParser::parseInstruction(unsigned &OpCode, unsigned &Flags) {
Flags |= MachineInstr::IsExact;
if (Token.is(MIToken::kw_nofpexcept))
Flags |= MachineInstr::NoFPExcept;
+ if (Token.is(MIToken::kw_nomerge))
+ Flags |= MachineInstr::NoMerge;
if (Token.is(MIToken::kw_unpredictable))
Flags |= MachineInstr::Unpredictable;
if (Token.is(MIToken::kw_noconvergent))
diff --git a/llvm/lib/CodeGen/MachineCSE.cpp b/llvm/lib/CodeGen/MachineCSE.cpp
index 3d0ac171b9fe4..e11f49715af9a 100644
--- a/llvm/lib/CodeGen/MachineCSE.cpp
+++ b/llvm/lib/CodeGen/MachineCSE.cpp
@@ -395,7 +395,7 @@ bool MachineCSEImpl::PhysRegDefsReach(MachineInstr *CSMI, MachineInstr *MI,
bool MachineCSEImpl::isCSECandidate(MachineInstr *MI) {
if (MI->isPosition() || MI->isPHI() || MI->isImplicitDef() || MI->isKill() ||
MI->isInlineAsm() || MI->isDebugInstr() || MI->isJumpTableDebugInfo() ||
- MI->isFakeUse())
+ MI->isFakeUse() || MI->getFlag(MachineInstr::NoMerge))
return false;
// Ignore copies.
diff --git a/llvm/lib/CodeGen/MachineLICM.cpp b/llvm/lib/CodeGen/MachineLICM.cpp
index 6435bbb9382c1..9a8a90c830025 100644
--- a/llvm/lib/CodeGen/MachineLICM.cpp
+++ b/llvm/lib/CodeGen/MachineLICM.cpp
@@ -1507,9 +1507,15 @@ void MachineLICMImpl::InitializeLoadsHoistableLoops() {
MachineInstr *
MachineLICMImpl::LookForDuplicate(const MachineInstr *MI,
std::vector<MachineInstr *> &PrevMIs) {
- for (MachineInstr *PrevMI : PrevMIs)
+ if (MI->getFlag(MachineInstr::NoMerge))
+ return nullptr;
+
+ for (MachineInstr *PrevMI : PrevMIs) {
+ if (PrevMI->getFlag(MachineInstr::NoMerge))
+ continue;
if (TII->produceSameValue(*MI, *PrevMI, (PreRegAlloc ? MRI : nullptr)))
return PrevMI;
+ }
return nullptr;
}
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUAttributor.cpp b/llvm/lib/Target/AMDGPU/AMDGPUAttributor.cpp
index 630ffad96e451..a2c7e03d046a2 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUAttributor.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUAttributor.cpp
@@ -1360,6 +1360,12 @@ struct AAAMDGPUMinAGPRAlloc
return true;
}
+ case Intrinsic::amdgcn_regalloc_handoff:
+ // Mode 2 forces the result into an AGPR, but the number of mode 2
+ // values simultaneously live during register allocation is unknown
+ // here. Do not manifest a bounded AGPR allocation for this function
+ // or its callers.
+ return !cast<ConstantInt>(CB.getArgOperand(1))->equalsInt(2);
// Trap-like intrinsics such as llvm.trap and llvm.debugtrap do not have
// the nocallback attribute, so the AMDGPU attributor can conservatively
// drop all implicitly-known inputs and AGPR allocation information. Make
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUISelDAGToDAG.cpp b/llvm/lib/Target/AMDGPU/AMDGPUISelDAGToDAG.cpp
index 744bde043e16c..e350bb97e391b 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUISelDAGToDAG.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUISelDAGToDAG.cpp
@@ -26,6 +26,7 @@
#include "llvm/CodeGen/SelectionDAG.h"
#include "llvm/CodeGen/SelectionDAGISel.h"
#include "llvm/CodeGen/SelectionDAGNodes.h"
+#include "llvm/IR/DiagnosticInfo.h"
#include "llvm/IR/IntrinsicsAMDGPU.h"
#include "llvm/InitializePasses.h"
#include "llvm/Support/ErrorHandling.h"
@@ -697,6 +698,13 @@ void AMDGPUDAGToDAGISel::Select(SDNode *N) {
switch (Opc) {
default:
break;
+ case ISD::SRCVALUE:
+ // A regalloc handoff retains this non-emitted identity operand on its
+ // selected machine node to prevent MorphNodeTo from CSE'ing distinct IR
+ // calls. Morph the same value to a non-emitted machine node. The unused
+ // glue result makes the machine node unique as well.
+ CurDAG->SelectNodeTo(N, TargetOpcode::IMPLICIT_DEF, MVT::Other, MVT::Glue);
+ return;
case ISD::UADDO_CARRY:
case ISD::USUBO_CARRY:
if (N->getValueType(0) == MVT::i64) {
@@ -3303,6 +3311,48 @@ void AMDGPUDAGToDAGISel::SelectINTRINSIC_WO_CHAIN(SDNode *N) {
ConvGlueNode = nullptr;
}
switch (IntrID) {
+ case Intrinsic::amdgcn_regalloc_handoff: {
+ unsigned Mode = N->getConstantOperandVal(2);
+ unsigned HandoffOpcode;
+ switch (Mode) {
+ case 0:
+ HandoffOpcode = AMDGPU::REGALLOC_HANDOFF_AV;
+ break;
+ case 1:
+ HandoffOpcode = AMDGPU::REGALLOC_HANDOFF_VGPR;
+ break;
+ case 2:
+ if (!Subtarget->hasMAIInsts()) {
+ const Function &Fn = CurDAG->getMachineFunction().getFunction();
+ Fn.getContext().diagnose(DiagnosticInfoUnsupported(
+ Fn,
+ "llvm.amdgcn.regalloc.handoff mode 2 requires matrix instruction "
+ "support",
+ N->getDebugLoc()));
+ HandoffOpcode = AMDGPU::REGALLOC_HANDOFF_VGPR;
+ } else {
+ HandoffOpcode = AMDGPU::REGALLOC_HANDOFF_AGPR;
+ }
+ break;
+ default:
+ llvm_unreachable("invalid llvm.amdgcn.regalloc.handoff mode");
+ }
+
+ SDLoc DL(N);
+ SDValue AVClass =
+ CurDAG->getTargetConstant(AMDGPU::AV_32RegClassID, DL, MVT::i32);
+ SDValue Src(CurDAG->getMachineNode(TargetOpcode::COPY_TO_REGCLASS, DL,
+ N->getValueType(0), N->getOperand(1),
+ AVClass),
+ 0);
+ SmallVector<SDValue, 3> Ops = {Src, N->getOperand(3)};
+ if (ConvGlueNode)
+ Ops.push_back(SDValue(ConvGlueNode, 0));
+ SDNode *Selected =
+ CurDAG->SelectNodeTo(N, HandoffOpcode, N->getVTList(), Ops);
+ CurDAG->addNoMergeSiteInfo(Selected, true);
+ return;
+ }
case Intrinsic::amdgcn_wqm:
Opcode = AMDGPU::WQM;
break;
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUInstructionSelector.cpp b/llvm/lib/Target/AMDGPU/AMDGPUInstructionSelector.cpp
index 353506378b229..577bb65cacb2e 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUInstructionSelector.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUInstructionSelector.cpp
@@ -1241,6 +1241,61 @@ bool AMDGPUInstructionSelector::selectDivScale(MachineInstr &MI) const {
bool AMDGPUInstructionSelector::selectG_INTRINSIC(MachineInstr &I) const {
Intrinsic::ID IntrinsicID = cast<GIntrinsic>(I).getIntrinsicID();
switch (IntrinsicID) {
+ case Intrinsic::amdgcn_regalloc_handoff: {
+ if (I.getNumOperands() != 4 || !I.getOperand(0).isReg() ||
+ !I.getOperand(2).isReg() || !I.getOperand(3).isImm())
+ return false;
+
+ unsigned Opcode;
+ switch (I.getOperand(3).getImm()) {
+ case 0:
+ Opcode = AMDGPU::REGALLOC_HANDOFF_AV;
+ break;
+ case 1:
+ Opcode = AMDGPU::REGALLOC_HANDOFF_VGPR;
+ break;
+ case 2:
+ if (!STI.hasMAIInsts()) {
+ const Function &Fn = I.getMF()->getFunction();
+ Fn.getContext().diagnose(DiagnosticInfoUnsupported(
+ Fn,
+ "llvm.amdgcn.regalloc.handoff mode 2 requires matrix instruction "
+ "support",
+ I.getDebugLoc()));
+ Opcode = AMDGPU::REGALLOC_HANDOFF_VGPR;
+ break;
+ }
+ Opcode = AMDGPU::REGALLOC_HANDOFF_AGPR;
+ break;
+ default:
+ return false;
+ }
+
+ bool IsAuto = Opcode == AMDGPU::REGALLOC_HANDOFF_AV;
+ Register Dst = I.getOperand(0).getReg();
+ if (IsAuto) {
+ const RegisterBank *DstRB = RBI.getRegBank(Dst, *MRI, TRI);
+ const TargetRegisterClass &DstRC = DstRB->getID() == AMDGPU::AGPRRegBankID
+ ? AMDGPU::AGPR_32RegClass
+ : AMDGPU::VGPR_32RegClass;
+ if (!RBI.constrainGenericRegister(Dst, DstRC, *MRI))
+ return false;
+ }
+
+ MachineInstrBuilder Handoff =
+ BuildMI(*I.getParent(), &I, I.getDebugLoc(), TII.get(Opcode))
+ .addDef(Dst)
+ .add(I.getOperand(2));
+ Handoff->setFlag(MachineInstr::NoMerge);
+ constrainSelectedInstRegOperands(*Handoff, TII, TRI, RBI);
+ I.eraseFromParent();
+ // RegBankSelect needed a concrete bank for the generic definition. Widen
+ // it back to AV_32 unless an already-selected use requires a narrower
+ // class.
+ if (IsAuto)
+ MRI->recomputeRegClass(Dst);
+ return true;
+ }
case Intrinsic::amdgcn_if_break: {
MachineBasicBlock *BB = I.getParent();
diff --git a/llvm/lib/Target/AMDGPU/AMDGPURegBankLegalizeRules.cpp b/llvm/lib/Target/AMDGPU/AMDGPURegBankLegalizeRules.cpp
index 0072ef76464cc..3993800132d4d 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPURegBankLegalizeRules.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPURegBankLegalizeRules.cpp
@@ -1777,6 +1777,9 @@ RegBankLegalizeRules::RegBankLegalizeRules(const GCNSubtarget &_ST,
addRulesForIOpcs({returnaddress}).Any({{UniP0}, {{SgprP0}, {}}});
+ addRulesForIOpcs({amdgcn_regalloc_handoff})
+ .Any({{B32, _, B32, _}, {{None}, {IntrId, None, Imm}}});
+
addRulesForIOpcs({amdgcn_s_getpc}).Any({{UniS64, _}, {{Sgpr64}, {}}});
addRulesForIOpcs({amdgcn_s_getreg}).Any({{}, {{Sgpr32}, {IntrId}}});
diff --git a/llvm/lib/Target/AMDGPU/AMDGPURegBankSelect.cpp b/llvm/lib/Target/AMDGPU/AMDGPURegBankSelect.cpp
index f55f183b92eee..43c62b71d4187 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPURegBankSelect.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPURegBankSelect.cpp
@@ -18,11 +18,15 @@
#include "AMDGPU.h"
#include "AMDGPUGlobalISelUtils.h"
#include "GCNSubtarget.h"
+#include "llvm/ADT/SmallVector.h"
#include "llvm/CodeGen/GlobalISel/CSEInfo.h"
#include "llvm/CodeGen/GlobalISel/CSEMIRBuilder.h"
+#include "llvm/CodeGen/GlobalISel/GenericMachineInstrs.h"
#include "llvm/CodeGen/MachineUniformityAnalysis.h"
#include "llvm/CodeGen/TargetPassConfig.h"
+#include "llvm/IR/IntrinsicsAMDGPU.h"
#include "llvm/InitializePasses.h"
+#include <utility>
#define DEBUG_TYPE "amdgpu-reg-bank-select"
@@ -81,6 +85,7 @@ class RegBankSelectHelper {
AMDGPU::IntrinsicLaneMaskAnalyzer &ILMA;
const MachineUniformityInfo &MUI;
const SIRegisterInfo &TRI;
+ const RegisterBankInfo &RBI;
const RegisterBank *SgprRB;
const RegisterBank *VgprRB;
const RegisterBank *VccRB;
@@ -90,7 +95,7 @@ class RegBankSelectHelper {
AMDGPU::IntrinsicLaneMaskAnalyzer &ILMA,
const MachineUniformityInfo &MUI,
const SIRegisterInfo &TRI, const RegisterBankInfo &RBI)
- : B(B), MRI(*B.getMRI()), ILMA(ILMA), MUI(MUI), TRI(TRI),
+ : B(B), MRI(*B.getMRI()), ILMA(ILMA), MUI(MUI), TRI(TRI), RBI(RBI),
SgprRB(&RBI.getRegBank(AMDGPU::SGPRRegBankID)),
VgprRB(&RBI.getRegBank(AMDGPU::VGPRRegBankID)),
VccRB(&RBI.getRegBank(AMDGPU::VCCRegBankID)) {}
@@ -183,6 +188,71 @@ class RegBankSelectHelper {
B.buildCopy(NewReg, Reg);
}
+
+ bool isUniformPhiHandoffUse(MachineInstr &MI, MachineOperand &UseOP) {
+ assert(MI.isPHI());
+ Register PhiReg = MI.getOperand(0).getReg();
+ Register UseReg = UseOP.getReg();
+ if (MRI.getRegBankOrNull(PhiReg) != SgprRB)
+ return false;
+
+ auto *GI = dyn_cast_or_null<GIntrinsic>(MRI.getVRegDef(UseReg));
+ return GI && GI->getIntrinsicID() == Intrinsic::amdgcn_regalloc_handoff;
+ }
+
+ void reconcileUniformPhiHandoffUse(MachineInstr &MI, MachineOperand &UseOP) {
+ assert(isUniformPhiHandoffUse(MI, UseOP));
+ Register UseReg = UseOP.getReg();
+ const RegisterBank *UseRB = MRI.getRegBankOrNull(UseReg);
+ assert(UseRB && UseRB != SgprRB);
+
+ MachineBasicBlock &IncomingMBB =
+ *MI.getOperand(UseOP.getOperandNo() + 1).getMBB();
+ B.setInsertPt(IncomingMBB, IncomingMBB.getFirstTerminator());
+
+ LLT Ty = MRI.getType(UseReg);
+ Register VgprUse = UseReg;
+ if (UseRB->getID() == AMDGPU::AGPRRegBankID) {
+ VgprUse = MRI.createVirtualRegister({VgprRB, Ty});
+ B.buildCopy(VgprUse, UseReg);
+ } else {
+ assert(UseRB == VgprRB);
+ }
+
+ Register SgprUse = MRI.createVirtualRegister({SgprRB, Ty});
+ buildReadFirstLane(B, SgprUse, VgprUse, RBI);
+ UseOP.setReg(SgprUse);
+ }
+
+ bool applyRegallocHandoffMapping(
+ MachineInstr &MI, const RegisterBankInfo::InstructionMapping &Mapping) {
+ if (!Mapping.isValid() || Mapping.getNumOperands() != MI.getNumOperands())
+ return false;
+
+ for (unsigned OpIdx : {0u, 2u}) {
+ MachineOperand &Op = MI.getOperand(OpIdx);
+ const RegisterBankInfo::ValueMapping &ValMapping =
+ Mapping.getOperandMapping(OpIdx);
+ if (!Op.isReg() || !Op.getReg().isVirtual() || !ValMapping.isValid() ||
+ ValMapping.NumBreakDowns != 1)
+ return false;
+
+ const RegisterBank *RB = ValMapping.BreakDown[0].RegBank;
+ Register Reg = Op.getReg();
+ if (Op.isDef()) {
+ if (MRI.getRegClassOrNull(Reg))
+ reAssignRegBankOnDef(MI, Op, RB);
+ else {
+ assert(!MRI.getRegBankOrNull(Reg));
+ MRI.setRegBank(Reg, *RB);
+ }
+ } else if (RBI.getRegBank(Reg, MRI, TRI) != RB) {
+ constrainRegBankUse(MI, Op, RB);
+ }
+ }
+
+ return true;
+ }
};
static Register getVReg(MachineOperand &Op) {
@@ -223,6 +293,7 @@ bool AMDGPURegBankSelect::runOnMachineFunction(MachineFunction &MF) {
const GCNSubtarget &ST = MF.getSubtarget<GCNSubtarget>();
RegBankSelectHelper RBSHelper(B, ILMA, MUI, *ST.getRegisterInfo(),
*ST.getRegBankInfo());
+ SmallVector<std::pair<MachineInstr *, unsigned>, 4> UniformPhiHandoffUses;
// Virtual registers at this point don't have register banks.
// Virtual registers in def and use operands of already inst-selected
// instruction have register class.
@@ -244,6 +315,17 @@ bool AMDGPURegBankSelect::runOnMachineFunction(MachineFunction &MF) {
if (!MI.isPreISelOpcode())
continue;
+ if (auto *GI = dyn_cast<GIntrinsic>(&MI);
+ GI && GI->getIntrinsicID() == Intrinsic::amdgcn_regalloc_handoff) {
+ const RegisterBankInfo::InstructionMapping &Mapping =
+ ST.getRegBankInfo()->getInstrMapping(MI);
+ if (!RBSHelper.applyRegallocHandoffMapping(MI, Mapping)) {
+ MF.getProperties().setFailedISel();
+ return true;
+ }
+ continue;
+ }
+
// Vregs in def and use operands of G_ instructions need to have register
// banks assigned. Before this loop possible case are
// - (1) vreg without register class or bank in def or use operand
@@ -275,6 +357,11 @@ bool AMDGPURegBankSelect::runOnMachineFunction(MachineFunction &MF) {
if (!UseReg.isValid())
continue;
+ if (MI.isPHI() && RBSHelper.isUniformPhiHandoffUse(MI, UseOP)) {
+ UniformPhiHandoffUses.emplace_back(&MI, UseOP.getOperandNo());
+ continue;
+ }
+
// Skip case (3).
if (!MRI.getRegClassOrNull(UseReg) ||
MRI.getVRegDef(UseReg)->isPreISelOpcode())
@@ -287,5 +374,13 @@ bool AMDGPURegBankSelect::runOnMachineFunction(MachineFunction &MF) {
}
...
[truncated]
``````````
</details>
https://github.com/llvm/llvm-project/pull/219842
More information about the llvm-commits
mailing list