[llvm] [AMDGPU] Add register allocation handoff intrinsic (PR #219842)

via llvm-commits llvm-commits at lists.llvm.org
Sun Aug 30 12:52:53 PDT 2026


llvmorg-github-actions[bot] wrote:


<!--LLVM PR SUMMARY COMMENT-->

@llvm/pr-subscribers-llvm-ir

Author: Bangtian Liu (bangtianliu)

<details>
<summary>Changes</summary>

Add llvm.amdgcn.regalloc.handoff to carry register allocation requirements through SelectionDAG and GlobalISel.

Preserve nomerge semantics across IR and machine optimizations, teach the AMDGPU Attributor and register-bank pipeline about the handoff, and add end-to-end regression coverage.

Fixes #<!-- -->219377

---

Patch is 86.28 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/219842.diff


45 Files Affected:

- (modified) llvm/docs/AMDGPUUsage.rst (+33) 
- (modified) llvm/include/llvm/IR/IntrinsicsAMDGPU.td (+8) 
- (modified) llvm/lib/CodeGen/BranchFolding.cpp (+4) 
- (modified) llvm/lib/CodeGen/EarlyIfConversion.cpp (+4) 
- (modified) llvm/lib/CodeGen/MIRParser/MILexer.cpp (+1) 
- (modified) llvm/lib/CodeGen/MIRParser/MILexer.h (+1) 
- (modified) llvm/lib/CodeGen/MIRParser/MIParser.cpp (+3) 
- (modified) llvm/lib/CodeGen/MachineCSE.cpp (+1-1) 
- (modified) llvm/lib/CodeGen/MachineLICM.cpp (+7-1) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPUAttributor.cpp (+6) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPUISelDAGToDAG.cpp (+50) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPUInstructionSelector.cpp (+55) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPURegBankLegalizeRules.cpp (+3) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPURegBankSelect.cpp (+96-1) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPURegisterBankInfo.cpp (+39) 
- (modified) llvm/lib/Target/AMDGPU/SIISelLowering.cpp (+50) 
- (modified) llvm/lib/Target/AMDGPU/SIInstrInfo.cpp (+11) 
- (modified) llvm/lib/Target/AMDGPU/SIInstructions.td (+13) 
- (modified) llvm/lib/Transforms/Scalar/EarlyCSE.cpp (+4-1) 
- (modified) llvm/lib/Transforms/Scalar/GVN.cpp (+5) 
- (modified) llvm/lib/Transforms/Scalar/GVNSink.cpp (+5) 
- (modified) llvm/lib/Transforms/Scalar/NewGVN.cpp (+3) 
- (added) llvm/test/CodeGen/AMDGPU/GlobalISel/inst-select-amdgcn.regalloc.handoff.mir (+74) 
- (added) llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.regalloc.handoff-hard-vgpr.ll (+19) 
- (added) llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.regalloc.handoff-phi-loop.ll (+101) 
- (added) llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.regalloc.handoff-phi.ll (+95) 
- (added) llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.regalloc.handoff-unsupported.ll (+24) 
- (added) llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.regalloc.handoff.ll (+66) 
- (added) llvm/test/CodeGen/AMDGPU/GlobalISel/regbankselect-amdgcn.regalloc.handoff.mir (+194) 
- (modified) llvm/test/CodeGen/AMDGPU/amdgpu-attributor-min-agpr-alloc.ll (+42) 
- (added) llvm/test/CodeGen/AMDGPU/amdgpu-attributor-regalloc-handoff.ll (+21) 
- (added) llvm/test/CodeGen/AMDGPU/branch-folder-nomerge.mir (+92) 
- (added) llvm/test/CodeGen/AMDGPU/early-ifcvt-nomerge.mir (+99) 
- (added) llvm/test/CodeGen/AMDGPU/llvm.amdgcn.regalloc.handoff.ll (+112) 
- (added) llvm/test/CodeGen/AMDGPU/machine-cse-nomerge.mir (+20) 
- (added) llvm/test/CodeGen/AMDGPU/machine-licm-nomerge.mir (+78) 
- (added) llvm/test/CodeGen/AMDGPU/regalloc-handoff-machine-passes.mir (+32) 
- (added) llvm/test/CodeGen/AMDGPU/regalloc-handoff-postra.mir (+55) 
- (added) llvm/test/CodeGen/AMDGPU/regalloc-handoff-stress-regalloc.mir (+109) 
- (added) llvm/test/Transforms/EarlyCSE/nomerge.ll (+31) 
- (added) llvm/test/Transforms/GVN/nomerge.ll (+17) 
- (added) llvm/test/Transforms/GVNSink/nomerge.ll (+42) 
- (added) llvm/test/Transforms/InstCombine/AMDGPU/regalloc-handoff.ll (+23) 
- (added) llvm/test/Transforms/NewGVN/nomerge.ll (+17) 
- (added) llvm/test/Verifier/AMDGPU/regalloc-handoff.ll (+40) 


``````````diff
diff --git a/llvm/docs/AMDGPUUsage.rst b/llvm/docs/AMDGPUUsage.rst
index 40a2820b9e04f..3a3c32dfb1ae0 100644
--- a/llvm/docs/AMDGPUUsage.rst
+++ b/llvm/docs/AMDGPUUsage.rst
@@ -2120,6 +2120,39 @@ The AMDGPU backend implements the following LLVM IR intrinsics.
 
    List AMDGPU intrinsics.
 
+.. _amdgpu-regalloc-handoff:
+
+'``llvm.amdgcn.regalloc.handoff``' Intrinsic
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+The ``llvm.amdgcn.regalloc.handoff`` intrinsic creates a local register
+allocation boundary around one 32-bit value:
+
+.. code-block:: llvm
+
+   i32 @llvm.amdgcn.regalloc.handoff(i32 %value, i32 immarg %mode)
+
+The result is equal to ``%value``, but the incoming and outgoing values have
+separate register-allocation intervals. The immediate mode controls outgoing
+register allocation:
+
+* ``0`` uses automatic AV outgoing allocation with a source allocation hint,
+  preferring the incoming value's allocation. Scalar and otherwise unhintable
+  values follow the normal AV allocation policy.
+* ``1`` selects VGPR.
+* ``2`` selects AGPR. This mode is unsupported on targets without matrix
+  instruction support.
+
+The intrinsic is pure. An unused call can be removed, and a live call can move
+like an ordinary data-dependent operation, but separate live occurrences are
+not merged or duplicated. It is not a scheduling, memory, inline-assembly, or
+``EXEC`` barrier.
+
+Register allocation may assign compatible incoming and outgoing intervals to
+the same physical register, in which case the handoff emits no instruction.
+Otherwise post-register-allocation expansion uses the normal AMDGPU physical
+copy lowering, including the required VGPR/AGPR transition when banks differ.
+
 .. _amdgpu-av-load-store:
 
 '``llvm.amdgcn.av``' Intrinsics
diff --git a/llvm/include/llvm/IR/IntrinsicsAMDGPU.td b/llvm/include/llvm/IR/IntrinsicsAMDGPU.td
index 037630ea86904..936ed86e22647 100644
--- a/llvm/include/llvm/IR/IntrinsicsAMDGPU.td
+++ b/llvm/include/llvm/IR/IntrinsicsAMDGPU.td
@@ -159,6 +159,14 @@ let TargetPrefix = "amdgcn" in {
 // ABI Special Intrinsics
 //===----------------------------------------------------------------------===//
 
+// Preserve a local register-allocation boundary around a 32-bit value. The
+// mode is immediate: 0 uses automatic AV allocation with a source allocation
+// preference, 1 selects VGPR, and 2 selects AGPR.
+def int_amdgcn_regalloc_handoff : PureIntrinsic<
+    [llvm_i32_ty], [llvm_i32_ty, llvm_i32_ty],
+    [ImmArg<ArgIndex<1>>, Range<ArgIndex<1>, 0, 3>,
+     IntrNoMerge, IntrNoDuplicate]>;
+
 defm int_amdgcn_workitem_id
     : AMDGPUReadPreloadRegisterIntrinsic_xyz_named<
           "__builtin_amdgcn_workitem_id", [Range<RetIndex, 0, 1024>]>;
diff --git a/llvm/lib/CodeGen/BranchFolding.cpp b/llvm/lib/CodeGen/BranchFolding.cpp
index eb36cad8ab3db..ce5f1578a8a96 100644
--- a/llvm/lib/CodeGen/BranchFolding.cpp
+++ b/llvm/lib/CodeGen/BranchFolding.cpp
@@ -2057,6 +2057,10 @@ bool BranchFolder::HoistCommonCodeInSuccs(MachineBasicBlock *MBB) {
     if (TIB == TIE || FIB == FIE)
       break;
 
+    if (TIB->getFlag(MachineInstr::NoMerge) ||
+        FIB->getFlag(MachineInstr::NoMerge))
+      break;
+
     if (!TIB->isIdenticalTo(*FIB, MachineInstr::CheckKillDead))
       break;
 
diff --git a/llvm/lib/CodeGen/EarlyIfConversion.cpp b/llvm/lib/CodeGen/EarlyIfConversion.cpp
index bf2664fd4225c..2175ce53f39bd 100644
--- a/llvm/lib/CodeGen/EarlyIfConversion.cpp
+++ b/llvm/lib/CodeGen/EarlyIfConversion.cpp
@@ -610,6 +610,10 @@ static bool hasSameValue(const MachineRegisterInfo &MRI,
   if (!TDef || !FDef)
     return false;
 
+  if (TDef->getFlag(MachineInstr::NoMerge) ||
+      FDef->getFlag(MachineInstr::NoMerge))
+    return false;
+
   // If there are side-effects, all bets are off.
   if (TDef->hasUnmodeledSideEffects())
     return false;
diff --git a/llvm/lib/CodeGen/MIRParser/MILexer.cpp b/llvm/lib/CodeGen/MIRParser/MILexer.cpp
index 67fbc00edd9db..d3aa96fe94c24 100644
--- a/llvm/lib/CodeGen/MIRParser/MILexer.cpp
+++ b/llvm/lib/CodeGen/MIRParser/MILexer.cpp
@@ -299,6 +299,7 @@ static MIToken::TokenKind getIdentifierKind(StringRef Identifier) {
       .Case("machine-block-address-taken",
             MIToken::kw_machine_block_address_taken)
       .Case("call-frame-size", MIToken::kw_call_frame_size)
+      .Case("nomerge", MIToken::kw_nomerge)
       .Case("noconvergent", MIToken::kw_noconvergent)
       .Case("mmra", MIToken::kw_mmra)
       .Case("lr-split", MIToken::kw_lr_split)
diff --git a/llvm/lib/CodeGen/MIRParser/MILexer.h b/llvm/lib/CodeGen/MIRParser/MILexer.h
index f5947bfe59b9d..fd50316a662e1 100644
--- a/llvm/lib/CodeGen/MIRParser/MILexer.h
+++ b/llvm/lib/CodeGen/MIRParser/MILexer.h
@@ -151,6 +151,7 @@ struct MIToken {
     kw_ir_block_address_taken,
     kw_machine_block_address_taken,
     kw_call_frame_size,
+    kw_nomerge,
     kw_noconvergent,
     kw_mmra,
     kw_lr_split,
diff --git a/llvm/lib/CodeGen/MIRParser/MIParser.cpp b/llvm/lib/CodeGen/MIRParser/MIParser.cpp
index bb0b87cc042d0..77ddcb6c0ccbc 100644
--- a/llvm/lib/CodeGen/MIRParser/MIParser.cpp
+++ b/llvm/lib/CodeGen/MIRParser/MIParser.cpp
@@ -1530,6 +1530,7 @@ bool MIParser::parseInstruction(unsigned &OpCode, unsigned &Flags) {
          Token.is(MIToken::kw_nsw) ||
          Token.is(MIToken::kw_exact) ||
          Token.is(MIToken::kw_nofpexcept) ||
+         Token.is(MIToken::kw_nomerge) ||
          Token.is(MIToken::kw_noconvergent) ||
          Token.is(MIToken::kw_unpredictable) ||
          Token.is(MIToken::kw_nneg) ||
@@ -1566,6 +1567,8 @@ bool MIParser::parseInstruction(unsigned &OpCode, unsigned &Flags) {
       Flags |= MachineInstr::IsExact;
     if (Token.is(MIToken::kw_nofpexcept))
       Flags |= MachineInstr::NoFPExcept;
+    if (Token.is(MIToken::kw_nomerge))
+      Flags |= MachineInstr::NoMerge;
     if (Token.is(MIToken::kw_unpredictable))
       Flags |= MachineInstr::Unpredictable;
     if (Token.is(MIToken::kw_noconvergent))
diff --git a/llvm/lib/CodeGen/MachineCSE.cpp b/llvm/lib/CodeGen/MachineCSE.cpp
index 3d0ac171b9fe4..e11f49715af9a 100644
--- a/llvm/lib/CodeGen/MachineCSE.cpp
+++ b/llvm/lib/CodeGen/MachineCSE.cpp
@@ -395,7 +395,7 @@ bool MachineCSEImpl::PhysRegDefsReach(MachineInstr *CSMI, MachineInstr *MI,
 bool MachineCSEImpl::isCSECandidate(MachineInstr *MI) {
   if (MI->isPosition() || MI->isPHI() || MI->isImplicitDef() || MI->isKill() ||
       MI->isInlineAsm() || MI->isDebugInstr() || MI->isJumpTableDebugInfo() ||
-      MI->isFakeUse())
+      MI->isFakeUse() || MI->getFlag(MachineInstr::NoMerge))
     return false;
 
   // Ignore copies.
diff --git a/llvm/lib/CodeGen/MachineLICM.cpp b/llvm/lib/CodeGen/MachineLICM.cpp
index 6435bbb9382c1..9a8a90c830025 100644
--- a/llvm/lib/CodeGen/MachineLICM.cpp
+++ b/llvm/lib/CodeGen/MachineLICM.cpp
@@ -1507,9 +1507,15 @@ void MachineLICMImpl::InitializeLoadsHoistableLoops() {
 MachineInstr *
 MachineLICMImpl::LookForDuplicate(const MachineInstr *MI,
                                   std::vector<MachineInstr *> &PrevMIs) {
-  for (MachineInstr *PrevMI : PrevMIs)
+  if (MI->getFlag(MachineInstr::NoMerge))
+    return nullptr;
+
+  for (MachineInstr *PrevMI : PrevMIs) {
+    if (PrevMI->getFlag(MachineInstr::NoMerge))
+      continue;
     if (TII->produceSameValue(*MI, *PrevMI, (PreRegAlloc ? MRI : nullptr)))
       return PrevMI;
+  }
 
   return nullptr;
 }
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUAttributor.cpp b/llvm/lib/Target/AMDGPU/AMDGPUAttributor.cpp
index 630ffad96e451..a2c7e03d046a2 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUAttributor.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUAttributor.cpp
@@ -1360,6 +1360,12 @@ struct AAAMDGPUMinAGPRAlloc
 
         return true;
       }
+      case Intrinsic::amdgcn_regalloc_handoff:
+        // Mode 2 forces the result into an AGPR, but the number of mode 2
+        // values simultaneously live during register allocation is unknown
+        // here. Do not manifest a bounded AGPR allocation for this function
+        // or its callers.
+        return !cast<ConstantInt>(CB.getArgOperand(1))->equalsInt(2);
       // Trap-like intrinsics such as llvm.trap and llvm.debugtrap do not have
       // the nocallback attribute, so the AMDGPU attributor can conservatively
       // drop all implicitly-known inputs and AGPR allocation information. Make
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUISelDAGToDAG.cpp b/llvm/lib/Target/AMDGPU/AMDGPUISelDAGToDAG.cpp
index 744bde043e16c..e350bb97e391b 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUISelDAGToDAG.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUISelDAGToDAG.cpp
@@ -26,6 +26,7 @@
 #include "llvm/CodeGen/SelectionDAG.h"
 #include "llvm/CodeGen/SelectionDAGISel.h"
 #include "llvm/CodeGen/SelectionDAGNodes.h"
+#include "llvm/IR/DiagnosticInfo.h"
 #include "llvm/IR/IntrinsicsAMDGPU.h"
 #include "llvm/InitializePasses.h"
 #include "llvm/Support/ErrorHandling.h"
@@ -697,6 +698,13 @@ void AMDGPUDAGToDAGISel::Select(SDNode *N) {
   switch (Opc) {
   default:
     break;
+  case ISD::SRCVALUE:
+    // A regalloc handoff retains this non-emitted identity operand on its
+    // selected machine node to prevent MorphNodeTo from CSE'ing distinct IR
+    // calls. Morph the same value to a non-emitted machine node. The unused
+    // glue result makes the machine node unique as well.
+    CurDAG->SelectNodeTo(N, TargetOpcode::IMPLICIT_DEF, MVT::Other, MVT::Glue);
+    return;
   case ISD::UADDO_CARRY:
   case ISD::USUBO_CARRY:
     if (N->getValueType(0) == MVT::i64) {
@@ -3303,6 +3311,48 @@ void AMDGPUDAGToDAGISel::SelectINTRINSIC_WO_CHAIN(SDNode *N) {
     ConvGlueNode = nullptr;
   }
   switch (IntrID) {
+  case Intrinsic::amdgcn_regalloc_handoff: {
+    unsigned Mode = N->getConstantOperandVal(2);
+    unsigned HandoffOpcode;
+    switch (Mode) {
+    case 0:
+      HandoffOpcode = AMDGPU::REGALLOC_HANDOFF_AV;
+      break;
+    case 1:
+      HandoffOpcode = AMDGPU::REGALLOC_HANDOFF_VGPR;
+      break;
+    case 2:
+      if (!Subtarget->hasMAIInsts()) {
+        const Function &Fn = CurDAG->getMachineFunction().getFunction();
+        Fn.getContext().diagnose(DiagnosticInfoUnsupported(
+            Fn,
+            "llvm.amdgcn.regalloc.handoff mode 2 requires matrix instruction "
+            "support",
+            N->getDebugLoc()));
+        HandoffOpcode = AMDGPU::REGALLOC_HANDOFF_VGPR;
+      } else {
+        HandoffOpcode = AMDGPU::REGALLOC_HANDOFF_AGPR;
+      }
+      break;
+    default:
+      llvm_unreachable("invalid llvm.amdgcn.regalloc.handoff mode");
+    }
+
+    SDLoc DL(N);
+    SDValue AVClass =
+        CurDAG->getTargetConstant(AMDGPU::AV_32RegClassID, DL, MVT::i32);
+    SDValue Src(CurDAG->getMachineNode(TargetOpcode::COPY_TO_REGCLASS, DL,
+                                       N->getValueType(0), N->getOperand(1),
+                                       AVClass),
+                0);
+    SmallVector<SDValue, 3> Ops = {Src, N->getOperand(3)};
+    if (ConvGlueNode)
+      Ops.push_back(SDValue(ConvGlueNode, 0));
+    SDNode *Selected =
+        CurDAG->SelectNodeTo(N, HandoffOpcode, N->getVTList(), Ops);
+    CurDAG->addNoMergeSiteInfo(Selected, true);
+    return;
+  }
   case Intrinsic::amdgcn_wqm:
     Opcode = AMDGPU::WQM;
     break;
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUInstructionSelector.cpp b/llvm/lib/Target/AMDGPU/AMDGPUInstructionSelector.cpp
index 353506378b229..577bb65cacb2e 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUInstructionSelector.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUInstructionSelector.cpp
@@ -1241,6 +1241,61 @@ bool AMDGPUInstructionSelector::selectDivScale(MachineInstr &MI) const {
 bool AMDGPUInstructionSelector::selectG_INTRINSIC(MachineInstr &I) const {
   Intrinsic::ID IntrinsicID = cast<GIntrinsic>(I).getIntrinsicID();
   switch (IntrinsicID) {
+  case Intrinsic::amdgcn_regalloc_handoff: {
+    if (I.getNumOperands() != 4 || !I.getOperand(0).isReg() ||
+        !I.getOperand(2).isReg() || !I.getOperand(3).isImm())
+      return false;
+
+    unsigned Opcode;
+    switch (I.getOperand(3).getImm()) {
+    case 0:
+      Opcode = AMDGPU::REGALLOC_HANDOFF_AV;
+      break;
+    case 1:
+      Opcode = AMDGPU::REGALLOC_HANDOFF_VGPR;
+      break;
+    case 2:
+      if (!STI.hasMAIInsts()) {
+        const Function &Fn = I.getMF()->getFunction();
+        Fn.getContext().diagnose(DiagnosticInfoUnsupported(
+            Fn,
+            "llvm.amdgcn.regalloc.handoff mode 2 requires matrix instruction "
+            "support",
+            I.getDebugLoc()));
+        Opcode = AMDGPU::REGALLOC_HANDOFF_VGPR;
+        break;
+      }
+      Opcode = AMDGPU::REGALLOC_HANDOFF_AGPR;
+      break;
+    default:
+      return false;
+    }
+
+    bool IsAuto = Opcode == AMDGPU::REGALLOC_HANDOFF_AV;
+    Register Dst = I.getOperand(0).getReg();
+    if (IsAuto) {
+      const RegisterBank *DstRB = RBI.getRegBank(Dst, *MRI, TRI);
+      const TargetRegisterClass &DstRC = DstRB->getID() == AMDGPU::AGPRRegBankID
+                                             ? AMDGPU::AGPR_32RegClass
+                                             : AMDGPU::VGPR_32RegClass;
+      if (!RBI.constrainGenericRegister(Dst, DstRC, *MRI))
+        return false;
+    }
+
+    MachineInstrBuilder Handoff =
+        BuildMI(*I.getParent(), &I, I.getDebugLoc(), TII.get(Opcode))
+            .addDef(Dst)
+            .add(I.getOperand(2));
+    Handoff->setFlag(MachineInstr::NoMerge);
+    constrainSelectedInstRegOperands(*Handoff, TII, TRI, RBI);
+    I.eraseFromParent();
+    // RegBankSelect needed a concrete bank for the generic definition. Widen
+    // it back to AV_32 unless an already-selected use requires a narrower
+    // class.
+    if (IsAuto)
+      MRI->recomputeRegClass(Dst);
+    return true;
+  }
   case Intrinsic::amdgcn_if_break: {
     MachineBasicBlock *BB = I.getParent();
 
diff --git a/llvm/lib/Target/AMDGPU/AMDGPURegBankLegalizeRules.cpp b/llvm/lib/Target/AMDGPU/AMDGPURegBankLegalizeRules.cpp
index 0072ef76464cc..3993800132d4d 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPURegBankLegalizeRules.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPURegBankLegalizeRules.cpp
@@ -1777,6 +1777,9 @@ RegBankLegalizeRules::RegBankLegalizeRules(const GCNSubtarget &_ST,
 
   addRulesForIOpcs({returnaddress}).Any({{UniP0}, {{SgprP0}, {}}});
 
+  addRulesForIOpcs({amdgcn_regalloc_handoff})
+      .Any({{B32, _, B32, _}, {{None}, {IntrId, None, Imm}}});
+
   addRulesForIOpcs({amdgcn_s_getpc}).Any({{UniS64, _}, {{Sgpr64}, {}}});
 
   addRulesForIOpcs({amdgcn_s_getreg}).Any({{}, {{Sgpr32}, {IntrId}}});
diff --git a/llvm/lib/Target/AMDGPU/AMDGPURegBankSelect.cpp b/llvm/lib/Target/AMDGPU/AMDGPURegBankSelect.cpp
index f55f183b92eee..43c62b71d4187 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPURegBankSelect.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPURegBankSelect.cpp
@@ -18,11 +18,15 @@
 #include "AMDGPU.h"
 #include "AMDGPUGlobalISelUtils.h"
 #include "GCNSubtarget.h"
+#include "llvm/ADT/SmallVector.h"
 #include "llvm/CodeGen/GlobalISel/CSEInfo.h"
 #include "llvm/CodeGen/GlobalISel/CSEMIRBuilder.h"
+#include "llvm/CodeGen/GlobalISel/GenericMachineInstrs.h"
 #include "llvm/CodeGen/MachineUniformityAnalysis.h"
 #include "llvm/CodeGen/TargetPassConfig.h"
+#include "llvm/IR/IntrinsicsAMDGPU.h"
 #include "llvm/InitializePasses.h"
+#include <utility>
 
 #define DEBUG_TYPE "amdgpu-reg-bank-select"
 
@@ -81,6 +85,7 @@ class RegBankSelectHelper {
   AMDGPU::IntrinsicLaneMaskAnalyzer &ILMA;
   const MachineUniformityInfo &MUI;
   const SIRegisterInfo &TRI;
+  const RegisterBankInfo &RBI;
   const RegisterBank *SgprRB;
   const RegisterBank *VgprRB;
   const RegisterBank *VccRB;
@@ -90,7 +95,7 @@ class RegBankSelectHelper {
                       AMDGPU::IntrinsicLaneMaskAnalyzer &ILMA,
                       const MachineUniformityInfo &MUI,
                       const SIRegisterInfo &TRI, const RegisterBankInfo &RBI)
-      : B(B), MRI(*B.getMRI()), ILMA(ILMA), MUI(MUI), TRI(TRI),
+      : B(B), MRI(*B.getMRI()), ILMA(ILMA), MUI(MUI), TRI(TRI), RBI(RBI),
         SgprRB(&RBI.getRegBank(AMDGPU::SGPRRegBankID)),
         VgprRB(&RBI.getRegBank(AMDGPU::VGPRRegBankID)),
         VccRB(&RBI.getRegBank(AMDGPU::VCCRegBankID)) {}
@@ -183,6 +188,71 @@ class RegBankSelectHelper {
 
     B.buildCopy(NewReg, Reg);
   }
+
+  bool isUniformPhiHandoffUse(MachineInstr &MI, MachineOperand &UseOP) {
+    assert(MI.isPHI());
+    Register PhiReg = MI.getOperand(0).getReg();
+    Register UseReg = UseOP.getReg();
+    if (MRI.getRegBankOrNull(PhiReg) != SgprRB)
+      return false;
+
+    auto *GI = dyn_cast_or_null<GIntrinsic>(MRI.getVRegDef(UseReg));
+    return GI && GI->getIntrinsicID() == Intrinsic::amdgcn_regalloc_handoff;
+  }
+
+  void reconcileUniformPhiHandoffUse(MachineInstr &MI, MachineOperand &UseOP) {
+    assert(isUniformPhiHandoffUse(MI, UseOP));
+    Register UseReg = UseOP.getReg();
+    const RegisterBank *UseRB = MRI.getRegBankOrNull(UseReg);
+    assert(UseRB && UseRB != SgprRB);
+
+    MachineBasicBlock &IncomingMBB =
+        *MI.getOperand(UseOP.getOperandNo() + 1).getMBB();
+    B.setInsertPt(IncomingMBB, IncomingMBB.getFirstTerminator());
+
+    LLT Ty = MRI.getType(UseReg);
+    Register VgprUse = UseReg;
+    if (UseRB->getID() == AMDGPU::AGPRRegBankID) {
+      VgprUse = MRI.createVirtualRegister({VgprRB, Ty});
+      B.buildCopy(VgprUse, UseReg);
+    } else {
+      assert(UseRB == VgprRB);
+    }
+
+    Register SgprUse = MRI.createVirtualRegister({SgprRB, Ty});
+    buildReadFirstLane(B, SgprUse, VgprUse, RBI);
+    UseOP.setReg(SgprUse);
+  }
+
+  bool applyRegallocHandoffMapping(
+      MachineInstr &MI, const RegisterBankInfo::InstructionMapping &Mapping) {
+    if (!Mapping.isValid() || Mapping.getNumOperands() != MI.getNumOperands())
+      return false;
+
+    for (unsigned OpIdx : {0u, 2u}) {
+      MachineOperand &Op = MI.getOperand(OpIdx);
+      const RegisterBankInfo::ValueMapping &ValMapping =
+          Mapping.getOperandMapping(OpIdx);
+      if (!Op.isReg() || !Op.getReg().isVirtual() || !ValMapping.isValid() ||
+          ValMapping.NumBreakDowns != 1)
+        return false;
+
+      const RegisterBank *RB = ValMapping.BreakDown[0].RegBank;
+      Register Reg = Op.getReg();
+      if (Op.isDef()) {
+        if (MRI.getRegClassOrNull(Reg))
+          reAssignRegBankOnDef(MI, Op, RB);
+        else {
+          assert(!MRI.getRegBankOrNull(Reg));
+          MRI.setRegBank(Reg, *RB);
+        }
+      } else if (RBI.getRegBank(Reg, MRI, TRI) != RB) {
+        constrainRegBankUse(MI, Op, RB);
+      }
+    }
+
+    return true;
+  }
 };
 
 static Register getVReg(MachineOperand &Op) {
@@ -223,6 +293,7 @@ bool AMDGPURegBankSelect::runOnMachineFunction(MachineFunction &MF) {
   const GCNSubtarget &ST = MF.getSubtarget<GCNSubtarget>();
   RegBankSelectHelper RBSHelper(B, ILMA, MUI, *ST.getRegisterInfo(),
                                 *ST.getRegBankInfo());
+  SmallVector<std::pair<MachineInstr *, unsigned>, 4> UniformPhiHandoffUses;
   // Virtual registers at this point don't have register banks.
   // Virtual registers in def and use operands of already inst-selected
   // instruction have register class.
@@ -244,6 +315,17 @@ bool AMDGPURegBankSelect::runOnMachineFunction(MachineFunction &MF) {
       if (!MI.isPreISelOpcode())
         continue;
 
+      if (auto *GI = dyn_cast<GIntrinsic>(&MI);
+          GI && GI->getIntrinsicID() == Intrinsic::amdgcn_regalloc_handoff) {
+        const RegisterBankInfo::InstructionMapping &Mapping =
+            ST.getRegBankInfo()->getInstrMapping(MI);
+        if (!RBSHelper.applyRegallocHandoffMapping(MI, Mapping)) {
+          MF.getProperties().setFailedISel();
+          return true;
+        }
+        continue;
+      }
+
       // Vregs in def and use operands of G_ instructions need to have register
       // banks assigned. Before this loop possible case are
       // - (1) vreg without register class or bank in def or use operand
@@ -275,6 +357,11 @@ bool AMDGPURegBankSelect::runOnMachineFunction(MachineFunction &MF) {
         if (!UseReg.isValid())
           continue;
 
+        if (MI.isPHI() && RBSHelper.isUniformPhiHandoffUse(MI, UseOP)) {
+          UniformPhiHandoffUses.emplace_back(&MI, UseOP.getOperandNo());
+          continue;
+        }
+
         // Skip case (3).
         if (!MRI.getRegClassOrNull(UseReg) ||
             MRI.getVRegDef(UseReg)->isPreISelOpcode())
@@ -287,5 +374,13 @@ bool AMDGPURegBankSelect::runOnMachineFunction(MachineFunction &MF) {
     }
...
[truncated]

``````````

</details>


https://github.com/llvm/llvm-project/pull/219842


More information about the llvm-commits mailing list