[llvm] [AMDGPU][GFX12/GFX13] Add cdbg branch support and lower llvm.is.debugging.enabled (PR #219178)

via llvm-commits llvm-commits at lists.llvm.org
Thu Aug 27 04:45:15 PDT 2026


Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>,
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>,
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>,
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>
Message-ID:
In-Reply-To: <llvm.org/llvm/llvm-project/pull/219178 at github.com>


llvmorg-github-actions[bot] wrote:


<!--LLVM PR SUMMARY COMMENT-->

@llvm/pr-subscribers-backend-amdgpu

Author: Alexander Hück (ahueck)

<details>
<summary>Changes</summary>

This is PR 3 of 4 in a series adding `llvm.is.debugging.enabled` intrinsic to LLVM IR and lowering this intrinsic for AMDGPU targets, , see [[[RFC] Introduce `llvm.is.debugging.enabled` intrinsic](https://discourse.llvm.org/t/rfc-introduce-llvm-is-debugging-enabled-intrinsic/91676)

## Changes of this PR
- Opt AMDGPU into lowering `llvm.is.debugging.enabled`. Unsupported subtargets replace the intrinsic with false.
- Add cdbg conditional branch instruction support for GFX12 and GFX13.
- On supported subtargets, SelectionDAG and GlobalISel lowers the intrinsic by either reading the generation-specific debugging-state bits with `s_getreg` or fuse safe single-branch uses into `s_cbranch_cdbgsys_or_user`. This includes negated branch conditions.
- Mark materialized and fused observations NoMerge to keep calls distinct.

PR 4 uses this support to implement AMDPAL lowering of `llvm.debugtrap`.

## PR stack
1.	[[AMDGPU][GlobalISel][SelectionDAG] Refactor control-flow intrinsic branch matching](https://github.com/llvm/llvm-project/pull/219173)
2.	[[IR][CodeGen] Add llvm.is.debugging.enabled intrinsic](https://github.com/llvm/llvm-project/pull/219175)
3.	[[AMDGPU][GFX12/GFX13] Add cdbg branch support and lower llvm.is.debugging.enabled](https://github.com/llvm/llvm-project/pull/219178)
4.	[[AMDGPU] Add AMDPAL support for llvm.debugtrap](https://github.com/llvm/llvm-project/pull/219179)

*Disclaimer*: Agentic AI assisted development.

---

Patch is 87.61 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/219178.diff


32 Files Affected:

- (modified) llvm/docs/AMDGPUUsage.rst (+46) 
- (modified) llvm/docs/LangRef.md (+30) 
- (modified) llvm/include/llvm/IR/Intrinsics.td (+6) 
- (modified) llvm/include/llvm/Target/TargetMachine.h (+7) 
- (modified) llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp (+8) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPU.td (+11-1) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp (+109-63) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp (+16) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPUTargetMachine.h (+2) 
- (modified) llvm/lib/Target/AMDGPU/SIISelLowering.cpp (+131-42) 
- (modified) llvm/lib/Target/AMDGPU/SOPInstructions.td (+12-8) 
- (modified) llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.cpp (+10) 
- (modified) llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.h (+4) 
- (modified) llvm/test/Analysis/UniformityAnalysis/AMDGPU/always_uniform.ll (+28) 
- (added) llvm/test/Assembler/is-debugging-enabled.ll (+11) 
- (modified) llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-amdgcn.if-invalid.mir (+26-1) 
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled-divergent-exec-guard.ll (+236) 
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled-shapes.ll (+449) 
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled-unsupported.ll (+32) 
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled.ll (+257) 
- (added) llvm/test/CodeGen/NVPTX/is-debugging-enabled.ll (+11) 
- (added) llvm/test/CodeGen/X86/is-debugging-enabled.ll (+12) 
- (modified) llvm/test/MC/AMDGPU/gfx12_asm_sopp.s (+24) 
- (modified) llvm/test/MC/AMDGPU/gfx12_unsupported.s (-12) 
- (modified) llvm/test/MC/AMDGPU/gfx13_asm_sopp.s (+24) 
- (added) llvm/test/Transforms/DCE/is-debugging-enabled.ll (+12) 
- (added) llvm/test/Transforms/EarlyCSE/is-debugging-enabled.ll (+17) 
- (added) llvm/test/Transforms/LICM/is-debugging-enabled.ll (+25) 
- (added) llvm/test/Transforms/LoopUnroll/is-debugging-enabled.ll (+30) 
- (added) llvm/test/Transforms/PreISelIntrinsicLowering/is-debugging-enabled.ll (+17) 
- (added) llvm/test/Transforms/SimplifyCFG/is-debugging-enabled.ll (+58) 
- (added) llvm/test/Verifier/is-debugging-enabled.ll (+15) 


``````````diff
diff --git a/llvm/docs/AMDGPUUsage.rst b/llvm/docs/AMDGPUUsage.rst
index 40a2820b9e04f..e82af4d16c9df 100644
--- a/llvm/docs/AMDGPUUsage.rst
+++ b/llvm/docs/AMDGPUUsage.rst
@@ -1709,6 +1709,52 @@ The AMDGPU backend implements the following LLVM IR intrinsics.
                                                    The format is a 64-bit concatenation of the MODE and TRAPSTS registers.
 
   :ref:`llvm.set.fpenv<int_set_fpenv>`             Sets the floating point environment to the specified state.
+
+  :ref:`llvm.is.debugging.enabled <llvm.is.debugging.enabled>`
+                                                   Supported on GFX11.5, GFX12, and GFX13 targets. Other
+                                                   subtargets lower the result to ``false``.
+
+                                                   The target-defined execution context is the current wave.
+                                                   The result is uniform across the active lanes of that
+                                                   wave, including when the intrinsic is executed in
+                                                   divergent control flow. Each call remains a distinct
+                                                   observation of the wave's debugging-enabled state.
+
+                                                   A wave executes in debugging mode when either the
+                                                   ``COND_DBG_SYS`` or ``COND_DBG_USER`` bit is set.
+                                                   ``COND_DBG_SYS`` reflects a system-wide debugger
+                                                   attach, while ``COND_DBG_USER`` reflects per-dispatch
+                                                   launch control, configured through
+                                                   :ref:`CDBG_USER <amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx13-table>`
+                                                   in the kernel descriptor. The driver sets these
+                                                   bits, and provides an interface allowing the user
+                                                   mode runtime or an external debugger to control
+                                                   the setting.
+
+                                                   The intrinsic lowers in one of two forms.
+                                                   When the query feeds a single conditional branch in the
+                                                   same basic block, with no intervening observable
+                                                   operations, it may fuse into a single
+                                                   ``s_cbranch_cdbgsys_or_user``. Branch-hint and
+                                                   negated-condition forms are supported; other uses of
+                                                   the query value prevent fusion.
+
+                                                   Any other use materializes an ``i1`` by reading the
+                                                   adjacent ``COND_DBG_USER`` and ``COND_DBG_SYS`` bits
+                                                   and testing them against zero. The register holding
+                                                   them differs by generation:
+
+                                                   .. code-block:: none
+
+                                                     ; GFX11.5
+                                                     s_getreg_b32 s0, hwreg(HW_REG_STATUS, 20, 2)
+                                                     ; GFX12
+                                                     s_getreg_b32 s0, hwreg(HW_REG_WAVE_STATE_PRIV, 16, 2)
+                                                     ; GFX13
+                                                     s_getreg_b32 s0, hwreg(HW_REG_WAVE_STATUS, 20, 2)
+                                                     ; all generations
+                                                     s_cmp_lg_u32 s0, 0
+
   llvm.amdgcn.readfirstlane                        Provides direct access to v_readfirstlane_b32. Returns the value in
                                                    the lowest active lane of the input operand. Currently implemented
                                                    for i16, i32, float, half, bfloat, <2 x i16>, <2 x half>, <2 x bfloat>,
diff --git a/llvm/docs/LangRef.md b/llvm/docs/LangRef.md
index ee1600b96f1dd..9f2ab993f3a80 100644
--- a/llvm/docs/LangRef.md
+++ b/llvm/docs/LangRef.md
@@ -26263,6 +26263,36 @@ This intrinsic is lowered to code which is intended to cause an
 execution trap with the intention of requesting the attention of a
 debugger.
 
+(llvm.is.debugging.enabled)=
+
+#### '`llvm.is.debugging.enabled`' Intrinsic
+
+##### Syntax:
+
+```llvm
+declare noundef i1 @llvm.is.debugging.enabled() nomerge memory(inaccessiblemem: readwrite)
+```
+
+##### Overview:
+
+The '`llvm.is.debugging.enabled`' intrinsic returns whether debugging is enabled
+for the target-defined execution context of the current invocation.
+
+##### Arguments:
+
+None.
+
+##### Semantics:
+
+Each call observes the current debugging-enabled state. Calls are distinct
+observations: LLVM must not assume separate calls return the same value, and a
+call may not be removed when its result is unused, commoned with another call,
+or reordered with respect to one.
+
+On targets that do not support querying debugging-enabled state, the intrinsic
+returns `false` and performs no observation.
+
+
 (llvm.ubsantrap)=
 
 #### '`llvm.ubsantrap`' Intrinsic
diff --git a/llvm/include/llvm/IR/Intrinsics.td b/llvm/include/llvm/IR/Intrinsics.td
index c20ea64c7eef4..b14375e277e7a 100644
--- a/llvm/include/llvm/IR/Intrinsics.td
+++ b/llvm/include/llvm/IR/Intrinsics.td
@@ -2059,6 +2059,12 @@ def int_trap : Intrinsic<[], [],
                ClangBuiltin<"__builtin_trap">;
 def int_debugtrap : Intrinsic<[]>,
                     ClangBuiltin<"__builtin_debugtrap">;
+// Query whether debugging is enabled for the current execution context.
+def int_is_debugging_enabled :
+  DefaultAttrsIntrinsic<[llvm_i1_ty], [],
+                        [IntrInaccessibleMemOnly,
+                         IntrNoMerge,
+                         NoUndef<RetIndex>]>;
 def int_ubsantrap : Intrinsic<[], [llvm_i8_ty],
                               [IntrNoReturn, IntrCold, ImmArg<ArgIndex<0>>,
                                IntrInaccessibleMemOnly, IntrWriteMem]>;
diff --git a/llvm/include/llvm/Target/TargetMachine.h b/llvm/include/llvm/Target/TargetMachine.h
index a6b73d636dc27..ea26be5e48bfb 100644
--- a/llvm/include/llvm/Target/TargetMachine.h
+++ b/llvm/include/llvm/Target/TargetMachine.h
@@ -567,6 +567,13 @@ class LLVM_ABI TargetMachine {
   /// this function returns false, the intrinsic will be supported generically
   /// but without loop detection support.
   virtual bool canLowerCondLoop() const { return false; }
+
+  /// Returns whether this target takes responsibility for lowering
+  /// llvm.is.debugging.enabled. If false, generic lowering replaces the intrinsic
+  /// with false. If true, the target must handle every supported subtarget,
+  /// including replacing the intrinsic with false on subtargets without native
+  /// lowering.
+  virtual bool canLowerIsDebuggingEnabled() const { return false; }
 };
 
 } // end namespace llvm
diff --git a/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp b/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp
index 24788148b9f15..4ec866d7b973f 100644
--- a/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp
+++ b/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp
@@ -815,6 +815,14 @@ bool PreISelIntrinsicLowering::lowerIntrinsics(Module &M) const {
     case Intrinsic::protected_field_ptr:
       Changed |= expandProtectedFieldPtr(F);
       break;
+    case Intrinsic::is_debugging_enabled:
+      if (!TM || !TM->canLowerIsDebuggingEnabled())
+        Changed |= forEachCall(F, [](CallInst *CI) {
+          CI->replaceAllUsesWith(ConstantInt::getFalse(CI->getContext()));
+          CI->eraseFromParent();
+          return true;
+        });
+      break;
     case Intrinsic::cond_loop:
       if (!TM->canLowerCondLoop())
         Changed |= expandCondLoop(F);
diff --git a/llvm/lib/Target/AMDGPU/AMDGPU.td b/llvm/lib/Target/AMDGPU/AMDGPU.td
index ed2ee1ff70f4f..17b22beebcbb9 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPU.td
+++ b/llvm/lib/Target/AMDGPU/AMDGPU.td
@@ -149,6 +149,12 @@ defm TrapHandler: AMDGPUSubtargetFeature<"trap-handler",
   /*GenPredicate=*/0, /*GenAssemblerPredicate=*/1, [], InlineIgnore
 >;
 
+defm CDBGSysOrUserBranch : AMDGPUSubtargetFeature<
+  "cdbg-sys-or-user-branch",
+  "Support querying debugging state with s_cbranch_cdbgsys_or_user",
+  /*GenPredicate=*/1, /*GenAssemblerPredicate=*/0
+>;
+
 defm UnalignedScratchAccess : AMDGPUSubtargetFeature<"unaligned-scratch-access",
   "Support unaligned scratch loads and stores",
   /*GenPredicate=*/1, /*GenAssemblerPredicate=*/1, [], InlineIgnore
@@ -2232,7 +2238,8 @@ def FeatureISAVersion11_0_3 : FeatureSet<
 
 def FeatureISAVersion11_5_Common : FeatureSet<
   !listconcat(FeatureISAVersion11_Common.Features,
-    [FeatureSALUFloatInsts,
+    [FeatureCDBGSysOrUserBranch,
+     FeatureSALUFloatInsts,
      FeatureDPPSrc1SGPR,
      FeatureRequiredExportPriority,
      FeatureDot5Insts,
@@ -2275,6 +2282,7 @@ def FeatureISAVersion11_7_Generic: FeatureSet<
 def FeatureISAVersion12 : FeatureSet<
   [FeatureGFX12,
    FeatureSupportsWave64, FeatureSupportsWGP,
+   FeatureCDBGSysOrUserBranch,
    FeatureBackOffBarrier,
    FeatureAddressableLocalMemorySize65536,
    FeatureHalfAddressablePhysicalLocalMemory,
@@ -2343,6 +2351,7 @@ def FeatureISAVersion12 : FeatureSet<
 def FeatureISAVersion12_50_Common : FeatureSet<
   [FeatureGFX12,
    FeatureGFX1250Insts,
+   FeatureCDBGSysOrUserBranch,
    FeatureBackOffBarrier,
    FeatureRequiresAlignedVGPRs,
    FeatureCuMode,
@@ -2536,6 +2545,7 @@ def FeatureISAVersion12_5_Generic: FeatureSet<
 def FeatureISAVersion13 : FeatureSet<
   [FeatureGFX13,
    FeatureGFX1250Insts,
+   FeatureCDBGSysOrUserBranch,
    FeatureAddressableLocalMemorySize196608,
    Feature64BitLiterals,
    FeatureLDSBankCount32,
diff --git a/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp b/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
index 7c1a26f761c96..c42060e69a274 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
@@ -4804,47 +4804,77 @@ static bool isNot(const MachineRegisterInfo &MRI, const MachineInstr &MI) {
   return ConstVal == -1;
 }
 
-// Return the use branch instruction, otherwise null if the usage is invalid.
-static MachineInstr *
-verifyCFIntrinsic(MachineInstr &MI, MachineRegisterInfo &MRI, MachineInstr *&Br,
-                  MachineBasicBlock *&UncondBrTarget, bool &Negated) {
+namespace {
+struct CFIntrinsicBranchMatch {
+  MachineInstr *Negation = nullptr;
+  MachineInstr *CondBr = nullptr;
+  MachineInstr *UncondBr = nullptr;
+
+  MachineBasicBlock *ConditionTrueTarget = nullptr;
+  MachineBasicBlock *ConditionFalseTarget = nullptr;
+
+  bool isNegated() const { return Negation != nullptr; }
+
+  // Retarget the explicit branch for the fallthrough edge, or materialize
+  // the branch when that edge was represented by layout fallthrough. The
+  // builder must be positioned at the replacement branch.
+  void redirectFallthroughEdge(MachineIRBuilder &B,
+                               MachineBasicBlock &Target) const {
+    if (UncondBr)
+      UncondBr->getOperand(0).setMBB(&Target);
+    else
+      B.buildBr(Target);
+  }
+
+  void eraseDeadNegation(MachineRegisterInfo &MRI) {
+    if (Negation)
+      eraseInstr(*Negation, MRI);
+  }
+};
+} // namespace
+
+static std::optional<CFIntrinsicBranchMatch>
+matchCFIntrinsicBranchUse(MachineInstr &MI, MachineRegisterInfo &MRI) {
   Register CondDef = MI.getOperand(0).getReg();
   if (!MRI.hasOneNonDBGUse(CondDef))
-    return nullptr;
+    return std::nullopt;
 
   MachineBasicBlock *Parent = MI.getParent();
   MachineInstr *UseMI = &*MRI.use_instr_nodbg_begin(CondDef);
+  MachineInstr *Negation = nullptr;
 
   if (isNot(MRI, *UseMI)) {
+    Negation = UseMI;
     Register NegatedCond = UseMI->getOperand(0).getReg();
     if (!MRI.hasOneNonDBGUse(NegatedCond))
-      return nullptr;
-
-    // We're deleting the def of this value, so we need to remove it.
-    eraseInstr(*UseMI, MRI);
+      return std::nullopt;
 
     UseMI = &*MRI.use_instr_nodbg_begin(NegatedCond);
-    Negated = true;
   }
 
   if (UseMI->getParent() != Parent || UseMI->getOpcode() != AMDGPU::G_BRCOND)
-    return nullptr;
+    return std::nullopt;
 
   // Make sure the cond br is followed by a G_BR, or is the last instruction.
+  MachineInstr *UncondBr = nullptr;
+  MachineBasicBlock *OtherTarget = nullptr;
   MachineBasicBlock::iterator Next = std::next(UseMI->getIterator());
   if (Next == Parent->end()) {
     MachineFunction::iterator NextMBB = std::next(Parent->getIterator());
     if (NextMBB == Parent->getParent()->end()) // Illegal intrinsic use.
-      return nullptr;
-    UncondBrTarget = &*NextMBB;
+      return std::nullopt;
+    OtherTarget = &*NextMBB;
   } else {
     if (Next->getOpcode() != AMDGPU::G_BR)
-      return nullptr;
-    Br = &*Next;
-    UncondBrTarget = Br->getOperand(0).getMBB();
+      return std::nullopt;
+    UncondBr = &*Next;
+    OtherTarget = UncondBr->getOperand(0).getMBB();
   }
 
-  return UseMI;
+  MachineBasicBlock *TakenTarget = UseMI->getOperand(1).getMBB();
+  return CFIntrinsicBranchMatch{Negation, UseMI, UncondBr,
+                                 Negation ? OtherTarget : TakenTarget,
+                                 Negation ? TakenTarget : OtherTarget};
 }
 
 void AMDGPULegalizerInfo::buildLoadInputValue(Register DstReg,
@@ -8185,6 +8215,44 @@ bool AMDGPULegalizerInfo::legalizeIntrinsic(LegalizerHelper &Helper,
   // Replace the use G_BRCOND with the exec manipulate and branch pseudos.
   auto IntrID = cast<GIntrinsic>(MI).getIntrinsicID();
   switch (IntrID) {
+  case Intrinsic::is_debugging_enabled: {
+    auto Match = matchCFIntrinsicBranchUse(MI, MRI);
+    bool CannotFuse =
+        !Match || any_of(make_range(std::next(MI.getIterator()),
+                                    Match->CondBr->getIterator()),
+                         [](const MachineInstr &Between) {
+                           return !Between.isMetaInstruction() &&
+                                  (Between.mayLoadOrStore() ||
+                                   Between.hasUnmodeledSideEffects());
+                         });
+    if (CannotFuse) {
+      auto Bits =
+          B.buildIntrinsic(Intrinsic::amdgcn_s_getreg, {LLT::scalar(32)})
+              .addImm(AMDGPU::Hwreg::getDebuggingEnabledHwregImm(ST));
+      Bits->setFlag(MachineInstr::NoMerge);
+      B.buildICmp(CmpInst::ICMP_NE, MI.getOperand(0).getReg(), Bits.getReg(0),
+                  B.buildConstant(LLT::scalar(32), 0));
+      MI.eraseFromParent();
+      return true;
+    }
+
+    B.setInsertPt(*Match->CondBr->getParent(), Match->CondBr->getIterator());
+    B.setDebugLoc(Match->CondBr->getDebugLoc());
+    MachineInstrBuilder CDBGBranch =
+        B.buildInstr(AMDGPU::S_CBRANCH_CDBGSYS_OR_USER)
+            .addMBB(Match->ConditionTrueTarget);
+    CDBGBranch->setFlag(MachineInstr::NoMerge);
+
+    if (Match->isNegated())
+      Match->redirectFallthroughEdge(B, *Match->ConditionFalseTarget);
+
+    Register Cond = MI.getOperand(0).getReg();
+    MRI.markUsesInDebugValueAsUndef(Cond);
+    Match->eraseDeadNegation(MRI);
+    MI.eraseFromParent();
+    Match->CondBr->eraseFromParent();
+    return true;
+  }
   case Intrinsic::amdgcn_icmp: {
     // amdgcn.icmp(i1 src0, i1 0, NE) -> ballot(src0)
     // This is the only valid form of amdgcn.icmp with i1 inputs.
@@ -8230,80 +8298,58 @@ bool AMDGPULegalizerInfo::legalizeIntrinsic(LegalizerHelper &Helper,
     return true;
   case Intrinsic::amdgcn_if:
   case Intrinsic::amdgcn_else: {
-    MachineInstr *Br = nullptr;
-    MachineBasicBlock *UncondBrTarget = nullptr;
-    bool Negated = false;
-    if (MachineInstr *BrCond =
-            verifyCFIntrinsic(MI, MRI, Br, UncondBrTarget, Negated)) {
-      const SIRegisterInfo *TRI
-        = static_cast<const SIRegisterInfo *>(MRI.getTargetRegisterInfo());
+    if (auto Match = matchCFIntrinsicBranchUse(MI, MRI)) {
+      const SIRegisterInfo *TRI =
+          static_cast<const SIRegisterInfo *>(MRI.getTargetRegisterInfo());
 
       Register Def = MI.getOperand(1).getReg();
       Register Use = MI.getOperand(3).getReg();
 
-      MachineBasicBlock *CondBrTarget = BrCond->getOperand(1).getMBB();
-
-      if (Negated)
-        std::swap(CondBrTarget, UncondBrTarget);
-
-      B.setInsertPt(B.getMBB(), BrCond->getIterator());
+      B.setInsertPt(B.getMBB(), Match->CondBr->getIterator());
       if (IntrID == Intrinsic::amdgcn_if) {
         B.buildInstr(AMDGPU::SI_IF)
-          .addDef(Def)
-          .addUse(Use)
-          .addMBB(UncondBrTarget);
+            .addDef(Def)
+            .addUse(Use)
+            .addMBB(Match->ConditionFalseTarget);
       } else {
         B.buildInstr(AMDGPU::SI_ELSE)
             .addDef(Def)
             .addUse(Use)
-            .addMBB(UncondBrTarget);
+            .addMBB(Match->ConditionFalseTarget);
       }
 
-      if (Br) {
-        Br->getOperand(0).setMBB(CondBrTarget);
-      } else {
-        // The IRTranslator skips inserting the G_BR for fallthrough cases, but
-        // since we're swapping branch targets it needs to be reinserted.
-        // FIXME: IRTranslator should probably not do this
-        B.buildBr(*CondBrTarget);
-      }
+      // The IRTranslator skips inserting the G_BR for fallthrough cases, but
+      // since we're swapping branch targets it needs to be reinserted.
+      // FIXME: IRTranslator should probably not do this
+      Match->redirectFallthroughEdge(B, *Match->ConditionTrueTarget);
 
       MRI.setRegClass(Def, TRI->getWaveMaskRegClass());
       MRI.setRegClass(Use, TRI->getWaveMaskRegClass());
+      Match->eraseDeadNegation(MRI);
       MI.eraseFromParent();
-      BrCond->eraseFromParent();
+      Match->CondBr->eraseFromParent();
       return true;
     }
 
     return false;
   }
   case Intrinsic::amdgcn_loop: {
-    MachineInstr *Br = nullptr;
-    MachineBasicBlock *UncondBrTarget = nullptr;
-    bool Negated = false;
-    if (MachineInstr *BrCond =
-            verifyCFIntrinsic(MI, MRI, Br, UncondBrTarget, Negated)) {
-      const SIRegisterInfo *TRI
-        = static_cast<const SIRegisterInfo *>(MRI.getTargetRegisterInfo());
-
-      MachineBasicBlock *CondBrTarget = BrCond->getOperand(1).getMBB();
-      Register Reg = MI.getOperand(2).getReg();
+    if (auto Match = matchCFIntrinsicBranchUse(MI, MRI)) {
+      const SIRegisterInfo *TRI =
+          static_cast<const SIRegisterInfo *>(MRI.getTargetRegisterInfo());
 
-      if (Negated)
-        std::swap(CondBrTarget, UncondBrTarget);
+      Register Reg = MI.getOperand(2).getReg();
 
-      B.setInsertPt(B.getMBB(), BrCond->getIterator());
+      B.setInsertPt(B.getMBB(), Match->CondBr->getIterator());
       B.buildInstr(AMDGPU::SI_LOOP)
-        .addUse(Reg)
-        .addMBB(UncondBrTarget);
+          .addUse(Reg)
+          .addMBB(Match->ConditionFalseTarget);
 
-      if (Br)
-        Br->getOperand(0).setMBB(CondBrTarget);
-      else
-        B.buildBr(*CondBrTarget);
+      Match->redirectFallthroughEdge(B, *Match->ConditionTrueTarget);
 
+      Match->eraseDeadNegation(MRI);
       MI.eraseFromParent();
-      BrCond->eraseFromParent();
+      Match->CondBr->eraseFromParent();
       MRI.setRegClass(Reg, TRI->getWaveMaskRegClass());
       return true;
     }
diff --git a/llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp b/llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp
index 123d9930fc49e..274cba75e5254 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp
@@ -37,6 +37,7 @@ class AMDGPULowerIntrinsicsImpl {
   bool run();
 
 private:
+  bool visitIsDebuggingEnabled(IntrinsicInst &I);
   bool visitBarrier(IntrinsicInst &I);
   bool visitPtrSBufferLoad(IntrinsicInst &I);
 };
@@ -63,6 +64,16 @@ template <class T> static void forEachCall(Function &Intrin, T Callback) {
 
 } // anonymous namespace
 
+bool AMDGPULowerIntrinsicsIm...
[truncated]

``````````

</details>


https://github.com/llvm/llvm-project/pull/219178


More information about the llvm-commits mailing list