[llvm] [AMDGPU] Add AMDPAL support for llvm.debugtrap (PR #219179)

via llvm-commits llvm-commits at lists.llvm.org
Thu Aug 27 04:45:41 PDT 2026


Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>,
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>,
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>,
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>,
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>
Message-ID:
In-Reply-To: <llvm.org/llvm/llvm-project/pull/219179 at github.com>


llvmorg-github-actions[bot] wrote:


<!--LLVM PR SUMMARY COMMENT-->

@llvm/pr-subscribers-llvm-globalisel

Author: Alexander Hück (ahueck)

<details>
<summary>Changes</summary>

This is PR 4 of 4 in a series adding the `llvm.is.debugging.enabled` intrinsic, lowering it for AMDGPU targets, and enabling debug traps for AMDPAL, see [[[RFC] Introduce `llvm.is.debugging.enabled` intrinsic](https://discourse.llvm.org/t/rfc-introduce-llvm-is-debugging-enabled-intrinsic/91676)


## Changes of this PR
- Add an AMDPAL trap-handler ABI and lower `llvm.debugtrap` to s_trap 3 when the trap-handler target feature is enabled.
- Preserve the existing warning and omit the trap instruction when trap-handler support is disabled. `llvm.trap` continues to lower to s_endpgm.

## PR stack
1.	[[AMDGPU][GlobalISel][SelectionDAG] Refactor control-flow intrinsic branch matching](https://github.com/llvm/llvm-project/pull/219173)
2.	[[IR][CodeGen] Add llvm.is.debugging.enabled intrinsic](https://github.com/llvm/llvm-project/pull/219175)
3.	[[AMDGPU][GFX12/GFX13] Add cdbg branch support and lower llvm.is.debugging.enabled](https://github.com/llvm/llvm-project/pull/219178)
4.	[[AMDGPU] Add AMDPAL support for llvm.debugtrap](https://github.com/llvm/llvm-project/pull/219179)

*Disclaimer*: Agentic AI assisted development.



---

Patch is 95.34 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/219179.diff


34 Files Affected:

- (modified) llvm/docs/AMDGPUUsage.rst (+78-3) 
- (modified) llvm/docs/LangRef.md (+30) 
- (modified) llvm/include/llvm/IR/Intrinsics.td (+6) 
- (modified) llvm/include/llvm/Target/TargetMachine.h (+7) 
- (modified) llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp (+8) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPU.td (+11-1) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp (+114-66) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp (+16) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPUTargetMachine.h (+2) 
- (modified) llvm/lib/Target/AMDGPU/GCNSubtarget.h (+7-1) 
- (modified) llvm/lib/Target/AMDGPU/SIISelLowering.cpp (+134-43) 
- (modified) llvm/lib/Target/AMDGPU/SOPInstructions.td (+12-8) 
- (modified) llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.cpp (+10) 
- (modified) llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.h (+4) 
- (modified) llvm/test/Analysis/UniformityAnalysis/AMDGPU/always_uniform.ll (+28) 
- (added) llvm/test/Assembler/is-debugging-enabled.ll (+11) 
- (modified) llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-amdgcn.if-invalid.mir (+26-1) 
- (added) llvm/test/CodeGen/AMDGPU/debugtrap-amdpal.ll (+54) 
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled-divergent-exec-guard.ll (+240) 
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled-shapes.ll (+449) 
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled-unsupported.ll (+32) 
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled.ll (+257) 
- (added) llvm/test/CodeGen/NVPTX/is-debugging-enabled.ll (+11) 
- (added) llvm/test/CodeGen/X86/is-debugging-enabled.ll (+12) 
- (modified) llvm/test/MC/AMDGPU/gfx12_asm_sopp.s (+24) 
- (modified) llvm/test/MC/AMDGPU/gfx12_unsupported.s (-12) 
- (modified) llvm/test/MC/AMDGPU/gfx13_asm_sopp.s (+24) 
- (added) llvm/test/Transforms/DCE/is-debugging-enabled.ll (+12) 
- (added) llvm/test/Transforms/EarlyCSE/is-debugging-enabled.ll (+17) 
- (added) llvm/test/Transforms/LICM/is-debugging-enabled.ll (+25) 
- (added) llvm/test/Transforms/LoopUnroll/is-debugging-enabled.ll (+30) 
- (added) llvm/test/Transforms/PreISelIntrinsicLowering/is-debugging-enabled.ll (+17) 
- (added) llvm/test/Transforms/SimplifyCFG/is-debugging-enabled.ll (+58) 
- (added) llvm/test/Verifier/is-debugging-enabled.ll (+15) 


``````````diff
diff --git a/llvm/docs/AMDGPUUsage.rst b/llvm/docs/AMDGPUUsage.rst
index 40a2820b9e04f..f066f48f96b38 100644
--- a/llvm/docs/AMDGPUUsage.rst
+++ b/llvm/docs/AMDGPUUsage.rst
@@ -1709,6 +1709,52 @@ The AMDGPU backend implements the following LLVM IR intrinsics.
                                                    The format is a 64-bit concatenation of the MODE and TRAPSTS registers.
 
   :ref:`llvm.set.fpenv<int_set_fpenv>`             Sets the floating point environment to the specified state.
+
+  :ref:`llvm.is.debugging.enabled <llvm.is.debugging.enabled>`
+                                                   Supported on GFX11.5, GFX12, and GFX13 targets. Other
+                                                   subtargets lower the result to ``false``.
+
+                                                   The target-defined execution context is the current wave.
+                                                   The result is uniform across the active lanes of that
+                                                   wave, including when the intrinsic is executed in
+                                                   divergent control flow. Each call remains a distinct
+                                                   observation of the wave's debugging-enabled state.
+
+                                                   A wave executes in debugging mode when either the
+                                                   ``COND_DBG_SYS`` or ``COND_DBG_USER`` bit is set.
+                                                   ``COND_DBG_SYS`` reflects a system-wide debugger
+                                                   attach, while ``COND_DBG_USER`` reflects per-dispatch
+                                                   launch control, configured through
+                                                   :ref:`CDBG_USER <amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx13-table>`
+                                                   in the kernel descriptor. The driver sets these
+                                                   bits, and provides an interface allowing the user
+                                                   mode runtime or an external debugger to control
+                                                   the setting.
+
+                                                   The intrinsic lowers in one of two forms.
+                                                   When the query feeds a single conditional branch in the
+                                                   same basic block, with no intervening observable
+                                                   operations, it may fuse into a single
+                                                   ``s_cbranch_cdbgsys_or_user``. Branch-hint and
+                                                   negated-condition forms are supported; other uses of
+                                                   the query value prevent fusion.
+
+                                                   Any other use materializes an ``i1`` by reading the
+                                                   adjacent ``COND_DBG_USER`` and ``COND_DBG_SYS`` bits
+                                                   and testing them against zero. The register holding
+                                                   them differs by generation:
+
+                                                   .. code-block:: none
+
+                                                     ; GFX11.5
+                                                     s_getreg_b32 s0, hwreg(HW_REG_STATUS, 20, 2)
+                                                     ; GFX12
+                                                     s_getreg_b32 s0, hwreg(HW_REG_WAVE_STATE_PRIV, 16, 2)
+                                                     ; GFX13
+                                                     s_getreg_b32 s0, hwreg(HW_REG_WAVE_STATUS, 20, 2)
+                                                     ; all generations
+                                                     s_cmp_lg_u32 s0, 0
+
   llvm.amdgcn.readfirstlane                        Provides direct access to v_readfirstlane_b32. Returns the value in
                                                    the lowest active lane of the input operand. Currently implemented
                                                    for i16, i32, float, half, bfloat, <2 x i16>, <2 x half>, <2 x bfloat>,
@@ -20990,6 +21036,35 @@ address. The *spill table* itself represents a set of 32-bit values
 managed by the PAL runtime in GPU-accessible memory that can be made
 indirectly accessible to a hardware shader.
 
+.. _amdgpu-amdpal-trap-handler-abi:
+
+Trap Handler ABI
+~~~~~~~~~~~~~~~~
+
+For code objects generated for the AMDPAL OS, the runtime installs a trap
+handler that supports the ``s_trap`` instruction only when the ``trap-handler``
+target feature is enabled. The handler must be resumable and must define trap
+ID 3 as the LLVM debug trap. For usage see
+:ref:`amdgpu-trap-handler-for-amdpal-os-table`.
+
+  .. table:: AMDGPU Trap Handler for AMDPAL OS
+     :name: amdgpu-trap-handler-for-amdpal-os-table
+
+     ================== =============== =============== ======================================
+     Usage              Code Sequence   Trap Handler    Description
+                                        Inputs
+     ================== =============== =============== ======================================
+     ``llvm.trap``      ``s_endpgm``    *none*          Causes the wavefront to be terminated.
+     ``llvm.debugtrap`` ``s_trap 0x03`` *none*          Causes the wave to enter the PAL debug
+                                                        trap handler. Execution resumes after
+                                                        the configured debug action completes.
+     ================== =============== =============== ======================================
+
+The ``trap-handler`` feature is not enabled by the AMDPAL target triple; a
+frontend must enable it only when the PAL runtime implements this ABI. If the
+feature is disabled, ``llvm.debugtrap`` produces a compiler warning and no trap
+instruction.
+
 Unspecified OS
 --------------
 
@@ -20999,9 +21074,9 @@ empty (see :ref:`amdgpu-target-triples`).
 Trap Handler ABI
 ~~~~~~~~~~~~~~~~
 
-For code objects generated by AMDGPU backend for non-amdhsa OS, the runtime does
-not install a trap handler. The ``llvm.trap`` and ``llvm.debugtrap``
-instructions are handled as follows:
+For code objects whose target OS has no recognized trap-handler ABI, the
+runtime does not install a trap handler. The ``llvm.trap`` and
+``llvm.debugtrap`` instructions are handled as follows:
 
   .. table:: AMDGPU Trap Handler for Non-AMDHSA OS
      :name: amdgpu-trap-handler-for-non-amdhsa-os-table
diff --git a/llvm/docs/LangRef.md b/llvm/docs/LangRef.md
index ee1600b96f1dd..9f2ab993f3a80 100644
--- a/llvm/docs/LangRef.md
+++ b/llvm/docs/LangRef.md
@@ -26263,6 +26263,36 @@ This intrinsic is lowered to code which is intended to cause an
 execution trap with the intention of requesting the attention of a
 debugger.
 
+(llvm.is.debugging.enabled)=
+
+#### '`llvm.is.debugging.enabled`' Intrinsic
+
+##### Syntax:
+
+```llvm
+declare noundef i1 @llvm.is.debugging.enabled() nomerge memory(inaccessiblemem: readwrite)
+```
+
+##### Overview:
+
+The '`llvm.is.debugging.enabled`' intrinsic returns whether debugging is enabled
+for the target-defined execution context of the current invocation.
+
+##### Arguments:
+
+None.
+
+##### Semantics:
+
+Each call observes the current debugging-enabled state. Calls are distinct
+observations: LLVM must not assume separate calls return the same value, and a
+call may not be removed when its result is unused, commoned with another call,
+or reordered with respect to one.
+
+On targets that do not support querying debugging-enabled state, the intrinsic
+returns `false` and performs no observation.
+
+
 (llvm.ubsantrap)=
 
 #### '`llvm.ubsantrap`' Intrinsic
diff --git a/llvm/include/llvm/IR/Intrinsics.td b/llvm/include/llvm/IR/Intrinsics.td
index c20ea64c7eef4..b14375e277e7a 100644
--- a/llvm/include/llvm/IR/Intrinsics.td
+++ b/llvm/include/llvm/IR/Intrinsics.td
@@ -2059,6 +2059,12 @@ def int_trap : Intrinsic<[], [],
                ClangBuiltin<"__builtin_trap">;
 def int_debugtrap : Intrinsic<[]>,
                     ClangBuiltin<"__builtin_debugtrap">;
+// Query whether debugging is enabled for the current execution context.
+def int_is_debugging_enabled :
+  DefaultAttrsIntrinsic<[llvm_i1_ty], [],
+                        [IntrInaccessibleMemOnly,
+                         IntrNoMerge,
+                         NoUndef<RetIndex>]>;
 def int_ubsantrap : Intrinsic<[], [llvm_i8_ty],
                               [IntrNoReturn, IntrCold, ImmArg<ArgIndex<0>>,
                                IntrInaccessibleMemOnly, IntrWriteMem]>;
diff --git a/llvm/include/llvm/Target/TargetMachine.h b/llvm/include/llvm/Target/TargetMachine.h
index a6b73d636dc27..ea26be5e48bfb 100644
--- a/llvm/include/llvm/Target/TargetMachine.h
+++ b/llvm/include/llvm/Target/TargetMachine.h
@@ -567,6 +567,13 @@ class LLVM_ABI TargetMachine {
   /// this function returns false, the intrinsic will be supported generically
   /// but without loop detection support.
   virtual bool canLowerCondLoop() const { return false; }
+
+  /// Returns whether this target takes responsibility for lowering
+  /// llvm.is.debugging.enabled. If false, generic lowering replaces the intrinsic
+  /// with false. If true, the target must handle every supported subtarget,
+  /// including replacing the intrinsic with false on subtargets without native
+  /// lowering.
+  virtual bool canLowerIsDebuggingEnabled() const { return false; }
 };
 
 } // end namespace llvm
diff --git a/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp b/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp
index 24788148b9f15..4ec866d7b973f 100644
--- a/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp
+++ b/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp
@@ -815,6 +815,14 @@ bool PreISelIntrinsicLowering::lowerIntrinsics(Module &M) const {
     case Intrinsic::protected_field_ptr:
       Changed |= expandProtectedFieldPtr(F);
       break;
+    case Intrinsic::is_debugging_enabled:
+      if (!TM || !TM->canLowerIsDebuggingEnabled())
+        Changed |= forEachCall(F, [](CallInst *CI) {
+          CI->replaceAllUsesWith(ConstantInt::getFalse(CI->getContext()));
+          CI->eraseFromParent();
+          return true;
+        });
+      break;
     case Intrinsic::cond_loop:
       if (!TM->canLowerCondLoop())
         Changed |= expandCondLoop(F);
diff --git a/llvm/lib/Target/AMDGPU/AMDGPU.td b/llvm/lib/Target/AMDGPU/AMDGPU.td
index ed2ee1ff70f4f..17b22beebcbb9 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPU.td
+++ b/llvm/lib/Target/AMDGPU/AMDGPU.td
@@ -149,6 +149,12 @@ defm TrapHandler: AMDGPUSubtargetFeature<"trap-handler",
   /*GenPredicate=*/0, /*GenAssemblerPredicate=*/1, [], InlineIgnore
 >;
 
+defm CDBGSysOrUserBranch : AMDGPUSubtargetFeature<
+  "cdbg-sys-or-user-branch",
+  "Support querying debugging state with s_cbranch_cdbgsys_or_user",
+  /*GenPredicate=*/1, /*GenAssemblerPredicate=*/0
+>;
+
 defm UnalignedScratchAccess : AMDGPUSubtargetFeature<"unaligned-scratch-access",
   "Support unaligned scratch loads and stores",
   /*GenPredicate=*/1, /*GenAssemblerPredicate=*/1, [], InlineIgnore
@@ -2232,7 +2238,8 @@ def FeatureISAVersion11_0_3 : FeatureSet<
 
 def FeatureISAVersion11_5_Common : FeatureSet<
   !listconcat(FeatureISAVersion11_Common.Features,
-    [FeatureSALUFloatInsts,
+    [FeatureCDBGSysOrUserBranch,
+     FeatureSALUFloatInsts,
      FeatureDPPSrc1SGPR,
      FeatureRequiredExportPriority,
      FeatureDot5Insts,
@@ -2275,6 +2282,7 @@ def FeatureISAVersion11_7_Generic: FeatureSet<
 def FeatureISAVersion12 : FeatureSet<
   [FeatureGFX12,
    FeatureSupportsWave64, FeatureSupportsWGP,
+   FeatureCDBGSysOrUserBranch,
    FeatureBackOffBarrier,
    FeatureAddressableLocalMemorySize65536,
    FeatureHalfAddressablePhysicalLocalMemory,
@@ -2343,6 +2351,7 @@ def FeatureISAVersion12 : FeatureSet<
 def FeatureISAVersion12_50_Common : FeatureSet<
   [FeatureGFX12,
    FeatureGFX1250Insts,
+   FeatureCDBGSysOrUserBranch,
    FeatureBackOffBarrier,
    FeatureRequiresAlignedVGPRs,
    FeatureCuMode,
@@ -2536,6 +2545,7 @@ def FeatureISAVersion12_5_Generic: FeatureSet<
 def FeatureISAVersion13 : FeatureSet<
   [FeatureGFX13,
    FeatureGFX1250Insts,
+   FeatureCDBGSysOrUserBranch,
    FeatureAddressableLocalMemorySize196608,
    Feature64BitLiterals,
    FeatureLDSBankCount32,
diff --git a/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp b/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
index 7c1a26f761c96..29a13e9a97c23 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
@@ -4804,47 +4804,77 @@ static bool isNot(const MachineRegisterInfo &MRI, const MachineInstr &MI) {
   return ConstVal == -1;
 }
 
-// Return the use branch instruction, otherwise null if the usage is invalid.
-static MachineInstr *
-verifyCFIntrinsic(MachineInstr &MI, MachineRegisterInfo &MRI, MachineInstr *&Br,
-                  MachineBasicBlock *&UncondBrTarget, bool &Negated) {
+namespace {
+struct CFIntrinsicBranchMatch {
+  MachineInstr *Negation = nullptr;
+  MachineInstr *CondBr = nullptr;
+  MachineInstr *UncondBr = nullptr;
+
+  MachineBasicBlock *ConditionTrueTarget = nullptr;
+  MachineBasicBlock *ConditionFalseTarget = nullptr;
+
+  bool isNegated() const { return Negation != nullptr; }
+
+  // Retarget the explicit branch for the fallthrough edge, or materialize
+  // the branch when that edge was represented by layout fallthrough. The
+  // builder must be positioned at the replacement branch.
+  void redirectFallthroughEdge(MachineIRBuilder &B,
+                               MachineBasicBlock &Target) const {
+    if (UncondBr)
+      UncondBr->getOperand(0).setMBB(&Target);
+    else
+      B.buildBr(Target);
+  }
+
+  void eraseDeadNegation(MachineRegisterInfo &MRI) {
+    if (Negation)
+      eraseInstr(*Negation, MRI);
+  }
+};
+} // namespace
+
+static std::optional<CFIntrinsicBranchMatch>
+matchCFIntrinsicBranchUse(MachineInstr &MI, MachineRegisterInfo &MRI) {
   Register CondDef = MI.getOperand(0).getReg();
   if (!MRI.hasOneNonDBGUse(CondDef))
-    return nullptr;
+    return std::nullopt;
 
   MachineBasicBlock *Parent = MI.getParent();
   MachineInstr *UseMI = &*MRI.use_instr_nodbg_begin(CondDef);
+  MachineInstr *Negation = nullptr;
 
   if (isNot(MRI, *UseMI)) {
+    Negation = UseMI;
     Register NegatedCond = UseMI->getOperand(0).getReg();
     if (!MRI.hasOneNonDBGUse(NegatedCond))
-      return nullptr;
-
-    // We're deleting the def of this value, so we need to remove it.
-    eraseInstr(*UseMI, MRI);
+      return std::nullopt;
 
     UseMI = &*MRI.use_instr_nodbg_begin(NegatedCond);
-    Negated = true;
   }
 
   if (UseMI->getParent() != Parent || UseMI->getOpcode() != AMDGPU::G_BRCOND)
-    return nullptr;
+    return std::nullopt;
 
   // Make sure the cond br is followed by a G_BR, or is the last instruction.
+  MachineInstr *UncondBr = nullptr;
+  MachineBasicBlock *OtherTarget = nullptr;
   MachineBasicBlock::iterator Next = std::next(UseMI->getIterator());
   if (Next == Parent->end()) {
     MachineFunction::iterator NextMBB = std::next(Parent->getIterator());
     if (NextMBB == Parent->getParent()->end()) // Illegal intrinsic use.
-      return nullptr;
-    UncondBrTarget = &*NextMBB;
+      return std::nullopt;
+    OtherTarget = &*NextMBB;
   } else {
     if (Next->getOpcode() != AMDGPU::G_BR)
-      return nullptr;
-    Br = &*Next;
-    UncondBrTarget = Br->getOperand(0).getMBB();
+      return std::nullopt;
+    UncondBr = &*Next;
+    OtherTarget = UncondBr->getOperand(0).getMBB();
   }
 
-  return UseMI;
+  MachineBasicBlock *TakenTarget = UseMI->getOperand(1).getMBB();
+  return CFIntrinsicBranchMatch{Negation, UseMI, UncondBr,
+                                 Negation ? OtherTarget : TakenTarget,
+                                 Negation ? TakenTarget : OtherTarget};
 }
 
 void AMDGPULegalizerInfo::buildLoadInputValue(Register DstReg,
@@ -7867,10 +7897,12 @@ bool AMDGPULegalizerInfo::legalizeTrapHsa(MachineInstr &MI,
 bool AMDGPULegalizerInfo::legalizeDebugTrap(MachineInstr &MI,
                                             MachineRegisterInfo &MRI,
                                             MachineIRBuilder &B) const {
-  // Is non-HSA path or trap-handler disabled? Then, report a warning
-  // accordingly
+  // Is this an unsupported ABI, or is the trap handler disabled? Then report
+  // a warning accordingly.
+  const GCNSubtarget::TrapHandlerAbi Abi = ST.getTrapHandlerAbi();
   if (!ST.hasTrapHandler() ||
-      ST.getTrapHandlerAbi() != GCNSubtarget::TrapHandlerAbi::AMDHSA) {
+      (Abi != GCNSubtarget::TrapHandlerAbi::AMDHSA &&
+       Abi != GCNSubtarget::TrapHandlerAbi::AMDPAL)) {
     Function &Fn = B.getMF().getFunction();
     Fn.getContext().diagnose(DiagnosticInfoUnsupported(
         Fn, "debugtrap handler not supported", MI.getDebugLoc(), DS_Warning));
@@ -8185,6 +8217,44 @@ bool AMDGPULegalizerInfo::legalizeIntrinsic(LegalizerHelper &Helper,
   // Replace the use G_BRCOND with the exec manipulate and branch pseudos.
   auto IntrID = cast<GIntrinsic>(MI).getIntrinsicID();
   switch (IntrID) {
+  case Intrinsic::is_debugging_enabled: {
+    auto Match = matchCFIntrinsicBranchUse(MI, MRI);
+    bool CannotFuse =
+        !Match || any_of(make_range(std::next(MI.getIterator()),
+                                    Match->CondBr->getIterator()),
+                         [](const MachineInstr &Between) {
+                           return !Between.isMetaInstruction() &&
+                                  (Between.mayLoadOrStore() ||
+                                   Between.hasUnmodeledSideEffects());
+                         });
+    if (CannotFuse) {
+      auto Bits =
+          B.buildIntrinsic(Intrinsic::amdgcn_s_getreg, {LLT::scalar(32)})
+              .addImm(AMDGPU::Hwreg::getDebuggingEnabledHwregImm(ST));
+      Bits->setFlag(MachineInstr::NoMerge);
+      B.buildICmp(CmpInst::ICMP_NE, MI.getOperand(0).getReg(), Bits.getReg(0),
+                  B.buildConstant(LLT::scalar(32), 0));
+      MI.eraseFromParent();
+      return true;
+    }
+
+    B.setInsertPt(*Match->CondBr->getParent(), Match->CondBr->getIterator());
+    B.setDebugLoc(Match->CondBr->getDebugLoc());
+    MachineInstrBuilder CDBGBranch =
+        B.buildInstr(AMDGPU::S_CBRANCH_CDBGSYS_OR_USER)
+            .addMBB(Match->ConditionTrueTarget);
+    CDBGBranch->setFlag(MachineInstr::NoMerge);
+
+    if (Match->isNegated())
+      Match->redirectFallthroughEdge(B, *Match->ConditionFalseTarget);
+
+    Register Cond = MI.getOperand(0).getReg();
+    MRI.markUsesInDebugValueAsUndef(Cond);
+    Match->eraseDeadNegation(MRI);
+    MI.eraseFromParent();
+    Match->CondBr->eraseFromParent();
+    return true;
+  }
   case Intrinsic::amdgcn_icmp: {
     // amdgcn.icmp(i1 src0, i1 0, NE) -> ballot(src0)
     // This is the only valid form of amdgcn.icmp with i1 inputs.
@@ -8230,80 +8300,58 @@ bool AMDGPULegalizerInfo::legalizeIntrinsic(LegalizerHelper &Helper,
     return true;
   case Intrinsic::amdgcn_if:
   case Intrinsic::amdgcn_else: {
-    MachineInstr *Br = nullptr;
-    MachineBasicBlock *UncondBrTarget = nullptr;
-    bool Negated = false;
-    if (MachineInstr *BrCond =
-            verifyCFIntrinsic(MI, MRI, Br, UncondBrTarget, Negated)) {
-      const SIRegisterInfo *TRI
-        = static_cast<const SIRegisterInfo *>(MRI.getTargetRegisterInfo());
+    if (auto Match = matchCFIntrinsicBranchUse(MI, MRI)) {
+      const SIRegisterInfo *TRI =
+          static_cast<const SIRegisterInfo *>(MRI.getTargetRegisterInfo());
 
       Register Def = MI.getOperand(1).getReg();
       Register Use = MI.getOperand(3).getReg();
 
-      MachineBasicBlock *CondBrTarget = BrCond->getOperand(1).getMBB();
-
-      if (Negated)
-        std::swap(CondBrTarget, UncondBrTarget);
-
-      B.setInsertPt(B.getMBB(), BrCond->getIterator());
+      B.setInsertPt(B.getMBB(), Match->CondBr->getIterator());
       if (IntrID == Intrinsic::amdgcn_if) {
         B.buildInstr(AMDGPU::SI_IF)
-          .addDef(Def)
-          .addUse(Use)
-          .addMBB(UncondBrTarget);
+        ...
[truncated]

``````````

</details>


https://github.com/llvm/llvm-project/pull/219179


More information about the llvm-commits mailing list