[llvm] [AMDGPU] Add AMDPAL support for llvm.debugtrap (PR #219179)
via llvm-commits
llvm-commits at lists.llvm.org
Thu Aug 27 04:45:41 PDT 2026
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>,
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>,
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>,
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>,
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>
Message-ID:
In-Reply-To: <llvm.org/llvm/llvm-project/pull/219179 at github.com>
llvmorg-github-actions[bot] wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-llvm-globalisel
Author: Alexander Hück (ahueck)
<details>
<summary>Changes</summary>
This is PR 4 of 4 in a series adding the `llvm.is.debugging.enabled` intrinsic, lowering it for AMDGPU targets, and enabling debug traps for AMDPAL, see [[[RFC] Introduce `llvm.is.debugging.enabled` intrinsic](https://discourse.llvm.org/t/rfc-introduce-llvm-is-debugging-enabled-intrinsic/91676)
## Changes of this PR
- Add an AMDPAL trap-handler ABI and lower `llvm.debugtrap` to s_trap 3 when the trap-handler target feature is enabled.
- Preserve the existing warning and omit the trap instruction when trap-handler support is disabled. `llvm.trap` continues to lower to s_endpgm.
## PR stack
1. [[AMDGPU][GlobalISel][SelectionDAG] Refactor control-flow intrinsic branch matching](https://github.com/llvm/llvm-project/pull/219173)
2. [[IR][CodeGen] Add llvm.is.debugging.enabled intrinsic](https://github.com/llvm/llvm-project/pull/219175)
3. [[AMDGPU][GFX12/GFX13] Add cdbg branch support and lower llvm.is.debugging.enabled](https://github.com/llvm/llvm-project/pull/219178)
4. [[AMDGPU] Add AMDPAL support for llvm.debugtrap](https://github.com/llvm/llvm-project/pull/219179)
*Disclaimer*: Agentic AI assisted development.
---
Patch is 95.34 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/219179.diff
34 Files Affected:
- (modified) llvm/docs/AMDGPUUsage.rst (+78-3)
- (modified) llvm/docs/LangRef.md (+30)
- (modified) llvm/include/llvm/IR/Intrinsics.td (+6)
- (modified) llvm/include/llvm/Target/TargetMachine.h (+7)
- (modified) llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp (+8)
- (modified) llvm/lib/Target/AMDGPU/AMDGPU.td (+11-1)
- (modified) llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp (+114-66)
- (modified) llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp (+16)
- (modified) llvm/lib/Target/AMDGPU/AMDGPUTargetMachine.h (+2)
- (modified) llvm/lib/Target/AMDGPU/GCNSubtarget.h (+7-1)
- (modified) llvm/lib/Target/AMDGPU/SIISelLowering.cpp (+134-43)
- (modified) llvm/lib/Target/AMDGPU/SOPInstructions.td (+12-8)
- (modified) llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.cpp (+10)
- (modified) llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.h (+4)
- (modified) llvm/test/Analysis/UniformityAnalysis/AMDGPU/always_uniform.ll (+28)
- (added) llvm/test/Assembler/is-debugging-enabled.ll (+11)
- (modified) llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-amdgcn.if-invalid.mir (+26-1)
- (added) llvm/test/CodeGen/AMDGPU/debugtrap-amdpal.ll (+54)
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled-divergent-exec-guard.ll (+240)
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled-shapes.ll (+449)
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled-unsupported.ll (+32)
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled.ll (+257)
- (added) llvm/test/CodeGen/NVPTX/is-debugging-enabled.ll (+11)
- (added) llvm/test/CodeGen/X86/is-debugging-enabled.ll (+12)
- (modified) llvm/test/MC/AMDGPU/gfx12_asm_sopp.s (+24)
- (modified) llvm/test/MC/AMDGPU/gfx12_unsupported.s (-12)
- (modified) llvm/test/MC/AMDGPU/gfx13_asm_sopp.s (+24)
- (added) llvm/test/Transforms/DCE/is-debugging-enabled.ll (+12)
- (added) llvm/test/Transforms/EarlyCSE/is-debugging-enabled.ll (+17)
- (added) llvm/test/Transforms/LICM/is-debugging-enabled.ll (+25)
- (added) llvm/test/Transforms/LoopUnroll/is-debugging-enabled.ll (+30)
- (added) llvm/test/Transforms/PreISelIntrinsicLowering/is-debugging-enabled.ll (+17)
- (added) llvm/test/Transforms/SimplifyCFG/is-debugging-enabled.ll (+58)
- (added) llvm/test/Verifier/is-debugging-enabled.ll (+15)
``````````diff
diff --git a/llvm/docs/AMDGPUUsage.rst b/llvm/docs/AMDGPUUsage.rst
index 40a2820b9e04f..f066f48f96b38 100644
--- a/llvm/docs/AMDGPUUsage.rst
+++ b/llvm/docs/AMDGPUUsage.rst
@@ -1709,6 +1709,52 @@ The AMDGPU backend implements the following LLVM IR intrinsics.
The format is a 64-bit concatenation of the MODE and TRAPSTS registers.
:ref:`llvm.set.fpenv<int_set_fpenv>` Sets the floating point environment to the specified state.
+
+ :ref:`llvm.is.debugging.enabled <llvm.is.debugging.enabled>`
+ Supported on GFX11.5, GFX12, and GFX13 targets. Other
+ subtargets lower the result to ``false``.
+
+ The target-defined execution context is the current wave.
+ The result is uniform across the active lanes of that
+ wave, including when the intrinsic is executed in
+ divergent control flow. Each call remains a distinct
+ observation of the wave's debugging-enabled state.
+
+ A wave executes in debugging mode when either the
+ ``COND_DBG_SYS`` or ``COND_DBG_USER`` bit is set.
+ ``COND_DBG_SYS`` reflects a system-wide debugger
+ attach, while ``COND_DBG_USER`` reflects per-dispatch
+ launch control, configured through
+ :ref:`CDBG_USER <amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx13-table>`
+ in the kernel descriptor. The driver sets these
+ bits, and provides an interface allowing the user
+ mode runtime or an external debugger to control
+ the setting.
+
+ The intrinsic lowers in one of two forms.
+ When the query feeds a single conditional branch in the
+ same basic block, with no intervening observable
+ operations, it may fuse into a single
+ ``s_cbranch_cdbgsys_or_user``. Branch-hint and
+ negated-condition forms are supported; other uses of
+ the query value prevent fusion.
+
+ Any other use materializes an ``i1`` by reading the
+ adjacent ``COND_DBG_USER`` and ``COND_DBG_SYS`` bits
+ and testing them against zero. The register holding
+ them differs by generation:
+
+ .. code-block:: none
+
+ ; GFX11.5
+ s_getreg_b32 s0, hwreg(HW_REG_STATUS, 20, 2)
+ ; GFX12
+ s_getreg_b32 s0, hwreg(HW_REG_WAVE_STATE_PRIV, 16, 2)
+ ; GFX13
+ s_getreg_b32 s0, hwreg(HW_REG_WAVE_STATUS, 20, 2)
+ ; all generations
+ s_cmp_lg_u32 s0, 0
+
llvm.amdgcn.readfirstlane Provides direct access to v_readfirstlane_b32. Returns the value in
the lowest active lane of the input operand. Currently implemented
for i16, i32, float, half, bfloat, <2 x i16>, <2 x half>, <2 x bfloat>,
@@ -20990,6 +21036,35 @@ address. The *spill table* itself represents a set of 32-bit values
managed by the PAL runtime in GPU-accessible memory that can be made
indirectly accessible to a hardware shader.
+.. _amdgpu-amdpal-trap-handler-abi:
+
+Trap Handler ABI
+~~~~~~~~~~~~~~~~
+
+For code objects generated for the AMDPAL OS, the runtime installs a trap
+handler that supports the ``s_trap`` instruction only when the ``trap-handler``
+target feature is enabled. The handler must be resumable and must define trap
+ID 3 as the LLVM debug trap. For usage see
+:ref:`amdgpu-trap-handler-for-amdpal-os-table`.
+
+ .. table:: AMDGPU Trap Handler for AMDPAL OS
+ :name: amdgpu-trap-handler-for-amdpal-os-table
+
+ ================== =============== =============== ======================================
+ Usage Code Sequence Trap Handler Description
+ Inputs
+ ================== =============== =============== ======================================
+ ``llvm.trap`` ``s_endpgm`` *none* Causes the wavefront to be terminated.
+ ``llvm.debugtrap`` ``s_trap 0x03`` *none* Causes the wave to enter the PAL debug
+ trap handler. Execution resumes after
+ the configured debug action completes.
+ ================== =============== =============== ======================================
+
+The ``trap-handler`` feature is not enabled by the AMDPAL target triple; a
+frontend must enable it only when the PAL runtime implements this ABI. If the
+feature is disabled, ``llvm.debugtrap`` produces a compiler warning and no trap
+instruction.
+
Unspecified OS
--------------
@@ -20999,9 +21074,9 @@ empty (see :ref:`amdgpu-target-triples`).
Trap Handler ABI
~~~~~~~~~~~~~~~~
-For code objects generated by AMDGPU backend for non-amdhsa OS, the runtime does
-not install a trap handler. The ``llvm.trap`` and ``llvm.debugtrap``
-instructions are handled as follows:
+For code objects whose target OS has no recognized trap-handler ABI, the
+runtime does not install a trap handler. The ``llvm.trap`` and
+``llvm.debugtrap`` instructions are handled as follows:
.. table:: AMDGPU Trap Handler for Non-AMDHSA OS
:name: amdgpu-trap-handler-for-non-amdhsa-os-table
diff --git a/llvm/docs/LangRef.md b/llvm/docs/LangRef.md
index ee1600b96f1dd..9f2ab993f3a80 100644
--- a/llvm/docs/LangRef.md
+++ b/llvm/docs/LangRef.md
@@ -26263,6 +26263,36 @@ This intrinsic is lowered to code which is intended to cause an
execution trap with the intention of requesting the attention of a
debugger.
+(llvm.is.debugging.enabled)=
+
+#### '`llvm.is.debugging.enabled`' Intrinsic
+
+##### Syntax:
+
+```llvm
+declare noundef i1 @llvm.is.debugging.enabled() nomerge memory(inaccessiblemem: readwrite)
+```
+
+##### Overview:
+
+The '`llvm.is.debugging.enabled`' intrinsic returns whether debugging is enabled
+for the target-defined execution context of the current invocation.
+
+##### Arguments:
+
+None.
+
+##### Semantics:
+
+Each call observes the current debugging-enabled state. Calls are distinct
+observations: LLVM must not assume separate calls return the same value, and a
+call may not be removed when its result is unused, commoned with another call,
+or reordered with respect to one.
+
+On targets that do not support querying debugging-enabled state, the intrinsic
+returns `false` and performs no observation.
+
+
(llvm.ubsantrap)=
#### '`llvm.ubsantrap`' Intrinsic
diff --git a/llvm/include/llvm/IR/Intrinsics.td b/llvm/include/llvm/IR/Intrinsics.td
index c20ea64c7eef4..b14375e277e7a 100644
--- a/llvm/include/llvm/IR/Intrinsics.td
+++ b/llvm/include/llvm/IR/Intrinsics.td
@@ -2059,6 +2059,12 @@ def int_trap : Intrinsic<[], [],
ClangBuiltin<"__builtin_trap">;
def int_debugtrap : Intrinsic<[]>,
ClangBuiltin<"__builtin_debugtrap">;
+// Query whether debugging is enabled for the current execution context.
+def int_is_debugging_enabled :
+ DefaultAttrsIntrinsic<[llvm_i1_ty], [],
+ [IntrInaccessibleMemOnly,
+ IntrNoMerge,
+ NoUndef<RetIndex>]>;
def int_ubsantrap : Intrinsic<[], [llvm_i8_ty],
[IntrNoReturn, IntrCold, ImmArg<ArgIndex<0>>,
IntrInaccessibleMemOnly, IntrWriteMem]>;
diff --git a/llvm/include/llvm/Target/TargetMachine.h b/llvm/include/llvm/Target/TargetMachine.h
index a6b73d636dc27..ea26be5e48bfb 100644
--- a/llvm/include/llvm/Target/TargetMachine.h
+++ b/llvm/include/llvm/Target/TargetMachine.h
@@ -567,6 +567,13 @@ class LLVM_ABI TargetMachine {
/// this function returns false, the intrinsic will be supported generically
/// but without loop detection support.
virtual bool canLowerCondLoop() const { return false; }
+
+ /// Returns whether this target takes responsibility for lowering
+ /// llvm.is.debugging.enabled. If false, generic lowering replaces the intrinsic
+ /// with false. If true, the target must handle every supported subtarget,
+ /// including replacing the intrinsic with false on subtargets without native
+ /// lowering.
+ virtual bool canLowerIsDebuggingEnabled() const { return false; }
};
} // end namespace llvm
diff --git a/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp b/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp
index 24788148b9f15..4ec866d7b973f 100644
--- a/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp
+++ b/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp
@@ -815,6 +815,14 @@ bool PreISelIntrinsicLowering::lowerIntrinsics(Module &M) const {
case Intrinsic::protected_field_ptr:
Changed |= expandProtectedFieldPtr(F);
break;
+ case Intrinsic::is_debugging_enabled:
+ if (!TM || !TM->canLowerIsDebuggingEnabled())
+ Changed |= forEachCall(F, [](CallInst *CI) {
+ CI->replaceAllUsesWith(ConstantInt::getFalse(CI->getContext()));
+ CI->eraseFromParent();
+ return true;
+ });
+ break;
case Intrinsic::cond_loop:
if (!TM->canLowerCondLoop())
Changed |= expandCondLoop(F);
diff --git a/llvm/lib/Target/AMDGPU/AMDGPU.td b/llvm/lib/Target/AMDGPU/AMDGPU.td
index ed2ee1ff70f4f..17b22beebcbb9 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPU.td
+++ b/llvm/lib/Target/AMDGPU/AMDGPU.td
@@ -149,6 +149,12 @@ defm TrapHandler: AMDGPUSubtargetFeature<"trap-handler",
/*GenPredicate=*/0, /*GenAssemblerPredicate=*/1, [], InlineIgnore
>;
+defm CDBGSysOrUserBranch : AMDGPUSubtargetFeature<
+ "cdbg-sys-or-user-branch",
+ "Support querying debugging state with s_cbranch_cdbgsys_or_user",
+ /*GenPredicate=*/1, /*GenAssemblerPredicate=*/0
+>;
+
defm UnalignedScratchAccess : AMDGPUSubtargetFeature<"unaligned-scratch-access",
"Support unaligned scratch loads and stores",
/*GenPredicate=*/1, /*GenAssemblerPredicate=*/1, [], InlineIgnore
@@ -2232,7 +2238,8 @@ def FeatureISAVersion11_0_3 : FeatureSet<
def FeatureISAVersion11_5_Common : FeatureSet<
!listconcat(FeatureISAVersion11_Common.Features,
- [FeatureSALUFloatInsts,
+ [FeatureCDBGSysOrUserBranch,
+ FeatureSALUFloatInsts,
FeatureDPPSrc1SGPR,
FeatureRequiredExportPriority,
FeatureDot5Insts,
@@ -2275,6 +2282,7 @@ def FeatureISAVersion11_7_Generic: FeatureSet<
def FeatureISAVersion12 : FeatureSet<
[FeatureGFX12,
FeatureSupportsWave64, FeatureSupportsWGP,
+ FeatureCDBGSysOrUserBranch,
FeatureBackOffBarrier,
FeatureAddressableLocalMemorySize65536,
FeatureHalfAddressablePhysicalLocalMemory,
@@ -2343,6 +2351,7 @@ def FeatureISAVersion12 : FeatureSet<
def FeatureISAVersion12_50_Common : FeatureSet<
[FeatureGFX12,
FeatureGFX1250Insts,
+ FeatureCDBGSysOrUserBranch,
FeatureBackOffBarrier,
FeatureRequiresAlignedVGPRs,
FeatureCuMode,
@@ -2536,6 +2545,7 @@ def FeatureISAVersion12_5_Generic: FeatureSet<
def FeatureISAVersion13 : FeatureSet<
[FeatureGFX13,
FeatureGFX1250Insts,
+ FeatureCDBGSysOrUserBranch,
FeatureAddressableLocalMemorySize196608,
Feature64BitLiterals,
FeatureLDSBankCount32,
diff --git a/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp b/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
index 7c1a26f761c96..29a13e9a97c23 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
@@ -4804,47 +4804,77 @@ static bool isNot(const MachineRegisterInfo &MRI, const MachineInstr &MI) {
return ConstVal == -1;
}
-// Return the use branch instruction, otherwise null if the usage is invalid.
-static MachineInstr *
-verifyCFIntrinsic(MachineInstr &MI, MachineRegisterInfo &MRI, MachineInstr *&Br,
- MachineBasicBlock *&UncondBrTarget, bool &Negated) {
+namespace {
+struct CFIntrinsicBranchMatch {
+ MachineInstr *Negation = nullptr;
+ MachineInstr *CondBr = nullptr;
+ MachineInstr *UncondBr = nullptr;
+
+ MachineBasicBlock *ConditionTrueTarget = nullptr;
+ MachineBasicBlock *ConditionFalseTarget = nullptr;
+
+ bool isNegated() const { return Negation != nullptr; }
+
+ // Retarget the explicit branch for the fallthrough edge, or materialize
+ // the branch when that edge was represented by layout fallthrough. The
+ // builder must be positioned at the replacement branch.
+ void redirectFallthroughEdge(MachineIRBuilder &B,
+ MachineBasicBlock &Target) const {
+ if (UncondBr)
+ UncondBr->getOperand(0).setMBB(&Target);
+ else
+ B.buildBr(Target);
+ }
+
+ void eraseDeadNegation(MachineRegisterInfo &MRI) {
+ if (Negation)
+ eraseInstr(*Negation, MRI);
+ }
+};
+} // namespace
+
+static std::optional<CFIntrinsicBranchMatch>
+matchCFIntrinsicBranchUse(MachineInstr &MI, MachineRegisterInfo &MRI) {
Register CondDef = MI.getOperand(0).getReg();
if (!MRI.hasOneNonDBGUse(CondDef))
- return nullptr;
+ return std::nullopt;
MachineBasicBlock *Parent = MI.getParent();
MachineInstr *UseMI = &*MRI.use_instr_nodbg_begin(CondDef);
+ MachineInstr *Negation = nullptr;
if (isNot(MRI, *UseMI)) {
+ Negation = UseMI;
Register NegatedCond = UseMI->getOperand(0).getReg();
if (!MRI.hasOneNonDBGUse(NegatedCond))
- return nullptr;
-
- // We're deleting the def of this value, so we need to remove it.
- eraseInstr(*UseMI, MRI);
+ return std::nullopt;
UseMI = &*MRI.use_instr_nodbg_begin(NegatedCond);
- Negated = true;
}
if (UseMI->getParent() != Parent || UseMI->getOpcode() != AMDGPU::G_BRCOND)
- return nullptr;
+ return std::nullopt;
// Make sure the cond br is followed by a G_BR, or is the last instruction.
+ MachineInstr *UncondBr = nullptr;
+ MachineBasicBlock *OtherTarget = nullptr;
MachineBasicBlock::iterator Next = std::next(UseMI->getIterator());
if (Next == Parent->end()) {
MachineFunction::iterator NextMBB = std::next(Parent->getIterator());
if (NextMBB == Parent->getParent()->end()) // Illegal intrinsic use.
- return nullptr;
- UncondBrTarget = &*NextMBB;
+ return std::nullopt;
+ OtherTarget = &*NextMBB;
} else {
if (Next->getOpcode() != AMDGPU::G_BR)
- return nullptr;
- Br = &*Next;
- UncondBrTarget = Br->getOperand(0).getMBB();
+ return std::nullopt;
+ UncondBr = &*Next;
+ OtherTarget = UncondBr->getOperand(0).getMBB();
}
- return UseMI;
+ MachineBasicBlock *TakenTarget = UseMI->getOperand(1).getMBB();
+ return CFIntrinsicBranchMatch{Negation, UseMI, UncondBr,
+ Negation ? OtherTarget : TakenTarget,
+ Negation ? TakenTarget : OtherTarget};
}
void AMDGPULegalizerInfo::buildLoadInputValue(Register DstReg,
@@ -7867,10 +7897,12 @@ bool AMDGPULegalizerInfo::legalizeTrapHsa(MachineInstr &MI,
bool AMDGPULegalizerInfo::legalizeDebugTrap(MachineInstr &MI,
MachineRegisterInfo &MRI,
MachineIRBuilder &B) const {
- // Is non-HSA path or trap-handler disabled? Then, report a warning
- // accordingly
+ // Is this an unsupported ABI, or is the trap handler disabled? Then report
+ // a warning accordingly.
+ const GCNSubtarget::TrapHandlerAbi Abi = ST.getTrapHandlerAbi();
if (!ST.hasTrapHandler() ||
- ST.getTrapHandlerAbi() != GCNSubtarget::TrapHandlerAbi::AMDHSA) {
+ (Abi != GCNSubtarget::TrapHandlerAbi::AMDHSA &&
+ Abi != GCNSubtarget::TrapHandlerAbi::AMDPAL)) {
Function &Fn = B.getMF().getFunction();
Fn.getContext().diagnose(DiagnosticInfoUnsupported(
Fn, "debugtrap handler not supported", MI.getDebugLoc(), DS_Warning));
@@ -8185,6 +8217,44 @@ bool AMDGPULegalizerInfo::legalizeIntrinsic(LegalizerHelper &Helper,
// Replace the use G_BRCOND with the exec manipulate and branch pseudos.
auto IntrID = cast<GIntrinsic>(MI).getIntrinsicID();
switch (IntrID) {
+ case Intrinsic::is_debugging_enabled: {
+ auto Match = matchCFIntrinsicBranchUse(MI, MRI);
+ bool CannotFuse =
+ !Match || any_of(make_range(std::next(MI.getIterator()),
+ Match->CondBr->getIterator()),
+ [](const MachineInstr &Between) {
+ return !Between.isMetaInstruction() &&
+ (Between.mayLoadOrStore() ||
+ Between.hasUnmodeledSideEffects());
+ });
+ if (CannotFuse) {
+ auto Bits =
+ B.buildIntrinsic(Intrinsic::amdgcn_s_getreg, {LLT::scalar(32)})
+ .addImm(AMDGPU::Hwreg::getDebuggingEnabledHwregImm(ST));
+ Bits->setFlag(MachineInstr::NoMerge);
+ B.buildICmp(CmpInst::ICMP_NE, MI.getOperand(0).getReg(), Bits.getReg(0),
+ B.buildConstant(LLT::scalar(32), 0));
+ MI.eraseFromParent();
+ return true;
+ }
+
+ B.setInsertPt(*Match->CondBr->getParent(), Match->CondBr->getIterator());
+ B.setDebugLoc(Match->CondBr->getDebugLoc());
+ MachineInstrBuilder CDBGBranch =
+ B.buildInstr(AMDGPU::S_CBRANCH_CDBGSYS_OR_USER)
+ .addMBB(Match->ConditionTrueTarget);
+ CDBGBranch->setFlag(MachineInstr::NoMerge);
+
+ if (Match->isNegated())
+ Match->redirectFallthroughEdge(B, *Match->ConditionFalseTarget);
+
+ Register Cond = MI.getOperand(0).getReg();
+ MRI.markUsesInDebugValueAsUndef(Cond);
+ Match->eraseDeadNegation(MRI);
+ MI.eraseFromParent();
+ Match->CondBr->eraseFromParent();
+ return true;
+ }
case Intrinsic::amdgcn_icmp: {
// amdgcn.icmp(i1 src0, i1 0, NE) -> ballot(src0)
// This is the only valid form of amdgcn.icmp with i1 inputs.
@@ -8230,80 +8300,58 @@ bool AMDGPULegalizerInfo::legalizeIntrinsic(LegalizerHelper &Helper,
return true;
case Intrinsic::amdgcn_if:
case Intrinsic::amdgcn_else: {
- MachineInstr *Br = nullptr;
- MachineBasicBlock *UncondBrTarget = nullptr;
- bool Negated = false;
- if (MachineInstr *BrCond =
- verifyCFIntrinsic(MI, MRI, Br, UncondBrTarget, Negated)) {
- const SIRegisterInfo *TRI
- = static_cast<const SIRegisterInfo *>(MRI.getTargetRegisterInfo());
+ if (auto Match = matchCFIntrinsicBranchUse(MI, MRI)) {
+ const SIRegisterInfo *TRI =
+ static_cast<const SIRegisterInfo *>(MRI.getTargetRegisterInfo());
Register Def = MI.getOperand(1).getReg();
Register Use = MI.getOperand(3).getReg();
- MachineBasicBlock *CondBrTarget = BrCond->getOperand(1).getMBB();
-
- if (Negated)
- std::swap(CondBrTarget, UncondBrTarget);
-
- B.setInsertPt(B.getMBB(), BrCond->getIterator());
+ B.setInsertPt(B.getMBB(), Match->CondBr->getIterator());
if (IntrID == Intrinsic::amdgcn_if) {
B.buildInstr(AMDGPU::SI_IF)
- .addDef(Def)
- .addUse(Use)
- .addMBB(UncondBrTarget);
+ ...
[truncated]
``````````
</details>
https://github.com/llvm/llvm-project/pull/219179
More information about the llvm-commits
mailing list