[llvm] [AMDGPU][GFX12/GFX13] Add cdbg branch support and lower llvm.is.debugging.enabled (PR #219178)
via llvm-commits
llvm-commits at lists.llvm.org
Thu Aug 27 04:45:14 PDT 2026
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>,
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>,
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>,
Alexander =?utf-8?q?Hück?= <alexander.huck at amd.com>
Message-ID:
In-Reply-To: <llvm.org/llvm/llvm-project/pull/219178 at github.com>
llvmorg-github-actions[bot] wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-backend-nvptx
Author: Alexander Hück (ahueck)
<details>
<summary>Changes</summary>
This is PR 3 of 4 in a series adding `llvm.is.debugging.enabled` intrinsic to LLVM IR and lowering this intrinsic for AMDGPU targets, , see [[[RFC] Introduce `llvm.is.debugging.enabled` intrinsic](https://discourse.llvm.org/t/rfc-introduce-llvm-is-debugging-enabled-intrinsic/91676)
## Changes of this PR
- Opt AMDGPU into lowering `llvm.is.debugging.enabled`. Unsupported subtargets replace the intrinsic with false.
- Add cdbg conditional branch instruction support for GFX12 and GFX13.
- On supported subtargets, SelectionDAG and GlobalISel lowers the intrinsic by either reading the generation-specific debugging-state bits with `s_getreg` or fuse safe single-branch uses into `s_cbranch_cdbgsys_or_user`. This includes negated branch conditions.
- Mark materialized and fused observations NoMerge to keep calls distinct.
PR 4 uses this support to implement AMDPAL lowering of `llvm.debugtrap`.
## PR stack
1. [[AMDGPU][GlobalISel][SelectionDAG] Refactor control-flow intrinsic branch matching](https://github.com/llvm/llvm-project/pull/219173)
2. [[IR][CodeGen] Add llvm.is.debugging.enabled intrinsic](https://github.com/llvm/llvm-project/pull/219175)
3. [[AMDGPU][GFX12/GFX13] Add cdbg branch support and lower llvm.is.debugging.enabled](https://github.com/llvm/llvm-project/pull/219178)
4. [[AMDGPU] Add AMDPAL support for llvm.debugtrap](https://github.com/llvm/llvm-project/pull/219179)
*Disclaimer*: Agentic AI assisted development.
---
Patch is 87.61 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/219178.diff
32 Files Affected:
- (modified) llvm/docs/AMDGPUUsage.rst (+46)
- (modified) llvm/docs/LangRef.md (+30)
- (modified) llvm/include/llvm/IR/Intrinsics.td (+6)
- (modified) llvm/include/llvm/Target/TargetMachine.h (+7)
- (modified) llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp (+8)
- (modified) llvm/lib/Target/AMDGPU/AMDGPU.td (+11-1)
- (modified) llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp (+109-63)
- (modified) llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp (+16)
- (modified) llvm/lib/Target/AMDGPU/AMDGPUTargetMachine.h (+2)
- (modified) llvm/lib/Target/AMDGPU/SIISelLowering.cpp (+131-42)
- (modified) llvm/lib/Target/AMDGPU/SOPInstructions.td (+12-8)
- (modified) llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.cpp (+10)
- (modified) llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.h (+4)
- (modified) llvm/test/Analysis/UniformityAnalysis/AMDGPU/always_uniform.ll (+28)
- (added) llvm/test/Assembler/is-debugging-enabled.ll (+11)
- (modified) llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-amdgcn.if-invalid.mir (+26-1)
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled-divergent-exec-guard.ll (+236)
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled-shapes.ll (+449)
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled-unsupported.ll (+32)
- (added) llvm/test/CodeGen/AMDGPU/is-debugging-enabled.ll (+257)
- (added) llvm/test/CodeGen/NVPTX/is-debugging-enabled.ll (+11)
- (added) llvm/test/CodeGen/X86/is-debugging-enabled.ll (+12)
- (modified) llvm/test/MC/AMDGPU/gfx12_asm_sopp.s (+24)
- (modified) llvm/test/MC/AMDGPU/gfx12_unsupported.s (-12)
- (modified) llvm/test/MC/AMDGPU/gfx13_asm_sopp.s (+24)
- (added) llvm/test/Transforms/DCE/is-debugging-enabled.ll (+12)
- (added) llvm/test/Transforms/EarlyCSE/is-debugging-enabled.ll (+17)
- (added) llvm/test/Transforms/LICM/is-debugging-enabled.ll (+25)
- (added) llvm/test/Transforms/LoopUnroll/is-debugging-enabled.ll (+30)
- (added) llvm/test/Transforms/PreISelIntrinsicLowering/is-debugging-enabled.ll (+17)
- (added) llvm/test/Transforms/SimplifyCFG/is-debugging-enabled.ll (+58)
- (added) llvm/test/Verifier/is-debugging-enabled.ll (+15)
``````````diff
diff --git a/llvm/docs/AMDGPUUsage.rst b/llvm/docs/AMDGPUUsage.rst
index 40a2820b9e04f..e82af4d16c9df 100644
--- a/llvm/docs/AMDGPUUsage.rst
+++ b/llvm/docs/AMDGPUUsage.rst
@@ -1709,6 +1709,52 @@ The AMDGPU backend implements the following LLVM IR intrinsics.
The format is a 64-bit concatenation of the MODE and TRAPSTS registers.
:ref:`llvm.set.fpenv<int_set_fpenv>` Sets the floating point environment to the specified state.
+
+ :ref:`llvm.is.debugging.enabled <llvm.is.debugging.enabled>`
+ Supported on GFX11.5, GFX12, and GFX13 targets. Other
+ subtargets lower the result to ``false``.
+
+ The target-defined execution context is the current wave.
+ The result is uniform across the active lanes of that
+ wave, including when the intrinsic is executed in
+ divergent control flow. Each call remains a distinct
+ observation of the wave's debugging-enabled state.
+
+ A wave executes in debugging mode when either the
+ ``COND_DBG_SYS`` or ``COND_DBG_USER`` bit is set.
+ ``COND_DBG_SYS`` reflects a system-wide debugger
+ attach, while ``COND_DBG_USER`` reflects per-dispatch
+ launch control, configured through
+ :ref:`CDBG_USER <amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx13-table>`
+ in the kernel descriptor. The driver sets these
+ bits, and provides an interface allowing the user
+ mode runtime or an external debugger to control
+ the setting.
+
+ The intrinsic lowers in one of two forms.
+ When the query feeds a single conditional branch in the
+ same basic block, with no intervening observable
+ operations, it may fuse into a single
+ ``s_cbranch_cdbgsys_or_user``. Branch-hint and
+ negated-condition forms are supported; other uses of
+ the query value prevent fusion.
+
+ Any other use materializes an ``i1`` by reading the
+ adjacent ``COND_DBG_USER`` and ``COND_DBG_SYS`` bits
+ and testing them against zero. The register holding
+ them differs by generation:
+
+ .. code-block:: none
+
+ ; GFX11.5
+ s_getreg_b32 s0, hwreg(HW_REG_STATUS, 20, 2)
+ ; GFX12
+ s_getreg_b32 s0, hwreg(HW_REG_WAVE_STATE_PRIV, 16, 2)
+ ; GFX13
+ s_getreg_b32 s0, hwreg(HW_REG_WAVE_STATUS, 20, 2)
+ ; all generations
+ s_cmp_lg_u32 s0, 0
+
llvm.amdgcn.readfirstlane Provides direct access to v_readfirstlane_b32. Returns the value in
the lowest active lane of the input operand. Currently implemented
for i16, i32, float, half, bfloat, <2 x i16>, <2 x half>, <2 x bfloat>,
diff --git a/llvm/docs/LangRef.md b/llvm/docs/LangRef.md
index ee1600b96f1dd..9f2ab993f3a80 100644
--- a/llvm/docs/LangRef.md
+++ b/llvm/docs/LangRef.md
@@ -26263,6 +26263,36 @@ This intrinsic is lowered to code which is intended to cause an
execution trap with the intention of requesting the attention of a
debugger.
+(llvm.is.debugging.enabled)=
+
+#### '`llvm.is.debugging.enabled`' Intrinsic
+
+##### Syntax:
+
+```llvm
+declare noundef i1 @llvm.is.debugging.enabled() nomerge memory(inaccessiblemem: readwrite)
+```
+
+##### Overview:
+
+The '`llvm.is.debugging.enabled`' intrinsic returns whether debugging is enabled
+for the target-defined execution context of the current invocation.
+
+##### Arguments:
+
+None.
+
+##### Semantics:
+
+Each call observes the current debugging-enabled state. Calls are distinct
+observations: LLVM must not assume separate calls return the same value, and a
+call may not be removed when its result is unused, commoned with another call,
+or reordered with respect to one.
+
+On targets that do not support querying debugging-enabled state, the intrinsic
+returns `false` and performs no observation.
+
+
(llvm.ubsantrap)=
#### '`llvm.ubsantrap`' Intrinsic
diff --git a/llvm/include/llvm/IR/Intrinsics.td b/llvm/include/llvm/IR/Intrinsics.td
index c20ea64c7eef4..b14375e277e7a 100644
--- a/llvm/include/llvm/IR/Intrinsics.td
+++ b/llvm/include/llvm/IR/Intrinsics.td
@@ -2059,6 +2059,12 @@ def int_trap : Intrinsic<[], [],
ClangBuiltin<"__builtin_trap">;
def int_debugtrap : Intrinsic<[]>,
ClangBuiltin<"__builtin_debugtrap">;
+// Query whether debugging is enabled for the current execution context.
+def int_is_debugging_enabled :
+ DefaultAttrsIntrinsic<[llvm_i1_ty], [],
+ [IntrInaccessibleMemOnly,
+ IntrNoMerge,
+ NoUndef<RetIndex>]>;
def int_ubsantrap : Intrinsic<[], [llvm_i8_ty],
[IntrNoReturn, IntrCold, ImmArg<ArgIndex<0>>,
IntrInaccessibleMemOnly, IntrWriteMem]>;
diff --git a/llvm/include/llvm/Target/TargetMachine.h b/llvm/include/llvm/Target/TargetMachine.h
index a6b73d636dc27..ea26be5e48bfb 100644
--- a/llvm/include/llvm/Target/TargetMachine.h
+++ b/llvm/include/llvm/Target/TargetMachine.h
@@ -567,6 +567,13 @@ class LLVM_ABI TargetMachine {
/// this function returns false, the intrinsic will be supported generically
/// but without loop detection support.
virtual bool canLowerCondLoop() const { return false; }
+
+ /// Returns whether this target takes responsibility for lowering
+ /// llvm.is.debugging.enabled. If false, generic lowering replaces the intrinsic
+ /// with false. If true, the target must handle every supported subtarget,
+ /// including replacing the intrinsic with false on subtargets without native
+ /// lowering.
+ virtual bool canLowerIsDebuggingEnabled() const { return false; }
};
} // end namespace llvm
diff --git a/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp b/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp
index 24788148b9f15..4ec866d7b973f 100644
--- a/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp
+++ b/llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp
@@ -815,6 +815,14 @@ bool PreISelIntrinsicLowering::lowerIntrinsics(Module &M) const {
case Intrinsic::protected_field_ptr:
Changed |= expandProtectedFieldPtr(F);
break;
+ case Intrinsic::is_debugging_enabled:
+ if (!TM || !TM->canLowerIsDebuggingEnabled())
+ Changed |= forEachCall(F, [](CallInst *CI) {
+ CI->replaceAllUsesWith(ConstantInt::getFalse(CI->getContext()));
+ CI->eraseFromParent();
+ return true;
+ });
+ break;
case Intrinsic::cond_loop:
if (!TM->canLowerCondLoop())
Changed |= expandCondLoop(F);
diff --git a/llvm/lib/Target/AMDGPU/AMDGPU.td b/llvm/lib/Target/AMDGPU/AMDGPU.td
index ed2ee1ff70f4f..17b22beebcbb9 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPU.td
+++ b/llvm/lib/Target/AMDGPU/AMDGPU.td
@@ -149,6 +149,12 @@ defm TrapHandler: AMDGPUSubtargetFeature<"trap-handler",
/*GenPredicate=*/0, /*GenAssemblerPredicate=*/1, [], InlineIgnore
>;
+defm CDBGSysOrUserBranch : AMDGPUSubtargetFeature<
+ "cdbg-sys-or-user-branch",
+ "Support querying debugging state with s_cbranch_cdbgsys_or_user",
+ /*GenPredicate=*/1, /*GenAssemblerPredicate=*/0
+>;
+
defm UnalignedScratchAccess : AMDGPUSubtargetFeature<"unaligned-scratch-access",
"Support unaligned scratch loads and stores",
/*GenPredicate=*/1, /*GenAssemblerPredicate=*/1, [], InlineIgnore
@@ -2232,7 +2238,8 @@ def FeatureISAVersion11_0_3 : FeatureSet<
def FeatureISAVersion11_5_Common : FeatureSet<
!listconcat(FeatureISAVersion11_Common.Features,
- [FeatureSALUFloatInsts,
+ [FeatureCDBGSysOrUserBranch,
+ FeatureSALUFloatInsts,
FeatureDPPSrc1SGPR,
FeatureRequiredExportPriority,
FeatureDot5Insts,
@@ -2275,6 +2282,7 @@ def FeatureISAVersion11_7_Generic: FeatureSet<
def FeatureISAVersion12 : FeatureSet<
[FeatureGFX12,
FeatureSupportsWave64, FeatureSupportsWGP,
+ FeatureCDBGSysOrUserBranch,
FeatureBackOffBarrier,
FeatureAddressableLocalMemorySize65536,
FeatureHalfAddressablePhysicalLocalMemory,
@@ -2343,6 +2351,7 @@ def FeatureISAVersion12 : FeatureSet<
def FeatureISAVersion12_50_Common : FeatureSet<
[FeatureGFX12,
FeatureGFX1250Insts,
+ FeatureCDBGSysOrUserBranch,
FeatureBackOffBarrier,
FeatureRequiresAlignedVGPRs,
FeatureCuMode,
@@ -2536,6 +2545,7 @@ def FeatureISAVersion12_5_Generic: FeatureSet<
def FeatureISAVersion13 : FeatureSet<
[FeatureGFX13,
FeatureGFX1250Insts,
+ FeatureCDBGSysOrUserBranch,
FeatureAddressableLocalMemorySize196608,
Feature64BitLiterals,
FeatureLDSBankCount32,
diff --git a/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp b/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
index 7c1a26f761c96..c42060e69a274 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
@@ -4804,47 +4804,77 @@ static bool isNot(const MachineRegisterInfo &MRI, const MachineInstr &MI) {
return ConstVal == -1;
}
-// Return the use branch instruction, otherwise null if the usage is invalid.
-static MachineInstr *
-verifyCFIntrinsic(MachineInstr &MI, MachineRegisterInfo &MRI, MachineInstr *&Br,
- MachineBasicBlock *&UncondBrTarget, bool &Negated) {
+namespace {
+struct CFIntrinsicBranchMatch {
+ MachineInstr *Negation = nullptr;
+ MachineInstr *CondBr = nullptr;
+ MachineInstr *UncondBr = nullptr;
+
+ MachineBasicBlock *ConditionTrueTarget = nullptr;
+ MachineBasicBlock *ConditionFalseTarget = nullptr;
+
+ bool isNegated() const { return Negation != nullptr; }
+
+ // Retarget the explicit branch for the fallthrough edge, or materialize
+ // the branch when that edge was represented by layout fallthrough. The
+ // builder must be positioned at the replacement branch.
+ void redirectFallthroughEdge(MachineIRBuilder &B,
+ MachineBasicBlock &Target) const {
+ if (UncondBr)
+ UncondBr->getOperand(0).setMBB(&Target);
+ else
+ B.buildBr(Target);
+ }
+
+ void eraseDeadNegation(MachineRegisterInfo &MRI) {
+ if (Negation)
+ eraseInstr(*Negation, MRI);
+ }
+};
+} // namespace
+
+static std::optional<CFIntrinsicBranchMatch>
+matchCFIntrinsicBranchUse(MachineInstr &MI, MachineRegisterInfo &MRI) {
Register CondDef = MI.getOperand(0).getReg();
if (!MRI.hasOneNonDBGUse(CondDef))
- return nullptr;
+ return std::nullopt;
MachineBasicBlock *Parent = MI.getParent();
MachineInstr *UseMI = &*MRI.use_instr_nodbg_begin(CondDef);
+ MachineInstr *Negation = nullptr;
if (isNot(MRI, *UseMI)) {
+ Negation = UseMI;
Register NegatedCond = UseMI->getOperand(0).getReg();
if (!MRI.hasOneNonDBGUse(NegatedCond))
- return nullptr;
-
- // We're deleting the def of this value, so we need to remove it.
- eraseInstr(*UseMI, MRI);
+ return std::nullopt;
UseMI = &*MRI.use_instr_nodbg_begin(NegatedCond);
- Negated = true;
}
if (UseMI->getParent() != Parent || UseMI->getOpcode() != AMDGPU::G_BRCOND)
- return nullptr;
+ return std::nullopt;
// Make sure the cond br is followed by a G_BR, or is the last instruction.
+ MachineInstr *UncondBr = nullptr;
+ MachineBasicBlock *OtherTarget = nullptr;
MachineBasicBlock::iterator Next = std::next(UseMI->getIterator());
if (Next == Parent->end()) {
MachineFunction::iterator NextMBB = std::next(Parent->getIterator());
if (NextMBB == Parent->getParent()->end()) // Illegal intrinsic use.
- return nullptr;
- UncondBrTarget = &*NextMBB;
+ return std::nullopt;
+ OtherTarget = &*NextMBB;
} else {
if (Next->getOpcode() != AMDGPU::G_BR)
- return nullptr;
- Br = &*Next;
- UncondBrTarget = Br->getOperand(0).getMBB();
+ return std::nullopt;
+ UncondBr = &*Next;
+ OtherTarget = UncondBr->getOperand(0).getMBB();
}
- return UseMI;
+ MachineBasicBlock *TakenTarget = UseMI->getOperand(1).getMBB();
+ return CFIntrinsicBranchMatch{Negation, UseMI, UncondBr,
+ Negation ? OtherTarget : TakenTarget,
+ Negation ? TakenTarget : OtherTarget};
}
void AMDGPULegalizerInfo::buildLoadInputValue(Register DstReg,
@@ -8185,6 +8215,44 @@ bool AMDGPULegalizerInfo::legalizeIntrinsic(LegalizerHelper &Helper,
// Replace the use G_BRCOND with the exec manipulate and branch pseudos.
auto IntrID = cast<GIntrinsic>(MI).getIntrinsicID();
switch (IntrID) {
+ case Intrinsic::is_debugging_enabled: {
+ auto Match = matchCFIntrinsicBranchUse(MI, MRI);
+ bool CannotFuse =
+ !Match || any_of(make_range(std::next(MI.getIterator()),
+ Match->CondBr->getIterator()),
+ [](const MachineInstr &Between) {
+ return !Between.isMetaInstruction() &&
+ (Between.mayLoadOrStore() ||
+ Between.hasUnmodeledSideEffects());
+ });
+ if (CannotFuse) {
+ auto Bits =
+ B.buildIntrinsic(Intrinsic::amdgcn_s_getreg, {LLT::scalar(32)})
+ .addImm(AMDGPU::Hwreg::getDebuggingEnabledHwregImm(ST));
+ Bits->setFlag(MachineInstr::NoMerge);
+ B.buildICmp(CmpInst::ICMP_NE, MI.getOperand(0).getReg(), Bits.getReg(0),
+ B.buildConstant(LLT::scalar(32), 0));
+ MI.eraseFromParent();
+ return true;
+ }
+
+ B.setInsertPt(*Match->CondBr->getParent(), Match->CondBr->getIterator());
+ B.setDebugLoc(Match->CondBr->getDebugLoc());
+ MachineInstrBuilder CDBGBranch =
+ B.buildInstr(AMDGPU::S_CBRANCH_CDBGSYS_OR_USER)
+ .addMBB(Match->ConditionTrueTarget);
+ CDBGBranch->setFlag(MachineInstr::NoMerge);
+
+ if (Match->isNegated())
+ Match->redirectFallthroughEdge(B, *Match->ConditionFalseTarget);
+
+ Register Cond = MI.getOperand(0).getReg();
+ MRI.markUsesInDebugValueAsUndef(Cond);
+ Match->eraseDeadNegation(MRI);
+ MI.eraseFromParent();
+ Match->CondBr->eraseFromParent();
+ return true;
+ }
case Intrinsic::amdgcn_icmp: {
// amdgcn.icmp(i1 src0, i1 0, NE) -> ballot(src0)
// This is the only valid form of amdgcn.icmp with i1 inputs.
@@ -8230,80 +8298,58 @@ bool AMDGPULegalizerInfo::legalizeIntrinsic(LegalizerHelper &Helper,
return true;
case Intrinsic::amdgcn_if:
case Intrinsic::amdgcn_else: {
- MachineInstr *Br = nullptr;
- MachineBasicBlock *UncondBrTarget = nullptr;
- bool Negated = false;
- if (MachineInstr *BrCond =
- verifyCFIntrinsic(MI, MRI, Br, UncondBrTarget, Negated)) {
- const SIRegisterInfo *TRI
- = static_cast<const SIRegisterInfo *>(MRI.getTargetRegisterInfo());
+ if (auto Match = matchCFIntrinsicBranchUse(MI, MRI)) {
+ const SIRegisterInfo *TRI =
+ static_cast<const SIRegisterInfo *>(MRI.getTargetRegisterInfo());
Register Def = MI.getOperand(1).getReg();
Register Use = MI.getOperand(3).getReg();
- MachineBasicBlock *CondBrTarget = BrCond->getOperand(1).getMBB();
-
- if (Negated)
- std::swap(CondBrTarget, UncondBrTarget);
-
- B.setInsertPt(B.getMBB(), BrCond->getIterator());
+ B.setInsertPt(B.getMBB(), Match->CondBr->getIterator());
if (IntrID == Intrinsic::amdgcn_if) {
B.buildInstr(AMDGPU::SI_IF)
- .addDef(Def)
- .addUse(Use)
- .addMBB(UncondBrTarget);
+ .addDef(Def)
+ .addUse(Use)
+ .addMBB(Match->ConditionFalseTarget);
} else {
B.buildInstr(AMDGPU::SI_ELSE)
.addDef(Def)
.addUse(Use)
- .addMBB(UncondBrTarget);
+ .addMBB(Match->ConditionFalseTarget);
}
- if (Br) {
- Br->getOperand(0).setMBB(CondBrTarget);
- } else {
- // The IRTranslator skips inserting the G_BR for fallthrough cases, but
- // since we're swapping branch targets it needs to be reinserted.
- // FIXME: IRTranslator should probably not do this
- B.buildBr(*CondBrTarget);
- }
+ // The IRTranslator skips inserting the G_BR for fallthrough cases, but
+ // since we're swapping branch targets it needs to be reinserted.
+ // FIXME: IRTranslator should probably not do this
+ Match->redirectFallthroughEdge(B, *Match->ConditionTrueTarget);
MRI.setRegClass(Def, TRI->getWaveMaskRegClass());
MRI.setRegClass(Use, TRI->getWaveMaskRegClass());
+ Match->eraseDeadNegation(MRI);
MI.eraseFromParent();
- BrCond->eraseFromParent();
+ Match->CondBr->eraseFromParent();
return true;
}
return false;
}
case Intrinsic::amdgcn_loop: {
- MachineInstr *Br = nullptr;
- MachineBasicBlock *UncondBrTarget = nullptr;
- bool Negated = false;
- if (MachineInstr *BrCond =
- verifyCFIntrinsic(MI, MRI, Br, UncondBrTarget, Negated)) {
- const SIRegisterInfo *TRI
- = static_cast<const SIRegisterInfo *>(MRI.getTargetRegisterInfo());
-
- MachineBasicBlock *CondBrTarget = BrCond->getOperand(1).getMBB();
- Register Reg = MI.getOperand(2).getReg();
+ if (auto Match = matchCFIntrinsicBranchUse(MI, MRI)) {
+ const SIRegisterInfo *TRI =
+ static_cast<const SIRegisterInfo *>(MRI.getTargetRegisterInfo());
- if (Negated)
- std::swap(CondBrTarget, UncondBrTarget);
+ Register Reg = MI.getOperand(2).getReg();
- B.setInsertPt(B.getMBB(), BrCond->getIterator());
+ B.setInsertPt(B.getMBB(), Match->CondBr->getIterator());
B.buildInstr(AMDGPU::SI_LOOP)
- .addUse(Reg)
- .addMBB(UncondBrTarget);
+ .addUse(Reg)
+ .addMBB(Match->ConditionFalseTarget);
- if (Br)
- Br->getOperand(0).setMBB(CondBrTarget);
- else
- B.buildBr(*CondBrTarget);
+ Match->redirectFallthroughEdge(B, *Match->ConditionTrueTarget);
+ Match->eraseDeadNegation(MRI);
MI.eraseFromParent();
- BrCond->eraseFromParent();
+ Match->CondBr->eraseFromParent();
MRI.setRegClass(Reg, TRI->getWaveMaskRegClass());
return true;
}
diff --git a/llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp b/llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp
index 123d9930fc49e..274cba75e5254 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp
@@ -37,6 +37,7 @@ class AMDGPULowerIntrinsicsImpl {
bool run();
private:
+ bool visitIsDebuggingEnabled(IntrinsicInst &I);
bool visitBarrier(IntrinsicInst &I);
bool visitPtrSBufferLoad(IntrinsicInst &I);
};
@@ -63,6 +64,16 @@ template <class T> static void forEachCall(Function &Intrin, T Callback) {
} // anonymous namespace
+bool AMDGPULowerIntrinsicsIm...
[truncated]
``````````
</details>
https://github.com/llvm/llvm-project/pull/219178
More information about the llvm-commits
mailing list