[llvm] [AMDGPU] Combine redundant ballot intrinsic calls (PR #218357)
Sameer Sahasrabuddhe via llvm-commits
llvm-commits at lists.llvm.org
Wed Aug 26 08:09:05 PDT 2026
================
@@ -135,8 +145,154 @@ static bool optimizeUniformIntrinsic(IntrinsicInst &II,
return false;
}
+/// Maximum number of basic blocks inspected while proving that exec is
+/// invariant between two ballots. Keeps the walk below linear-per-pair in
+/// pathological CFGs.
+static constexpr unsigned MaxExecInvarianceBlocks = 100;
+
+/// Returns true if \p I may change exec, i.e. the set of lanes that are active
+/// when the following instructions execute.
+static bool isExecModifyingInst(const Instruction &I) {
----------------
ssahasra wrote:
@arsenm on a topic this subtle, fraught with historic footguns, a more detailed explanation would be highly appreciated.
https://github.com/llvm/llvm-project/pull/218357
More information about the llvm-commits
mailing list