[llvm] [AMDGPU] Optimize block count calculations to the new ABI (PR #174112)
Juan Manuel Martinez CaamaƱo via llvm-commits
llvm-commits at lists.llvm.org
Fri Jan 2 00:54:22 PST 2026
================
@@ -323,6 +326,50 @@ static bool processUse(CallInst *CI, bool IsV5OrAbove) {
}
}
+ // Upgrade the old method of calculating the block size using the grid size.
+ // We pattern match any case where the implicit argument group size is the
+ // divisor to a dispatch packet grid size read of the same dimension.
+ if (IsV5OrAbove) {
+ for (int I = 0; I < 3; I++) {
+ Value *GroupSize = GroupSizes[I];
+ if (!GroupSize)
+ continue;
+
+ for (User *U : GroupSize->users()) {
+ Instruction *Inst = cast<Instruction>(U);
+ if (isa<ZExtInst>(Inst) && !Inst->use_empty())
+ Inst = dyn_cast<Instruction>(*Inst->user_begin());
----------------
jmmartinez wrote:
Maybe use `cast<Instruction>(...)` to hit the assertion early if for whatever strange reason `*Inst->user_begin()` is not an instruction.
https://github.com/llvm/llvm-project/pull/174112
More information about the llvm-commits
mailing list