[llvm] [AMDGPU] Fix incorrect grid_dims constant folding for reqd_work_group_size (PR #211285)
Wooseok Lee via llvm-commits
llvm-commits at lists.llvm.org
Fri Jul 24 07:32:53 PDT 2026
wooseoklee wrote:
I can see that this approach improves the conservative one from https://github.com/llvm/llvm-project/pull/205866 as follows.
Z != 1 → fold to constant 3 (still a constant fold)
Y != 1, Z == 1 → range [2, 4) metadata only
Y == 1, Z == 1 → range [1, 4) metadata only
https://registry.khronos.org/OpenCL/specs/unified/refpages/man/html/optionalAttributeQualifiers.html
If Z is one, the work_dim argument to clEnqueueNDRangeKernel can be 2 or 3. If Y and Z are one, the work_dim argument to clEnqueueNDRangeKernel can be 1, 2 or 3.
https://github.com/llvm/llvm-project/pull/211285
More information about the llvm-commits
mailing list