[llvm] [openmp] [OpenMPOpt] Ask the runtime how many of a block's threads can be workers (PR #221450)
Larry Meadows via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 8 11:01:31 PDT 2026
================
@@ -39,10 +39,8 @@ define weak amdgpu_kernel void @visible_body(ptr %dyn, ptr %payload) #0 {
; CHECK-NEXT: [[THREAD_IS_WORKER:%.*]] = icmp ne i32 [[TMP0]], -1
; CHECK-NEXT: br i1 [[THREAD_IS_WORKER]], label %[[IS_WORKER_CHECK:.*]], label %[[THREAD_USER_CODE_CHECK:.*]]
; CHECK: [[IS_WORKER_CHECK]]:
-; CHECK-NEXT: [[BLOCK_HW_SIZE:%.*]] = call i32 @__kmpc_get_hardware_num_threads_in_block()
-; CHECK-NEXT: [[WARP_SIZE:%.*]] = call i32 @__kmpc_get_warp_size()
-; CHECK-NEXT: [[BLOCK_SIZE:%.*]] = sub i32 [[BLOCK_HW_SIZE]], [[WARP_SIZE]]
-; CHECK-NEXT: [[THREAD_IS_MAIN_OR_WORKER:%.*]] = icmp slt i32 [[TMP0]], [[BLOCK_SIZE]]
+; CHECK-NEXT: [[MAX_TEAM_THREADS:%.*]] = call i32 @__kmpc_get_max_team_threads()
----------------
lfmeadow wrote:
You are right, it is not guaranteed. `IsSPMDMode` is a `Local<int>`, so it lives
in shared memory, and `mapping::init()` has only the initial thread write it. In
generic mode `__kmpc_target_init()` deliberately returns to the workers without a
barrier -- "No need to wait since only the main threads will execute user code and
workers will run into a barrier right away" -- and with a custom state machine the
barrier they run into comes after this check rather than before it, so the write
would not yet be theirs to read.
The runtime's own copy of this gate, `shouldEnterStateMachine()`, sidesteps that by
taking `IsSPMD` as a parameter instead of calling `mapping::isSPMDMode()`, so I have
done the same. `__kmpc_get_max_team_threads()` now takes the mode, and the pass
passes a constant false, which is always right because a custom state machine is
only ever built for a generic-mode kernel. Thanks for catching this.
https://github.com/llvm/llvm-project/pull/221450
More information about the llvm-commits
mailing list