[llvm-branch-commits] [compiler-rt] [llvm] [PGO] Add GPU wave counters (PR #225589)

Joseph Huber via llvm-branch-commits llvm-branch-commits at lists.llvm.org
Thu Sep 24 12:49:58 PDT 2026


================
@@ -29,17 +29,19 @@ static int is_uniform(uint64_t mask) {
 
 // Wave-cooperative counter increment. The instrumentation pass emits calls to
 // this in place of the default non-atomic load/add/store or atomicrmw sequence.
-// The optional uniform counter allows calculating wave uniformity if present.
+// The uniform counter is optional; the wave counter records every wave visit.
 COMPILER_RT_VISIBILITY void INSTR_PROF_INSTRUMENT_GPU_FUNC(uint64_t *counter,
                                                            uint64_t *uniform,
-                                                           uint64_t step) {
+                                                           uint64_t step,
+                                                           uint64_t *wave) {
   uint64_t mask = __gpu_lane_mask();
   if (__gpu_is_first_in_lane(mask)) {
     __scoped_atomic_fetch_add(counter, step * __builtin_popcountg(mask),
                               __ATOMIC_RELAXED, __MEMORY_SCOPE_DEVICE);
     if (uniform && is_uniform(mask))
       __scoped_atomic_fetch_add(uniform, step * __builtin_popcountg(mask),
                                 __ATOMIC_RELAXED, __MEMORY_SCOPE_DEVICE);
+    __scoped_atomic_fetch_add(wave, 1, __ATOMIC_RELAXED, __MEMORY_SCOPE_DEVICE);
----------------
jhuber6 wrote:

Does this not need to use `step` at all? I'm a little hesitant to change the format again for something that seems pretty minor. I've always wondered if we could just change the uniformity to a count against full occupancy. I.e. we know if we executed w/ full occupancy we'd have N, but we observed M. But I don't r know what kind of optimizations PGO on HIP is enabling. What does this extra information let us do?

https://github.com/llvm/llvm-project/pull/225589


More information about the llvm-branch-commits mailing list