[all-commits] [llvm/llvm-project] 2cab66: [clang][OpenMP] Improve loop structure for distrib...

Robert Imschweiler via All-commits all-commits at lists.llvm.org
Thu Jun 4 09:46:42 PDT 2026


  Branch: refs/heads/users/ro-i/xteam-red-codegen
  Home:   https://github.com/llvm/llvm-project
  Commit: 2cab66c114b2c50e49562356e10563f0622438dc
      https://github.com/llvm/llvm-project/commit/2cab66c114b2c50e49562356e10563f0622438dc
  Author: Robert Imschweiler <robert.imschweiler at amd.com>
  Date:   2026-06-04 (Thu, 04 Jun 2026)

  Changed paths:
    M clang/include/clang/Basic/OpenMPKinds.h
    M clang/lib/CodeGen/CGOpenMPRuntime.cpp
    M clang/lib/CodeGen/CGStmtOpenMP.cpp
    M clang/test/OpenMP/amdgcn_target_device_vla.cpp
    M clang/test/OpenMP/amdgpu_target_with_aligned_attribute.c
    M clang/test/OpenMP/metadirective_device_arch_codegen.cpp
    M clang/test/OpenMP/nvptx_SPMD_codegen.cpp
    M clang/test/OpenMP/nvptx_distribute_parallel_generic_mode_codegen.cpp
    M clang/test/OpenMP/nvptx_target_teams_distribute_parallel_for_codegen.cpp
    M clang/test/OpenMP/nvptx_target_teams_distribute_parallel_for_simd_codegen.cpp
    M clang/test/OpenMP/nvptx_target_teams_generic_loop_codegen.cpp
    M clang/test/OpenMP/nvptx_target_teams_generic_loop_generic_mode_codegen.cpp
    M clang/test/OpenMP/target_teams_generic_loop_codegen.cpp
    M clang/test/OpenMP/target_teams_generic_loop_codegen_as_distribute.cpp
    M clang/test/OpenMP/target_teams_generic_loop_codegen_as_parallel_for.cpp
    M offload/test/offloading/gpupgo/pgo_atomic_teams.c

  Log Message:
  -----------
  [clang][OpenMP] Improve loop structure for distributed loops

This is a part of a series of patches that rework OpenMP cross-team
reductions.

This patches wires the existing
`kmp_sched_distr_static_chunk_sched_static_chunkone` to be used by
CodeGen.

Example of the intended change of this patch:
```
target teams distribute parallel for reduction(+:s)
  for (i = 0; i < N; i++) s += a[i];
```

Before:
```
__kmpc_distribute_static_init(91)
for (team_lb = team*nthreads; team_lb < N; team_lb += nteams*nthreads) {
  __kmpc_for_static_init(33)
  for (iv = team_lb + tid; iv < team_lb + nthreads; iv += nthreads) {
    priv += a[iv];
  }
  __kmpc_nvptx_parallel_reduce_nowait_v2
}
__kmpc_nvptx_teams_reduce_nowait_v2
```

After:
```
__kmpc_for_static_init(93)
for (iv = team*nthreads + tid;
     iv < N;
     iv += nteams*nthreads) {
    priv += a[iv];
}
__kmpc_nvptx_parallel_reduce_nowait_v2
__kmpc_nvptx_teams_reduce_nowait_v2
```

Performance:
All performance tests can be reproduced with
https://github.com/ro-i/xteam-test @ commit
6025e5afc14dd6e65ee2658e5001c16e9b9245ff. To reproduce, simply create a
`local.mk` file in the cloned directory with a suitable `OFFLOAD_ARCH`
for your machine and `CXX_trunk` + `CXX_trunk_cg` set to the paths of
the clang++ binaries for llvm/main and this patch. (llvm/main should
best be at the commit that is currently the base for this PR. At the
moment, this is fe87c971bf07eb38af97ca96bf2810e94e7549dc). Then, run
`make trunk trunk_cg` to build the benchmark binaries for 208 and 10400
teams. Run them with `./run_bench.sh -rq -n10 red_trunk_208
red_trunk_cg_208 red_trunk_10400 red_trunk_cg_10400` to get the avg
performance numbers over 10 rounds. This tests multiple reduction
workloads, including reductions that run in the Generic-SPMD mode, with
208 teams and with 10400 teams, both à 512 threads, and with a reduction
array size of 177,777,777. I tested on a gfx942 and found the following
numbers showing the performance of this patch relative to the baseline:

```
red_comb_sep_arr_32    double   change for 208 teams:    +0.04%   change for 10400 teams:    +5.33%
red_sum_arr_32         double   change for 208 teams:  +569.81%   change for 10400 teams:    -3.11%
red_comb               double   change for 208 teams:  +350.47%   change for 10400 teams:    +2.81%
red_comb_sep           double   change for 208 teams:    +4.83%   change for 10400 teams:    +2.56%
red_dot                double   change for 208 teams:  +202.26%   change for 10400 teams:    +5.48%
red_indirect           double   change for 208 teams:  +239.29%   change for 10400 teams:    +6.39%
red_kernel_part        double   change for 208 teams:    +3.28%   change for 10400 teams:    +1.39%
red_max                double   change for 208 teams:  +273.82%   change for 10400 teams:    +7.00%
red_mult               double   change for 208 teams:  +239.93%   change for 10400 teams:    +6.84%
red_sum                double   change for 208 teams:  +239.90%   change for 10400 teams:    +6.66%
red_pi                 double   change for 208 teams:   +90.06%   change for 10400 teams:   +78.64%
red_comb_sep_arr_32    uint     change for 208 teams:    -0.01%   change for 10400 teams:   +27.70%
red_sum_arr_32         uint     change for 208 teams:  +138.69%   change for 10400 teams:   -15.32%
red_dot                uint     change for 208 teams:  +202.50%   change for 10400 teams:    +6.39%
red_max                uint     change for 208 teams:  +221.10%   change for 10400 teams:    +6.91%
red_sum                uint     change for 208 teams:  +221.17%   change for 10400 teams:    +8.91%
red_comb_sep_arr_32    ulong    change for 208 teams:    -0.00%   change for 10400 teams:    +5.58%
red_sum_arr_32         ulong    change for 208 teams:  +523.94%   change for 10400 teams:    -2.91%
red_dot                ulong    change for 208 teams:  +232.92%   change for 10400 teams:    +4.38%
red_max                ulong    change for 208 teams:  +279.70%   change for 10400 teams:    +6.51%
red_sum                ulong    change for 208 teams:  +261.71%   change for 10400 teams:    +5.77%
red_comb_sep_arr_32    Value    change for 208 teams:    +0.07%   change for 10400 teams:    +0.12%
red_sum_arr_32         Value    change for 208 teams:  +423.06%   change for 10400 teams:    +9.81%
red_dot                Value    change for 208 teams:  +154.16%   change for 10400 teams:    -1.91%
red_max                Value    change for 208 teams: +1105.51%   change for 10400 teams:  +260.17%
red_sum                Value    change for 208 teams:  +360.80%   change for 10400 teams:   +17.99%
```



To unsubscribe from these emails, change your notification settings at https://github.com/llvm/llvm-project/settings/notifications


More information about the All-commits mailing list