[all-commits] [llvm/llvm-project] 2cab66: [clang][OpenMP] Improve loop structure for distrib...
Robert Imschweiler via All-commits
all-commits at lists.llvm.org
Thu Jun 4 09:46:42 PDT 2026
Branch: refs/heads/users/ro-i/xteam-red-codegen
Home: https://github.com/llvm/llvm-project
Commit: 2cab66c114b2c50e49562356e10563f0622438dc
https://github.com/llvm/llvm-project/commit/2cab66c114b2c50e49562356e10563f0622438dc
Author: Robert Imschweiler <robert.imschweiler at amd.com>
Date: 2026-06-04 (Thu, 04 Jun 2026)
Changed paths:
M clang/include/clang/Basic/OpenMPKinds.h
M clang/lib/CodeGen/CGOpenMPRuntime.cpp
M clang/lib/CodeGen/CGStmtOpenMP.cpp
M clang/test/OpenMP/amdgcn_target_device_vla.cpp
M clang/test/OpenMP/amdgpu_target_with_aligned_attribute.c
M clang/test/OpenMP/metadirective_device_arch_codegen.cpp
M clang/test/OpenMP/nvptx_SPMD_codegen.cpp
M clang/test/OpenMP/nvptx_distribute_parallel_generic_mode_codegen.cpp
M clang/test/OpenMP/nvptx_target_teams_distribute_parallel_for_codegen.cpp
M clang/test/OpenMP/nvptx_target_teams_distribute_parallel_for_simd_codegen.cpp
M clang/test/OpenMP/nvptx_target_teams_generic_loop_codegen.cpp
M clang/test/OpenMP/nvptx_target_teams_generic_loop_generic_mode_codegen.cpp
M clang/test/OpenMP/target_teams_generic_loop_codegen.cpp
M clang/test/OpenMP/target_teams_generic_loop_codegen_as_distribute.cpp
M clang/test/OpenMP/target_teams_generic_loop_codegen_as_parallel_for.cpp
M offload/test/offloading/gpupgo/pgo_atomic_teams.c
Log Message:
-----------
[clang][OpenMP] Improve loop structure for distributed loops
This is a part of a series of patches that rework OpenMP cross-team
reductions.
This patches wires the existing
`kmp_sched_distr_static_chunk_sched_static_chunkone` to be used by
CodeGen.
Example of the intended change of this patch:
```
target teams distribute parallel for reduction(+:s)
for (i = 0; i < N; i++) s += a[i];
```
Before:
```
__kmpc_distribute_static_init(91)
for (team_lb = team*nthreads; team_lb < N; team_lb += nteams*nthreads) {
__kmpc_for_static_init(33)
for (iv = team_lb + tid; iv < team_lb + nthreads; iv += nthreads) {
priv += a[iv];
}
__kmpc_nvptx_parallel_reduce_nowait_v2
}
__kmpc_nvptx_teams_reduce_nowait_v2
```
After:
```
__kmpc_for_static_init(93)
for (iv = team*nthreads + tid;
iv < N;
iv += nteams*nthreads) {
priv += a[iv];
}
__kmpc_nvptx_parallel_reduce_nowait_v2
__kmpc_nvptx_teams_reduce_nowait_v2
```
Performance:
All performance tests can be reproduced with
https://github.com/ro-i/xteam-test @ commit
6025e5afc14dd6e65ee2658e5001c16e9b9245ff. To reproduce, simply create a
`local.mk` file in the cloned directory with a suitable `OFFLOAD_ARCH`
for your machine and `CXX_trunk` + `CXX_trunk_cg` set to the paths of
the clang++ binaries for llvm/main and this patch. (llvm/main should
best be at the commit that is currently the base for this PR. At the
moment, this is fe87c971bf07eb38af97ca96bf2810e94e7549dc). Then, run
`make trunk trunk_cg` to build the benchmark binaries for 208 and 10400
teams. Run them with `./run_bench.sh -rq -n10 red_trunk_208
red_trunk_cg_208 red_trunk_10400 red_trunk_cg_10400` to get the avg
performance numbers over 10 rounds. This tests multiple reduction
workloads, including reductions that run in the Generic-SPMD mode, with
208 teams and with 10400 teams, both à 512 threads, and with a reduction
array size of 177,777,777. I tested on a gfx942 and found the following
numbers showing the performance of this patch relative to the baseline:
```
red_comb_sep_arr_32 double change for 208 teams: +0.04% change for 10400 teams: +5.33%
red_sum_arr_32 double change for 208 teams: +569.81% change for 10400 teams: -3.11%
red_comb double change for 208 teams: +350.47% change for 10400 teams: +2.81%
red_comb_sep double change for 208 teams: +4.83% change for 10400 teams: +2.56%
red_dot double change for 208 teams: +202.26% change for 10400 teams: +5.48%
red_indirect double change for 208 teams: +239.29% change for 10400 teams: +6.39%
red_kernel_part double change for 208 teams: +3.28% change for 10400 teams: +1.39%
red_max double change for 208 teams: +273.82% change for 10400 teams: +7.00%
red_mult double change for 208 teams: +239.93% change for 10400 teams: +6.84%
red_sum double change for 208 teams: +239.90% change for 10400 teams: +6.66%
red_pi double change for 208 teams: +90.06% change for 10400 teams: +78.64%
red_comb_sep_arr_32 uint change for 208 teams: -0.01% change for 10400 teams: +27.70%
red_sum_arr_32 uint change for 208 teams: +138.69% change for 10400 teams: -15.32%
red_dot uint change for 208 teams: +202.50% change for 10400 teams: +6.39%
red_max uint change for 208 teams: +221.10% change for 10400 teams: +6.91%
red_sum uint change for 208 teams: +221.17% change for 10400 teams: +8.91%
red_comb_sep_arr_32 ulong change for 208 teams: -0.00% change for 10400 teams: +5.58%
red_sum_arr_32 ulong change for 208 teams: +523.94% change for 10400 teams: -2.91%
red_dot ulong change for 208 teams: +232.92% change for 10400 teams: +4.38%
red_max ulong change for 208 teams: +279.70% change for 10400 teams: +6.51%
red_sum ulong change for 208 teams: +261.71% change for 10400 teams: +5.77%
red_comb_sep_arr_32 Value change for 208 teams: +0.07% change for 10400 teams: +0.12%
red_sum_arr_32 Value change for 208 teams: +423.06% change for 10400 teams: +9.81%
red_dot Value change for 208 teams: +154.16% change for 10400 teams: -1.91%
red_max Value change for 208 teams: +1105.51% change for 10400 teams: +260.17%
red_sum Value change for 208 teams: +360.80% change for 10400 teams: +17.99%
```
To unsubscribe from these emails, change your notification settings at https://github.com/llvm/llvm-project/settings/notifications
More information about the All-commits
mailing list