[llvm] [NFC][AMDGPU] Inline `getEffectiveWavesPerEU` (PR #221600)
Shilei Tian via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 8 06:20:46 PDT 2026
================
@@ -218,13 +192,29 @@ AMDGPUSubtarget::getWavesPerEU(const Function &F) const {
std::pair<unsigned, unsigned>
AMDGPUSubtarget::getWavesPerEU(std::pair<unsigned, unsigned> FlatWorkGroupSizes,
unsigned LDSBytes, const Function &F) const {
- // Default minimum/maximum number of waves per execution unit.
- std::pair<unsigned, unsigned> Default(1, getMaxWavesPerEU());
+ // Default minimum/maximum number of waves per EU. The range of flat workgroup
+ // sizes limits the achievable maximum, and we aim to support enough waves per
+ // EU so that we can concurrently execute all waves of a single workgroup of
+ // maximum size on a CU.
+ std::pair<unsigned, unsigned> Default = {
+ getWavesPerEUForWorkGroup(FlatWorkGroupSizes.second),
+ getOccupancyWithWorkGroupSizes(LDSBytes, FlatWorkGroupSizes).second};
+ Default.first = std::min(Default.first, Default.second);
// Requested minimum/maximum number of waves per execution unit.
- std::pair<unsigned, unsigned> Requested =
- AMDGPU::getIntegerPairAttribute(F, "amdgpu-waves-per-eu", Default, true);
- return getEffectiveWavesPerEU(Requested, FlatWorkGroupSizes, LDSBytes);
+ std::pair<unsigned, unsigned> Requested = AMDGPU::getIntegerPairAttribute(
+ F, "amdgpu-waves-per-eu", {1, getMaxWavesPerEU()}, true);
+
+ // Make sure requested minimum is within the default range and lower than the
+ // requested maximum. The latter must not violate target specification.
+ if (Requested.first < Default.first || Requested.first > Default.second ||
----------------
shiltian wrote:
Agreed. Since waves-per-eu is not an ABI, to have it clamped also conforms with its semantics.
https://github.com/llvm/llvm-project/pull/221600
More information about the llvm-commits
mailing list