[llvm] [NFC][AMDGPU] Inline `getEffectiveWavesPerEU` (PR #221600)

Shilei Tian via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 8 06:20:46 PDT 2026


================
@@ -218,13 +192,29 @@ AMDGPUSubtarget::getWavesPerEU(const Function &F) const {
 std::pair<unsigned, unsigned>
 AMDGPUSubtarget::getWavesPerEU(std::pair<unsigned, unsigned> FlatWorkGroupSizes,
                                unsigned LDSBytes, const Function &F) const {
-  // Default minimum/maximum number of waves per execution unit.
-  std::pair<unsigned, unsigned> Default(1, getMaxWavesPerEU());
+  // Default minimum/maximum number of waves per EU. The range of flat workgroup
+  // sizes limits the achievable maximum, and we aim to support enough waves per
+  // EU so that we can concurrently execute all waves of a single workgroup of
+  // maximum size on a CU.
+  std::pair<unsigned, unsigned> Default = {
+      getWavesPerEUForWorkGroup(FlatWorkGroupSizes.second),
+      getOccupancyWithWorkGroupSizes(LDSBytes, FlatWorkGroupSizes).second};
+  Default.first = std::min(Default.first, Default.second);
 
   // Requested minimum/maximum number of waves per execution unit.
-  std::pair<unsigned, unsigned> Requested =
-      AMDGPU::getIntegerPairAttribute(F, "amdgpu-waves-per-eu", Default, true);
-  return getEffectiveWavesPerEU(Requested, FlatWorkGroupSizes, LDSBytes);
+  std::pair<unsigned, unsigned> Requested = AMDGPU::getIntegerPairAttribute(
+      F, "amdgpu-waves-per-eu", {1, getMaxWavesPerEU()}, true);
+
+  // Make sure requested minimum is within the default range and lower than the
+  // requested maximum. The latter must not violate target specification.
+  if (Requested.first < Default.first || Requested.first > Default.second ||
----------------
shiltian wrote:

Agreed. Since waves-per-eu is not an ABI, to have it clamped also conforms with its semantics.

https://github.com/llvm/llvm-project/pull/221600


More information about the llvm-commits mailing list