[llvm] [AMDGPU] Add waterfall intrinsics (PR #192409)

via llvm-commits llvm-commits at lists.llvm.org
Wed Jun 3 05:08:48 PDT 2026


================
@@ -1959,6 +1959,122 @@ enabled this should map to ``scope:SCOPE_SE``.
 
 **Note:** Cache control bits for Store are not affected by WGP mode.
 
+'``llvm.amdgcn.waterfall``' Intrinsics
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+The ``llvm.amdgcn.waterfall`` :ref:`family of intrinsics<amdgpu-waterfall-intrinsics-table>`
+describes how to generate a waterfall loop around a region of code.
+
+:Background:
+
+A waterfall loop handles the case where an operation that requires a uniform
+operand (e.g., in an SGPR) is applied to a non-uniform operand (held in a
+VGPR, with values varying per lane).
+
+Each iteration of the waterfall loop activates a subset of lanes that share
+the same value of the non-uniform operand (the value in the first active
+lane). The operation is then executed using that value as the uniform
+operand.
+
+If the operand is already uniform, the waterfall loop executes only
+once. The worst case for a waterfall loop is one iteration per lane (all
+lanes have different values of the operand), but this is not common in
+practice.
+
+:Motivation:
+
+The ``llvm.amdgcn.waterfall.*`` intrinsics provide a practical way for a
+frontend to combine several operations into a single waterfall loop with an
+efficient iteration strategy over multiple non-uniform operands.
+
+In particular, a frontend can specify a set of values that
+is sufficient to use as the "index" to loop over, i.e., lanes that have the
+same index values must have the same values for all
+non-uniform operands that need to be uniform. It is hard for the
+backend to find a fitting index, but it can be easy
+for a frontend.
+
+:Implementation:
+
+A group of waterfall intrinsics that depend on the same token defines a
+single waterfall loop. The group identifies a region of code, the
+non-uniform operands that need to be made uniform, and the values
+to use as the "index" of the loop.
+
+A waterfall group must contain at least one
+``waterfall.begin``, at least one ``waterfall.readfirstlane``, and at least one
+``waterfall.end`` intrinsic.
----------------
gretay-amd wrote:

> Shouldn't there be exactly one waterfall.end? Why would more than one be allowed?

`waterfall.end` is per value, not per region. Each `waterfall.end` marks one value that needs to  be available after the waterfall loop and gives it a new name used outside waterfall region. 
There are dedicated tests for it, for example, `test_waterfall_multi_end_struct` in
`llvm/test/CodeGen/AMDGPU/llvm.amdgcn.waterfall.ll`.

I add a note about it to the description of `waterfall.end` in the docs.

https://github.com/llvm/llvm-project/pull/192409


More information about the llvm-commits mailing list