[llvm] [AMDGPU] Add partial unroll threshold function attribute (PR #223291)

via llvm-commits llvm-commits at lists.llvm.org
Sun Sep 13 20:13:20 PDT 2026


llvmorg-github-actions[bot] wrote:


<!--LLVM PR SUMMARY COMMENT-->

@llvm/pr-subscribers-backend-amdgpu

Author: nina-zhang-0

<details>
<summary>Changes</summary>

This change adds an `amdgpu-partial-unroll-threshold` function attribute for controlling `TargetTransformInfo::UnrollingPreferences::PartialThreshold` on a per-function basis.

The existing `amdgpu-unroll-threshold` function attribute initializes `UP.Threshold`, which is used for decisions about full
unrolling. However, there is currently no corresponding AMDGPU function attribute for configuring `UP.PartialThreshold`
independently. As a result, clients of the AMDGPU backend cannot provide an independent per-function cost threshold for partial
and runtime unrolling.

When present, the new attribute sets the base value of `UP.PartialThreshold` for loops in that function. This allows the cost
threshold for partial and runtime unrolling to be configured independently of the threshold used for full unrolling. Functions
that do not use the new attribute retain their existing behavior.

The AMDGPU function attribute documentation is updated accordingly. A regression test sets `amdgpu-unroll-threshold` to zero and `amdgpu-partial-unroll-threshold` to 30, then verifies that the loop is partially unrolled by a factor of two without being
fully unrolled.

Testing: `Transforms/LoopUnroll/AMDGPU/unroll-threshold.ll`

---
Full diff: https://github.com/llvm/llvm-project/pull/223291.diff


3 Files Affected:

- (modified) llvm/docs/AMDGPUUsage.rst (+4) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp (+2) 
- (modified) llvm/test/Transforms/LoopUnroll/AMDGPU/unroll-threshold.ll (+34) 


``````````diff
diff --git a/llvm/docs/AMDGPUUsage.rst b/llvm/docs/AMDGPUUsage.rst
index 80126312738635..8916b721ab0f6f 100644
--- a/llvm/docs/AMDGPUUsage.rst
+++ b/llvm/docs/AMDGPUUsage.rst
@@ -2732,6 +2732,10 @@ The AMDGPU backend supports the following LLVM IR attributes.
                                                       default is 300. Actual threshold may be varied by per-loop metadata or
                                                       reduced by heuristics.
 
+     "amdgpu-partial-unroll-threshold"                Set base cost threshold preference for partial and runtime loop unrolling
+                                                      within this function, default is 150. This does not change the threshold
+                                                      used for full unrolling.
+
      "amdgpu-max-num-workgroups"="x,y,z"              Specify the maximum number of work groups for the kernel dispatch in the
                                                       X, Y, and Z dimensions. Each number must be >= 1. Generated by the
                                                       ``amdgpu_max_num_work_groups`` CLANG attribute [CLANG-ATTR]_. Clang only
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp b/llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
index a7556278b7e0de..fbf4d6ed528884 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
@@ -117,6 +117,8 @@ void AMDGPUTTIImpl::getUnrollingPreferences(
   const Function &F = *L->getHeader()->getParent();
   UP.Threshold =
       F.getFnAttributeAsParsedInteger("amdgpu-unroll-threshold", 300);
+  UP.PartialThreshold = F.getFnAttributeAsParsedInteger(
+      "amdgpu-partial-unroll-threshold", UP.PartialThreshold);
   UP.MaxCount = std::numeric_limits<unsigned>::max();
   UP.Partial = true;
 
diff --git a/llvm/test/Transforms/LoopUnroll/AMDGPU/unroll-threshold.ll b/llvm/test/Transforms/LoopUnroll/AMDGPU/unroll-threshold.ll
index af437ba1d8a6a8..a41a13e2c1d5a6 100644
--- a/llvm/test/Transforms/LoopUnroll/AMDGPU/unroll-threshold.ll
+++ b/llvm/test/Transforms/LoopUnroll/AMDGPU/unroll-threshold.ll
@@ -104,8 +104,42 @@ do.end:                                           ; preds = %do.body
   ret void
 }
 
+; Check that the threshold used for partial unrolling is independent of the
+; threshold used for full unrolling. A partial threshold of 30 produces two
+; copies of the loop body while the threshold for full unrolling remains zero.
+; CHECK-LABEL: @partial_unroll_threshold(
+; CHECK: partial.body:
+; CHECK: store i32
+; CHECK: br i1
+; CHECK: partial.body.1:
+; CHECK: store i32
+; CHECK-NOT: partial.body.2:
+; CHECK: ret void
+
+define void @partial_unroll_threshold(ptr addrspace(1) %a,
+                                      ptr addrspace(1) %b) #2 {
+entry:
+  br label %partial.body
+
+partial.body:                                     ; preds = %entry, %partial.body
+  %iv = phi i64 [ 1, %entry ], [ %iv.next, %partial.body ]
+  %src = getelementptr inbounds i32, ptr addrspace(1) %b, i64 %iv
+  %value = load i32, ptr addrspace(1) %src, align 4
+  %index = sext i32 %value to i64
+  %dst = getelementptr inbounds i32, ptr addrspace(1) %a, i64 %index
+  %stored = trunc i64 %iv to i32
+  store i32 %stored, ptr addrspace(1) %dst, align 4
+  %iv.next = add nuw nsw i64 %iv, 1
+  %exitcond = icmp eq i64 %iv.next, 20
+  br i1 %exitcond, label %exit, label %partial.body
+
+exit:                                             ; preds = %partial.body
+  ret void
+}
+
 attributes #0 = { "amdgpu-unroll-threshold"="1000" }
 attributes #1 = { "amdgpu-unroll-threshold"="100" }
+attributes #2 = { "amdgpu-unroll-threshold"="0" "amdgpu-partial-unroll-threshold"="30" }
 
 !1 = !{!1, !2}
 !2 = !{!"amdgpu.loop.unroll.threshold", i32 1000}

``````````

</details>


https://github.com/llvm/llvm-project/pull/223291


More information about the llvm-commits mailing list