[llvm] [AMDGPU] Gate runtime unroll of LDS loops instead of overriding it (PR #222853)

Matt Arsenault via llvm-commits llvm-commits at lists.llvm.org
Fri Sep 11 01:15:50 PDT 2026


================
@@ -0,0 +1,415 @@
+; RUN: opt -mtriple=amdgcn-amd-amdhsa -mcpu=gfx950 -passes=loop-unroll \
+; RUN:     -amdgpu-unroll-runtime-local=false -S %s | FileCheck %s --check-prefix=NOLOCAL
+;
+; Reduced triton bf16 GEMM K-loop (gfx950): convergent MFMA body reading LDS
+; (addrspace(3) @global_smem). Runtime unrolling here adds a by-2 epilogue that
+; doubles the MFMA body and regresses occupancy. With
+; -amdgpu-unroll-runtime-local off, this LDS K-loop must not be runtime-unrolled.
+;
+; NOLOCAL-LABEL: @_gemm_a16_w16_kernel
+; NOLOCAL-NOT: xtraiter
+; NOLOCAL-NOT: .epil
+; NOLOCAL-NOT: unr-lcssa
+
+ at global_smem = external addrspace(3) global [0 x i8], align 16
+
+define amdgpu_kernel void @_gemm_a16_w16_kernel_BLOCK_SIZE_M_16_BLOCK_SIZE_N_16_BLOCK_SIZE_K_256_GROUP_SIZE_M_1_NUM_KSPLIT_1_SPLITK_BLOCK_SIZE_4096_EVEN_K_1_EVEN_MN_0_cache_modifier_CG_activation_NONE_use_activation_0_ADD_BIAS_0_SKIP_REDUCE_0(ptr addrspace(1) inreg nofree readonly captures(none) %0, ptr addrspace(1) inreg nofree readonly captures(none) %1, ptr addrspace(1) inreg nofree writeonly captures(none) %2, i32 inreg %3, i32 inreg %4, i32 inreg %5, i32 inreg %6, i32 inreg %7, i32 inreg %8, i32 inreg %9, ptr addrspace(1) inreg nofree readnone captures(none) %10, ptr addrspace(1) inreg nofree readnone captures(none) %11) {
+  %13 = tail call i32 @llvm.amdgcn.workitem.id.x()
----------------
arsenm wrote:

Use named values in test, and can this be simplified? There are an awful lot of arguments here 

https://github.com/llvm/llvm-project/pull/222853


More information about the llvm-commits mailing list