[llvm] [AMDGPU] PromoteAlloca: flatten homogeneous structs to vectors (PR #217055)

Domenic Nutile via llvm-commits llvm-commits at lists.llvm.org
Wed Aug 19 10:24:20 PDT 2026


================
@@ -0,0 +1,59 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 5
+; RUN: opt -S -mtriple=amdgcn-amd-amdhsa -passes=sroa,amdgpu-promote-alloca < %s | FileCheck %s
+
+%wrapper = type { [1 x i32] }
+%pair = type { i32, i32 }
+%nested = type { %pair, %pair }
+
+; A single-field wrapper struct: [4 x { [1 x i32] }] is really <4 x i32>.
+define amdgpu_kernel void @wrapper_struct(ptr addrspace(1) %out, i32 %idx) {
+; CHECK-LABEL: define amdgpu_kernel void @wrapper_struct(
+; CHECK-SAME: ptr addrspace(1) [[OUT:%.*]], i32 [[IDX:%.*]]) {
+; CHECK-NEXT:    [[ALLOCA:%.*]] = freeze <4 x i32> poison
+; CHECK-NEXT:    [[TMP1:%.*]] = insertelement <4 x i32> [[ALLOCA]], i32 1, i32 [[IDX]]
+; CHECK-NEXT:    store i32 1, ptr addrspace(1) [[OUT]], align 4
+; CHECK-NEXT:    ret void
+;
+  %alloca = alloca [4 x %wrapper], align 16, addrspace(5)
+  %gep = getelementptr [4 x %wrapper], ptr addrspace(5) %alloca, i32 0, i32 %idx
+  store i32 1, ptr addrspace(5) %gep, align 4
----------------
saxlungs wrote:

I made the test cases a bit more complex to avoid the fold into store of constant to make sure it wasn't hiding anything

https://github.com/llvm/llvm-project/pull/217055


More information about the llvm-commits mailing list