[llvm] [AMDGPU][TTI] Refine gfx9 packed FP32 SLP costs for pair formation and shuffles (PR #208572)

Akash Dutta via llvm-commits llvm-commits at lists.llvm.org
Mon Sep 21 10:48:06 PDT 2026


================
@@ -543,8 +543,8 @@ define amdgpu_kernel void @test_single_exp_hreduction(
 ; GCN-NEXT:  [[ENTRY:.*:]]
 ; GCN-NEXT:    [[P0:%.*]] = getelementptr float, ptr addrspace(1) [[INPUT]], i64 0
 ; GCN-NEXT:    [[TMP0:%.*]] = load <4 x float>, ptr addrspace(1) [[P0]], align 4
-; GCN-NEXT:    [[TMP1:%.*]] = call fast float @llvm.vector.reduce.fadd.v4f32(float 0.000000e+00, <4 x float> [[TMP0]])
-; GCN-NEXT:    [[EXP0:%.*]] = tail call float @llvm.amdgcn.exp2.f32(float [[TMP1]])
+; GCN-NEXT:    [[SUM:%.*]] = call fast float @llvm.vector.reduce.fadd.v4f32(float 0.000000e+00, <4 x float> [[TMP0]])
+; GCN-NEXT:    [[EXP0:%.*]] = tail call float @llvm.amdgcn.exp2.f32(float [[SUM]])
----------------
akadutta wrote:

reverted

https://github.com/llvm/llvm-project/pull/208572


More information about the llvm-commits mailing list