[llvm] [AMDGPU][TTI] Refine gfx9 packed FP32 SLP costs for pair formation and shuffles (PR #208572)

Krzysztof Drewniak via llvm-commits llvm-commits at lists.llvm.org
Thu Aug 20 12:03:36 PDT 2026


================
@@ -1046,6 +1046,18 @@ InstructionCost GCNTTIImpl::getVectorInstrCost(
                                        VIC);
     }
 
+    // Gfx9 packed f32 pair formation for v_pk_*_f32: load-fed inserts are free;
+    // compute-fed inserts cost scales with legalization (wide vectors split to
+    // native <2 x f32>).
+    if (Opcode == Instruction::InsertElement && EltSize == 32 &&
+        ST->hasPackedFP32Ops())
+      if (auto *VecTy = dyn_cast<FixedVectorType>(ValTy))
+        if (VecTy->getElementType()->isFloatTy()) {
----------------
krzysz00 wrote:

1. I'm more making the general point that feeding 32 bit values of any kind into a register with a load is basically free, no? I still don't see why this is restricted to floats
2. Yeah, makes sense

https://github.com/llvm/llvm-project/pull/208572


More information about the llvm-commits mailing list