[llvm] [AMDGPU][TTI] Refine gfx9 packed FP32 SLP costs for pair formation and shuffles (PR #208572)
Krzysztof Drewniak via llvm-commits
llvm-commits at lists.llvm.org
Thu Aug 20 12:03:36 PDT 2026
================
@@ -1046,6 +1046,18 @@ InstructionCost GCNTTIImpl::getVectorInstrCost(
VIC);
}
+ // Gfx9 packed f32 pair formation for v_pk_*_f32: load-fed inserts are free;
+ // compute-fed inserts cost scales with legalization (wide vectors split to
+ // native <2 x f32>).
+ if (Opcode == Instruction::InsertElement && EltSize == 32 &&
+ ST->hasPackedFP32Ops())
+ if (auto *VecTy = dyn_cast<FixedVectorType>(ValTy))
+ if (VecTy->getElementType()->isFloatTy()) {
----------------
krzysz00 wrote:
1. I'm more making the general point that feeding 32 bit values of any kind into a register with a load is basically free, no? I still don't see why this is restricted to floats
2. Yeah, makes sense
https://github.com/llvm/llvm-project/pull/208572
More information about the llvm-commits
mailing list