[llvm] [X86][CostModel] Add per-shape gather/scatter cost tables for AMD znver4+ (PR #199488)
via llvm-commits
llvm-commits at lists.llvm.org
Wed Sep 2 19:49:29 PDT 2026
================
@@ -6659,24 +6814,66 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
std::pair<InstructionCost, MVT> SrcLT = getTypeLegalizationCost(SrcVTy);
InstructionCost::CostType SplitFactor =
std::max(IdxsLT.first, SrcLT.first).getValue();
+ const bool IsLoad = Opcode == Instruction::Load;
+
+ // A split op (e.g. 16-wide with 64-bit indices that don't fit one zmm)
+ // issues SplitFactor sub-ops; for code size that is just the sub-op count.
+ if (CostKind == TTI::TCK_CodeSize)
+ return SplitFactor;
+
+ // AMD znver4+ (TuningPreferGSCostTable): overhead + schedule-model body, both
+ // keyed on the *native* shape, so a shape fitting one register is charged
+ // once however legalisation splits it (narrow- and wide-index forms then
+ // report alike). Wider-than-native shapes tile into GSMul native ops.
+ if (CostKind == TTI::TCK_RecipThroughput && ST->hasPreferGSCostTable() &&
+ ST->hasAVX512()) {
+ Type *EltTy = SrcVTy->getScalarType();
+ // Take the width from the legalized type, not the IR one: a pointer
+ // element gathers as its integer equivalent (the same VPGATHERQQ), but
+ // reports a primitive size of 0.
+ EVT SrcEVT = TLI->getValueType(DL, SrcVTy);
+ unsigned EltBits = SrcEVT.isVector() ? SrcEVT.getScalarSizeInBits() : 0;
+ if (EltBits == 32 || EltBits == 64) {
+ unsigned NativeMaxElts = 512 / EltBits;
+ // Split only for genuinely-wider-than-native shapes, and only when the
+ // width is an exact multiple so the native op tiles cleanly.
+ unsigned NativeVF = std::min(VF, NativeMaxElts);
+ if (VF % NativeVF == 0) {
+ unsigned GSMul = VF / NativeVF;
+ auto *NativeVTy = FixedVectorType::get(EltTy, NativeVF);
+ std::optional<unsigned> Overhead = getZenGSOverhead(IsLoad, NativeVTy);
+ std::optional<unsigned> Body =
+ getModeledGSInstrCost(IsLoad, NativeVTy, CostKind);
+ if (Overhead && Body)
+ return GSMul * (*Overhead + *Body);
+ }
+ }
+ }
+
+ // Otherwise (a non-enumerated VF, or a target without
+ // TuningPreferGSCostTable): flat break-even overhead plus body. The overhead
+ // belongs to the pre-split shape and is paid once; only the body
+ // (VF * scalar-memory-op) scales with the split factor. Not guaranteed
+ // monotonic against the calibrated table - v6i32 can read cheaper than
+ // v4i32's measured total - but such VFs are rare and the table above remains
+ // the source of truth.
+ const int GSOverhead = IsLoad ? getGatherOverhead() : getScatterOverhead();
----------------
MattPD wrote:
Confirmed: The broad generic overhead change and its effects on Intel vector costs were removed.
https://github.com/llvm/llvm-project/pull/199488
More information about the llvm-commits
mailing list