[llvm] [X86][CostModel] Add per-shape gather/scatter cost tables for AMD znver4+ (PR #199488)

via llvm-commits llvm-commits at lists.llvm.org
Wed Sep 2 19:49:29 PDT 2026


================
@@ -6659,24 +6814,66 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
   std::pair<InstructionCost, MVT> SrcLT = getTypeLegalizationCost(SrcVTy);
   InstructionCost::CostType SplitFactor =
       std::max(IdxsLT.first, SrcLT.first).getValue();
+  const bool IsLoad = Opcode == Instruction::Load;
+
+  // A split op (e.g. 16-wide with 64-bit indices that don't fit one zmm)
+  // issues SplitFactor sub-ops; for code size that is just the sub-op count.
+  if (CostKind == TTI::TCK_CodeSize)
+    return SplitFactor;
+
+  // AMD znver4+ (TuningPreferGSCostTable): overhead + schedule-model body, both
+  // keyed on the *native* shape, so a shape fitting one register is charged
+  // once however legalisation splits it (narrow- and wide-index forms then
+  // report alike). Wider-than-native shapes tile into GSMul native ops.
+  if (CostKind == TTI::TCK_RecipThroughput && ST->hasPreferGSCostTable() &&
+      ST->hasAVX512()) {
+    Type *EltTy = SrcVTy->getScalarType();
+    // Take the width from the legalized type, not the IR one: a pointer
+    // element gathers as its integer equivalent (the same VPGATHERQQ), but
+    // reports a primitive size of 0.
+    EVT SrcEVT = TLI->getValueType(DL, SrcVTy);
+    unsigned EltBits = SrcEVT.isVector() ? SrcEVT.getScalarSizeInBits() : 0;
+    if (EltBits == 32 || EltBits == 64) {
+      unsigned NativeMaxElts = 512 / EltBits;
+      // Split only for genuinely-wider-than-native shapes, and only when the
+      // width is an exact multiple so the native op tiles cleanly.
+      unsigned NativeVF = std::min(VF, NativeMaxElts);
+      if (VF % NativeVF == 0) {
+        unsigned GSMul = VF / NativeVF;
+        auto *NativeVTy = FixedVectorType::get(EltTy, NativeVF);
+        std::optional<unsigned> Overhead = getZenGSOverhead(IsLoad, NativeVTy);
+        std::optional<unsigned> Body =
+            getModeledGSInstrCost(IsLoad, NativeVTy, CostKind);
+        if (Overhead && Body)
+          return GSMul * (*Overhead + *Body);
+      }
+    }
+  }
+
+  // Otherwise (a non-enumerated VF, or a target without
+  // TuningPreferGSCostTable): flat break-even overhead plus body. The overhead
+  // belongs to the pre-split shape and is paid once; only the body
+  // (VF * scalar-memory-op) scales with the split factor. Not guaranteed
+  // monotonic against the calibrated table - v6i32 can read cheaper than
+  // v4i32's measured total - but such VFs are rare and the table above remains
+  // the source of truth.
+  const int GSOverhead = IsLoad ? getGatherOverhead() : getScatterOverhead();
----------------
MattPD wrote:

Confirmed: The broad generic overhead change and its effects on Intel vector costs were removed.

https://github.com/llvm/llvm-project/pull/199488


More information about the llvm-commits mailing list