[llvm] [X86][CostModel] Cost gathers by the instructions CodeGen emits (PR #220565)
Sumukh J Bharadwaj via llvm-commits
llvm-commits at lists.llvm.org
Thu Sep 17 01:46:39 PDT 2026
================
@@ -6633,50 +6667,97 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
const Value *Ptrs = GEP->getPointerOperand();
if (Ptrs->getType()->isVectorTy() && !getSplatValue(Ptrs))
return IndexSize;
- for (unsigned I = 1, E = GEP->getNumOperands(); I != E; ++I) {
- if (isa<Constant>(GEP->getOperand(I)))
+ for (gep_type_iterator GTI = gep_type_begin(GEP), GTE = gep_type_end(GEP);
+ GTI != GTE; ++GTI) {
+ const Value *Operand = GTI.getOperand();
+ if (isa<Constant>(Operand))
continue;
- Type *IndxTy = GEP->getOperand(I)->getType();
+ Type *IndxTy = Operand->getType();
if (auto *IndexVTy = dyn_cast<VectorType>(IndxTy))
IndxTy = IndexVTy->getElementType();
- if ((IndxTy->getPrimitiveSizeInBits() == 64 &&
- !isa<SExtInst>(GEP->getOperand(I))) ||
+ if ((IndxTy->getPrimitiveSizeInBits() == 64 && !isa<SExtInst>(Operand)) ||
++NumOfVarIndices > 1)
return IndexSize; // 64
+ // The narrow index only reaches the instruction if the addressing mode
+ // can apply its stride as a scale, and the scale field encodes 1, 2, 4
+ // and 8. Any other stride has to be multiplied into the index first,
+ // and that product is pointer-width, so the operation ends up gathering
+ // with 64-bit indices however narrow the index started out.
+ TypeSize EltSize = DL.getTypeAllocSize(GTI.getIndexedType());
+ if (EltSize.isScalable())
+ return IndexSize; // 64
+ uint64_t Stride = EltSize.getFixedValue();
+ if (!isPowerOf2_64(Stride) || Stride > 8)
----------------
amd-subharad wrote:
Adopted. The stride no longer decides the width on its own.
https://github.com/llvm/llvm-project/pull/220565
More information about the llvm-commits
mailing list