[llvm] [X86][CostModel] Cost gathers by the instructions CodeGen emits (PR #220565)

Simon Pilgrim via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 22 06:38:44 PDT 2026


================
@@ -6633,50 +6668,160 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
     const Value *Ptrs = GEP->getPointerOperand();
     if (Ptrs->getType()->isVectorTy() && !getSplatValue(Ptrs))
       return IndexSize;
-    for (unsigned I = 1, E = GEP->getNumOperands(); I != E; ++I) {
-      if (isa<Constant>(GEP->getOperand(I)))
+    for (gep_type_iterator GTI = gep_type_begin(GEP), GTE = gep_type_end(GEP);
+         GTI != GTE; ++GTI) {
+      const Value *Operand = GTI.getOperand();
+
+      // An index that holds the same value in every lane contributes a fixed
+      // byte offset, which folds into the base pointer instead of reaching the
+      // index operand. That covers every struct member, named by a scalar
+      // constant, as well as a splat. A constant that is not a splat does
+      // reach the index, so it is examined like any other.
+      if (isa<Constant>(Operand) && (!Operand->getType()->isVectorTy() ||
+                                     getSplatValue(Operand) != nullptr))
         continue;
-      Type *IndxTy = GEP->getOperand(I)->getType();
-      if (auto *IndexVTy = dyn_cast<VectorType>(IndxTy))
-        IndxTy = IndexVTy->getElementType();
-      if ((IndxTy->getPrimitiveSizeInBits() == 64 &&
-           !isa<SExtInst>(GEP->getOperand(I))) ||
-          ++NumOfVarIndices > 1)
+
+      if (++NumOfVarIndices > 1)
+        return IndexSize; // 64
+
+      unsigned IndexBits;
+      if (Operand->getType()->isVectorTy()) {
+        // Whether the index reaches the instruction as a dword is a property
+        // of the values it takes, not of how it was written: CodeGen narrows
+        // it when it is representable in a signed dword, which is what the
+        // gather/scatter combine in X86ISelLowering tests with
+        // ComputeNumSignBits. Deciding on the declared width and the
+        // extension opcode instead misses in both directions -- a
+        // zero-extension from a narrow type is representable although it is
+        // not an SExtInst, and an index wider than a pointer is not
+        // representable although it is not exactly 64 bits -- and a constant
+        // vector was not examined at all.
+        IndexBits = ComputeMaxSignificantBits(Operand, DL);
+      } else {
+        // A scalar index reaches here from a caller that has not widened the
+        // address yet, so the component that varies across lanes is not
+        // visible and its range says nothing about the index the instruction
+        // will see. Keep the conservative width for those until the query
+        // carries the vector form.
+        IndexBits = Operand->getType()->getPrimitiveSizeInBits() == 64 &&
----------------
RKSimon wrote:

There's still some missed optimizations for handling unsigned indices in combineGatherScatter - see #163023

https://github.com/llvm/llvm-project/pull/220565


More information about the llvm-commits mailing list