[llvm] [AArch64] [CostModel] Improve costs for scalar inserts into fixed-length SVE constant vector (PR #223638)

Sander de Smalen via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 22 02:58:27 PDT 2026


================
@@ -4866,9 +4866,57 @@ InstructionCost AArch64TTIImpl::getScalarizationOverhead(
     TTI::VectorInstrContext VIC) const {
   if (isa<ScalableVectorType>(Ty))
     return InstructionCost::getInvalid();
-  if (Ty->getElementType()->isFloatingPointTy())
-    return BaseT::getScalarizationOverhead(Ty, DemandedElts, Insert, Extract,
-                                           CostKind);
+  if (Ty->getElementType()->isFloatingPointTy()) {
+    InstructionCost Cost = BaseT::getScalarizationOverhead(
+        Ty, DemandedElts, Insert, Extract, CostKind);
+
+    if (!Insert || VL.empty())
+      return Cost;
+
+    auto LT = getTypeLegalizationCost(Ty);
+    if (!ST->isNeonAvailable() || !LT.second.isFixedLengthVector() ||
+        LT.second.getFixedSizeInBits() <= 128)
+      return Cost;
+
+    auto HasNonUniformConstants = [&VL]() -> bool {
----------------
sdesmalen-arm wrote:

This PR is combines two different changes:
* It tries to return a more accurate cost for the specific case of "splat(constant) + inserts"
* It tries to return a more accurate cost for `BUILD_VECTOR`s (i.e. `Insert = true`) for wide vectors when using SVE.

I think each of them should be a PR on its own.

The code below also treats the "insert" and "extract" case in a similar way, although the splitting into 128-bit subvectors is only something we'd do for `BUILD_VECTOR` lowering and so specifically relates to inserts.

https://github.com/llvm/llvm-project/pull/223638


More information about the llvm-commits mailing list