[llvm] [AArch64] [CostModel] Improve costs for scalar inserts into fixed-length SVE constant vector (PR #223638)
Sander de Smalen via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 22 02:58:27 PDT 2026
================
@@ -4866,9 +4866,57 @@ InstructionCost AArch64TTIImpl::getScalarizationOverhead(
TTI::VectorInstrContext VIC) const {
if (isa<ScalableVectorType>(Ty))
return InstructionCost::getInvalid();
- if (Ty->getElementType()->isFloatingPointTy())
- return BaseT::getScalarizationOverhead(Ty, DemandedElts, Insert, Extract,
- CostKind);
+ if (Ty->getElementType()->isFloatingPointTy()) {
+ InstructionCost Cost = BaseT::getScalarizationOverhead(
+ Ty, DemandedElts, Insert, Extract, CostKind);
+
+ if (!Insert || VL.empty())
+ return Cost;
+
+ auto LT = getTypeLegalizationCost(Ty);
+ if (!ST->isNeonAvailable() || !LT.second.isFixedLengthVector() ||
+ LT.second.getFixedSizeInBits() <= 128)
+ return Cost;
+
+ auto HasNonUniformConstants = [&VL]() -> bool {
----------------
sdesmalen-arm wrote:
This PR is combines two different changes:
* It tries to return a more accurate cost for the specific case of "splat(constant) + inserts"
* It tries to return a more accurate cost for `BUILD_VECTOR`s (i.e. `Insert = true`) for wide vectors when using SVE.
I think each of them should be a PR on its own.
The code below also treats the "insert" and "extract" case in a similar way, although the splitting into 128-bit subvectors is only something we'd do for `BUILD_VECTOR` lowering and so specifically relates to inserts.
https://github.com/llvm/llvm-project/pull/223638
More information about the llvm-commits
mailing list