[llvm] bd439d5 - [RISCV] Don't overcost wide load in optimized segment load/store (#207146)
via llvm-commits
llvm-commits at lists.llvm.org
Thu Aug 20 18:50:24 PDT 2026
Author: Luke Lau
Date: 2026-08-21T01:50:19Z
New Revision: bd439d56eca9a2a95e6f3747fb25200dd39c7e6f
URL: https://github.com/llvm/llvm-project/commit/bd439d56eca9a2a95e6f3747fb25200dd39c7e6f
DIFF: https://github.com/llvm/llvm-project/commit/bd439d56eca9a2a95e6f3747fb25200dd39c7e6f.diff
LOG: [RISCV] Don't overcost wide load in optimized segment load/store (#207146)
With the +optimized-nfX-segment-load-store tuning flag, we cost a
segmented store as a single wide load + some shuffle ops.
However for e.g. a `<vscale x 5 x i64>` Factor=5 segmented load, a wide
`<vscale x 5 x i64>` load gets costed as a full LMUL 8 load.
>From what I can see on
https://camel-cdr.github.io/rvv-bench-results/spacemit_x100/index.html
and on my own measurements on the spacemit-x60, uarchs likely don't do a
full LMUL 8 load under the hood and instead dispatch the minimum number
of DLEN sized ops needed for the full segment.
This changes the wide load cost to be divideCeil(vector size, DLEN) ops
so we don't overcost it.
Whilst we're here, this also removes the LT.first legalization
multiplier. We're computing the cost in terms of the unlegalized type so
we shouldn't be scaling it by the legalization cost.
Added:
Modified:
llvm/lib/Target/RISCV/RISCVTargetTransformInfo.cpp
llvm/test/Transforms/LoopVectorize/RISCV/interleaved-cost.ll
Removed:
################################################################################
diff --git a/llvm/lib/Target/RISCV/RISCVTargetTransformInfo.cpp b/llvm/lib/Target/RISCV/RISCVTargetTransformInfo.cpp
index 3598835e2c26b..970b1605cc610 100644
--- a/llvm/lib/Target/RISCV/RISCVTargetTransformInfo.cpp
+++ b/llvm/lib/Target/RISCV/RISCVTargetTransformInfo.cpp
@@ -1130,14 +1130,18 @@ InstructionCost RISCVTTIImpl::getInterleavedMemoryOpCost(
TLI->isLegalInterleavedAccessType(SubVecTy, Factor, Alignment,
AddressSpace, DL)) {
- // Some processors optimize segment loads/stores as one wide memory op +
- // Factor * LMUL shuffle ops.
+ // Some processors optimize segment loads/stores as N * DLEN sized
+ // load ops + Factor * LMUL shuffle ops.
if (ST->hasOptimizedSegmentLoadStore(Factor)) {
- InstructionCost Cost =
- getMemoryOpCost(Opcode, VTy, Alignment, AddressSpace, CostKind);
+ unsigned VecSizeInBits =
+ getEstimatedVLFor(VTy) * VTy->getScalarSizeInBits();
+ unsigned VLENForTuning =
+ *getVScaleForTuning() * RISCV::RVVBitsPerBlock;
+ unsigned DLENForTuning = VLENForTuning / ST->getDLenFactor();
+ InstructionCost Cost = divideCeil(VecSizeInBits, DLENForTuning);
MVT SubVecVT = getTLI()->getValueType(DL, SubVecTy).getSimpleVT();
Cost += Factor * TLI->getLMULCost(SubVecVT);
- return LT.first * Cost;
+ return Cost;
}
// Otherwise, the cost is proportional to the number of elements (VL *
diff --git a/llvm/test/Transforms/LoopVectorize/RISCV/interleaved-cost.ll b/llvm/test/Transforms/LoopVectorize/RISCV/interleaved-cost.ll
index e8f98e6ead8b4..29c767728d1d3 100644
--- a/llvm/test/Transforms/LoopVectorize/RISCV/interleaved-cost.ll
+++ b/llvm/test/Transforms/LoopVectorize/RISCV/interleaved-cost.ll
@@ -109,10 +109,10 @@ define void @i8_factor_3(ptr %data, i64 %n) {
; OPT: Cost of 4 for VF vscale x 2: INTERLEAVE-GROUP with factor 3, ir<%p0>, vp<%evl>
; OPT: Cost of 5 for VF vscale x 4: INTERLEAVE-GROUP with factor 3, ir<%p0>, vp<%evl>
; OPT: Cost of 5 for VF vscale x 4: INTERLEAVE-GROUP with factor 3, ir<%p0>, vp<%evl>
-; OPT: Cost of 7 for VF vscale x 8: INTERLEAVE-GROUP with factor 3, ir<%p0>, vp<%evl>
-; OPT: Cost of 7 for VF vscale x 8: INTERLEAVE-GROUP with factor 3, ir<%p0>, vp<%evl>
-; OPT: Cost of 14 for VF vscale x 16: INTERLEAVE-GROUP with factor 3, ir<%p0>, vp<%evl>
-; OPT: Cost of 14 for VF vscale x 16: INTERLEAVE-GROUP with factor 3, ir<%p0>, vp<%evl>
+; OPT: Cost of 6 for VF vscale x 8: INTERLEAVE-GROUP with factor 3, ir<%p0>, vp<%evl>
+; OPT: Cost of 6 for VF vscale x 8: INTERLEAVE-GROUP with factor 3, ir<%p0>, vp<%evl>
+; OPT: Cost of 12 for VF vscale x 16: INTERLEAVE-GROUP with factor 3, ir<%p0>, vp<%evl>
+; OPT: Cost of 12 for VF vscale x 16: INTERLEAVE-GROUP with factor 3, ir<%p0>, vp<%evl>
;
; FIXED-NO-OPT-LABEL: 'i8_factor_3'
; FIXED-NO-OPT: Cost of 6 for VF 2: INTERLEAVE-GROUP with factor 3, ir<%p0>
@@ -133,10 +133,10 @@ define void @i8_factor_3(ptr %data, i64 %n) {
; FIXED-OPT: Cost of 4 for VF 4: INTERLEAVE-GROUP with factor 3, ir<%p0>
; FIXED-OPT: Cost of 5 for VF 8: INTERLEAVE-GROUP with factor 3, ir<%p0>
; FIXED-OPT: Cost of 5 for VF 8: INTERLEAVE-GROUP with factor 3, ir<%p0>
-; FIXED-OPT: Cost of 7 for VF 16: INTERLEAVE-GROUP with factor 3, ir<%p0>
-; FIXED-OPT: Cost of 7 for VF 16: INTERLEAVE-GROUP with factor 3, ir<%p0>
-; FIXED-OPT: Cost of 14 for VF 32: INTERLEAVE-GROUP with factor 3, ir<%p0>
-; FIXED-OPT: Cost of 14 for VF 32: INTERLEAVE-GROUP with factor 3, ir<%p0>
+; FIXED-OPT: Cost of 6 for VF 16: INTERLEAVE-GROUP with factor 3, ir<%p0>
+; FIXED-OPT: Cost of 6 for VF 16: INTERLEAVE-GROUP with factor 3, ir<%p0>
+; FIXED-OPT: Cost of 12 for VF 32: INTERLEAVE-GROUP with factor 3, ir<%p0>
+; FIXED-OPT: Cost of 12 for VF 32: INTERLEAVE-GROUP with factor 3, ir<%p0>
;
; MINSIZE-LABEL: 'i8_factor_3'
; MINSIZE: Cost of 1 for VF vscale x 1: INTERLEAVE-GROUP with factor 3, ir<%p0>, vp<%evl>
@@ -281,10 +281,10 @@ define void @i8_factor_5(ptr %data, i64 %n) {
; OPT: Cost of 6 for VF vscale x 1: INTERLEAVE-GROUP with factor 5, ir<%p0>, vp<%evl>
; OPT: Cost of 7 for VF vscale x 2: INTERLEAVE-GROUP with factor 5, ir<%p0>, vp<%evl>
; OPT: Cost of 7 for VF vscale x 2: INTERLEAVE-GROUP with factor 5, ir<%p0>, vp<%evl>
-; OPT: Cost of 9 for VF vscale x 4: INTERLEAVE-GROUP with factor 5, ir<%p0>, vp<%evl>
-; OPT: Cost of 9 for VF vscale x 4: INTERLEAVE-GROUP with factor 5, ir<%p0>, vp<%evl>
-; OPT: Cost of 13 for VF vscale x 8: INTERLEAVE-GROUP with factor 5, ir<%p0>, vp<%evl>
-; OPT: Cost of 13 for VF vscale x 8: INTERLEAVE-GROUP with factor 5, ir<%p0>, vp<%evl>
+; OPT: Cost of 8 for VF vscale x 4: INTERLEAVE-GROUP with factor 5, ir<%p0>, vp<%evl>
+; OPT: Cost of 8 for VF vscale x 4: INTERLEAVE-GROUP with factor 5, ir<%p0>, vp<%evl>
+; OPT: Cost of 10 for VF vscale x 8: INTERLEAVE-GROUP with factor 5, ir<%p0>, vp<%evl>
+; OPT: Cost of 10 for VF vscale x 8: INTERLEAVE-GROUP with factor 5, ir<%p0>, vp<%evl>
;
; FIXED-NO-OPT-LABEL: 'i8_factor_5'
; FIXED-NO-OPT: Cost of 10 for VF 2: INTERLEAVE-GROUP with factor 5, ir<%p0>
@@ -301,10 +301,10 @@ define void @i8_factor_5(ptr %data, i64 %n) {
; FIXED-OPT: Cost of 6 for VF 2: INTERLEAVE-GROUP with factor 5, ir<%p0>
; FIXED-OPT: Cost of 7 for VF 4: INTERLEAVE-GROUP with factor 5, ir<%p0>
; FIXED-OPT: Cost of 7 for VF 4: INTERLEAVE-GROUP with factor 5, ir<%p0>
-; FIXED-OPT: Cost of 9 for VF 8: INTERLEAVE-GROUP with factor 5, ir<%p0>
-; FIXED-OPT: Cost of 9 for VF 8: INTERLEAVE-GROUP with factor 5, ir<%p0>
-; FIXED-OPT: Cost of 13 for VF 16: INTERLEAVE-GROUP with factor 5, ir<%p0>
-; FIXED-OPT: Cost of 13 for VF 16: INTERLEAVE-GROUP with factor 5, ir<%p0>
+; FIXED-OPT: Cost of 8 for VF 8: INTERLEAVE-GROUP with factor 5, ir<%p0>
+; FIXED-OPT: Cost of 8 for VF 8: INTERLEAVE-GROUP with factor 5, ir<%p0>
+; FIXED-OPT: Cost of 10 for VF 16: INTERLEAVE-GROUP with factor 5, ir<%p0>
+; FIXED-OPT: Cost of 10 for VF 16: INTERLEAVE-GROUP with factor 5, ir<%p0>
;
; MINSIZE-LABEL: 'i8_factor_5'
; MINSIZE: Cost of 1 for VF vscale x 1: INTERLEAVE-GROUP with factor 5, ir<%p0>, vp<%evl>
@@ -367,10 +367,10 @@ define void @i8_factor_6(ptr %data, i64 %n) {
; OPT: Cost of 7 for VF vscale x 1: INTERLEAVE-GROUP with factor 6, ir<%p0>, vp<%evl>
; OPT: Cost of 8 for VF vscale x 2: INTERLEAVE-GROUP with factor 6, ir<%p0>, vp<%evl>
; OPT: Cost of 8 for VF vscale x 2: INTERLEAVE-GROUP with factor 6, ir<%p0>, vp<%evl>
-; OPT: Cost of 10 for VF vscale x 4: INTERLEAVE-GROUP with factor 6, ir<%p0>, vp<%evl>
-; OPT: Cost of 10 for VF vscale x 4: INTERLEAVE-GROUP with factor 6, ir<%p0>, vp<%evl>
-; OPT: Cost of 14 for VF vscale x 8: INTERLEAVE-GROUP with factor 6, ir<%p0>, vp<%evl>
-; OPT: Cost of 14 for VF vscale x 8: INTERLEAVE-GROUP with factor 6, ir<%p0>, vp<%evl>
+; OPT: Cost of 9 for VF vscale x 4: INTERLEAVE-GROUP with factor 6, ir<%p0>, vp<%evl>
+; OPT: Cost of 9 for VF vscale x 4: INTERLEAVE-GROUP with factor 6, ir<%p0>, vp<%evl>
+; OPT: Cost of 12 for VF vscale x 8: INTERLEAVE-GROUP with factor 6, ir<%p0>, vp<%evl>
+; OPT: Cost of 12 for VF vscale x 8: INTERLEAVE-GROUP with factor 6, ir<%p0>, vp<%evl>
;
; FIXED-NO-OPT-LABEL: 'i8_factor_6'
; FIXED-NO-OPT: Cost of 12 for VF 2: INTERLEAVE-GROUP with factor 6, ir<%p0>
@@ -387,10 +387,10 @@ define void @i8_factor_6(ptr %data, i64 %n) {
; FIXED-OPT: Cost of 7 for VF 2: INTERLEAVE-GROUP with factor 6, ir<%p0>
; FIXED-OPT: Cost of 8 for VF 4: INTERLEAVE-GROUP with factor 6, ir<%p0>
; FIXED-OPT: Cost of 8 for VF 4: INTERLEAVE-GROUP with factor 6, ir<%p0>
-; FIXED-OPT: Cost of 10 for VF 8: INTERLEAVE-GROUP with factor 6, ir<%p0>
-; FIXED-OPT: Cost of 10 for VF 8: INTERLEAVE-GROUP with factor 6, ir<%p0>
-; FIXED-OPT: Cost of 14 for VF 16: INTERLEAVE-GROUP with factor 6, ir<%p0>
-; FIXED-OPT: Cost of 14 for VF 16: INTERLEAVE-GROUP with factor 6, ir<%p0>
+; FIXED-OPT: Cost of 9 for VF 8: INTERLEAVE-GROUP with factor 6, ir<%p0>
+; FIXED-OPT: Cost of 9 for VF 8: INTERLEAVE-GROUP with factor 6, ir<%p0>
+; FIXED-OPT: Cost of 12 for VF 16: INTERLEAVE-GROUP with factor 6, ir<%p0>
+; FIXED-OPT: Cost of 12 for VF 16: INTERLEAVE-GROUP with factor 6, ir<%p0>
;
; MINSIZE-LABEL: 'i8_factor_6'
; MINSIZE: Cost of 1 for VF vscale x 1: INTERLEAVE-GROUP with factor 6, ir<%p0>, vp<%evl>
@@ -459,8 +459,8 @@ define void @i8_factor_7(ptr %data, i64 %n) {
; OPT: Cost of 9 for VF vscale x 2: INTERLEAVE-GROUP with factor 7, ir<%p0>, vp<%evl>
; OPT: Cost of 11 for VF vscale x 4: INTERLEAVE-GROUP with factor 7, ir<%p0>, vp<%evl>
; OPT: Cost of 11 for VF vscale x 4: INTERLEAVE-GROUP with factor 7, ir<%p0>, vp<%evl>
-; OPT: Cost of 15 for VF vscale x 8: INTERLEAVE-GROUP with factor 7, ir<%p0>, vp<%evl>
-; OPT: Cost of 15 for VF vscale x 8: INTERLEAVE-GROUP with factor 7, ir<%p0>, vp<%evl>
+; OPT: Cost of 14 for VF vscale x 8: INTERLEAVE-GROUP with factor 7, ir<%p0>, vp<%evl>
+; OPT: Cost of 14 for VF vscale x 8: INTERLEAVE-GROUP with factor 7, ir<%p0>, vp<%evl>
;
; FIXED-NO-OPT-LABEL: 'i8_factor_7'
; FIXED-NO-OPT: Cost of 14 for VF 2: INTERLEAVE-GROUP with factor 7, ir<%p0>
@@ -479,8 +479,8 @@ define void @i8_factor_7(ptr %data, i64 %n) {
; FIXED-OPT: Cost of 9 for VF 4: INTERLEAVE-GROUP with factor 7, ir<%p0>
; FIXED-OPT: Cost of 11 for VF 8: INTERLEAVE-GROUP with factor 7, ir<%p0>
; FIXED-OPT: Cost of 11 for VF 8: INTERLEAVE-GROUP with factor 7, ir<%p0>
-; FIXED-OPT: Cost of 15 for VF 16: INTERLEAVE-GROUP with factor 7, ir<%p0>
-; FIXED-OPT: Cost of 15 for VF 16: INTERLEAVE-GROUP with factor 7, ir<%p0>
+; FIXED-OPT: Cost of 14 for VF 16: INTERLEAVE-GROUP with factor 7, ir<%p0>
+; FIXED-OPT: Cost of 14 for VF 16: INTERLEAVE-GROUP with factor 7, ir<%p0>
;
; MINSIZE-LABEL: 'i8_factor_7'
; MINSIZE: Cost of 1 for VF vscale x 1: INTERLEAVE-GROUP with factor 7, ir<%p0>, vp<%evl>
More information about the llvm-commits
mailing list