[llvm] [LV] Support tail-folded epilogue loops (PR #208764)
David Sherwood via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 7 06:56:51 PDT 2026
================
@@ -5851,16 +6026,20 @@ LoopVectorizationPlanner::computeBestVF() {
LoopVectorizationPlanner::LoopVectorizationPlanner(
Loop *L, LoopInfo *LI, DominatorTree *DT, const TargetLibraryInfo *TLI,
const TargetTransformInfo &TTI, LoopVectorizationLegality *Legal,
- std::unique_ptr<LoopVectorizationCostModel> CM, VFSelectionContext &Config,
- InterleavedAccessInfo &IAI, PredicatedScalarEvolution &PSE,
- OptimizationRemarkEmitter *ORE)
+ std::unique_ptr<LoopVectorizationCostModel> CM,
----------------
david-arm wrote:
Sorry, I can't leave a comment in the exact place in the code, but while building the LLVM test suite with the flags `-O3 -mcpu=native -Wno-unused-command-line-argument -mllvm -epilogue-tail-folding-policy=prefer-fold-tail -mllvm -force-vector-width="vscale x 4" -mllvm -epilogue-vectorization-force-VF="vscale x 2"` I hit this assert in computeMaxVF:
```
Assertion `VPlans[0]->getSingleVF() == UserVF && "expected second plan to be for the forced UserVF"' failed.
```
It's a bit unusual, but a lot of tests in the LLVM test suite want to disable vectorisation and this conflicts with epilogue vectorisation. Here's the reproducer:
```
target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i8:8:32-i16:16:32-i64:64-i128:128-n32:64-S128-Fn32"
target triple = "aarch64-unknown-linux-gnu"
define i32 @foo(i64 %wide.trip.count.i) {
entry:
br label %for.body.i
for.body.i:
%indvars.iv.i = phi i64 [ 0, %entry ], [ %indvars.iv.next.i, %for.body.i ]
%arrayidx3.i = getelementptr [4 x i8], ptr null, i64 %indvars.iv.i
%0 = load i32, ptr %arrayidx3.i, align 4
%indvars.iv.next.i = add i64 %indvars.iv.i, 1
%exitcond.not.i = icmp eq i64 %indvars.iv.i, %wide.trip.count.i
br i1 %exitcond.not.i, label %exit, label %for.body.i, !llvm.loop !0
exit:
ret i32 %0
}
!0 = distinct !{!0, !1, !2, !3}
!1 = !{!"llvm.loop.mustprogress"}
!2 = !{!"llvm.loop.vectorize.width", i32 1}
!3 = !{!"llvm.loop.interleave.count", i32 1}
```
and the command I used was `opt -p loop-vectorize -mcpu=native -epilogue-tail-folding-policy=prefer-fold-tail -force-vector-width="vscale x 4" -epilogue-vectorization-force-VF="vscale x 2" -S < /tmp/reduced.ll`
Can you add a similar test and bail out early to avoid the assert?
https://github.com/llvm/llvm-project/pull/208764
More information about the llvm-commits
mailing list