[llvm] [LoopVectorize] Improve Vectorization of Low Trip Count Loops (PR #195823)

Sander de Smalen via llvm-commits llvm-commits at lists.llvm.org
Thu Aug 20 01:29:12 PDT 2026


================
@@ -3055,17 +3057,41 @@ LoopVectorizationCostModel::computeMaxVF(ElementCount UserVF, unsigned UserIC) {
   if (ExpectedTC && ExpectedTC->isFixed() &&
       ExpectedTC->getFixedValue() <=
           TTI.getMinTripCountTailFoldingThreshold()) {
-    if (MaxPowerOf2RuntimeVF > 0u) {
+    if (MaxPowerOf2RuntimeVF > 0u &&
+        EpilogueLoweringStatus == CM_EpilogueNotAllowedLowTripLoop) {
       // If we have a low-trip-count, and the fixed-width VF is known to divide
       // the trip count but the scalable factor does not, use the fixed-width
       // factor in preference to allow the generation of a non-predicated loop.
-      if (EpilogueLoweringStatus == CM_EpilogueNotAllowedLowTripLoop &&
-          NoScalarEpilogueNeeded(MaxFactors.FixedVF.getFixedValue())) {
+      if (const auto *Rem =
+              GetLoopRemainder(MaxFactors.FixedVF.getFixedValue(), EffectiveIC);
+          Rem && Rem->isZero()) {
         LLVM_DEBUG(dbgs() << "LV: Picking a fixed-width so that no tail will "
                              "remain for any chosen VF.\n");
         MaxFactors.ScalableVF = ElementCount::getScalable(0);
         return MaxFactors;
       }
+      // Allow cases where the ExactTC == (VF * IC) + 1.
+      //
+      // This produces 1 vector iteration, and 1 scalar iteration with
+      // no remainder. Later passes will eliminate the loop and leave
+      // straight-line code as the both iteration counts are statically known.
+      //
+      // If a function is marked as minsize/optsize or OptForSize is set, do not
+      // allow this form of transformation as this will increase CodeSize.
+      ElementCount ExactTC = getSmallConstantTripCount(PSE.getSE(), TheLoop);
+      if (ExactTC.getFixedValue() > 1 && !TheFunction->hasOptSize() &&
+          !Config.OptForSize) {
+        unsigned TC = ExactTC.getFixedValue();
+        unsigned VF = (TC - 1) / EffectiveIC;
+        if (const auto *Rem = GetLoopRemainder(VF, EffectiveIC);
----------------
sdesmalen-arm wrote:

My apologies, after you made this change I realised my suggestion for reusing `NoScalarEpilogueNeeded/GetLoopRemainder` was unnecessary.

I think you can just set the Max FixedVF to `1ULL << Log2_32(ExactTC.getFixedValue())` if this is a valid MaxVF (i.e. it's lower than what `computeFeasibleMaxVF` returned) and if it doesn't result in more than 1 scalar iteration, because that would mean selecting the MaxVF as the largest power-of-2 VF that fits the tripcount of the loop (which ensures there will only ever be 1 vector iteration).

https://github.com/llvm/llvm-project/pull/195823


More information about the llvm-commits mailing list