[llvm] [LV] Enable wide active lane masks when tail-folding if preferred (PR #193757)

Kerry McLaughlin via llvm-commits llvm-commits at lists.llvm.org
Fri Jul 10 09:46:31 PDT 2026


================
@@ -4650,6 +4664,33 @@ LoopVectorizationCostModel::getMemoryInstructionCost(Instruction *I,
   return getWideningCost(I, VF);
 }
 
+bool LoopVectorizationCostModel::shouldWidenActiveLaneMask(ElementCount VF,
+                                                           unsigned IC) {
+  if (EnableWideActiveLaneMask.getValue() == WideActiveLaneMask::Force ||
+      ForceTargetInstructionCost.getNumOccurrences() > 0)
+    return true;
+
+  LLVMContext &Ctx = TheFunction->getContext();
+  Type *ArgTy = Type::getInt64Ty(Ctx);
+
+  // Compare the cost of one narrow mask per part vs one wide lane mask
+  // with extracts.
+  Type *ResTy = VectorType::get(Type::getInt1Ty(Ctx), VF);
+  InstructionCost ALMCost =
+      TTI.getActiveLaneMaskCost(ResTy, ArgTy, FastMathFlags(),
+                                TTI::TCK_RecipThroughput, 1) *
+      IC;
+
+  ResTy = VectorType::get(Type::getInt1Ty(Ctx), VF * IC);
+  InstructionCost WideALMCost = TTI.getActiveLaneMaskCost(
+      ResTy, ArgTy, FastMathFlags(), TTI::TCK_RecipThroughput, IC);
----------------
kmclaughlin-arm wrote:

The reason for adding the TTI hook was for AArch64, where we may be able to use the whilelo (predicate pair) instructions depending on the features available on the target. This instruction returns two results, so the cost is different as NumResults can be halved and this would also not require any extracts.

I will remove the cost model changes from this PR when making wide lane masks the canonical form though. I am planning for interleaving of tail-folded loops to only be allowed when forced, and it will be up to each target to choose the best form (e.g. https://github.com/llvm/llvm-project/pull/202909 for AArch64).

https://github.com/llvm/llvm-project/pull/193757


More information about the llvm-commits mailing list