[llvm] [SLP] Add -slp-use-vplan-codegen to emit vector code via VPlan. (POC) (PR #226843)
Florian Hahn via llvm-commits
llvm-commits at lists.llvm.org
Sun Sep 27 15:10:42 PDT 2026
https://github.com/fhahn created https://github.com/llvm/llvm-project/pull/226843
Add an off-by-default option to generate the vector code for an SLP tree
through VPlan instead of the hand-rolled emitter in vectorizeTree().
This patch is intended as proof-of-concept, showing how VPlan could be
used as codegen backend for SLP, following a similar path like the
initial VPlan bring-up in LoopVectorize.
Moving the codegen to VPlan could allow moving some SLP functionality to
be VPlan based (e.g. simplifications, codegen-optimizations), helping
modularizing the code.
I have bigger prototype, which can handle about 80% of cases in the
VPlan path on a large test set of workloads. Compile-time impact looks
neutral, so I would not expect that to become a blocker.
Aided by Opus 5
>From 5091a7c9eec0a7c78c4355b173a6659b6df63e78 Mon Sep 17 00:00:00 2001
From: Florian Hahn <flo at fhahn.com>
Date: Sat, 1 Aug 2026 09:57:27 +0100
Subject: [PATCH 1/3] [VPlan] Don't populate the scalar header in the
BasicBlock constructor
VPlan(BasicBlock *, Type *) created a VPIRInstruction for every instruction in
the scalar header block. Neither of its two users wants that: the VPlan unit
tests pass a block containing only a terminator, and the SLP vectorizer only
needs the block as the plan's scalar header and never looks at its recipes.
Building a plan was therefore O(size of the containing block). SLP builds one
plan per vectorized tree, so on a function with a single large block this was
quadratic: compiling a 4000-group straight-line block with -O3 and VPlan-based
SLP codegen spent 71s, against 39s for the existing codegen path. Using
createEmptyVPIRBasicBlock() brings that to 49s, and the rest is addressed
separately.
The loop-based constructor, which does need the scalar loop header's recipes,
is unchanged.
---
llvm/lib/Transforms/Vectorize/VPlan.h | 6 ++++--
1 file changed, 4 insertions(+), 2 deletions(-)
diff --git a/llvm/lib/Transforms/Vectorize/VPlan.h b/llvm/lib/Transforms/Vectorize/VPlan.h
index 746c0231f6fcf..a5776a8e6f635 100644
--- a/llvm/lib/Transforms/Vectorize/VPlan.h
+++ b/llvm/lib/Transforms/Vectorize/VPlan.h
@@ -4890,11 +4890,13 @@ class VPlan {
VPlan(Loop *L, Type *IdxTy);
/// Construct a VPlan with a new VPBasicBlock as entry, a VPIRBasicBlock
- /// wrapping \p ScalarHeaderBB and vector loop index of type \p IdxTy.
+ /// wrapping \p ScalarHeaderBB and vector loop index of type \p IdxTy. The
+ /// scalar header is left empty; callers that need recipes for the
+ /// instructions in \p ScalarHeaderBB must create them themselves.
VPlan(BasicBlock *ScalarHeaderBB, Type *IdxTy)
: VectorTripCount(IdxTy), VF(IdxTy), UF(IdxTy), VFxUF(IdxTy) {
setEntry(createVPBasicBlock("preheader"));
- ScalarHeader = createVPIRBasicBlock(ScalarHeaderBB);
+ ScalarHeader = createEmptyVPIRBasicBlock(ScalarHeaderBB);
}
LLVM_ABI_FOR_TEST ~VPlan();
>From cf749a5e793771cd34ef052a8a4620ea5db03495 Mon Sep 17 00:00:00 2001
From: Florian Hahn <flo at fhahn.com>
Date: Sat, 1 Aug 2026 12:56:56 +0100
Subject: [PATCH 2/3] [VectorUtils] Add a getMetadataToPropagate overload for a
bundle (NFC)
Split the metadata merging out of propagateMetadata() so it can be used by
callers that do not have the combined instruction yet, and only want the
metadata that would be valid for it. propagateMetadata() becomes a thin
wrapper. The representative instruction is passed separately, as it is needed
to decide which metadata kinds the combined instruction can carry.
---
llvm/include/llvm/Analysis/VectorUtils.h | 9 ++++++++
llvm/lib/Analysis/VectorUtils.cpp | 29 ++++++++++++++++--------
2 files changed, 28 insertions(+), 10 deletions(-)
diff --git a/llvm/include/llvm/Analysis/VectorUtils.h b/llvm/include/llvm/Analysis/VectorUtils.h
index b177d9eec2189..d47bc503fe898 100644
--- a/llvm/include/llvm/Analysis/VectorUtils.h
+++ b/llvm/include/llvm/Analysis/VectorUtils.h
@@ -378,6 +378,15 @@ LLVM_ABI void getMetadataToPropagate(
Instruction *Inst,
SmallVectorImpl<std::pair<unsigned, MDNode *>> &Metadata);
+/// Add the metadata that can be preserved when combining all of \p VL into a
+/// single instruction to \p Metadata. \p Repr is a representative for the
+/// combined instruction, used to determine which metadata kinds it can carry.
+/// Callers that already have the combined instruction should use
+/// propagateMetadata() instead.
+LLVM_ABI void getMetadataToPropagate(
+ const Instruction *Repr, ArrayRef<Value *> VL,
+ SmallVectorImpl<std::pair<unsigned, MDNode *>> &Metadata);
+
/// Specifically, let Kinds = [MD_tbaa, MD_alias_scope, MD_noalias, MD_fpmath,
/// MD_nontemporal, MD_access_group, MD_mmra].
/// For K in Kinds, we get the MDNode for K from each of the
diff --git a/llvm/lib/Analysis/VectorUtils.cpp b/llvm/lib/Analysis/VectorUtils.cpp
index 1c105ebb772b3..8b5a0cbf0220e 100644
--- a/llvm/lib/Analysis/VectorUtils.cpp
+++ b/llvm/lib/Analysis/VectorUtils.cpp
@@ -1073,17 +1073,21 @@ void llvm::getMetadataToPropagate(
}
}
-/// \returns \p I after propagating metadata from \p VL.
-Instruction *llvm::propagateMetadata(Instruction *Inst, ArrayRef<Value *> VL) {
+/// Add metadata from all of \p VL to \p Metadata, if it can be preserved after
+/// combining them into \p Repr.
+void llvm::getMetadataToPropagate(
+ const Instruction *Repr, ArrayRef<Value *> VL,
+ SmallVectorImpl<std::pair<unsigned, MDNode *>> &Metadata) {
if (VL.empty())
- return Inst;
- SmallVector<std::pair<unsigned, MDNode *>> Metadata;
+ return;
getMetadataToPropagate(cast<Instruction>(VL[0]), Metadata);
for (auto &[Kind, MD] : Metadata) {
- // Skip MMRA metadata if the instruction cannot have it.
- if (Kind == LLVMContext::MD_mmra && !canInstructionHaveMMRAs(*Inst))
+ // Drop MMRA metadata if the combined instruction cannot have it.
+ if (Kind == LLVMContext::MD_mmra && !canInstructionHaveMMRAs(*Repr)) {
+ MD = nullptr;
continue;
+ }
for (int J = 1, E = VL.size(); MD && J != E; ++J) {
const Instruction *IJ = cast<Instruction>(VL[J]);
@@ -1091,7 +1095,7 @@ Instruction *llvm::propagateMetadata(Instruction *Inst, ArrayRef<Value *> VL) {
switch (Kind) {
case LLVMContext::MD_mmra: {
- MD = MMRAMetadata::combine(Inst->getContext(), MD, IMD);
+ MD = MMRAMetadata::combine(Repr->getContext(), MD, IMD);
break;
}
case LLVMContext::MD_tbaa:
@@ -1109,16 +1113,21 @@ Instruction *llvm::propagateMetadata(Instruction *Inst, ArrayRef<Value *> VL) {
MD = MDNode::intersect(MD, IMD);
break;
case LLVMContext::MD_access_group:
- MD = intersectAccessGroups(Inst, IJ);
+ MD = intersectAccessGroups(Repr, IJ);
break;
default:
llvm_unreachable("unhandled metadata");
}
}
-
- Inst->setMetadata(Kind, MD);
}
+}
+/// \returns \p Inst after propagating metadata from \p VL.
+Instruction *llvm::propagateMetadata(Instruction *Inst, ArrayRef<Value *> VL) {
+ SmallVector<std::pair<unsigned, MDNode *>> Metadata;
+ getMetadataToPropagate(Inst, VL, Metadata);
+ for (auto &[Kind, MD] : Metadata)
+ Inst->setMetadata(Kind, MD);
return Inst;
}
>From 748424937d46a647b5df020ef88185dac3fcce38 Mon Sep 17 00:00:00 2001
From: Florian Hahn <flo at fhahn.com>
Date: Tue, 22 Sep 2026 10:09:36 +0100
Subject: [PATCH 3/3] [SLP] Add -slp-use-vplan-codegen to emit vector code via
VPlan. (POC)
Add an off-by-default option to generate the vector code for an SLP tree
through VPlan instead of the hand-rolled emitter in vectorizeTree().
This patch is intended as proof-of-concept, showing how VPlan could be
used as codegen backend for SLP, following a similar path like the
initial VPlan bring-up in LoopVectorize.
Moving the codegen to VPlan could allow moving some SLP functionality to
be VPlan based (e.g. simplifications, codegen-optimizations), helping
modularizing the code.
I have bigger prototype, which can handle about 80% of cases in the
VPlan path on a large test set of workloads. Compile-time impact looks
neutral, so I would not expect that to become a blocker.
Aided by Opus 5
---
llvm/lib/Transforms/Vectorize/CMakeLists.txt | 1 +
.../Transforms/Vectorize/SLPVPlanCodegen.cpp | 83 +++++++++
.../Transforms/Vectorize/SLPVPlanCodegen.h | 51 ++++++
.../Transforms/Vectorize/SLPVectorizer.cpp | 171 +++++++++++++++++-
llvm/lib/Transforms/Vectorize/VPlan.h | 8 +
.../WebAssembly/simd-min-vec-reg-32.ll | 1 +
.../PhaseOrdering/AArch64/interleave_vec.ll | 1 +
.../SLPVectorizer/AArch64/32-bit.ll | 1 +
.../AArch64/masked-div-rem-non-pow2.ll | 1 +
.../AArch64/memory-runtime-checks.ll | 1 +
.../SLPVectorizer/AArch64/nontemporal.ll | 1 +
.../SLPVectorizer/AArch64/slp-frem.ll | 1 +
.../spillcost-call-between-operands.ll | 1 +
.../SLPVectorizer/AArch64/spillcost-di.ll | 1 +
.../SLPVectorizer/AArch64/store-ptr.ll | 1 +
.../SLPVectorizer/AArch64/vec3-base.ll | 2 +
.../AArch64/vplan-codegen-load-sinking.ll | 133 ++++++++++++++
.../AMDGPU/elementwise-fma-operand1.ll | 3 +
.../AMDGPU/fma-operand-contract-selection.ll | 2 +
.../AMDGPU/invariant-load-no-alias-store.ll | 1 +
.../SLPVectorizer/AMDGPU/packed-math.ll | 2 +
.../SLPVectorizer/AMDGPU/slp-v2f16.ll | 4 +
.../RISCV/basic-strided-loads.ll | 1 +
.../RISCV/basic-strided-stores.ll | 1 +
.../SLPVectorizer/RISCV/external.ll | 1 +
.../SLPVectorizer/RISCV/floating-point.ll | 5 +
.../Transforms/SLPVectorizer/RISCV/gep.ll | 2 +
.../SLPVectorizer/RISCV/load-binop-store.ll | 3 +
.../SLPVectorizer/RISCV/load-store.ll | 3 +
.../RISCV/rotated-strided-loads.ll | 1 +
.../RISCV/runtime-strided-stores.ll | 1 +
.../SLPVectorizer/RISCV/test-delete-tree.ll | 1 +
.../SLPVectorizer/RISCV/vec3-base.ll | 2 +
.../SLPVectorizer/VE/disable_slp.ll | 1 +
.../Transforms/SLPVectorizer/X86/align.ll | 1 +
.../SLPVectorizer/X86/arith-add-load.ll | 4 +
.../Transforms/SLPVectorizer/X86/arith-add.ll | 13 ++
.../SLPVectorizer/X86/arith-mul-load.ll | 4 +
.../Transforms/SLPVectorizer/X86/arith-mul.ll | 13 ++
.../Transforms/SLPVectorizer/X86/arith-sub.ll | 13 ++
.../SLPVectorizer/X86/continue_vectorizing.ll | 1 +
.../SLPVectorizer/X86/control-dependence.ll | 1 +
.../X86/fmuladd-copyable-add-part.ll | 2 +
.../X86/fsub-fmul-rhs-combine.ll | 1 +
.../Transforms/SLPVectorizer/X86/lookahead.ll | 2 +
.../Transforms/SLPVectorizer/X86/metadata.ll | 1 +
.../SLPVectorizer/X86/operandorder.ll | 2 +
llvm/test/Transforms/SLPVectorizer/X86/opt.ll | 1 +
.../SLPVectorizer/X86/reassociate-ops.ll | 1 +
.../SLPVectorizer/X86/runtime-alias-checks.ll | 1 +
.../SLPVectorizer/X86/schedule_budget.ll | 2 +
.../X86/schedule_budget_debug_info.ll | 2 +
.../SLPVectorizer/X86/shift-ashr.ll | 11 ++
.../SLPVectorizer/X86/shift-lshr.ll | 11 ++
.../Transforms/SLPVectorizer/X86/shift-shl.ll | 11 ++
.../Transforms/SLPVectorizer/X86/simplebb.ll | 1 +
.../Transforms/SLPVectorizer/X86/tiny-tree.ll | 1 +
.../SLPVectorizer/consecutive-access.ll | 2 +
.../SLPVectorizer/int_sideeffect.ll | 1 +
59 files changed, 591 insertions(+), 5 deletions(-)
create mode 100644 llvm/lib/Transforms/Vectorize/SLPVPlanCodegen.cpp
create mode 100644 llvm/lib/Transforms/Vectorize/SLPVPlanCodegen.h
create mode 100644 llvm/test/Transforms/SLPVectorizer/AArch64/vplan-codegen-load-sinking.ll
diff --git a/llvm/lib/Transforms/Vectorize/CMakeLists.txt b/llvm/lib/Transforms/Vectorize/CMakeLists.txt
index 9073211280886..80f7bb81b8a45 100644
--- a/llvm/lib/Transforms/Vectorize/CMakeLists.txt
+++ b/llvm/lib/Transforms/Vectorize/CMakeLists.txt
@@ -30,6 +30,7 @@ add_llvm_component_library(LLVMVectorize
SLPVectorizer/SLPTypeUtils.cpp
SLPVectorizer/SLPUtils.cpp
SLPVectorizer.cpp
+ SLPVPlanCodegen.cpp
Vectorize.cpp
VectorCombine.cpp
VPlan.cpp
diff --git a/llvm/lib/Transforms/Vectorize/SLPVPlanCodegen.cpp b/llvm/lib/Transforms/Vectorize/SLPVPlanCodegen.cpp
new file mode 100644
index 0000000000000..1f7f5c30cdeae
--- /dev/null
+++ b/llvm/lib/Transforms/Vectorize/SLPVPlanCodegen.cpp
@@ -0,0 +1,83 @@
+//===- SLPVPlanCodegen.cpp - VPlan-based codegen for SLP ------------------===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+#include "SLPVPlanCodegen.h"
+#include "LoopVectorizationPlanner.h"
+#include "VPlan.h"
+#include "VPlanHelpers.h"
+#include "llvm/ADT/STLExtras.h"
+#include "llvm/IR/IRBuilder.h"
+#include "llvm/IR/Instructions.h"
+
+using namespace llvm;
+
+bool slpvectorizer::isLiveInOperand(unsigned Opcode, unsigned J) {
+ // Loads and stores take their pointer operand as a live-in.
+ return (Opcode == Instruction::Load &&
+ J == LoadInst::getPointerOperandIndex()) ||
+ (Opcode == Instruction::Store &&
+ J == StoreInst::getPointerOperandIndex());
+}
+
+bool slpvectorizer::isSupportedVPlanCodegenOpcode(unsigned Opcode) {
+ if (Instruction::isBinaryOp(Opcode))
+ return true;
+ switch (Opcode) {
+ case Instruction::FNeg:
+ case Instruction::Load:
+ case Instruction::Store:
+ return true;
+ default:
+ return false;
+ }
+}
+
+/// Returns the IR flags common to all of \p Scalars, seeded from \p MainOp.
+static VPIRFlags computeIntersectedFlags(Instruction *MainOp,
+ ArrayRef<Value *> Scalars) {
+ VPIRFlags Flags(*MainOp);
+ for (Value *V : Scalars) {
+ // Drop all flags if a lane is not an instruction, e.g. poison.
+ auto *I = dyn_cast<Instruction>(V);
+ if (!I)
+ return VPIRFlags();
+ Flags.intersectFlags(VPIRFlags(*I));
+ }
+ return Flags;
+}
+
+VPValue *slpvectorizer::createRecipeForBundle(VPlan &Plan, VPBuilder &VPB,
+ Instruction *MainOp,
+ ArrayRef<Value *> Scalars,
+ ArrayRef<VPValue *> Ops) {
+ VPIRFlags Flags = computeIntersectedFlags(MainOp, Scalars);
+ // Only instructions carry metadata, so other lanes, e.g. poison, are skipped.
+ VPIRMetadata Metadata(
+ *MainOp, to_vector(make_filter_range(Scalars, IsaPred<Instruction>)));
+ DebugLoc DL = MainOp->getDebugLoc();
+
+ if (auto *LI = dyn_cast<LoadInst>(MainOp))
+ return VPB.createWidenLoad(
+ *LI, Plan.getOrAddLiveIn(LI->getPointerOperand()),
+ /*Mask=*/nullptr, /*Consecutive=*/true, Metadata, DL);
+ if (auto *SI = dyn_cast<StoreInst>(MainOp)) {
+ VPB.createWidenStore(*SI, Plan.getOrAddLiveIn(SI->getPointerOperand()),
+ Ops[0], /*Mask=*/nullptr, /*Consecutive=*/true,
+ Metadata, DL);
+ // Stores do not define a value.
+ return nullptr;
+ }
+ return VPB.insert(new VPWidenRecipe(*MainOp, Ops, Flags, Metadata, DL));
+}
+
+void slpvectorizer::executeSLPPlan(VPlan &Plan, VPTransformState &State) {
+ for (VPRecipeBase &R : *Plan.getEntry()) {
+ State.Builder.SetCurrentDebugLocation(R.getDebugLoc());
+ R.execute(State);
+ }
+}
diff --git a/llvm/lib/Transforms/Vectorize/SLPVPlanCodegen.h b/llvm/lib/Transforms/Vectorize/SLPVPlanCodegen.h
new file mode 100644
index 0000000000000..39807f14a7ede
--- /dev/null
+++ b/llvm/lib/Transforms/Vectorize/SLPVPlanCodegen.h
@@ -0,0 +1,51 @@
+//===- SLPVPlanCodegen.h - VPlan-based codegen for SLP ----------*- C++ -*-===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+//
+// Helpers building and executing the VPlan for an SLP tree that do not depend
+// on BoUpSLP.
+//
+//===----------------------------------------------------------------------===//
+
+#ifndef LLVM_LIB_TRANSFORMS_VECTORIZE_SLPVPLANCODEGEN_H
+#define LLVM_LIB_TRANSFORMS_VECTORIZE_SLPVPLANCODEGEN_H
+
+#include "llvm/ADT/ArrayRef.h"
+
+namespace llvm {
+class Instruction;
+class Value;
+struct VPBuilderDefaultInserter;
+template <typename InserterTy> class VPBuilderBase;
+using VPBuilder = VPBuilderBase<VPBuilderDefaultInserter>;
+class VPValue;
+class VPlan;
+struct VPTransformState;
+
+namespace slpvectorizer {
+
+/// Returns true if operand \p J of a bundle with main opcode \p Opcode is a
+/// plan live-in rather than being defined by another tree entry.
+bool isLiveInOperand(unsigned Opcode, unsigned J);
+
+/// Returns true if VPlan-based codegen can emit a recipe for a bundle with
+/// main opcode \p Opcode. Keep in sync with createRecipeForBundle().
+bool isSupportedVPlanCodegenOpcode(unsigned Opcode);
+
+/// Creates the recipe for the bundle \p Scalars with main instruction \p MainOp
+/// and vector operands \p Ops. Returns its value, or nullptr for stores.
+VPValue *createRecipeForBundle(VPlan &Plan, VPBuilder &VPB, Instruction *MainOp,
+ ArrayRef<Value *> Scalars,
+ ArrayRef<VPValue *> Ops);
+
+/// Executes the recipes of \p Plan, which are already in execution order.
+void executeSLPPlan(VPlan &Plan, VPTransformState &State);
+
+} // namespace slpvectorizer
+} // namespace llvm
+
+#endif // LLVM_LIB_TRANSFORMS_VECTORIZE_SLPVPLANCODEGEN_H
diff --git a/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp b/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
index f76a754418422..1ff91a3f5863a 100644
--- a/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+++ b/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
@@ -17,6 +17,8 @@
//===----------------------------------------------------------------------===//
#include "llvm/Transforms/Vectorize/SLPVectorizer.h"
+#include "LoopVectorizationPlanner.h"
+#include "SLPVPlanCodegen.h"
#include "SLPVectorizer/SLPCompatibilityAnalysis.h"
#include "SLPVectorizer/SLPCostAnalysis.h"
#include "SLPVectorizer/SLPMemoryUtils.h"
@@ -24,6 +26,8 @@
#include "SLPVectorizer/SLPShuffleAnalysis.h"
#include "SLPVectorizer/SLPTypeUtils.h"
#include "SLPVectorizer/SLPUtils.h"
+#include "VPlan.h"
+#include "VPlanHelpers.h"
#include "llvm/ADT/DenseMap.h"
#include "llvm/ADT/DenseSet.h"
#include "llvm/ADT/PriorityQueue.h"
@@ -337,6 +341,10 @@ static cl::opt<unsigned> SLPRuntimeAliasChecksMaxScalarCostPercent(
"guarded scalar region cost, before versioning is rejected to "
"avoid pessimizing the scalar fallback path."));
+static cl::opt<bool>
+ SLPUseVPlanCodegen("slp-use-vplan-codegen", cl::init(false), cl::Hidden,
+ cl::desc("Use VPlan-based codegen in SLP vectorizer"));
+
// Limit the number of alias checks. The limit is chosen so that
// it has no negative effect on the llvm benchmarks.
static const unsigned AliasedCheckLimit = 10;
@@ -457,6 +465,16 @@ class slpvectorizer::BoUpSLP {
Instruction *ReductionRoot = nullptr,
ArrayRef<ReductionVectorPart> VectorValuesAndScales = {});
+ /// Returns true if the current SLP tree is eligible for VPlan-based codegen.
+ /// Must be called after scheduling.
+ bool isVPlanEligible();
+
+ /// Build a VPlan for the current SLP tree.
+ std::unique_ptr<VPlan> buildVPlanForTree();
+
+ /// Execute \p Plan, generating vector IR for the SLP tree.
+ void executeVPlanForTree(VPlan &Plan);
+
/// \returns the cost incurred by unwanted spills and fills, caused by
/// holding live values over call sites.
InstructionCost getSpillCost();
@@ -25722,6 +25740,139 @@ void BoUpSLP::versionBlocksForRuntimeChecks() {
CFGChanged = true;
}
+bool BoUpSLP::isVPlanEligible() {
+ // External uses require extracts, which are not supported yet.
+ if (!ExternalUses.empty())
+ return false;
+
+ BasicBlock *RootBB =
+ cast<Instruction>(getRootNode().Scalars.front())->getParent();
+ Instruction *FirstLoad = nullptr;
+ for (const std::unique_ptr<TreeEntry> &TE : VectorizableTree) {
+ if (DeletedNodes.contains(TE.get()))
+ continue;
+
+ if (MinBWs.contains(TE.get()))
+ return false;
+
+ // Only handle plain Vectorize entries, i.e. no gathers or combined entries,
+ // without any lane permutation.
+ if (TE->State != TreeEntry::Vectorize || !TE->hasState() ||
+ TE->CombinedOp != TreeEntry::NotCombinedOp || TE->isAltShuffle() ||
+ !TE->ReorderIndices.empty() || !TE->ReuseShuffleIndices.empty())
+ return false;
+
+ unsigned Opcode = TE->getOpcode();
+ if (!isSupportedVPlanCodegenOpcode(Opcode))
+ return false;
+
+ // All recipes are emitted into the root's block.
+ if (TE->getMainOp()->getParent() != RootBB)
+ return false;
+
+ // The recipes assume scalar element types, so revec is not supported.
+ if (TE->Scalars.front()->getType()->isVectorTy())
+ return false;
+
+ // Lanes with a different opcode, e.g. copyables, need their IR flags
+ // adjusted, which is not supported.
+ if (TE->hasCopyableElements() || any_of(TE->Scalars, [&](Value *V) {
+ auto *I = dyn_cast<Instruction>(V);
+ return I && I->getOpcode() != Opcode;
+ }))
+ return false;
+
+ // A commutative sub (feeding icmp eq/ne 0 or abs) may have its operands
+ // swapped, which invalidates nuw and nsw.
+ if (Opcode == Instruction::Sub && any_of(TE->Scalars, [](Value *V) {
+ auto *I = dyn_cast<Instruction>(V);
+ return !I || isCommutative(I);
+ }))
+ return false;
+
+ // Non-power-of-2 div/rem emitted as masked intrinsics is not supported.
+ if (getMaskedDivRemCost(*TTI, SLPReVec, Opcode,
+ getValueType(TE->Scalars.front(), SLPReVec),
+ TE->Scalars.size(), TTI::TCK_RecipThroughput)
+ .isValid())
+ return false;
+
+ // All operands must be defined by entries that are still in the tree.
+ for (unsigned J : seq<unsigned>(TE->getNumOperands())) {
+ if (isLiveInOperand(Opcode, J))
+ continue;
+ TreeEntry *OpTE = OperandsToTreeEntry.lookup({TE.get(), J});
+ if (!OpTE || DeletedNodes.contains(OpTE))
+ return false;
+ }
+
+ // Bundles with extra operands, e.g. reassociated ones, are not supported.
+ if (TE->getNumOperands() != TE->getMainOp()->getNumOperands())
+ return false;
+
+ if (Opcode == Instruction::Load) {
+ Instruction *LastInst = &getLastInstructionInBundle(TE.get());
+ if (!FirstLoad || LastInst->comesBefore(FirstLoad))
+ FirstLoad = LastInst;
+ }
+ }
+
+ // Recipes are emitted at the root, so loads must not be sunk past a write.
+ if (!FirstLoad)
+ return true;
+ Instruction *RootInst = &getLastInstructionInBundle(&getRootNode());
+ return none_of(make_range(FirstLoad->getIterator(), RootInst->getIterator()),
+ [](Instruction &I) { return I.mayWriteToMemory(); });
+}
+
+std::unique_ptr<VPlan> BoUpSLP::buildVPlanForTree() {
+ // All recipes use the root's lane count as VF.
+ assert(all_of(VectorizableTree,
+ [&](const std::unique_ptr<TreeEntry> &TE) {
+ return DeletedNodes.contains(TE.get()) ||
+ TE->Scalars.size() == getRootNode().Scalars.size();
+ }) &&
+ "all entries must have the same number of lanes as the root");
+
+ auto Plan = std::make_unique<VPlan>(getRootNode().getMainOp()->getParent(),
+ Type::getInt32Ty(F->getContext()));
+ VPBuilder VPB(Plan->getEntry());
+
+ // Create the recipes in scheduled order, like the existing codegen, with
+ // operands first. Operand entries are never deleted, see isVPlanEligible().
+ DenseMap<const TreeEntry *, VPValue *> EntryToVPValue;
+ auto AddEntry = [&](TreeEntry *E, auto &Self) -> VPValue * {
+ if (auto It = EntryToVPValue.find(E); It != EntryToVPValue.end())
+ return It->second;
+ SmallVector<VPValue *> Ops;
+ for (unsigned J : seq<unsigned>(E->getNumOperands()))
+ if (!isLiveInOperand(E->getOpcode(), J))
+ Ops.push_back(Self(getOperandEntry(E, J), Self));
+ return EntryToVPValue[E] = createRecipeForBundle(*Plan, VPB, E->getMainOp(),
+ E->Scalars, Ops);
+ };
+ SmallVector<TreeEntry *> Entries;
+ for (const std::unique_ptr<TreeEntry> &TE : VectorizableTree)
+ if (!DeletedNodes.contains(TE.get()))
+ Entries.push_back(TE.get());
+ stable_sort(Entries, [&](const TreeEntry *A, const TreeEntry *B) {
+ return getLastInstructionInBundle(A).comesBefore(
+ &getLastInstructionInBundle(B));
+ });
+ for (TreeEntry *TE : Entries)
+ AddEntry(TE, AddEntry);
+
+ return Plan;
+}
+
+void BoUpSLP::executeVPlanForTree(VPlan &Plan) {
+ VPTransformState State(
+ TTI, ElementCount::getFixed(getRootNode().Scalars.size()), LI, DT, AC,
+ Builder, &Plan, /*CurrentParentLoop=*/nullptr);
+
+ executeSLPPlan(Plan, State);
+}
+
Value *BoUpSLP::vectorizeTree() {
ExtraValueToDebugLocsMap ExternallyUsedValues;
return vectorizeTree(ExternallyUsedValues);
@@ -25768,6 +25919,9 @@ BoUpSLP::vectorizeTree(const ExtraValueToDebugLocsMap &ExternallyUsedValues,
else
Builder.SetInsertPoint(&F->getEntryBlock(), F->getEntryBlock().begin());
+ // Reductions need the root's VectorizedValue, which VPlan does not set.
+ bool UsedVPlan = SLPUseVPlanCodegen && !ReductionRoot && isVPlanEligible();
+
// Vectorize gather operands of the nodes with the external uses only.
SmallVector<std::pair<TreeEntry *, Instruction *>> GatherEntries;
// Multiple gather TEs may share the same UserTE - cache the per-UserTE
@@ -25805,14 +25959,14 @@ BoUpSLP::vectorizeTree(const ExtraValueToDebugLocsMap &ExternallyUsedValues,
// the splat gathers can be emitted as their broadcasts. They go before the
// gathered loads, which skip entries that already have a vector value.
for (TreeEntry *TE : SplatGatheredScalarsRoots) {
- if (DeletedNodes.contains(TE) || TE->VectorizedValue)
+ if (UsedVPlan || DeletedNodes.contains(TE) || TE->VectorizedValue)
continue;
(void)vectorizeTree(TE);
}
// Emit gathered loads first to emit better code for the users of those
// gathered loads.
for (const std::unique_ptr<TreeEntry> &TE : VectorizableTree) {
- if (DeletedNodes.contains(TE.get()))
+ if (UsedVPlan || DeletedNodes.contains(TE.get()))
continue;
if (GatheredLoadsEntriesFirst.has_value() &&
TE->Idx >= *GatheredLoadsEntriesFirst && !TE->VectorizedValue &&
@@ -25823,7 +25977,13 @@ BoUpSLP::vectorizeTree(const ExtraValueToDebugLocsMap &ExternallyUsedValues,
(void)vectorizeTree(TE.get());
}
}
- (void)vectorizeTree(&getRootNode());
+ if (UsedVPlan) {
+ setInsertPointAfterBundle(&getRootNode());
+ std::unique_ptr<VPlan> Plan = buildVPlanForTree();
+ executeVPlanForTree(*Plan);
+ } else {
+ (void)vectorizeTree(&getRootNode());
+ }
// Run through the list of postponed gathers and emit them, replacing the temp
// emitted allocas with actual vector instructions.
ArrayRef<const TreeEntry *> PostponedNodes = PostponedGathers.getArrayRef();
@@ -26453,7 +26613,8 @@ BoUpSLP::vectorizeTree(const ExtraValueToDebugLocsMap &ExternallyUsedValues,
continue;
}
- assert(Entry->VectorizedValue && "Can't find vectorizable value");
+ assert((Entry->VectorizedValue || UsedVPlan) &&
+ "Can't find vectorizable value");
// For each lane:
for (int Lane = 0, LE = Entry->Scalars.size(); Lane != LE; ++Lane) {
@@ -26558,7 +26719,7 @@ BoUpSLP::vectorizeTree(const ExtraValueToDebugLocsMap &ExternallyUsedValues,
// Merge the DIAssignIDs from the about-to-be-deleted instructions into the
// new vector instruction.
- if (auto *V = dyn_cast<Instruction>(getRootNode().VectorizedValue))
+ if (auto *V = dyn_cast_if_present<Instruction>(getRootNode().VectorizedValue))
V->mergeDIAssignID(RemovedInsts);
// Clear up reduction references, if any.
diff --git a/llvm/lib/Transforms/Vectorize/VPlan.h b/llvm/lib/Transforms/Vectorize/VPlan.h
index a5776a8e6f635..c19c7c3170d48 100644
--- a/llvm/lib/Transforms/Vectorize/VPlan.h
+++ b/llvm/lib/Transforms/Vectorize/VPlan.h
@@ -1214,6 +1214,14 @@ class LLVM_ABI_FOR_TEST VPIRMetadata {
Metadata.emplace_back(LLVMContext::MD_prof, BW);
}
+ /// Adds the metadata that can be preserved when combining all of \p VL into
+ /// a single instruction represented by \p Repr.
+ VPIRMetadata(const Instruction &Repr, ArrayRef<Value *> VL) {
+ getMetadataToPropagate(&Repr, VL, Metadata);
+ // Drop the kinds the bundle does not agree on, which come back as null.
+ erase_if(Metadata, [](const auto &P) { return !P.second; });
+ }
+
/// Copy constructor for cloning.
VPIRMetadata(const VPIRMetadata &Other) = default;
diff --git a/llvm/test/CodeGen/WebAssembly/simd-min-vec-reg-32.ll b/llvm/test/CodeGen/WebAssembly/simd-min-vec-reg-32.ll
index ddcb81f49973a..36226e270fdb0 100644
--- a/llvm/test/CodeGen/WebAssembly/simd-min-vec-reg-32.ll
+++ b/llvm/test/CodeGen/WebAssembly/simd-min-vec-reg-32.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 6
; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=wasm32-unknown-unknown -mattr=+simd128 | llc --mtriple=wasm32-unknown-unknown -mattr=+simd128 -wasm-disable-explicit-locals -wasm-keep-registers | FileCheck %s
+; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=wasm32-unknown-unknown -mattr=+simd128 -slp-use-vplan-codegen | llc --mtriple=wasm32-unknown-unknown -mattr=+simd128 -wasm-disable-explicit-locals -wasm-keep-registers | FileCheck %s
; Test that SLP vectorizer can vectorize consecutive sub-128-bit operations
; into SIMD when getMinVectorRegisterBitWidth is overridden to 32.
diff --git a/llvm/test/Transforms/PhaseOrdering/AArch64/interleave_vec.ll b/llvm/test/Transforms/PhaseOrdering/AArch64/interleave_vec.ll
index 096581961b79a..01b12d8e0b222 100644
--- a/llvm/test/Transforms/PhaseOrdering/AArch64/interleave_vec.ll
+++ b/llvm/test/Transforms/PhaseOrdering/AArch64/interleave_vec.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 5
; RUN: opt -passes="default<O3>" -mcpu=neoverse-v2 -S < %s | FileCheck %s
+; RUN: opt -passes="default<O3>" -mcpu=neoverse-v2 -S -slp-use-vplan-codegen < %s | FileCheck %s
target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i8:8:32-i16:16:32-i64:64-i128:128-n32:64-S128-Fn32"
target triple = "aarch64"
diff --git a/llvm/test/Transforms/SLPVectorizer/AArch64/32-bit.ll b/llvm/test/Transforms/SLPVectorizer/AArch64/32-bit.ll
index bfa18f88a2467..0eb51ca824b29 100644
--- a/llvm/test/Transforms/SLPVectorizer/AArch64/32-bit.ll
+++ b/llvm/test/Transforms/SLPVectorizer/AArch64/32-bit.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
; RUN: opt -passes=slp-vectorizer -S < %s | FileCheck %s
+; RUN: opt -passes=slp-vectorizer -slp-use-vplan-codegen -S < %s | FileCheck %s
target datalayout = "e-m:e-i8:8:32-i16:16:32-i64:64-i128:128-n32:64-S128"
target triple = "aarch64-unknown-linux-gnu"
diff --git a/llvm/test/Transforms/SLPVectorizer/AArch64/masked-div-rem-non-pow2.ll b/llvm/test/Transforms/SLPVectorizer/AArch64/masked-div-rem-non-pow2.ll
index d7f61902c6753..e3fecc1ffb29b 100644
--- a/llvm/test/Transforms/SLPVectorizer/AArch64/masked-div-rem-non-pow2.ll
+++ b/llvm/test/Transforms/SLPVectorizer/AArch64/masked-div-rem-non-pow2.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
; RUN: opt -S --passes=slp-vectorizer -mtriple=aarch64-unknown-linux-gnu -mcpu=neoverse-512tvb -slp-vectorize-non-power-of-2 < %s | FileCheck %s
+; RUN: opt -S --passes=slp-vectorizer -mtriple=aarch64-unknown-linux-gnu -mcpu=neoverse-512tvb -slp-vectorize-non-power-of-2 -slp-use-vplan-codegen < %s | FileCheck %s
define void @udiv_v7i32(ptr noalias %dst, ptr noalias %x, ptr noalias %y) vscale_range(2,2) {
; CHECK-LABEL: define void @udiv_v7i32(
diff --git a/llvm/test/Transforms/SLPVectorizer/AArch64/memory-runtime-checks.ll b/llvm/test/Transforms/SLPVectorizer/AArch64/memory-runtime-checks.ll
index 44225f708b5bc..8b826904df891 100644
--- a/llvm/test/Transforms/SLPVectorizer/AArch64/memory-runtime-checks.ll
+++ b/llvm/test/Transforms/SLPVectorizer/AArch64/memory-runtime-checks.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt -aa-pipeline='basic-aa,scoped-noalias-aa' -passes=slp-vectorizer -mtriple=arm64-apple-darwin -S %s | FileCheck %s
+; RUN: opt -aa-pipeline='basic-aa,scoped-noalias-aa' -passes=slp-vectorizer -mtriple=arm64-apple-darwin -slp-use-vplan-codegen -S %s | FileCheck %s
define void @needs_versioning_not_profitable(ptr %dst, ptr %src) {
; CHECK-LABEL: @needs_versioning_not_profitable(
diff --git a/llvm/test/Transforms/SLPVectorizer/AArch64/nontemporal.ll b/llvm/test/Transforms/SLPVectorizer/AArch64/nontemporal.ll
index c53de872ebdcf..abee63bcc892d 100644
--- a/llvm/test/Transforms/SLPVectorizer/AArch64/nontemporal.ll
+++ b/llvm/test/Transforms/SLPVectorizer/AArch64/nontemporal.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt -S -passes=slp-vectorizer,dce < %s | FileCheck %s
+; RUN: opt -S -passes=slp-vectorizer,dce -slp-use-vplan-codegen < %s | FileCheck %s
target datalayout = "e-m:o-i64:64-i128:128-n32:64-S128"
target triple = "arm64-apple-ios5.0.0"
diff --git a/llvm/test/Transforms/SLPVectorizer/AArch64/slp-frem.ll b/llvm/test/Transforms/SLPVectorizer/AArch64/slp-frem.ll
index a38f4bdc4640e..87ff2916d42a3 100644
--- a/llvm/test/Transforms/SLPVectorizer/AArch64/slp-frem.ll
+++ b/llvm/test/Transforms/SLPVectorizer/AArch64/slp-frem.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 4
; RUN: opt < %s -S -mtriple=aarch64 -vector-library=ArmPL -passes=slp-vectorizer | FileCheck %s
+; RUN: opt < %s -S -mtriple=aarch64 -vector-library=ArmPL -passes=slp-vectorizer -slp-use-vplan-codegen | FileCheck %s
@a = common global ptr null, align 8
diff --git a/llvm/test/Transforms/SLPVectorizer/AArch64/spillcost-call-between-operands.ll b/llvm/test/Transforms/SLPVectorizer/AArch64/spillcost-call-between-operands.ll
index 0de6be442daa4..a47da9069dade 100644
--- a/llvm/test/Transforms/SLPVectorizer/AArch64/spillcost-call-between-operands.ll
+++ b/llvm/test/Transforms/SLPVectorizer/AArch64/spillcost-call-between-operands.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
; RUN: opt -S -passes=slp-vectorizer -mtriple=aarch64 -slp-threshold=-1 < %s | FileCheck %s
+; RUN: opt -S -passes=slp-vectorizer -mtriple=aarch64 -slp-threshold=-1 -slp-use-vplan-codegen < %s | FileCheck %s
declare void @external()
diff --git a/llvm/test/Transforms/SLPVectorizer/AArch64/spillcost-di.ll b/llvm/test/Transforms/SLPVectorizer/AArch64/spillcost-di.ll
index 76b1d18fdc0a8..0ff87415518b0 100644
--- a/llvm/test/Transforms/SLPVectorizer/AArch64/spillcost-di.ll
+++ b/llvm/test/Transforms/SLPVectorizer/AArch64/spillcost-di.ll
@@ -1,6 +1,7 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
; Debug informations shouldn't affect spill cost.
; RUN: opt -S -passes=slp-vectorizer %s -o - | FileCheck %s
+; RUN: opt -S -passes=slp-vectorizer %s -o - -slp-use-vplan-codegen | FileCheck %s
target triple = "aarch64"
diff --git a/llvm/test/Transforms/SLPVectorizer/AArch64/store-ptr.ll b/llvm/test/Transforms/SLPVectorizer/AArch64/store-ptr.ll
index 2b6a41403fb48..0284156afe53f 100644
--- a/llvm/test/Transforms/SLPVectorizer/AArch64/store-ptr.ll
+++ b/llvm/test/Transforms/SLPVectorizer/AArch64/store-ptr.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
; RUN: opt < %s -passes=slp-vectorizer -S | FileCheck %s
+; RUN: opt < %s -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s
target datalayout = "e-m:e-i8:8:32-i16:16:32-i64:64-i128:128-n32:64-S128"
target triple = "aarch64"
diff --git a/llvm/test/Transforms/SLPVectorizer/AArch64/vec3-base.ll b/llvm/test/Transforms/SLPVectorizer/AArch64/vec3-base.ll
index 21ce38cd2bd8c..c34cfba06e4b2 100644
--- a/llvm/test/Transforms/SLPVectorizer/AArch64/vec3-base.ll
+++ b/llvm/test/Transforms/SLPVectorizer/AArch64/vec3-base.ll
@@ -1,6 +1,8 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt -passes=slp-vectorizer -slp-vectorize-non-power-of-2 -mtriple=arm64-apple-ios -S %s | FileCheck --check-prefixes=CHECK,NON-POW2 %s
+; RUN: opt -passes=slp-vectorizer -slp-vectorize-non-power-of-2 -mtriple=arm64-apple-ios -slp-use-vplan-codegen -S %s | FileCheck --check-prefixes=CHECK,NON-POW2 %s
; RUN: opt -passes=slp-vectorizer -slp-vectorize-non-power-of-2=false -mtriple=arm64-apple-ios -S %s | FileCheck --check-prefixes=CHECK,POW2-ONLY %s
+; RUN: opt -passes=slp-vectorizer -slp-vectorize-non-power-of-2=false -mtriple=arm64-apple-ios -slp-use-vplan-codegen -S %s | FileCheck --check-prefixes=CHECK,POW2-ONLY %s
define void @v3_load_i32_mul_by_constant_store(ptr %src, ptr %dst) {
; NON-POW2-LABEL: @v3_load_i32_mul_by_constant_store(
diff --git a/llvm/test/Transforms/SLPVectorizer/AArch64/vplan-codegen-load-sinking.ll b/llvm/test/Transforms/SLPVectorizer/AArch64/vplan-codegen-load-sinking.ll
new file mode 100644
index 0000000000000..5517b5f36ac1c
--- /dev/null
+++ b/llvm/test/Transforms/SLPVectorizer/AArch64/vplan-codegen-load-sinking.ll
@@ -0,0 +1,133 @@
+; RUN: opt -S -passes=slp-vectorizer -mtriple=aarch64 -slp-threshold=-1 -slp-use-vplan-codegen < %s | FileCheck %s
+
+declare void @may_write()
+declare void @read_only() memory(read)
+
+; The loads must not be sunk past a call that may write memory.
+define void @call_between_loads_and_root(ptr %p, ptr %q, ptr %r) {
+; CHECK-LABEL: @call_between_loads_and_root(
+; CHECK-NEXT: entry:
+; CHECK-NEXT: [[TMP0:%.*]] = load <2 x double>, ptr [[P:%.*]], align 8
+; CHECK-NEXT: [[TMP1:%.*]] = load <2 x double>, ptr [[Q:%.*]], align 8
+; CHECK-NEXT: [[TMP2:%.*]] = fadd <2 x double> [[TMP0]], [[TMP1]]
+; CHECK-NEXT: call void @may_write()
+; CHECK-NEXT: store <2 x double> [[TMP2]], ptr [[R:%.*]], align 8
+; CHECK-NEXT: ret void
+;
+entry:
+ %a0 = load double, ptr %p, align 8
+ %p1 = getelementptr inbounds double, ptr %p, i64 1
+ %a1 = load double, ptr %p1, align 8
+ %b0 = load double, ptr %q, align 8
+ %q1 = getelementptr inbounds double, ptr %q, i64 1
+ %b1 = load double, ptr %q1, align 8
+ %s0 = fadd double %a0, %b0
+ %s1 = fadd double %a1, %b1
+ call void @may_write()
+ store double %s0, ptr %r, align 8
+ %r1 = getelementptr inbounds double, ptr %r, i64 1
+ store double %s1, ptr %r1, align 8
+ ret void
+}
+
+; Same for a store that may alias the loads.
+define void @store_between_loads_and_root(ptr %p, ptr %q, ptr %r, ptr %x) {
+; CHECK-LABEL: @store_between_loads_and_root(
+; CHECK-NEXT: entry:
+; CHECK-NEXT: [[TMP0:%.*]] = load <2 x double>, ptr [[P:%.*]], align 8
+; CHECK-NEXT: [[TMP1:%.*]] = load <2 x double>, ptr [[Q:%.*]], align 8
+; CHECK-NEXT: [[TMP2:%.*]] = fadd <2 x double> [[TMP0]], [[TMP1]]
+; CHECK-NEXT: store i8 0, ptr [[X:%.*]], align 1
+; CHECK-NEXT: store <2 x double> [[TMP2]], ptr [[R:%.*]], align 8
+; CHECK-NEXT: ret void
+;
+entry:
+ %a0 = load double, ptr %p, align 8
+ %p1 = getelementptr inbounds double, ptr %p, i64 1
+ %a1 = load double, ptr %p1, align 8
+ %b0 = load double, ptr %q, align 8
+ %q1 = getelementptr inbounds double, ptr %q, i64 1
+ %b1 = load double, ptr %q1, align 8
+ %s0 = fadd double %a0, %b0
+ %s1 = fadd double %a1, %b1
+ store i8 0, ptr %x, align 1
+ store double %s0, ptr %r, align 8
+ %r1 = getelementptr inbounds double, ptr %r, i64 1
+ store double %s1, ptr %r1, align 8
+ ret void
+}
+
+; A read-only call can be crossed.
+define void @read_only_call_between_loads_and_root(ptr %p, ptr %q, ptr %r) {
+; CHECK-LABEL: @read_only_call_between_loads_and_root(
+; CHECK-NEXT: entry:
+; CHECK-NEXT: call void @read_only()
+; CHECK-NEXT: %wide.load = load <2 x double>, ptr %p, align 8
+; CHECK-NEXT: %wide.load1 = load <2 x double>, ptr %q, align 8
+; CHECK-NEXT: [[TMP0:%.*]] = fadd <2 x double> %wide.load, %wide.load1
+; CHECK-NEXT: store <2 x double> [[TMP0]], ptr %r, align 8
+; CHECK-NEXT: ret void
+;
+entry:
+ %a0 = load double, ptr %p, align 8
+ %p1 = getelementptr inbounds double, ptr %p, i64 1
+ %a1 = load double, ptr %p1, align 8
+ %b0 = load double, ptr %q, align 8
+ %q1 = getelementptr inbounds double, ptr %q, i64 1
+ %b1 = load double, ptr %q1, align 8
+ %s0 = fadd double %a0, %b0
+ %s1 = fadd double %a1, %b1
+ call void @read_only()
+ store double %s0, ptr %r, align 8
+ %r1 = getelementptr inbounds double, ptr %r, i64 1
+ store double %s1, ptr %r1, align 8
+ ret void
+}
+
+; A call after the root is not crossed.
+define void @call_after_root(ptr %p, ptr %q, ptr %r) {
+; CHECK-LABEL: @call_after_root(
+; CHECK-NEXT: entry:
+; CHECK-NEXT: %wide.load = load <2 x double>, ptr %p, align 8
+; CHECK-NEXT: %wide.load1 = load <2 x double>, ptr %q, align 8
+; CHECK-NEXT: [[TMP0:%.*]] = fadd <2 x double> %wide.load, %wide.load1
+; CHECK-NEXT: store <2 x double> [[TMP0]], ptr %r, align 8
+; CHECK-NEXT: call void @may_write()
+; CHECK-NEXT: ret void
+;
+entry:
+ %a0 = load double, ptr %p, align 8
+ %p1 = getelementptr inbounds double, ptr %p, i64 1
+ %a1 = load double, ptr %p1, align 8
+ %b0 = load double, ptr %q, align 8
+ %q1 = getelementptr inbounds double, ptr %q, i64 1
+ %b1 = load double, ptr %q1, align 8
+ %s0 = fadd double %a0, %b0
+ %s1 = fadd double %a1, %b1
+ store double %s0, ptr %r, align 8
+ %r1 = getelementptr inbounds double, ptr %r, i64 1
+ store double %s1, ptr %r1, align 8
+ call void @may_write()
+ ret void
+}
+
+; Scheduling moves the interleaved root stores after the loads.
+define void @interleaved_root_stores(ptr noalias %p, ptr noalias %r) {
+; CHECK-LABEL: @interleaved_root_stores(
+; CHECK-NEXT: entry:
+; CHECK-NEXT: %wide.load = load <2 x double>, ptr %p, align 8
+; CHECK-NEXT: [[TMP0:%.*]] = fneg <2 x double> %wide.load
+; CHECK-NEXT: store <2 x double> [[TMP0]], ptr %r, align 8
+; CHECK-NEXT: ret void
+;
+entry:
+ %a0 = load double, ptr %p, align 8
+ %n0 = fneg double %a0
+ store double %n0, ptr %r, align 8
+ %p1 = getelementptr inbounds double, ptr %p, i64 1
+ %a1 = load double, ptr %p1, align 8
+ %n1 = fneg double %a1
+ %r1 = getelementptr inbounds double, ptr %r, i64 1
+ store double %n1, ptr %r1, align 8
+ ret void
+}
diff --git a/llvm/test/Transforms/SLPVectorizer/AMDGPU/elementwise-fma-operand1.ll b/llvm/test/Transforms/SLPVectorizer/AMDGPU/elementwise-fma-operand1.ll
index 05e5e5f29d51b..b089e8ef3e50c 100644
--- a/llvm/test/Transforms/SLPVectorizer/AMDGPU/elementwise-fma-operand1.ll
+++ b/llvm/test/Transforms/SLPVectorizer/AMDGPU/elementwise-fma-operand1.ll
@@ -3,8 +3,11 @@
; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgcn-amd-amdhsa -mcpu=gfx942 -slp-threshold=14 < %s | FileCheck %s
; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgcn-amd-amdhsa -mcpu=gfx950 -slp-threshold=14 < %s | FileCheck %s
; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgcn-amd-amdhsa -mcpu=gfx90a -slp-threshold=12 < %s | FileCheck %s --check-prefix=THR12
+; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgcn-amd-amdhsa -mcpu=gfx90a -slp-threshold=12 -slp-use-vplan-codegen < %s | FileCheck %s --check-prefix=THR12
; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgcn-amd-amdhsa -mcpu=gfx942 -slp-threshold=12 < %s | FileCheck %s --check-prefix=THR12
+; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgcn-amd-amdhsa -mcpu=gfx942 -slp-threshold=12 -slp-use-vplan-codegen < %s | FileCheck %s --check-prefix=THR12
; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgcn-amd-amdhsa -mcpu=gfx950 -slp-threshold=12 < %s | FileCheck %s --check-prefix=THR12
+; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgcn-amd-amdhsa -mcpu=gfx950 -slp-threshold=12 -slp-use-vplan-codegen < %s | FileCheck %s --check-prefix=THR12
; Elementwise d = c + a * b, where the fmul is operand 1 of the fadd. These
; targets halve the cost of a packed fmul, so SLP is tempted to vectorize and
diff --git a/llvm/test/Transforms/SLPVectorizer/AMDGPU/fma-operand-contract-selection.ll b/llvm/test/Transforms/SLPVectorizer/AMDGPU/fma-operand-contract-selection.ll
index 9b37c518981e3..6d9a2d5770609 100644
--- a/llvm/test/Transforms/SLPVectorizer/AMDGPU/fma-operand-contract-selection.ll
+++ b/llvm/test/Transforms/SLPVectorizer/AMDGPU/fma-operand-contract-selection.ll
@@ -1,6 +1,8 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgcn-amd-amdhsa -mcpu=gfx90a < %s | FileCheck %s
+; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgcn-amd-amdhsa -mcpu=gfx90a -slp-use-vplan-codegen < %s | FileCheck %s
; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgcn-amd-amdhsa -mcpu=gfx90a -slp-threshold=15 < %s | FileCheck %s --check-prefix=THR15
+; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgcn-amd-amdhsa -mcpu=gfx90a -slp-threshold=15 -slp-use-vplan-codegen < %s | FileCheck %s --check-prefix=THR15
; The fma check has to read the multiply's own fast math flags. Reading the
; add's instead makes every multiply look contractable, so a fusion that the
diff --git a/llvm/test/Transforms/SLPVectorizer/AMDGPU/invariant-load-no-alias-store.ll b/llvm/test/Transforms/SLPVectorizer/AMDGPU/invariant-load-no-alias-store.ll
index 3781a57d4ff83..3a7dec583e568 100644
--- a/llvm/test/Transforms/SLPVectorizer/AMDGPU/invariant-load-no-alias-store.ll
+++ b/llvm/test/Transforms/SLPVectorizer/AMDGPU/invariant-load-no-alias-store.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
; RUN: opt -passes="function(slp-vectorizer)" -mtriple=amdgpu12.00-amd-amdhsa %s -S | FileCheck %s
+; RUN: opt -passes="function(slp-vectorizer)" -mtriple=amdgpu12.00-amd-amdhsa -slp-use-vplan-codegen %s -S | FileCheck %s
define void @test(ptr addrspace(1) %base, ptr addrspace(1) %otherA, ptr addrspace(1) %otherB) #0 {
; CHECK-LABEL: define void @test(
diff --git a/llvm/test/Transforms/SLPVectorizer/AMDGPU/packed-math.ll b/llvm/test/Transforms/SLPVectorizer/AMDGPU/packed-math.ll
index 49d7ee569c871..a2e8fa269af44 100644
--- a/llvm/test/Transforms/SLPVectorizer/AMDGPU/packed-math.ll
+++ b/llvm/test/Transforms/SLPVectorizer/AMDGPU/packed-math.ll
@@ -1,6 +1,8 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt -S -mtriple=amdgpu9.00-amd-amdhsa -passes=slp-vectorizer,dce < %s | FileCheck -check-prefixes=GCN,GFX9 %s
+; RUN: opt -S -mtriple=amdgpu9.00-amd-amdhsa -passes=slp-vectorizer,dce -slp-use-vplan-codegen < %s | FileCheck -check-prefixes=GCN,GFX9 %s
; RUN: opt -S -mtriple=amdgpu8.03-amd-amdhsa -passes=slp-vectorizer,dce < %s | FileCheck -check-prefixes=GCN,VI %s
+; RUN: opt -S -mtriple=amdgpu8.03-amd-amdhsa -passes=slp-vectorizer,dce -slp-use-vplan-codegen < %s | FileCheck -check-prefixes=GCN,VI %s
; FIXME: Should still like to vectorize the memory operations for VI
diff --git a/llvm/test/Transforms/SLPVectorizer/AMDGPU/slp-v2f16.ll b/llvm/test/Transforms/SLPVectorizer/AMDGPU/slp-v2f16.ll
index 31f5314bba8c7..2431d3538ca9d 100644
--- a/llvm/test/Transforms/SLPVectorizer/AMDGPU/slp-v2f16.ll
+++ b/llvm/test/Transforms/SLPVectorizer/AMDGPU/slp-v2f16.ll
@@ -1,8 +1,12 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --prefix-filecheck-ir-name I --version 6
; RUN: opt -S -mtriple=amdgpu8.03-amd-amdhsa -passes=slp-vectorizer < %s | FileCheck -check-prefixes=GCN,GFX8 %s
+; RUN: opt -S -mtriple=amdgpu8.03-amd-amdhsa -passes=slp-vectorizer -slp-use-vplan-codegen < %s | FileCheck -check-prefixes=GCN,GFX8 %s
; RUN: opt -S -mtriple=amdgpu9.08-amd-amdhsa -passes=slp-vectorizer < %s | FileCheck -check-prefixes=GCN,GFX9 %s
+; RUN: opt -S -mtriple=amdgpu9.08-amd-amdhsa -passes=slp-vectorizer -slp-use-vplan-codegen < %s | FileCheck -check-prefixes=GCN,GFX9 %s
; RUN: opt -S -mtriple=amdgpu9.0a-amd-amdhsa -passes=slp-vectorizer < %s | FileCheck -check-prefixes=GCN,GFX90A %s
+; RUN: opt -S -mtriple=amdgpu9.0a-amd-amdhsa -passes=slp-vectorizer -slp-use-vplan-codegen < %s | FileCheck -check-prefixes=GCN,GFX90A %s
; RUN: opt -S -mtriple=amdgpu10.30-amd-amdhsa -passes=slp-vectorizer < %s | FileCheck -check-prefixes=GCN,GFX9 %s
+; RUN: opt -S -mtriple=amdgpu10.30-amd-amdhsa -passes=slp-vectorizer -slp-use-vplan-codegen < %s | FileCheck -check-prefixes=GCN,GFX9 %s
; FIXME: Should not vectorize on gfx8
diff --git a/llvm/test/Transforms/SLPVectorizer/RISCV/basic-strided-loads.ll b/llvm/test/Transforms/SLPVectorizer/RISCV/basic-strided-loads.ll
index b2729cba17a01..3fd00b097123c 100644
--- a/llvm/test/Transforms/SLPVectorizer/RISCV/basic-strided-loads.ll
+++ b/llvm/test/Transforms/SLPVectorizer/RISCV/basic-strided-loads.ll
@@ -1,6 +1,7 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 5
; RUN: opt -mtriple=riscv64 -mattr=+m,+v,+unaligned-vector-mem -passes=slp-vectorizer -S < %s | FileCheck %s
+; RUN: opt -mtriple=riscv64 -mattr=+m,+v,+unaligned-vector-mem -passes=slp-vectorizer -slp-use-vplan-codegen -S < %s | FileCheck %s
define void @const_stride_1_no_reordering(ptr %pl, ptr %ps) {
; CHECK-LABEL: define void @const_stride_1_no_reordering(
diff --git a/llvm/test/Transforms/SLPVectorizer/RISCV/basic-strided-stores.ll b/llvm/test/Transforms/SLPVectorizer/RISCV/basic-strided-stores.ll
index a738dd27bcaf7..a5be71a221d1f 100644
--- a/llvm/test/Transforms/SLPVectorizer/RISCV/basic-strided-stores.ll
+++ b/llvm/test/Transforms/SLPVectorizer/RISCV/basic-strided-stores.ll
@@ -1,6 +1,7 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 5
; RUN: opt -mtriple=riscv64 -mattr=+m,+v -passes=slp-vectorizer -S -slp-enable-strided-stores < %s | FileCheck %s
+; RUN: opt -mtriple=riscv64 -mattr=+m,+v -passes=slp-vectorizer -S -slp-enable-strided-stores -slp-use-vplan-codegen < %s | FileCheck %s
define void @constant_stride_2(ptr %pl, ptr %ps) {
; CHECK-LABEL: define void @constant_stride_2(
diff --git a/llvm/test/Transforms/SLPVectorizer/RISCV/external.ll b/llvm/test/Transforms/SLPVectorizer/RISCV/external.ll
index 1863de244204d..b3b64437c9e0a 100644
--- a/llvm/test/Transforms/SLPVectorizer/RISCV/external.ll
+++ b/llvm/test/Transforms/SLPVectorizer/RISCV/external.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
; RUN: opt < %s -passes=slp-vectorizer -mtriple=riscv64 -mattr=+v -S | FileCheck %s --check-prefixes=DEFAULT
+; RUN: opt < %s -passes=slp-vectorizer -mtriple=riscv64 -mattr=+v -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=DEFAULT
define void @external(ptr %dest, ptr %dest2, ptr %src) {
; DEFAULT-LABEL: define void @external(
diff --git a/llvm/test/Transforms/SLPVectorizer/RISCV/floating-point.ll b/llvm/test/Transforms/SLPVectorizer/RISCV/floating-point.ll
index b5c8bf81a1921..85b2adb192f66 100644
--- a/llvm/test/Transforms/SLPVectorizer/RISCV/floating-point.ll
+++ b/llvm/test/Transforms/SLPVectorizer/RISCV/floating-point.ll
@@ -3,7 +3,12 @@
; RUN: -riscv-v-vector-bits-min=-1 -riscv-v-slp-max-vf=0 \
; RUN: | FileCheck %s
; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=riscv64 -mattr=+v,+f \
+; RUN: -riscv-v-vector-bits-min=-1 -riscv-v-slp-max-vf=0 \
+; RUN: -slp-use-vplan-codegen | FileCheck %s
+; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=riscv64 -mattr=+v,+f \
; RUN: | FileCheck %s
+; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=riscv64 -mattr=+v,+f \
+; RUN: -slp-use-vplan-codegen | FileCheck %s
define void @fp_add(ptr %dst, ptr %p, ptr %q) {
; CHECK-LABEL: define void @fp_add
diff --git a/llvm/test/Transforms/SLPVectorizer/RISCV/gep.ll b/llvm/test/Transforms/SLPVectorizer/RISCV/gep.ll
index 39d8676e1917f..af92d2c78904f 100644
--- a/llvm/test/Transforms/SLPVectorizer/RISCV/gep.ll
+++ b/llvm/test/Transforms/SLPVectorizer/RISCV/gep.ll
@@ -1,6 +1,8 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -passes=slp-vectorizer -mtriple=riscv64 -mattr=+v \
; RUN: -riscv-v-slp-max-vf=0 -S | FileCheck %s
+; RUN: opt < %s -passes=slp-vectorizer -mtriple=riscv64 -mattr=+v \
+; RUN: -riscv-v-slp-max-vf=0 -S -slp-use-vplan-codegen | FileCheck %s
; This should not be vectorized, as the cost of computing the offsets nullifies
; the benefits of vectorizing:
diff --git a/llvm/test/Transforms/SLPVectorizer/RISCV/load-binop-store.ll b/llvm/test/Transforms/SLPVectorizer/RISCV/load-binop-store.ll
index b285ef3da6735..26dd4bd558d86 100644
--- a/llvm/test/Transforms/SLPVectorizer/RISCV/load-binop-store.ll
+++ b/llvm/test/Transforms/SLPVectorizer/RISCV/load-binop-store.ll
@@ -1,7 +1,10 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -passes=slp-vectorizer -mtriple=riscv64 -mattr=+m,+v \
; RUN: -riscv-v-vector-bits-min=-1 -riscv-v-slp-max-vf=0 -S | FileCheck %s
+; RUN: opt < %s -passes=slp-vectorizer -mtriple=riscv64 -mattr=+m,+v \
+; RUN: -riscv-v-vector-bits-min=-1 -riscv-v-slp-max-vf=0 -S -slp-use-vplan-codegen | FileCheck %s
; RUN: opt < %s -passes=slp-vectorizer -mtriple=riscv64 -mattr=+m,+v -S | FileCheck %s
+; RUN: opt < %s -passes=slp-vectorizer -mtriple=riscv64 -mattr=+m,+v -S -slp-use-vplan-codegen | FileCheck %s
define void @vec_add(ptr %dest, ptr %p) {
; CHECK-LABEL: @vec_add(
diff --git a/llvm/test/Transforms/SLPVectorizer/RISCV/load-store.ll b/llvm/test/Transforms/SLPVectorizer/RISCV/load-store.ll
index 604695cb05b3a..ff244870e006d 100644
--- a/llvm/test/Transforms/SLPVectorizer/RISCV/load-store.ll
+++ b/llvm/test/Transforms/SLPVectorizer/RISCV/load-store.ll
@@ -1,7 +1,10 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -passes=slp-vectorizer -mtriple=riscv64 -mattr=+v \
; RUN: -riscv-v-vector-bits-min=-1 -riscv-v-slp-max-vf=0 -S | FileCheck %s
+; RUN: opt < %s -passes=slp-vectorizer -mtriple=riscv64 -mattr=+v \
+; RUN: -riscv-v-vector-bits-min=-1 -riscv-v-slp-max-vf=0 -S -slp-use-vplan-codegen | FileCheck %s
; RUN: opt < %s -passes=slp-vectorizer -mtriple=riscv64 -mattr=+v -S | FileCheck %s
+; RUN: opt < %s -passes=slp-vectorizer -mtriple=riscv64 -mattr=+v -S -slp-use-vplan-codegen | FileCheck %s
define void @simple_copy(ptr %dest, ptr %p) {
diff --git a/llvm/test/Transforms/SLPVectorizer/RISCV/rotated-strided-loads.ll b/llvm/test/Transforms/SLPVectorizer/RISCV/rotated-strided-loads.ll
index dcd3d830428de..5de9a2dffef42 100644
--- a/llvm/test/Transforms/SLPVectorizer/RISCV/rotated-strided-loads.ll
+++ b/llvm/test/Transforms/SLPVectorizer/RISCV/rotated-strided-loads.ll
@@ -1,6 +1,7 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 5
; RUN: opt -mtriple=riscv64 -mattr=+m,+v,+unaligned-vector-mem -passes=slp-vectorizer -S < %s | FileCheck %s
+; RUN: opt -mtriple=riscv64 -mattr=+m,+v,+unaligned-vector-mem -passes=slp-vectorizer -S -slp-use-vplan-codegen < %s | FileCheck %s
define void @constant_stride_widen_rotated1(ptr %pl, i64 %stride, ptr %ps) {
; %gep_l0 = getelementptr inbounds i8, ptr %pl, i64 0
diff --git a/llvm/test/Transforms/SLPVectorizer/RISCV/runtime-strided-stores.ll b/llvm/test/Transforms/SLPVectorizer/RISCV/runtime-strided-stores.ll
index 9c759761b14f8..2785f3eb56ac4 100644
--- a/llvm/test/Transforms/SLPVectorizer/RISCV/runtime-strided-stores.ll
+++ b/llvm/test/Transforms/SLPVectorizer/RISCV/runtime-strided-stores.ll
@@ -1,6 +1,7 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 5
; RUN: opt -mtriple=riscv64 -mattr=+m,+v -passes=slp-vectorizer -S -slp-enable-strided-stores -slp-threshold=-100 -slp-revec < %s | FileCheck %s
+; RUN: opt -mtriple=riscv64 -mattr=+m,+v -passes=slp-vectorizer -S -slp-enable-strided-stores -slp-threshold=-100 -slp-revec -slp-use-vplan-codegen < %s | FileCheck %s
define void @runtime_stride(ptr %pl, ptr %ps, i64 %stride) {
; CHECK-LABEL: define void @runtime_stride(
diff --git a/llvm/test/Transforms/SLPVectorizer/RISCV/test-delete-tree.ll b/llvm/test/Transforms/SLPVectorizer/RISCV/test-delete-tree.ll
index c4e6c4e5d5db5..5154f4c46979e 100644
--- a/llvm/test/Transforms/SLPVectorizer/RISCV/test-delete-tree.ll
+++ b/llvm/test/Transforms/SLPVectorizer/RISCV/test-delete-tree.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
; RUN: opt -mtriple=riscv64 -mattr=+m,+v -passes=slp-vectorizer -S < %s | FileCheck %s
+; RUN: opt -mtriple=riscv64 -mattr=+m,+v -passes=slp-vectorizer -slp-use-vplan-codegen -S < %s | FileCheck %s
; CHECK-NOT: TreeEntryToStridedPtrInfoMap is not cleared
define void @const_stride_1_no_reordering(ptr %pl, ptr %ps) {
diff --git a/llvm/test/Transforms/SLPVectorizer/RISCV/vec3-base.ll b/llvm/test/Transforms/SLPVectorizer/RISCV/vec3-base.ll
index d9a2a495cd55d..5352c647994b3 100644
--- a/llvm/test/Transforms/SLPVectorizer/RISCV/vec3-base.ll
+++ b/llvm/test/Transforms/SLPVectorizer/RISCV/vec3-base.ll
@@ -1,6 +1,8 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt -passes=slp-vectorizer -slp-vectorize-non-power-of-2 -mtriple=riscv64 -mattr=+m,+v -S %s | FileCheck --check-prefixes=CHECK,NON-POW2 %s
+; RUN: opt -passes=slp-vectorizer -slp-vectorize-non-power-of-2 -mtriple=riscv64 -mattr=+m,+v -slp-use-vplan-codegen -S %s | FileCheck --check-prefixes=CHECK,NON-POW2 %s
; RUN: opt -passes=slp-vectorizer -slp-vectorize-non-power-of-2=false -mtriple=riscv64 -mattr=+m,+v -S %s | FileCheck --check-prefixes=CHECK,POW2-ONLY %s
+; RUN: opt -passes=slp-vectorizer -slp-vectorize-non-power-of-2=false -mtriple=riscv64 -mattr=+m,+v -slp-use-vplan-codegen -S %s | FileCheck --check-prefixes=CHECK,POW2-ONLY %s
define void @v3_load_i32_mul_by_constant_store(ptr %src, ptr %dst) {
; NON-POW2-LABEL: @v3_load_i32_mul_by_constant_store(
diff --git a/llvm/test/Transforms/SLPVectorizer/VE/disable_slp.ll b/llvm/test/Transforms/SLPVectorizer/VE/disable_slp.ll
index c1868c54f268d..3d9ee16cff588 100644
--- a/llvm/test/Transforms/SLPVectorizer/VE/disable_slp.ll
+++ b/llvm/test/Transforms/SLPVectorizer/VE/disable_slp.ll
@@ -1,5 +1,6 @@
; RUN: opt < %s -passes=slp-vectorizer -mtriple=ve-linux -S | FileCheck %s -check-prefix=VE
; RUN: opt < %s -passes=slp-vectorizer -mtriple=x86_64-pc_linux -mcpu=core-avx2 -S | FileCheck %s -check-prefix=SSE
+; RUN: opt < %s -passes=slp-vectorizer -mtriple=x86_64-pc_linux -mcpu=core-avx2 -S -slp-use-vplan-codegen | FileCheck %s -check-prefix=SSE
; Make sure SLP does not trigger for VE on an appealing set of combinable loads
; and stores that vectorizes for x86 SSE.
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/align.ll b/llvm/test/Transforms/SLPVectorizer/X86/align.ll
index 58bcc3f372cfb..156a0741c469a 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/align.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/align.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx | FileCheck %s
+; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx -slp-use-vplan-codegen | FileCheck %s
target datalayout = "e-p:64:64:64-i1:8:8-i8:8:8-i16:16:16-i32:32:32-i64:64:64-f32:32:32-f64:64:64-v64:64:64-v128:128:128-a0:0:64-s0:64:64-f80:128:128-n8:16:32:64-S128"
target triple = "x86_64-apple-macosx10.8.0"
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/arith-add-load.ll b/llvm/test/Transforms/SLPVectorizer/X86/arith-add-load.ll
index dd0471b9f0fcf..e528f26d5247c 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/arith-add-load.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/arith-add-load.ll
@@ -1,8 +1,12 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64 -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=CHECK,SSE
+; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64 -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=CHECK,SSE
; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64-v2 -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=CHECK,SSE
+; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64-v2 -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=CHECK,SSE
; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64-v3 -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=CHECK,AVX
+; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64-v3 -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=CHECK,AVX
; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64-v4 -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=CHECK,AVX
+; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64-v4 -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=CHECK,AVX
; // PR47491
; void pr(char* r, char* a){
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/arith-add.ll b/llvm/test/Transforms/SLPVectorizer/X86/arith-add.ll
index 61e08822004d4..8275abed13231 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/arith-add.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/arith-add.ll
@@ -1,17 +1,30 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -mtriple=x86_64-unknown -passes=slp-vectorizer -S | FileCheck %s --check-prefix=SSE
+; RUN: opt < %s -mtriple=x86_64-unknown -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=SSE
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=slm -passes=slp-vectorizer -S | FileCheck %s --check-prefix=SLM
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=slm -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=SLM
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -mattr=-prefer-128-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -mattr=-prefer-128-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -mattr=+prefer-128-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=SSE
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -mattr=+prefer-128-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=SSE
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -mattr=-prefer-128-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -mattr=-prefer-128-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -mattr=+prefer-128-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=SSE
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -mattr=+prefer-128-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=SSE
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=knl -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=knl -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX
@a64 = common global [8 x i64] zeroinitializer, align 64
@b64 = common global [8 x i64] zeroinitializer, align 64
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/arith-mul-load.ll b/llvm/test/Transforms/SLPVectorizer/X86/arith-mul-load.ll
index c24bc54753476..836a9fd4b9fee 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/arith-mul-load.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/arith-mul-load.ll
@@ -1,8 +1,12 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64 -S | FileCheck %s --check-prefixes=CHECK,SSE
+; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64 -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=CHECK,SSE
; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64-v2 -S | FileCheck %s --check-prefixes=CHECK,SSE
+; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64-v2 -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=CHECK,SSE
; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64-v3 -S | FileCheck %s --check-prefixes=CHECK,AVX
+; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64-v3 -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=CHECK,AVX
; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64-v4 -S | FileCheck %s --check-prefixes=CHECK,AVX
+; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-unknown -mcpu=x86-64-v4 -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=CHECK,AVX
; // PR47491
; void pr(char* r, char* a){
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/arith-mul.ll b/llvm/test/Transforms/SLPVectorizer/X86/arith-mul.ll
index 2d846d30d819d..7aa71d39ef0e2 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/arith-mul.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/arith-mul.ll
@@ -1,17 +1,30 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -mtriple=x86_64-unknown -passes=slp-vectorizer -S | FileCheck %s --check-prefix=SSE
+; RUN: opt < %s -mtriple=x86_64-unknown -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=SSE
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=slm -passes=slp-vectorizer -S | FileCheck %s --check-prefix=SLM
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=slm -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=SLM
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -mattr=+prefer-128-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX128
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -mattr=+prefer-128-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX128
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -mattr=-prefer-128-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX256
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -mattr=-prefer-128-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX256
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -mattr=+prefer-128-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX128
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -mattr=+prefer-128-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX128
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -mattr=-prefer-128-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX256
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -mattr=-prefer-128-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX256
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=knl -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=knl -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX256
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX256
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX256
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX256
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX256
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX256
@a64 = common global [8 x i64] zeroinitializer, align 64
@b64 = common global [8 x i64] zeroinitializer, align 64
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/arith-sub.ll b/llvm/test/Transforms/SLPVectorizer/X86/arith-sub.ll
index 019cd3ac965fe..c75fec6c03a5f 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/arith-sub.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/arith-sub.ll
@@ -1,17 +1,30 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -mtriple=x86_64-unknown -passes=slp-vectorizer -S | FileCheck %s --check-prefix=SSE
+; RUN: opt < %s -mtriple=x86_64-unknown -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=SSE
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=slm -passes=slp-vectorizer -S | FileCheck %s --check-prefix=SLM
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=slm -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=SLM
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -mattr=-prefer-128-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -mattr=-prefer-128-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -mattr=+prefer-128-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=SSE
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -mattr=+prefer-128-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=SSE
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -mattr=-prefer-128-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -mattr=-prefer-128-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -mattr=+prefer-128-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=SSE
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -mattr=+prefer-128-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=SSE
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=knl -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=knl -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX
@a64 = common global [8 x i64] zeroinitializer, align 64
@b64 = common global [8 x i64] zeroinitializer, align 64
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/continue_vectorizing.ll b/llvm/test/Transforms/SLPVectorizer/X86/continue_vectorizing.ll
index fca26cbd1c459..7bc233a37ce9a 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/continue_vectorizing.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/continue_vectorizing.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx | FileCheck %s
+; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx -slp-use-vplan-codegen | FileCheck %s
target datalayout = "e-p:64:64:64-i1:8:8-i8:8:8-i16:16:16-i32:32:32-i64:64:64-f32:32:32-f64:64:64-v64:64:64-v128:128:128-a0:0:64-s0:64:64-f80:128:128-n8:16:32:64-S128"
target triple = "x86_64-apple-macosx10.8.0"
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/control-dependence.ll b/llvm/test/Transforms/SLPVectorizer/X86/control-dependence.ll
index 27817ea7f3b31..4a04f458163a0 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/control-dependence.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/control-dependence.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt -passes=slp-vectorizer -slp-threshold=-999 -S -mtriple=x86_64-unknown-linux-gnu -mcpu=skylake < %s | FileCheck %s
+; RUN: opt -passes=slp-vectorizer -slp-threshold=-999 -S -mtriple=x86_64-unknown-linux-gnu -mcpu=skylake -slp-use-vplan-codegen < %s | FileCheck %s
declare i64 @may_inf_loop_ro() nounwind readonly
declare i64 @may_inf_loop_rw() nounwind
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/fmuladd-copyable-add-part.ll b/llvm/test/Transforms/SLPVectorizer/X86/fmuladd-copyable-add-part.ll
index 8feb518f4e823..0ecd96b2eaa3f 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/fmuladd-copyable-add-part.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/fmuladd-copyable-add-part.ll
@@ -1,6 +1,8 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
; RUN: opt -passes=slp-vectorizer -slp-threshold=-99999 -S -mtriple=x86_64-unknown-linux-gnu < %s | FileCheck %s --check-prefixes=CHECK,ENABLED
+; RUN: opt -passes=slp-vectorizer -slp-threshold=-99999 -S -mtriple=x86_64-unknown-linux-gnu -slp-use-vplan-codegen < %s | FileCheck %s --check-prefixes=CHECK,ENABLED
; RUN: opt -passes=slp-vectorizer -slp-threshold=-99999 -slp-copyable-elements=false -S -mtriple=x86_64-unknown-linux-gnu < %s | FileCheck %s --check-prefixes=CHECK,DISABLED
+; RUN: opt -passes=slp-vectorizer -slp-threshold=-99999 -slp-copyable-elements=false -S -mtriple=x86_64-unknown-linux-gnu -slp-use-vplan-codegen < %s | FileCheck %s --check-prefixes=CHECK,DISABLED
declare double @llvm.fmuladd.f64(double, double, double)
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/fsub-fmul-rhs-combine.ll b/llvm/test/Transforms/SLPVectorizer/X86/fsub-fmul-rhs-combine.ll
index b766d92bd13fb..40de554b09a1b 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/fsub-fmul-rhs-combine.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/fsub-fmul-rhs-combine.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 5
; RUN: opt -S -passes=slp-vectorizer -mtriple=x86_64-unknown-linux-gnu -mcpu=x86-64-v3 < %s | FileCheck %s
+; RUN: opt -S -passes=slp-vectorizer -mtriple=x86_64-unknown-linux-gnu -mcpu=x86-64-v3 -slp-use-vplan-codegen < %s | FileCheck %s
; c - a*b with the fmul in the subtrahend. The backend fuses it into an fmsub
; just like a*b - c, so the fmul/fsub pair forms the combined fmuladd node for
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/lookahead.ll b/llvm/test/Transforms/SLPVectorizer/X86/lookahead.ll
index 81c6b60b3fdde..e79a00ae557ae 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/lookahead.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/lookahead.ll
@@ -1,6 +1,8 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt -passes=slp-vectorizer -S < %s -mtriple=x86_64-unknown-linux -mattr=+sse2 | FileCheck %s --check-prefixes=CHECK,SSE
+; RUN: opt -passes=slp-vectorizer -S < %s -mtriple=x86_64-unknown-linux -mattr=+sse2 -slp-use-vplan-codegen | FileCheck %s --check-prefixes=CHECK,SSE
; RUN: opt -passes=slp-vectorizer -S < %s -mtriple=x86_64-unknown-linux -mcpu=corei7-avx | FileCheck %s --check-prefixes=CHECK,AVX
+; RUN: opt -passes=slp-vectorizer -S < %s -mtriple=x86_64-unknown-linux -mcpu=corei7-avx -slp-use-vplan-codegen | FileCheck %s --check-prefixes=CHECK,AVX
;
; This file tests the look-ahead operand reordering heuristic.
;
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/metadata.ll b/llvm/test/Transforms/SLPVectorizer/X86/metadata.ll
index c11e5d16b4f07..6378facfba2af 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/metadata.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/metadata.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
; RUN: opt < %s -passes=slp-vectorizer,dce -S -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx | FileCheck %s
+; RUN: opt < %s -passes=slp-vectorizer,dce -S -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx -slp-use-vplan-codegen | FileCheck %s
target datalayout = "e-p:64:64:64-i1:8:8-i8:8:8-i16:16:16-i32:32:32-i64:64:64-f32:32:32-f64:64:64-v64:64:64-v128:128:128-a0:0:64-s0:64:64-f80:128:128-n8:16:32:64-S128"
target triple = "x86_64-apple-macosx10.8.0"
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/operandorder.ll b/llvm/test/Transforms/SLPVectorizer/X86/operandorder.ll
index 93ffc2985dd1f..97b25f2fe8997 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/operandorder.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/operandorder.ll
@@ -1,6 +1,8 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -passes=slp-vectorizer,dce -slp-threshold=-100 -S -mtriple=i386-apple-macosx10.8.0 -mcpu=corei7-avx | FileCheck %s
+; RUN: opt < %s -passes=slp-vectorizer,dce -slp-threshold=-100 -S -mtriple=i386-apple-macosx10.8.0 -mcpu=corei7-avx -slp-use-vplan-codegen | FileCheck %s
; RUN: opt < %s -passes=slp-vectorizer,dce -slp-threshold=-100 -S -mtriple=i386-apple-macosx10.8.0 -mattr=+sse2 | FileCheck %s --check-prefix=SSE2
+; RUN: opt < %s -passes=slp-vectorizer,dce -slp-threshold=-100 -S -mtriple=i386-apple-macosx10.8.0 -mattr=+sse2 -slp-use-vplan-codegen | FileCheck %s --check-prefix=SSE2
target datalayout = "e-p:32:32:32-i1:8:8-i8:8:8-i16:16:16-i32:32:32-i64:32:64-f32:32:32-f64:32:64-v64:64:64-v128:128:128-a0:0:64-f80:128:128-n8:16:32-S128"
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/opt.ll b/llvm/test/Transforms/SLPVectorizer/X86/opt.ll
index bfe93dc7f88b7..4a64f38b91266 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/opt.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/opt.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -O3 -S -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx | FileCheck %s --check-prefix=SLP
+; RUN: opt < %s -O3 -S -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx -slp-use-vplan-codegen | FileCheck %s --check-prefix=SLP
; RUN: opt < %s -O3 -vectorize-slp=false -S -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx | FileCheck %s --check-prefix=NOSLP
target datalayout = "e-p:64:64:64-i1:8:8-i8:8:8-i16:16:16-i32:32:32-i64:64:64-f32:32:32-f64:64:64-v64:64:64-v128:128:128-a0:0:64-s0:64:64-f80:128:128-n8:16:32:64-S128"
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/reassociate-ops.ll b/llvm/test/Transforms/SLPVectorizer/X86/reassociate-ops.ll
index 46ea5799b5b95..c2971854b9695 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/reassociate-ops.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/reassociate-ops.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
; RUN: opt -passes=slp-vectorizer -S < %s -mtriple=x86_64-unknown-linux -mcpu=corei7-avx | FileCheck %s
+; RUN: opt -passes=slp-vectorizer -S < %s -mtriple=x86_64-unknown-linux -mcpu=corei7-avx -slp-use-vplan-codegen | FileCheck %s
;
; Lanes 0 and 1 use the same 3 terms {A, B, C}, just nested/paired
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/runtime-alias-checks.ll b/llvm/test/Transforms/SLPVectorizer/X86/runtime-alias-checks.ll
index 2ae0b7fc51bda..05c9c063afb4b 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/runtime-alias-checks.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/runtime-alias-checks.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
; RUN: opt -passes=slp-vectorizer -mtriple=x86_64-unknown-linux-gnu -mcpu=skylake-avx512 -S < %s | FileCheck %s
+; RUN: opt -passes=slp-vectorizer -mtriple=x86_64-unknown-linux-gnu -mcpu=skylake-avx512 -slp-use-vplan-codegen -S < %s | FileCheck %s
; RUN: opt -passes=slp-vectorizer -mtriple=x86_64-unknown-linux-gnu -mcpu=skylake-avx512 -slp-vectorize-with-runtime-alias-checks=false -S < %s | FileCheck %s --check-prefix=NOCHK
declare void @nodup() noduplicate
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/schedule_budget.ll b/llvm/test/Transforms/SLPVectorizer/X86/schedule_budget.ll
index d9c8a76d208c1..b2b2acb2298e8 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/schedule_budget.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/schedule_budget.ll
@@ -1,6 +1,8 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -passes=slp-vectorizer -S -slp-schedule-budget=16 -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx | FileCheck --check-prefix LOBUDGET %s
+; RUN: opt < %s -passes=slp-vectorizer -S -slp-schedule-budget=16 -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx -slp-use-vplan-codegen | FileCheck --check-prefix LOBUDGET %s
; RUN: opt < %s -passes=slp-vectorizer -S -slp-schedule-budget=32 -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx | FileCheck --check-prefix HIBUDGET %s
+; RUN: opt < %s -passes=slp-vectorizer -S -slp-schedule-budget=32 -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx -slp-use-vplan-codegen | FileCheck --check-prefix HIBUDGET %s
target datalayout = "e-m:o-i64:64-f80:128-n8:16:32:64-S128"
target triple = "x86_64-apple-macosx10.9.0"
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/schedule_budget_debug_info.ll b/llvm/test/Transforms/SLPVectorizer/X86/schedule_budget_debug_info.ll
index 031cf87863e00..6e49df73963be 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/schedule_budget_debug_info.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/schedule_budget_debug_info.ll
@@ -1,6 +1,8 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -passes=slp-vectorizer -S -slp-schedule-budget=3 -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx | FileCheck %s -check-prefix VECTOR_DBG
+; RUN: opt < %s -passes=slp-vectorizer -S -slp-schedule-budget=3 -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx -slp-use-vplan-codegen | FileCheck %s -check-prefix VECTOR_DBG
; RUN: opt < %s -strip-debug -passes=slp-vectorizer -S -slp-schedule-budget=3 -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx | FileCheck %s -check-prefix VECTOR_NODBG
+; RUN: opt < %s -strip-debug -passes=slp-vectorizer -S -slp-schedule-budget=3 -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx -slp-use-vplan-codegen | FileCheck %s -check-prefix VECTOR_NODBG
target datalayout = "e-m:o-i64:64-f80:128-n8:16:32:64-S128"
target triple = "x86_64-apple-macosx10.9.0"
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/shift-ashr.ll b/llvm/test/Transforms/SLPVectorizer/X86/shift-ashr.ll
index e4f5e23f13279..9426540a99620 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/shift-ashr.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/shift-ashr.ll
@@ -1,15 +1,26 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -mtriple=x86_64-unknown -passes=slp-vectorizer -S | FileCheck %s --check-prefix=SSE
+; RUN: opt < %s -mtriple=x86_64-unknown -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=SSE
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=AVX,AVX1
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=AVX,AVX1
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=AVX,AVX2
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=AVX,AVX2
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=knl -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=knl -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=AVX,AVX2
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=AVX,AVX2
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=AVX,AVX2
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=AVX,AVX2
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=AVX,AVX2
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=AVX,AVX2
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=bdver4 -passes=slp-vectorizer -S | FileCheck %s --check-prefix=XOP
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=bdver4 -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=XOP
@a64 = common global [8 x i64] zeroinitializer, align 64
@b64 = common global [8 x i64] zeroinitializer, align 64
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/shift-lshr.ll b/llvm/test/Transforms/SLPVectorizer/X86/shift-lshr.ll
index b859ea57f16e0..4a059c4ca6e51 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/shift-lshr.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/shift-lshr.ll
@@ -1,15 +1,26 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -mtriple=x86_64-unknown -passes=slp-vectorizer -S | FileCheck %s --check-prefix=SSE
+; RUN: opt < %s -mtriple=x86_64-unknown -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=SSE
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=knl -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=knl -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=bdver4 -passes=slp-vectorizer -S | FileCheck %s --check-prefix=XOP
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=bdver4 -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=XOP
@a64 = common global [8 x i64] zeroinitializer, align 64
@b64 = common global [8 x i64] zeroinitializer, align 64
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/shift-shl.ll b/llvm/test/Transforms/SLPVectorizer/X86/shift-shl.ll
index c518a066c4f89..1aab9bdb46b54 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/shift-shl.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/shift-shl.ll
@@ -1,15 +1,26 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -mtriple=x86_64-unknown -passes=slp-vectorizer -S | FileCheck %s --check-prefix=SSE
+; RUN: opt < %s -mtriple=x86_64-unknown -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=SSE
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=corei7-avx -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=core-avx2 -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=knl -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=knl -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=skx -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m7 -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=-prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefix=AVX512
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=-prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=AVX512
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=+prefer-256-bit -passes=slp-vectorizer -S | FileCheck %s --check-prefixes=AVX
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=c86-4g-m8 -mattr=+prefer-256-bit -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefixes=AVX
; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=bdver4 -passes=slp-vectorizer -S | FileCheck %s --check-prefix=XOP
+; RUN: opt < %s -mtriple=x86_64-unknown -mcpu=bdver4 -passes=slp-vectorizer -S -slp-use-vplan-codegen | FileCheck %s --check-prefix=XOP
@a64 = common global [8 x i64] zeroinitializer, align 64
@b64 = common global [8 x i64] zeroinitializer, align 64
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/simplebb.ll b/llvm/test/Transforms/SLPVectorizer/X86/simplebb.ll
index 7782cc34125f7..1bf04dda5d037 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/simplebb.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/simplebb.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -passes=slp-vectorizer,dce -S -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx | FileCheck %s
+; RUN: opt < %s -passes=slp-vectorizer,dce -S -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx -slp-use-vplan-codegen | FileCheck %s
target datalayout = "e-p:64:64:64-i1:8:8-i8:8:8-i16:16:16-i32:32:32-i64:64:64-f32:32:32-f64:64:64-v64:64:64-v128:128:128-a0:0:64-s0:64:64-f80:128:128-n8:16:32:64-S128"
target triple = "x86_64-apple-macosx10.8.0"
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/tiny-tree.ll b/llvm/test/Transforms/SLPVectorizer/X86/tiny-tree.ll
index d592c49725483..850b4e243165b 100644
--- a/llvm/test/Transforms/SLPVectorizer/X86/tiny-tree.ll
+++ b/llvm/test/Transforms/SLPVectorizer/X86/tiny-tree.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx | FileCheck %s
+; RUN: opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-apple-macosx10.8.0 -mcpu=corei7-avx -slp-use-vplan-codegen | FileCheck %s
define void @tiny_tree_fully_vectorizable(ptr noalias nocapture %dst, ptr noalias nocapture readonly %src, i64 %count) #0 {
; CHECK-LABEL: @tiny_tree_fully_vectorizable(
diff --git a/llvm/test/Transforms/SLPVectorizer/consecutive-access.ll b/llvm/test/Transforms/SLPVectorizer/consecutive-access.ll
index 8347d5c8094b1..15968bd7aea06 100644
--- a/llvm/test/Transforms/SLPVectorizer/consecutive-access.ll
+++ b/llvm/test/Transforms/SLPVectorizer/consecutive-access.ll
@@ -1,6 +1,8 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: %if x86-registered-target %{ opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-apple-macosx10.9.0 | FileCheck %s --check-prefixes=CHECK,X86 %}
+; RUN: %if x86-registered-target %{ opt < %s -passes=slp-vectorizer -S -mtriple=x86_64-apple-macosx10.9.0 -slp-use-vplan-codegen | FileCheck %s --check-prefixes=CHECK,X86 %}
; RUN: %if aarch64-registered-target %{ opt < %s -passes=slp-vectorizer -S -mtriple=aarch64-unknown-linux-gnu | FileCheck %s --check-prefixes=CHECK,AARCH64 %}
+; RUN: %if aarch64-registered-target %{ opt < %s -passes=slp-vectorizer -S -mtriple=aarch64-unknown-linux-gnu -slp-use-vplan-codegen | FileCheck %s --check-prefixes=CHECK,AARCH64 %}
@A = common global [2000 x double] zeroinitializer, align 16
@B = common global [2000 x double] zeroinitializer, align 16
diff --git a/llvm/test/Transforms/SLPVectorizer/int_sideeffect.ll b/llvm/test/Transforms/SLPVectorizer/int_sideeffect.ll
index e075e9291639b..1e64a91bc7664 100644
--- a/llvm/test/Transforms/SLPVectorizer/int_sideeffect.ll
+++ b/llvm/test/Transforms/SLPVectorizer/int_sideeffect.ll
@@ -1,5 +1,6 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
; RUN: opt -S < %s -passes=slp-vectorizer -slp-max-reg-size=128 -slp-min-reg-size=128 | FileCheck %s
+; RUN: opt -S < %s -passes=slp-vectorizer -slp-max-reg-size=128 -slp-min-reg-size=128 -slp-use-vplan-codegen | FileCheck %s
declare void @llvm.sideeffect()
More information about the llvm-commits
mailing list