[llvm] [RISCV] Lower VECTOR_INTERLEAVE 4 and 8 with Zvzip (PR #222432)
Luke Lau via llvm-commits
llvm-commits at lists.llvm.org
Thu Sep 10 21:58:02 PDT 2026
================
@@ -14641,6 +14641,50 @@ SDValue RISCVTargetLowering::lowerVECTOR_INTERLEAVE(SDValue Op,
SDValue Interleaved;
+ if (Subtarget.hasStdExtZvzip() && (Factor == 4 || Factor == 8)) {
+ // Interleave by a tree of vzip.vv instructions.
+ SmallVector<SDValue, 8> Operands(Op->op_values());
+ // First, reorder the operands.
+ // For Factor=4, given the original operand order `ABCD`, we need
+ // to reorder it into `ACBD`.
+ // For Factor=8, given the original operand order `ABCDEFGH`, the new
+ // order should be `AECGBFDH`.
+ // So the rule here is that for every operands with an odd index `I`, swap
+ // it with the operand of index `I + (Factor / 2 - 1)`.
+ for (unsigned I = 1U, HalfFactor = Factor / 2; I < HalfFactor; I += 2)
+ std::swap(Operands[I], Operands[I + (HalfFactor - 1)]);
+
+ for (unsigned CurrFactor = Factor; CurrFactor > 1; CurrFactor /= 2) {
+ // Generate a vzip.vv for every two operands.
+ for (unsigned I = 0U; I < CurrFactor; I += 2) {
+ assert(Operands[I].getValueType().isSimple() &&
+ isLegalVTForZvzipOperand(Operands[I].getSimpleValueType(),
+ Subtarget));
+ // Do not generate VECTOR_INTERLEAVE2 + CONCAT_VECTORS here. Because
+ // when those two nodes are subsequently lowered, there will
+ // be a bunch of insert_subvector and extract_subvector generated.
+ // Though most of them can be combined, since we're not
+ // running DAGCombiner in between different operations' lowering,
+ // some of the insert/extract subvectors will be turned into
+ // VSLIDEUP/DOWN_VL right away and stay thru the rest of the codegen.
+ // We could write additional combining rules for those VSLIDEUP/DOWN_VL
+ // but it'll probably be a lot easier to just not generate
+ // VECTOR_INTERLEAVE2 + CONCAT_VECTORS in the first place here.
+ Operands[I / 2] =
+ lowerZvzipVZIP(Operands[I], Operands[I + 1], DL, DAG, Subtarget);
----------------
lukel97 wrote:
Can we just do this as a combine on isd::vector_interleave instead of during lowering so we have a chance to eliminate the slides/extracts? I feel like we should already have enough generic combines to avoid the need for #222750.
https://github.com/llvm/llvm-project/pull/222432
More information about the llvm-commits
mailing list