[llvm] [IR] Add llvm.vector.shuffle intrinsic for vector shuffles with runtime masks Part 1. (PR #219259)
Oscar Smith via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 22 22:32:26 PDT 2026
================
@@ -2911,6 +2921,88 @@ void DAGTypeLegalizer::SplitVecRes_VECTOR_COMPRESS(SDNode *N, SDValue &Lo,
std::tie(Lo, Hi) = DAG.SplitVector(Compressed, DL);
}
+void DAGTypeLegalizer::SplitVecRes_VECTOR_SHUFFLE_VAR(SDNode *N, SDValue &Lo,
+ SDValue &Hi) {
+ SDLoc DL(N);
+ SDValue Src = N->getOperand(0), Mask = N->getOperand(1);
+ EVT MaskVT = Mask.getValueType();
+ unsigned MaskEltBits = MaskVT.getScalarSizeInBits();
+
+ SDValue SrcLo, SrcHi;
+ if (getTypeAction(Src.getValueType()) == TargetLowering::TypeSplitVector)
+ GetSplitVector(Src, SrcLo, SrcHi);
+ else
+ std::tie(SrcLo, SrcHi) = DAG.SplitVector(Src, DL);
+
+ SDValue MaskLo, MaskHi;
+ if (getTypeAction(MaskVT) == TargetLowering::TypeSplitVector)
+ GetSplitVector(Mask, MaskLo, MaskHi);
+ else
+ std::tie(MaskLo, MaskHi) = DAG.SplitVector(Mask, DL);
+
+ EVT HalfVT = SrcLo.getValueType();
+ EVT HalfMaskVT = MaskLo.getValueType();
+ ElementCount HalfEC = HalfVT.getVectorElementCount();
+ APInt MaxIdx = APInt::getMaxValue(MaskEltBits);
+
+ // Every index the mask element type can hold lands in the low half of the
+ // source, so the high half is unreachable and each result half is just a
+ // shuffle of the low source half.
+ if (MaxIdx.ult(HalfEC.getKnownMinValue())) {
+ Lo = DAG.getNode(ISD::VECTOR_SHUFFLE_VAR, DL, HalfVT, SrcLo, MaskLo);
+ Hi = DAG.getNode(ISD::VECTOR_SHUFFLE_VAR, DL, HalfVT, SrcLo, MaskHi);
+ return;
+ }
+
+ // Otherwise shuffle both source halves with the same indices and select per
+ // lane on which half the index falls in:
+ // Half[i] = Idx[i] <u HalfElts ? SrcLo[Idx[i]] : SrcHi[Idx[i] - HalfElts]
+ // Indices past the end of the whole source are out of range in the SrcHi
+ // shuffle too, so they stay poison. This needs HalfElts to be representable
+ // in the mask element type; for scalable vectors that is only known when the
+ // elements are wide enough to hold any element count a target can produce.
+ // A fixed HalfElts is representable: the early return above took every case
+ // where the mask element type cannot reach it.
+ if (HalfEC.isScalable() && MaskEltBits < 32) {
+ SDValue Expanded = TLI.expandVECTOR_SHUFFLE_VAR(N, DAG);
+ std::tie(Lo, Hi) = DAG.SplitVector(Expanded, DL);
+ return;
+ }
+
+ // Give the half masks the result half's integer element type, so that the
+ // select's condition type matches its value type as targets expect. Indices
----------------
oscardssmith wrote:
Not target-specific. per-half select needs a condition that matches the value type, but the compare producing it is on the mask, so the compare result is converted to the result's condition type. The `split_v8i32_v8i8` covers this (and fails otherwise) and I assume other backends do as well.
https://github.com/llvm/llvm-project/pull/219259
More information about the llvm-commits
mailing list