[llvm] [AArch64] Add DAG combine to widen v3i8/v4i8 VECTOR_MATCH needles (PR #218972)
Benjamin Maxwell via llvm-commits
llvm-commits at lists.llvm.org
Thu Aug 27 07:19:08 PDT 2026
================
@@ -31480,6 +31457,40 @@ static SDValue performPredicateLoadCombine(SDNode *N,
return LoadPred;
}
+static SDValue performVectorMatchCombine(SDNode *N,
+ TargetLowering::DAGCombinerInfo &DCI,
+ SelectionDAG &DAG) {
+ // Widen v3i8/v4i8 match needles to v8i8. For needles >= 2 elements it's
+ // generally better to widen the needle and lower to a `match` (rather than
+ // expanding). This is handled by type legalization for all other types (> 2
+ // elements) but v4i8 is promoted rather than widened, hence this combine.
+ SDValue Needle = N->getOperand(1);
+ EVT NeedleVT = Needle.getValueType();
+
+ if (!DCI.isBeforeLegalize() ||
+ (NeedleVT != MVT::v3i8 && NeedleVT != MVT::v4i8))
+ return SDValue();
+
+ SDLoc DL(N);
+ if (NeedleVT == MVT::v3i8) {
+ // Pad a v3i8 needle to v4i8.
+ SDValue Pad = DAG.getExtractVectorElt(DL, MVT::i8, Needle, 0);
+ Needle = DAG.getInsertSubvector(DL, DAG.getPOISON(MVT::v4i8), Needle, 0);
+ Needle = DAG.getInsertVectorElt(DL, Needle, Pad, 3);
+ }
+
+ // Splat the needle to a full scalable vector (in such a way that the bitcasts
+ // and extracts should fold away).
+ Needle = DAG.getBitcast(MVT::v1i32, Needle);
+ Needle = DAG.getExtractVectorElt(DL, MVT::i32, Needle, 0);
+ Needle = DAG.getSplatVector(MVT::nxv4i32, DL, Needle);
+ Needle = DAG.getBitcast(MVT::nxv16i8, Needle);
+ Needle = DAG.getExtractSubvector(DL, MVT::v16i8, Needle, 0);
----------------
MacDue wrote:
I've simplified this to just bitcast, splat, bitcast. The anticipating the promotion of the needle to scalable vectors was not necessary.
https://github.com/llvm/llvm-project/pull/218972
More information about the llvm-commits
mailing list