[llvm] [AArch64] Add DAG combine to widen v3i8/v4i8 VECTOR_MATCH needles (PR #218972)

Paul Walker via llvm-commits llvm-commits at lists.llvm.org
Thu Aug 27 05:01:40 PDT 2026


================
@@ -31480,6 +31457,40 @@ static SDValue performPredicateLoadCombine(SDNode *N,
   return LoadPred;
 }
 
+static SDValue performVectorMatchCombine(SDNode *N,
+                                         TargetLowering::DAGCombinerInfo &DCI,
+                                         SelectionDAG &DAG) {
+  // Widen v3i8/v4i8 match needles to v8i8. For needles >= 2 elements it's
+  // generally better to widen the needle and lower to a `match` (rather than
+  // expanding). This is handled by type legalization for all other types (> 2
+  // elements) but v4i8 is promoted rather than widened, hence this combine.
+  SDValue Needle = N->getOperand(1);
+  EVT NeedleVT = Needle.getValueType();
+
+  if (!DCI.isBeforeLegalize() ||
+      (NeedleVT != MVT::v3i8 && NeedleVT != MVT::v4i8))
+    return SDValue();
+
+  SDLoc DL(N);
+  if (NeedleVT == MVT::v3i8) {
+    // Pad a v3i8 needle to v4i8.
+    SDValue Pad = DAG.getExtractVectorElt(DL, MVT::i8, Needle, 0);
+    Needle = DAG.getInsertSubvector(DL, DAG.getPOISON(MVT::v4i8), Needle, 0);
+    Needle = DAG.getInsertVectorElt(DL, Needle, Pad, 3);
+  }
+
+  // Splat the needle to a full scalable vector (in such a way that the bitcasts
+  // and extracts should fold away).
+  Needle = DAG.getBitcast(MVT::v1i32, Needle);
+  Needle = DAG.getExtractVectorElt(DL, MVT::i32, Needle, 0);
+  Needle = DAG.getSplatVector(MVT::nxv4i32, DL, Needle);
+  Needle = DAG.getBitcast(MVT::nxv16i8, Needle);
+  Needle = DAG.getExtractSubvector(DL, MVT::v16i8, Needle, 0);
----------------
paulwalker-arm wrote:

This looks like you're second guessing future lowering and are thus building a DAG that will optimise in the direction you want, which I don't like. I assume we're not just missing some post legalisation combines?

Alternatively, once implemented as lowering you'll know the result type is legal (because result types are legalised before operand types and don't directly call the lower#### functions) so perhaps it's easier to simply lower the whole operation to SVE? 

https://github.com/llvm/llvm-project/pull/218972


More information about the llvm-commits mailing list