[llvm] [AArch64][GlobalISel] Add known bits for G_DUP, G_VASHR and G_VLSHR (PR #210503)

Joel Walker via llvm-commits llvm-commits at lists.llvm.org
Sun Sep 13 16:03:21 PDT 2026


================
@@ -3087,6 +3087,46 @@ unsigned AArch64TargetLowering::ComputeNumSignBitsForTargetNode(
   return 1;
 }
 
+void AArch64TargetLowering::computeKnownBitsForTargetInstr(
+    GISelValueTracking &Analysis, Register R, KnownBits &Known,
+    const APInt &DemandedElts, const MachineRegisterInfo &MRI,
+    unsigned Depth) const {
+  const MachineInstr *MI = MRI.getVRegDef(R);
+  switch (MI->getOpcode()) {
+  case AArch64::G_DUP: {
+    Register Src = MI->getOperand(1).getReg();
+    unsigned SrcBits = MRI.getType(Src).getSizeInBits();
+    KnownBits SrcKnown(SrcBits);
+    Analysis.computeKnownBitsImpl(Src, SrcKnown, APInt(1, 1), Depth + 1);
+    unsigned EltBits = MRI.getType(R).getScalarSizeInBits();
+    if (SrcBits != EltBits) {
+      assert(SrcBits > EltBits && "Expected DUP implicit truncation");
+      SrcKnown = SrcKnown.trunc(EltBits);
+    }
+    Known = SrcKnown;
+    break;
+  }
+  case AArch64::G_VASHR:
+  case AArch64::G_VLSHR: {
+    unsigned BitWidth = MRI.getType(R).getScalarSizeInBits();
+    uint64_t Shift = MI->getOperand(2).getImm();
+    // A shift by the full element width is legal for these instructions, but
+    // KnownBits::ashr/lshr model IR shifts, for which it is poison. Leave the
+    // result unknown in that case.
+    if (Shift >= BitWidth)
+      break;
----------------
Joel-Wwalker wrote:

The DAG never forms a full-width VASHR/VLSHR: `LowerVectorSRA_SRL_SHL` requires `Cnt < EltSize`, and `performVectorShiftCombine` asserts `OpScalarSize > ShiftImm`. The GlobalISel matcher (`isVShiftRImm` in AArch64PostLegalizerLowering) accepts `Cnt <= ElementBits`, so at -O0 `ashr <4 x i32> %x, splat (i32 32)` becomes `G_VASHR %x, 32` and then `sshr v0.4s, v0.4s, #32` (at -O2 the combiner turns the poison shift into `G_IMPLICIT_DEF` first). `KnownBits::ashr` treats that amount as poison and returns all-zero, while `sshr` of a negative lane produces all-ones; that is what the guard was avoiding.

The shift amount is an immediate operand here rather than a register, so `makeConstant` is the equivalent of the DAG calling `computeKnownBits` on its TargetConstant.

Rather than giving up on that case, the update clamps the `sshr` amount to `BitWidth - 1`, which yields the same lane, and leaves `ushr` alone since `KnownBits::lshr` already returns zero for a shift by the width. This matches the existing G_VASHR sign-bits case, which already treats `Imm == width` as sign-filling. Tests added for both; also rebased since arm64-neon-aba-abd.ll had drifted.

Assisted by Claude (Anthropic).


https://github.com/llvm/llvm-project/pull/210503


More information about the llvm-commits mailing list