[llvm] [AArch64][GlobalISel] Add known bits for G_DUP, G_VASHR and G_VLSHR (PR #210503)
Joel Walker via llvm-commits
llvm-commits at lists.llvm.org
Sun Sep 13 16:03:21 PDT 2026
================
@@ -3087,6 +3087,46 @@ unsigned AArch64TargetLowering::ComputeNumSignBitsForTargetNode(
return 1;
}
+void AArch64TargetLowering::computeKnownBitsForTargetInstr(
+ GISelValueTracking &Analysis, Register R, KnownBits &Known,
+ const APInt &DemandedElts, const MachineRegisterInfo &MRI,
+ unsigned Depth) const {
+ const MachineInstr *MI = MRI.getVRegDef(R);
+ switch (MI->getOpcode()) {
+ case AArch64::G_DUP: {
+ Register Src = MI->getOperand(1).getReg();
+ unsigned SrcBits = MRI.getType(Src).getSizeInBits();
+ KnownBits SrcKnown(SrcBits);
+ Analysis.computeKnownBitsImpl(Src, SrcKnown, APInt(1, 1), Depth + 1);
+ unsigned EltBits = MRI.getType(R).getScalarSizeInBits();
+ if (SrcBits != EltBits) {
+ assert(SrcBits > EltBits && "Expected DUP implicit truncation");
+ SrcKnown = SrcKnown.trunc(EltBits);
+ }
+ Known = SrcKnown;
+ break;
+ }
+ case AArch64::G_VASHR:
+ case AArch64::G_VLSHR: {
+ unsigned BitWidth = MRI.getType(R).getScalarSizeInBits();
+ uint64_t Shift = MI->getOperand(2).getImm();
+ // A shift by the full element width is legal for these instructions, but
+ // KnownBits::ashr/lshr model IR shifts, for which it is poison. Leave the
+ // result unknown in that case.
+ if (Shift >= BitWidth)
+ break;
----------------
Joel-Wwalker wrote:
The DAG never forms a full-width VASHR/VLSHR: `LowerVectorSRA_SRL_SHL` requires `Cnt < EltSize`, and `performVectorShiftCombine` asserts `OpScalarSize > ShiftImm`. The GlobalISel matcher (`isVShiftRImm` in AArch64PostLegalizerLowering) accepts `Cnt <= ElementBits`, so at -O0 `ashr <4 x i32> %x, splat (i32 32)` becomes `G_VASHR %x, 32` and then `sshr v0.4s, v0.4s, #32` (at -O2 the combiner turns the poison shift into `G_IMPLICIT_DEF` first). `KnownBits::ashr` treats that amount as poison and returns all-zero, while `sshr` of a negative lane produces all-ones; that is what the guard was avoiding.
The shift amount is an immediate operand here rather than a register, so `makeConstant` is the equivalent of the DAG calling `computeKnownBits` on its TargetConstant.
Rather than giving up on that case, the update clamps the `sshr` amount to `BitWidth - 1`, which yields the same lane, and leaves `ushr` alone since `KnownBits::lshr` already returns zero for a shift by the width. This matches the existing G_VASHR sign-bits case, which already treats `Imm == width` as sign-filling. Tests added for both; also rebased since arm64-neon-aba-abd.ll had drifted.
Assisted by Claude (Anthropic).
https://github.com/llvm/llvm-project/pull/210503
More information about the llvm-commits
mailing list