[llvm] [SelectionDAG] optimize sdiv with positive divisor and positive magic (PR #189287)
Takashi Idobe via llvm-commits
llvm-commits at lists.llvm.org
Tue Apr 28 05:22:37 PDT 2026
================
@@ -6772,8 +6780,19 @@ SDValue TargetLowering::BuildSDIV(SDNode *N, SelectionDAG &DAG,
Q = DAG.getNode(ISD::SRA, dl, VT, Q, Shift);
Created.push_back(Q.getNode());
- // Extract the sign bit, mask it and add it to the quotient.
+ // Extract the sign bit, mask it and add/subtract it from the quotient.
SDValue SignShift = DAG.getConstant(EltBits - 1, dl, ShVT);
+
+ // SRA replicates the sign bit to fill the register width, so this works
+ // correctly for promoted types (where N0's upper bits may be set from
+ // sign-extension) without needing an extra masking AND.
+ // vector targets may have fused shift-accumulate instructions that make
+ // SRL+ADD cheaper, so gate to scalar only.
+ if (UseInputSign && !VT.isVector()) {
----------------
Takashiidobe wrote:
arm neon has SSRA, so the SRA + SUB path is two instructions compared to the one that neon could emit. I assume this fold would be a pessimization in that case, at least in terms of op count, and possibly in terms of latency. I'm unsure if other targets have similar instructions.
I think the best way to handle this is to have a hook to opt in or opt out, but that will take a bit more research so I tried to gate to scalar only for now.
https://github.com/llvm/llvm-project/pull/189287
More information about the llvm-commits
mailing list