[llvm] [SelectionDAG] optimize sdiv with positive divisor and positive magic (PR #189287)

Takashi Idobe via llvm-commits llvm-commits at lists.llvm.org
Tue Apr 28 05:22:37 PDT 2026


================
@@ -6772,8 +6780,19 @@ SDValue TargetLowering::BuildSDIV(SDNode *N, SelectionDAG &DAG,
   Q = DAG.getNode(ISD::SRA, dl, VT, Q, Shift);
   Created.push_back(Q.getNode());
 
-  // Extract the sign bit, mask it and add it to the quotient.
+  // Extract the sign bit, mask it and add/subtract it from the quotient.
   SDValue SignShift = DAG.getConstant(EltBits - 1, dl, ShVT);
+
+  // SRA replicates the sign bit to fill the register width, so this works
+  // correctly for promoted types (where N0's upper bits may be set from
+  // sign-extension) without needing an extra masking AND.
+  // vector targets may have fused shift-accumulate instructions that make
+  // SRL+ADD cheaper, so gate to scalar only.
+  if (UseInputSign && !VT.isVector()) {
----------------
Takashiidobe wrote:

arm neon has SSRA, so the SRA + SUB path is two instructions compared to the one that neon could emit. I assume this fold would be a pessimization in that case, at least in terms of op count, and possibly in terms of latency. I'm unsure if other targets have similar instructions. 

I think the best way to handle this is to have a hook to opt in or opt out, but that will take a bit more research so I tried to gate to scalar only for now.

https://github.com/llvm/llvm-project/pull/189287


More information about the llvm-commits mailing list