[llvm] [SelectionDAG] Widen even 33-bit-magic udiv on free-zext targets (PR #207634)
MITSUNARI Shigeo via llvm-commits
llvm-commits at lists.llvm.org
Tue Jul 7 17:50:20 PDT 2026
================
@@ -6921,6 +6921,12 @@ SDValue TargetLowering::BuildUDIV(SDNode *N, SelectionDAG &DAG,
isOperationLegalOrCustom(ISD::UMUL_LOHI, WideSVT, IsAfterLegalization);
const bool AllowWiden = (HasWideMULHU || HasWideUMUL_LOHI);
+ // For even divisors with a 33-bit magic number, the widened high-multiply
+ // path is only worthwhile over the even-divisor rewrite on targets that
+ // zero-extend i32 to i64 for free (e.g. x86-64 and AArch64). Elsewhere (e.g.
+ // RISC-V) keep the even-divisor rewrite, which avoids the explicit extension.
+ const bool AllowEvenToWiden = AllowWiden && isZExtFree(MVT::i32, WideSVT);
----------------
herumi wrote:
Let me explain why I gated the change on RISC-V with `isZExtFree()`.
Consider `uint32_t udiv14(uint32_t x) { return x / 14; }`.
On RISC-V, i32 values are produced and passed in sign-extended form, so the upper 32 bits of register `a0` are not guaranteed to be zero.
The widened `mulhu` path is a 64x64 multiply, so it needs to explicitly clear the upper bits of the dividend before multiplying.
Before (the existing even-divisor rewrite, `-mattr=+m`):
```asm
udiv14:
lui a1, 299593
slli a1, a1, 1
srliw a0, a0, 1
addi a1, a1, 1171
mul a0, a0, a1
srli a0, a0, 34
ret
```
After (the widened path, with the `isZExtFree()` condition removed):
```asm
udiv14:
lui a1, %hi(.LCPI0_0)
ld a1, %lo(.LCPI0_0)(a1)
slli a0, a0, 32 ; zero-extension
srli a0, a0, 32 ; zero-extension
mulhu a0, a0, a1
ret
```
The `slli a0, a0, 32` / `srli a0, a0, 32` pair is the explicit zero-extension.
On free-zext targets such as x86-64 and AArch64, this extension is absorbed into an instruction that runs anyway, but on RISC-V it costs two extra instructions.
That is why I conservatively added the `isZExtFree()` condition and kept the existing rewrite on RISC-V.
However, the situation changes when the division runs multiple times, e.g. inside a loop.
The magic constant is hoisted out of the loop, and if the dividend already arrives zero-extended (for example, from a `lwu` load), the explicit zero-extension disappears.
As a result the loop body shrinks from the three instructions `srli`/`mul`/`srli` to a single `mulhu`, so the widened path produces clearly better code on RISC-V as well.
Loop body comparison (`sum ^= p[i] / 14;`):
```asm
; Before (magic constant already hoisted)
lwu a5, 0(a5)
srli a5, a5, 1
mul a5, a5, a4
srli a5, a5, 34
; After
lwu a5, 0(a5)
mulhu a5, a5, a4
```
So in this case it would be preferable to drop the `isZExtFree()` condition and widen on RISC-V too.
Which would you prefer: keeping the conservative guard, or dropping it and always widening on RISC-V?
https://github.com/llvm/llvm-project/pull/207634
More information about the llvm-commits
mailing list