[llvm] [SelectionDAG] Widen even 33-bit-magic udiv on free-zext targets (PR #207634)

MITSUNARI Shigeo via llvm-commits llvm-commits at lists.llvm.org
Tue Jul 7 17:50:20 PDT 2026


================
@@ -6921,6 +6921,12 @@ SDValue TargetLowering::BuildUDIV(SDNode *N, SelectionDAG &DAG,
       isOperationLegalOrCustom(ISD::UMUL_LOHI, WideSVT, IsAfterLegalization);
   const bool AllowWiden = (HasWideMULHU || HasWideUMUL_LOHI);
 
+  // For even divisors with a 33-bit magic number, the widened high-multiply
+  // path is only worthwhile over the even-divisor rewrite on targets that
+  // zero-extend i32 to i64 for free (e.g. x86-64 and AArch64). Elsewhere (e.g.
+  // RISC-V) keep the even-divisor rewrite, which avoids the explicit extension.
+  const bool AllowEvenToWiden = AllowWiden && isZExtFree(MVT::i32, WideSVT);
----------------
herumi wrote:

Let me explain why I gated the change on RISC-V with `isZExtFree()`.

Consider `uint32_t udiv14(uint32_t x) { return x / 14; }`.
On RISC-V, i32 values are produced and passed in sign-extended form, so the upper 32 bits of register `a0` are not guaranteed to be zero.
The widened `mulhu` path is a 64x64 multiply, so it needs to explicitly clear the upper bits of the dividend before multiplying.

Before (the existing even-divisor rewrite, `-mattr=+m`):

```asm
udiv14:
    lui   a1, 299593
    slli  a1, a1, 1
    srliw a0, a0, 1
    addi  a1, a1, 1171
    mul   a0, a0, a1
    srli  a0, a0, 34
    ret
```

After (the widened path, with the `isZExtFree()` condition removed):

```asm
udiv14:
    lui   a1, %hi(.LCPI0_0)
    ld    a1, %lo(.LCPI0_0)(a1)
    slli  a0, a0, 32 ; zero-extension
    srli  a0, a0, 32 ; zero-extension
    mulhu a0, a0, a1
    ret
```

The `slli a0, a0, 32` / `srli a0, a0, 32` pair is the explicit zero-extension.
On free-zext targets such as x86-64 and AArch64, this extension is absorbed into an instruction that runs anyway, but on RISC-V it costs two extra instructions.
That is why I conservatively added the `isZExtFree()` condition and kept the existing rewrite on RISC-V.

However, the situation changes when the division runs multiple times, e.g. inside a loop.
The magic constant is hoisted out of the loop, and if the dividend already arrives zero-extended (for example, from a `lwu` load), the explicit zero-extension disappears.
As a result the loop body shrinks from the three instructions `srli`/`mul`/`srli` to a single `mulhu`, so the widened path produces clearly better code on RISC-V as well.

Loop body comparison (`sum ^= p[i] / 14;`):

```asm
; Before (magic constant already hoisted)
    lwu   a5, 0(a5)
    srli  a5, a5, 1
    mul   a5, a5, a4
    srli  a5, a5, 34

; After
    lwu   a5, 0(a5)
    mulhu a5, a5, a4
```

So in this case it would be preferable to drop the `isZExtFree()` condition and widen on RISC-V too.
Which would you prefer: keeping the conservative guard, or dropping it and always widening on RISC-V?


https://github.com/llvm/llvm-project/pull/207634


More information about the llvm-commits mailing list