[llvm] [SelectionDAG] optimize sdiv with positive divisor and positive magic (PR #189287)
Craig Topper via llvm-commits
llvm-commits at lists.llvm.org
Wed Apr 29 09:56:46 PDT 2026
topperc wrote:
> > It looks like this is only better on X86 because it avoids a single `mov`, because X86 only has two-operand shift instructions and because we already happened to have a copy of the input lying around in `edi`.
> > Does it show any benefit on other archs? I would expect in some cases it could even be worse, since register pressure is higher at the point of the `imul` (both `rax` and `edi` are live).
>
> For aarch64: I believe this is the same as x86, main uses an extra reg x9, the branch reuses w0.
>
> main:
>
> ```assembly
> sdiv3: // @sdiv3
> mov w8, #21846
> movk w8, #21845, lsl #16
> smull x8, w0, w8
> lsr x9, x8, #32
> add x0, x9, x8, lsr #63
> ret
> ```
>
> this branch:
>
> ```assembly
> sdiv3: // @sdiv3
> mov w8, #21846
> movk w8, #21845, lsl #16
> smull x8, w0, w8
> lsr x8, x8, #32
> sub w0, w8, w0, asr #31
> ret
> ```
>
> For risc-v with +m: The branch uses (a0, a1, a2) and main uses (a0, a1), so extra register pressure.
>
> main:
>
> ```assembly
> sdiv3: # @sdiv3
> sext.w a0, a0
> lui a1, 349525
> addi a1, a1, 1366
> mul a0, a0, a1
> srli a1, a0, 63
> srli a0, a0, 32
> addw a0, a0, a1
> ret
> ```
>
> branch:
>
> ```assembly
> sdiv3: # @sdiv3
> sext.w a1, a0
> lui a2, 349525
> addi a2, a2, 1366
> mul a1, a1, a2
> srli a1, a1, 32
> sraiw a0, a0, 31
> subw a0, a1, a0
> ret
> ```
>
> For ppc: mulhw removes the need for an and or mov here so the change is srwi + add turning into srawi + sub. This is at best neutral so not worth it.
>
> main:
>
> ```assembly
> sdiv3: # @sdiv3
> lis 4, 21845
> ori 4, 4, 21846
> mulhw 3, 3, 4
> srwi 4, 3, 31
> add 3, 3, 4
> blr
> ```
>
> branch:
>
> ```assembly
> sdiv3: # @sdiv3
> lis 4, 21845
> ori 4, 4, 21846
> mulhw 4, 3, 4
> srawi 3, 3, 31
> sub 3, 4, 3
> blr
> ```
>
> arm32 uses smmul for the same pattern as ppc. x86_32 emits worse code for this branch cause it emits an extra mov, unlike x86_64.
>
> Looks like it's only useful for x86_64, and for aarch64 without neon, so I need to narrow this down.
For ppc, the shift and mulhw from this PR can be done in parallel which shortens the critical path. For AArch64, RISC-V, X86-64, we're emulating an i32 mulh with a sext before the mul and a shift after. That final shift can be done in parallel with the shift by 63 from the old code so we don't shorten the critical path.
https://github.com/llvm/llvm-project/pull/189287
More information about the llvm-commits
mailing list