[llvm] [SelectionDAG] optimize sdiv with positive divisor and positive magic (PR #189287)

Craig Topper via llvm-commits llvm-commits at lists.llvm.org
Wed Apr 29 09:56:46 PDT 2026


topperc wrote:





> > It looks like this is only better on X86 because it avoids a single `mov`, because X86 only has two-operand shift instructions and because we already happened to have a copy of the input lying around in `edi`.
> > Does it show any benefit on other archs? I would expect in some cases it could even be worse, since register pressure is higher at the point of the `imul` (both `rax` and `edi` are live).
> 
> For aarch64: I believe this is the same as x86, main uses an extra reg x9, the branch reuses w0.
> 
> main:
> 
> ```assembly
> sdiv3:                                  // @sdiv3
> 	mov	w8, #21846
> 	movk	w8, #21845, lsl #16
> 	smull	x8, w0, w8
> 	lsr	x9, x8, #32
> 	add	x0, x9, x8, lsr #63
> 	ret
> ```
> 
> this branch:
> 
> ```assembly
> sdiv3:                                  // @sdiv3
> 	mov	w8, #21846
> 	movk	w8, #21845, lsl #16
> 	smull	x8, w0, w8
> 	lsr	x8, x8, #32
> 	sub	w0, w8, w0, asr #31
> 	ret
> ```
> 
> For risc-v with +m: The branch uses (a0, a1, a2) and main uses (a0, a1), so extra register pressure.
> 
> main:
> 
> ```assembly
> sdiv3:                                  # @sdiv3
> 	sext.w	a0, a0
> 	lui	a1, 349525
> 	addi	a1, a1, 1366
> 	mul	a0, a0, a1
> 	srli	a1, a0, 63
> 	srli	a0, a0, 32
> 	addw	a0, a0, a1
> 	ret
> ```
> 
> branch:
> 
> ```assembly
> sdiv3:                                  # @sdiv3
> 	sext.w	a1, a0
> 	lui	a2, 349525
> 	addi	a2, a2, 1366
> 	mul	a1, a1, a2
> 	srli	a1, a1, 32
> 	sraiw	a0, a0, 31
> 	subw	a0, a1, a0
> 	ret
> ```
> 
> For ppc: mulhw removes the need for an and or mov here so the change is srwi + add turning into srawi + sub. This is at best neutral so not worth it.
> 
> main:
> 
> ```assembly
> sdiv3:                                  # @sdiv3
> 	lis 4, 21845
> 	ori 4, 4, 21846
> 	mulhw 3, 3, 4
> 	srwi 4, 3, 31
> 	add 3, 3, 4
> 	blr
> ```
> 
> branch:
> 
> ```assembly
> sdiv3:                                  # @sdiv3
> 	lis 4, 21845
> 	ori 4, 4, 21846
> 	mulhw 4, 3, 4
> 	srawi 3, 3, 31
> 	sub	3, 4, 3
> 	blr
> ```
> 
> arm32 uses smmul for the same pattern as ppc. x86_32 emits worse code for this branch cause it emits an extra mov, unlike x86_64.
> 
> Looks like it's only useful for x86_64, and for aarch64 without neon, so I need to narrow this down.

For ppc, the shift and mulhw from this PR can be done in parallel which shortens the critical path. For AArch64, RISC-V, X86-64, we're emulating an i32 mulh with a sext before the mul and a shift after. That final shift can be done in parallel with the shift by 63 from the old code so we don't shorten the critical path.

https://github.com/llvm/llvm-project/pull/189287


More information about the llvm-commits mailing list