[llvm] [AArch64MIPeepholeOpt] Coalesce sibling base-address materializations (PR #223682)

via llvm-commits llvm-commits at lists.llvm.org
Mon Sep 21 23:09:14 PDT 2026


whokeke wrote:

> We prefer aligned common bases for large positive gep bases, why do we not do the same for large negative bases? Would that cover most of this?

Thanks for raising this — we looked into CodeGenPrepare::splitLargeGEPOffsets (the pass that handles the positive-offset aligned-common-base case) and tried relaxing its ConstantOffset > 0 gate to != 0 to see if it would cover the negative-offset case this series handles.
Instruction count, dependency chain, and register pressure. For a single pair (e.g. p[-100]/p[-99]), relaxing CGP's gate plus #223684 (a separate patch that inserts a base-adjust ADDXri in AArch64LoadStoreOptimizer when a pair's offset is otherwise out of LDP/STP's range) produces:
sub x8, x0, #1, lsl #12   ; CGP's coarse rebase
add x10, x8, #3696        ; base-adjust's fine rebase
ldp w9, w8, [x10]
— 3 instructions, a 3-deep dependency chain, and 2 live address registers, versus shareBaseAddresses's:
sub x8, x0, #400
ldp w9, w8, [x8]
— 2 instructions, 2-deep chain, 1 register.
This is because getPreferredLargeGEPBaseOffset rebases to a 4096-byte-aligned high part (needed for LDR/STR's own 12-bit range), which is much coarser than LDP/STP's pairing window (a few hundred bytes depending on type). Relaxing CGP alone doesn't achieve pairing by itself — the residual offset after CGP's rebase (3696 bytes here) is still out of LDP's range. Pairing only happens once base-adjust does a second, finer rebase on top. shareBaseAddresses avoids the extra step by picking the exact minimum-offset sibling as the base directly, needing only one rebase. We checked this on both a small and a large negative offset and saw the same pattern both times.
So relaxing the CGP gate would let pairing happen for the negative case, but at a worse instruction count / dependency depth than the targeted fix — since pairing exists to save an instruction in the first place, the extra ADD here eats into that saving.
On the relationship between this series and #223684: they address complementary cases rather than overlapping ones. This series (shareBaseAddresses, pre-RA in AArch64MIPeepholeOpt) coalesces sibling base materializations when ISel has already produced different base registers for negative far offsets it can't fold into a single immediate. #223684 (post-RA, in AArch64LoadStoreOptimizer) handles the opposite starting point: accesses that already share a base register but whose offset delta is too large for LDP/STP's 7-bit field, inserting a base-adjust to bring it in range. Since #223684 is also under review right now, feedback there on the base-adjust approach would be very welcome too, if you have time to take a look.
One thing we couldn't pin down: we didn't find the history/rationale for why the gate is > 0 rather than != 0 (nothing beyond the surrounding comment describing the "large struct, positive field offsets" scenario). Since CodeGenPrepare is shared across all targets, do you happen to know if there was a reason to exclude negative offsets there, or was it simply not the motivating case at the time?

https://github.com/llvm/llvm-project/pull/223682


More information about the llvm-commits mailing list