[llvm] [AArch64][NeoverseV1] Load pair immed offset LDPXi, LDNPXi, LDPX<pre|… (PR #221718)

Asher Dobrescu via llvm-commits llvm-commits at lists.llvm.org
Wed Sep 9 03:57:27 PDT 2026


https://github.com/Asher8118 requested changes to this pull request.

>Instruction | uop/Latency/RThroughput
ldp (X-form, offset) | 3/4/1.00
ldpsw (offset) | 3/5/0.50
ldp (X-form, pre-post) | 4/4/1.00
ldpsw (pre-post) | 4/5/0.75

Hello, I am afraid that this is not the correct throughput of these instructions in Neoverse V1. For example, I have measured the below snippet on Neoverse V1:
```asm
.Lloop:
  .rept 15
    ldpsw x0,  x1,  [x20]
    ldpsw x2,  x3,  [x20]
    ldpsw x4,  x5,  [x20]
    ldpsw x6,  x7,  [x20]
    ldpsw x10, x11, [x20]
    ldpsw x12, x13, [x20]
    ldpsw x14, x15, [x20]
    ldpsw x16, x17, [x20]
  .endr
  subs x9, x9, #1
  b.ne .Lloop
```
With 1,000,000 iterations, five runs of `perf stat -r 5 -e cycles,instructions` averaged ≈ 80 million cycles:
```
taskset -c 0 perf stat -r 5 -e cycles,instructions ./ldpsw

 Performance counter stats for './ldpsw' (5 runs):

          80682861      cycles                                                        ( +-  0.01% )
         122764949      instructions              #    1.52  insn per cycle           ( +-  0.00% )

         0.0313851 +- 0.0000101 seconds time elapsed  ( +-  0.03% )
```
With writeback this actually correlates to about 1 instructions per cycle, averaged 120,824,397 cycles for 120 million writeback `LDPSW` instructions, giving ≈ 0.99 instruction/cycle.

Switching to X-reg `LDP`, this PR does not take into account the fact that X-register `LDP` can execute in parallel with other loads and stores. When that is taken into account, a similar unrolled loop tested with `perf` on Neoverse V1 shows a throughput of 1:
```
  .Lloop:
    .rept 15
      ldp x0,  x1,  [x20]
      ldp x2,  x3,  [x20]
      ldp x4,  x5,  [x20]
      ldp x6,  x7,  [x20]
      ldp x10, x11, [x20]
      ldp x12, x13, [x20]
      ldp x14, x15, [x20]
      ldp x16, x17, [x20]
    .endr

    subs x9, x9, #1
    b.ne .Lloop
```
I think the only change that could come from this is changing the throughput of `LDPSW` to 3/2 when it does not have writeback.

https://github.com/llvm/llvm-project/pull/221718


More information about the llvm-commits mailing list