[llvm] [AArch64] Reorder ZPR stack spills to maximize ld1b/st1b pairings (PR #218950)

Kieran B via llvm-commits llvm-commits at lists.llvm.org
Wed Sep 30 01:47:34 PDT 2026


================
@@ -150,6 +151,74 @@ define void @fbyte(<vscale x 16 x i8> %v){
 ; PAIR-NEXT:    addvl sp, sp, #18
 ; PAIR-NEXT:    ldp x29, x30, [sp], #16 // 16-byte Folded Reload
 ; PAIR-NEXT:    ret
+;
+; SPLIT-PAIR-LABEL: fbyte:
+; SPLIT-PAIR:       // %bb.0:
+; SPLIT-PAIR-NEXT:    stp x29, x30, [sp, #-16]! // 16-byte Folded Spill
+; SPLIT-PAIR-NEXT:    addvl sp, sp, #-2
+; SPLIT-PAIR-NEXT:    str p15, [sp, #4, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT:    str p14, [sp, #5, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT:    str p13, [sp, #6, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT:    str p12, [sp, #7, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT:    str p11, [sp, #8, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT:    str p10, [sp, #9, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT:    str p9, [sp, #10, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT:    str p8, [sp, #11, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT:    str p7, [sp, #12, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT:    str p6, [sp, #13, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT:    str p5, [sp, #14, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT:    str p4, [sp, #15, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT:    sub sp, sp, #16
+; SPLIT-PAIR-NEXT:    addvl sp, sp, #-16
+; SPLIT-PAIR-NEXT:    ptrue pn8.b
+; SPLIT-PAIR-NEXT:    st1b { z22.b, z23.b }, pn8, [sp] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT:    st1b { z20.b, z21.b }, pn8, [sp, #2, mul vl] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT:    st1b { z18.b, z19.b }, pn8, [sp, #4, mul vl] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT:    st1b { z16.b, z17.b }, pn8, [sp, #6, mul vl] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT:    st1b { z14.b, z15.b }, pn8, [sp, #8, mul vl] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT:    st1b { z12.b, z13.b }, pn8, [sp, #10, mul vl] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT:    st1b { z10.b, z11.b }, pn8, [sp, #12, mul vl] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT:    st1b { z8.b, z9.b }, pn8, [sp, #14, mul vl] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT:    sub sp, sp, #16
+; SPLIT-PAIR-NEXT:    .cfi_escape 0x0f, 0x0a, 0x8f, 0x30, 0x92, 0x2e, 0x00, 0x11, 0x90, 0x01, 0x1e, 0x22 // sp + 48 + 144 * VG
+; SPLIT-PAIR-NEXT:    .cfi_offset w30, -8
+; SPLIT-PAIR-NEXT:    .cfi_offset w29, -16
+; SPLIT-PAIR-NEXT:    .cfi_escape 0x10, 0x48, 0x0a, 0x92, 0x2e, 0x00, 0x11, 0x68, 0x1e, 0x22, 0x11, 0x60, 0x22 // $d8 @ cfa - 24 * VG - 32
----------------
kieroxide wrote:

I had a look at this and it seems this has been an issue with pair loads/stores before this change. This change definitely made the issue more common however.

[Here](https://godbolt.org/z/9bb5GzM7W) we can see the issue before this change.

I think this is purely because the pair/quad loads break the order assumption. I will try to fix this in the other [PR](https://github.com/llvm/llvm-project/pull/225100).

https://github.com/llvm/llvm-project/pull/218950


More information about the llvm-commits mailing list