[llvm] [AArch64] Reorder ZPR stack spills to maximize ld1b/st1b pairings (PR #218950)
Kieran B via llvm-commits
llvm-commits at lists.llvm.org
Wed Sep 30 01:47:34 PDT 2026
================
@@ -150,6 +151,74 @@ define void @fbyte(<vscale x 16 x i8> %v){
; PAIR-NEXT: addvl sp, sp, #18
; PAIR-NEXT: ldp x29, x30, [sp], #16 // 16-byte Folded Reload
; PAIR-NEXT: ret
+;
+; SPLIT-PAIR-LABEL: fbyte:
+; SPLIT-PAIR: // %bb.0:
+; SPLIT-PAIR-NEXT: stp x29, x30, [sp, #-16]! // 16-byte Folded Spill
+; SPLIT-PAIR-NEXT: addvl sp, sp, #-2
+; SPLIT-PAIR-NEXT: str p15, [sp, #4, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT: str p14, [sp, #5, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT: str p13, [sp, #6, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT: str p12, [sp, #7, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT: str p11, [sp, #8, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT: str p10, [sp, #9, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT: str p9, [sp, #10, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT: str p8, [sp, #11, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT: str p7, [sp, #12, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT: str p6, [sp, #13, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT: str p5, [sp, #14, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT: str p4, [sp, #15, mul vl] // 2-byte Spill
+; SPLIT-PAIR-NEXT: sub sp, sp, #16
+; SPLIT-PAIR-NEXT: addvl sp, sp, #-16
+; SPLIT-PAIR-NEXT: ptrue pn8.b
+; SPLIT-PAIR-NEXT: st1b { z22.b, z23.b }, pn8, [sp] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT: st1b { z20.b, z21.b }, pn8, [sp, #2, mul vl] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT: st1b { z18.b, z19.b }, pn8, [sp, #4, mul vl] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT: st1b { z16.b, z17.b }, pn8, [sp, #6, mul vl] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT: st1b { z14.b, z15.b }, pn8, [sp, #8, mul vl] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT: st1b { z12.b, z13.b }, pn8, [sp, #10, mul vl] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT: st1b { z10.b, z11.b }, pn8, [sp, #12, mul vl] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT: st1b { z8.b, z9.b }, pn8, [sp, #14, mul vl] // 32-byte Folded Spill
+; SPLIT-PAIR-NEXT: sub sp, sp, #16
+; SPLIT-PAIR-NEXT: .cfi_escape 0x0f, 0x0a, 0x8f, 0x30, 0x92, 0x2e, 0x00, 0x11, 0x90, 0x01, 0x1e, 0x22 // sp + 48 + 144 * VG
+; SPLIT-PAIR-NEXT: .cfi_offset w30, -8
+; SPLIT-PAIR-NEXT: .cfi_offset w29, -16
+; SPLIT-PAIR-NEXT: .cfi_escape 0x10, 0x48, 0x0a, 0x92, 0x2e, 0x00, 0x11, 0x68, 0x1e, 0x22, 0x11, 0x60, 0x22 // $d8 @ cfa - 24 * VG - 32
----------------
kieroxide wrote:
I had a look at this and it seems this has been an issue with pair loads/stores before this change. This change definitely made the issue more common however.
[Here](https://godbolt.org/z/9bb5GzM7W) we can see the issue before this change.
I think this is purely because the pair/quad loads break the order assumption. I will try to fix this in the other [PR](https://github.com/llvm/llvm-project/pull/225100).
https://github.com/llvm/llvm-project/pull/218950
More information about the llvm-commits
mailing list