[llvm-branch-commits] [llvm] [CodeGen][AMDGPU] Add opt-in partial SGPR spills (PR #225220)
Yaxun Liu via llvm-branch-commits
llvm-branch-commits at lists.llvm.org
Tue Sep 29 05:19:14 PDT 2026
yxsamliu wrote:
@arsenm This is not a primitive version of #175002: the two PRs solve different issues with different approaches.
This PR fixes stores that overwrite valid data in a shared spill slot. For example:
```text
Slot before: [A, B, C, D]
Split register: [?, B2, ?, ?]
Full-width spill: [?, B2, ?, ?]
This PR's spill: [A, B2, C, D]
```
Here `?` means an undefined word. The full-width store overwrites valid data; we saw this corrupt a buffer descriptor and cause a GPU fault. This PR stores only `B2` at its original offset, keeping the other words intact.
#175002 improves reloads. If an instruction needs only `B`, it loads that word into a 32-bit register and adjusts the uses, avoiding a full 128-bit reload. It does not fix the overwrite above.
This PR also avoids unnecessary stores. Standalone CK measurements showed 4.45% lower runtime with normal compilation.
The two PRs can work together. Either can land first, and the other can adapt any shared code changes.
https://github.com/llvm/llvm-project/pull/225220
More information about the llvm-branch-commits
mailing list