[llvm] [AMDGPU][SIMemoryLegalizer] Ensure LDS -> DMA hazards respect fences (PR #220288)
Krzysztof Drewniak via llvm-commits
llvm-commits at lists.llvm.org
Fri Sep 11 14:50:14 PDT 2026
https://github.com/krzysz00 commented:
Ok, I'm not sure _release_mark is the right name, but the name of this operation is rough. But that's not a huge thing.
I do want to check, though, and this is where @ssahasra's earlier notes about fences come in.
I have the sense that, if I do a normal load from LDS and want that to be ordered before my upcoming async operation, it's not unreasonable to have a programming model where I have to do *something* to indicate that I care. I think a workgroup-scoped release fence is a good way to do that, and it's possible that what this patch is revealing is that the workgroup => wavefront demotion was *unsound* in the presence of async operations.
One model of this is that, even if you only have one wavefront in your workgroup, you can think of the async DMA engine as a second participant in your workgroup that'll touch LDS at arbitrary times. So the assumption that, if you've marked your workgroup as having one wave ... it only has one wave for synchronization purposes (aka there's nothing to race on LDS *with*) is false at the level of memory modeling.
I don't think that that realization fundamentally blocks this PR and un-reverting that demotion code ... but it may be worth being clear that a wavefront scope fence *doesn't* synchronize with DMA but that a workgroup scope one *does* ... and then this special flavor of soft wait becomes part of an optimization rather than something we're promising.
(That is, having typed all this out, I think I object to the documentation change semantically)
https://github.com/llvm/llvm-project/pull/220288
More information about the llvm-commits
mailing list