[llvm] [AMDGPU] Preserve LDS waits when demoting single-wave workgroup scope to wavefront (PR #207473)
Justin Rosner via llvm-commits
llvm-commits at lists.llvm.org
Wed Jul 8 10:57:45 PDT 2026
justinrosner wrote:
> Do you have an example of an LLVM IR snippet that is wrongly lowered with this demotion? It sounds like a problem for wavefront-scope fences in general if there is a wait missing that is necessary for correctness of accesses within the same wave.
I've attached an example here (GitHub won't allow .ll file types):
[intermittent-failure.txt](https://github.com/user-attachments/files/29812220/intermittent-failure.txt)
The important snippet is:
```LLVM
tail call void @llvm.amdgcn.raw.ptr.buffer.load.async.lds(ptr addrspace(8) nonnull %59, ptr addrspace(3) nonnull getelementptr inbounds nuw (i8, ptr addrspace(3) @global_smem, i32 1024), i32 4, i32 %67, i32 0, i32 0, i32 0), !alias.scope !2, !noalias !7
tail call void @llvm.amdgcn.raw.ptr.buffer.load.async.lds(ptr addrspace(8) nonnull %59, ptr addrspace(3) nonnull getelementptr inbounds nuw (i8, ptr addrspace(3) @global_smem, i32 1280), i32 4, i32 %75, i32 0, i32 0, i32 0), !alias.scope !2, !noalias !7
tail call void @llvm.amdgcn.asyncmark()
...
tail call void @llvm.amdgcn.raw.ptr.buffer.load.async.lds(ptr addrspace(8) nonnull %76, ptr addrspace(3) @global_smem, i32 4, i32 %79, i32 0, i32 0, i32 0), !alias.scope !10, !noalias !11
tail call void @llvm.amdgcn.raw.ptr.buffer.load.async.lds(ptr addrspace(8) nonnull %76, ptr addrspace(3) nonnull getelementptr inbounds nuw (i8, ptr addrspace(3) @global_smem, i32 256), i32 4, i32 %82, i32 0, i32 0, i32 0), !alias.scope !10, !noalias !11
tail call void @llvm.amdgcn.asyncmark()
...
tail call void @llvm.amdgcn.wait.asyncmark(i16 0)
fence syncscope("workgroup") release, !mmra !12
tail call void @llvm.amdgcn.s.barrier()
fence syncscope("workgroup") acquire, !mmra !12
```
https://github.com/llvm/llvm-project/pull/207473
More information about the llvm-commits
mailing list