[llvm] [AMDGPU] Fix missing waitcnt after buffer_wbl2 (PR #178316)
Vigneshwar Jayakumar via llvm-commits
llvm-commits at lists.llvm.org
Thu Jan 29 08:57:10 PST 2026
VigneshwarJ wrote:
> I thought the issue was that we should always include the waitcnt because the wb increments vmcnt. If there is a count the wait shouldn't be optimized away. Do we need to model this (wb increasing vmcnt) anywhere else?
Yes, There are many current wbl2 instruction based tests (25 tests tbp) that fail because just increasing the counter for every WBL2 instruction makes the soft_waitcnt(0) inserted by SIMemoryLegalizer becomes a proper waitcnt at every place where wbl2 is present. I think that is very conservative because not always the wbl2 instruction is going to writeback from vector memory that requires vmcnt wait.
I am afraid doing that would cause any performance regression.
But if its the right way to do it, I will push that change. Right now, I am checking if there are any previous vmem instructions that writes or may store, only then I increment the counter for wbl2.
https://github.com/llvm/llvm-project/pull/178316
More information about the llvm-commits
mailing list