[llvm] [AMDGPU] Balance VM_CNT histories across branches (PR #221115)

Pierre van Houtryve via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 8 01:54:39 PDT 2026


https://github.com/Pierre-vh requested changes to this pull request.

400 lines of only tangentially related code is not a good fit for the InsertWaitcnt pass, it should be in a separate file at the bare minimum, and broken up into multiple PRs (at least one for pre-commit test and one for declaring that `buffer_inv 0` is a "no-op", with associated documentation in AMDGPUUsage).  

I also question whether this is the right approach. There are several red flags here to me:

- Only one specific kernel mentioned as the driver for the change 
- Change is costly, lots of new code, and adds multiple new dependencies to the pass that were previously optimized out.
- Super specific control knob to serve one user because we don't know if this is actually a useful thing in the grand scheme of things.

Can't this be done in the kernel directly ?  Or, can we help differently ? e.g. by adding a `__pad_vmcnt` built-in that the kernel can use and that'd inserts those "no-ops" (which also seem very hacky), but without burdening us with all the analysis ?

Side note, I think this class of issue would likely be fixed by the currently-in-research-phase InsertWaitcnt rewrite. One of the things I am looking at is keeping the entire instruction timeline, and more history across CFG edges. In a future version of InsertWaitcnt, maybe this kind of optimization will make a lot more sense, but it doesn't make sense in the current version.

https://github.com/llvm/llvm-project/pull/221115


More information about the llvm-commits mailing list