[llvm] [AMDGPU] Balance VM_CNT histories across branches (PR #221115)

Pierre van Houtryve via llvm-commits llvm-commits at lists.llvm.org
Wed Sep 9 02:57:30 PDT 2026


Pierre-vh wrote:

> Just curious for my own learning (not trying to be nitty)- can you clarify the dependencies you're referring to?

All the new dependencies in the pass pipeline. MDL and loop info can be expensive I believe.

> As (I think) you inferred, this isn't possible from the kernel side today, since there's no way to emit buffer_inv 0 in a way that SIInsertWaitcnts recognizes (it doesn't parse through inline asm to update its scorecard). Introducing something like __pad_vmcnt could do the trick and avoid the analysis this PR introduces.
> 
> Do you think introducing __builtin_amdgcn_pad_vmcnt() / llvm.amdgcn.pad.vmcnt() or __builtin_amdgcn_buffer_inv(arg) / llvm.amdgcn.buffer.inv(arg) is better? I'm leaning towards the latter, but you'd know which is best.
> 
> For the sake of making sure I understand correctly (again not nitting, just for learning): we don't want to put the analysis/optimization in the AMDGPU backend to prevent introducing new complexity, right? If so, it appears that we're suggesting the tradeoff of pushing this logic to triton (or the kernel author in HIP/CK world) is worth it?

It's not that straightforward. A balance has to be found really. This is why this kind of issue has to be carefully researched and discussed with a wider group of people.
Generally it's better if we can do an optimization that works for every target and is safe. It's less of a bother for the users. But if this is a super niche case (which is what I feel like it is, because of the limited examples provided), then adding some specific built-in instead *might* be a better idea as it has less surface area in its implementation and is easier to rip out/deprecate later.

> Just as devil's advocate (again for the sake of learning): I originally aimed at a backend pass because this trick is tied to specifically gfx942/gfx950, and is generic across any frontend. It sounds like this isn't enough to warrant the new analysis, is that right? I guess the risk here is a frontend/kernel author missing this opportunity because they don't know about the trick. I don't have the hardware in front of me, but gfx12+/CDNA5 appears it might have a similar opportunity, so the surface for the trick might expand (both in the sense of more hardware and more counters on CDNA5). As long as this is the right view of the world for you I'm happy to abide 

To be frank, the fact that you refer to it such a complex patch as "just" a trick is a red flag to me. Either this is a proven optimization that can be generalized to a wider set of targets, or this is just a hack for specific Triton kernels on specific targets and we can try to provide the minimum surface area needed to implement it, even though the compiler isn't exactly responsible for that class of problem.

*For example*: We already have intrinsics like `__builtin_amdgcn_buffer_wbinvl1` - I think providing a version for the affected targets so you can do as you please with it would be non-controversial.

> (to show my cards: I want explicit confirmation to refer to this thread in an inevitable conversation I'll have in a triton PR).

Sure, but again, this is something that needs wider discussion with more people (e.g. in a Github issue perhaps, or some email chain). I don't think it's productive to pass around *my* opinion and take it as truth. We need a consensus.

> This is great to hear! Is the rewrite public somewhere that I can track?

Not yet now, it's in the earty design/POC stages.


https://github.com/llvm/llvm-project/pull/221115


More information about the llvm-commits mailing list