[llvm] [AMDGPU] Balance VM_CNT histories across branches (PR #221115)

Dan Zimmerman via llvm-commits llvm-commits at lists.llvm.org
Mon Sep 14 09:47:39 PDT 2026


danzimm wrote:

@arsenm and I had a quick chat about this PR on Friday to figure out the right path forward. Firstly, regardless of this change I'll open a new PR to introduce the `buffer_inv` intrinsic. Secondly, Matt pointed out that this optimization can't be correctly implemented at the triton level, since, even with the intrinsic, inserting `buffer_inv` would depend on downstream implementation details.

After the conversation I think I understand the issues with this PR as it stands now and will push a few commits this week to address the issues. I'll also look to gather more perf data from non-triton kernels as well. 

On the topic of heuristics: I think Matt suggested (correct me if I'm wrong) that we should initially implement the new optimization without heuristics, assuming the simplest implementation is good until proven otherwise.

https://github.com/llvm/llvm-project/pull/221115


More information about the llvm-commits mailing list