[llvm] [AMDGPU] Balance VM_CNT histories across branches (PR #221115)
Dan Zimmerman via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 14 09:47:39 PDT 2026
danzimm wrote:
@arsenm and I had a quick chat about this PR on Friday to figure out the right path forward. Firstly, regardless of this change I'll open a new PR to introduce the `buffer_inv` intrinsic. Secondly, Matt pointed out that this optimization can't be correctly implemented at the triton level, since, even with the intrinsic, inserting `buffer_inv` would depend on downstream implementation details.
After the conversation I think I understand the issues with this PR as it stands now and will push a few commits this week to address the issues. I'll also look to gather more perf data from non-triton kernels as well.
On the topic of heuristics: I think Matt suggested (correct me if I'm wrong) that we should initially implement the new optimization without heuristics, assuming the simplest implementation is good until proven otherwise.
https://github.com/llvm/llvm-project/pull/221115
More information about the llvm-commits
mailing list