[llvm] [AMDGPU] Balance VM_CNT histories across branches (PR #221115)
Dan Zimmerman via llvm-commits
llvm-commits at lists.llvm.org
Fri Sep 11 06:59:14 PDT 2026
danzimm wrote:
@krzysz00 that's exactly right. More generally, if there are branches which contain any wait-required-ops, then the sort of optimization in this PR could prove worthwhile, assuming:
- one of the aforementioned ops appears prior to the branch, and the program wants to wait on *only* these initial ops after the branch (as occurs in a pipelined loop)
- the aforementioned op has an associated nop, akin to `buffer_inv 0`-- this is where hardware specificity enters the conversation.
The idea of a heuristic, and specifically the ideas you mention, seems reasonable. I suspect this would be a great thing to discuss more in depth.
The intrinsic path definitely works. I'm putting together changeset with @CRobeck to make sure the new PR doesn't cause as much friction as this one did (slow moving because of some family illness this week).
On the triton side, after quickly looking I think
- for masked loads I we can insert the intrinsics in a single place within [`MaskedOpsToLLVM`](https://github.com/triton-lang/triton/blob/main/third_party/amd/lib/TritonAMDGPUToLLVM/MaskedOpsToLLVM.cpp#L100)
- for lowering load buffer ops with non-zero `other` we can focus on [`emitBranch`](https://github.com/triton-lang/triton/blob/main/third_party/amd/lib/TritonAMDGPUToLLVM/LoadStoreOpToLLVM.cpp#L188) in `LoadStoreOpToLLVM`, but I'd need to look a little closer to confirm.
How do you think we can investigate if this optimization wants to be its own pass in the backend? As with other comments: this is for my own learning, not as an attempt to sway the status of this PR.
https://github.com/llvm/llvm-project/pull/221115
More information about the llvm-commits
mailing list