[llvm] [NVPTX] Avoid nonlocal MachineCSE of special register reads. (PR #228303)
Nick Riasanovsky via llvm-commits
llvm-commits at lists.llvm.org
Fri Oct 2 07:05:12 PDT 2026
================
@@ -5778,6 +5779,7 @@ class PTX_READ_SREG_R64<string regname, Intrinsic intop, list<Predicate> Preds=[
[(set i64:$d, (intop))]>,
Requires<Preds>;
+let isAsCheapAsAMove = true in
----------------
njriasan wrote:
Definitely not broad enough performance gains. Our motivating case was a WarpSpecialized attention kernel written in TLX where CSE pulled the ctaid.x computation before the warp specialized region. This is definitely not the average kernel, so I think we need more exhaustive benchmarking.
Could you point me to a potential benchmarking suite? I should be able to run things on A100, H100, GH200, B200, GB200, or GB300 if necessary.
https://github.com/llvm/llvm-project/pull/228303
More information about the llvm-commits
mailing list