[llvm] [NVPTX] Avoid nonlocal MachineCSE of special register reads. (PR #228303)

Nick Riasanovsky via llvm-commits llvm-commits at lists.llvm.org
Fri Oct 2 07:05:12 PDT 2026


================
@@ -5778,6 +5779,7 @@ class PTX_READ_SREG_R64<string regname, Intrinsic intop, list<Predicate> Preds=[
               [(set i64:$d, (intop))]>,
     Requires<Preds>;
 
+let isAsCheapAsAMove = true in
----------------
njriasan wrote:

Definitely not broad enough performance gains. Our motivating case was a WarpSpecialized attention kernel written in TLX where CSE pulled the ctaid.x computation before the warp specialized region. This is definitely not the average kernel, so I think we need more exhaustive benchmarking.

Could you point me to a potential benchmarking suite? I should be able to run things on A100, H100, GH200, B200, GB200, or GB300 if necessary.

https://github.com/llvm/llvm-project/pull/228303


More information about the llvm-commits mailing list