[llvm-branch-commits] [llvm] [AMDGPU][InstCombine] Fold constant add/sub into the dot accumulator (PR #225002)

Harrison Hao via llvm-branch-commits llvm-branch-commits at lists.llvm.org
Mon Sep 21 07:17:41 PDT 2026


harrisonGPU wrote:

> The solution looks alright to me, but I wonder if this could be implemented by hoisting the accumulator instead and then relying on the existing constant folding to bring the constants down. That is:
> 
> ```
> x := amdgcn.{s,u}dot{2,4,8}(a, b, C)
> y := x +/- K
> ->
> x := C +/- K
> y := dot(a, b, x)
> ```
> 
> This would avoid us needing to compute +/- and instead just hoist for all OpCodes we know are well-behaved w.r.t. this, so it would be extendible in the future (if it would ever make sense.)
> 
> I suppose the biggest question (aside from correctness) is whether this would make sense to do always, or only if we know both sides are constants.

Yes, I’m planning to implement this optimization as well. For this first step, I wanted to limit the change to constant folding and keep the patch small and easy to review. As a follow up, I plan to handle non constant cases when the transformation is profitable, such as folding an add into the accumulator when the original accumulator is zero.

https://github.com/llvm/llvm-project/pull/225002


More information about the llvm-branch-commits mailing list