[llvm] [NFC][SLP][AMDGPU] Precommit a cross-block fmul fadd contraction test (PR #226609)

Alexey Bataev via llvm-commits llvm-commits at lists.llvm.org
Fri Sep 25 18:29:50 PDT 2026


================
@@ -0,0 +1,124 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
+; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgpu9.0a-amd-amdhsa < %s | FileCheck %s
+; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgpu9.42-amd-amdhsa < %s | FileCheck %s
+; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgpu12.50-amd-amdhsa < %s | FileCheck %s
+
+; Two products are computed up front and each one feeds a contractable fadd in
+; a different successor. The backend sinks such a single use fmul into the
+; block of its fadd and fuses the pair, and it merges the adjacent scalar
+; loads on its own, so the scalar code is one fma per path. Pairing the
+; multiplies keeps the loads and trades the fma of each path for a packed
+; multiply before the branch plus an add on the path.
+
+define amdgpu_kernel void @cross_block_fmul_lhs(ptr addrspace(1) %p, ptr addrspace(1) %q, ptr addrspace(1) %r, ptr addrspace(1) %r2, float %x, float %y, i1 %c) {
----------------
alexey-bataev wrote:

Drop amdgpu_kernel

https://github.com/llvm/llvm-project/pull/226609


More information about the llvm-commits mailing list