[llvm] [SLP]Skip gathers with erased scalars in the cast context (PR #228395)

Alexey Bataev via llvm-commits llvm-commits at lists.llvm.org
Fri Oct 2 04:13:41 PDT 2026


https://github.com/alexey-bataev created https://github.com/llvm/llvm-project/pull/228395

The cast context of the reduction root is requested after the tree is
vectorized, when the scalars of the gathered root with the vectorized
subvector are already erased and have no operands, causes the crash.

Fixes https://github.com/llvm/llvm-project/pull/224919#issuecomment-5944875384


>From bc00dc6f824edbffff053ce9b279b58782146c5a Mon Sep 17 00:00:00 2001
From: Alexey Bataev <a.bataev at outlook.com>
Date: Fri, 2 Oct 2026 04:13:15 -0700
Subject: [PATCH] =?UTF-8?q?[=F0=9D=98=80=F0=9D=97=BD=F0=9D=97=BF]=20initia?=
 =?UTF-8?q?l=20version?=
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit

Created using spr 1.3.7
---
 .../Transforms/Vectorize/SLPVectorizer.cpp    |  6 ++--
 .../reduction-gathered-root-erased-scalars.ll | 35 +++++++++++++++++++
 2 files changed, 39 insertions(+), 2 deletions(-)
 create mode 100644 llvm/test/Transforms/SLPVectorizer/X86/reduction-gathered-root-erased-scalars.ll

diff --git a/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp b/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
index c78cb9863e8a1f..1f8c90aa6573e0 100644
--- a/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+++ b/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
@@ -16043,8 +16043,10 @@ TTI::CastContextHint BoUpSLP::getCastContextHint(const TreeEntry &TE) const {
       return TTI::CastContextHint::Reversed;
   }
   // A gather of extracted sub-fields inherits the context of the entry
-  // vectorizing the common source scalar, or of the source scalar load.
-  if (TE.isGather())
+  // vectorizing the common source scalar, or of the source scalar load. The
+  // scalars erased by the vectorization have no operands to match.
+  if (TE.isGather() && none_of(make_isa_range<Instruction>(TE.Scalars),
+                               [&](Instruction *I) { return isDeleted(I); }))
     if (std::optional<std::tuple<Value *, unsigned, SmallVector<int>>> Fields =
             matchGatheredExtractedFields(TE.Scalars, *DL)) {
       Value *Src = std::get<0>(*Fields);
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/reduction-gathered-root-erased-scalars.ll b/llvm/test/Transforms/SLPVectorizer/X86/reduction-gathered-root-erased-scalars.ll
new file mode 100644
index 00000000000000..7cd248a54348f2
--- /dev/null
+++ b/llvm/test/Transforms/SLPVectorizer/X86/reduction-gathered-root-erased-scalars.ll
@@ -0,0 +1,35 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
+; RUN: opt -passes=slp-vectorizer -S -mtriple=x86_64-unknown-linux-gnu -mattr=+sse4.1 < %s | FileCheck %s
+
+; The reduction root is a gather with the vectorized subvector. The scalars of
+; the subvector are erased with the tree and must not be matched afterwards.
+define i64 @test(i64 %x, ptr %p) {
+; CHECK-LABEL: define i64 @test(
+; CHECK-SAME: i64 [[X:%.*]], ptr [[P:%.*]]) #[[ATTR0:[0-9]+]] {
+; CHECK-NEXT:    [[SHL:%.*]] = shl i64 [[X]], 1
+; CHECK-NEXT:    [[OR:%.*]] = or i64 0, [[SHL]]
+; CHECK-NEXT:    store i64 [[OR]], ptr [[P]], align 8
+; CHECK-NEXT:    [[TMP1:%.*]] = insertelement <8 x i64> poison, i64 [[SHL]], i64 6
+; CHECK-NEXT:    [[TMP2:%.*]] = insertelement <8 x i64> [[TMP1]], i64 [[OR]], i64 7
+; CHECK-NEXT:    [[TMP3:%.*]] = shufflevector <8 x i64> [[TMP2]], <8 x i64> <i64 0, i64 0, i64 0, i64 0, i64 0, i64 0, i64 undef, i64 undef>, <8 x i32> <i32 8, i32 9, i32 10, i32 11, i32 12, i32 13, i32 6, i32 7>
+; CHECK-NEXT:    [[TMP4:%.*]] = call i64 @llvm.vector.reduce.or.v8i64(<8 x i64> [[TMP3]])
+; CHECK-NEXT:    ret i64 [[TMP4]]
+;
+  %shl = shl i64 %x, 1
+  %or = or i64 0, %shl
+  store i64 %or, ptr %p
+  %and0 = and i64 0, 0
+  %r0 = or i64 %and0, %or
+  %and1 = and i64 0, 0
+  %r1 = or i64 %and1, %r0
+  %and2 = and i64 0, 0
+  %r2 = or i64 %and2, %r1
+  %and3 = and i64 0, 0
+  %r3 = or i64 %and3, %r2
+  %and4 = and i64 0, 0
+  %r4 = or i64 %and4, %r3
+  %and5 = and i64 0, 0
+  %r5 = or i64 %and5, %r4
+  %rs = or i64 %r5, %shl
+  ret i64 %rs
+}



More information about the llvm-commits mailing list