[llvm] [X86][CostModel] Fix lane and index accounting in split gather/scatter (PR #220565)
Sumukh J Bharadwaj via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 8 04:20:01 PDT 2026
https://github.com/amd-subharad updated https://github.com/llvm/llvm-project/pull/220565
>From 9eced47eafbd084676b739a3b07b8d6378b52615 Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Wed, 2 Sep 2026 17:11:18 +0530
Subject: [PATCH 1/3] [X86][CostModel] Fix lane and index accounting in split
gather/scatter
When a masked gather or scatter needs more than one register's worth of
pointers, getGSVectorCost recursed on a vector of VF / SplitFactor elements
and multiplied the result. Two things were wrong with that.
The division truncates, so any vector length that is not a multiple of its
split factor lost the remainder: a v9i32 gather was costed as eight lanes,
making its body term as cheap as v8i32's.
The recursive call also recomputed the index width from the shortened vector.
Narrowing a 64-bit index to 32 bits requires a minimum vector length, so an
operation long enough to qualify was still priced with pointer-width indices
once it had been split into parts that were individually too short. Every part
of an operation uses the index width the whole operation qualifies for, so
this overstated the cost of the wide case.
Compute the split cost directly instead: one overhead per part, plus one
scalar memory op for every lane of the original vector, with the index width
chosen once for the operation as a whole.
Both corrections move reported costs, on every subtarget and every cost kind.
On skylake-avx512 a v9i32 gather goes from 12 to 13, and a v24i32 gather
through a dword-index GEP goes from 32 to 28, the latter also dropping from
four parts to two. Lengths that do divide their split factor, and operations
that genuinely need 64-bit indices, are unchanged.
---
llvm/docs/ReleaseNotes.md | 10 ++
.../lib/Target/X86/X86TargetTransformInfo.cpp | 20 ++--
.../X86/masked-gather-scatter-split-cost.ll | 91 +++++++++++++++++++
.../CostModel/X86/masked-intrinsic-cost.ll | 20 +++-
4 files changed, 127 insertions(+), 14 deletions(-)
create mode 100644 llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
diff --git a/llvm/docs/ReleaseNotes.md b/llvm/docs/ReleaseNotes.md
index ef90a1f1f1c41..a6f73348d0d25 100644
--- a/llvm/docs/ReleaseNotes.md
+++ b/llvm/docs/ReleaseNotes.md
@@ -227,6 +227,16 @@ Makes programs 10x faster by doing Special New Thing.
### Changes to the X86 Backend
+* Masked gather/scatter operations that are split across several registers are
+ costed more accurately. Lanes in a vector whose length is not a multiple of
+ its split factor are no longer dropped, so a `v9i32` gather is costed as nine
+ lanes rather than eight. The index width is also chosen for the operation as
+ a whole rather than recomputed for each part, so a wide gather through a
+ 32-bit-index GEP is no longer priced as if it used 64-bit indices. Both
+ corrections apply to every X86 subtarget and every cost kind. Lengths that do
+ divide their split factor, and operations that genuinely need 64-bit indices,
+ are unchanged.
+
### Changes to the OCaml bindings
### Changes to the Python bindings
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
index 8e0cf1fc5a153..fc40b17f91e8e 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
@@ -6659,24 +6659,20 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
std::pair<InstructionCost, MVT> SrcLT = getTypeLegalizationCost(SrcVTy);
InstructionCost::CostType SplitFactor =
std::max(IdxsLT.first, SrcLT.first).getValue();
- if (SplitFactor > 1) {
- // Handle splitting of vector of pointers
- auto *SplitSrcTy =
- FixedVectorType::get(SrcVTy->getScalarType(), VF / SplitFactor);
- return SplitFactor * getGSVectorCost(Opcode, CostKind, SplitSrcTy, Ptr,
- Alignment, AddressSpace);
- }
-
- // If we didn't split, this will be a single gather/scatter instruction.
+ // A vector of pointers that does not fit one register is split into
+ // SplitFactor gather/scatter instructions, each paying the overhead.
if (CostKind == TTI::TCK_CodeSize)
- return 1;
+ return SplitFactor;
// The gather / scatter cost is given by Intel architects. It is a rough
// number since we are looking at one instruction in a time.
const int GSOverhead = (Opcode == Instruction::Load) ? getGatherOverhead()
: getScatterOverhead();
- return GSOverhead + VF * getMemoryOpCost(Opcode, SrcVTy->getScalarType(),
- Alignment, AddressSpace, CostKind);
+ // Charge every lane of the original vector. Descending into VF / SplitFactor
+ // dropped the remainder, so a v9 gather was costed as eight lanes.
+ return SplitFactor * GSOverhead +
+ VF * getMemoryOpCost(Opcode, SrcVTy->getScalarType(), Alignment,
+ AddressSpace, CostKind);
}
/// Calculate the cost of Gather / Scatter operation
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
new file mode 100644
index 0000000000000..26954697c2b3f
--- /dev/null
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
@@ -0,0 +1,91 @@
+; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --version 6
+; Costs for masked gather/scatter operations that need more than one register's
+; worth of pointers, across every cost kind.
+;
+; Two properties are pinned here. A vector length that is not a multiple of its
+; split factor must still be charged for all of its lanes. And the index width
+; is chosen once for the whole operation, so a length wide enough to qualify for
+; narrowing keeps the narrow index in each of its parts, rather than reverting
+; to pointer width because an individual part is too short to qualify.
+
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=haswell | FileCheck %s --check-prefix=AVX2
+
+; A length that divides its split factor: unchanged, and the reference point for
+; the two cases below.
+define <8 x i32> @gather_v8i32(<8 x ptr> %ptrs, <8 x i1> %mask) {
+; SKX-LABEL: 'gather_v8i32'
+; SKX-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <8 x i32> %v
+;
+; AVX2-LABEL: 'gather_v8i32'
+; AVX2-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <8 x i32> %v
+;
+ %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x i32> poison)
+ ret <8 x i32> %v
+}
+
+; Remainder lanes: v9 splits into two parts but is not a multiple of two, so the
+; ninth lane must not be dropped from the body term.
+define <9 x i32> @gather_v9i32(<9 x ptr> %ptrs, <9 x i1> %mask) {
+; SKX-LABEL: 'gather_v9i32'
+; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 4 %ptrs, <9 x i1> %mask, <9 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <9 x i32> %v
+;
+; AVX2-LABEL: 'gather_v9i32'
+; AVX2-NEXT: Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 4 %ptrs, <9 x i1> %mask, <9 x i32> poison)
+; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <9 x i32> %v
+;
+ %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> %ptrs, i32 4, <9 x i1> %mask, <9 x i32> poison)
+ ret <9 x i32> %v
+}
+
+define void @scatter_v9i32(<9 x i32> %val, <9 x ptr> %ptrs, <9 x i1> %mask) {
+; SKX-LABEL: 'scatter_v9i32'
+; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> align 4 %ptrs, <9 x i1> %mask)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+; AVX2-LABEL: 'scatter_v9i32'
+; AVX2-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> align 4 %ptrs, <9 x i1> %mask)
+; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+ call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> %ptrs, i32 4, <9 x i1> %mask)
+ ret void
+}
+
+; Index width across a split: the GEP indices are 32-bit, and v24 is wide enough
+; to qualify for narrowing, so all three parts are priced with a dword index
+; even though a single part on its own would not qualify.
+define <24 x i32> @gather_v24i32_dword_index(ptr %base, <24 x i32> %idx, <24 x i1> %mask) {
+; SKX-LABEL: 'gather_v24i32_dword_index'
+; SKX-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
+; SKX-NEXT: Cost Model: Found costs of RThru:28 CodeSize:2 Lat:100 SizeLat:28 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
+;
+; AVX2-LABEL: 'gather_v24i32_dword_index'
+; AVX2-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
+; AVX2-NEXT: Cost Model: Found costs of RThru:82 CodeSize:106 Lat:178 SizeLat:106 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
+;
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> %mask, <24 x i32> poison)
+ ret <24 x i32> %v
+}
+
+; Control for the above: genuinely 64-bit indices cannot narrow.
+define <24 x i32> @gather_v24i32_qword_index(ptr %base, <24 x i64> %idx, <24 x i1> %mask) {
+; SKX-LABEL: 'gather_v24i32_qword_index'
+; SKX-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+; SKX-NEXT: Cost Model: Found costs of RThru:32 CodeSize:4 Lat:104 SizeLat:32 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
+;
+; AVX2-LABEL: 'gather_v24i32_qword_index'
+; AVX2-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+; AVX2-NEXT: Cost Model: Found costs of RThru:82 CodeSize:106 Lat:178 SizeLat:106 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
+;
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> %mask, <24 x i32> poison)
+ ret <24 x i32> %v
+}
diff --git a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
index ed1b534fac8f8..c02dd6db8b99f 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
@@ -730,7 +730,7 @@ define i32 @masked_store(<1 x i1> %m1, <2 x i1> %m2, <3 x i1> %m3, <4 x i1> %m4,
ret i32 0
}
-define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8, <16 x i1> %m16, <32 x i1> %m32, <64 x i1> %m64) {
+define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8, <9 x i1> %m9, <16 x i1> %m16, <32 x i1> %m32, <64 x i1> %m64) {
; SSE2-LABEL: 'masked_gather'
; SSE2-NEXT: Cost Model: Found costs of RThru:29 CodeSize:37 Lat:61 SizeLat:37 for: %V8F64 = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x double> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:15 CodeSize:19 Lat:31 SizeLat:19 for: %V4F64 = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x double> undef)
@@ -745,6 +745,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:18 SizeLat:12 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:77 CodeSize:93 Lat:141 SizeLat:93 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SSE2-NEXT: Cost Model: Found costs of RThru:49 CodeSize:58 Lat:85 SizeLat:58 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:39 CodeSize:47 Lat:71 SizeLat:47 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:20 CodeSize:24 Lat:36 SizeLat:24 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:18 SizeLat:12 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -772,6 +773,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:49 CodeSize:65 Lat:113 SizeLat:65 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SSE42-NEXT: Cost Model: Found costs of RThru:28 CodeSize:37 Lat:64 SizeLat:37 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:25 CodeSize:33 Lat:57 SizeLat:33 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:13 CodeSize:17 Lat:29 SizeLat:17 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -799,6 +801,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:119 SizeLat:71 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; AVX1-NEXT: Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:30 SizeLat:18 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -826,6 +829,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:119 SizeLat:71 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; AVX2-NEXT: Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:30 SizeLat:18 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -853,6 +857,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SKL-NEXT: Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:17 CodeSize:4 Lat:44 SizeLat:17 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -880,6 +885,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; KNL-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:17 CodeSize:21 Lat:33 SizeLat:21 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -907,6 +913,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -936,6 +943,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
%V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> undef, i32 1, <1 x i1> %m1, <1 x i64> undef)
%V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> undef, i32 1, <16 x i1> %m16, <16 x i32> undef)
+ %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> undef, i32 1, <9 x i1> %m9, <9 x i32> undef)
%V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> undef, i32 1, <8 x i1> %m8, <8 x i32> undef)
%V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> undef, i32 1, <4 x i1> %m4, <4 x i32> undef)
%V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> undef, i32 1, <2 x i1> %m2, <2 x i32> undef)
@@ -953,7 +961,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
ret i32 0
}
-define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8, <16 x i1> %m16, <32 x i1> %m32, <64 x i1> %m64) {
+define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8, <9 x i1> %m9, <16 x i1> %m16, <32 x i1> %m32, <64 x i1> %m64) {
; SSE2-LABEL: 'masked_scatter'
; SSE2-NEXT: Cost Model: Found costs of RThru:29 CodeSize:37 Lat:37 SizeLat:37 for: call void @llvm.masked.scatter.v8f64.v8p0(<8 x double> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SSE2-NEXT: Cost Model: Found costs of RThru:15 CodeSize:19 Lat:19 SizeLat:19 for: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
@@ -968,6 +976,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:12 SizeLat:12 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SSE2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SSE2-NEXT: Cost Model: Found costs of RThru:77 CodeSize:93 Lat:93 SizeLat:93 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SSE2-NEXT: Cost Model: Found costs of RThru:43 CodeSize:52 Lat:52 SizeLat:52 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
; SSE2-NEXT: Cost Model: Found costs of RThru:39 CodeSize:47 Lat:47 SizeLat:47 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SSE2-NEXT: Cost Model: Found costs of RThru:20 CodeSize:24 Lat:24 SizeLat:24 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:12 SizeLat:12 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -995,6 +1004,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SSE42-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SSE42-NEXT: Cost Model: Found costs of RThru:49 CodeSize:65 Lat:65 SizeLat:65 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SSE42-NEXT: Cost Model: Found costs of RThru:28 CodeSize:37 Lat:37 SizeLat:37 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
; SSE42-NEXT: Cost Model: Found costs of RThru:25 CodeSize:33 Lat:33 SizeLat:33 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SSE42-NEXT: Cost Model: Found costs of RThru:13 CodeSize:17 Lat:17 SizeLat:17 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1022,6 +1032,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; AVX1-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; AVX1-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; AVX1-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
; AVX1-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; AVX1-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1049,6 +1060,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; AVX2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; AVX2-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; AVX2-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
; AVX2-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; AVX2-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1076,6 +1088,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SKL-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SKL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SKL-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SKL-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
; SKL-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SKL-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SKL-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1103,6 +1116,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; KNL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; KNL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; KNL-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
; KNL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; KNL-NEXT: Cost Model: Found costs of RThru:17 CodeSize:21 Lat:21 SizeLat:21 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1130,6 +1144,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SKX-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
; SKX-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SKX-NEXT: Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1159,6 +1174,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> undef, i32 1, <1 x i1> %m1)
call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> undef, i32 1, <16 x i1> %m16)
+ call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> undef, i32 1, <9 x i1> %m9)
call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> undef, i32 1, <8 x i1> %m8)
call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> undef, i32 1, <4 x i1> %m4)
call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> undef, i32 1, <2 x i1> %m2)
>From e2d0dcfeea5f857070e75070fb30c14542efaeba Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Thu, 3 Sep 2026 10:39:50 +0530
Subject: [PATCH 2/3] [X86][CostModel] Take the gather/scatter index width from
the pointer's address space
getGSVectorCost started the index width at DL.getPointerSizeInBits(), which
answers for address space 0 regardless of where the gathered pointers actually
live. On x86, address spaces 270 and 271 hold 32-bit pointers, so a vector of
them is half the width of the same number of ordinary pointers.
Costing them at 64 bits made a <16 x ptr addrspace(270)> gather look as though
its indices needed two ZMMs, so it was charged for a split it does not need:
on skylake-avx512 it cost 20 with two parts rather than 18 with one. A GEP
whose indices were 32-bit already reached the right answer through the index
analysis below, so only pointers used directly, without such a GEP, were
mispriced.
Take the width from the pointer operand's own address space, falling back to
the caller's address space when there is no operand to ask. Address space 0 is
unaffected, as is address space 272, whose pointers really are 64-bit.
---
llvm/docs/ReleaseNotes.md | 6 ++
.../lib/Target/X86/X86TargetTransformInfo.cpp | 17 +++--
.../masked-gather-scatter-address-space.ll | 70 +++++++++++++++++++
3 files changed, 88 insertions(+), 5 deletions(-)
create mode 100644 llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
diff --git a/llvm/docs/ReleaseNotes.md b/llvm/docs/ReleaseNotes.md
index a6f73348d0d25..f5d1e54133887 100644
--- a/llvm/docs/ReleaseNotes.md
+++ b/llvm/docs/ReleaseNotes.md
@@ -237,6 +237,12 @@ Makes programs 10x faster by doing Special New Thing.
divide their split factor, and operations that genuinely need 64-bit indices,
are unchanged.
+* Masked gather/scatter costs now take the index width from the address space
+ their pointers live in, instead of assuming the width of address space 0. On
+ x86 this matters for the 32-bit address spaces 270 and 271: a vector of those
+ pointers is half as wide as the same count of ordinary pointers, so it is no
+ longer costed as though it had to be split in two.
+
### Changes to the OCaml bindings
### Changes to the Python bindings
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
index fc40b17f91e8e..db2c0d1849d72 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
@@ -6623,8 +6623,16 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
// operation will use 16 x 64 indices which do not fit in a zmm and needs
// to split. Also check that the base pointer is the same for all lanes,
// and that there's at most one variable index.
- auto getIndexSizeInBits = [](const Value *Ptr, const DataLayout &DL) {
- unsigned IndexSize = DL.getPointerSizeInBits();
+ // The index starts out as wide as the pointers being gathered, which is a
+ // property of the address space they live in. Asking the data layout without
+ // one answers for address space 0, which is not necessarily theirs.
+ unsigned PtrSizeInBits = DL.getPointerSizeInBits(
+ Ptr && Ptr->getType()->isPtrOrPtrVectorTy()
+ ? Ptr->getType()->getScalarType()->getPointerAddressSpace()
+ : AddressSpace);
+
+ auto getIndexSizeInBits = [PtrSizeInBits](const Value *Ptr) {
+ unsigned IndexSize = PtrSizeInBits;
const GetElementPtrInst *GEP = dyn_cast_or_null<GetElementPtrInst>(Ptr);
if (IndexSize < 64 || !GEP)
return IndexSize;
@@ -6649,9 +6657,8 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
// Trying to reduce IndexSize to 32 bits for vector 16.
// By default the IndexSize is equal to pointer size.
- unsigned IndexSize = (ST->hasAVX512() && VF >= 16)
- ? getIndexSizeInBits(Ptr, DL)
- : DL.getPointerSizeInBits();
+ unsigned IndexSize =
+ (ST->hasAVX512() && VF >= 16) ? getIndexSizeInBits(Ptr) : PtrSizeInBits;
auto *IndexVTy = FixedVectorType::get(
IntegerType::get(SrcVTy->getContext(), IndexSize), VF);
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
new file mode 100644
index 0000000000000..0125c7975b7c5
--- /dev/null
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
@@ -0,0 +1,70 @@
+; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --version 6
+; Gathers and scatters over pointers that are not in address space 0. On x86,
+; address spaces 270 and 271 hold 32-bit pointers, so a vector of them is half
+; the width of the same number of address-space-0 pointers, needs no split, and
+; is indexed with a dword rather than a qword index.
+
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX
+
+target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i64:64-i128:128-f80:128-n8:16:32:64-S128"
+
+; A vector of sixteen 32-bit pointers occupies one ZMM, so this is a single
+; dword-indexed gather, not a split qword-indexed pair.
+define <16 x i32> @gather_v16i32_p270(<16 x ptr addrspace(270)> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'gather_v16i32_p270'
+; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+ %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+ ret <16 x i32> %v
+}
+
+; The same shape in address space 0 does need 64-bit indices, and does split.
+define <16 x i32> @gather_v16i32_p0(<16 x ptr> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'gather_v16i32_p0'
+; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+ %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+ ret <16 x i32> %v
+}
+
+; Reaching the same conclusion through a GEP rather than from the pointer type.
+define <16 x i32> @gather_v16i32_p270_gep(ptr addrspace(270) %base, <16 x i32> %idx, <16 x i1> %mask) {
+; SKX-LABEL: 'gather_v16i32_p270_gep'
+; SKX-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr addrspace(270) %base, <16 x i32> %idx
+; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+ %ptrs = getelementptr inbounds i32, ptr addrspace(270) %base, <16 x i32> %idx
+ %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+ ret <16 x i32> %v
+}
+
+define void @scatter_v16i32_p270(<16 x i32> %val, <16 x ptr addrspace(270)> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'scatter_v16i32_p270'
+; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p270(<16 x i32> %val, <16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+ call void @llvm.masked.scatter.v16i32.v16p270(<16 x i32> %val, <16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask)
+ ret void
+}
+
+define void @scatter_v16i32_p0(<16 x i32> %val, <16 x ptr> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'scatter_v16i32_p0'
+; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %val, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+ call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %val, <16 x ptr> %ptrs, i32 4, <16 x i1> %mask)
+ ret void
+}
+
+; Address space 272 holds 64-bit pointers, so it must behave like space 0.
+define <16 x i32> @gather_v16i32_p272(<16 x ptr addrspace(272)> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'gather_v16i32_p272'
+; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p272(<16 x ptr addrspace(272)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+ %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p272(<16 x ptr addrspace(272)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+ ret <16 x i32> %v
+}
>From 19d41d50a2d77d82cf5c3da84565c3c5fb6c2abc Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Tue, 8 Sep 2026 01:59:02 +0530
Subject: [PATCH 3/3] [X86][CostModel] Cost gathers by the instructions CodeGen
emits
The code-size cost of a hardware masked gather/scatter is its part
count, so it can be read directly against CodeGen. It did not match.
Three things made it differ, each with its own boundary.
The part count came from the legalized register count of the widest
operand. Legalization first widens the vector length to a power of two,
so that count exceeds the instructions emitted whenever the length does
not fill each of its parts: a 24-lane qword-indexed gather spreads its
indices over four registers but leaves the fourth with no live lane, and
becomes three instructions. Count instead the lanes one instruction
covers -- the wider of the data and the index bounds it, since a dword
gather indexed by qwords fills a register with indices before it fills
one with data -- and divide the length by that. This affects every
subtarget with a hardware gather or scatter, and it is the term that
sets the code-size cost, so code-size costs move too.
The index width was derived only on AVX-512 at VF >= 16. Outside that
window a gather through a 32-bit-index GEP was priced as though its
indices were pointer-width, tying it with the qword-indexed form that
CodeGen splits into roughly twice as many instructions. Derive the width
whenever a hardware gather/scatter is priced. This affects non-AVX-512
targets that have a gather, such as skylake, and AVX-512 targets at
shorter lengths; operations that genuinely need 64-bit indices are
unchanged.
The mask was not consulted at all. A mask known at compile time says
which parts survive: a part with no live lane is folded away, and an
operation with no live lane at all disappears. A 24-lane qword-indexed
gather with only its first eight lanes live is one instruction, not
three, and only its live lanes are charged for memory accesses. Where
the mask is unknown the two directions differ, so the opcode decides: a
gather drops the tail that widening added, because its result goes
unused, while a scatter keeps it, because it survives as a store under a
zeroed mask that CodeGen still emits. How far that tail reaches follows
the predicate rather than the data: AVX512BW holds the mask in a
k-register as wide as the legalized vector, so the length rounds up to a
power of two, while without it the mask is broken into 16-lane pieces
and the length rounds up only to the next of those. A 48-lane
qword-indexed scatter is eight instructions on skylake-avx512 and six on
knl; the legalized register count says eight for both.
Verified by comparing the code-size cost against the instructions llc
emits, over four data types, both index widths, 21 vector lengths from 2
to 64 and 15 subtargets: 3200 of 3200 hardware cases agree, against 2530
before. Past 64 lanes, and for a compile-time mask carrying undef or
poison lanes, the count is an upper bound rather than exact: CodeGen may
fold parts there in ways that do not follow from the length, and an
undef lane admits either answer. Both overcount rather than under, and
neither shape is one a vectorizer produces. A second sweep holding those
fixed and varying the GEP the index comes from, the shape of a
compile-time mask, the vector-width function attributes and the address
space agrees on 1962 of 2023, against 1216 before; the remainder are
vp.gather/vp.scatter, which never reach this code because X86 routes
only the masked_* intrinsics to it. The costs the sweeps cannot read
against CodeGen were checked against the invariants they must obey
instead: over four cost kinds and five subtargets, no cost is invalid or
negative, an empty mask is free, killing lanes never costs more than
keeping them, and cost never falls as the vector grows.
masked-gather-scatter-part-count.ll pins that agreement, pairing each
cost against an llc instruction count in the same module, for both index
widths on an AVX-512, an AVX2 and a no-AVX512BW target, over gathers,
variable-mask scatters, all-lanes-active scatters, and constant masks
that are a prefix, a suffix, non-contiguous and empty. It also covers
gathering pointers, where the width that sets the lane count comes from
the pointer rather than from an integer type, and which no X86 cost
model test exercised. A throughput run pins the per-lane term that code
size does not expose: the remainder lane of a length that does not fill
its parts, and the lanes a constant mask kills. Undoing any one of the
behaviours above, including the predicate width, makes the test fail.
The address-space test gains an AVX2 run, where the narrower pointer
also changes the cost reached through a GEP, and the address space 271
cases it described but did not cover.
One vectorizer decision changes: in cast-costs.ll a v8i32 scatter with
dword indices was costed as two instructions where llc emits one, and
the loop now vectorizes at VF 8 rather than VF 4.
---
llvm/docs/ReleaseNotes.md | 57 ++-
.../lib/Target/X86/X86TargetTransformInfo.cpp | 117 ++++-
llvm/lib/Target/X86/X86TargetTransformInfo.h | 3 +-
.../masked-gather-scatter-address-space.ll | 62 ++-
.../X86/masked-gather-scatter-part-count.ll | 402 ++++++++++++++++++
.../X86/masked-gather-scatter-split-cost.ll | 24 +-
.../X86/masked-intrinsic-cost-inseltpoison.ll | 6 +-
.../CostModel/X86/masked-intrinsic-cost.ll | 38 +-
.../X86/CostModel/gather-i32-with-i8-index.ll | 6 +-
.../masked-gather-i32-with-i8-index.ll | 6 +-
.../LoopVectorize/X86/cast-costs.ll | 24 +-
11 files changed, 659 insertions(+), 86 deletions(-)
create mode 100644 llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll
diff --git a/llvm/docs/ReleaseNotes.md b/llvm/docs/ReleaseNotes.md
index f5d1e54133887..00309bd43f2b5 100644
--- a/llvm/docs/ReleaseNotes.md
+++ b/llvm/docs/ReleaseNotes.md
@@ -227,21 +227,52 @@ Makes programs 10x faster by doing Special New Thing.
### Changes to the X86 Backend
-* Masked gather/scatter operations that are split across several registers are
- costed more accurately. Lanes in a vector whose length is not a multiple of
- its split factor are no longer dropped, so a `v9i32` gather is costed as nine
- lanes rather than eight. The index width is also chosen for the operation as
- a whole rather than recomputed for each part, so a wide gather through a
- 32-bit-index GEP is no longer priced as if it used 64-bit indices. Both
- corrections apply to every X86 subtarget and every cost kind. Lengths that do
- divide their split factor, and operations that genuinely need 64-bit indices,
+* A masked gather/scatter that is split across several instructions is now
+ costed by how many instructions CodeGen emits, rather than by how many
+ legalized registers its operands occupy. The two differ whenever the vector
+ length is not a multiple of the lanes one instruction covers, because
+ legalization first widens the length to a power of two: a 24-lane
+ qword-indexed gather spreads its indices over four registers but leaves the
+ fourth with no live lane, and is now costed as the three instructions it
+ becomes. This affects every subtarget that has a hardware gather or scatter,
+ and it is the term that sets the code-size cost, so code-size costs move as
+ well. Lengths that already fill each of their parts are unchanged.
+
+* All lanes of the original vector are now charged. Dividing the length by the
+ number of parts discarded the remainder, so a `v9i32` gather was costed as
+ eight lanes. This affects any length that is not a multiple of its part
+ count on any subtarget with a hardware gather or scatter. It does not change
+ code-size costs, which count instructions rather than lanes.
+
+* The index width of a masked gather/scatter is now derived whenever one is
+ priced, instead of only on AVX-512 targets at vector lengths of 16 or more.
+ Outside that window a gather through a 32-bit-index GEP was priced as though
+ its indices were pointer-width, tying it with the qword-indexed form that
+ CodeGen splits into roughly twice as many instructions. Non-AVX-512 targets
+ with a hardware gather, such as `skylake`, and AVX-512 targets at shorter
+ lengths, are the ones affected; operations that genuinely need 64-bit indices
are unchanged.
-* Masked gather/scatter costs now take the index width from the address space
- their pointers live in, instead of assuming the width of address space 0. On
- x86 this matters for the 32-bit address spaces 270 and 271: a vector of those
- pointers is half as wide as the same count of ordinary pointers, so it is no
- longer costed as though it had to be split in two.
+* That index width is also taken from the address space the pointers live in,
+ instead of assuming the width of address space 0. On x86 this matters for the
+ 32-bit address spaces 270 and 271: a vector of those pointers is half as wide
+ as the same count of ordinary pointers, so it is no longer costed as though it
+ had to be split in two.
+
+* A mask known at compile time now decides which parts are counted, so a part
+ holding no live lane is not charged and an operation with no live lane at all
+ is free. A 24-lane qword-indexed gather with only its first eight lanes live
+ becomes one instruction rather than three, and only the live lanes are charged
+ for their memory accesses. Where the mask is not known, a gather still drops
+ the tail that legalization's widening added, because its result goes unused,
+ while a scatter counts it, because it survives as a store under a zeroed mask
+ that CodeGen still emits. How far that tail reaches follows the predicate:
+ with AVX512BW the mask fills a k-register as wide as the legalized vector, and
+ without it the mask is broken into 16-lane pieces and the tail stops at the
+ next of those.
+
+ Because these are the costs the loop vectorizer compares, some loops
+ containing a gather or scatter now vectorize at a wider factor than before.
### Changes to the OCaml bindings
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
index db2c0d1849d72..827d7eec37ed7 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
@@ -6608,12 +6608,37 @@ int X86TTIImpl::getScatterOverhead() const {
return 1024;
}
+// Count the lanes a compile-time constant mask leaves live, and the groups of
+// GroupSize consecutive lanes that hold at least one of them. Returns nullopt
+// when the mask is not known at compile time, which is the common case.
+static std::optional<std::pair<unsigned, unsigned>>
+getActiveGSLanesAndGroups(const Value *Mask, unsigned VF, unsigned GroupSize) {
+ const auto *CMask = dyn_cast_or_null<Constant>(Mask);
+ if (!CMask || !GroupSize)
+ return std::nullopt;
+ unsigned Lanes = 0, Groups = 0;
+ for (unsigned Base = 0; Base < VF; Base += GroupSize) {
+ bool GroupLive = false;
+ for (unsigned I = Base, E = std::min(Base + GroupSize, VF); I != E; ++I) {
+ // An element that cannot be read, or that is not a known-zero bit, has
+ // to be counted as active.
+ const Constant *Elt = CMask->getAggregateElement(I);
+ if (!Elt || !Elt->isNullValue()) {
+ ++Lanes;
+ GroupLive = true;
+ }
+ }
+ if (GroupLive)
+ ++Groups;
+ }
+ return std::make_pair(Lanes, Groups);
+}
+
// Return an average cost of Gather / Scatter instruction, maybe improved later.
-InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
- TTI::TargetCostKind CostKind,
- Type *SrcVTy, const Value *Ptr,
- Align Alignment,
- unsigned AddressSpace) const {
+InstructionCost
+X86TTIImpl::getGSVectorCost(unsigned Opcode, TTI::TargetCostKind CostKind,
+ Type *SrcVTy, const Value *Ptr, Align Alignment,
+ unsigned AddressSpace, const Value *Mask) const {
assert(isa<VectorType>(SrcVTy) && "Unexpected type in getGSVectorCost");
unsigned VF = cast<FixedVectorType>(SrcVTy)->getNumElements();
@@ -6655,31 +6680,71 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
return (unsigned)32;
};
- // Trying to reduce IndexSize to 32 bits for vector 16.
- // By default the IndexSize is equal to pointer size.
- unsigned IndexSize =
- (ST->hasAVX512() && VF >= 16) ? getIndexSizeInBits(Ptr) : PtrSizeInBits;
+ // The index width belongs to the operation as a whole, not to any one part,
+ // so derive it whenever a hardware gather/scatter is being priced. Asking
+ // only for AVX-512 at VF >= 16 left a dword-indexed gather costed as though
+ // its indices were pointer-width everywhere else, so it tied with the
+ // qword-indexed form that CodeGen splits into twice as many instructions.
+ unsigned IndexSize = getIndexSizeInBits(Ptr);
auto *IndexVTy = FixedVectorType::get(
IntegerType::get(SrcVTy->getContext(), IndexSize), VF);
std::pair<InstructionCost, MVT> IdxsLT = getTypeLegalizationCost(IndexVTy);
std::pair<InstructionCost, MVT> SrcLT = getTypeLegalizationCost(SrcVTy);
- InstructionCost::CostType SplitFactor =
- std::max(IdxsLT.first, SrcLT.first).getValue();
- // A vector of pointers that does not fit one register is split into
- // SplitFactor gather/scatter instructions, each paying the overhead.
+
+ // One instruction covers as many lanes as fit in a legal register once both
+ // the data and the indices are held, so the wider of the two bounds the lane
+ // count: a dword gather indexed by qwords fills a register with indices
+ // before it fills one with data. Counting legalized registers instead
+ // overcounts, because legalization first widens to a power of two -- a
+ // 24-lane qword-indexed gather occupies four index registers, but only three
+ // of them hold live lanes and CodeGen emits three instructions.
+ auto LanesOf = [](MVT VT) {
+ return VT.isVector() ? VT.getVectorNumElements() : 1U;
+ };
+ unsigned LanesPerPart =
+ std::max(std::min(LanesOf(IdxsLT.second), LanesOf(SrcLT.second)), 1U);
+
+ // Lanes the operation is declared over, and the parts they occupy.
+ InstructionCost::CostType ChargedLanes = VF;
+ InstructionCost::CostType EmittedParts = divideCeil(VF, LanesPerPart);
+
+ if (auto Active = getActiveGSLanesAndGroups(Mask, VF, LanesPerPart)) {
+ // A compile-time mask says which parts survive: one with no live lane
+ // gathers nothing, and CodeGen folds it away rather than emitting it.
+ ChargedLanes = Active->first;
+ EmittedParts = Active->second;
+ } else if (Opcode != Instruction::Load) {
+ // Without a compile-time mask the tail that legalization's widening added
+ // cannot be proved dead. A gather's tail produces an unused result and is
+ // dropped, but a scatter's survives as a store under a zeroed mask, which
+ // is still an instruction. How far the widening reaches is set by the
+ // predicate: AVX512BW holds it in a k-register as wide as the legalized
+ // vector, so the length rounds up to a power of two, while without it the
+ // mask is broken into 16-lane pieces and only rounds up to one of those.
+ unsigned WidenedLanes = PowerOf2Ceil(VF);
+ if (!ST->hasBWI())
+ WidenedLanes = std::min(WidenedLanes, alignTo(VF, 16u));
+ EmittedParts = divideCeil(WidenedLanes, LanesPerPart);
+ }
+
+ // Nothing survives, so no instruction is emitted.
+ if (EmittedParts == 0)
+ return 0;
+
+ // Each emitted instruction pays the overhead once.
if (CostKind == TTI::TCK_CodeSize)
- return SplitFactor;
+ return EmittedParts;
// The gather / scatter cost is given by Intel architects. It is a rough
// number since we are looking at one instruction in a time.
const int GSOverhead = (Opcode == Instruction::Load) ? getGatherOverhead()
: getScatterOverhead();
- // Charge every lane of the original vector. Descending into VF / SplitFactor
- // dropped the remainder, so a v9 gather was costed as eight lanes.
- return SplitFactor * GSOverhead +
- VF * getMemoryOpCost(Opcode, SrcVTy->getScalarType(), Alignment,
- AddressSpace, CostKind);
+ // Charge every lane that is actually accessed. Dividing the length by the
+ // part count dropped the remainder, so a v9 gather was costed as eight lanes.
+ return EmittedParts * GSOverhead +
+ ChargedLanes * getMemoryOpCost(Opcode, SrcVTy->getScalarType(),
+ Alignment, AddressSpace, CostKind);
}
/// Calculate the cost of Gather / Scatter operation
@@ -6704,8 +6769,18 @@ X86TTIImpl::getGatherScatterOpCost(const MemIntrinsicCostAttributes &MICA,
assert(SrcVTy->isVectorTy() && "Unexpected data type for Gather/Scatter");
unsigned AddressSpace = MICA.getAddressSpace();
- return getGSVectorCost(Opcode, CostKind, SrcVTy, Ptr, Alignment,
- AddressSpace);
+
+ // A mask known at compile time tells us which parts CodeGen keeps. Read it
+ // only from a real gather/scatter call: the context instruction can also be a
+ // plain load or store being considered for the transform, whose operands mean
+ // something else entirely.
+ const Value *Mask = nullptr;
+ if (const auto *II = dyn_cast_or_null<IntrinsicInst>(MICA.getInst()))
+ if (II->getIntrinsicID() == MICA.getID())
+ Mask = II->getArgOperand(IsLoad ? 1 : 2);
+
+ return getGSVectorCost(Opcode, CostKind, SrcVTy, Ptr, Alignment, AddressSpace,
+ Mask);
}
bool X86TTIImpl::isLSRCostLess(const TargetTransformInfo::LSRCost &C1,
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.h b/llvm/lib/Target/X86/X86TargetTransformInfo.h
index 422e4aad2316c..a130835173b5d 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.h
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.h
@@ -271,7 +271,8 @@ class X86TTIImpl final : public BasicTTIImplBase<X86TTIImpl> {
bool supportsGather() const;
InstructionCost getGSVectorCost(unsigned Opcode, TTI::TargetCostKind CostKind,
Type *DataTy, const Value *Ptr,
- Align Alignment, unsigned AddressSpace) const;
+ Align Alignment, unsigned AddressSpace,
+ const Value *Mask = nullptr) const;
int getGatherOverhead() const;
int getScatterOverhead() const;
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
index 0125c7975b7c5..2d904412f5d52 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
@@ -1,10 +1,16 @@
; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --version 6
; Gathers and scatters over pointers that are not in address space 0. On x86,
; address spaces 270 and 271 hold 32-bit pointers, so a vector of them is half
-; the width of the same number of address-space-0 pointers, needs no split, and
-; is indexed with a dword rather than a qword index.
+; the width of the same number of address-space-0 pointers, needs fewer parts,
+; and is indexed with a dword rather than a qword index.
+;
+; Both targets are covered because the two differ in what they expose. On the
+; AVX-512 target the narrower pointer removes a split that address space 0 still
+; needs, while on the AVX2 target it also changes the cost reached through a GEP,
+; whose index width is derived from the pointer's own address space.
; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake | FileCheck %s --check-prefix=SKL
target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i64:64-i128:128-f80:128-n8:16:32:64-S128"
@@ -14,6 +20,10 @@ define <16 x i32> @gather_v16i32_p270(<16 x ptr addrspace(270)> %ptrs, <16 x i1>
; SKX-LABEL: 'gather_v16i32_p270'
; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+; SKL-LABEL: 'gather_v16i32_p270'
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
;
%v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
ret <16 x i32> %v
@@ -24,6 +34,10 @@ define <16 x i32> @gather_v16i32_p0(<16 x ptr> %ptrs, <16 x i1> %mask) {
; SKX-LABEL: 'gather_v16i32_p0'
; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+; SKL-LABEL: 'gather_v16i32_p0'
+; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
;
%v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
ret <16 x i32> %v
@@ -35,6 +49,11 @@ define <16 x i32> @gather_v16i32_p270_gep(ptr addrspace(270) %base, <16 x i32> %
; SKX-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr addrspace(270) %base, <16 x i32> %idx
; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+; SKL-LABEL: 'gather_v16i32_p270_gep'
+; SKL-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr addrspace(270) %base, <16 x i32> %idx
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
;
%ptrs = getelementptr inbounds i32, ptr addrspace(270) %base, <16 x i32> %idx
%v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
@@ -45,6 +64,10 @@ define void @scatter_v16i32_p270(<16 x i32> %val, <16 x ptr addrspace(270)> %ptr
; SKX-LABEL: 'scatter_v16i32_p270'
; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p270(<16 x i32> %val, <16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask)
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+; SKL-LABEL: 'scatter_v16i32_p270'
+; SKL-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p270(<16 x i32> %val, <16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
;
call void @llvm.masked.scatter.v16i32.v16p270(<16 x i32> %val, <16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask)
ret void
@@ -54,16 +77,51 @@ define void @scatter_v16i32_p0(<16 x i32> %val, <16 x ptr> %ptrs, <16 x i1> %mas
; SKX-LABEL: 'scatter_v16i32_p0'
; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %val, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+; SKL-LABEL: 'scatter_v16i32_p0'
+; SKL-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %val, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
;
call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %val, <16 x ptr> %ptrs, i32 4, <16 x i1> %mask)
ret void
}
+; Address space 271 holds 32-bit pointers too, and must be treated like 270.
+define <16 x i32> @gather_v16i32_p271(<16 x ptr addrspace(271)> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'gather_v16i32_p271'
+; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p271(<16 x ptr addrspace(271)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+; SKL-LABEL: 'gather_v16i32_p271'
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p271(<16 x ptr addrspace(271)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+ %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p271(<16 x ptr addrspace(271)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+ ret <16 x i32> %v
+}
+
+define void @scatter_v16i32_p271(<16 x i32> %val, <16 x ptr addrspace(271)> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'scatter_v16i32_p271'
+; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p271(<16 x i32> %val, <16 x ptr addrspace(271)> align 4 %ptrs, <16 x i1> %mask)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+; SKL-LABEL: 'scatter_v16i32_p271'
+; SKL-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p271(<16 x i32> %val, <16 x ptr addrspace(271)> align 4 %ptrs, <16 x i1> %mask)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+ call void @llvm.masked.scatter.v16i32.v16p271(<16 x i32> %val, <16 x ptr addrspace(271)> %ptrs, i32 4, <16 x i1> %mask)
+ ret void
+}
+
; Address space 272 holds 64-bit pointers, so it must behave like space 0.
define <16 x i32> @gather_v16i32_p272(<16 x ptr addrspace(272)> %ptrs, <16 x i1> %mask) {
; SKX-LABEL: 'gather_v16i32_p272'
; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p272(<16 x ptr addrspace(272)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+; SKL-LABEL: 'gather_v16i32_p272'
+; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p272(<16 x ptr addrspace(272)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
;
%v = call <16 x i32> @llvm.masked.gather.v16i32.v16p272(<16 x ptr addrspace(272)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
ret <16 x i32> %v
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll
new file mode 100644
index 0000000000000..6c30f32f1c97e
--- /dev/null
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll
@@ -0,0 +1,402 @@
+; Pins the cost model's part count against the number of gather/scatter
+; instructions CodeGen actually emits, for both index widths on an AVX-512 and
+; an AVX2 target.
+;
+; The pairing is the point of the test: the code-size cost of a hardware
+; gather/scatter is its part count, so every cost check below has a matching
+; instruction count taken from the same module. Three properties would
+; otherwise drift unnoticed. A part count taken from the legalized register
+; count overstates the work when legalization widens to a power of two; an
+; index width derived only for wide AVX-512 vectors leaves a dword-indexed
+; operation tied with the qword-indexed form that CodeGen splits into more
+; instructions; and the tail a scatter still emits under a zeroed mask reaches
+; only as far as the predicate widens, which depends on AVX512BW.
+;
+; The throughput run pins the other half of the cost, the per-lane term, which
+; code size does not expose: it is charged for the lanes that are really
+; accessed, so it must follow the remainder lane of a length that does not fill
+; its parts, and must drop the lanes a compile-time mask kills.
+;
+; These checks are maintained by hand rather than by
+; update_analyze_test_checks.py, which does not know about the paired llc runs.
+
+; RUN: opt < %s -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" -cost-kind=code-size -disable-output -mcpu=skylake-avx512 2>&1 | FileCheck %s --check-prefix=SKX-COST
+; RUN: opt < %s -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" -cost-kind=throughput -disable-output -mcpu=skylake-avx512 2>&1 | FileCheck %s --check-prefix=SKX-TPUT
+; RUN: llc < %s -mtriple=x86_64-unknown-linux-gnu -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX-ASM
+; RUN: opt < %s -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" -cost-kind=code-size -disable-output -mcpu=skylake 2>&1 | FileCheck %s --check-prefix=AVX2-COST
+; RUN: llc < %s -mtriple=x86_64-unknown-linux-gnu -mcpu=skylake | FileCheck %s --check-prefix=AVX2-ASM
+; RUN: opt < %s -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" -cost-kind=code-size -disable-output -mcpu=knl 2>&1 | FileCheck %s --check-prefix=KNL-COST
+; RUN: llc < %s -mtriple=x86_64-unknown-linux-gnu -mcpu=knl | FileCheck %s --check-prefix=KNL-ASM
+
+; A length that fills its single part exactly, to sit against the nine-lane form
+; below: both are one instruction, but the ninth lane is still accessed and has
+; to be charged, so the throughput costs must differ by exactly one lane.
+; SKX-COST-LABEL: 'gather_v8i32_dword_index'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v8i32_dword_index'
+; SKX-TPUT: cost of 10 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v8i32_dword_index:
+; SKX-ASM-COUNT-1: vpgatherdd
+; SKX-ASM-NOT: vpgather
+define <8 x i32> @gather_v8i32_dword_index(ptr %base, <8 x i32> %idx, <8 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <8 x i32> %idx
+ %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x i32> poison)
+ ret <8 x i32> %v
+}
+
+; A vector length that is not a multiple of its part count: the remainder lane
+; still needs a part of its own once the index no longer fits alongside the rest.
+; SKX-COST-LABEL: 'gather_v9i32_dword_index'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v9i32_dword_index'
+; SKX-TPUT: cost of 11 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v9i32_dword_index:
+; SKX-ASM-COUNT-1: vpgatherdd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v9i32_dword_index'
+; AVX2-COST: cost of 2 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v9i32_dword_index:
+; AVX2-ASM-COUNT-2: vpgatherdd
+; AVX2-ASM-NOT: vpgather
+define <9 x i32> @gather_v9i32_dword_index(ptr %base, <9 x i32> %idx, <9 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <9 x i32> %idx
+ %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> %ptrs, i32 4, <9 x i1> %mask, <9 x i32> poison)
+ ret <9 x i32> %v
+}
+
+; SKX-COST-LABEL: 'gather_v9i32_qword_index'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v9i32_qword_index:
+; SKX-ASM-COUNT-2: vpgatherqd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v9i32_qword_index'
+; AVX2-COST: cost of 3 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v9i32_qword_index:
+; AVX2-ASM-COUNT-3: vpgatherqd
+; AVX2-ASM-NOT: vpgather
+define <9 x i32> @gather_v9i32_qword_index(ptr %base, <9 x i64> %idx, <9 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <9 x i64> %idx
+ %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> %ptrs, i32 4, <9 x i1> %mask, <9 x i32> poison)
+ ret <9 x i32> %v
+}
+
+; A length below the width at which the index used to be examined, so both index
+; widths would otherwise be priced the same on either target.
+; SKX-COST-LABEL: 'gather_v17i32_dword_index'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v17i32_dword_index:
+; SKX-ASM-COUNT-2: vpgatherdd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v17i32_dword_index'
+; AVX2-COST: cost of 3 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v17i32_dword_index:
+; AVX2-ASM-COUNT-3: vpgatherdd
+; AVX2-ASM-NOT: vpgather
+define <17 x i32> @gather_v17i32_dword_index(ptr %base, <17 x i32> %idx, <17 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <17 x i32> %idx
+ %v = call <17 x i32> @llvm.masked.gather.v17i32.v17p0(<17 x ptr> %ptrs, i32 4, <17 x i1> %mask, <17 x i32> poison)
+ ret <17 x i32> %v
+}
+
+; SKX-COST-LABEL: 'gather_v17i32_qword_index'
+; SKX-COST: cost of 3 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v17i32_qword_index:
+; SKX-ASM-COUNT-3: vpgatherqd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v17i32_qword_index'
+; AVX2-COST: cost of 5 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v17i32_qword_index:
+; AVX2-ASM-COUNT-5: vpgatherqd
+; AVX2-ASM-NOT: vpgather
+define <17 x i32> @gather_v17i32_qword_index(ptr %base, <17 x i64> %idx, <17 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <17 x i64> %idx
+ %v = call <17 x i32> @llvm.masked.gather.v17i32.v17p0(<17 x ptr> %ptrs, i32 4, <17 x i1> %mask, <17 x i32> poison)
+ ret <17 x i32> %v
+}
+
+; A length whose qword indices occupy four legal registers but fill only three
+; of them with live lanes.
+; SKX-COST-LABEL: 'gather_v24i32_dword_index'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_dword_index:
+; SKX-ASM-COUNT-2: vpgatherdd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v24i32_dword_index'
+; AVX2-COST: cost of 3 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v24i32_dword_index:
+; AVX2-ASM-COUNT-3: vpgatherdd
+; AVX2-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_dword_index(ptr %base, <24 x i32> %idx, <24 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> %mask, <24 x i32> poison)
+ ret <24 x i32> %v
+}
+
+; SKX-COST-LABEL: 'gather_v24i32_qword_index'
+; SKX-COST: cost of 3 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v24i32_qword_index'
+; SKX-TPUT: cost of 30 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_qword_index:
+; SKX-ASM-COUNT-3: vpgatherqd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v24i32_qword_index'
+; AVX2-COST: cost of 6 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v24i32_qword_index:
+; AVX2-ASM-COUNT-6: vpgatherqd
+; AVX2-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_qword_index(ptr %base, <24 x i64> %idx, <24 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> %mask, <24 x i32> poison)
+ ret <24 x i32> %v
+}
+
+; Scatters with a variable mask. AVX2 has no hardware scatter, so its cost comes
+; from scalarizing instead of from a part count, and it emits no scatter at all.
+; SKX-COST-LABEL: 'scatter_v9i32_dword_index_var_mask'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v9i32_dword_index_var_mask:
+; SKX-ASM-COUNT-1: vpscatterdd
+; SKX-ASM-NOT: vpscatter
+; AVX2-ASM-LABEL: scatter_v9i32_dword_index_var_mask:
+; AVX2-ASM-NOT: vpscatter
+define void @scatter_v9i32_dword_index_var_mask(ptr %base, <9 x i32> %idx, <9 x i32> %val, <9 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <9 x i32> %idx
+ call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> %ptrs, i32 4, <9 x i1> %mask)
+ ret void
+}
+
+; SKX-COST-LABEL: 'scatter_v9i32_qword_index_var_mask'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v9i32_qword_index_var_mask:
+; SKX-ASM-COUNT-2: vpscatterqd
+; SKX-ASM-NOT: vpscatter
+; AVX2-ASM-LABEL: scatter_v9i32_qword_index_var_mask:
+; AVX2-ASM-NOT: vpscatter
+define void @scatter_v9i32_qword_index_var_mask(ptr %base, <9 x i64> %idx, <9 x i32> %val, <9 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <9 x i64> %idx
+ call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> %ptrs, i32 4, <9 x i1> %mask)
+ ret void
+}
+
+; Scatters with every lane active: an all-ones mask leaves the part count alone.
+; SKX-COST-LABEL: 'scatter_v24i32_dword_index_all_active'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v24i32_dword_index_all_active:
+; SKX-ASM-COUNT-2: vpscatterdd
+; SKX-ASM-NOT: vpscatter
+; AVX2-ASM-LABEL: scatter_v24i32_dword_index_all_active:
+; AVX2-ASM-NOT: vpscatter
+define void @scatter_v24i32_dword_index_all_active(ptr %base, <24 x i32> %idx, <24 x i32> %val) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
+ call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %val, <24 x ptr> %ptrs, i32 4, <24 x i1> splat (i1 true))
+ ret void
+}
+
+; SKX-COST-LABEL: 'scatter_v24i32_qword_index_all_active'
+; SKX-COST: cost of 3 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v24i32_qword_index_all_active:
+; SKX-ASM-COUNT-3: vpscatterqd
+; SKX-ASM-NOT: vpscatter
+; AVX2-ASM-LABEL: scatter_v24i32_qword_index_all_active:
+; AVX2-ASM-NOT: vpscatter
+define void @scatter_v24i32_qword_index_all_active(ptr %base, <24 x i64> %idx, <24 x i32> %val) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %val, <24 x ptr> %ptrs, i32 4, <24 x i1> splat (i1 true))
+ ret void
+}
+
+; A variable mask on a length that legalization widens past a part boundary. The
+; widened tail cannot be proved dead, and it survives as a store under a zeroed
+; mask, so CodeGen emits a fourth instruction where the all-ones form above
+; emits three. That store is counted, since it is still an instruction.
+; SKX-COST-LABEL: 'scatter_v24i32_qword_index_var_mask'
+; SKX-COST: cost of 4 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v24i32_qword_index_var_mask:
+; SKX-ASM-COUNT-3: vpscatterqd
+; SKX-ASM: kxor
+; SKX-ASM-COUNT-1: vpscatterqd
+; SKX-ASM-NOT: vpscatter
+; AVX2-ASM-LABEL: scatter_v24i32_qword_index_var_mask:
+; AVX2-ASM-NOT: vpscatter
+define void @scatter_v24i32_qword_index_var_mask(ptr %base, <24 x i64> %idx, <24 x i32> %val, <24 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %val, <24 x ptr> %ptrs, i32 4, <24 x i1> %mask)
+ ret void
+}
+
+; A compile-time mask says which parts survive. With only the first eight lanes
+; live, two of these three parts hold nothing and are folded away, leaving one
+; instruction rather than the three the declared length would suggest.
+; The dead lanes are not charged either: against the 30 the same shape costs
+; with an unknown mask above, this is the cost of the eight lanes it reaches.
+; SKX-COST-LABEL: 'gather_v24i32_qword_index_first8_mask'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v24i32_qword_index_first8_mask'
+; SKX-TPUT: cost of 10 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_qword_index_first8_mask:
+; SKX-ASM-COUNT-1: vpgatherqd
+; SKX-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_qword_index_first8_mask(ptr %base, <24 x i64> %idx) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> <i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false>, <24 x i32> poison)
+ ret <24 x i32> %v
+}
+
+; The same for a scatter, where the dead parts would otherwise be the zeroed
+; stores counted above.
+; SKX-COST-LABEL: 'scatter_v24i32_qword_index_first8_mask'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v24i32_qword_index_first8_mask:
+; SKX-ASM-COUNT-1: vpscatterqd
+; SKX-ASM-NOT: vpscatter
+define void @scatter_v24i32_qword_index_first8_mask(ptr %base, <24 x i64> %idx, <24 x i32> %val) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %val, <24 x ptr> %ptrs, i32 4, <24 x i1> <i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false>)
+ ret void
+}
+
+; A mask with no live lane leaves nothing to do, and CodeGen emits no gather at
+; all, so the operation is free rather than costed as three parts.
+; SKX-COST-LABEL: 'gather_v24i32_qword_index_zero_mask'
+; SKX-COST: cost of 0 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v24i32_qword_index_zero_mask'
+; SKX-TPUT: cost of 0 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_qword_index_zero_mask:
+; SKX-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_qword_index_zero_mask(ptr %base, <24 x i64> %idx) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> zeroinitializer, <24 x i32> poison)
+ ret <24 x i32> %v
+}
+
+; The same for a scatter, which unlike a gather would otherwise still be charged
+; for the tail it emits under a zeroed mask.
+; SKX-COST-LABEL: 'scatter_v24i32_qword_index_zero_mask'
+; SKX-COST: cost of 0 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v24i32_qword_index_zero_mask:
+; SKX-ASM-NOT: vpscatter
+define void @scatter_v24i32_qword_index_zero_mask(ptr %base, <24 x i64> %idx, <24 x i32> %val) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %val, <24 x ptr> %ptrs, i32 4, <24 x i1> zeroinitializer)
+ ret void
+}
+
+; The live lanes need not be a prefix. Here they sit in the first and last part
+; with a dead one between, so it is the parts holding a live lane that are
+; counted, not the number of live lanes rounded up to a part.
+; SKX-COST-LABEL: 'gather_v24i32_qword_index_gapped_mask'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v24i32_qword_index_gapped_mask'
+; SKX-TPUT: cost of 6 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_qword_index_gapped_mask:
+; SKX-ASM-COUNT-2: vpgatherqd
+; SKX-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_qword_index_gapped_mask(ptr %base, <24 x i64> %idx) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> <i1 true, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 true, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false>, <24 x i32> poison)
+ ret <24 x i32> %v
+}
+
+; Live lanes confined to the last part, the mirror of the first-eight case, so
+; that a part is not counted merely because earlier parts were.
+; SKX-COST-LABEL: 'gather_v24i32_qword_index_last8_mask'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_qword_index_last8_mask:
+; SKX-ASM-COUNT-1: vpgatherqd
+; SKX-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_qword_index_last8_mask(ptr %base, <24 x i64> %idx) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> <i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true>, <24 x i32> poison)
+ ret <24 x i32> %v
+}
+
+; How far that widened tail reaches is set by the predicate rather than by the
+; data. AVX512BW holds the mask in a k-register as wide as the legalized
+; vector, so a 48-lane length rounds up to 64 and eight parts are emitted;
+; without it the mask is broken into 16-lane pieces, the length rounds up only
+; to 48, and six are. Counting legal registers would give eight on both.
+; SKX-COST-LABEL: 'scatter_v48i32_qword_index_var_mask'
+; SKX-COST: cost of 8 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v48i32_qword_index_var_mask:
+; SKX-ASM-COUNT-8: vpscatterqd
+; SKX-ASM-NOT: vpscatter
+; KNL-COST-LABEL: 'scatter_v48i32_qword_index_var_mask'
+; KNL-COST: cost of 6 for instruction: {{.*}}masked.scatter
+; KNL-ASM-LABEL: scatter_v48i32_qword_index_var_mask:
+; KNL-ASM-COUNT-6: vpscatterqd
+; KNL-ASM-NOT: vpscatter
+define void @scatter_v48i32_qword_index_var_mask(ptr %base, <48 x i64> %idx, <48 x i32> %val, <48 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <48 x i64> %idx
+ call void @llvm.masked.scatter.v48i32.v48p0(<48 x i32> %val, <48 x ptr> %ptrs, i32 4, <48 x i1> %mask)
+ ret void
+}
+
+; The same split with dword indices, where a part covers sixteen lanes instead
+; of eight: four parts with the wide predicate, three without it.
+; SKX-COST-LABEL: 'scatter_v48i32_dword_index_var_mask'
+; SKX-COST: cost of 4 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v48i32_dword_index_var_mask:
+; SKX-ASM-COUNT-4: vpscatterdd
+; SKX-ASM-NOT: vpscatter
+; KNL-COST-LABEL: 'scatter_v48i32_dword_index_var_mask'
+; KNL-COST: cost of 3 for instruction: {{.*}}masked.scatter
+; KNL-ASM-LABEL: scatter_v48i32_dword_index_var_mask:
+; KNL-ASM-COUNT-3: vpscatterdd
+; KNL-ASM-NOT: vpscatter
+define void @scatter_v48i32_dword_index_var_mask(ptr %base, <48 x i32> %idx, <48 x i32> %val, <48 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <48 x i32> %idx
+ call void @llvm.masked.scatter.v48i32.v48p0(<48 x i32> %val, <48 x ptr> %ptrs, i32 4, <48 x i1> %mask)
+ ret void
+}
+
+; Pointers are gathered as often as integers are, and the element size that
+; decides how many lanes a part holds has to come from the pointer's own width
+; rather than from an integer type. Eight pointer-sized lanes fill one AVX-512
+; part and two AVX2 parts.
+; SKX-COST-LABEL: 'gather_v8ptr_qword_index'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v8ptr_qword_index'
+; SKX-TPUT: cost of 10 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v8ptr_qword_index:
+; SKX-ASM-COUNT-1: vpgatherqq
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v8ptr_qword_index'
+; AVX2-COST: cost of 2 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v8ptr_qword_index:
+; AVX2-ASM-COUNT-2: vpgatherqq
+; AVX2-ASM-NOT: vpgatherqq
+define <8 x ptr> @gather_v8ptr_qword_index(ptr %base, <8 x i64> %idx, <8 x i1> %mask) {
+ %ptrs = getelementptr inbounds ptr, ptr %base, <8 x i64> %idx
+ %r = call <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr> %ptrs, i32 8, <8 x i1> %mask, <8 x ptr> poison)
+ ret <8 x ptr> %r
+}
+
+; The same gather with only its first two lanes left live by a compile-time
+; mask. The second AVX2 part holds no live lane and is not emitted, and both
+; targets charge two lanes rather than eight.
+; SKX-COST-LABEL: 'gather_v8ptr_qword_index_first2_mask'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v8ptr_qword_index_first2_mask'
+; SKX-TPUT: cost of 4 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v8ptr_qword_index_first2_mask:
+; SKX-ASM-COUNT-1: vpgatherqq
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v8ptr_qword_index_first2_mask'
+; AVX2-COST: cost of 1 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v8ptr_qword_index_first2_mask:
+; AVX2-ASM-COUNT-1: vpgatherqq
+; AVX2-ASM-NOT: vpgatherqq
+define <8 x ptr> @gather_v8ptr_qword_index_first2_mask(ptr %base, <8 x i64> %idx) {
+ %ptrs = getelementptr inbounds ptr, ptr %base, <8 x i64> %idx
+ %r = call <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr> %ptrs, i32 8, <8 x i1> <i1 1, i1 1, i1 0, i1 0, i1 0, i1 0, i1 0, i1 0>, <8 x ptr> poison)
+ ret <8 x ptr> %r
+}
+
+declare <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr>, i32, <8 x i1>, <8 x ptr>)
+declare <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr>, i32, <8 x i1>, <8 x i32>)
+declare <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr>, i32, <9 x i1>, <9 x i32>)
+declare <17 x i32> @llvm.masked.gather.v17i32.v17p0(<17 x ptr>, i32, <17 x i1>, <17 x i32>)
+declare <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr>, i32, <24 x i1>, <24 x i32>)
+declare void @llvm.masked.scatter.v9i32.v9p0(<9 x i32>, <9 x ptr>, i32, <9 x i1>)
+declare void @llvm.masked.scatter.v24i32.v24p0(<24 x i32>, <24 x ptr>, i32, <24 x i1>)
+declare void @llvm.masked.scatter.v48i32.v48p0(<48 x i32>, <48 x ptr>, i32, <48 x i1>)
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
index 26954697c2b3f..f98799b982839 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
@@ -3,10 +3,13 @@
; worth of pointers, across every cost kind.
;
; Two properties are pinned here. A vector length that is not a multiple of its
-; split factor must still be charged for all of its lanes. And the index width
-; is chosen once for the whole operation, so a length wide enough to qualify for
-; narrowing keeps the narrow index in each of its parts, rather than reverting
-; to pointer width because an individual part is too short to qualify.
+; part count must still be charged for all of its lanes. And the index width is
+; chosen once for the whole operation, from the indices themselves, so the same
+; data type is priced differently depending on how wide its indices are.
+;
+; The code-size cost of a hardware gather/scatter is its part count, so the
+; numbers below can be read against CodeGen directly; the paired instruction
+; counts live in masked-gather-scatter-part-count.ll.
; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX
; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=haswell | FileCheck %s --check-prefix=AVX2
@@ -54,9 +57,9 @@ define void @scatter_v9i32(<9 x i32> %val, <9 x ptr> %ptrs, <9 x i1> %mask) {
ret void
}
-; Index width across a split: the GEP indices are 32-bit, and v24 is wide enough
-; to qualify for narrowing, so all three parts are priced with a dword index
-; even though a single part on its own would not qualify.
+; Index width across a split: the GEP indices are 32-bit, so the operation is
+; priced with a dword index. Sixteen dword-indexed lanes fill a part, so these
+; 24 lanes take two, and CodeGen emits two vpgatherdd.
define <24 x i32> @gather_v24i32_dword_index(ptr %base, <24 x i32> %idx, <24 x i1> %mask) {
; SKX-LABEL: 'gather_v24i32_dword_index'
; SKX-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
@@ -73,11 +76,14 @@ define <24 x i32> @gather_v24i32_dword_index(ptr %base, <24 x i32> %idx, <24 x i
ret <24 x i32> %v
}
-; Control for the above: genuinely 64-bit indices cannot narrow.
+; Control for the above: genuinely 64-bit indices cannot narrow, so only eight
+; lanes fill a part and the same 24 lanes take three vpgatherqd. Their indices
+; occupy four legal registers, one of which holds no live lane, so a part count
+; taken from the register count would overstate the work by one.
define <24 x i32> @gather_v24i32_qword_index(ptr %base, <24 x i64> %idx, <24 x i1> %mask) {
; SKX-LABEL: 'gather_v24i32_qword_index'
; SKX-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
-; SKX-NEXT: Cost Model: Found costs of RThru:32 CodeSize:4 Lat:104 SizeLat:32 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:30 CodeSize:3 Lat:102 SizeLat:30 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
;
; AVX2-LABEL: 'gather_v24i32_qword_index'
diff --git a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll
index f835a8aedbbcc..bf19e53e1a374 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll
@@ -1909,7 +1909,7 @@ define <16 x float> @test_gather_16f32_const_mask(ptr %base, <16 x i32> %ind) {
; SKL-LABEL: 'test_gather_16f32_const_mask'
; SKL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SKL-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX512-LABEL: 'test_gather_16f32_const_mask'
@@ -1953,7 +1953,7 @@ define <16 x float> @test_gather_16f32_var_mask(ptr %base, <16 x i32> %ind, <16
; SKL-LABEL: 'test_gather_16f32_var_mask'
; SKL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SKL-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX512-LABEL: 'test_gather_16f32_var_mask'
@@ -2051,7 +2051,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
; SKL-NEXT: Cost Model: Found costs of RThru:1 CodeSize:1 Lat:3 SizeLat:2 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> poison, <16 x i32> zeroinitializer
; SKL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SKL-NEXT: Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX512-LABEL: 'test_gather_16f32_const_mask2'
diff --git a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
index c02dd6db8b99f..787bb012f79bb 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
@@ -745,7 +745,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:18 SizeLat:12 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:77 CodeSize:93 Lat:141 SizeLat:93 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SSE2-NEXT: Cost Model: Found costs of RThru:49 CodeSize:58 Lat:85 SizeLat:58 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; SSE2-NEXT: Cost Model: Found costs of RThru:49 CodeSize:58 Lat:85 SizeLat:58 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
; SSE2-NEXT: Cost Model: Found costs of RThru:39 CodeSize:47 Lat:71 SizeLat:47 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:20 CodeSize:24 Lat:36 SizeLat:24 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:18 SizeLat:12 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -773,7 +773,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:49 CodeSize:65 Lat:113 SizeLat:65 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SSE42-NEXT: Cost Model: Found costs of RThru:28 CodeSize:37 Lat:64 SizeLat:37 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; SSE42-NEXT: Cost Model: Found costs of RThru:28 CodeSize:37 Lat:64 SizeLat:37 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
; SSE42-NEXT: Cost Model: Found costs of RThru:25 CodeSize:33 Lat:57 SizeLat:33 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:13 CodeSize:17 Lat:29 SizeLat:17 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -801,7 +801,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:119 SizeLat:71 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; AVX1-NEXT: Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; AVX1-NEXT: Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
; AVX1-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:30 SizeLat:18 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -829,7 +829,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:119 SizeLat:71 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; AVX2-NEXT: Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; AVX2-NEXT: Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
; AVX2-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:30 SizeLat:18 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -857,7 +857,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SKL-NEXT: Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SKL-NEXT: Cost Model: Found costs of RThru:17 CodeSize:4 Lat:44 SizeLat:17 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:15 CodeSize:3 Lat:42 SizeLat:15 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
; SKL-NEXT: Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -885,7 +885,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; KNL-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; KNL-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
; KNL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:17 CodeSize:21 Lat:33 SizeLat:21 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -913,7 +913,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
; SKX-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -943,7 +943,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
%V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> undef, i32 1, <1 x i1> %m1, <1 x i64> undef)
%V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> undef, i32 1, <16 x i1> %m16, <16 x i32> undef)
- %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> undef, i32 1, <9 x i1> %m9, <9 x i32> undef)
+ %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> poison, i32 1, <9 x i1> %m9, <9 x i32> poison)
%V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> undef, i32 1, <8 x i1> %m8, <8 x i32> undef)
%V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> undef, i32 1, <4 x i1> %m4, <4 x i32> undef)
%V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> undef, i32 1, <2 x i1> %m2, <2 x i32> undef)
@@ -976,7 +976,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:12 SizeLat:12 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SSE2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SSE2-NEXT: Cost Model: Found costs of RThru:77 CodeSize:93 Lat:93 SizeLat:93 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; SSE2-NEXT: Cost Model: Found costs of RThru:43 CodeSize:52 Lat:52 SizeLat:52 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; SSE2-NEXT: Cost Model: Found costs of RThru:43 CodeSize:52 Lat:52 SizeLat:52 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
; SSE2-NEXT: Cost Model: Found costs of RThru:39 CodeSize:47 Lat:47 SizeLat:47 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SSE2-NEXT: Cost Model: Found costs of RThru:20 CodeSize:24 Lat:24 SizeLat:24 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:12 SizeLat:12 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1004,7 +1004,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SSE42-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SSE42-NEXT: Cost Model: Found costs of RThru:49 CodeSize:65 Lat:65 SizeLat:65 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; SSE42-NEXT: Cost Model: Found costs of RThru:28 CodeSize:37 Lat:37 SizeLat:37 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; SSE42-NEXT: Cost Model: Found costs of RThru:28 CodeSize:37 Lat:37 SizeLat:37 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
; SSE42-NEXT: Cost Model: Found costs of RThru:25 CodeSize:33 Lat:33 SizeLat:33 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SSE42-NEXT: Cost Model: Found costs of RThru:13 CodeSize:17 Lat:17 SizeLat:17 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1032,7 +1032,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; AVX1-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; AVX1-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; AVX1-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; AVX1-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
; AVX1-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; AVX1-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1060,7 +1060,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; AVX2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; AVX2-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; AVX2-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; AVX2-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
; AVX2-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; AVX2-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1088,7 +1088,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SKL-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SKL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SKL-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; SKL-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; SKL-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
; SKL-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SKL-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SKL-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1116,7 +1116,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; KNL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; KNL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; KNL-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; KNL-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
; KNL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; KNL-NEXT: Cost Model: Found costs of RThru:17 CodeSize:21 Lat:21 SizeLat:21 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1144,7 +1144,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SKX-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
; SKX-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SKX-NEXT: Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1174,7 +1174,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> undef, i32 1, <1 x i1> %m1)
call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> undef, i32 1, <16 x i1> %m16)
- call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> undef, i32 1, <9 x i1> %m9)
+ call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> poison, i32 1, <9 x i1> %m9)
call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> undef, i32 1, <8 x i1> %m8)
call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> undef, i32 1, <4 x i1> %m4)
call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> undef, i32 1, <2 x i1> %m2)
@@ -1925,7 +1925,7 @@ define <16 x float> @test_gather_16f32_const_mask(ptr %base, <16 x i32> %ind) {
; SKL-LABEL: 'test_gather_16f32_const_mask'
; SKL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SKL-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX512-LABEL: 'test_gather_16f32_const_mask'
@@ -1969,7 +1969,7 @@ define <16 x float> @test_gather_16f32_var_mask(ptr %base, <16 x i32> %ind, <16
; SKL-LABEL: 'test_gather_16f32_var_mask'
; SKL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SKL-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX512-LABEL: 'test_gather_16f32_var_mask'
@@ -2067,7 +2067,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
; SKL-NEXT: Cost Model: Found costs of RThru:1 CodeSize:1 Lat:3 SizeLat:2 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> undef, <16 x i32> zeroinitializer
; SKL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SKL-NEXT: Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX512-LABEL: 'test_gather_16f32_const_mask2'
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll
index 2d5a30019bacd..371338b6bcb78 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll
@@ -50,9 +50,9 @@ define void @test() {
; AVX2-FASTGATHER: LV: Found an estimated cost of 1 for VF 1 For instruction: %valB = load i32, ptr %inB, align 4
; AVX2-FASTGATHER: Cost of 4 for VF 2: WIDEN ir<%valB> = load ir<%inB>
; AVX2-FASTGATHER: Cost of 6 for VF 4: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER: Cost of 12 for VF 8: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER: Cost of 24 for VF 16: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER: Cost of 48 for VF 32: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER: Cost of 10 for VF 8: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER: Cost of 20 for VF 16: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER: Cost of 40 for VF 32: WIDEN ir<%valB> = load ir<%inB>
;
; AVX512-LABEL: 'test'
; AVX512: LV: Found an estimated cost of 1 for VF 1 For instruction: %valB = load i32, ptr %inB, align 4
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll
index 8f5da77027970..0cf2b6b9c12b2 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll
@@ -43,9 +43,9 @@ define void @test() {
; AVX2-FASTGATHER: LV: Found an estimated cost of 1 for VF 1 For instruction: %valB.loaded = load i32, ptr %inB, align 4
; AVX2-FASTGATHER: Cost of 4 for VF 2: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
; AVX2-FASTGATHER: Cost of 6 for VF 4: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER: Cost of 12 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER: Cost of 24 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER: Cost of 48 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER: Cost of 10 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER: Cost of 20 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER: Cost of 40 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
;
; AVX512-LABEL: 'test'
; AVX512: LV: Found an estimated cost of 1 for VF 1 For instruction: %valB.loaded = load i32, ptr %inB, align 4
diff --git a/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll b/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll
index 7f0ba19d3e7a8..bc2bea9708227 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll
@@ -123,28 +123,28 @@ define void @replicate_sext(i32 %N, ptr %dst, ptr %src) #0 {
; CHECK-NEXT: [[FOUND_CONFLICT:%.*]] = and i1 [[BOUND0]], [[BOUND1]]
; CHECK-NEXT: br i1 [[FOUND_CONFLICT]], label %[[SCALAR_PH]], label %[[VECTOR_PH:.*]]
; CHECK: [[VECTOR_PH]]:
-; CHECK-NEXT: [[N_MOD_VF:%.*]] = and i32 [[TMP0]], 3
+; CHECK-NEXT: [[N_MOD_VF:%.*]] = and i32 [[TMP0]], 7
; CHECK-NEXT: [[TMP18:%.*]] = icmp eq i32 [[N_MOD_VF]], 0
-; CHECK-NEXT: [[TMP19:%.*]] = select i1 [[TMP18]], i32 4, i32 [[N_MOD_VF]]
+; CHECK-NEXT: [[TMP19:%.*]] = select i1 [[TMP18]], i32 8, i32 [[N_MOD_VF]]
; CHECK-NEXT: [[N_VEC:%.*]] = sub i32 [[TMP0]], [[TMP19]]
; CHECK-NEXT: [[TMP20:%.*]] = shl i32 [[N_VEC]], 2
; CHECK-NEXT: [[TMP21:%.*]] = mul i32 [[N_VEC]], 3
; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK: [[VECTOR_BODY]]:
; CHECK-NEXT: [[INDEX:%.*]] = phi i32 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[VEC_IND:%.*]] = phi <4 x i32> [ <i32 0, i32 3, i32 6, i32 9>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_IND:%.*]] = phi <8 x i32> [ <i32 0, i32 3, i32 6, i32 9, i32 12, i32 15, i32 18, i32 21>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
; CHECK-NEXT: [[OFFSET_IDX:%.*]] = shl i32 [[INDEX]], 2
; CHECK-NEXT: [[TMP22:%.*]] = sext i32 [[OFFSET_IDX]] to i64
; CHECK-NEXT: [[TMP23:%.*]] = getelementptr nusw i32, ptr [[SRC]], i64 [[TMP22]]
-; CHECK-NEXT: [[WIDE_VEC:%.*]] = load <16 x i32>, ptr [[TMP23]], align 4, !alias.scope [[META4:![0-9]+]]
-; CHECK-NEXT: [[STRIDED_VEC:%.*]] = shufflevector <16 x i32> [[WIDE_VEC]], <16 x i32> poison, <4 x i32> <i32 0, i32 4, i32 8, i32 12>
-; CHECK-NEXT: [[STRIDED_VEC9:%.*]] = shufflevector <16 x i32> [[WIDE_VEC]], <16 x i32> poison, <4 x i32> <i32 1, i32 5, i32 9, i32 13>
-; CHECK-NEXT: [[TMP25:%.*]] = sext <4 x i32> [[VEC_IND]] to <4 x i64>
-; CHECK-NEXT: [[TMP27:%.*]] = getelementptr i32, ptr [[DST]], <4 x i64> [[TMP25]]
-; CHECK-NEXT: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[STRIDED_VEC]], <4 x ptr> align 4 [[TMP27]], <4 x i1> splat (i1 true)), !alias.scope [[META7:![0-9]+]], !noalias [[META4]]
-; CHECK-NEXT: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[STRIDED_VEC9]], <4 x ptr> align 4 [[TMP27]], <4 x i1> splat (i1 true)), !alias.scope [[META7]], !noalias [[META4]]
-; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 4
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add <4 x i32> [[VEC_IND]], splat (i32 12)
+; CHECK-NEXT: [[WIDE_VEC:%.*]] = load <32 x i32>, ptr [[TMP23]], align 4, !alias.scope [[META4:![0-9]+]]
+; CHECK-NEXT: [[STRIDED_VEC:%.*]] = shufflevector <32 x i32> [[WIDE_VEC]], <32 x i32> poison, <8 x i32> <i32 0, i32 4, i32 8, i32 12, i32 16, i32 20, i32 24, i32 28>
+; CHECK-NEXT: [[STRIDED_VEC2:%.*]] = shufflevector <32 x i32> [[WIDE_VEC]], <32 x i32> poison, <8 x i32> <i32 1, i32 5, i32 9, i32 13, i32 17, i32 21, i32 25, i32 29>
+; CHECK-NEXT: [[TMP24:%.*]] = sext <8 x i32> [[VEC_IND]] to <8 x i64>
+; CHECK-NEXT: [[WIDE_GEP:%.*]] = getelementptr i32, ptr [[DST]], <8 x i64> [[TMP24]]
+; CHECK-NEXT: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[STRIDED_VEC]], <8 x ptr> align 4 [[WIDE_GEP]], <8 x i1> splat (i1 true)), !alias.scope [[META7:![0-9]+]], !noalias [[META4]]
+; CHECK-NEXT: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[STRIDED_VEC2]], <8 x ptr> align 4 [[WIDE_GEP]], <8 x i1> splat (i1 true)), !alias.scope [[META7]], !noalias [[META4]]
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 8
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <8 x i32> [[VEC_IND]], splat (i32 24)
; CHECK-NEXT: [[TMP26:%.*]] = icmp eq i32 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-NEXT: br i1 [[TMP26]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP9:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
More information about the llvm-commits
mailing list