[llvm] [X86][CostModel] Fix lane and index accounting in split gather/scatter (PR #220565)
Sumukh J Bharadwaj via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 8 05:00:09 PDT 2026
https://github.com/amd-subharad updated https://github.com/llvm/llvm-project/pull/220565
>From 3338e80ae011a806499d035e05ec6bee41a4c17b Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Wed, 2 Sep 2026 17:11:18 +0530
Subject: [PATCH 1/3] [X86][CostModel] Fix lane and index accounting in split
gather/scatter
When a masked gather or scatter needs more than one register's worth of
pointers, getGSVectorCost recursed on a vector of VF / SplitFactor elements
and multiplied the result. Two things were wrong with that.
The division truncates, so any vector length that is not a multiple of its
split factor lost the remainder: a v9i32 gather was costed as eight lanes,
making its body term as cheap as v8i32's.
The recursive call also recomputed the index width from the shortened vector.
Narrowing a 64-bit index to 32 bits requires a minimum vector length, so an
operation long enough to qualify was still priced with pointer-width indices
once it had been split into parts that were individually too short. Every part
of an operation uses the index width the whole operation qualifies for, so
this overstated the cost of the wide case.
Compute the split cost directly instead: one overhead per part, plus one
scalar memory op for every lane of the original vector, with the index width
chosen once for the operation as a whole.
Both corrections move reported costs, on every subtarget and every cost kind.
On skylake-avx512 a v9i32 gather goes from 12 to 13, and a v24i32 gather
through a dword-index GEP goes from 32 to 28, the latter also dropping from
four parts to two. Lengths that do divide their split factor, and operations
that genuinely need 64-bit indices, are unchanged.
---
llvm/docs/ReleaseNotes.md | 10 ++
.../lib/Target/X86/X86TargetTransformInfo.cpp | 20 ++--
.../X86/masked-gather-scatter-split-cost.ll | 91 +++++++++++++++++++
.../CostModel/X86/masked-intrinsic-cost.ll | 20 +++-
4 files changed, 127 insertions(+), 14 deletions(-)
create mode 100644 llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
diff --git a/llvm/docs/ReleaseNotes.md b/llvm/docs/ReleaseNotes.md
index 9f1611e4cab7f..4b863fb5c62bc 100644
--- a/llvm/docs/ReleaseNotes.md
+++ b/llvm/docs/ReleaseNotes.md
@@ -230,6 +230,16 @@ Makes programs 10x faster by doing Special New Thing.
### Changes to the X86 Backend
+* Masked gather/scatter operations that are split across several registers are
+ costed more accurately. Lanes in a vector whose length is not a multiple of
+ its split factor are no longer dropped, so a `v9i32` gather is costed as nine
+ lanes rather than eight. The index width is also chosen for the operation as
+ a whole rather than recomputed for each part, so a wide gather through a
+ 32-bit-index GEP is no longer priced as if it used 64-bit indices. Both
+ corrections apply to every X86 subtarget and every cost kind. Lengths that do
+ divide their split factor, and operations that genuinely need 64-bit indices,
+ are unchanged.
+
### Changes to the OCaml bindings
### Changes to the Python bindings
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
index 8e0cf1fc5a153..fc40b17f91e8e 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
@@ -6659,24 +6659,20 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
std::pair<InstructionCost, MVT> SrcLT = getTypeLegalizationCost(SrcVTy);
InstructionCost::CostType SplitFactor =
std::max(IdxsLT.first, SrcLT.first).getValue();
- if (SplitFactor > 1) {
- // Handle splitting of vector of pointers
- auto *SplitSrcTy =
- FixedVectorType::get(SrcVTy->getScalarType(), VF / SplitFactor);
- return SplitFactor * getGSVectorCost(Opcode, CostKind, SplitSrcTy, Ptr,
- Alignment, AddressSpace);
- }
-
- // If we didn't split, this will be a single gather/scatter instruction.
+ // A vector of pointers that does not fit one register is split into
+ // SplitFactor gather/scatter instructions, each paying the overhead.
if (CostKind == TTI::TCK_CodeSize)
- return 1;
+ return SplitFactor;
// The gather / scatter cost is given by Intel architects. It is a rough
// number since we are looking at one instruction in a time.
const int GSOverhead = (Opcode == Instruction::Load) ? getGatherOverhead()
: getScatterOverhead();
- return GSOverhead + VF * getMemoryOpCost(Opcode, SrcVTy->getScalarType(),
- Alignment, AddressSpace, CostKind);
+ // Charge every lane of the original vector. Descending into VF / SplitFactor
+ // dropped the remainder, so a v9 gather was costed as eight lanes.
+ return SplitFactor * GSOverhead +
+ VF * getMemoryOpCost(Opcode, SrcVTy->getScalarType(), Alignment,
+ AddressSpace, CostKind);
}
/// Calculate the cost of Gather / Scatter operation
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
new file mode 100644
index 0000000000000..26954697c2b3f
--- /dev/null
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
@@ -0,0 +1,91 @@
+; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --version 6
+; Costs for masked gather/scatter operations that need more than one register's
+; worth of pointers, across every cost kind.
+;
+; Two properties are pinned here. A vector length that is not a multiple of its
+; split factor must still be charged for all of its lanes. And the index width
+; is chosen once for the whole operation, so a length wide enough to qualify for
+; narrowing keeps the narrow index in each of its parts, rather than reverting
+; to pointer width because an individual part is too short to qualify.
+
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=haswell | FileCheck %s --check-prefix=AVX2
+
+; A length that divides its split factor: unchanged, and the reference point for
+; the two cases below.
+define <8 x i32> @gather_v8i32(<8 x ptr> %ptrs, <8 x i1> %mask) {
+; SKX-LABEL: 'gather_v8i32'
+; SKX-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <8 x i32> %v
+;
+; AVX2-LABEL: 'gather_v8i32'
+; AVX2-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <8 x i32> %v
+;
+ %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x i32> poison)
+ ret <8 x i32> %v
+}
+
+; Remainder lanes: v9 splits into two parts but is not a multiple of two, so the
+; ninth lane must not be dropped from the body term.
+define <9 x i32> @gather_v9i32(<9 x ptr> %ptrs, <9 x i1> %mask) {
+; SKX-LABEL: 'gather_v9i32'
+; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 4 %ptrs, <9 x i1> %mask, <9 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <9 x i32> %v
+;
+; AVX2-LABEL: 'gather_v9i32'
+; AVX2-NEXT: Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 4 %ptrs, <9 x i1> %mask, <9 x i32> poison)
+; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <9 x i32> %v
+;
+ %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> %ptrs, i32 4, <9 x i1> %mask, <9 x i32> poison)
+ ret <9 x i32> %v
+}
+
+define void @scatter_v9i32(<9 x i32> %val, <9 x ptr> %ptrs, <9 x i1> %mask) {
+; SKX-LABEL: 'scatter_v9i32'
+; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> align 4 %ptrs, <9 x i1> %mask)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+; AVX2-LABEL: 'scatter_v9i32'
+; AVX2-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> align 4 %ptrs, <9 x i1> %mask)
+; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+ call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> %ptrs, i32 4, <9 x i1> %mask)
+ ret void
+}
+
+; Index width across a split: the GEP indices are 32-bit, and v24 is wide enough
+; to qualify for narrowing, so all three parts are priced with a dword index
+; even though a single part on its own would not qualify.
+define <24 x i32> @gather_v24i32_dword_index(ptr %base, <24 x i32> %idx, <24 x i1> %mask) {
+; SKX-LABEL: 'gather_v24i32_dword_index'
+; SKX-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
+; SKX-NEXT: Cost Model: Found costs of RThru:28 CodeSize:2 Lat:100 SizeLat:28 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
+;
+; AVX2-LABEL: 'gather_v24i32_dword_index'
+; AVX2-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
+; AVX2-NEXT: Cost Model: Found costs of RThru:82 CodeSize:106 Lat:178 SizeLat:106 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
+;
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> %mask, <24 x i32> poison)
+ ret <24 x i32> %v
+}
+
+; Control for the above: genuinely 64-bit indices cannot narrow.
+define <24 x i32> @gather_v24i32_qword_index(ptr %base, <24 x i64> %idx, <24 x i1> %mask) {
+; SKX-LABEL: 'gather_v24i32_qword_index'
+; SKX-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+; SKX-NEXT: Cost Model: Found costs of RThru:32 CodeSize:4 Lat:104 SizeLat:32 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
+;
+; AVX2-LABEL: 'gather_v24i32_qword_index'
+; AVX2-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+; AVX2-NEXT: Cost Model: Found costs of RThru:82 CodeSize:106 Lat:178 SizeLat:106 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
+;
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> %mask, <24 x i32> poison)
+ ret <24 x i32> %v
+}
diff --git a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
index ed1b534fac8f8..c02dd6db8b99f 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
@@ -730,7 +730,7 @@ define i32 @masked_store(<1 x i1> %m1, <2 x i1> %m2, <3 x i1> %m3, <4 x i1> %m4,
ret i32 0
}
-define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8, <16 x i1> %m16, <32 x i1> %m32, <64 x i1> %m64) {
+define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8, <9 x i1> %m9, <16 x i1> %m16, <32 x i1> %m32, <64 x i1> %m64) {
; SSE2-LABEL: 'masked_gather'
; SSE2-NEXT: Cost Model: Found costs of RThru:29 CodeSize:37 Lat:61 SizeLat:37 for: %V8F64 = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x double> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:15 CodeSize:19 Lat:31 SizeLat:19 for: %V4F64 = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x double> undef)
@@ -745,6 +745,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:18 SizeLat:12 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:77 CodeSize:93 Lat:141 SizeLat:93 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SSE2-NEXT: Cost Model: Found costs of RThru:49 CodeSize:58 Lat:85 SizeLat:58 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:39 CodeSize:47 Lat:71 SizeLat:47 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:20 CodeSize:24 Lat:36 SizeLat:24 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:18 SizeLat:12 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -772,6 +773,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:49 CodeSize:65 Lat:113 SizeLat:65 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SSE42-NEXT: Cost Model: Found costs of RThru:28 CodeSize:37 Lat:64 SizeLat:37 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:25 CodeSize:33 Lat:57 SizeLat:33 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:13 CodeSize:17 Lat:29 SizeLat:17 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -799,6 +801,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:119 SizeLat:71 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; AVX1-NEXT: Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:30 SizeLat:18 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -826,6 +829,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:119 SizeLat:71 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; AVX2-NEXT: Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:30 SizeLat:18 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -853,6 +857,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SKL-NEXT: Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:17 CodeSize:4 Lat:44 SizeLat:17 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -880,6 +885,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; KNL-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:17 CodeSize:21 Lat:33 SizeLat:21 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -907,6 +913,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -936,6 +943,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
%V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> undef, i32 1, <1 x i1> %m1, <1 x i64> undef)
%V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> undef, i32 1, <16 x i1> %m16, <16 x i32> undef)
+ %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> undef, i32 1, <9 x i1> %m9, <9 x i32> undef)
%V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> undef, i32 1, <8 x i1> %m8, <8 x i32> undef)
%V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> undef, i32 1, <4 x i1> %m4, <4 x i32> undef)
%V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> undef, i32 1, <2 x i1> %m2, <2 x i32> undef)
@@ -953,7 +961,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
ret i32 0
}
-define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8, <16 x i1> %m16, <32 x i1> %m32, <64 x i1> %m64) {
+define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8, <9 x i1> %m9, <16 x i1> %m16, <32 x i1> %m32, <64 x i1> %m64) {
; SSE2-LABEL: 'masked_scatter'
; SSE2-NEXT: Cost Model: Found costs of RThru:29 CodeSize:37 Lat:37 SizeLat:37 for: call void @llvm.masked.scatter.v8f64.v8p0(<8 x double> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SSE2-NEXT: Cost Model: Found costs of RThru:15 CodeSize:19 Lat:19 SizeLat:19 for: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
@@ -968,6 +976,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:12 SizeLat:12 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SSE2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SSE2-NEXT: Cost Model: Found costs of RThru:77 CodeSize:93 Lat:93 SizeLat:93 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SSE2-NEXT: Cost Model: Found costs of RThru:43 CodeSize:52 Lat:52 SizeLat:52 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
; SSE2-NEXT: Cost Model: Found costs of RThru:39 CodeSize:47 Lat:47 SizeLat:47 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SSE2-NEXT: Cost Model: Found costs of RThru:20 CodeSize:24 Lat:24 SizeLat:24 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:12 SizeLat:12 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -995,6 +1004,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SSE42-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SSE42-NEXT: Cost Model: Found costs of RThru:49 CodeSize:65 Lat:65 SizeLat:65 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SSE42-NEXT: Cost Model: Found costs of RThru:28 CodeSize:37 Lat:37 SizeLat:37 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
; SSE42-NEXT: Cost Model: Found costs of RThru:25 CodeSize:33 Lat:33 SizeLat:33 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SSE42-NEXT: Cost Model: Found costs of RThru:13 CodeSize:17 Lat:17 SizeLat:17 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1022,6 +1032,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; AVX1-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; AVX1-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; AVX1-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
; AVX1-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; AVX1-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1049,6 +1060,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; AVX2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; AVX2-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; AVX2-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
; AVX2-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; AVX2-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1076,6 +1088,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SKL-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SKL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SKL-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SKL-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
; SKL-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SKL-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SKL-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1103,6 +1116,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; KNL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; KNL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; KNL-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
; KNL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; KNL-NEXT: Cost Model: Found costs of RThru:17 CodeSize:21 Lat:21 SizeLat:21 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1130,6 +1144,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SKX-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
; SKX-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SKX-NEXT: Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1159,6 +1174,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> undef, i32 1, <1 x i1> %m1)
call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> undef, i32 1, <16 x i1> %m16)
+ call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> undef, i32 1, <9 x i1> %m9)
call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> undef, i32 1, <8 x i1> %m8)
call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> undef, i32 1, <4 x i1> %m4)
call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> undef, i32 1, <2 x i1> %m2)
>From 16d10bd16b1057d7f0f1b15db60846197b9f61cd Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Thu, 3 Sep 2026 10:39:50 +0530
Subject: [PATCH 2/3] [X86][CostModel] Take the gather/scatter index width from
the pointer's address space
getGSVectorCost started the index width at DL.getPointerSizeInBits(), which
answers for address space 0 regardless of where the gathered pointers actually
live. On x86, address spaces 270 and 271 hold 32-bit pointers, so a vector of
them is half the width of the same number of ordinary pointers.
Costing them at 64 bits made a <16 x ptr addrspace(270)> gather look as though
its indices needed two ZMMs, so it was charged for a split it does not need:
on skylake-avx512 it cost 20 with two parts rather than 18 with one. A GEP
whose indices were 32-bit already reached the right answer through the index
analysis below, so only pointers used directly, without such a GEP, were
mispriced.
Take the width from the pointer operand's own address space, falling back to
the caller's address space when there is no operand to ask. Address space 0 is
unaffected, as is address space 272, whose pointers really are 64-bit.
---
llvm/docs/ReleaseNotes.md | 6 ++
.../lib/Target/X86/X86TargetTransformInfo.cpp | 17 +++--
.../masked-gather-scatter-address-space.ll | 70 +++++++++++++++++++
3 files changed, 88 insertions(+), 5 deletions(-)
create mode 100644 llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
diff --git a/llvm/docs/ReleaseNotes.md b/llvm/docs/ReleaseNotes.md
index 4b863fb5c62bc..98cfd6fd0e075 100644
--- a/llvm/docs/ReleaseNotes.md
+++ b/llvm/docs/ReleaseNotes.md
@@ -240,6 +240,12 @@ Makes programs 10x faster by doing Special New Thing.
divide their split factor, and operations that genuinely need 64-bit indices,
are unchanged.
+* Masked gather/scatter costs now take the index width from the address space
+ their pointers live in, instead of assuming the width of address space 0. On
+ x86 this matters for the 32-bit address spaces 270 and 271: a vector of those
+ pointers is half as wide as the same count of ordinary pointers, so it is no
+ longer costed as though it had to be split in two.
+
### Changes to the OCaml bindings
### Changes to the Python bindings
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
index fc40b17f91e8e..db2c0d1849d72 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
@@ -6623,8 +6623,16 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
// operation will use 16 x 64 indices which do not fit in a zmm and needs
// to split. Also check that the base pointer is the same for all lanes,
// and that there's at most one variable index.
- auto getIndexSizeInBits = [](const Value *Ptr, const DataLayout &DL) {
- unsigned IndexSize = DL.getPointerSizeInBits();
+ // The index starts out as wide as the pointers being gathered, which is a
+ // property of the address space they live in. Asking the data layout without
+ // one answers for address space 0, which is not necessarily theirs.
+ unsigned PtrSizeInBits = DL.getPointerSizeInBits(
+ Ptr && Ptr->getType()->isPtrOrPtrVectorTy()
+ ? Ptr->getType()->getScalarType()->getPointerAddressSpace()
+ : AddressSpace);
+
+ auto getIndexSizeInBits = [PtrSizeInBits](const Value *Ptr) {
+ unsigned IndexSize = PtrSizeInBits;
const GetElementPtrInst *GEP = dyn_cast_or_null<GetElementPtrInst>(Ptr);
if (IndexSize < 64 || !GEP)
return IndexSize;
@@ -6649,9 +6657,8 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
// Trying to reduce IndexSize to 32 bits for vector 16.
// By default the IndexSize is equal to pointer size.
- unsigned IndexSize = (ST->hasAVX512() && VF >= 16)
- ? getIndexSizeInBits(Ptr, DL)
- : DL.getPointerSizeInBits();
+ unsigned IndexSize =
+ (ST->hasAVX512() && VF >= 16) ? getIndexSizeInBits(Ptr) : PtrSizeInBits;
auto *IndexVTy = FixedVectorType::get(
IntegerType::get(SrcVTy->getContext(), IndexSize), VF);
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
new file mode 100644
index 0000000000000..0125c7975b7c5
--- /dev/null
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
@@ -0,0 +1,70 @@
+; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --version 6
+; Gathers and scatters over pointers that are not in address space 0. On x86,
+; address spaces 270 and 271 hold 32-bit pointers, so a vector of them is half
+; the width of the same number of address-space-0 pointers, needs no split, and
+; is indexed with a dword rather than a qword index.
+
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX
+
+target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i64:64-i128:128-f80:128-n8:16:32:64-S128"
+
+; A vector of sixteen 32-bit pointers occupies one ZMM, so this is a single
+; dword-indexed gather, not a split qword-indexed pair.
+define <16 x i32> @gather_v16i32_p270(<16 x ptr addrspace(270)> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'gather_v16i32_p270'
+; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+ %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+ ret <16 x i32> %v
+}
+
+; The same shape in address space 0 does need 64-bit indices, and does split.
+define <16 x i32> @gather_v16i32_p0(<16 x ptr> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'gather_v16i32_p0'
+; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+ %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+ ret <16 x i32> %v
+}
+
+; Reaching the same conclusion through a GEP rather than from the pointer type.
+define <16 x i32> @gather_v16i32_p270_gep(ptr addrspace(270) %base, <16 x i32> %idx, <16 x i1> %mask) {
+; SKX-LABEL: 'gather_v16i32_p270_gep'
+; SKX-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr addrspace(270) %base, <16 x i32> %idx
+; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+ %ptrs = getelementptr inbounds i32, ptr addrspace(270) %base, <16 x i32> %idx
+ %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+ ret <16 x i32> %v
+}
+
+define void @scatter_v16i32_p270(<16 x i32> %val, <16 x ptr addrspace(270)> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'scatter_v16i32_p270'
+; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p270(<16 x i32> %val, <16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+ call void @llvm.masked.scatter.v16i32.v16p270(<16 x i32> %val, <16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask)
+ ret void
+}
+
+define void @scatter_v16i32_p0(<16 x i32> %val, <16 x ptr> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'scatter_v16i32_p0'
+; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %val, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+ call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %val, <16 x ptr> %ptrs, i32 4, <16 x i1> %mask)
+ ret void
+}
+
+; Address space 272 holds 64-bit pointers, so it must behave like space 0.
+define <16 x i32> @gather_v16i32_p272(<16 x ptr addrspace(272)> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'gather_v16i32_p272'
+; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p272(<16 x ptr addrspace(272)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+ %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p272(<16 x ptr addrspace(272)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+ ret <16 x i32> %v
+}
>From 9e8334ab6d31abd71e7885bd92e466797445de49 Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Tue, 8 Sep 2026 01:59:02 +0530
Subject: [PATCH 3/3] [X86][CostModel] Cost gathers by the instructions CodeGen
emits
The code-size cost of a hardware masked gather/scatter is its part
count, so it can be read directly against CodeGen. It did not match.
Four things made it differ, each with its own boundary.
The part count came from the legalized register count of the widest
operand. Legalization first widens the vector length to a power of two,
so that count exceeds the instructions emitted whenever that rounding
adds a whole part with no live lane in it: a 24-lane qword-indexed
gather covers eight lanes per instruction but legalizes to a 32-lane,
four-register form whose fourth register holds nothing live, and becomes
three instructions. Count instead the lanes one instruction covers --
the wider of the data and the index bounds it, since a dword gather
indexed by qwords fills a register with indices before it fills one with
data -- and divide the length by that. This affects every subtarget with
a hardware gather or scatter, and it is the term that sets the code-size
cost, so code-size costs move too.
The index width was derived only on AVX-512 at VF >= 16. Outside that
window a gather through a 32-bit-index GEP was priced as though its
indices were pointer-width, tying it with the qword-indexed form that
CodeGen splits into roughly twice as many instructions. Derive the width
whenever a hardware gather/scatter is priced.
Deriving it more often exposed that the derivation ignored the GEP's
stride. A narrow index only survives into the instruction if the
addressing mode can apply that stride as a scale, and the scale field
encodes 1, 2, 4 and 8; any other stride is multiplied into the index
first, and the product is pointer-width. Without that check a gather
over an array of structures -- the form a vectorized `a[idx[i]].f`
takes -- was costed at half the instructions CodeGen emits. Check it, so
that both spellings are priced as CodeGen compiles them.
The mask was not consulted at all. A mask known at compile time says
which parts survive: a part with no live lane is folded away, and an
operation with no live lane at all disappears. A 24-lane qword-indexed
gather with only its first eight lanes live is one instruction, not
three, and only its live lanes are charged for memory accesses. Where
the mask is unknown the two directions differ, so the opcode decides: a
gather drops the tail that widening added, because its result goes
unused, while a scatter keeps it, because it survives as a store under a
zeroed mask that CodeGen still emits. How far that tail reaches follows
the predicate rather than the data: AVX512BW holds the mask in a
k-register as wide as the legalized vector, so the length rounds up to a
power of two, while without it the mask is broken into 16-lane pieces
and the length rounds up only to the next of those. A 48-lane
qword-indexed scatter is eight instructions on skylake-avx512 and six on
knl; the legalized register count says eight for both.
Verified by comparing the code-size cost against the instructions llc
emits, over four data types, both index widths, 21 vector lengths from 2
to 64 and 15 subtargets: 3200 of 3200 hardware cases agree, against 2530
before. A separate sweep over GEP strides from 1 to 32 bytes agrees on
all of them, on an AVX-512 and an AVX2 target. Past 64 lanes, and for a
compile-time mask carrying undef or poison lanes, the count is an upper
bound rather than exact: CodeGen may fold parts there in ways that do
not follow from the length, and an undef lane admits either answer. Both
overcount rather than under, and neither shape is one a vectorizer
produces. A second sweep holding those fixed and varying the GEP the
index comes from, the shape of a compile-time mask, the vector-width
function attributes and the address space agrees on 1962 of 2023,
against 1216 before; the remainder are vp.gather/vp.scatter, which never
reach this code because X86 routes only the masked_* intrinsics to it.
The costs the sweeps cannot read against CodeGen were checked against
the invariants they must obey instead: over four cost kinds and five
subtargets, no cost is invalid or negative, an empty mask is free,
killing lanes never costs more than keeping them, and cost never falls
as the vector grows.
masked-gather-scatter-part-count.ll pins that agreement, pairing each
cost against an llc instruction count in the same module, for both index
widths on an AVX-512, an AVX2 and a no-AVX512BW target, over gathers,
variable-mask scatters, all-lanes-active scatters, and constant masks
that are a prefix, a suffix, non-contiguous and empty. It also covers a
stride the scale field cannot encode against one it can, and gathering
pointers, where the width that sets the lane count comes from the
pointer rather than from an integer type, and which no X86 cost model
test exercised. A throughput run pins the per-lane term that code size
does not expose: the remainder lane of a length that does not fill its
parts, and the lanes a constant mask kills. Undoing any one of the
behaviours above makes the test fail.
masked-gather-scatter-split-cost.ll takes skylake as its AVX2 target
rather than haswell, which has no hardware gather at all, so that the
numbers it pins are part counts that can be read against CodeGen rather
than the cost of scalarizing.
The address-space test gains an AVX2 run, where the narrower pointer
also changes the cost, and the address space 271 cases it described but
did not cover.
One vectorizer decision changes: in cast-costs.ll a v8i32 scatter with
dword indices was costed as two instructions where llc emits one, and
the loop now vectorizes at VF 8 rather than VF 4.
---
llvm/docs/ReleaseNotes.md | 66 ++-
.../lib/Target/X86/X86TargetTransformInfo.cpp | 142 ++++--
llvm/lib/Target/X86/X86TargetTransformInfo.h | 3 +-
.../masked-gather-scatter-address-space.ll | 62 ++-
.../X86/masked-gather-scatter-part-count.ll | 445 ++++++++++++++++++
.../X86/masked-gather-scatter-split-cost.ll | 40 +-
.../X86/masked-intrinsic-cost-inseltpoison.ll | 6 +-
.../CostModel/X86/masked-intrinsic-cost.ll | 74 +--
.../X86/CostModel/gather-i32-with-i8-index.ll | 6 +-
.../masked-gather-i32-with-i8-index.ll | 6 +-
.../LoopVectorize/X86/cast-costs.ll | 24 +-
11 files changed, 757 insertions(+), 117 deletions(-)
create mode 100644 llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll
diff --git a/llvm/docs/ReleaseNotes.md b/llvm/docs/ReleaseNotes.md
index 98cfd6fd0e075..ffdf8c1fc458a 100644
--- a/llvm/docs/ReleaseNotes.md
+++ b/llvm/docs/ReleaseNotes.md
@@ -230,21 +230,61 @@ Makes programs 10x faster by doing Special New Thing.
### Changes to the X86 Backend
-* Masked gather/scatter operations that are split across several registers are
- costed more accurately. Lanes in a vector whose length is not a multiple of
- its split factor are no longer dropped, so a `v9i32` gather is costed as nine
- lanes rather than eight. The index width is also chosen for the operation as
- a whole rather than recomputed for each part, so a wide gather through a
- 32-bit-index GEP is no longer priced as if it used 64-bit indices. Both
- corrections apply to every X86 subtarget and every cost kind. Lengths that do
- divide their split factor, and operations that genuinely need 64-bit indices,
+* A masked gather/scatter that is split across several instructions is now
+ costed by how many instructions CodeGen emits, rather than by how many
+ legalized registers its operands occupy. The two differ when rounding the
+ length up to a power of two, which is what legalization does first, adds at
+ least one whole part with no live lane in it: a 24-lane qword-indexed gather
+ covers eight lanes per instruction, but legalizes to a 32-lane, four-register
+ form whose fourth register holds nothing live, so it is now costed as the
+ three instructions it becomes. This affects every subtarget that has a
+ hardware gather or scatter, at the lengths where that widening adds a dead
+ part, and it is the term that sets the code-size cost, so code-size costs
+ move as well. Lengths whose widening adds no whole dead part are unchanged.
+
+* All lanes of the original vector are now charged. Dividing the length by the
+ number of parts discarded the remainder, so a `v9i32` gather was costed as
+ eight lanes. This affects any length that is not a multiple of its part
+ count on any subtarget with a hardware gather or scatter. It does not change
+ code-size costs, which count instructions rather than lanes.
+
+* The index width of a masked gather/scatter is now derived whenever one is
+ priced, instead of only on AVX-512 targets at vector lengths of 16 or more.
+ Outside that window a gather through a 32-bit-index GEP was priced as though
+ its indices were pointer-width, tying it with the qword-indexed form that
+ CodeGen splits into roughly twice as many instructions. Non-AVX-512 targets
+ with a hardware gather, such as `skylake`, and AVX-512 targets at shorter
+ lengths, are the ones affected; operations that genuinely need 64-bit indices
are unchanged.
-* Masked gather/scatter costs now take the index width from the address space
- their pointers live in, instead of assuming the width of address space 0. On
- x86 this matters for the 32-bit address spaces 270 and 271: a vector of those
- pointers is half as wide as the same count of ordinary pointers, so it is no
- longer costed as though it had to be split in two.
+ A narrow index only survives into the instruction if the addressing mode can
+ apply the GEP's stride as a scale, and the scale field encodes 1, 2, 4 and 8.
+ Any other stride is multiplied into the index first, and the product is
+ pointer-width, so such a gather is now costed with 64-bit indices. This
+ corrects a gather over an array of structures, the form a vectorized
+ `a[idx[i]].f` takes, which was previously costed at half the instructions
+ CodeGen emits.
+
+* That index width is also taken from the address space the pointers live in,
+ instead of assuming the width of address space 0. On x86 this matters for the
+ 32-bit address spaces 270 and 271: a vector of those pointers is half as wide
+ as the same count of ordinary pointers, so it is no longer costed as though it
+ had to be split in two.
+
+* A mask known at compile time now decides which parts are counted, so a part
+ holding no live lane is not charged and an operation with no live lane at all
+ is free. A 24-lane qword-indexed gather with only its first eight lanes live
+ becomes one instruction rather than three, and only the live lanes are charged
+ for their memory accesses. Where the mask is not known, a gather still drops
+ the tail that legalization's widening added, because its result goes unused,
+ while a scatter counts it, because it survives as a store under a zeroed mask
+ that CodeGen still emits. How far that tail reaches follows the predicate:
+ with AVX512BW the mask fills a k-register as wide as the legalized vector, and
+ without it the mask is broken into 16-lane pieces and the tail stops at the
+ next of those.
+
+ Because these are the costs the loop vectorizer compares, some loops
+ containing a gather or scatter now vectorize at a wider factor than before.
### Changes to the OCaml bindings
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
index db2c0d1849d72..9cebe1c422daf 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
@@ -55,6 +55,7 @@
#include "llvm/CodeGen/BasicTTIImpl.h"
#include "llvm/CodeGen/CostTable.h"
#include "llvm/CodeGen/TargetLowering.h"
+#include "llvm/IR/GetElementPtrTypeIterator.h"
#include "llvm/IR/InstIterator.h"
#include "llvm/IR/IntrinsicInst.h"
#include <optional>
@@ -6608,12 +6609,37 @@ int X86TTIImpl::getScatterOverhead() const {
return 1024;
}
+// Count the lanes a compile-time constant mask leaves live, and the groups of
+// GroupSize consecutive lanes that hold at least one of them. Returns nullopt
+// when the mask is not known at compile time, which is the common case.
+static std::optional<std::pair<unsigned, unsigned>>
+getActiveGSLanesAndGroups(const Value *Mask, unsigned VF, unsigned GroupSize) {
+ const auto *CMask = dyn_cast_or_null<Constant>(Mask);
+ if (!CMask || !GroupSize)
+ return std::nullopt;
+ unsigned Lanes = 0, Groups = 0;
+ for (unsigned Base = 0; Base < VF; Base += GroupSize) {
+ bool GroupLive = false;
+ for (unsigned I = Base, E = std::min(Base + GroupSize, VF); I != E; ++I) {
+ // An element that cannot be read, or that is not a known-zero bit, has
+ // to be counted as active.
+ const Constant *Elt = CMask->getAggregateElement(I);
+ if (!Elt || !Elt->isNullValue()) {
+ ++Lanes;
+ GroupLive = true;
+ }
+ }
+ if (GroupLive)
+ ++Groups;
+ }
+ return std::make_pair(Lanes, Groups);
+}
+
// Return an average cost of Gather / Scatter instruction, maybe improved later.
-InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
- TTI::TargetCostKind CostKind,
- Type *SrcVTy, const Value *Ptr,
- Align Alignment,
- unsigned AddressSpace) const {
+InstructionCost
+X86TTIImpl::getGSVectorCost(unsigned Opcode, TTI::TargetCostKind CostKind,
+ Type *SrcVTy, const Value *Ptr, Align Alignment,
+ unsigned AddressSpace, const Value *Mask) const {
assert(isa<VectorType>(SrcVTy) && "Unexpected type in getGSVectorCost");
unsigned VF = cast<FixedVectorType>(SrcVTy)->getNumElements();
@@ -6631,7 +6657,7 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
? Ptr->getType()->getScalarType()->getPointerAddressSpace()
: AddressSpace);
- auto getIndexSizeInBits = [PtrSizeInBits](const Value *Ptr) {
+ auto getIndexSizeInBits = [&](const Value *Ptr) {
unsigned IndexSize = PtrSizeInBits;
const GetElementPtrInst *GEP = dyn_cast_or_null<GetElementPtrInst>(Ptr);
if (IndexSize < 64 || !GEP)
@@ -6641,45 +6667,97 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
const Value *Ptrs = GEP->getPointerOperand();
if (Ptrs->getType()->isVectorTy() && !getSplatValue(Ptrs))
return IndexSize;
- for (unsigned I = 1, E = GEP->getNumOperands(); I != E; ++I) {
- if (isa<Constant>(GEP->getOperand(I)))
+ for (gep_type_iterator GTI = gep_type_begin(GEP), GTE = gep_type_end(GEP);
+ GTI != GTE; ++GTI) {
+ const Value *Operand = GTI.getOperand();
+ if (isa<Constant>(Operand))
continue;
- Type *IndxTy = GEP->getOperand(I)->getType();
+ Type *IndxTy = Operand->getType();
if (auto *IndexVTy = dyn_cast<VectorType>(IndxTy))
IndxTy = IndexVTy->getElementType();
- if ((IndxTy->getPrimitiveSizeInBits() == 64 &&
- !isa<SExtInst>(GEP->getOperand(I))) ||
+ if ((IndxTy->getPrimitiveSizeInBits() == 64 && !isa<SExtInst>(Operand)) ||
++NumOfVarIndices > 1)
return IndexSize; // 64
+ // The narrow index only reaches the instruction if the addressing mode
+ // can apply its stride as a scale, and the scale field encodes 1, 2, 4
+ // and 8. Any other stride has to be multiplied into the index first,
+ // and that product is pointer-width, so the operation ends up gathering
+ // with 64-bit indices however narrow the index started out.
+ TypeSize EltSize = DL.getTypeAllocSize(GTI.getIndexedType());
+ if (EltSize.isScalable())
+ return IndexSize; // 64
+ uint64_t Stride = EltSize.getFixedValue();
+ if (!isPowerOf2_64(Stride) || Stride > 8)
+ return IndexSize; // 64
}
return (unsigned)32;
};
- // Trying to reduce IndexSize to 32 bits for vector 16.
- // By default the IndexSize is equal to pointer size.
- unsigned IndexSize =
- (ST->hasAVX512() && VF >= 16) ? getIndexSizeInBits(Ptr) : PtrSizeInBits;
+ // The index width belongs to the operation as a whole, not to any one part,
+ // so derive it whenever a hardware gather/scatter is being priced. Asking
+ // only for AVX-512 at VF >= 16 left a dword-indexed gather costed as though
+ // its indices were pointer-width everywhere else, so it tied with the
+ // qword-indexed form that CodeGen splits into twice as many instructions.
+ unsigned IndexSize = getIndexSizeInBits(Ptr);
auto *IndexVTy = FixedVectorType::get(
IntegerType::get(SrcVTy->getContext(), IndexSize), VF);
std::pair<InstructionCost, MVT> IdxsLT = getTypeLegalizationCost(IndexVTy);
std::pair<InstructionCost, MVT> SrcLT = getTypeLegalizationCost(SrcVTy);
- InstructionCost::CostType SplitFactor =
- std::max(IdxsLT.first, SrcLT.first).getValue();
- // A vector of pointers that does not fit one register is split into
- // SplitFactor gather/scatter instructions, each paying the overhead.
+
+ // One instruction covers as many lanes as fit in a legal register once both
+ // the data and the indices are held, so the wider of the two bounds the lane
+ // count: a dword gather indexed by qwords fills a register with indices
+ // before it fills one with data. Counting legalized registers instead
+ // overcounts, because legalization first widens to a power of two -- a
+ // 24-lane qword-indexed gather occupies four index registers, but only three
+ // of them hold live lanes and CodeGen emits three instructions.
+ auto LanesOf = [](MVT VT) {
+ return VT.isVector() ? VT.getVectorNumElements() : 1U;
+ };
+ unsigned LanesPerPart =
+ std::max(std::min(LanesOf(IdxsLT.second), LanesOf(SrcLT.second)), 1U);
+
+ // Lanes the operation is declared over, and the parts they occupy.
+ InstructionCost::CostType ChargedLanes = VF;
+ InstructionCost::CostType EmittedParts = divideCeil(VF, LanesPerPart);
+
+ if (auto Active = getActiveGSLanesAndGroups(Mask, VF, LanesPerPart)) {
+ // A compile-time mask says which parts survive: one with no live lane
+ // gathers nothing, and CodeGen folds it away rather than emitting it.
+ ChargedLanes = Active->first;
+ EmittedParts = Active->second;
+ } else if (Opcode != Instruction::Load) {
+ // Without a compile-time mask the tail that legalization's widening added
+ // cannot be proved dead. A gather's tail produces an unused result and is
+ // dropped, but a scatter's survives as a store under a zeroed mask, which
+ // is still an instruction. How far the widening reaches is set by the
+ // predicate: AVX512BW holds it in a k-register as wide as the legalized
+ // vector, so the length rounds up to a power of two, while without it the
+ // mask is broken into 16-lane pieces and only rounds up to one of those.
+ unsigned WidenedLanes = PowerOf2Ceil(VF);
+ if (!ST->hasBWI())
+ WidenedLanes = std::min(WidenedLanes, alignTo(VF, 16u));
+ EmittedParts = divideCeil(WidenedLanes, LanesPerPart);
+ }
+
+ // Nothing survives, so no instruction is emitted.
+ if (EmittedParts == 0)
+ return 0;
+
+ // Each emitted instruction pays the overhead once.
if (CostKind == TTI::TCK_CodeSize)
- return SplitFactor;
+ return EmittedParts;
// The gather / scatter cost is given by Intel architects. It is a rough
// number since we are looking at one instruction in a time.
const int GSOverhead = (Opcode == Instruction::Load) ? getGatherOverhead()
: getScatterOverhead();
- // Charge every lane of the original vector. Descending into VF / SplitFactor
- // dropped the remainder, so a v9 gather was costed as eight lanes.
- return SplitFactor * GSOverhead +
- VF * getMemoryOpCost(Opcode, SrcVTy->getScalarType(), Alignment,
- AddressSpace, CostKind);
+ // Charge every lane that is actually accessed. Dividing the length by the
+ // part count dropped the remainder, so a v9 gather was costed as eight lanes.
+ return EmittedParts * GSOverhead +
+ ChargedLanes * getMemoryOpCost(Opcode, SrcVTy->getScalarType(),
+ Alignment, AddressSpace, CostKind);
}
/// Calculate the cost of Gather / Scatter operation
@@ -6704,8 +6782,18 @@ X86TTIImpl::getGatherScatterOpCost(const MemIntrinsicCostAttributes &MICA,
assert(SrcVTy->isVectorTy() && "Unexpected data type for Gather/Scatter");
unsigned AddressSpace = MICA.getAddressSpace();
- return getGSVectorCost(Opcode, CostKind, SrcVTy, Ptr, Alignment,
- AddressSpace);
+
+ // A mask known at compile time tells us which parts CodeGen keeps. Read it
+ // only from a real gather/scatter call: the context instruction can also be a
+ // plain load or store being considered for the transform, whose operands mean
+ // something else entirely.
+ const Value *Mask = nullptr;
+ if (const auto *II = dyn_cast_or_null<IntrinsicInst>(MICA.getInst()))
+ if (II->getIntrinsicID() == MICA.getID())
+ Mask = II->getArgOperand(IsLoad ? 1 : 2);
+
+ return getGSVectorCost(Opcode, CostKind, SrcVTy, Ptr, Alignment, AddressSpace,
+ Mask);
}
bool X86TTIImpl::isLSRCostLess(const TargetTransformInfo::LSRCost &C1,
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.h b/llvm/lib/Target/X86/X86TargetTransformInfo.h
index 422e4aad2316c..a130835173b5d 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.h
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.h
@@ -271,7 +271,8 @@ class X86TTIImpl final : public BasicTTIImplBase<X86TTIImpl> {
bool supportsGather() const;
InstructionCost getGSVectorCost(unsigned Opcode, TTI::TargetCostKind CostKind,
Type *DataTy, const Value *Ptr,
- Align Alignment, unsigned AddressSpace) const;
+ Align Alignment, unsigned AddressSpace,
+ const Value *Mask = nullptr) const;
int getGatherOverhead() const;
int getScatterOverhead() const;
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
index 0125c7975b7c5..2d904412f5d52 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
@@ -1,10 +1,16 @@
; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --version 6
; Gathers and scatters over pointers that are not in address space 0. On x86,
; address spaces 270 and 271 hold 32-bit pointers, so a vector of them is half
-; the width of the same number of address-space-0 pointers, needs no split, and
-; is indexed with a dword rather than a qword index.
+; the width of the same number of address-space-0 pointers, needs fewer parts,
+; and is indexed with a dword rather than a qword index.
+;
+; Both targets are covered because the two differ in what they expose. On the
+; AVX-512 target the narrower pointer removes a split that address space 0 still
+; needs, while on the AVX2 target it also changes the cost reached through a GEP,
+; whose index width is derived from the pointer's own address space.
; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake | FileCheck %s --check-prefix=SKL
target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i64:64-i128:128-f80:128-n8:16:32:64-S128"
@@ -14,6 +20,10 @@ define <16 x i32> @gather_v16i32_p270(<16 x ptr addrspace(270)> %ptrs, <16 x i1>
; SKX-LABEL: 'gather_v16i32_p270'
; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+; SKL-LABEL: 'gather_v16i32_p270'
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
;
%v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
ret <16 x i32> %v
@@ -24,6 +34,10 @@ define <16 x i32> @gather_v16i32_p0(<16 x ptr> %ptrs, <16 x i1> %mask) {
; SKX-LABEL: 'gather_v16i32_p0'
; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+; SKL-LABEL: 'gather_v16i32_p0'
+; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
;
%v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
ret <16 x i32> %v
@@ -35,6 +49,11 @@ define <16 x i32> @gather_v16i32_p270_gep(ptr addrspace(270) %base, <16 x i32> %
; SKX-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr addrspace(270) %base, <16 x i32> %idx
; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+; SKL-LABEL: 'gather_v16i32_p270_gep'
+; SKL-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr addrspace(270) %base, <16 x i32> %idx
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
;
%ptrs = getelementptr inbounds i32, ptr addrspace(270) %base, <16 x i32> %idx
%v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
@@ -45,6 +64,10 @@ define void @scatter_v16i32_p270(<16 x i32> %val, <16 x ptr addrspace(270)> %ptr
; SKX-LABEL: 'scatter_v16i32_p270'
; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p270(<16 x i32> %val, <16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask)
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+; SKL-LABEL: 'scatter_v16i32_p270'
+; SKL-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p270(<16 x i32> %val, <16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
;
call void @llvm.masked.scatter.v16i32.v16p270(<16 x i32> %val, <16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask)
ret void
@@ -54,16 +77,51 @@ define void @scatter_v16i32_p0(<16 x i32> %val, <16 x ptr> %ptrs, <16 x i1> %mas
; SKX-LABEL: 'scatter_v16i32_p0'
; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %val, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+; SKL-LABEL: 'scatter_v16i32_p0'
+; SKL-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %val, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
;
call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %val, <16 x ptr> %ptrs, i32 4, <16 x i1> %mask)
ret void
}
+; Address space 271 holds 32-bit pointers too, and must be treated like 270.
+define <16 x i32> @gather_v16i32_p271(<16 x ptr addrspace(271)> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'gather_v16i32_p271'
+; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p271(<16 x ptr addrspace(271)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+; SKL-LABEL: 'gather_v16i32_p271'
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p271(<16 x ptr addrspace(271)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+ %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p271(<16 x ptr addrspace(271)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+ ret <16 x i32> %v
+}
+
+define void @scatter_v16i32_p271(<16 x i32> %val, <16 x ptr addrspace(271)> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'scatter_v16i32_p271'
+; SKX-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p271(<16 x i32> %val, <16 x ptr addrspace(271)> align 4 %ptrs, <16 x i1> %mask)
+; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+; SKL-LABEL: 'scatter_v16i32_p271'
+; SKL-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p271(<16 x i32> %val, <16 x ptr addrspace(271)> align 4 %ptrs, <16 x i1> %mask)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+ call void @llvm.masked.scatter.v16i32.v16p271(<16 x i32> %val, <16 x ptr addrspace(271)> %ptrs, i32 4, <16 x i1> %mask)
+ ret void
+}
+
; Address space 272 holds 64-bit pointers, so it must behave like space 0.
define <16 x i32> @gather_v16i32_p272(<16 x ptr addrspace(272)> %ptrs, <16 x i1> %mask) {
; SKX-LABEL: 'gather_v16i32_p272'
; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p272(<16 x ptr addrspace(272)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+; SKL-LABEL: 'gather_v16i32_p272'
+; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p272(<16 x ptr addrspace(272)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
;
%v = call <16 x i32> @llvm.masked.gather.v16i32.v16p272(<16 x ptr addrspace(272)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
ret <16 x i32> %v
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll
new file mode 100644
index 0000000000000..fd6374b595275
--- /dev/null
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll
@@ -0,0 +1,445 @@
+; Pins the cost model's part count against the number of gather/scatter
+; instructions CodeGen actually emits, for both index widths on an AVX-512 and
+; an AVX2 target.
+;
+; The pairing is the point of the test: the code-size cost of a hardware
+; gather/scatter is its part count, so every cost check below has a matching
+; instruction count taken from the same module. Three properties would
+; otherwise drift unnoticed. A part count taken from the legalized register
+; count overstates the work when legalization widens to a power of two; an
+; index width derived only for wide AVX-512 vectors leaves a dword-indexed
+; operation tied with the qword-indexed form that CodeGen splits into more
+; instructions; and the tail a scatter still emits under a zeroed mask reaches
+; only as far as the predicate widens, which depends on AVX512BW.
+;
+; The throughput run pins the other half of the cost, the per-lane term, which
+; code size does not expose: it is charged for the lanes that are really
+; accessed, so it must follow the remainder lane of a length that does not fill
+; its parts, and must drop the lanes a compile-time mask kills.
+;
+; These checks are maintained by hand rather than by
+; update_analyze_test_checks.py, which does not know about the paired llc runs.
+
+; RUN: opt < %s -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" -cost-kind=code-size -disable-output -mcpu=skylake-avx512 2>&1 | FileCheck %s --check-prefix=SKX-COST
+; RUN: opt < %s -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" -cost-kind=throughput -disable-output -mcpu=skylake-avx512 2>&1 | FileCheck %s --check-prefix=SKX-TPUT
+; RUN: llc < %s -mtriple=x86_64-unknown-linux-gnu -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX-ASM
+; RUN: opt < %s -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" -cost-kind=code-size -disable-output -mcpu=skylake 2>&1 | FileCheck %s --check-prefix=AVX2-COST
+; RUN: llc < %s -mtriple=x86_64-unknown-linux-gnu -mcpu=skylake | FileCheck %s --check-prefix=AVX2-ASM
+; RUN: opt < %s -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" -cost-kind=code-size -disable-output -mcpu=knl 2>&1 | FileCheck %s --check-prefix=KNL-COST
+; RUN: llc < %s -mtriple=x86_64-unknown-linux-gnu -mcpu=knl | FileCheck %s --check-prefix=KNL-ASM
+
+; A length that fills its single part exactly, to sit against the nine-lane form
+; below: both are one instruction, but the ninth lane is still accessed and has
+; to be charged, so the throughput costs must differ by exactly one lane.
+; SKX-COST-LABEL: 'gather_v8i32_dword_index'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v8i32_dword_index'
+; SKX-TPUT: cost of 10 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v8i32_dword_index:
+; SKX-ASM-COUNT-1: vpgatherdd
+; SKX-ASM-NOT: vpgather
+define <8 x i32> @gather_v8i32_dword_index(ptr %base, <8 x i32> %idx, <8 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <8 x i32> %idx
+ %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x i32> poison)
+ ret <8 x i32> %v
+}
+
+; A vector length that is not a multiple of its part count: the remainder lane
+; still needs a part of its own once the index no longer fits alongside the rest.
+; SKX-COST-LABEL: 'gather_v9i32_dword_index'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v9i32_dword_index'
+; SKX-TPUT: cost of 11 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v9i32_dword_index:
+; SKX-ASM-COUNT-1: vpgatherdd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v9i32_dword_index'
+; AVX2-COST: cost of 2 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v9i32_dword_index:
+; AVX2-ASM-COUNT-2: vpgatherdd
+; AVX2-ASM-NOT: vpgather
+define <9 x i32> @gather_v9i32_dword_index(ptr %base, <9 x i32> %idx, <9 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <9 x i32> %idx
+ %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> %ptrs, i32 4, <9 x i1> %mask, <9 x i32> poison)
+ ret <9 x i32> %v
+}
+
+; SKX-COST-LABEL: 'gather_v9i32_qword_index'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v9i32_qword_index:
+; SKX-ASM-COUNT-2: vpgatherqd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v9i32_qword_index'
+; AVX2-COST: cost of 3 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v9i32_qword_index:
+; AVX2-ASM-COUNT-3: vpgatherqd
+; AVX2-ASM-NOT: vpgather
+define <9 x i32> @gather_v9i32_qword_index(ptr %base, <9 x i64> %idx, <9 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <9 x i64> %idx
+ %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> %ptrs, i32 4, <9 x i1> %mask, <9 x i32> poison)
+ ret <9 x i32> %v
+}
+
+; A length below the width at which the index used to be examined, so both index
+; widths would otherwise be priced the same on either target.
+; SKX-COST-LABEL: 'gather_v17i32_dword_index'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v17i32_dword_index:
+; SKX-ASM-COUNT-2: vpgatherdd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v17i32_dword_index'
+; AVX2-COST: cost of 3 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v17i32_dword_index:
+; AVX2-ASM-COUNT-3: vpgatherdd
+; AVX2-ASM-NOT: vpgather
+define <17 x i32> @gather_v17i32_dword_index(ptr %base, <17 x i32> %idx, <17 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <17 x i32> %idx
+ %v = call <17 x i32> @llvm.masked.gather.v17i32.v17p0(<17 x ptr> %ptrs, i32 4, <17 x i1> %mask, <17 x i32> poison)
+ ret <17 x i32> %v
+}
+
+; SKX-COST-LABEL: 'gather_v17i32_qword_index'
+; SKX-COST: cost of 3 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v17i32_qword_index:
+; SKX-ASM-COUNT-3: vpgatherqd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v17i32_qword_index'
+; AVX2-COST: cost of 5 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v17i32_qword_index:
+; AVX2-ASM-COUNT-5: vpgatherqd
+; AVX2-ASM-NOT: vpgather
+define <17 x i32> @gather_v17i32_qword_index(ptr %base, <17 x i64> %idx, <17 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <17 x i64> %idx
+ %v = call <17 x i32> @llvm.masked.gather.v17i32.v17p0(<17 x ptr> %ptrs, i32 4, <17 x i1> %mask, <17 x i32> poison)
+ ret <17 x i32> %v
+}
+
+; A length whose qword indices occupy four legal registers but fill only three
+; of them with live lanes.
+; SKX-COST-LABEL: 'gather_v24i32_dword_index'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_dword_index:
+; SKX-ASM-COUNT-2: vpgatherdd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v24i32_dword_index'
+; AVX2-COST: cost of 3 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v24i32_dword_index:
+; AVX2-ASM-COUNT-3: vpgatherdd
+; AVX2-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_dword_index(ptr %base, <24 x i32> %idx, <24 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> %mask, <24 x i32> poison)
+ ret <24 x i32> %v
+}
+
+; SKX-COST-LABEL: 'gather_v24i32_qword_index'
+; SKX-COST: cost of 3 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v24i32_qword_index'
+; SKX-TPUT: cost of 30 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_qword_index:
+; SKX-ASM-COUNT-3: vpgatherqd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v24i32_qword_index'
+; AVX2-COST: cost of 6 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v24i32_qword_index:
+; AVX2-ASM-COUNT-6: vpgatherqd
+; AVX2-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_qword_index(ptr %base, <24 x i64> %idx, <24 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> %mask, <24 x i32> poison)
+ ret <24 x i32> %v
+}
+
+; Scatters with a variable mask. AVX2 has no hardware scatter, so its cost comes
+; from scalarizing instead of from a part count, and it emits no scatter at all.
+; SKX-COST-LABEL: 'scatter_v9i32_dword_index_var_mask'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v9i32_dword_index_var_mask:
+; SKX-ASM-COUNT-1: vpscatterdd
+; SKX-ASM-NOT: vpscatter
+; AVX2-ASM-LABEL: scatter_v9i32_dword_index_var_mask:
+; AVX2-ASM-NOT: vpscatter
+define void @scatter_v9i32_dword_index_var_mask(ptr %base, <9 x i32> %idx, <9 x i32> %val, <9 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <9 x i32> %idx
+ call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> %ptrs, i32 4, <9 x i1> %mask)
+ ret void
+}
+
+; SKX-COST-LABEL: 'scatter_v9i32_qword_index_var_mask'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v9i32_qword_index_var_mask:
+; SKX-ASM-COUNT-2: vpscatterqd
+; SKX-ASM-NOT: vpscatter
+; AVX2-ASM-LABEL: scatter_v9i32_qword_index_var_mask:
+; AVX2-ASM-NOT: vpscatter
+define void @scatter_v9i32_qword_index_var_mask(ptr %base, <9 x i64> %idx, <9 x i32> %val, <9 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <9 x i64> %idx
+ call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> %ptrs, i32 4, <9 x i1> %mask)
+ ret void
+}
+
+; Scatters with every lane active: an all-ones mask leaves the part count alone.
+; SKX-COST-LABEL: 'scatter_v24i32_dword_index_all_active'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v24i32_dword_index_all_active:
+; SKX-ASM-COUNT-2: vpscatterdd
+; SKX-ASM-NOT: vpscatter
+; AVX2-ASM-LABEL: scatter_v24i32_dword_index_all_active:
+; AVX2-ASM-NOT: vpscatter
+define void @scatter_v24i32_dword_index_all_active(ptr %base, <24 x i32> %idx, <24 x i32> %val) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
+ call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %val, <24 x ptr> %ptrs, i32 4, <24 x i1> splat (i1 true))
+ ret void
+}
+
+; SKX-COST-LABEL: 'scatter_v24i32_qword_index_all_active'
+; SKX-COST: cost of 3 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v24i32_qword_index_all_active:
+; SKX-ASM-COUNT-3: vpscatterqd
+; SKX-ASM-NOT: vpscatter
+; AVX2-ASM-LABEL: scatter_v24i32_qword_index_all_active:
+; AVX2-ASM-NOT: vpscatter
+define void @scatter_v24i32_qword_index_all_active(ptr %base, <24 x i64> %idx, <24 x i32> %val) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %val, <24 x ptr> %ptrs, i32 4, <24 x i1> splat (i1 true))
+ ret void
+}
+
+; A variable mask on a length that legalization widens past a part boundary. The
+; widened tail cannot be proved dead, and it survives as a store under a zeroed
+; mask, so CodeGen emits a fourth instruction where the all-ones form above
+; emits three. That store is counted, since it is still an instruction.
+; SKX-COST-LABEL: 'scatter_v24i32_qword_index_var_mask'
+; SKX-COST: cost of 4 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v24i32_qword_index_var_mask:
+; SKX-ASM-COUNT-3: vpscatterqd
+; SKX-ASM: kxor
+; SKX-ASM-COUNT-1: vpscatterqd
+; SKX-ASM-NOT: vpscatter
+; AVX2-ASM-LABEL: scatter_v24i32_qword_index_var_mask:
+; AVX2-ASM-NOT: vpscatter
+define void @scatter_v24i32_qword_index_var_mask(ptr %base, <24 x i64> %idx, <24 x i32> %val, <24 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %val, <24 x ptr> %ptrs, i32 4, <24 x i1> %mask)
+ ret void
+}
+
+; A compile-time mask says which parts survive. With only the first eight lanes
+; live, two of these three parts hold nothing and are folded away, leaving one
+; instruction rather than the three the declared length would suggest.
+; The dead lanes are not charged either: against the 30 the same shape costs
+; with an unknown mask above, this is the cost of the eight lanes it reaches.
+; SKX-COST-LABEL: 'gather_v24i32_qword_index_first8_mask'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v24i32_qword_index_first8_mask'
+; SKX-TPUT: cost of 10 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_qword_index_first8_mask:
+; SKX-ASM-COUNT-1: vpgatherqd
+; SKX-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_qword_index_first8_mask(ptr %base, <24 x i64> %idx) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> <i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false>, <24 x i32> poison)
+ ret <24 x i32> %v
+}
+
+; The same for a scatter, where the dead parts would otherwise be the zeroed
+; stores counted above.
+; SKX-COST-LABEL: 'scatter_v24i32_qword_index_first8_mask'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v24i32_qword_index_first8_mask:
+; SKX-ASM-COUNT-1: vpscatterqd
+; SKX-ASM-NOT: vpscatter
+define void @scatter_v24i32_qword_index_first8_mask(ptr %base, <24 x i64> %idx, <24 x i32> %val) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %val, <24 x ptr> %ptrs, i32 4, <24 x i1> <i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false>)
+ ret void
+}
+
+; A mask with no live lane leaves nothing to do, and CodeGen emits no gather at
+; all, so the operation is free rather than costed as three parts.
+; SKX-COST-LABEL: 'gather_v24i32_qword_index_zero_mask'
+; SKX-COST: cost of 0 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v24i32_qword_index_zero_mask'
+; SKX-TPUT: cost of 0 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_qword_index_zero_mask:
+; SKX-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_qword_index_zero_mask(ptr %base, <24 x i64> %idx) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> zeroinitializer, <24 x i32> poison)
+ ret <24 x i32> %v
+}
+
+; The same for a scatter, which unlike a gather would otherwise still be charged
+; for the tail it emits under a zeroed mask.
+; SKX-COST-LABEL: 'scatter_v24i32_qword_index_zero_mask'
+; SKX-COST: cost of 0 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v24i32_qword_index_zero_mask:
+; SKX-ASM-NOT: vpscatter
+define void @scatter_v24i32_qword_index_zero_mask(ptr %base, <24 x i64> %idx, <24 x i32> %val) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %val, <24 x ptr> %ptrs, i32 4, <24 x i1> zeroinitializer)
+ ret void
+}
+
+; The live lanes need not be a prefix. Here they sit in the first and last part
+; with a dead one between, so it is the parts holding a live lane that are
+; counted, not the number of live lanes rounded up to a part.
+; SKX-COST-LABEL: 'gather_v24i32_qword_index_gapped_mask'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v24i32_qword_index_gapped_mask'
+; SKX-TPUT: cost of 6 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_qword_index_gapped_mask:
+; SKX-ASM-COUNT-2: vpgatherqd
+; SKX-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_qword_index_gapped_mask(ptr %base, <24 x i64> %idx) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> <i1 true, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 true, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false>, <24 x i32> poison)
+ ret <24 x i32> %v
+}
+
+; Live lanes confined to the last part, the mirror of the first-eight case, so
+; that a part is not counted merely because earlier parts were.
+; SKX-COST-LABEL: 'gather_v24i32_qword_index_last8_mask'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_qword_index_last8_mask:
+; SKX-ASM-COUNT-1: vpgatherqd
+; SKX-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_qword_index_last8_mask(ptr %base, <24 x i64> %idx) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+ %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> <i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true>, <24 x i32> poison)
+ ret <24 x i32> %v
+}
+
+; How far that widened tail reaches is set by the predicate rather than by the
+; data. AVX512BW holds the mask in a k-register as wide as the legalized
+; vector, so a 48-lane length rounds up to 64 and eight parts are emitted;
+; without it the mask is broken into 16-lane pieces, the length rounds up only
+; to 48, and six are. Counting legal registers would give eight on both.
+; SKX-COST-LABEL: 'scatter_v48i32_qword_index_var_mask'
+; SKX-COST: cost of 8 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v48i32_qword_index_var_mask:
+; SKX-ASM-COUNT-8: vpscatterqd
+; SKX-ASM-NOT: vpscatter
+; KNL-COST-LABEL: 'scatter_v48i32_qword_index_var_mask'
+; KNL-COST: cost of 6 for instruction: {{.*}}masked.scatter
+; KNL-ASM-LABEL: scatter_v48i32_qword_index_var_mask:
+; KNL-ASM-COUNT-6: vpscatterqd
+; KNL-ASM-NOT: vpscatter
+define void @scatter_v48i32_qword_index_var_mask(ptr %base, <48 x i64> %idx, <48 x i32> %val, <48 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <48 x i64> %idx
+ call void @llvm.masked.scatter.v48i32.v48p0(<48 x i32> %val, <48 x ptr> %ptrs, i32 4, <48 x i1> %mask)
+ ret void
+}
+
+; The same split with dword indices, where a part covers sixteen lanes instead
+; of eight: four parts with the wide predicate, three without it.
+; SKX-COST-LABEL: 'scatter_v48i32_dword_index_var_mask'
+; SKX-COST: cost of 4 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v48i32_dword_index_var_mask:
+; SKX-ASM-COUNT-4: vpscatterdd
+; SKX-ASM-NOT: vpscatter
+; KNL-COST-LABEL: 'scatter_v48i32_dword_index_var_mask'
+; KNL-COST: cost of 3 for instruction: {{.*}}masked.scatter
+; KNL-ASM-LABEL: scatter_v48i32_dword_index_var_mask:
+; KNL-ASM-COUNT-3: vpscatterdd
+; KNL-ASM-NOT: vpscatter
+define void @scatter_v48i32_dword_index_var_mask(ptr %base, <48 x i32> %idx, <48 x i32> %val, <48 x i1> %mask) {
+ %ptrs = getelementptr inbounds i32, ptr %base, <48 x i32> %idx
+ call void @llvm.masked.scatter.v48i32.v48p0(<48 x i32> %val, <48 x ptr> %ptrs, i32 4, <48 x i1> %mask)
+ ret void
+}
+
+; Pointers are gathered as often as integers are, and the element size that
+; decides how many lanes a part holds has to come from the pointer's own width
+; rather than from an integer type. Eight pointer-sized lanes fill one AVX-512
+; part and two AVX2 parts.
+; SKX-COST-LABEL: 'gather_v8ptr_qword_index'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v8ptr_qword_index'
+; SKX-TPUT: cost of 10 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v8ptr_qword_index:
+; SKX-ASM-COUNT-1: vpgatherqq
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v8ptr_qword_index'
+; AVX2-COST: cost of 2 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v8ptr_qword_index:
+; AVX2-ASM-COUNT-2: vpgatherqq
+; AVX2-ASM-NOT: vpgatherqq
+define <8 x ptr> @gather_v8ptr_qword_index(ptr %base, <8 x i64> %idx, <8 x i1> %mask) {
+ %ptrs = getelementptr inbounds ptr, ptr %base, <8 x i64> %idx
+ %r = call <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr> %ptrs, i32 8, <8 x i1> %mask, <8 x ptr> poison)
+ ret <8 x ptr> %r
+}
+
+; The same gather with only its first two lanes left live by a compile-time
+; mask. The second AVX2 part holds no live lane and is not emitted, and both
+; targets charge two lanes rather than eight.
+; SKX-COST-LABEL: 'gather_v8ptr_qword_index_first2_mask'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v8ptr_qword_index_first2_mask'
+; SKX-TPUT: cost of 4 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v8ptr_qword_index_first2_mask:
+; SKX-ASM-COUNT-1: vpgatherqq
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v8ptr_qword_index_first2_mask'
+; AVX2-COST: cost of 1 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v8ptr_qword_index_first2_mask:
+; AVX2-ASM-COUNT-1: vpgatherqq
+; AVX2-ASM-NOT: vpgatherqq
+define <8 x ptr> @gather_v8ptr_qword_index_first2_mask(ptr %base, <8 x i64> %idx) {
+ %ptrs = getelementptr inbounds ptr, ptr %base, <8 x i64> %idx
+ %r = call <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr> %ptrs, i32 8, <8 x i1> <i1 1, i1 1, i1 0, i1 0, i1 0, i1 0, i1 0, i1 0>, <8 x ptr> poison)
+ ret <8 x ptr> %r
+}
+
+; A narrow index only stays narrow if the addressing mode can apply its stride
+; as a scale, and the scale field encodes 1, 2, 4 and 8. This GEP walks an
+; array of three-word structures, the form a loop vectorizer emits for
+; a[idx[i]].f, so its stride is twelve: the multiply is folded into the index
+; and the gather ends up indexed by qwords, taking twice the instructions of
+; the four-byte-stride control below.
+; SKX-COST-LABEL: 'gather_v16i32_struct_stride'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v16i32_struct_stride:
+; SKX-ASM-COUNT-2: vpgatherqd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v16i32_struct_stride'
+; AVX2-COST: cost of 4 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v16i32_struct_stride:
+; AVX2-ASM-COUNT-4: vpgatherqd
+; AVX2-ASM-NOT: vpgatherqd
+define <16 x i32> @gather_v16i32_struct_stride(ptr %base, <16 x i32> %idx, <16 x i1> %mask) {
+ %sext = sext <16 x i32> %idx to <16 x i64>
+ %ptrs = getelementptr inbounds {i32, i32, i32}, ptr %base, <16 x i64> %sext, i32 0
+ %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+ ret <16 x i32> %v
+}
+
+; The same gather over a four-byte stride, which the scale field does encode,
+; so the index stays a dword and one instruction covers twice the lanes.
+; SKX-COST-LABEL: 'gather_v16i32_scaled_stride'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v16i32_scaled_stride:
+; SKX-ASM-COUNT-1: vpgatherdd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v16i32_scaled_stride'
+; AVX2-COST: cost of 2 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v16i32_scaled_stride:
+; AVX2-ASM-COUNT-2: vpgatherdd
+; AVX2-ASM-NOT: vpgatherdd
+define <16 x i32> @gather_v16i32_scaled_stride(ptr %base, <16 x i32> %idx, <16 x i1> %mask) {
+ %sext = sext <16 x i32> %idx to <16 x i64>
+ %ptrs = getelementptr inbounds i32, ptr %base, <16 x i64> %sext
+ %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+ ret <16 x i32> %v
+}
+
+declare <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr>, i32, <16 x i1>, <16 x i32>)
+declare <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr>, i32, <8 x i1>, <8 x ptr>)
+declare <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr>, i32, <8 x i1>, <8 x i32>)
+declare <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr>, i32, <9 x i1>, <9 x i32>)
+declare <17 x i32> @llvm.masked.gather.v17i32.v17p0(<17 x ptr>, i32, <17 x i1>, <17 x i32>)
+declare <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr>, i32, <24 x i1>, <24 x i32>)
+declare void @llvm.masked.scatter.v9i32.v9p0(<9 x i32>, <9 x ptr>, i32, <9 x i1>)
+declare void @llvm.masked.scatter.v24i32.v24p0(<24 x i32>, <24 x ptr>, i32, <24 x i1>)
+declare void @llvm.masked.scatter.v48i32.v48p0(<48 x i32>, <48 x ptr>, i32, <48 x i1>)
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
index 26954697c2b3f..28a80f6d8f84a 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
@@ -1,15 +1,20 @@
; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --version 6
-; Costs for masked gather/scatter operations that need more than one register's
-; worth of pointers, across every cost kind.
+; Costs for masked gather/scatter operations that need one register's worth of
+; pointers or more, across every cost kind.
;
; Two properties are pinned here. A vector length that is not a multiple of its
-; split factor must still be charged for all of its lanes. And the index width
-; is chosen once for the whole operation, so a length wide enough to qualify for
-; narrowing keeps the narrow index in each of its parts, rather than reverting
-; to pointer width because an individual part is too short to qualify.
+; part count must still be charged for all of its lanes. And the index width is
+; chosen once for the whole operation, from the indices themselves, so the same
+; data type is priced differently depending on how wide its indices are.
+;
+; The code-size cost of a hardware gather/scatter is its part count, so those
+; numbers can be read against CodeGen directly; the paired instruction counts
+; live in masked-gather-scatter-part-count.ll. Both targets gather in
+; hardware, but only the AVX-512 one scatters, so the AVX2 scatter below is
+; the cost of scalarizing rather than a part count.
; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX
-; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=haswell | FileCheck %s --check-prefix=AVX2
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake | FileCheck %s --check-prefix=AVX2
; A length that divides its split factor: unchanged, and the reference point for
; the two cases below.
@@ -19,7 +24,7 @@ define <8 x i32> @gather_v8i32(<8 x ptr> %ptrs, <8 x i1> %mask) {
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <8 x i32> %v
;
; AVX2-LABEL: 'gather_v8i32'
-; AVX2-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; AVX2-NEXT: Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <8 x i32> %v
;
%v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x i32> poison)
@@ -34,7 +39,7 @@ define <9 x i32> @gather_v9i32(<9 x ptr> %ptrs, <9 x i1> %mask) {
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <9 x i32> %v
;
; AVX2-LABEL: 'gather_v9i32'
-; AVX2-NEXT: Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 4 %ptrs, <9 x i1> %mask, <9 x i32> poison)
+; AVX2-NEXT: Cost Model: Found costs of RThru:15 CodeSize:3 Lat:42 SizeLat:15 for: %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 4 %ptrs, <9 x i1> %mask, <9 x i32> poison)
; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <9 x i32> %v
;
%v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> %ptrs, i32 4, <9 x i1> %mask, <9 x i32> poison)
@@ -54,9 +59,9 @@ define void @scatter_v9i32(<9 x i32> %val, <9 x ptr> %ptrs, <9 x i1> %mask) {
ret void
}
-; Index width across a split: the GEP indices are 32-bit, and v24 is wide enough
-; to qualify for narrowing, so all three parts are priced with a dword index
-; even though a single part on its own would not qualify.
+; Index width across a split: the GEP indices are 32-bit, so the operation is
+; priced with a dword index. Sixteen dword-indexed lanes fill a part, so these
+; 24 lanes take two, and CodeGen emits two vpgatherdd.
define <24 x i32> @gather_v24i32_dword_index(ptr %base, <24 x i32> %idx, <24 x i1> %mask) {
; SKX-LABEL: 'gather_v24i32_dword_index'
; SKX-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
@@ -65,7 +70,7 @@ define <24 x i32> @gather_v24i32_dword_index(ptr %base, <24 x i32> %idx, <24 x i
;
; AVX2-LABEL: 'gather_v24i32_dword_index'
; AVX2-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
-; AVX2-NEXT: Cost Model: Found costs of RThru:82 CodeSize:106 Lat:178 SizeLat:106 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; AVX2-NEXT: Cost Model: Found costs of RThru:30 CodeSize:3 Lat:102 SizeLat:30 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
;
%ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
@@ -73,16 +78,19 @@ define <24 x i32> @gather_v24i32_dword_index(ptr %base, <24 x i32> %idx, <24 x i
ret <24 x i32> %v
}
-; Control for the above: genuinely 64-bit indices cannot narrow.
+; Control for the above: genuinely 64-bit indices cannot narrow, so only eight
+; lanes fill a part and the same 24 lanes take three vpgatherqd. Their indices
+; occupy four legal registers, one of which holds no live lane, so a part count
+; taken from the register count would overstate the work by one.
define <24 x i32> @gather_v24i32_qword_index(ptr %base, <24 x i64> %idx, <24 x i1> %mask) {
; SKX-LABEL: 'gather_v24i32_qword_index'
; SKX-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
-; SKX-NEXT: Cost Model: Found costs of RThru:32 CodeSize:4 Lat:104 SizeLat:32 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; SKX-NEXT: Cost Model: Found costs of RThru:30 CodeSize:3 Lat:102 SizeLat:30 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
; SKX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
;
; AVX2-LABEL: 'gather_v24i32_qword_index'
; AVX2-NEXT: Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
-; AVX2-NEXT: Cost Model: Found costs of RThru:82 CodeSize:106 Lat:178 SizeLat:106 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; AVX2-NEXT: Cost Model: Found costs of RThru:36 CodeSize:6 Lat:108 SizeLat:36 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
;
%ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
diff --git a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll
index f835a8aedbbcc..bf19e53e1a374 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll
@@ -1909,7 +1909,7 @@ define <16 x float> @test_gather_16f32_const_mask(ptr %base, <16 x i32> %ind) {
; SKL-LABEL: 'test_gather_16f32_const_mask'
; SKL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SKL-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX512-LABEL: 'test_gather_16f32_const_mask'
@@ -1953,7 +1953,7 @@ define <16 x float> @test_gather_16f32_var_mask(ptr %base, <16 x i32> %ind, <16
; SKL-LABEL: 'test_gather_16f32_var_mask'
; SKL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SKL-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX512-LABEL: 'test_gather_16f32_var_mask'
@@ -2051,7 +2051,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
; SKL-NEXT: Cost Model: Found costs of RThru:1 CodeSize:1 Lat:3 SizeLat:2 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> poison, <16 x i32> zeroinitializer
; SKL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SKL-NEXT: Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX512-LABEL: 'test_gather_16f32_const_mask2'
diff --git a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
index c02dd6db8b99f..5797e9bd31017 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
@@ -745,7 +745,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:18 SizeLat:12 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:77 CodeSize:93 Lat:141 SizeLat:93 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SSE2-NEXT: Cost Model: Found costs of RThru:49 CodeSize:58 Lat:85 SizeLat:58 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; SSE2-NEXT: Cost Model: Found costs of RThru:49 CodeSize:58 Lat:85 SizeLat:58 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
; SSE2-NEXT: Cost Model: Found costs of RThru:39 CodeSize:47 Lat:71 SizeLat:47 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:20 CodeSize:24 Lat:36 SizeLat:24 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:18 SizeLat:12 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -773,7 +773,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:49 CodeSize:65 Lat:113 SizeLat:65 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SSE42-NEXT: Cost Model: Found costs of RThru:28 CodeSize:37 Lat:64 SizeLat:37 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; SSE42-NEXT: Cost Model: Found costs of RThru:28 CodeSize:37 Lat:64 SizeLat:37 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
; SSE42-NEXT: Cost Model: Found costs of RThru:25 CodeSize:33 Lat:57 SizeLat:33 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:13 CodeSize:17 Lat:29 SizeLat:17 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -801,7 +801,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:119 SizeLat:71 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; AVX1-NEXT: Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; AVX1-NEXT: Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
; AVX1-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:30 SizeLat:18 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -829,7 +829,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:119 SizeLat:71 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; AVX2-NEXT: Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; AVX2-NEXT: Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
; AVX2-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:30 SizeLat:18 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -857,7 +857,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SKL-NEXT: Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SKL-NEXT: Cost Model: Found costs of RThru:17 CodeSize:4 Lat:44 SizeLat:17 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:15 CodeSize:3 Lat:42 SizeLat:15 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
; SKL-NEXT: Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SKL-NEXT: Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -885,7 +885,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; KNL-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; KNL-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
; KNL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:17 CodeSize:21 Lat:33 SizeLat:21 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -913,7 +913,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
; SKX-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -943,7 +943,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
%V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> undef, i32 1, <1 x i1> %m1, <1 x i64> undef)
%V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> undef, i32 1, <16 x i1> %m16, <16 x i32> undef)
- %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> undef, i32 1, <9 x i1> %m9, <9 x i32> undef)
+ %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> poison, i32 1, <9 x i1> %m9, <9 x i32> poison)
%V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> undef, i32 1, <8 x i1> %m8, <8 x i32> undef)
%V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> undef, i32 1, <4 x i1> %m4, <4 x i32> undef)
%V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> undef, i32 1, <2 x i1> %m2, <2 x i32> undef)
@@ -976,7 +976,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:12 SizeLat:12 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SSE2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SSE2-NEXT: Cost Model: Found costs of RThru:77 CodeSize:93 Lat:93 SizeLat:93 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; SSE2-NEXT: Cost Model: Found costs of RThru:43 CodeSize:52 Lat:52 SizeLat:52 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; SSE2-NEXT: Cost Model: Found costs of RThru:43 CodeSize:52 Lat:52 SizeLat:52 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
; SSE2-NEXT: Cost Model: Found costs of RThru:39 CodeSize:47 Lat:47 SizeLat:47 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SSE2-NEXT: Cost Model: Found costs of RThru:20 CodeSize:24 Lat:24 SizeLat:24 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SSE2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:12 Lat:12 SizeLat:12 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1004,7 +1004,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SSE42-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SSE42-NEXT: Cost Model: Found costs of RThru:49 CodeSize:65 Lat:65 SizeLat:65 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; SSE42-NEXT: Cost Model: Found costs of RThru:28 CodeSize:37 Lat:37 SizeLat:37 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; SSE42-NEXT: Cost Model: Found costs of RThru:28 CodeSize:37 Lat:37 SizeLat:37 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
; SSE42-NEXT: Cost Model: Found costs of RThru:25 CodeSize:33 Lat:33 SizeLat:33 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SSE42-NEXT: Cost Model: Found costs of RThru:13 CodeSize:17 Lat:17 SizeLat:17 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SSE42-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1032,7 +1032,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; AVX1-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; AVX1-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; AVX1-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; AVX1-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
; AVX1-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; AVX1-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; AVX1-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1060,7 +1060,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; AVX2-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; AVX2-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; AVX2-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; AVX2-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
; AVX2-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; AVX2-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; AVX2-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1088,7 +1088,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SKL-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SKL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SKL-NEXT: Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; SKL-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; SKL-NEXT: Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
; SKL-NEXT: Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SKL-NEXT: Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SKL-NEXT: Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1116,7 +1116,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; KNL-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; KNL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; KNL-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; KNL-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
; KNL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; KNL-NEXT: Cost Model: Found costs of RThru:17 CodeSize:21 Lat:21 SizeLat:21 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; KNL-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1144,7 +1144,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
; SKX-NEXT: Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
; SKX-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; SKX-NEXT: Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
; SKX-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
; SKX-NEXT: Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
; SKX-NEXT: Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1174,7 +1174,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> undef, i32 1, <1 x i1> %m1)
call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> undef, i32 1, <16 x i1> %m16)
- call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> undef, i32 1, <9 x i1> %m9)
+ call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> poison, i32 1, <9 x i1> %m9)
call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> undef, i32 1, <8 x i1> %m8)
call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> undef, i32 1, <4 x i1> %m4)
call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> undef, i32 1, <2 x i1> %m2)
@@ -1901,43 +1901,43 @@ define <16 x float> @test_gather_16f32_const_mask(ptr %base, <16 x i32> %ind) {
; SSE2-LABEL: 'test_gather_16f32_const_mask'
; SSE2-NEXT: Cost Model: Found costs of RThru:16 CodeSize:8 Lat:8 SizeLat:8 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SSE2-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SSE2-NEXT: Cost Model: Found costs of RThru:60 CodeSize:60 Lat:108 SizeLat:60 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; SSE2-NEXT: Cost Model: Found costs of RThru:60 CodeSize:60 Lat:108 SizeLat:60 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> poison)
; SSE2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; SSE42-LABEL: 'test_gather_16f32_const_mask'
; SSE42-NEXT: Cost Model: Found costs of 8 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SSE42-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SSE42-NEXT: Cost Model: Found costs of RThru:44 CodeSize:44 Lat:92 SizeLat:44 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; SSE42-NEXT: Cost Model: Found costs of RThru:44 CodeSize:44 Lat:92 SizeLat:44 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> poison)
; SSE42-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX1-LABEL: 'test_gather_16f32_const_mask'
; AVX1-NEXT: Cost Model: Found costs of RThru:14 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; AVX1-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; AVX1-NEXT: Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; AVX1-NEXT: Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> poison)
; AVX1-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX2-LABEL: 'test_gather_16f32_const_mask'
; AVX2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; AVX2-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; AVX2-NEXT: Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; AVX2-NEXT: Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> poison)
; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; SKL-LABEL: 'test_gather_16f32_const_mask'
; SKL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SKL-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> poison)
; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX512-LABEL: 'test_gather_16f32_const_mask'
; AVX512-NEXT: Cost Model: Found costs of RThru:3 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; AVX512-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; AVX512-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; AVX512-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> poison)
; AVX512-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
%sext_ind = sext <16 x i32> %ind to <16 x i64>
%gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
- %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> %gep.v, i32 4, <16 x i1> <i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true>, <16 x float> undef)
+ %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> %gep.v, i32 4, <16 x i1> <i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true>, <16 x float> poison)
ret <16 x float>%res
}
@@ -1945,43 +1945,43 @@ define <16 x float> @test_gather_16f32_var_mask(ptr %base, <16 x i32> %ind, <16
; SSE2-LABEL: 'test_gather_16f32_var_mask'
; SSE2-NEXT: Cost Model: Found costs of RThru:16 CodeSize:8 Lat:8 SizeLat:8 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SSE2-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SSE2-NEXT: Cost Model: Found costs of RThru:61 CodeSize:77 Lat:125 SizeLat:77 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SSE2-NEXT: Cost Model: Found costs of RThru:61 CodeSize:77 Lat:125 SizeLat:77 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> poison)
; SSE2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; SSE42-LABEL: 'test_gather_16f32_var_mask'
; SSE42-NEXT: Cost Model: Found costs of 8 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SSE42-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SSE42-NEXT: Cost Model: Found costs of RThru:45 CodeSize:61 Lat:109 SizeLat:61 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SSE42-NEXT: Cost Model: Found costs of RThru:45 CodeSize:61 Lat:109 SizeLat:61 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> poison)
; SSE42-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX1-LABEL: 'test_gather_16f32_var_mask'
; AVX1-NEXT: Cost Model: Found costs of RThru:14 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; AVX1-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; AVX1-NEXT: Cost Model: Found costs of RThru:51 CodeSize:67 Lat:115 SizeLat:67 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; AVX1-NEXT: Cost Model: Found costs of RThru:51 CodeSize:67 Lat:115 SizeLat:67 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> poison)
; AVX1-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX2-LABEL: 'test_gather_16f32_var_mask'
; AVX2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; AVX2-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; AVX2-NEXT: Cost Model: Found costs of RThru:51 CodeSize:67 Lat:115 SizeLat:67 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; AVX2-NEXT: Cost Model: Found costs of RThru:51 CodeSize:67 Lat:115 SizeLat:67 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> poison)
; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; SKL-LABEL: 'test_gather_16f32_var_mask'
; SKL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SKL-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> poison)
; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX512-LABEL: 'test_gather_16f32_var_mask'
; AVX512-NEXT: Cost Model: Found costs of RThru:3 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; AVX512-NEXT: Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; AVX512-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; AVX512-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> poison)
; AVX512-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
%sext_ind = sext <16 x i32> %ind to <16 x i64>
%gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
- %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> %gep.v, i32 4, <16 x i1> %mask, <16 x float> undef)
+ %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> %gep.v, i32 4, <16 x i1> %mask, <16 x float> poison)
ret <16 x float>%res
}
@@ -2035,7 +2035,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
; SSE2-NEXT: Cost Model: Found costs of 1 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> undef, <16 x i32> zeroinitializer
; SSE2-NEXT: Cost Model: Found costs of RThru:16 CodeSize:8 Lat:8 SizeLat:8 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SSE2-NEXT: Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; SSE2-NEXT: Cost Model: Found costs of RThru:60 CodeSize:60 Lat:108 SizeLat:60 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; SSE2-NEXT: Cost Model: Found costs of RThru:60 CodeSize:60 Lat:108 SizeLat:60 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> poison)
; SSE2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; SSE42-LABEL: 'test_gather_16f32_const_mask2'
@@ -2043,7 +2043,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
; SSE42-NEXT: Cost Model: Found costs of 1 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> undef, <16 x i32> zeroinitializer
; SSE42-NEXT: Cost Model: Found costs of 8 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SSE42-NEXT: Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; SSE42-NEXT: Cost Model: Found costs of RThru:44 CodeSize:44 Lat:92 SizeLat:44 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; SSE42-NEXT: Cost Model: Found costs of RThru:44 CodeSize:44 Lat:92 SizeLat:44 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> poison)
; SSE42-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX1-LABEL: 'test_gather_16f32_const_mask2'
@@ -2051,7 +2051,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
; AVX1-NEXT: Cost Model: Found costs of RThru:2 CodeSize:2 Lat:3 SizeLat:3 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> undef, <16 x i32> zeroinitializer
; AVX1-NEXT: Cost Model: Found costs of RThru:14 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; AVX1-NEXT: Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; AVX1-NEXT: Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; AVX1-NEXT: Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> poison)
; AVX1-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX2-LABEL: 'test_gather_16f32_const_mask2'
@@ -2059,7 +2059,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
; AVX2-NEXT: Cost Model: Found costs of RThru:1 CodeSize:1 Lat:3 SizeLat:2 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> undef, <16 x i32> zeroinitializer
; AVX2-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; AVX2-NEXT: Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; AVX2-NEXT: Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; AVX2-NEXT: Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> poison)
; AVX2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; SKL-LABEL: 'test_gather_16f32_const_mask2'
@@ -2067,7 +2067,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
; SKL-NEXT: Cost Model: Found costs of RThru:1 CodeSize:1 Lat:3 SizeLat:2 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> undef, <16 x i32> zeroinitializer
; SKL-NEXT: Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; SKL-NEXT: Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; SKL-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT: Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> poison)
; SKL-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
; AVX512-LABEL: 'test_gather_16f32_const_mask2'
@@ -2075,7 +2075,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
; AVX512-NEXT: Cost Model: Found costs of RThru:1 CodeSize:1 Lat:3 SizeLat:1 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> undef, <16 x i32> zeroinitializer
; AVX512-NEXT: Cost Model: Found costs of RThru:3 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
; AVX512-NEXT: Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; AVX512-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; AVX512-NEXT: Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> poison)
; AVX512-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
;
%broadcast.splatinsert = insertelement <16 x ptr> undef, ptr %base, i32 0
@@ -2084,7 +2084,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
%sext_ind = sext <16 x i32> %ind to <16 x i64>
%gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
- %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> %gep.random, i32 4, <16 x i1> <i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true>, <16 x float> undef)
+ %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> %gep.random, i32 4, <16 x i1> <i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true>, <16 x float> poison)
ret <16 x float>%res
}
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll
index 2d5a30019bacd..371338b6bcb78 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll
@@ -50,9 +50,9 @@ define void @test() {
; AVX2-FASTGATHER: LV: Found an estimated cost of 1 for VF 1 For instruction: %valB = load i32, ptr %inB, align 4
; AVX2-FASTGATHER: Cost of 4 for VF 2: WIDEN ir<%valB> = load ir<%inB>
; AVX2-FASTGATHER: Cost of 6 for VF 4: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER: Cost of 12 for VF 8: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER: Cost of 24 for VF 16: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER: Cost of 48 for VF 32: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER: Cost of 10 for VF 8: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER: Cost of 20 for VF 16: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER: Cost of 40 for VF 32: WIDEN ir<%valB> = load ir<%inB>
;
; AVX512-LABEL: 'test'
; AVX512: LV: Found an estimated cost of 1 for VF 1 For instruction: %valB = load i32, ptr %inB, align 4
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll
index 8f5da77027970..0cf2b6b9c12b2 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll
@@ -43,9 +43,9 @@ define void @test() {
; AVX2-FASTGATHER: LV: Found an estimated cost of 1 for VF 1 For instruction: %valB.loaded = load i32, ptr %inB, align 4
; AVX2-FASTGATHER: Cost of 4 for VF 2: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
; AVX2-FASTGATHER: Cost of 6 for VF 4: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER: Cost of 12 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER: Cost of 24 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER: Cost of 48 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER: Cost of 10 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER: Cost of 20 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER: Cost of 40 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
;
; AVX512-LABEL: 'test'
; AVX512: LV: Found an estimated cost of 1 for VF 1 For instruction: %valB.loaded = load i32, ptr %inB, align 4
diff --git a/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll b/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll
index 7f0ba19d3e7a8..bc2bea9708227 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll
@@ -123,28 +123,28 @@ define void @replicate_sext(i32 %N, ptr %dst, ptr %src) #0 {
; CHECK-NEXT: [[FOUND_CONFLICT:%.*]] = and i1 [[BOUND0]], [[BOUND1]]
; CHECK-NEXT: br i1 [[FOUND_CONFLICT]], label %[[SCALAR_PH]], label %[[VECTOR_PH:.*]]
; CHECK: [[VECTOR_PH]]:
-; CHECK-NEXT: [[N_MOD_VF:%.*]] = and i32 [[TMP0]], 3
+; CHECK-NEXT: [[N_MOD_VF:%.*]] = and i32 [[TMP0]], 7
; CHECK-NEXT: [[TMP18:%.*]] = icmp eq i32 [[N_MOD_VF]], 0
-; CHECK-NEXT: [[TMP19:%.*]] = select i1 [[TMP18]], i32 4, i32 [[N_MOD_VF]]
+; CHECK-NEXT: [[TMP19:%.*]] = select i1 [[TMP18]], i32 8, i32 [[N_MOD_VF]]
; CHECK-NEXT: [[N_VEC:%.*]] = sub i32 [[TMP0]], [[TMP19]]
; CHECK-NEXT: [[TMP20:%.*]] = shl i32 [[N_VEC]], 2
; CHECK-NEXT: [[TMP21:%.*]] = mul i32 [[N_VEC]], 3
; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK: [[VECTOR_BODY]]:
; CHECK-NEXT: [[INDEX:%.*]] = phi i32 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[VEC_IND:%.*]] = phi <4 x i32> [ <i32 0, i32 3, i32 6, i32 9>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_IND:%.*]] = phi <8 x i32> [ <i32 0, i32 3, i32 6, i32 9, i32 12, i32 15, i32 18, i32 21>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
; CHECK-NEXT: [[OFFSET_IDX:%.*]] = shl i32 [[INDEX]], 2
; CHECK-NEXT: [[TMP22:%.*]] = sext i32 [[OFFSET_IDX]] to i64
; CHECK-NEXT: [[TMP23:%.*]] = getelementptr nusw i32, ptr [[SRC]], i64 [[TMP22]]
-; CHECK-NEXT: [[WIDE_VEC:%.*]] = load <16 x i32>, ptr [[TMP23]], align 4, !alias.scope [[META4:![0-9]+]]
-; CHECK-NEXT: [[STRIDED_VEC:%.*]] = shufflevector <16 x i32> [[WIDE_VEC]], <16 x i32> poison, <4 x i32> <i32 0, i32 4, i32 8, i32 12>
-; CHECK-NEXT: [[STRIDED_VEC9:%.*]] = shufflevector <16 x i32> [[WIDE_VEC]], <16 x i32> poison, <4 x i32> <i32 1, i32 5, i32 9, i32 13>
-; CHECK-NEXT: [[TMP25:%.*]] = sext <4 x i32> [[VEC_IND]] to <4 x i64>
-; CHECK-NEXT: [[TMP27:%.*]] = getelementptr i32, ptr [[DST]], <4 x i64> [[TMP25]]
-; CHECK-NEXT: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[STRIDED_VEC]], <4 x ptr> align 4 [[TMP27]], <4 x i1> splat (i1 true)), !alias.scope [[META7:![0-9]+]], !noalias [[META4]]
-; CHECK-NEXT: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[STRIDED_VEC9]], <4 x ptr> align 4 [[TMP27]], <4 x i1> splat (i1 true)), !alias.scope [[META7]], !noalias [[META4]]
-; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 4
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add <4 x i32> [[VEC_IND]], splat (i32 12)
+; CHECK-NEXT: [[WIDE_VEC:%.*]] = load <32 x i32>, ptr [[TMP23]], align 4, !alias.scope [[META4:![0-9]+]]
+; CHECK-NEXT: [[STRIDED_VEC:%.*]] = shufflevector <32 x i32> [[WIDE_VEC]], <32 x i32> poison, <8 x i32> <i32 0, i32 4, i32 8, i32 12, i32 16, i32 20, i32 24, i32 28>
+; CHECK-NEXT: [[STRIDED_VEC2:%.*]] = shufflevector <32 x i32> [[WIDE_VEC]], <32 x i32> poison, <8 x i32> <i32 1, i32 5, i32 9, i32 13, i32 17, i32 21, i32 25, i32 29>
+; CHECK-NEXT: [[TMP24:%.*]] = sext <8 x i32> [[VEC_IND]] to <8 x i64>
+; CHECK-NEXT: [[WIDE_GEP:%.*]] = getelementptr i32, ptr [[DST]], <8 x i64> [[TMP24]]
+; CHECK-NEXT: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[STRIDED_VEC]], <8 x ptr> align 4 [[WIDE_GEP]], <8 x i1> splat (i1 true)), !alias.scope [[META7:![0-9]+]], !noalias [[META4]]
+; CHECK-NEXT: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[STRIDED_VEC2]], <8 x ptr> align 4 [[WIDE_GEP]], <8 x i1> splat (i1 true)), !alias.scope [[META7]], !noalias [[META4]]
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 8
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <8 x i32> [[VEC_IND]], splat (i32 24)
; CHECK-NEXT: [[TMP26:%.*]] = icmp eq i32 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-NEXT: br i1 [[TMP26]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP9:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
More information about the llvm-commits
mailing list