[llvm] [X86][CostModel] Cost gathers by the instructions CodeGen emits (PR #220565)

Sumukh J Bharadwaj via llvm-commits llvm-commits at lists.llvm.org
Thu Sep 17 04:00:04 PDT 2026


https://github.com/amd-subharad updated https://github.com/llvm/llvm-project/pull/220565

>From 3338e80ae011a806499d035e05ec6bee41a4c17b Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Wed, 2 Sep 2026 17:11:18 +0530
Subject: [PATCH 1/5] [X86][CostModel] Fix lane and index accounting in split
 gather/scatter

When a masked gather or scatter needs more than one register's worth of
pointers, getGSVectorCost recursed on a vector of VF / SplitFactor elements
and multiplied the result. Two things were wrong with that.

The division truncates, so any vector length that is not a multiple of its
split factor lost the remainder: a v9i32 gather was costed as eight lanes,
making its body term as cheap as v8i32's.

The recursive call also recomputed the index width from the shortened vector.
Narrowing a 64-bit index to 32 bits requires a minimum vector length, so an
operation long enough to qualify was still priced with pointer-width indices
once it had been split into parts that were individually too short. Every part
of an operation uses the index width the whole operation qualifies for, so
this overstated the cost of the wide case.

Compute the split cost directly instead: one overhead per part, plus one
scalar memory op for every lane of the original vector, with the index width
chosen once for the operation as a whole.

Both corrections move reported costs, on every subtarget and every cost kind.
On skylake-avx512 a v9i32 gather goes from 12 to 13, and a v24i32 gather
through a dword-index GEP goes from 32 to 28, the latter also dropping from
four parts to two. Lengths that do divide their split factor, and operations
that genuinely need 64-bit indices, are unchanged.
---
 llvm/docs/ReleaseNotes.md                     | 10 ++
 .../lib/Target/X86/X86TargetTransformInfo.cpp | 20 ++--
 .../X86/masked-gather-scatter-split-cost.ll   | 91 +++++++++++++++++++
 .../CostModel/X86/masked-intrinsic-cost.ll    | 20 +++-
 4 files changed, 127 insertions(+), 14 deletions(-)
 create mode 100644 llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll

diff --git a/llvm/docs/ReleaseNotes.md b/llvm/docs/ReleaseNotes.md
index 9f1611e4cab7f..4b863fb5c62bc 100644
--- a/llvm/docs/ReleaseNotes.md
+++ b/llvm/docs/ReleaseNotes.md
@@ -230,6 +230,16 @@ Makes programs 10x faster by doing Special New Thing.
 
 ### Changes to the X86 Backend
 
+* Masked gather/scatter operations that are split across several registers are
+  costed more accurately. Lanes in a vector whose length is not a multiple of
+  its split factor are no longer dropped, so a `v9i32` gather is costed as nine
+  lanes rather than eight. The index width is also chosen for the operation as
+  a whole rather than recomputed for each part, so a wide gather through a
+  32-bit-index GEP is no longer priced as if it used 64-bit indices. Both
+  corrections apply to every X86 subtarget and every cost kind. Lengths that do
+  divide their split factor, and operations that genuinely need 64-bit indices,
+  are unchanged.
+
 ### Changes to the OCaml bindings
 
 ### Changes to the Python bindings
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
index 8e0cf1fc5a153..fc40b17f91e8e 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
@@ -6659,24 +6659,20 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
   std::pair<InstructionCost, MVT> SrcLT = getTypeLegalizationCost(SrcVTy);
   InstructionCost::CostType SplitFactor =
       std::max(IdxsLT.first, SrcLT.first).getValue();
-  if (SplitFactor > 1) {
-    // Handle splitting of vector of pointers
-    auto *SplitSrcTy =
-        FixedVectorType::get(SrcVTy->getScalarType(), VF / SplitFactor);
-    return SplitFactor * getGSVectorCost(Opcode, CostKind, SplitSrcTy, Ptr,
-                                         Alignment, AddressSpace);
-  }
-
-  // If we didn't split, this will be a single gather/scatter instruction.
+  // A vector of pointers that does not fit one register is split into
+  // SplitFactor gather/scatter instructions, each paying the overhead.
   if (CostKind == TTI::TCK_CodeSize)
-    return 1;
+    return SplitFactor;
 
   // The gather / scatter cost is given by Intel architects. It is a rough
   // number since we are looking at one instruction in a time.
   const int GSOverhead = (Opcode == Instruction::Load) ? getGatherOverhead()
                                                        : getScatterOverhead();
-  return GSOverhead + VF * getMemoryOpCost(Opcode, SrcVTy->getScalarType(),
-                                           Alignment, AddressSpace, CostKind);
+  // Charge every lane of the original vector. Descending into VF / SplitFactor
+  // dropped the remainder, so a v9 gather was costed as eight lanes.
+  return SplitFactor * GSOverhead +
+         VF * getMemoryOpCost(Opcode, SrcVTy->getScalarType(), Alignment,
+                              AddressSpace, CostKind);
 }
 
 /// Calculate the cost of Gather / Scatter operation
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
new file mode 100644
index 0000000000000..26954697c2b3f
--- /dev/null
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
@@ -0,0 +1,91 @@
+; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --version 6
+; Costs for masked gather/scatter operations that need more than one register's
+; worth of pointers, across every cost kind.
+;
+; Two properties are pinned here. A vector length that is not a multiple of its
+; split factor must still be charged for all of its lanes. And the index width
+; is chosen once for the whole operation, so a length wide enough to qualify for
+; narrowing keeps the narrow index in each of its parts, rather than reverting
+; to pointer width because an individual part is too short to qualify.
+
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=haswell | FileCheck %s --check-prefix=AVX2
+
+; A length that divides its split factor: unchanged, and the reference point for
+; the two cases below.
+define <8 x i32> @gather_v8i32(<8 x ptr> %ptrs, <8 x i1> %mask) {
+; SKX-LABEL: 'gather_v8i32'
+; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <8 x i32> %v
+;
+; AVX2-LABEL: 'gather_v8i32'
+; AVX2-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <8 x i32> %v
+;
+  %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x i32> poison)
+  ret <8 x i32> %v
+}
+
+; Remainder lanes: v9 splits into two parts but is not a multiple of two, so the
+; ninth lane must not be dropped from the body term.
+define <9 x i32> @gather_v9i32(<9 x ptr> %ptrs, <9 x i1> %mask) {
+; SKX-LABEL: 'gather_v9i32'
+; SKX-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 4 %ptrs, <9 x i1> %mask, <9 x i32> poison)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <9 x i32> %v
+;
+; AVX2-LABEL: 'gather_v9i32'
+; AVX2-NEXT:  Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 4 %ptrs, <9 x i1> %mask, <9 x i32> poison)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <9 x i32> %v
+;
+  %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> %ptrs, i32 4, <9 x i1> %mask, <9 x i32> poison)
+  ret <9 x i32> %v
+}
+
+define void @scatter_v9i32(<9 x i32> %val, <9 x ptr> %ptrs, <9 x i1> %mask) {
+; SKX-LABEL: 'scatter_v9i32'
+; SKX-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> align 4 %ptrs, <9 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+; AVX2-LABEL: 'scatter_v9i32'
+; AVX2-NEXT:  Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> align 4 %ptrs, <9 x i1> %mask)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+  call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> %ptrs, i32 4, <9 x i1> %mask)
+  ret void
+}
+
+; Index width across a split: the GEP indices are 32-bit, and v24 is wide enough
+; to qualify for narrowing, so all three parts are priced with a dword index
+; even though a single part on its own would not qualify.
+define <24 x i32> @gather_v24i32_dword_index(ptr %base, <24 x i32> %idx, <24 x i1> %mask) {
+; SKX-LABEL: 'gather_v24i32_dword_index'
+; SKX-NEXT:  Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
+; SKX-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:2 Lat:100 SizeLat:28 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
+;
+; AVX2-LABEL: 'gather_v24i32_dword_index'
+; AVX2-NEXT:  Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
+; AVX2-NEXT:  Cost Model: Found costs of RThru:82 CodeSize:106 Lat:178 SizeLat:106 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
+;
+  %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
+  %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> %mask, <24 x i32> poison)
+  ret <24 x i32> %v
+}
+
+; Control for the above: genuinely 64-bit indices cannot narrow.
+define <24 x i32> @gather_v24i32_qword_index(ptr %base, <24 x i64> %idx, <24 x i1> %mask) {
+; SKX-LABEL: 'gather_v24i32_qword_index'
+; SKX-NEXT:  Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+; SKX-NEXT:  Cost Model: Found costs of RThru:32 CodeSize:4 Lat:104 SizeLat:32 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
+;
+; AVX2-LABEL: 'gather_v24i32_qword_index'
+; AVX2-NEXT:  Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+; AVX2-NEXT:  Cost Model: Found costs of RThru:82 CodeSize:106 Lat:178 SizeLat:106 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
+;
+  %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+  %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> %mask, <24 x i32> poison)
+  ret <24 x i32> %v
+}
diff --git a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
index ed1b534fac8f8..c02dd6db8b99f 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
@@ -730,7 +730,7 @@ define i32 @masked_store(<1 x i1> %m1, <2 x i1> %m2, <3 x i1> %m3, <4 x i1> %m4,
   ret i32 0
 }
 
-define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8, <16 x i1> %m16, <32 x i1> %m32, <64 x i1> %m64) {
+define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8, <9 x i1> %m9, <16 x i1> %m16, <32 x i1> %m32, <64 x i1> %m64) {
 ; SSE2-LABEL: 'masked_gather'
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:29 CodeSize:37 Lat:61 SizeLat:37 for: %V8F64 = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x double> undef)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:15 CodeSize:19 Lat:31 SizeLat:19 for: %V4F64 = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x double> undef)
@@ -745,6 +745,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:12 Lat:18 SizeLat:12 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:77 CodeSize:93 Lat:141 SizeLat:93 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SSE2-NEXT:  Cost Model: Found costs of RThru:49 CodeSize:58 Lat:85 SizeLat:58 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:39 CodeSize:47 Lat:71 SizeLat:47 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:24 Lat:36 SizeLat:24 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:12 Lat:18 SizeLat:12 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -772,6 +773,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:49 CodeSize:65 Lat:113 SizeLat:65 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SSE42-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:37 Lat:64 SizeLat:37 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:25 CodeSize:33 Lat:57 SizeLat:33 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:17 Lat:29 SizeLat:17 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -799,6 +801,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:55 CodeSize:71 Lat:119 SizeLat:71 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; AVX1-NEXT:  Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:14 CodeSize:18 Lat:30 SizeLat:18 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -826,6 +829,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:55 CodeSize:71 Lat:119 SizeLat:71 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:14 CodeSize:18 Lat:30 SizeLat:18 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -853,6 +857,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:17 CodeSize:4 Lat:44 SizeLat:17 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -880,6 +885,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; KNL-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:17 CodeSize:21 Lat:33 SizeLat:21 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -907,6 +913,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SKX-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -936,6 +943,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
   %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> undef, i32 1, <1 x i1> %m1, <1 x i64> undef)
 
   %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> undef, i32 1, <16 x i1> %m16, <16 x i32> undef)
+  %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> undef, i32 1, <9 x i1> %m9, <9 x i32> undef)
   %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> undef, i32 1, <8 x i1> %m8, <8 x i32> undef)
   %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> undef, i32 1, <4 x i1> %m4, <4 x i32> undef)
   %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> undef, i32 1, <2 x i1> %m2, <2 x i32> undef)
@@ -953,7 +961,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
   ret i32 0
 }
 
-define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8, <16 x i1> %m16, <32 x i1> %m32, <64 x i1> %m64) {
+define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8, <9 x i1> %m9, <16 x i1> %m16, <32 x i1> %m32, <64 x i1> %m64) {
 ; SSE2-LABEL: 'masked_scatter'
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:29 CodeSize:37 Lat:37 SizeLat:37 for: call void @llvm.masked.scatter.v8f64.v8p0(<8 x double> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:15 CodeSize:19 Lat:19 SizeLat:19 for: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
@@ -968,6 +976,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:12 Lat:12 SizeLat:12 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:77 CodeSize:93 Lat:93 SizeLat:93 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SSE2-NEXT:  Cost Model: Found costs of RThru:43 CodeSize:52 Lat:52 SizeLat:52 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:39 CodeSize:47 Lat:47 SizeLat:47 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:24 Lat:24 SizeLat:24 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:12 Lat:12 SizeLat:12 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -995,6 +1004,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:49 CodeSize:65 Lat:65 SizeLat:65 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SSE42-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:37 Lat:37 SizeLat:37 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:25 CodeSize:33 Lat:33 SizeLat:33 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:17 Lat:17 SizeLat:17 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1022,6 +1032,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; AVX1-NEXT:  Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1049,6 +1060,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1076,6 +1088,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SKL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SKL-NEXT:  Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1103,6 +1116,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; KNL-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:17 CodeSize:21 Lat:21 SizeLat:21 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1130,6 +1144,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SKX-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1159,6 +1174,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
   call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> undef, i32 1, <1 x i1> %m1)
 
   call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> undef, i32 1, <16 x i1> %m16)
+  call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> undef, i32 1, <9 x i1> %m9)
   call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> undef, i32 1, <8 x i1> %m8)
   call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> undef, i32 1, <4 x i1> %m4)
   call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> undef, i32 1, <2 x i1> %m2)

>From 16d10bd16b1057d7f0f1b15db60846197b9f61cd Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Thu, 3 Sep 2026 10:39:50 +0530
Subject: [PATCH 2/5] [X86][CostModel] Take the gather/scatter index width from
 the pointer's address space

getGSVectorCost started the index width at DL.getPointerSizeInBits(), which
answers for address space 0 regardless of where the gathered pointers actually
live. On x86, address spaces 270 and 271 hold 32-bit pointers, so a vector of
them is half the width of the same number of ordinary pointers.

Costing them at 64 bits made a <16 x ptr addrspace(270)> gather look as though
its indices needed two ZMMs, so it was charged for a split it does not need:
on skylake-avx512 it cost 20 with two parts rather than 18 with one. A GEP
whose indices were 32-bit already reached the right answer through the index
analysis below, so only pointers used directly, without such a GEP, were
mispriced.

Take the width from the pointer operand's own address space, falling back to
the caller's address space when there is no operand to ask. Address space 0 is
unaffected, as is address space 272, whose pointers really are 64-bit.
---
 llvm/docs/ReleaseNotes.md                     |  6 ++
 .../lib/Target/X86/X86TargetTransformInfo.cpp | 17 +++--
 .../masked-gather-scatter-address-space.ll    | 70 +++++++++++++++++++
 3 files changed, 88 insertions(+), 5 deletions(-)
 create mode 100644 llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll

diff --git a/llvm/docs/ReleaseNotes.md b/llvm/docs/ReleaseNotes.md
index 4b863fb5c62bc..98cfd6fd0e075 100644
--- a/llvm/docs/ReleaseNotes.md
+++ b/llvm/docs/ReleaseNotes.md
@@ -240,6 +240,12 @@ Makes programs 10x faster by doing Special New Thing.
   divide their split factor, and operations that genuinely need 64-bit indices,
   are unchanged.
 
+* Masked gather/scatter costs now take the index width from the address space
+  their pointers live in, instead of assuming the width of address space 0. On
+  x86 this matters for the 32-bit address spaces 270 and 271: a vector of those
+  pointers is half as wide as the same count of ordinary pointers, so it is no
+  longer costed as though it had to be split in two.
+
 ### Changes to the OCaml bindings
 
 ### Changes to the Python bindings
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
index fc40b17f91e8e..db2c0d1849d72 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
@@ -6623,8 +6623,16 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
   // operation will use 16 x 64 indices which do not fit in a zmm and needs
   // to split. Also check that the base pointer is the same for all lanes,
   // and that there's at most one variable index.
-  auto getIndexSizeInBits = [](const Value *Ptr, const DataLayout &DL) {
-    unsigned IndexSize = DL.getPointerSizeInBits();
+  // The index starts out as wide as the pointers being gathered, which is a
+  // property of the address space they live in. Asking the data layout without
+  // one answers for address space 0, which is not necessarily theirs.
+  unsigned PtrSizeInBits = DL.getPointerSizeInBits(
+      Ptr && Ptr->getType()->isPtrOrPtrVectorTy()
+          ? Ptr->getType()->getScalarType()->getPointerAddressSpace()
+          : AddressSpace);
+
+  auto getIndexSizeInBits = [PtrSizeInBits](const Value *Ptr) {
+    unsigned IndexSize = PtrSizeInBits;
     const GetElementPtrInst *GEP = dyn_cast_or_null<GetElementPtrInst>(Ptr);
     if (IndexSize < 64 || !GEP)
       return IndexSize;
@@ -6649,9 +6657,8 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
 
   // Trying to reduce IndexSize to 32 bits for vector 16.
   // By default the IndexSize is equal to pointer size.
-  unsigned IndexSize = (ST->hasAVX512() && VF >= 16)
-                           ? getIndexSizeInBits(Ptr, DL)
-                           : DL.getPointerSizeInBits();
+  unsigned IndexSize =
+      (ST->hasAVX512() && VF >= 16) ? getIndexSizeInBits(Ptr) : PtrSizeInBits;
 
   auto *IndexVTy = FixedVectorType::get(
       IntegerType::get(SrcVTy->getContext(), IndexSize), VF);
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
new file mode 100644
index 0000000000000..0125c7975b7c5
--- /dev/null
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
@@ -0,0 +1,70 @@
+; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --version 6
+; Gathers and scatters over pointers that are not in address space 0. On x86,
+; address spaces 270 and 271 hold 32-bit pointers, so a vector of them is half
+; the width of the same number of address-space-0 pointers, needs no split, and
+; is indexed with a dword rather than a qword index.
+
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX
+
+target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i64:64-i128:128-f80:128-n8:16:32:64-S128"
+
+; A vector of sixteen 32-bit pointers occupies one ZMM, so this is a single
+; dword-indexed gather, not a split qword-indexed pair.
+define <16 x i32> @gather_v16i32_p270(<16 x ptr addrspace(270)> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'gather_v16i32_p270'
+; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+  %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+  ret <16 x i32> %v
+}
+
+; The same shape in address space 0 does need 64-bit indices, and does split.
+define <16 x i32> @gather_v16i32_p0(<16 x ptr> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'gather_v16i32_p0'
+; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+  %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+  ret <16 x i32> %v
+}
+
+; Reaching the same conclusion through a GEP rather than from the pointer type.
+define <16 x i32> @gather_v16i32_p270_gep(ptr addrspace(270) %base, <16 x i32> %idx, <16 x i1> %mask) {
+; SKX-LABEL: 'gather_v16i32_p270_gep'
+; SKX-NEXT:  Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr addrspace(270) %base, <16 x i32> %idx
+; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+  %ptrs = getelementptr inbounds i32, ptr addrspace(270) %base, <16 x i32> %idx
+  %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+  ret <16 x i32> %v
+}
+
+define void @scatter_v16i32_p270(<16 x i32> %val, <16 x ptr addrspace(270)> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'scatter_v16i32_p270'
+; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:1 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p270(<16 x i32> %val, <16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+  call void @llvm.masked.scatter.v16i32.v16p270(<16 x i32> %val, <16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask)
+  ret void
+}
+
+define void @scatter_v16i32_p0(<16 x i32> %val, <16 x ptr> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'scatter_v16i32_p0'
+; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %val, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+  call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %val, <16 x ptr> %ptrs, i32 4, <16 x i1> %mask)
+  ret void
+}
+
+; Address space 272 holds 64-bit pointers, so it must behave like space 0.
+define <16 x i32> @gather_v16i32_p272(<16 x ptr addrspace(272)> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'gather_v16i32_p272'
+; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p272(<16 x ptr addrspace(272)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+  %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p272(<16 x ptr addrspace(272)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+  ret <16 x i32> %v
+}

>From 9e8334ab6d31abd71e7885bd92e466797445de49 Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Tue, 8 Sep 2026 01:59:02 +0530
Subject: [PATCH 3/5] [X86][CostModel] Cost gathers by the instructions CodeGen
 emits

The code-size cost of a hardware masked gather/scatter is its part
count, so it can be read directly against CodeGen. It did not match.
Four things made it differ, each with its own boundary.

The part count came from the legalized register count of the widest
operand. Legalization first widens the vector length to a power of two,
so that count exceeds the instructions emitted whenever that rounding
adds a whole part with no live lane in it: a 24-lane qword-indexed
gather covers eight lanes per instruction but legalizes to a 32-lane,
four-register form whose fourth register holds nothing live, and becomes
three instructions. Count instead the lanes one instruction covers --
the wider of the data and the index bounds it, since a dword gather
indexed by qwords fills a register with indices before it fills one with
data -- and divide the length by that. This affects every subtarget with
a hardware gather or scatter, and it is the term that sets the code-size
cost, so code-size costs move too.

The index width was derived only on AVX-512 at VF >= 16. Outside that
window a gather through a 32-bit-index GEP was priced as though its
indices were pointer-width, tying it with the qword-indexed form that
CodeGen splits into roughly twice as many instructions. Derive the width
whenever a hardware gather/scatter is priced.

Deriving it more often exposed that the derivation ignored the GEP's
stride. A narrow index only survives into the instruction if the
addressing mode can apply that stride as a scale, and the scale field
encodes 1, 2, 4 and 8; any other stride is multiplied into the index
first, and the product is pointer-width. Without that check a gather
over an array of structures -- the form a vectorized `a[idx[i]].f`
takes -- was costed at half the instructions CodeGen emits. Check it, so
that both spellings are priced as CodeGen compiles them.

The mask was not consulted at all. A mask known at compile time says
which parts survive: a part with no live lane is folded away, and an
operation with no live lane at all disappears. A 24-lane qword-indexed
gather with only its first eight lanes live is one instruction, not
three, and only its live lanes are charged for memory accesses. Where
the mask is unknown the two directions differ, so the opcode decides: a
gather drops the tail that widening added, because its result goes
unused, while a scatter keeps it, because it survives as a store under a
zeroed mask that CodeGen still emits. How far that tail reaches follows
the predicate rather than the data: AVX512BW holds the mask in a
k-register as wide as the legalized vector, so the length rounds up to a
power of two, while without it the mask is broken into 16-lane pieces
and the length rounds up only to the next of those. A 48-lane
qword-indexed scatter is eight instructions on skylake-avx512 and six on
knl; the legalized register count says eight for both.

Verified by comparing the code-size cost against the instructions llc
emits, over four data types, both index widths, 21 vector lengths from 2
to 64 and 15 subtargets: 3200 of 3200 hardware cases agree, against 2530
before. A separate sweep over GEP strides from 1 to 32 bytes agrees on
all of them, on an AVX-512 and an AVX2 target. Past 64 lanes, and for a
compile-time mask carrying undef or poison lanes, the count is an upper
bound rather than exact: CodeGen may fold parts there in ways that do
not follow from the length, and an undef lane admits either answer. Both
overcount rather than under, and neither shape is one a vectorizer
produces. A second sweep holding those fixed and varying the GEP the
index comes from, the shape of a compile-time mask, the vector-width
function attributes and the address space agrees on 1962 of 2023,
against 1216 before; the remainder are vp.gather/vp.scatter, which never
reach this code because X86 routes only the masked_* intrinsics to it.
The costs the sweeps cannot read against CodeGen were checked against
the invariants they must obey instead: over four cost kinds and five
subtargets, no cost is invalid or negative, an empty mask is free,
killing lanes never costs more than keeping them, and cost never falls
as the vector grows.

masked-gather-scatter-part-count.ll pins that agreement, pairing each
cost against an llc instruction count in the same module, for both index
widths on an AVX-512, an AVX2 and a no-AVX512BW target, over gathers,
variable-mask scatters, all-lanes-active scatters, and constant masks
that are a prefix, a suffix, non-contiguous and empty. It also covers a
stride the scale field cannot encode against one it can, and gathering
pointers, where the width that sets the lane count comes from the
pointer rather than from an integer type, and which no X86 cost model
test exercised. A throughput run pins the per-lane term that code size
does not expose: the remainder lane of a length that does not fill its
parts, and the lanes a constant mask kills. Undoing any one of the
behaviours above makes the test fail.

masked-gather-scatter-split-cost.ll takes skylake as its AVX2 target
rather than haswell, which has no hardware gather at all, so that the
numbers it pins are part counts that can be read against CodeGen rather
than the cost of scalarizing.

The address-space test gains an AVX2 run, where the narrower pointer
also changes the cost, and the address space 271 cases it described but
did not cover.

One vectorizer decision changes: in cast-costs.ll a v8i32 scatter with
dword indices was costed as two instructions where llc emits one, and
the loop now vectorizes at VF 8 rather than VF 4.
---
 llvm/docs/ReleaseNotes.md                     |  66 ++-
 .../lib/Target/X86/X86TargetTransformInfo.cpp | 142 ++++--
 llvm/lib/Target/X86/X86TargetTransformInfo.h  |   3 +-
 .../masked-gather-scatter-address-space.ll    |  62 ++-
 .../X86/masked-gather-scatter-part-count.ll   | 445 ++++++++++++++++++
 .../X86/masked-gather-scatter-split-cost.ll   |  40 +-
 .../X86/masked-intrinsic-cost-inseltpoison.ll |   6 +-
 .../CostModel/X86/masked-intrinsic-cost.ll    |  74 +--
 .../X86/CostModel/gather-i32-with-i8-index.ll |   6 +-
 .../masked-gather-i32-with-i8-index.ll        |   6 +-
 .../LoopVectorize/X86/cast-costs.ll           |  24 +-
 11 files changed, 757 insertions(+), 117 deletions(-)
 create mode 100644 llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll

diff --git a/llvm/docs/ReleaseNotes.md b/llvm/docs/ReleaseNotes.md
index 98cfd6fd0e075..ffdf8c1fc458a 100644
--- a/llvm/docs/ReleaseNotes.md
+++ b/llvm/docs/ReleaseNotes.md
@@ -230,21 +230,61 @@ Makes programs 10x faster by doing Special New Thing.
 
 ### Changes to the X86 Backend
 
-* Masked gather/scatter operations that are split across several registers are
-  costed more accurately. Lanes in a vector whose length is not a multiple of
-  its split factor are no longer dropped, so a `v9i32` gather is costed as nine
-  lanes rather than eight. The index width is also chosen for the operation as
-  a whole rather than recomputed for each part, so a wide gather through a
-  32-bit-index GEP is no longer priced as if it used 64-bit indices. Both
-  corrections apply to every X86 subtarget and every cost kind. Lengths that do
-  divide their split factor, and operations that genuinely need 64-bit indices,
+* A masked gather/scatter that is split across several instructions is now
+  costed by how many instructions CodeGen emits, rather than by how many
+  legalized registers its operands occupy. The two differ when rounding the
+  length up to a power of two, which is what legalization does first, adds at
+  least one whole part with no live lane in it: a 24-lane qword-indexed gather
+  covers eight lanes per instruction, but legalizes to a 32-lane, four-register
+  form whose fourth register holds nothing live, so it is now costed as the
+  three instructions it becomes. This affects every subtarget that has a
+  hardware gather or scatter, at the lengths where that widening adds a dead
+  part, and it is the term that sets the code-size cost, so code-size costs
+  move as well. Lengths whose widening adds no whole dead part are unchanged.
+
+* All lanes of the original vector are now charged. Dividing the length by the
+  number of parts discarded the remainder, so a `v9i32` gather was costed as
+  eight lanes. This affects any length that is not a multiple of its part
+  count on any subtarget with a hardware gather or scatter. It does not change
+  code-size costs, which count instructions rather than lanes.
+
+* The index width of a masked gather/scatter is now derived whenever one is
+  priced, instead of only on AVX-512 targets at vector lengths of 16 or more.
+  Outside that window a gather through a 32-bit-index GEP was priced as though
+  its indices were pointer-width, tying it with the qword-indexed form that
+  CodeGen splits into roughly twice as many instructions. Non-AVX-512 targets
+  with a hardware gather, such as `skylake`, and AVX-512 targets at shorter
+  lengths, are the ones affected; operations that genuinely need 64-bit indices
   are unchanged.
 
-* Masked gather/scatter costs now take the index width from the address space
-  their pointers live in, instead of assuming the width of address space 0. On
-  x86 this matters for the 32-bit address spaces 270 and 271: a vector of those
-  pointers is half as wide as the same count of ordinary pointers, so it is no
-  longer costed as though it had to be split in two.
+  A narrow index only survives into the instruction if the addressing mode can
+  apply the GEP's stride as a scale, and the scale field encodes 1, 2, 4 and 8.
+  Any other stride is multiplied into the index first, and the product is
+  pointer-width, so such a gather is now costed with 64-bit indices. This
+  corrects a gather over an array of structures, the form a vectorized
+  `a[idx[i]].f` takes, which was previously costed at half the instructions
+  CodeGen emits.
+
+* That index width is also taken from the address space the pointers live in,
+  instead of assuming the width of address space 0. On x86 this matters for the
+  32-bit address spaces 270 and 271: a vector of those pointers is half as wide
+  as the same count of ordinary pointers, so it is no longer costed as though it
+  had to be split in two.
+
+* A mask known at compile time now decides which parts are counted, so a part
+  holding no live lane is not charged and an operation with no live lane at all
+  is free. A 24-lane qword-indexed gather with only its first eight lanes live
+  becomes one instruction rather than three, and only the live lanes are charged
+  for their memory accesses. Where the mask is not known, a gather still drops
+  the tail that legalization's widening added, because its result goes unused,
+  while a scatter counts it, because it survives as a store under a zeroed mask
+  that CodeGen still emits. How far that tail reaches follows the predicate:
+  with AVX512BW the mask fills a k-register as wide as the legalized vector, and
+  without it the mask is broken into 16-lane pieces and the tail stops at the
+  next of those.
+
+  Because these are the costs the loop vectorizer compares, some loops
+  containing a gather or scatter now vectorize at a wider factor than before.
 
 ### Changes to the OCaml bindings
 
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
index db2c0d1849d72..9cebe1c422daf 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
@@ -55,6 +55,7 @@
 #include "llvm/CodeGen/BasicTTIImpl.h"
 #include "llvm/CodeGen/CostTable.h"
 #include "llvm/CodeGen/TargetLowering.h"
+#include "llvm/IR/GetElementPtrTypeIterator.h"
 #include "llvm/IR/InstIterator.h"
 #include "llvm/IR/IntrinsicInst.h"
 #include <optional>
@@ -6608,12 +6609,37 @@ int X86TTIImpl::getScatterOverhead() const {
   return 1024;
 }
 
+// Count the lanes a compile-time constant mask leaves live, and the groups of
+// GroupSize consecutive lanes that hold at least one of them. Returns nullopt
+// when the mask is not known at compile time, which is the common case.
+static std::optional<std::pair<unsigned, unsigned>>
+getActiveGSLanesAndGroups(const Value *Mask, unsigned VF, unsigned GroupSize) {
+  const auto *CMask = dyn_cast_or_null<Constant>(Mask);
+  if (!CMask || !GroupSize)
+    return std::nullopt;
+  unsigned Lanes = 0, Groups = 0;
+  for (unsigned Base = 0; Base < VF; Base += GroupSize) {
+    bool GroupLive = false;
+    for (unsigned I = Base, E = std::min(Base + GroupSize, VF); I != E; ++I) {
+      // An element that cannot be read, or that is not a known-zero bit, has
+      // to be counted as active.
+      const Constant *Elt = CMask->getAggregateElement(I);
+      if (!Elt || !Elt->isNullValue()) {
+        ++Lanes;
+        GroupLive = true;
+      }
+    }
+    if (GroupLive)
+      ++Groups;
+  }
+  return std::make_pair(Lanes, Groups);
+}
+
 // Return an average cost of Gather / Scatter instruction, maybe improved later.
-InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
-                                            TTI::TargetCostKind CostKind,
-                                            Type *SrcVTy, const Value *Ptr,
-                                            Align Alignment,
-                                            unsigned AddressSpace) const {
+InstructionCost
+X86TTIImpl::getGSVectorCost(unsigned Opcode, TTI::TargetCostKind CostKind,
+                            Type *SrcVTy, const Value *Ptr, Align Alignment,
+                            unsigned AddressSpace, const Value *Mask) const {
 
   assert(isa<VectorType>(SrcVTy) && "Unexpected type in getGSVectorCost");
   unsigned VF = cast<FixedVectorType>(SrcVTy)->getNumElements();
@@ -6631,7 +6657,7 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
           ? Ptr->getType()->getScalarType()->getPointerAddressSpace()
           : AddressSpace);
 
-  auto getIndexSizeInBits = [PtrSizeInBits](const Value *Ptr) {
+  auto getIndexSizeInBits = [&](const Value *Ptr) {
     unsigned IndexSize = PtrSizeInBits;
     const GetElementPtrInst *GEP = dyn_cast_or_null<GetElementPtrInst>(Ptr);
     if (IndexSize < 64 || !GEP)
@@ -6641,45 +6667,97 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
     const Value *Ptrs = GEP->getPointerOperand();
     if (Ptrs->getType()->isVectorTy() && !getSplatValue(Ptrs))
       return IndexSize;
-    for (unsigned I = 1, E = GEP->getNumOperands(); I != E; ++I) {
-      if (isa<Constant>(GEP->getOperand(I)))
+    for (gep_type_iterator GTI = gep_type_begin(GEP), GTE = gep_type_end(GEP);
+         GTI != GTE; ++GTI) {
+      const Value *Operand = GTI.getOperand();
+      if (isa<Constant>(Operand))
         continue;
-      Type *IndxTy = GEP->getOperand(I)->getType();
+      Type *IndxTy = Operand->getType();
       if (auto *IndexVTy = dyn_cast<VectorType>(IndxTy))
         IndxTy = IndexVTy->getElementType();
-      if ((IndxTy->getPrimitiveSizeInBits() == 64 &&
-           !isa<SExtInst>(GEP->getOperand(I))) ||
+      if ((IndxTy->getPrimitiveSizeInBits() == 64 && !isa<SExtInst>(Operand)) ||
           ++NumOfVarIndices > 1)
         return IndexSize; // 64
+      // The narrow index only reaches the instruction if the addressing mode
+      // can apply its stride as a scale, and the scale field encodes 1, 2, 4
+      // and 8. Any other stride has to be multiplied into the index first,
+      // and that product is pointer-width, so the operation ends up gathering
+      // with 64-bit indices however narrow the index started out.
+      TypeSize EltSize = DL.getTypeAllocSize(GTI.getIndexedType());
+      if (EltSize.isScalable())
+        return IndexSize; // 64
+      uint64_t Stride = EltSize.getFixedValue();
+      if (!isPowerOf2_64(Stride) || Stride > 8)
+        return IndexSize; // 64
     }
     return (unsigned)32;
   };
 
-  // Trying to reduce IndexSize to 32 bits for vector 16.
-  // By default the IndexSize is equal to pointer size.
-  unsigned IndexSize =
-      (ST->hasAVX512() && VF >= 16) ? getIndexSizeInBits(Ptr) : PtrSizeInBits;
+  // The index width belongs to the operation as a whole, not to any one part,
+  // so derive it whenever a hardware gather/scatter is being priced. Asking
+  // only for AVX-512 at VF >= 16 left a dword-indexed gather costed as though
+  // its indices were pointer-width everywhere else, so it tied with the
+  // qword-indexed form that CodeGen splits into twice as many instructions.
+  unsigned IndexSize = getIndexSizeInBits(Ptr);
 
   auto *IndexVTy = FixedVectorType::get(
       IntegerType::get(SrcVTy->getContext(), IndexSize), VF);
   std::pair<InstructionCost, MVT> IdxsLT = getTypeLegalizationCost(IndexVTy);
   std::pair<InstructionCost, MVT> SrcLT = getTypeLegalizationCost(SrcVTy);
-  InstructionCost::CostType SplitFactor =
-      std::max(IdxsLT.first, SrcLT.first).getValue();
-  // A vector of pointers that does not fit one register is split into
-  // SplitFactor gather/scatter instructions, each paying the overhead.
+
+  // One instruction covers as many lanes as fit in a legal register once both
+  // the data and the indices are held, so the wider of the two bounds the lane
+  // count: a dword gather indexed by qwords fills a register with indices
+  // before it fills one with data. Counting legalized registers instead
+  // overcounts, because legalization first widens to a power of two -- a
+  // 24-lane qword-indexed gather occupies four index registers, but only three
+  // of them hold live lanes and CodeGen emits three instructions.
+  auto LanesOf = [](MVT VT) {
+    return VT.isVector() ? VT.getVectorNumElements() : 1U;
+  };
+  unsigned LanesPerPart =
+      std::max(std::min(LanesOf(IdxsLT.second), LanesOf(SrcLT.second)), 1U);
+
+  // Lanes the operation is declared over, and the parts they occupy.
+  InstructionCost::CostType ChargedLanes = VF;
+  InstructionCost::CostType EmittedParts = divideCeil(VF, LanesPerPart);
+
+  if (auto Active = getActiveGSLanesAndGroups(Mask, VF, LanesPerPart)) {
+    // A compile-time mask says which parts survive: one with no live lane
+    // gathers nothing, and CodeGen folds it away rather than emitting it.
+    ChargedLanes = Active->first;
+    EmittedParts = Active->second;
+  } else if (Opcode != Instruction::Load) {
+    // Without a compile-time mask the tail that legalization's widening added
+    // cannot be proved dead. A gather's tail produces an unused result and is
+    // dropped, but a scatter's survives as a store under a zeroed mask, which
+    // is still an instruction. How far the widening reaches is set by the
+    // predicate: AVX512BW holds it in a k-register as wide as the legalized
+    // vector, so the length rounds up to a power of two, while without it the
+    // mask is broken into 16-lane pieces and only rounds up to one of those.
+    unsigned WidenedLanes = PowerOf2Ceil(VF);
+    if (!ST->hasBWI())
+      WidenedLanes = std::min(WidenedLanes, alignTo(VF, 16u));
+    EmittedParts = divideCeil(WidenedLanes, LanesPerPart);
+  }
+
+  // Nothing survives, so no instruction is emitted.
+  if (EmittedParts == 0)
+    return 0;
+
+  // Each emitted instruction pays the overhead once.
   if (CostKind == TTI::TCK_CodeSize)
-    return SplitFactor;
+    return EmittedParts;
 
   // The gather / scatter cost is given by Intel architects. It is a rough
   // number since we are looking at one instruction in a time.
   const int GSOverhead = (Opcode == Instruction::Load) ? getGatherOverhead()
                                                        : getScatterOverhead();
-  // Charge every lane of the original vector. Descending into VF / SplitFactor
-  // dropped the remainder, so a v9 gather was costed as eight lanes.
-  return SplitFactor * GSOverhead +
-         VF * getMemoryOpCost(Opcode, SrcVTy->getScalarType(), Alignment,
-                              AddressSpace, CostKind);
+  // Charge every lane that is actually accessed. Dividing the length by the
+  // part count dropped the remainder, so a v9 gather was costed as eight lanes.
+  return EmittedParts * GSOverhead +
+         ChargedLanes * getMemoryOpCost(Opcode, SrcVTy->getScalarType(),
+                                        Alignment, AddressSpace, CostKind);
 }
 
 /// Calculate the cost of Gather / Scatter operation
@@ -6704,8 +6782,18 @@ X86TTIImpl::getGatherScatterOpCost(const MemIntrinsicCostAttributes &MICA,
 
   assert(SrcVTy->isVectorTy() && "Unexpected data type for Gather/Scatter");
   unsigned AddressSpace = MICA.getAddressSpace();
-  return getGSVectorCost(Opcode, CostKind, SrcVTy, Ptr, Alignment,
-                         AddressSpace);
+
+  // A mask known at compile time tells us which parts CodeGen keeps. Read it
+  // only from a real gather/scatter call: the context instruction can also be a
+  // plain load or store being considered for the transform, whose operands mean
+  // something else entirely.
+  const Value *Mask = nullptr;
+  if (const auto *II = dyn_cast_or_null<IntrinsicInst>(MICA.getInst()))
+    if (II->getIntrinsicID() == MICA.getID())
+      Mask = II->getArgOperand(IsLoad ? 1 : 2);
+
+  return getGSVectorCost(Opcode, CostKind, SrcVTy, Ptr, Alignment, AddressSpace,
+                         Mask);
 }
 
 bool X86TTIImpl::isLSRCostLess(const TargetTransformInfo::LSRCost &C1,
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.h b/llvm/lib/Target/X86/X86TargetTransformInfo.h
index 422e4aad2316c..a130835173b5d 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.h
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.h
@@ -271,7 +271,8 @@ class X86TTIImpl final : public BasicTTIImplBase<X86TTIImpl> {
   bool supportsGather() const;
   InstructionCost getGSVectorCost(unsigned Opcode, TTI::TargetCostKind CostKind,
                                   Type *DataTy, const Value *Ptr,
-                                  Align Alignment, unsigned AddressSpace) const;
+                                  Align Alignment, unsigned AddressSpace,
+                                  const Value *Mask = nullptr) const;
 
   int getGatherOverhead() const;
   int getScatterOverhead() const;
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
index 0125c7975b7c5..2d904412f5d52 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-address-space.ll
@@ -1,10 +1,16 @@
 ; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --version 6
 ; Gathers and scatters over pointers that are not in address space 0. On x86,
 ; address spaces 270 and 271 hold 32-bit pointers, so a vector of them is half
-; the width of the same number of address-space-0 pointers, needs no split, and
-; is indexed with a dword rather than a qword index.
+; the width of the same number of address-space-0 pointers, needs fewer parts,
+; and is indexed with a dword rather than a qword index.
+;
+; Both targets are covered because the two differ in what they expose. On the
+; AVX-512 target the narrower pointer removes a split that address space 0 still
+; needs, while on the AVX2 target it also changes the cost reached through a GEP,
+; whose index width is derived from the pointer's own address space.
 
 ; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake | FileCheck %s --check-prefix=SKL
 
 target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i64:64-i128:128-f80:128-n8:16:32:64-S128"
 
@@ -14,6 +20,10 @@ define <16 x i32> @gather_v16i32_p270(<16 x ptr addrspace(270)> %ptrs, <16 x i1>
 ; SKX-LABEL: 'gather_v16i32_p270'
 ; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+; SKL-LABEL: 'gather_v16i32_p270'
+; SKL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
 ;
   %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
   ret <16 x i32> %v
@@ -24,6 +34,10 @@ define <16 x i32> @gather_v16i32_p0(<16 x ptr> %ptrs, <16 x i1> %mask) {
 ; SKX-LABEL: 'gather_v16i32_p0'
 ; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+; SKL-LABEL: 'gather_v16i32_p0'
+; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
 ;
   %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
   ret <16 x i32> %v
@@ -35,6 +49,11 @@ define <16 x i32> @gather_v16i32_p270_gep(ptr addrspace(270) %base, <16 x i32> %
 ; SKX-NEXT:  Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr addrspace(270) %base, <16 x i32> %idx
 ; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+; SKL-LABEL: 'gather_v16i32_p270_gep'
+; SKL-NEXT:  Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr addrspace(270) %base, <16 x i32> %idx
+; SKL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
 ;
   %ptrs = getelementptr inbounds i32, ptr addrspace(270) %base, <16 x i32> %idx
   %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p270(<16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
@@ -45,6 +64,10 @@ define void @scatter_v16i32_p270(<16 x i32> %val, <16 x ptr addrspace(270)> %ptr
 ; SKX-LABEL: 'scatter_v16i32_p270'
 ; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:1 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p270(<16 x i32> %val, <16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+; SKL-LABEL: 'scatter_v16i32_p270'
+; SKL-NEXT:  Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p270(<16 x i32> %val, <16 x ptr addrspace(270)> align 4 %ptrs, <16 x i1> %mask)
+; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
 ;
   call void @llvm.masked.scatter.v16i32.v16p270(<16 x i32> %val, <16 x ptr addrspace(270)> %ptrs, i32 4, <16 x i1> %mask)
   ret void
@@ -54,16 +77,51 @@ define void @scatter_v16i32_p0(<16 x i32> %val, <16 x ptr> %ptrs, <16 x i1> %mas
 ; SKX-LABEL: 'scatter_v16i32_p0'
 ; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %val, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+; SKL-LABEL: 'scatter_v16i32_p0'
+; SKL-NEXT:  Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %val, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
 ;
   call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %val, <16 x ptr> %ptrs, i32 4, <16 x i1> %mask)
   ret void
 }
 
+; Address space 271 holds 32-bit pointers too, and must be treated like 270.
+define <16 x i32> @gather_v16i32_p271(<16 x ptr addrspace(271)> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'gather_v16i32_p271'
+; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p271(<16 x ptr addrspace(271)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+; SKL-LABEL: 'gather_v16i32_p271'
+; SKL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p271(<16 x ptr addrspace(271)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+  %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p271(<16 x ptr addrspace(271)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+  ret <16 x i32> %v
+}
+
+define void @scatter_v16i32_p271(<16 x i32> %val, <16 x ptr addrspace(271)> %ptrs, <16 x i1> %mask) {
+; SKX-LABEL: 'scatter_v16i32_p271'
+; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:1 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p271(<16 x i32> %val, <16 x ptr addrspace(271)> align 4 %ptrs, <16 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+; SKL-LABEL: 'scatter_v16i32_p271'
+; SKL-NEXT:  Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p271(<16 x i32> %val, <16 x ptr addrspace(271)> align 4 %ptrs, <16 x i1> %mask)
+; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+  call void @llvm.masked.scatter.v16i32.v16p271(<16 x i32> %val, <16 x ptr addrspace(271)> %ptrs, i32 4, <16 x i1> %mask)
+  ret void
+}
+
 ; Address space 272 holds 64-bit pointers, so it must behave like space 0.
 define <16 x i32> @gather_v16i32_p272(<16 x ptr addrspace(272)> %ptrs, <16 x i1> %mask) {
 ; SKX-LABEL: 'gather_v16i32_p272'
 ; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p272(<16 x ptr addrspace(272)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
+;
+; SKL-LABEL: 'gather_v16i32_p272'
+; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p272(<16 x ptr addrspace(272)> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x i32> %v
 ;
   %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p272(<16 x ptr addrspace(272)> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
   ret <16 x i32> %v
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll
new file mode 100644
index 0000000000000..fd6374b595275
--- /dev/null
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll
@@ -0,0 +1,445 @@
+; Pins the cost model's part count against the number of gather/scatter
+; instructions CodeGen actually emits, for both index widths on an AVX-512 and
+; an AVX2 target.
+;
+; The pairing is the point of the test: the code-size cost of a hardware
+; gather/scatter is its part count, so every cost check below has a matching
+; instruction count taken from the same module. Three properties would
+; otherwise drift unnoticed. A part count taken from the legalized register
+; count overstates the work when legalization widens to a power of two; an
+; index width derived only for wide AVX-512 vectors leaves a dword-indexed
+; operation tied with the qword-indexed form that CodeGen splits into more
+; instructions; and the tail a scatter still emits under a zeroed mask reaches
+; only as far as the predicate widens, which depends on AVX512BW.
+;
+; The throughput run pins the other half of the cost, the per-lane term, which
+; code size does not expose: it is charged for the lanes that are really
+; accessed, so it must follow the remainder lane of a length that does not fill
+; its parts, and must drop the lanes a compile-time mask kills.
+;
+; These checks are maintained by hand rather than by
+; update_analyze_test_checks.py, which does not know about the paired llc runs.
+
+; RUN: opt < %s -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" -cost-kind=code-size -disable-output -mcpu=skylake-avx512 2>&1 | FileCheck %s --check-prefix=SKX-COST
+; RUN: opt < %s -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" -cost-kind=throughput -disable-output -mcpu=skylake-avx512 2>&1 | FileCheck %s --check-prefix=SKX-TPUT
+; RUN: llc < %s -mtriple=x86_64-unknown-linux-gnu -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX-ASM
+; RUN: opt < %s -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" -cost-kind=code-size -disable-output -mcpu=skylake 2>&1 | FileCheck %s --check-prefix=AVX2-COST
+; RUN: llc < %s -mtriple=x86_64-unknown-linux-gnu -mcpu=skylake | FileCheck %s --check-prefix=AVX2-ASM
+; RUN: opt < %s -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" -cost-kind=code-size -disable-output -mcpu=knl 2>&1 | FileCheck %s --check-prefix=KNL-COST
+; RUN: llc < %s -mtriple=x86_64-unknown-linux-gnu -mcpu=knl | FileCheck %s --check-prefix=KNL-ASM
+
+; A length that fills its single part exactly, to sit against the nine-lane form
+; below: both are one instruction, but the ninth lane is still accessed and has
+; to be charged, so the throughput costs must differ by exactly one lane.
+; SKX-COST-LABEL: 'gather_v8i32_dword_index'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v8i32_dword_index'
+; SKX-TPUT: cost of 10 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v8i32_dword_index:
+; SKX-ASM-COUNT-1: vpgatherdd
+; SKX-ASM-NOT: vpgather
+define <8 x i32> @gather_v8i32_dword_index(ptr %base, <8 x i32> %idx, <8 x i1> %mask) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <8 x i32> %idx
+  %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x i32> poison)
+  ret <8 x i32> %v
+}
+
+; A vector length that is not a multiple of its part count: the remainder lane
+; still needs a part of its own once the index no longer fits alongside the rest.
+; SKX-COST-LABEL: 'gather_v9i32_dword_index'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v9i32_dword_index'
+; SKX-TPUT: cost of 11 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v9i32_dword_index:
+; SKX-ASM-COUNT-1: vpgatherdd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v9i32_dword_index'
+; AVX2-COST: cost of 2 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v9i32_dword_index:
+; AVX2-ASM-COUNT-2: vpgatherdd
+; AVX2-ASM-NOT: vpgather
+define <9 x i32> @gather_v9i32_dword_index(ptr %base, <9 x i32> %idx, <9 x i1> %mask) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <9 x i32> %idx
+  %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> %ptrs, i32 4, <9 x i1> %mask, <9 x i32> poison)
+  ret <9 x i32> %v
+}
+
+; SKX-COST-LABEL: 'gather_v9i32_qword_index'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v9i32_qword_index:
+; SKX-ASM-COUNT-2: vpgatherqd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v9i32_qword_index'
+; AVX2-COST: cost of 3 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v9i32_qword_index:
+; AVX2-ASM-COUNT-3: vpgatherqd
+; AVX2-ASM-NOT: vpgather
+define <9 x i32> @gather_v9i32_qword_index(ptr %base, <9 x i64> %idx, <9 x i1> %mask) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <9 x i64> %idx
+  %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> %ptrs, i32 4, <9 x i1> %mask, <9 x i32> poison)
+  ret <9 x i32> %v
+}
+
+; A length below the width at which the index used to be examined, so both index
+; widths would otherwise be priced the same on either target.
+; SKX-COST-LABEL: 'gather_v17i32_dword_index'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v17i32_dword_index:
+; SKX-ASM-COUNT-2: vpgatherdd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v17i32_dword_index'
+; AVX2-COST: cost of 3 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v17i32_dword_index:
+; AVX2-ASM-COUNT-3: vpgatherdd
+; AVX2-ASM-NOT: vpgather
+define <17 x i32> @gather_v17i32_dword_index(ptr %base, <17 x i32> %idx, <17 x i1> %mask) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <17 x i32> %idx
+  %v = call <17 x i32> @llvm.masked.gather.v17i32.v17p0(<17 x ptr> %ptrs, i32 4, <17 x i1> %mask, <17 x i32> poison)
+  ret <17 x i32> %v
+}
+
+; SKX-COST-LABEL: 'gather_v17i32_qword_index'
+; SKX-COST: cost of 3 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v17i32_qword_index:
+; SKX-ASM-COUNT-3: vpgatherqd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v17i32_qword_index'
+; AVX2-COST: cost of 5 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v17i32_qword_index:
+; AVX2-ASM-COUNT-5: vpgatherqd
+; AVX2-ASM-NOT: vpgather
+define <17 x i32> @gather_v17i32_qword_index(ptr %base, <17 x i64> %idx, <17 x i1> %mask) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <17 x i64> %idx
+  %v = call <17 x i32> @llvm.masked.gather.v17i32.v17p0(<17 x ptr> %ptrs, i32 4, <17 x i1> %mask, <17 x i32> poison)
+  ret <17 x i32> %v
+}
+
+; A length whose qword indices occupy four legal registers but fill only three
+; of them with live lanes.
+; SKX-COST-LABEL: 'gather_v24i32_dword_index'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_dword_index:
+; SKX-ASM-COUNT-2: vpgatherdd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v24i32_dword_index'
+; AVX2-COST: cost of 3 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v24i32_dword_index:
+; AVX2-ASM-COUNT-3: vpgatherdd
+; AVX2-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_dword_index(ptr %base, <24 x i32> %idx, <24 x i1> %mask) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
+  %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> %mask, <24 x i32> poison)
+  ret <24 x i32> %v
+}
+
+; SKX-COST-LABEL: 'gather_v24i32_qword_index'
+; SKX-COST: cost of 3 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v24i32_qword_index'
+; SKX-TPUT: cost of 30 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_qword_index:
+; SKX-ASM-COUNT-3: vpgatherqd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v24i32_qword_index'
+; AVX2-COST: cost of 6 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v24i32_qword_index:
+; AVX2-ASM-COUNT-6: vpgatherqd
+; AVX2-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_qword_index(ptr %base, <24 x i64> %idx, <24 x i1> %mask) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+  %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> %mask, <24 x i32> poison)
+  ret <24 x i32> %v
+}
+
+; Scatters with a variable mask. AVX2 has no hardware scatter, so its cost comes
+; from scalarizing instead of from a part count, and it emits no scatter at all.
+; SKX-COST-LABEL: 'scatter_v9i32_dword_index_var_mask'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v9i32_dword_index_var_mask:
+; SKX-ASM-COUNT-1: vpscatterdd
+; SKX-ASM-NOT: vpscatter
+; AVX2-ASM-LABEL: scatter_v9i32_dword_index_var_mask:
+; AVX2-ASM-NOT: vpscatter
+define void @scatter_v9i32_dword_index_var_mask(ptr %base, <9 x i32> %idx, <9 x i32> %val, <9 x i1> %mask) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <9 x i32> %idx
+  call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> %ptrs, i32 4, <9 x i1> %mask)
+  ret void
+}
+
+; SKX-COST-LABEL: 'scatter_v9i32_qword_index_var_mask'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v9i32_qword_index_var_mask:
+; SKX-ASM-COUNT-2: vpscatterqd
+; SKX-ASM-NOT: vpscatter
+; AVX2-ASM-LABEL: scatter_v9i32_qword_index_var_mask:
+; AVX2-ASM-NOT: vpscatter
+define void @scatter_v9i32_qword_index_var_mask(ptr %base, <9 x i64> %idx, <9 x i32> %val, <9 x i1> %mask) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <9 x i64> %idx
+  call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> %ptrs, i32 4, <9 x i1> %mask)
+  ret void
+}
+
+; Scatters with every lane active: an all-ones mask leaves the part count alone.
+; SKX-COST-LABEL: 'scatter_v24i32_dword_index_all_active'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v24i32_dword_index_all_active:
+; SKX-ASM-COUNT-2: vpscatterdd
+; SKX-ASM-NOT: vpscatter
+; AVX2-ASM-LABEL: scatter_v24i32_dword_index_all_active:
+; AVX2-ASM-NOT: vpscatter
+define void @scatter_v24i32_dword_index_all_active(ptr %base, <24 x i32> %idx, <24 x i32> %val) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
+  call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %val, <24 x ptr> %ptrs, i32 4, <24 x i1> splat (i1 true))
+  ret void
+}
+
+; SKX-COST-LABEL: 'scatter_v24i32_qword_index_all_active'
+; SKX-COST: cost of 3 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v24i32_qword_index_all_active:
+; SKX-ASM-COUNT-3: vpscatterqd
+; SKX-ASM-NOT: vpscatter
+; AVX2-ASM-LABEL: scatter_v24i32_qword_index_all_active:
+; AVX2-ASM-NOT: vpscatter
+define void @scatter_v24i32_qword_index_all_active(ptr %base, <24 x i64> %idx, <24 x i32> %val) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+  call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %val, <24 x ptr> %ptrs, i32 4, <24 x i1> splat (i1 true))
+  ret void
+}
+
+; A variable mask on a length that legalization widens past a part boundary. The
+; widened tail cannot be proved dead, and it survives as a store under a zeroed
+; mask, so CodeGen emits a fourth instruction where the all-ones form above
+; emits three. That store is counted, since it is still an instruction.
+; SKX-COST-LABEL: 'scatter_v24i32_qword_index_var_mask'
+; SKX-COST: cost of 4 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v24i32_qword_index_var_mask:
+; SKX-ASM-COUNT-3: vpscatterqd
+; SKX-ASM: kxor
+; SKX-ASM-COUNT-1: vpscatterqd
+; SKX-ASM-NOT: vpscatter
+; AVX2-ASM-LABEL: scatter_v24i32_qword_index_var_mask:
+; AVX2-ASM-NOT: vpscatter
+define void @scatter_v24i32_qword_index_var_mask(ptr %base, <24 x i64> %idx, <24 x i32> %val, <24 x i1> %mask) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+  call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %val, <24 x ptr> %ptrs, i32 4, <24 x i1> %mask)
+  ret void
+}
+
+; A compile-time mask says which parts survive. With only the first eight lanes
+; live, two of these three parts hold nothing and are folded away, leaving one
+; instruction rather than the three the declared length would suggest.
+; The dead lanes are not charged either: against the 30 the same shape costs
+; with an unknown mask above, this is the cost of the eight lanes it reaches.
+; SKX-COST-LABEL: 'gather_v24i32_qword_index_first8_mask'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v24i32_qword_index_first8_mask'
+; SKX-TPUT: cost of 10 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_qword_index_first8_mask:
+; SKX-ASM-COUNT-1: vpgatherqd
+; SKX-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_qword_index_first8_mask(ptr %base, <24 x i64> %idx) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+  %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> <i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false>, <24 x i32> poison)
+  ret <24 x i32> %v
+}
+
+; The same for a scatter, where the dead parts would otherwise be the zeroed
+; stores counted above.
+; SKX-COST-LABEL: 'scatter_v24i32_qword_index_first8_mask'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v24i32_qword_index_first8_mask:
+; SKX-ASM-COUNT-1: vpscatterqd
+; SKX-ASM-NOT: vpscatter
+define void @scatter_v24i32_qword_index_first8_mask(ptr %base, <24 x i64> %idx, <24 x i32> %val) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+  call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %val, <24 x ptr> %ptrs, i32 4, <24 x i1> <i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false>)
+  ret void
+}
+
+; A mask with no live lane leaves nothing to do, and CodeGen emits no gather at
+; all, so the operation is free rather than costed as three parts.
+; SKX-COST-LABEL: 'gather_v24i32_qword_index_zero_mask'
+; SKX-COST: cost of 0 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v24i32_qword_index_zero_mask'
+; SKX-TPUT: cost of 0 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_qword_index_zero_mask:
+; SKX-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_qword_index_zero_mask(ptr %base, <24 x i64> %idx) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+  %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> zeroinitializer, <24 x i32> poison)
+  ret <24 x i32> %v
+}
+
+; The same for a scatter, which unlike a gather would otherwise still be charged
+; for the tail it emits under a zeroed mask.
+; SKX-COST-LABEL: 'scatter_v24i32_qword_index_zero_mask'
+; SKX-COST: cost of 0 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v24i32_qword_index_zero_mask:
+; SKX-ASM-NOT: vpscatter
+define void @scatter_v24i32_qword_index_zero_mask(ptr %base, <24 x i64> %idx, <24 x i32> %val) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+  call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %val, <24 x ptr> %ptrs, i32 4, <24 x i1> zeroinitializer)
+  ret void
+}
+
+; The live lanes need not be a prefix. Here they sit in the first and last part
+; with a dead one between, so it is the parts holding a live lane that are
+; counted, not the number of live lanes rounded up to a part.
+; SKX-COST-LABEL: 'gather_v24i32_qword_index_gapped_mask'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v24i32_qword_index_gapped_mask'
+; SKX-TPUT: cost of 6 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_qword_index_gapped_mask:
+; SKX-ASM-COUNT-2: vpgatherqd
+; SKX-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_qword_index_gapped_mask(ptr %base, <24 x i64> %idx) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+  %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> <i1 true, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 true, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false>, <24 x i32> poison)
+  ret <24 x i32> %v
+}
+
+; Live lanes confined to the last part, the mirror of the first-eight case, so
+; that a part is not counted merely because earlier parts were.
+; SKX-COST-LABEL: 'gather_v24i32_qword_index_last8_mask'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v24i32_qword_index_last8_mask:
+; SKX-ASM-COUNT-1: vpgatherqd
+; SKX-ASM-NOT: vpgather
+define <24 x i32> @gather_v24i32_qword_index_last8_mask(ptr %base, <24 x i64> %idx) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
+  %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> <i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 false, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true>, <24 x i32> poison)
+  ret <24 x i32> %v
+}
+
+; How far that widened tail reaches is set by the predicate rather than by the
+; data. AVX512BW holds the mask in a k-register as wide as the legalized
+; vector, so a 48-lane length rounds up to 64 and eight parts are emitted;
+; without it the mask is broken into 16-lane pieces, the length rounds up only
+; to 48, and six are. Counting legal registers would give eight on both.
+; SKX-COST-LABEL: 'scatter_v48i32_qword_index_var_mask'
+; SKX-COST: cost of 8 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v48i32_qword_index_var_mask:
+; SKX-ASM-COUNT-8: vpscatterqd
+; SKX-ASM-NOT: vpscatter
+; KNL-COST-LABEL: 'scatter_v48i32_qword_index_var_mask'
+; KNL-COST: cost of 6 for instruction: {{.*}}masked.scatter
+; KNL-ASM-LABEL: scatter_v48i32_qword_index_var_mask:
+; KNL-ASM-COUNT-6: vpscatterqd
+; KNL-ASM-NOT: vpscatter
+define void @scatter_v48i32_qword_index_var_mask(ptr %base, <48 x i64> %idx, <48 x i32> %val, <48 x i1> %mask) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <48 x i64> %idx
+  call void @llvm.masked.scatter.v48i32.v48p0(<48 x i32> %val, <48 x ptr> %ptrs, i32 4, <48 x i1> %mask)
+  ret void
+}
+
+; The same split with dword indices, where a part covers sixteen lanes instead
+; of eight: four parts with the wide predicate, three without it.
+; SKX-COST-LABEL: 'scatter_v48i32_dword_index_var_mask'
+; SKX-COST: cost of 4 for instruction: {{.*}}masked.scatter
+; SKX-ASM-LABEL: scatter_v48i32_dword_index_var_mask:
+; SKX-ASM-COUNT-4: vpscatterdd
+; SKX-ASM-NOT: vpscatter
+; KNL-COST-LABEL: 'scatter_v48i32_dword_index_var_mask'
+; KNL-COST: cost of 3 for instruction: {{.*}}masked.scatter
+; KNL-ASM-LABEL: scatter_v48i32_dword_index_var_mask:
+; KNL-ASM-COUNT-3: vpscatterdd
+; KNL-ASM-NOT: vpscatter
+define void @scatter_v48i32_dword_index_var_mask(ptr %base, <48 x i32> %idx, <48 x i32> %val, <48 x i1> %mask) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <48 x i32> %idx
+  call void @llvm.masked.scatter.v48i32.v48p0(<48 x i32> %val, <48 x ptr> %ptrs, i32 4, <48 x i1> %mask)
+  ret void
+}
+
+; Pointers are gathered as often as integers are, and the element size that
+; decides how many lanes a part holds has to come from the pointer's own width
+; rather than from an integer type. Eight pointer-sized lanes fill one AVX-512
+; part and two AVX2 parts.
+; SKX-COST-LABEL: 'gather_v8ptr_qword_index'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v8ptr_qword_index'
+; SKX-TPUT: cost of 10 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v8ptr_qword_index:
+; SKX-ASM-COUNT-1: vpgatherqq
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v8ptr_qword_index'
+; AVX2-COST: cost of 2 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v8ptr_qword_index:
+; AVX2-ASM-COUNT-2: vpgatherqq
+; AVX2-ASM-NOT: vpgatherqq
+define <8 x ptr> @gather_v8ptr_qword_index(ptr %base, <8 x i64> %idx, <8 x i1> %mask) {
+  %ptrs = getelementptr inbounds ptr, ptr %base, <8 x i64> %idx
+  %r = call <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr> %ptrs, i32 8, <8 x i1> %mask, <8 x ptr> poison)
+  ret <8 x ptr> %r
+}
+
+; The same gather with only its first two lanes left live by a compile-time
+; mask. The second AVX2 part holds no live lane and is not emitted, and both
+; targets charge two lanes rather than eight.
+; SKX-COST-LABEL: 'gather_v8ptr_qword_index_first2_mask'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-TPUT-LABEL: 'gather_v8ptr_qword_index_first2_mask'
+; SKX-TPUT: cost of 4 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v8ptr_qword_index_first2_mask:
+; SKX-ASM-COUNT-1: vpgatherqq
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v8ptr_qword_index_first2_mask'
+; AVX2-COST: cost of 1 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v8ptr_qword_index_first2_mask:
+; AVX2-ASM-COUNT-1: vpgatherqq
+; AVX2-ASM-NOT: vpgatherqq
+define <8 x ptr> @gather_v8ptr_qword_index_first2_mask(ptr %base, <8 x i64> %idx) {
+  %ptrs = getelementptr inbounds ptr, ptr %base, <8 x i64> %idx
+  %r = call <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr> %ptrs, i32 8, <8 x i1> <i1 1, i1 1, i1 0, i1 0, i1 0, i1 0, i1 0, i1 0>, <8 x ptr> poison)
+  ret <8 x ptr> %r
+}
+
+; A narrow index only stays narrow if the addressing mode can apply its stride
+; as a scale, and the scale field encodes 1, 2, 4 and 8. This GEP walks an
+; array of three-word structures, the form a loop vectorizer emits for
+; a[idx[i]].f, so its stride is twelve: the multiply is folded into the index
+; and the gather ends up indexed by qwords, taking twice the instructions of
+; the four-byte-stride control below.
+; SKX-COST-LABEL: 'gather_v16i32_struct_stride'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v16i32_struct_stride:
+; SKX-ASM-COUNT-2: vpgatherqd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v16i32_struct_stride'
+; AVX2-COST: cost of 4 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v16i32_struct_stride:
+; AVX2-ASM-COUNT-4: vpgatherqd
+; AVX2-ASM-NOT: vpgatherqd
+define <16 x i32> @gather_v16i32_struct_stride(ptr %base, <16 x i32> %idx, <16 x i1> %mask) {
+  %sext = sext <16 x i32> %idx to <16 x i64>
+  %ptrs = getelementptr inbounds {i32, i32, i32}, ptr %base, <16 x i64> %sext, i32 0
+  %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+  ret <16 x i32> %v
+}
+
+; The same gather over a four-byte stride, which the scale field does encode,
+; so the index stays a dword and one instruction covers twice the lanes.
+; SKX-COST-LABEL: 'gather_v16i32_scaled_stride'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v16i32_scaled_stride:
+; SKX-ASM-COUNT-1: vpgatherdd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v16i32_scaled_stride'
+; AVX2-COST: cost of 2 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v16i32_scaled_stride:
+; AVX2-ASM-COUNT-2: vpgatherdd
+; AVX2-ASM-NOT: vpgatherdd
+define <16 x i32> @gather_v16i32_scaled_stride(ptr %base, <16 x i32> %idx, <16 x i1> %mask) {
+  %sext = sext <16 x i32> %idx to <16 x i64>
+  %ptrs = getelementptr inbounds i32, ptr %base, <16 x i64> %sext
+  %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+  ret <16 x i32> %v
+}
+
+declare <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr>, i32, <16 x i1>, <16 x i32>)
+declare <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr>, i32, <8 x i1>, <8 x ptr>)
+declare <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr>, i32, <8 x i1>, <8 x i32>)
+declare <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr>, i32, <9 x i1>, <9 x i32>)
+declare <17 x i32> @llvm.masked.gather.v17i32.v17p0(<17 x ptr>, i32, <17 x i1>, <17 x i32>)
+declare <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr>, i32, <24 x i1>, <24 x i32>)
+declare void @llvm.masked.scatter.v9i32.v9p0(<9 x i32>, <9 x ptr>, i32, <9 x i1>)
+declare void @llvm.masked.scatter.v24i32.v24p0(<24 x i32>, <24 x ptr>, i32, <24 x i1>)
+declare void @llvm.masked.scatter.v48i32.v48p0(<48 x i32>, <48 x ptr>, i32, <48 x i1>)
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
index 26954697c2b3f..28a80f6d8f84a 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-split-cost.ll
@@ -1,15 +1,20 @@
 ; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --version 6
-; Costs for masked gather/scatter operations that need more than one register's
-; worth of pointers, across every cost kind.
+; Costs for masked gather/scatter operations that need one register's worth of
+; pointers or more, across every cost kind.
 ;
 ; Two properties are pinned here. A vector length that is not a multiple of its
-; split factor must still be charged for all of its lanes. And the index width
-; is chosen once for the whole operation, so a length wide enough to qualify for
-; narrowing keeps the narrow index in each of its parts, rather than reverting
-; to pointer width because an individual part is too short to qualify.
+; part count must still be charged for all of its lanes. And the index width is
+; chosen once for the whole operation, from the indices themselves, so the same
+; data type is priced differently depending on how wide its indices are.
+;
+; The code-size cost of a hardware gather/scatter is its part count, so those
+; numbers can be read against CodeGen directly; the paired instruction counts
+; live in masked-gather-scatter-part-count.ll. Both targets gather in
+; hardware, but only the AVX-512 one scatters, so the AVX2 scatter below is
+; the cost of scalarizing rather than a part count.
 
 ; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX
-; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=haswell | FileCheck %s --check-prefix=AVX2
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake | FileCheck %s --check-prefix=AVX2
 
 ; A length that divides its split factor: unchanged, and the reference point for
 ; the two cases below.
@@ -19,7 +24,7 @@ define <8 x i32> @gather_v8i32(<8 x ptr> %ptrs, <8 x i1> %mask) {
 ; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <8 x i32> %v
 ;
 ; AVX2-LABEL: 'gather_v8i32'
-; AVX2-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <8 x i32> %v
 ;
   %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x i32> poison)
@@ -34,7 +39,7 @@ define <9 x i32> @gather_v9i32(<9 x ptr> %ptrs, <9 x i1> %mask) {
 ; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <9 x i32> %v
 ;
 ; AVX2-LABEL: 'gather_v9i32'
-; AVX2-NEXT:  Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 4 %ptrs, <9 x i1> %mask, <9 x i32> poison)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:15 CodeSize:3 Lat:42 SizeLat:15 for: %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 4 %ptrs, <9 x i1> %mask, <9 x i32> poison)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <9 x i32> %v
 ;
   %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> %ptrs, i32 4, <9 x i1> %mask, <9 x i32> poison)
@@ -54,9 +59,9 @@ define void @scatter_v9i32(<9 x i32> %val, <9 x ptr> %ptrs, <9 x i1> %mask) {
   ret void
 }
 
-; Index width across a split: the GEP indices are 32-bit, and v24 is wide enough
-; to qualify for narrowing, so all three parts are priced with a dword index
-; even though a single part on its own would not qualify.
+; Index width across a split: the GEP indices are 32-bit, so the operation is
+; priced with a dword index. Sixteen dword-indexed lanes fill a part, so these
+; 24 lanes take two, and CodeGen emits two vpgatherdd.
 define <24 x i32> @gather_v24i32_dword_index(ptr %base, <24 x i32> %idx, <24 x i1> %mask) {
 ; SKX-LABEL: 'gather_v24i32_dword_index'
 ; SKX-NEXT:  Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
@@ -65,7 +70,7 @@ define <24 x i32> @gather_v24i32_dword_index(ptr %base, <24 x i32> %idx, <24 x i
 ;
 ; AVX2-LABEL: 'gather_v24i32_dword_index'
 ; AVX2-NEXT:  Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
-; AVX2-NEXT:  Cost Model: Found costs of RThru:82 CodeSize:106 Lat:178 SizeLat:106 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:30 CodeSize:3 Lat:102 SizeLat:30 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
 ;
   %ptrs = getelementptr inbounds i32, ptr %base, <24 x i32> %idx
@@ -73,16 +78,19 @@ define <24 x i32> @gather_v24i32_dword_index(ptr %base, <24 x i32> %idx, <24 x i
   ret <24 x i32> %v
 }
 
-; Control for the above: genuinely 64-bit indices cannot narrow.
+; Control for the above: genuinely 64-bit indices cannot narrow, so only eight
+; lanes fill a part and the same 24 lanes take three vpgatherqd. Their indices
+; occupy four legal registers, one of which holds no live lane, so a part count
+; taken from the register count would overstate the work by one.
 define <24 x i32> @gather_v24i32_qword_index(ptr %base, <24 x i64> %idx, <24 x i1> %mask) {
 ; SKX-LABEL: 'gather_v24i32_qword_index'
 ; SKX-NEXT:  Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
-; SKX-NEXT:  Cost Model: Found costs of RThru:32 CodeSize:4 Lat:104 SizeLat:32 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; SKX-NEXT:  Cost Model: Found costs of RThru:30 CodeSize:3 Lat:102 SizeLat:30 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
 ;
 ; AVX2-LABEL: 'gather_v24i32_qword_index'
 ; AVX2-NEXT:  Cost Model: Found costs of 0 for: %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
-; AVX2-NEXT:  Cost Model: Found costs of RThru:82 CodeSize:106 Lat:178 SizeLat:106 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:36 CodeSize:6 Lat:108 SizeLat:36 for: %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> align 4 %ptrs, <24 x i1> %mask, <24 x i32> poison)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <24 x i32> %v
 ;
   %ptrs = getelementptr inbounds i32, ptr %base, <24 x i64> %idx
diff --git a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll
index f835a8aedbbcc..bf19e53e1a374 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll
@@ -1909,7 +1909,7 @@ define <16 x float> @test_gather_16f32_const_mask(ptr %base, <16 x i32> %ind) {
 ; SKL-LABEL: 'test_gather_16f32_const_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_const_mask'
@@ -1953,7 +1953,7 @@ define <16 x float> @test_gather_16f32_var_mask(ptr %base, <16 x i32> %ind, <16
 ; SKL-LABEL: 'test_gather_16f32_var_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_var_mask'
@@ -2051,7 +2051,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
 ; SKL-NEXT:  Cost Model: Found costs of RThru:1 CodeSize:1 Lat:3 SizeLat:2 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> poison, <16 x i32> zeroinitializer
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_const_mask2'
diff --git a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
index c02dd6db8b99f..5797e9bd31017 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
@@ -745,7 +745,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:12 Lat:18 SizeLat:12 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:77 CodeSize:93 Lat:141 SizeLat:93 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SSE2-NEXT:  Cost Model: Found costs of RThru:49 CodeSize:58 Lat:85 SizeLat:58 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; SSE2-NEXT:  Cost Model: Found costs of RThru:49 CodeSize:58 Lat:85 SizeLat:58 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:39 CodeSize:47 Lat:71 SizeLat:47 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:24 Lat:36 SizeLat:24 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:12 Lat:18 SizeLat:12 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -773,7 +773,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:49 CodeSize:65 Lat:113 SizeLat:65 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SSE42-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:37 Lat:64 SizeLat:37 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; SSE42-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:37 Lat:64 SizeLat:37 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:25 CodeSize:33 Lat:57 SizeLat:33 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:17 Lat:29 SizeLat:17 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -801,7 +801,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:55 CodeSize:71 Lat:119 SizeLat:71 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; AVX1-NEXT:  Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; AVX1-NEXT:  Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:14 CodeSize:18 Lat:30 SizeLat:18 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -829,7 +829,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:55 CodeSize:71 Lat:119 SizeLat:71 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; AVX2-NEXT:  Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:14 CodeSize:18 Lat:30 SizeLat:18 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -857,7 +857,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:17 CodeSize:4 Lat:44 SizeLat:17 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:15 CodeSize:3 Lat:42 SizeLat:15 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -885,7 +885,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; KNL-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; KNL-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:17 CodeSize:21 Lat:33 SizeLat:21 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -913,7 +913,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SKX-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 undef, <9 x i1> %m9, <9 x i32> undef)
+; SKX-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 1 poison, <9 x i1> %m9, <9 x i32> poison)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -943,7 +943,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
   %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> undef, i32 1, <1 x i1> %m1, <1 x i64> undef)
 
   %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> undef, i32 1, <16 x i1> %m16, <16 x i32> undef)
-  %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> undef, i32 1, <9 x i1> %m9, <9 x i32> undef)
+  %V9I32 = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> poison, i32 1, <9 x i1> %m9, <9 x i32> poison)
   %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> undef, i32 1, <8 x i1> %m8, <8 x i32> undef)
   %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> undef, i32 1, <4 x i1> %m4, <4 x i32> undef)
   %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> undef, i32 1, <2 x i1> %m2, <2 x i32> undef)
@@ -976,7 +976,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:12 Lat:12 SizeLat:12 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:77 CodeSize:93 Lat:93 SizeLat:93 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; SSE2-NEXT:  Cost Model: Found costs of RThru:43 CodeSize:52 Lat:52 SizeLat:52 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; SSE2-NEXT:  Cost Model: Found costs of RThru:43 CodeSize:52 Lat:52 SizeLat:52 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:39 CodeSize:47 Lat:47 SizeLat:47 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:24 Lat:24 SizeLat:24 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:12 Lat:12 SizeLat:12 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1004,7 +1004,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:49 CodeSize:65 Lat:65 SizeLat:65 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; SSE42-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:37 Lat:37 SizeLat:37 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; SSE42-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:37 Lat:37 SizeLat:37 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:25 CodeSize:33 Lat:33 SizeLat:33 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:17 Lat:17 SizeLat:17 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1032,7 +1032,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; AVX1-NEXT:  Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; AVX1-NEXT:  Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1060,7 +1060,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; AVX2-NEXT:  Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1088,7 +1088,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SKL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:55 CodeSize:71 Lat:71 SizeLat:71 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; SKL-NEXT:  Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; SKL-NEXT:  Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:36 Lat:36 SizeLat:36 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:14 CodeSize:18 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1116,7 +1116,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; KNL-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; KNL-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:17 CodeSize:21 Lat:21 SizeLat:21 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1144,7 +1144,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
-; SKX-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> align 1 undef, <9 x i1> %m9)
+; SKX-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> align 1 poison, <9 x i1> %m9)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1174,7 +1174,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
   call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> undef, i32 1, <1 x i1> %m1)
 
   call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> undef, i32 1, <16 x i1> %m16)
-  call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> undef, <9 x ptr> undef, i32 1, <9 x i1> %m9)
+  call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> poison, <9 x ptr> poison, i32 1, <9 x i1> %m9)
   call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> undef, i32 1, <8 x i1> %m8)
   call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> undef, i32 1, <4 x i1> %m4)
   call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> undef, i32 1, <2 x i1> %m2)
@@ -1901,43 +1901,43 @@ define <16 x float> @test_gather_16f32_const_mask(ptr %base, <16 x i32> %ind) {
 ; SSE2-LABEL: 'test_gather_16f32_const_mask'
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:8 Lat:8 SizeLat:8 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SSE2-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SSE2-NEXT:  Cost Model: Found costs of RThru:60 CodeSize:60 Lat:108 SizeLat:60 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; SSE2-NEXT:  Cost Model: Found costs of RThru:60 CodeSize:60 Lat:108 SizeLat:60 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> poison)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; SSE42-LABEL: 'test_gather_16f32_const_mask'
 ; SSE42-NEXT:  Cost Model: Found costs of 8 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SSE42-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SSE42-NEXT:  Cost Model: Found costs of RThru:44 CodeSize:44 Lat:92 SizeLat:44 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; SSE42-NEXT:  Cost Model: Found costs of RThru:44 CodeSize:44 Lat:92 SizeLat:44 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> poison)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX1-LABEL: 'test_gather_16f32_const_mask'
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:14 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; AVX1-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; AVX1-NEXT:  Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; AVX1-NEXT:  Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> poison)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX2-LABEL: 'test_gather_16f32_const_mask'
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; AVX2-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; AVX2-NEXT:  Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> poison)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; SKL-LABEL: 'test_gather_16f32_const_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> poison)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_const_mask'
 ; AVX512-NEXT:  Cost Model: Found costs of RThru:3 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; AVX512-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; AVX512-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; AVX512-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> poison)
 ; AVX512-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
   %sext_ind = sext <16 x i32> %ind to <16 x i64>
   %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
 
-  %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> %gep.v, i32 4, <16 x i1> <i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true>, <16 x float> undef)
+  %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> %gep.v, i32 4, <16 x i1> <i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true>, <16 x float> poison)
   ret <16 x float>%res
 }
 
@@ -1945,43 +1945,43 @@ define <16 x float> @test_gather_16f32_var_mask(ptr %base, <16 x i32> %ind, <16
 ; SSE2-LABEL: 'test_gather_16f32_var_mask'
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:8 Lat:8 SizeLat:8 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SSE2-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SSE2-NEXT:  Cost Model: Found costs of RThru:61 CodeSize:77 Lat:125 SizeLat:77 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SSE2-NEXT:  Cost Model: Found costs of RThru:61 CodeSize:77 Lat:125 SizeLat:77 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> poison)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; SSE42-LABEL: 'test_gather_16f32_var_mask'
 ; SSE42-NEXT:  Cost Model: Found costs of 8 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SSE42-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SSE42-NEXT:  Cost Model: Found costs of RThru:45 CodeSize:61 Lat:109 SizeLat:61 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SSE42-NEXT:  Cost Model: Found costs of RThru:45 CodeSize:61 Lat:109 SizeLat:61 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> poison)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX1-LABEL: 'test_gather_16f32_var_mask'
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:14 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; AVX1-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; AVX1-NEXT:  Cost Model: Found costs of RThru:51 CodeSize:67 Lat:115 SizeLat:67 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; AVX1-NEXT:  Cost Model: Found costs of RThru:51 CodeSize:67 Lat:115 SizeLat:67 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> poison)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX2-LABEL: 'test_gather_16f32_var_mask'
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; AVX2-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; AVX2-NEXT:  Cost Model: Found costs of RThru:51 CodeSize:67 Lat:115 SizeLat:67 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:51 CodeSize:67 Lat:115 SizeLat:67 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> poison)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; SKL-LABEL: 'test_gather_16f32_var_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> poison)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_var_mask'
 ; AVX512-NEXT:  Cost Model: Found costs of RThru:3 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; AVX512-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; AVX512-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; AVX512-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> poison)
 ; AVX512-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
   %sext_ind = sext <16 x i32> %ind to <16 x i64>
   %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
 
-  %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> %gep.v, i32 4, <16 x i1> %mask, <16 x float> undef)
+  %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> %gep.v, i32 4, <16 x i1> %mask, <16 x float> poison)
   ret <16 x float>%res
 }
 
@@ -2035,7 +2035,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
 ; SSE2-NEXT:  Cost Model: Found costs of 1 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> undef, <16 x i32> zeroinitializer
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:8 Lat:8 SizeLat:8 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SSE2-NEXT:  Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; SSE2-NEXT:  Cost Model: Found costs of RThru:60 CodeSize:60 Lat:108 SizeLat:60 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; SSE2-NEXT:  Cost Model: Found costs of RThru:60 CodeSize:60 Lat:108 SizeLat:60 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> poison)
 ; SSE2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; SSE42-LABEL: 'test_gather_16f32_const_mask2'
@@ -2043,7 +2043,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
 ; SSE42-NEXT:  Cost Model: Found costs of 1 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> undef, <16 x i32> zeroinitializer
 ; SSE42-NEXT:  Cost Model: Found costs of 8 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SSE42-NEXT:  Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; SSE42-NEXT:  Cost Model: Found costs of RThru:44 CodeSize:44 Lat:92 SizeLat:44 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; SSE42-NEXT:  Cost Model: Found costs of RThru:44 CodeSize:44 Lat:92 SizeLat:44 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> poison)
 ; SSE42-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX1-LABEL: 'test_gather_16f32_const_mask2'
@@ -2051,7 +2051,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:2 Lat:3 SizeLat:3 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> undef, <16 x i32> zeroinitializer
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:14 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; AVX1-NEXT:  Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; AVX1-NEXT:  Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; AVX1-NEXT:  Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> poison)
 ; AVX1-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX2-LABEL: 'test_gather_16f32_const_mask2'
@@ -2059,7 +2059,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:1 CodeSize:1 Lat:3 SizeLat:2 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> undef, <16 x i32> zeroinitializer
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; AVX2-NEXT:  Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; AVX2-NEXT:  Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:50 CodeSize:50 Lat:98 SizeLat:50 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> poison)
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; SKL-LABEL: 'test_gather_16f32_const_mask2'
@@ -2067,7 +2067,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
 ; SKL-NEXT:  Cost Model: Found costs of RThru:1 CodeSize:1 Lat:3 SizeLat:2 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> undef, <16 x i32> zeroinitializer
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> poison)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_const_mask2'
@@ -2075,7 +2075,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
 ; AVX512-NEXT:  Cost Model: Found costs of RThru:1 CodeSize:1 Lat:3 SizeLat:1 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> undef, <16 x i32> zeroinitializer
 ; AVX512-NEXT:  Cost Model: Found costs of RThru:3 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; AVX512-NEXT:  Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; AVX512-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; AVX512-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:1 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> poison)
 ; AVX512-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
   %broadcast.splatinsert = insertelement <16 x ptr> undef, ptr %base, i32 0
@@ -2084,7 +2084,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
   %sext_ind = sext <16 x i32> %ind to <16 x i64>
   %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
 
-  %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> %gep.random, i32 4, <16 x i1> <i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true>, <16 x float> undef)
+  %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> %gep.random, i32 4, <16 x i1> <i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true, i1 true>, <16 x float> poison)
   ret <16 x float>%res
 }
 
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll
index 2d5a30019bacd..371338b6bcb78 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll
@@ -50,9 +50,9 @@ define void @test() {
 ; AVX2-FASTGATHER:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB = load i32, ptr %inB, align 4
 ; AVX2-FASTGATHER:  Cost of 4 for VF 2: WIDEN ir<%valB> = load ir<%inB>
 ; AVX2-FASTGATHER:  Cost of 6 for VF 4: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER:  Cost of 12 for VF 8: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER:  Cost of 24 for VF 16: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER:  Cost of 48 for VF 32: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER:  Cost of 10 for VF 8: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER:  Cost of 20 for VF 16: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER:  Cost of 40 for VF 32: WIDEN ir<%valB> = load ir<%inB>
 ;
 ; AVX512-LABEL: 'test'
 ; AVX512:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB = load i32, ptr %inB, align 4
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll
index 8f5da77027970..0cf2b6b9c12b2 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll
@@ -43,9 +43,9 @@ define void @test() {
 ; AVX2-FASTGATHER:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB.loaded = load i32, ptr %inB, align 4
 ; AVX2-FASTGATHER:  Cost of 4 for VF 2: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ; AVX2-FASTGATHER:  Cost of 6 for VF 4: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER:  Cost of 12 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER:  Cost of 24 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER:  Cost of 48 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER:  Cost of 10 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER:  Cost of 20 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER:  Cost of 40 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ;
 ; AVX512-LABEL: 'test'
 ; AVX512:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB.loaded = load i32, ptr %inB, align 4
diff --git a/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll b/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll
index 7f0ba19d3e7a8..bc2bea9708227 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll
@@ -123,28 +123,28 @@ define void @replicate_sext(i32 %N, ptr %dst, ptr %src) #0 {
 ; CHECK-NEXT:    [[FOUND_CONFLICT:%.*]] = and i1 [[BOUND0]], [[BOUND1]]
 ; CHECK-NEXT:    br i1 [[FOUND_CONFLICT]], label %[[SCALAR_PH]], label %[[VECTOR_PH:.*]]
 ; CHECK:       [[VECTOR_PH]]:
-; CHECK-NEXT:    [[N_MOD_VF:%.*]] = and i32 [[TMP0]], 3
+; CHECK-NEXT:    [[N_MOD_VF:%.*]] = and i32 [[TMP0]], 7
 ; CHECK-NEXT:    [[TMP18:%.*]] = icmp eq i32 [[N_MOD_VF]], 0
-; CHECK-NEXT:    [[TMP19:%.*]] = select i1 [[TMP18]], i32 4, i32 [[N_MOD_VF]]
+; CHECK-NEXT:    [[TMP19:%.*]] = select i1 [[TMP18]], i32 8, i32 [[N_MOD_VF]]
 ; CHECK-NEXT:    [[N_VEC:%.*]] = sub i32 [[TMP0]], [[TMP19]]
 ; CHECK-NEXT:    [[TMP20:%.*]] = shl i32 [[N_VEC]], 2
 ; CHECK-NEXT:    [[TMP21:%.*]] = mul i32 [[N_VEC]], 3
 ; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
 ; CHECK:       [[VECTOR_BODY]]:
 ; CHECK-NEXT:    [[INDEX:%.*]] = phi i32 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT:    [[VEC_IND:%.*]] = phi <4 x i32> [ <i32 0, i32 3, i32 6, i32 9>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[VEC_IND:%.*]] = phi <8 x i32> [ <i32 0, i32 3, i32 6, i32 9, i32 12, i32 15, i32 18, i32 21>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
 ; CHECK-NEXT:    [[OFFSET_IDX:%.*]] = shl i32 [[INDEX]], 2
 ; CHECK-NEXT:    [[TMP22:%.*]] = sext i32 [[OFFSET_IDX]] to i64
 ; CHECK-NEXT:    [[TMP23:%.*]] = getelementptr nusw i32, ptr [[SRC]], i64 [[TMP22]]
-; CHECK-NEXT:    [[WIDE_VEC:%.*]] = load <16 x i32>, ptr [[TMP23]], align 4, !alias.scope [[META4:![0-9]+]]
-; CHECK-NEXT:    [[STRIDED_VEC:%.*]] = shufflevector <16 x i32> [[WIDE_VEC]], <16 x i32> poison, <4 x i32> <i32 0, i32 4, i32 8, i32 12>
-; CHECK-NEXT:    [[STRIDED_VEC9:%.*]] = shufflevector <16 x i32> [[WIDE_VEC]], <16 x i32> poison, <4 x i32> <i32 1, i32 5, i32 9, i32 13>
-; CHECK-NEXT:    [[TMP25:%.*]] = sext <4 x i32> [[VEC_IND]] to <4 x i64>
-; CHECK-NEXT:    [[TMP27:%.*]] = getelementptr i32, ptr [[DST]], <4 x i64> [[TMP25]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[STRIDED_VEC]], <4 x ptr> align 4 [[TMP27]], <4 x i1> splat (i1 true)), !alias.scope [[META7:![0-9]+]], !noalias [[META4]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[STRIDED_VEC9]], <4 x ptr> align 4 [[TMP27]], <4 x i1> splat (i1 true)), !alias.scope [[META7]], !noalias [[META4]]
-; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 4
-; CHECK-NEXT:    [[VEC_IND_NEXT]] = add <4 x i32> [[VEC_IND]], splat (i32 12)
+; CHECK-NEXT:    [[WIDE_VEC:%.*]] = load <32 x i32>, ptr [[TMP23]], align 4, !alias.scope [[META4:![0-9]+]]
+; CHECK-NEXT:    [[STRIDED_VEC:%.*]] = shufflevector <32 x i32> [[WIDE_VEC]], <32 x i32> poison, <8 x i32> <i32 0, i32 4, i32 8, i32 12, i32 16, i32 20, i32 24, i32 28>
+; CHECK-NEXT:    [[STRIDED_VEC2:%.*]] = shufflevector <32 x i32> [[WIDE_VEC]], <32 x i32> poison, <8 x i32> <i32 1, i32 5, i32 9, i32 13, i32 17, i32 21, i32 25, i32 29>
+; CHECK-NEXT:    [[TMP24:%.*]] = sext <8 x i32> [[VEC_IND]] to <8 x i64>
+; CHECK-NEXT:    [[WIDE_GEP:%.*]] = getelementptr i32, ptr [[DST]], <8 x i64> [[TMP24]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[STRIDED_VEC]], <8 x ptr> align 4 [[WIDE_GEP]], <8 x i1> splat (i1 true)), !alias.scope [[META7:![0-9]+]], !noalias [[META4]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[STRIDED_VEC2]], <8 x ptr> align 4 [[WIDE_GEP]], <8 x i1> splat (i1 true)), !alias.scope [[META7]], !noalias [[META4]]
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 8
+; CHECK-NEXT:    [[VEC_IND_NEXT]] = add <8 x i32> [[VEC_IND]], splat (i32 24)
 ; CHECK-NEXT:    [[TMP26:%.*]] = icmp eq i32 [[INDEX_NEXT]], [[N_VEC]]
 ; CHECK-NEXT:    br i1 [[TMP26]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP9:![0-9]+]]
 ; CHECK:       [[MIDDLE_BLOCK]]:

>From 6ce86c55d8a8e962f2b4fcbe0a54a2f89d1a82a6 Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Fri, 11 Sep 2026 15:39:05 +0530
Subject: [PATCH 4/5] [X86][CostModel] Derive the gather index width from its
 value range

The index width was decided from how the index was written, its declared
type and its extension opcode, where CodeGen decides it from the values
the index carries and narrows to a dword whenever they fit one. The
gather/scatter combine in X86ISelLowering tests that with
ComputeNumSignBits, so ask the same question here.

Reading the spelling misses in both directions. A zero-extension from a
narrow type cannot leave the signed dword range and CodeGen narrows it,
but it is not an SExtInst and was priced as the qword form. An index
wider than a pointer is cut down to pointer width rather than to a dword,
but it is not exactly 64 bits and was priced as a dword. Zero-extending a
dword has to stay distinct from the narrow case, since its top bit is no
longer a sign, and the sign-bit count draws that boundary by itself: 32
significant bits for a sign-extended dword against 33 for a zero-extended
one. Constant vectors were not examined at all, so indices well outside
the range were taken for dwords.

A constant holding the same value in every lane is a fixed byte offset
and folds into the base, so struct members and splats are still skipped.
One that varies across lanes reaches the index operand, and is examined
like any other.

Scalar index operands keep the conservative width. LoopVectorize asks
with the original scalar address, where the range of an induction
variable says nothing about the index the widened form will see, and
reasoning from it moves vectorisation decisions rather than correcting
costs. Where no index varies across lanes at all, though, the component
that varies has to be the base pointer, so the widened form gathers from
a vector of pointers and is indexed at pointer width. That is what a
pointer chase such as p[i]->y needs and did not previously get.

The stride follows the same combine rather than standing apart from it. A
stride the scale field encodes is applied by the addressing mode, so the
index still reaches the combine as an extension and narrows on its range
alone. Any other stride is multiplied in, so the product must fit a
signed dword with the stride included, and it is no longer an extension
node -- the route the combine normally narrows through. Two routes
remain: a constant index narrows by being folded, and otherwise narrowing
happens only where it removes an illegal type. That last one is why the
subtargets disagree. Sixteen qword indices are illegal on both, but the
dword form they narrow to is legal only once a 512-bit register is
available, so the same stride-12 gather keeps qword indices on AVX2 and
narrows to dwords on AVX-512. Treating a nonencodable stride as
pointer-width regardless mispriced every narrow-index form on AVX-512.

Checked by pairing the code-size cost against the gather count llc emits
for the same function, over five index spellings, five strides, four data
types, seven vector lengths and seven subtargets: 1952 of 2050 agreed
before, 2050 of 2050 after. The remaining 98 were the two families this
corrects, a zero-extension too narrow to leave the dword range and a
narrow index under a stride the scale field cannot encode.
---
 .../lib/Target/X86/X86TargetTransformInfo.cpp |  88 ++++++++--
 .../X86/masked-gather-scatter-part-count.ll   | 154 +++++++++++++++++-
 .../LoopVectorize/X86/gather-index-width.ll   |  82 ++++++++++
 3 files changed, 303 insertions(+), 21 deletions(-)
 create mode 100644 llvm/test/Transforms/LoopVectorize/X86/gather-index-width.ll

diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
index 9cebe1c422daf..383490b1fc4b8 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
@@ -51,6 +51,7 @@
 #include "X86TargetTransformInfo.h"
 #include "llvm/ADT/SmallBitVector.h"
 #include "llvm/Analysis/TargetTransformInfo.h"
+#include "llvm/Analysis/ValueTracking.h"
 #include "llvm/CodeGen/Analysis.h"
 #include "llvm/CodeGen/BasicTTIImpl.h"
 #include "llvm/CodeGen/CostTable.h"
@@ -6670,26 +6671,89 @@ X86TTIImpl::getGSVectorCost(unsigned Opcode, TTI::TargetCostKind CostKind,
     for (gep_type_iterator GTI = gep_type_begin(GEP), GTE = gep_type_end(GEP);
          GTI != GTE; ++GTI) {
       const Value *Operand = GTI.getOperand();
-      if (isa<Constant>(Operand))
+
+      // An index that holds the same value in every lane contributes a fixed
+      // byte offset, which folds into the base pointer instead of reaching the
+      // index operand. That covers every struct member, named by a scalar
+      // constant, as well as a splat. A constant that is not a splat does
+      // reach the index, so it is examined like any other.
+      if (isa<Constant>(Operand) && (!Operand->getType()->isVectorTy() ||
+                                     getSplatValue(Operand) != nullptr))
         continue;
-      Type *IndxTy = Operand->getType();
-      if (auto *IndexVTy = dyn_cast<VectorType>(IndxTy))
-        IndxTy = IndexVTy->getElementType();
-      if ((IndxTy->getPrimitiveSizeInBits() == 64 && !isa<SExtInst>(Operand)) ||
-          ++NumOfVarIndices > 1)
+
+      if (++NumOfVarIndices > 1)
         return IndexSize; // 64
-      // The narrow index only reaches the instruction if the addressing mode
-      // can apply its stride as a scale, and the scale field encodes 1, 2, 4
-      // and 8. Any other stride has to be multiplied into the index first,
-      // and that product is pointer-width, so the operation ends up gathering
-      // with 64-bit indices however narrow the index started out.
+
+      unsigned IndexBits;
+      if (Operand->getType()->isVectorTy()) {
+        // Whether the index reaches the instruction as a dword is a property
+        // of the values it takes, not of how it was written: CodeGen narrows
+        // it when it is representable in a signed dword, which is what the
+        // gather/scatter combine in X86ISelLowering tests with
+        // ComputeNumSignBits. Deciding on the declared width and the
+        // extension opcode instead misses in both directions -- a
+        // zero-extension from a narrow type is representable although it is
+        // not an SExtInst, and an index wider than a pointer is not
+        // representable although it is not exactly 64 bits -- and a constant
+        // vector was not examined at all.
+        IndexBits = ComputeMaxSignificantBits(Operand, DL);
+      } else {
+        // A scalar index reaches here from a caller that has not widened the
+        // address yet, so the component that varies across lanes is not
+        // visible and its range says nothing about the index the instruction
+        // will see. Keep the conservative width for those until the query
+        // carries the vector form.
+        IndexBits = Operand->getType()->getPrimitiveSizeInBits() == 64 &&
+                            !isa<SExtInst>(Operand)
+                        ? 64
+                        : 32;
+      }
+      if (IndexBits > 32)
+        return IndexSize; // 64
+
       TypeSize EltSize = DL.getTypeAllocSize(GTI.getIndexedType());
       if (EltSize.isScalable())
         return IndexSize; // 64
       uint64_t Stride = EltSize.getFixedValue();
-      if (!isPowerOf2_64(Stride) || Stride > 8)
+
+      // A stride the scale field encodes -- 1, 2, 4 or 8 -- is applied by the
+      // addressing mode, so the index reaches the combine as the extension it
+      // was written as, and narrowing follows from its range alone.
+      if (isPowerOf2_64(Stride) && Stride <= 8)
+        continue;
+
+      // Any other stride is multiplied into the index first, and the product
+      // is what the combine sees. It has to fit a signed dword with the
+      // stride included, and it is no longer an extension node, which is the
+      // route the combine normally narrows through.
+      if (IndexBits + Log2_64_Ceil(Stride) > 32)
+        return IndexSize; // 64
+
+      // Two routes remain for that product. A constant index narrows by being
+      // folded to its truncated form. Otherwise the combine narrows only
+      // where doing so removes an illegal type, which depends on the widest
+      // vector the subtarget has: sixteen qword indices are illegal
+      // everywhere, but the dword form they would narrow to is itself legal
+      // only once a 512-bit register is available. The same gather therefore
+      // keeps qword indices on AVX2 and narrows to dwords on AVX-512.
+      if (isa<Constant>(Operand))
+        continue;
+      LLVMContext &Ctx = SrcVTy->getContext();
+      if (TLI->isTypeLegal(EVT::getVectorVT(Ctx, MVT::i64, VF)) ||
+          !TLI->isTypeLegal(EVT::getVectorVT(Ctx, MVT::i32, VF)))
         return IndexSize; // 64
     }
+
+    // Every index holds the same value in every lane, so the component that
+    // varies has to be the base pointer, and the operation gathers from a
+    // vector of pointers at pointer width. A caller that has already widened
+    // the address says as much with a vector base, which is answered above;
+    // LoopVectorize asks with the original scalar address, where a pointer
+    // chase like p[i]->y is otherwise indistinguishable from a uniform base
+    // reached through a narrow index.
+    if (NumOfVarIndices == 0)
+      return IndexSize; // 64
+
     return (unsigned)32;
   };
 
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll
index fd6374b595275..ab7bae63197eb 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-part-count.ll
@@ -4,13 +4,19 @@
 ;
 ; The pairing is the point of the test: the code-size cost of a hardware
 ; gather/scatter is its part count, so every cost check below has a matching
-; instruction count taken from the same module. Three properties would
+; instruction count taken from the same module. Four properties would
 ; otherwise drift unnoticed. A part count taken from the legalized register
 ; count overstates the work when legalization widens to a power of two; an
 ; index width derived only for wide AVX-512 vectors leaves a dword-indexed
 ; operation tied with the qword-indexed form that CodeGen splits into more
-; instructions; and the tail a scatter still emits under a zeroed mask reaches
-; only as far as the predicate widens, which depends on AVX512BW.
+; instructions; an index width read from the declared type and the extension
+; opcode rather than from the range of values the index carries misreads both
+; a zero-extension too narrow to leave that range and an index wider than a
+; pointer; a stride the scale field cannot encode does not by itself force a
+; qword index, since the product may still fit a dword and then narrows or not
+; according to which vector types the subtarget makes legal; and the tail a
+; scatter still emits under a zeroed mask reaches only as far as the predicate
+; widens, which depends on AVX512BW.
 ;
 ; The throughput run pins the other half of the cost, the per-lane term, which
 ; code size does not expose: it is charged for the lanes that are really
@@ -80,6 +86,26 @@ define <9 x i32> @gather_v9i32_qword_index(ptr %base, <9 x i64> %idx, <9 x i1> %
   ret <9 x i32> %v
 }
 
+; The same shape indexed more widely than a pointer, which is not a narrow
+; index: it is cut down to pointer width rather than to a dword, so it costs
+; what the qword form above costs. Asking whether the index type is exactly 64
+; bits wide would miss it and price it as a dword.
+; SKX-COST-LABEL: 'gather_v9i32_oversized_index'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v9i32_oversized_index:
+; SKX-ASM-COUNT-2: vpgatherqd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v9i32_oversized_index'
+; AVX2-COST: cost of 3 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v9i32_oversized_index:
+; AVX2-ASM-COUNT-3: vpgatherqd
+; AVX2-ASM-NOT: vpgatherqd
+define <9 x i32> @gather_v9i32_oversized_index(ptr %base, <9 x i128> %idx, <9 x i1> %mask) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <9 x i128> %idx
+  %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> %ptrs, i32 4, <9 x i1> %mask, <9 x i32> poison)
+  ret <9 x i32> %v
+}
+
 ; A length below the width at which the index used to be examined, so both index
 ; widths would otherwise be priced the same on either target.
 ; SKX-COST-LABEL: 'gather_v17i32_dword_index'
@@ -392,12 +418,13 @@ define <8 x ptr> @gather_v8ptr_qword_index_first2_mask(ptr %base, <8 x i64> %idx
   ret <8 x ptr> %r
 }
 
-; A narrow index only stays narrow if the addressing mode can apply its stride
-; as a scale, and the scale field encodes 1, 2, 4 and 8. This GEP walks an
-; array of three-word structures, the form a loop vectorizer emits for
-; a[idx[i]].f, so its stride is twelve: the multiply is folded into the index
-; and the gather ends up indexed by qwords, taking twice the instructions of
-; the four-byte-stride control below.
+; A stride the scale field cannot encode -- it encodes 1, 2, 4 and 8 -- is
+; multiplied into the index instead, and that product is what CodeGen narrows
+; from. This GEP walks an array of three-word structures, the form a loop
+; vectorizer emits for a[idx[i]].f, so its stride is twelve: a dword index
+; scaled by twelve needs thirty-six bits and no longer fits a signed dword, so
+; the gather ends up indexed by qwords, taking twice the instructions of the
+; four-byte-stride control below.
 ; SKX-COST-LABEL: 'gather_v16i32_struct_stride'
 ; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
 ; SKX-ASM-LABEL: gather_v16i32_struct_stride:
@@ -415,6 +442,35 @@ define <16 x i32> @gather_v16i32_struct_stride(ptr %base, <16 x i32> %idx, <16 x
   ret <16 x i32> %v
 }
 
+; The same stride over an index narrow enough that the product still fits a
+; signed dword, so a nonencodable stride does not by itself decide the width.
+; Fitting is not sufficient either: the multiply leaves the index as something
+; other than an extension node, and CodeGen then narrows it only where that
+; removes an illegal type. Sixteen qword indices are illegal everywhere, but
+; the dword form they would narrow to is itself legal only once a 512-bit
+; register is available, so the same gather keeps qword indices on AVX2 and
+; narrows to dwords on AVX-512.
+; SKX-COST-LABEL: 'gather_v16i32_struct_stride_narrow_index'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v16i32_struct_stride_narrow_index:
+; SKX-ASM-COUNT-1: vpgatherdd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v16i32_struct_stride_narrow_index'
+; AVX2-COST: cost of 4 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v16i32_struct_stride_narrow_index:
+; AVX2-ASM-COUNT-4: vpgatherqd
+; AVX2-ASM-NOT: vpgatherqd
+; KNL-COST-LABEL: 'gather_v16i32_struct_stride_narrow_index'
+; KNL-COST: cost of 1 for instruction: {{.*}}masked.gather
+; KNL-ASM-LABEL: gather_v16i32_struct_stride_narrow_index:
+; KNL-ASM-COUNT-1: vpgatherdd
+define <16 x i32> @gather_v16i32_struct_stride_narrow_index(ptr %base, <16 x i16> %idx, <16 x i1> %mask) {
+  %sext = sext <16 x i16> %idx to <16 x i64>
+  %ptrs = getelementptr inbounds {i32, i32, i32}, ptr %base, <16 x i64> %sext, i32 0
+  %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+  ret <16 x i32> %v
+}
+
 ; The same gather over a four-byte stride, which the scale field does encode,
 ; so the index stays a dword and one instruction covers twice the lanes.
 ; SKX-COST-LABEL: 'gather_v16i32_scaled_stride'
@@ -434,6 +490,86 @@ define <16 x i32> @gather_v16i32_scaled_stride(ptr %base, <16 x i32> %idx, <16 x
   ret <16 x i32> %v
 }
 
+; A zero-extension from a narrow type cannot reach the signed dword range, so
+; the index stays a dword and one instruction covers all sixteen lanes, exactly
+; as for the sign-extension above. Reading the extension opcode rather than the
+; range it produces would price this as the qword form it is not.
+; SKX-COST-LABEL: 'gather_v16i32_narrow_zext_index'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v16i32_narrow_zext_index:
+; SKX-ASM-COUNT-1: vpgatherdd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v16i32_narrow_zext_index'
+; AVX2-COST: cost of 2 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v16i32_narrow_zext_index:
+; AVX2-ASM-COUNT-2: vpgatherdd
+; AVX2-ASM-NOT: vpgatherdd
+define <16 x i32> @gather_v16i32_narrow_zext_index(ptr %base, <16 x i8> %idx, <16 x i1> %mask) {
+  %zext = zext <16 x i8> %idx to <16 x i64>
+  %ptrs = getelementptr inbounds i32, ptr %base, <16 x i64> %zext
+  %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+  ret <16 x i32> %v
+}
+
+; Zero-extending a dword is the case that has to stay distinct from it. The top
+; bit is no longer a sign, so the value can exceed what a signed dword index
+; addresses and the index stays a qword, taking twice the instructions of the
+; sign-extended form.
+; SKX-COST-LABEL: 'gather_v16i32_wide_zext_index'
+; SKX-COST: cost of 2 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v16i32_wide_zext_index:
+; SKX-ASM-COUNT-2: vpgatherqd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v16i32_wide_zext_index'
+; AVX2-COST: cost of 4 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v16i32_wide_zext_index:
+; AVX2-ASM-COUNT-4: vpgatherqd
+; AVX2-ASM-NOT: vpgatherqd
+define <16 x i32> @gather_v16i32_wide_zext_index(ptr %base, <16 x i32> %idx, <16 x i1> %mask) {
+  %zext = zext <16 x i32> %idx to <16 x i64>
+  %ptrs = getelementptr inbounds i32, ptr %base, <16 x i64> %zext
+  %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+  ret <16 x i32> %v
+}
+
+; A constant index that does fit a signed dword stays one, the control for the
+; out-of-range form below.
+; SKX-COST-LABEL: 'gather_v8i32_const_index'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v8i32_const_index:
+; SKX-ASM-COUNT-1: vpgatherdd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v8i32_const_index'
+; AVX2-COST: cost of 1 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v8i32_const_index:
+; AVX2-ASM-COUNT-1: vpgatherdd
+; AVX2-ASM-NOT: vpgatherdd
+define <8 x i32> @gather_v8i32_const_index(ptr %base, <8 x i1> %mask) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <8 x i64> <i64 0, i64 1, i64 2, i64 3, i64 4, i64 5, i64 6, i64 7>
+  %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x i32> poison)
+  ret <8 x i32> %v
+}
+
+; A constant index is an index like any other, and it has to be examined like
+; one: these values do not fit a signed dword, so the gather ends up indexed by
+; qwords. Eight qword indices fill one AVX-512 register, which emits a single
+; instruction at either width, so only AVX2 shows the split.
+; SKX-COST-LABEL: 'gather_v8i32_const_index_out_of_range'
+; SKX-COST: cost of 1 for instruction: {{.*}}masked.gather
+; SKX-ASM-LABEL: gather_v8i32_const_index_out_of_range:
+; SKX-ASM-COUNT-1: vpgatherqd
+; SKX-ASM-NOT: vpgather
+; AVX2-COST-LABEL: 'gather_v8i32_const_index_out_of_range'
+; AVX2-COST: cost of 2 for instruction: {{.*}}masked.gather
+; AVX2-ASM-LABEL: gather_v8i32_const_index_out_of_range:
+; AVX2-ASM-COUNT-2: vpgatherqd
+; AVX2-ASM-NOT: vpgatherqd
+define <8 x i32> @gather_v8i32_const_index_out_of_range(ptr %base, <8 x i1> %mask) {
+  %ptrs = getelementptr inbounds i32, ptr %base, <8 x i64> <i64 0, i64 2147483648, i64 4294967296, i64 6442450944, i64 8589934592, i64 10737418240, i64 12884901888, i64 15032385536>
+  %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x i32> poison)
+  ret <8 x i32> %v
+}
+
 declare <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr>, i32, <16 x i1>, <16 x i32>)
 declare <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr>, i32, <8 x i1>, <8 x ptr>)
 declare <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr>, i32, <8 x i1>, <8 x i32>)
diff --git a/llvm/test/Transforms/LoopVectorize/X86/gather-index-width.ll b/llvm/test/Transforms/LoopVectorize/X86/gather-index-width.ll
new file mode 100644
index 0000000000000..f8864fe157876
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/X86/gather-index-width.ll
@@ -0,0 +1,82 @@
+; Pins the index width a gather is costed at against the width CodeGen gives
+; it, for the consumer that asks before the address has been widened.
+;
+; LoopVectorize queries the cost of a widened load with the original scalar
+; address, where the component that varies across lanes is not spelled out: a
+; pointer chase and a uniform base reached through a narrow index both present
+; a GEP whose operands are scalars. The two lower differently, so each cost
+; check below is paired with the gather count llc emits for the same loop.
+;
+; These checks are maintained by hand rather than by
+; update_analyze_test_checks.py, which does not know about the paired llc run.
+
+; RUN: opt -passes=loop-vectorize -force-vector-width=8 -force-vector-interleave=1 -mtriple=x86_64-unknown-linux-gnu -mcpu=skylake -debug-only=loop-vectorize -disable-output < %s 2>&1 | FileCheck %s --check-prefix=COST
+; RUN: opt -passes=loop-vectorize -force-vector-width=8 -force-vector-interleave=1 -mtriple=x86_64-unknown-linux-gnu -mcpu=skylake -S < %s -o %t.ll
+; RUN: llc -mtriple=x86_64-unknown-linux-gnu -mcpu=skylake %t.ll -o - | FileCheck %s --check-prefix=ASM
+
+; REQUIRES: asserts
+
+%S = type { i32, i32 }
+
+ at A = global [1024 x i8] zeroinitializer, align 128
+ at B = global [1024 x i32] zeroinitializer, align 128
+
+; The base is loaded on each iteration and the field indices are constants, so
+; nothing in the index list varies across lanes and the widened form gathers
+; from a vector of pointers. That needs pointer-width indices, which hold half
+; as many lanes per register, so eight lanes take two instructions. Reading the
+; index list alone would find no varying index and price this as a dword form
+; covering all eight in one.
+; COST-LABEL: 'chase'
+; COST: Cost of 12 for VF 8: WIDEN ir<%y> = load ir<%field>
+; ASM-LABEL: chase:
+; ASM-COUNT-2: vpgatherqd
+; ASM-NOT: vpgather
+define i32 @chase(ptr noalias readonly %p) {
+entry:
+  br label %loop
+
+loop:
+  %iv = phi i64 [ 0, %entry ], [ %next, %loop ]
+  %sum = phi i32 [ 0, %entry ], [ %sum.next, %loop ]
+  %pi = getelementptr inbounds ptr, ptr %p, i64 %iv
+  %q = load ptr, ptr %pi, align 8
+  %field = getelementptr inbounds %S, ptr %q, i64 0, i32 1
+  %y = load i32, ptr %field, align 4
+  %sum.next = add i32 %sum, %y
+  %next = add nuw nsw i64 %iv, 1
+  %done = icmp eq i64 %next, 1024
+  br i1 %done, label %exit, label %loop
+
+exit:
+  ret i32 %sum.next
+}
+
+; The control, and the reason the case above cannot simply be assumed wide: the
+; base is a global and it is the index that varies, sign-extended from a byte
+; and so narrow enough to stay a dword. One instruction covers all eight lanes.
+; COST-LABEL: 'narrow_index'
+; COST: Cost of 10 for VF 8: WIDEN ir<%valB> = load ir<%inB>
+; ASM-LABEL: narrow_index:
+; ASM-COUNT-1: vpgatherdd
+; ASM-NOT: vpgather
+define i32 @narrow_index() {
+entry:
+  br label %loop
+
+loop:
+  %iv = phi i64 [ 0, %entry ], [ %next, %loop ]
+  %sum = phi i32 [ 0, %entry ], [ %sum.next, %loop ]
+  %inA = getelementptr inbounds [1024 x i8], ptr @A, i64 0, i64 %iv
+  %valA = load i8, ptr %inA, align 1
+  %ext = sext i8 %valA to i64
+  %inB = getelementptr inbounds [1024 x i32], ptr @B, i64 0, i64 %ext
+  %valB = load i32, ptr %inB, align 4
+  %sum.next = add i32 %sum, %valB
+  %next = add nuw nsw i64 %iv, 1
+  %done = icmp eq i64 %next, 1024
+  br i1 %done, label %exit, label %loop
+
+exit:
+  ret i32 %sum.next
+}

>From 9934a3d609a8d1b7be02d43df125c2c05f87d2a9 Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Thu, 17 Sep 2026 16:29:46 +0530
Subject: [PATCH 5/5] [X86][CostModel] Correct the release note on stride and
 index narrowing

The note still described the first, over-broad correction: that any
stride the scale field cannot encode makes the gather pointer-width.
The implementation no longer says that. A nonencodable stride is
multiplied into the index, and it is the product that has to fit a
signed dword, so a dword index scaled by twelve needs 36 bits and is
costed at pointer width while a narrower index under the same stride
still fits.

Whether an index that fits then narrows depends on the subtarget, which
the note did not mention at all: narrowing happens where it removes an
illegal type, and the dword form of sixteen indices is legal only once a
512-bit register is available. The same 16-lane stride-12 gather
therefore keeps qword indices on AVX2 and narrows to dwords on AVX-512,
as gather_v16i32_struct_stride_narrow_index pins.
---
 llvm/docs/ReleaseNotes.md | 18 +++++++++++++-----
 1 file changed, 13 insertions(+), 5 deletions(-)

diff --git a/llvm/docs/ReleaseNotes.md b/llvm/docs/ReleaseNotes.md
index ffdf8c1fc458a..ebb1fee0cd7f8 100644
--- a/llvm/docs/ReleaseNotes.md
+++ b/llvm/docs/ReleaseNotes.md
@@ -259,11 +259,19 @@ Makes programs 10x faster by doing Special New Thing.
 
   A narrow index only survives into the instruction if the addressing mode can
   apply the GEP's stride as a scale, and the scale field encodes 1, 2, 4 and 8.
-  Any other stride is multiplied into the index first, and the product is
-  pointer-width, so such a gather is now costed with 64-bit indices. This
-  corrects a gather over an array of structures, the form a vectorized
-  `a[idx[i]].f` takes, which was previously costed at half the instructions
-  CodeGen emits.
+  Any other stride is multiplied into the index first, and it is that product
+  which has to fit a signed dword. A dword index scaled by twelve needs 36 bits
+  and so is costed at pointer width, which corrects a gather over an array of
+  structures — the form a vectorized `a[idx[i]].f` takes — that was previously
+  costed at half the instructions CodeGen emits. A narrower index under the
+  same stride still fits, and is not forced to pointer width on that account.
+
+  Whether it then narrows depends on the subtarget. A constant index narrows by
+  being folded to its truncated form; otherwise narrowing happens only where it
+  removes an illegal type. Sixteen qword indices are illegal everywhere, but the
+  dword form they would narrow to is itself legal only once a 512-bit register
+  is available, so the same 16-lane stride-12 gather keeps qword indices on AVX2
+  and narrows to dwords on AVX-512, and is costed accordingly on each.
 
 * That index width is also taken from the address space the pointers live in,
   instead of assuming the width of address space 0. On x86 this matters for the



More information about the llvm-commits mailing list