[llvm] [X86][CostModel] Add per-shape gather/scatter cost tables for AMD znver4+ (PR #199488)

Sumukh J Bharadwaj via llvm-commits llvm-commits at lists.llvm.org
Wed Sep 2 02:08:37 PDT 2026


https://github.com/amd-subharad updated https://github.com/llvm/llvm-project/pull/199488

>From 7cf05669ac4f86e67ac7a8a1169d1f2b1545317b Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Wed, 2 Sep 2026 01:15:07 +0530
Subject: [PATCH 1/5] [X86][SchedModel] Add masked gather/scatter overrides to
 Znver4Model

Znver4Model had no overrides for the AVX-512 masked gather and scatter
instructions, so they fell back to the default vector load/store write:
one uop, latency 5, reciprocal throughput 0.33. On these parts they are
microcoded sequences costing tens of uops and tens of cycles, so the
default was wrong by more than an order of magnitude.

Add 21 SchedWriteRes overrides covering every EVEX masked gather
(VPGATHER* / VGATHER*) and scatter (VPSCATTER* / VSCATTER*) form, and
bind the AVX2 (VEX) gathers, which were also on the default write and
diverged from it by more than an order of magnitude. All but the 8 x i32
VEX form measure like their EVEX counterparts and reuse those entries;
VPGATHERDDY is genuinely cheaper and gets its own.

Agner and uops.info both show these sequences using the vector pipes
heavily, so every entry charges Zn4FPU0123. Only that resource carries a
fitted number, taken from the measured bottleneck as round(tput * 4);
every other resource carries its plain physical occupancy. Fitting each
resource independently let a non-limiting resource round past the
intended limiter -- the 2 x i64 scatter reported 2.7 cycles against an
intended 2.5 -- and deriving the bottleneck once removes that by
construction.

Throughput, uops and gather latency are measured on Znver4 (Ryzen 5
8645HS). Znver4Model also backs znver5 and znver6, which run these
sequences 25-55% faster and are therefore modeled pessimistically; a
per-part split is left as future work.

Entries are keyed by (#elements, element width, index width). Sharing one
entry between the dword- and qword-index encodings is wrong for six
shapes, worst at the 8 x i64 scatter: 13.94 cycles with a qword index
against 11.05 with a dword one, at equal uop counts over identical
address sets. The four 2-element shapes and both 4-element gathers agree
to within 0.03 cycles and keep a shared entry. The 2-element shapes are
force-scalarised by the auto-vectoriser but remain reachable through the
intrinsics, so they are measured rather than extrapolated.

Gather latency is measured from a dependency chain that feeds the loaded
data back into the index register, with the feedback op's own latency
subtracted. Scatter latency remains estimated, as there is no dependent
consumer to time.

The instruction destroys its own mask and so cannot be timed alone;
substituting an equal-cost kxnorw for the kmovd isolates the reload,
which only changes the total for the 2-element gathers (5.00 vs 3.98
cycles), so only those have it subtracted. Sub-cycle accounting stops
there: identical code in two processes agrees to a p90 of 0.24 cycles,
but the same instructions in a different schedule move by a p90 of 1.35,
so the table is accurate to about a cycle regardless.

llvm-mca now tracks hardware within 1.3% on every scatter shape and
within 1.1% on every gather of four elements or more. Add a CodeGen test
so the values are exercised by something other than the llvm-mca
resource tests, since these writes also feed the machine scheduler.
---
 llvm/lib/Target/X86/X86ScheduleZnver4.td      | 137 ++++++++++++++++++
 .../X86/znver4-gather-scatter-schedule.ll     |  36 +++++
 .../llvm-mca/X86/Znver4/resources-avx2.s      |  66 ++++-----
 .../llvm-mca/X86/Znver4/resources-avx512.s    |  66 ++++-----
 .../llvm-mca/X86/Znver4/resources-avx512vl.s  | 130 ++++++++---------
 5 files changed, 304 insertions(+), 131 deletions(-)
 create mode 100644 llvm/test/CodeGen/X86/znver4-gather-scatter-schedule.ll

diff --git a/llvm/lib/Target/X86/X86ScheduleZnver4.td b/llvm/lib/Target/X86/X86ScheduleZnver4.td
index 34e57f634a1e6..e3711a0701977 100644
--- a/llvm/lib/Target/X86/X86ScheduleZnver4.td
+++ b/llvm/lib/Target/X86/X86ScheduleZnver4.td
@@ -509,6 +509,143 @@ defm : Zn4WriteResInt<WriteLoad, [Zn4AGU012, Zn4Load], !add(Znver4Model.LoadLate
 // Does not cost anything by itself, only has latency, matching that of the WriteLoad,
 defm : Zn4WriteResInt<WriteVecMaskedGatherWriteback, [], !add(Znver4Model.LoadLatency, 1), [], 0>;
 
+// AVX-512 masked GATHER / SCATTER. Microcoded, so entries are keyed by
+// (#elements, element width, index width); the two index widths share an entry
+// only where they measure alike. Zn4FPU0123 carries the fitted bottleneck,
+// round(tput * 4), and every other resource its physical occupancy, so a
+// retune edits the Zn4FPU0123 number. Measured on Znver4 and accurate to about
+// a cycle; scatter latency is estimated rather than timed. Znver4Model also
+// backs znver5 and znver6, which run 25-55% faster and are pessimistic here.
+// The commit message describes how the numbers were measured.
+def Zn4WriteVPGATHERQDZ128 : SchedWriteRes<[Zn4AGU012, Zn4Load, Zn4FPLd01, Zn4FPU0123]> {
+  let ReleaseAtCycles = [2, 2, 2, 16]; let Latency = 13; let NumMicroOps = 18;
+}
+def : InstRW<[Zn4WriteVPGATHERQDZ128, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQDZ128rm, VGATHERQPSZ128rm)>;
+def Zn4WriteVPGATHERDDZ128 : SchedWriteRes<[Zn4AGU012, Zn4Load, Zn4FPLd01, Zn4FPU0123]> {
+  let ReleaseAtCycles = [4, 4, 4, 20]; let Latency = 15; let NumMicroOps = 24;
+}
+def : InstRW<[Zn4WriteVPGATHERDDZ128, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERDDZ128rm, VGATHERDPSZ128rm,
+                     VPGATHERQDZ256rm, VGATHERQPSZ256rm)>;
+def Zn4WriteVPGATHERDDZ256 : SchedWriteRes<[Zn4AGU012, Zn4Load, Zn4FPLd01, Zn4FPU0123]> {
+  let ReleaseAtCycles = [8, 8, 8, 37]; let Latency = 20; let NumMicroOps = 41;
+}
+def : InstRW<[Zn4WriteVPGATHERDDZ256, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERDDZ256rm, VGATHERDPSZ256rm)>;
+def Zn4WriteVPGATHERQDZ : SchedWriteRes<[Zn4AGU012, Zn4Load, Zn4FPLd01, Zn4FPU0123]> {
+  let ReleaseAtCycles = [8, 8, 8, 39]; let Latency = 24; let NumMicroOps = 46;
+}
+def : InstRW<[Zn4WriteVPGATHERQDZ, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQDZrm, VGATHERQPSZrm)>;
+def Zn4WriteVPGATHERDDZ : SchedWriteRes<[Zn4AGU012, Zn4Load, Zn4FPLd01, Zn4FPU0123]> {
+  let ReleaseAtCycles = [16, 16, 16, 70]; let Latency = 32; let NumMicroOps = 81;
+}
+def : InstRW<[Zn4WriteVPGATHERDDZ, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERDDZrm, VGATHERDPSZrm)>;
+def Zn4WriteVPGATHERQQZ128 : SchedWriteRes<[Zn4AGU012, Zn4Load, Zn4FPLd01, Zn4FPU0123]> {
+  let ReleaseAtCycles = [2, 2, 2, 16]; let Latency = 12; let NumMicroOps = 17;
+}
+def : InstRW<[Zn4WriteVPGATHERQQZ128, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQQZ128rm, VGATHERQPDZ128rm,
+                     VPGATHERDQZ128rm, VGATHERDPDZ128rm)>;
+def Zn4WriteVPGATHERQQZ256 : SchedWriteRes<[Zn4AGU012, Zn4Load, Zn4FPLd01, Zn4FPU0123]> {
+  let ReleaseAtCycles = [4, 4, 4, 20]; let Latency = 16; let NumMicroOps = 24;
+}
+def : InstRW<[Zn4WriteVPGATHERQQZ256, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQQZ256rm, VGATHERQPDZ256rm,
+                     VPGATHERDQZ256rm, VGATHERDPDZ256rm)>;
+def Zn4WriteVPGATHERQQZ : SchedWriteRes<[Zn4AGU012, Zn4Load, Zn4FPLd01, Zn4FPU0123]> {
+  let ReleaseAtCycles = [8, 8, 8, 39]; let Latency = 25; let NumMicroOps = 48;
+}
+def : InstRW<[Zn4WriteVPGATHERQQZ, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQQZrm, VGATHERQPDZrm)>;
+def Zn4WriteVPGATHERDQZ : SchedWriteRes<[Zn4AGU012, Zn4Load, Zn4FPLd01, Zn4FPU0123]> {
+  let ReleaseAtCycles = [8, 8, 8, 38]; let Latency = 22; let NumMicroOps = 46;
+}
+def : InstRW<[Zn4WriteVPGATHERDQZ, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERDQZrm, VGATHERDPDZrm)>;
+
+// AVX2 (VEX) gathers. AVX2 has no scatter and no VEX form exceeds 8 x i32. All
+// but 8 x i32 measure like the EVEX entry for their shape and reuse it; 8 x i32
+// is genuinely cheaper (8.35 vs 9.19 cyc) and gets its own entry. An InstRW
+// replaces the whole Sched<> list, so each one must restate the mask-writeback
+// def that X86InstrSSE.td attaches; it is the EVEX one, an AVX-512-latency
+// write, and a VEX-specific one is future work.
+def : InstRW<[Zn4WriteVPGATHERQDZ128, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQDrm, VGATHERQPSrm)>;
+def : InstRW<[Zn4WriteVPGATHERDDZ128, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERDDrm, VGATHERDPSrm,
+                     VPGATHERQDYrm, VGATHERQPSYrm)>;
+def Zn4WriteVPGATHERDDY : SchedWriteRes<[Zn4AGU012, Zn4Load, Zn4FPLd01, Zn4FPU0123]> {
+  let ReleaseAtCycles = [8, 8, 8, 33]; let Latency = 21; let NumMicroOps = 42;
+}
+def : InstRW<[Zn4WriteVPGATHERDDY, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERDDYrm, VGATHERDPSYrm)>;
+def : InstRW<[Zn4WriteVPGATHERQQZ128, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQQrm, VGATHERQPDrm,
+                     VPGATHERDQrm, VGATHERDPDrm)>;
+def : InstRW<[Zn4WriteVPGATHERQQZ256, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQQYrm, VGATHERQPDYrm,
+                     VPGATHERDQYrm, VGATHERDPDYrm)>;
+
+def Zn4WriteVPSCATTERQDZ128 : SchedWriteRes<[Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FPU0123]> {
+  let ReleaseAtCycles = [2, 2, 2, 16]; let Latency = 6; let NumMicroOps = 17;
+}
+def : InstRW<[Zn4WriteVPSCATTERQDZ128],
+             (instrs VPSCATTERQDZ128mr, VSCATTERQPSZ128mr)>;
+def Zn4WriteVPSCATTERDDZ128 : SchedWriteRes<[Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FPU0123]> {
+  let ReleaseAtCycles = [4, 4, 4, 33]; let Latency = 9; let NumMicroOps = 27;
+}
+def : InstRW<[Zn4WriteVPSCATTERDDZ128],
+             (instrs VPSCATTERDDZ128mr, VSCATTERDPSZ128mr)>;
+def Zn4WriteVPSCATTERQDZ256 : SchedWriteRes<[Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FPU0123]> {
+  let ReleaseAtCycles = [4, 4, 4, 26]; let Latency = 9; let NumMicroOps = 27;
+}
+def : InstRW<[Zn4WriteVPSCATTERQDZ256],
+             (instrs VPSCATTERQDZ256mr, VSCATTERQPSZ256mr)>;
+def Zn4WriteVPSCATTERDDZ256 : SchedWriteRes<[Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FPU0123]> {
+  let ReleaseAtCycles = [8, 8, 8, 49]; let Latency = 14; let NumMicroOps = 48;
+}
+def : InstRW<[Zn4WriteVPSCATTERDDZ256],
+             (instrs VPSCATTERDDZ256mr, VSCATTERDPSZ256mr)>;
+def Zn4WriteVPSCATTERQDZ : SchedWriteRes<[Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FPU0123]> {
+  let ReleaseAtCycles = [8, 8, 8, 46]; let Latency = 14; let NumMicroOps = 48;
+}
+def : InstRW<[Zn4WriteVPSCATTERQDZ],
+             (instrs VPSCATTERQDZmr, VSCATTERQPSZmr)>;
+def Zn4WriteVPSCATTERDDZ : SchedWriteRes<[Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FPU0123]> {
+  let ReleaseAtCycles = [16, 16, 16, 103]; let Latency = 24; let NumMicroOps = 89;
+}
+def : InstRW<[Zn4WriteVPSCATTERDDZ],
+             (instrs VPSCATTERDDZmr, VSCATTERDPSZmr)>;
+def Zn4WriteVPSCATTERQQZ128 : SchedWriteRes<[Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FPU0123]> {
+  let ReleaseAtCycles = [2, 2, 2, 16]; let Latency = 6; let NumMicroOps = 17;
+}
+def : InstRW<[Zn4WriteVPSCATTERQQZ128],
+             (instrs VPSCATTERQQZ128mr, VSCATTERQPDZ128mr,
+                     VPSCATTERDQZ128mr, VSCATTERDPDZ128mr)>;
+def Zn4WriteVPSCATTERQQZ256 : SchedWriteRes<[Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FPU0123]> {
+  let ReleaseAtCycles = [4, 4, 4, 26]; let Latency = 9; let NumMicroOps = 27;
+}
+def : InstRW<[Zn4WriteVPSCATTERQQZ256],
+             (instrs VPSCATTERQQZ256mr, VSCATTERQPDZ256mr)>;
+def Zn4WriteVPSCATTERDQZ256 : SchedWriteRes<[Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FPU0123]> {
+  let ReleaseAtCycles = [4, 4, 4, 24]; let Latency = 9; let NumMicroOps = 28;
+}
+def : InstRW<[Zn4WriteVPSCATTERDQZ256],
+             (instrs VPSCATTERDQZ256mr, VSCATTERDPDZ256mr)>;
+def Zn4WriteVPSCATTERQQZ : SchedWriteRes<[Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FPU0123]> {
+  let ReleaseAtCycles = [8, 8, 8, 56]; let Latency = 14; let NumMicroOps = 48;
+}
+def : InstRW<[Zn4WriteVPSCATTERQQZ],
+             (instrs VPSCATTERQQZmr, VSCATTERQPDZmr)>;
+def Zn4WriteVPSCATTERDQZ : SchedWriteRes<[Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FPU0123]> {
+  let ReleaseAtCycles = [8, 8, 8, 44]; let Latency = 14; let NumMicroOps = 48;
+}
+def : InstRW<[Zn4WriteVPSCATTERDQZ],
+             (instrs VPSCATTERDQZmr, VSCATTERDPDZmr)>;
+
 def Zn4WriteMOVSlow : SchedWriteRes<[Zn4AGU012, Zn4Load]> {
   let Latency = !add(Znver4Model.LoadLatency, 1);
   let ReleaseAtCycles = [3, 1];
diff --git a/llvm/test/CodeGen/X86/znver4-gather-scatter-schedule.ll b/llvm/test/CodeGen/X86/znver4-gather-scatter-schedule.ll
new file mode 100644
index 0000000000000..8a3f8d2ba15dc
--- /dev/null
+++ b/llvm/test/CodeGen/X86/znver4-gather-scatter-schedule.ll
@@ -0,0 +1,36 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py
+; RUN: llc < %s -mtriple=x86_64-unknown-unknown -mcpu=znver4 | FileCheck %s
+
+; The znver4 schedule model carries per-shape masked gather/scatter overrides
+; (see X86ScheduleZnver4.td). No other znver-tuned CodeGen test exercises these
+; instructions, so this pins which gather/scatter form znver4/znver5/znver6
+; select. It checks lowering only; the latency, throughput and uop counts of the
+; overrides are pinned by the llvm-mca tests under
+; test/tools/llvm-mca/X86/Znver4.
+
+declare <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr>, i32, <8 x i1>, <8 x i32>)
+declare void @llvm.masked.scatter.v8i32.v8p0(<8 x i32>, <8 x ptr>, i32, <8 x i1>)
+
+define <8 x i32> @gather_v8i32(<8 x ptr> %ptrs, <8 x i1> %mask, <8 x i32> %passthru) {
+; CHECK-LABEL: gather_v8i32:
+; CHECK:       # %bb.0:
+; CHECK-NEXT:    vpsllw $15, %xmm1, %xmm1
+; CHECK-NEXT:    vpmovw2m %xmm1, %k1
+; CHECK-NEXT:    vpgatherqd (,%zmm0), %ymm2 {%k1}
+; CHECK-NEXT:    vmovdqa %ymm2, %ymm0
+; CHECK-NEXT:    retq
+  %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x i32> %passthru)
+  ret <8 x i32> %v
+}
+
+define void @scatter_v8i32(<8 x ptr> %ptrs, <8 x i1> %mask, <8 x i32> %val) {
+; CHECK-LABEL: scatter_v8i32:
+; CHECK:       # %bb.0:
+; CHECK-NEXT:    vpsllw $15, %xmm1, %xmm1
+; CHECK-NEXT:    vpmovw2m %xmm1, %k1
+; CHECK-NEXT:    vpscatterqd %ymm2, (,%zmm0) {%k1}
+; CHECK-NEXT:    vzeroupper
+; CHECK-NEXT:    retq
+  call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> %val, <8 x ptr> %ptrs, i32 4, <8 x i1> %mask)
+  ret void
+}
diff --git a/llvm/test/tools/llvm-mca/X86/Znver4/resources-avx2.s b/llvm/test/tools/llvm-mca/X86/Znver4/resources-avx2.s
index 6c8fac4566498..69f8a4961dc92 100644
--- a/llvm/test/tools/llvm-mca/X86/Znver4/resources-avx2.s
+++ b/llvm/test/tools/llvm-mca/X86/Znver4/resources-avx2.s
@@ -465,14 +465,14 @@ vpxor           (%rax), %ymm1, %ymm2
 # CHECK-NEXT:  1      2     1.00                        vbroadcastss	%xmm0, %ymm0
 # CHECK-NEXT:  1      4     1.00                        vextracti128	$1, %ymm0, %xmm2
 # CHECK-NEXT:  2      11    1.00           *            vextracti128	$1, %ymm0, (%rax)
-# CHECK-NEXT:  1      5     0.33    *                   vgatherdpd	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT:  1      5     0.33    *                   vgatherdpd	%ymm0, (%rax,%xmm1,2), %ymm2
-# CHECK-NEXT:  1      5     0.33    *                   vgatherdps	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT:  1      5     0.33    *                   vgatherdps	%ymm0, (%rax,%ymm1,2), %ymm2
-# CHECK-NEXT:  1      5     0.33    *                   vgatherqpd	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT:  1      5     0.33    *                   vgatherqpd	%ymm0, (%rax,%ymm1,2), %ymm2
-# CHECK-NEXT:  1      5     0.33    *                   vgatherqps	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT:  1      5     0.33    *                   vgatherqps	%xmm0, (%rax,%ymm1,2), %xmm2
+# CHECK-NEXT:  17     12    4.00    *                   vgatherdpd	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT:  24     16    5.00    *                   vgatherdpd	%ymm0, (%rax,%xmm1,2), %ymm2
+# CHECK-NEXT:  24     15    5.00    *                   vgatherdps	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT:  42     21    8.25    *                   vgatherdps	%ymm0, (%rax,%ymm1,2), %ymm2
+# CHECK-NEXT:  17     12    4.00    *                   vgatherqpd	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT:  24     16    5.00    *                   vgatherqpd	%ymm0, (%rax,%ymm1,2), %ymm2
+# CHECK-NEXT:  18     13    4.00    *                   vgatherqps	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT:  24     15    5.00    *                   vgatherqps	%xmm0, (%rax,%ymm1,2), %xmm2
 # CHECK-NEXT:  1      1     1.00                        vinserti128	$1, %xmm0, %ymm1, %ymm2
 # CHECK-NEXT:  1      8     1.00    *                   vinserti128	$1, (%rax), %ymm1, %ymm2
 # CHECK-NEXT:  1      8     0.50    *                   vmovntdqa	(%rax), %ymm0
@@ -568,14 +568,14 @@ vpxor           (%rax), %ymm1, %ymm2
 # CHECK-NEXT:  1      11    1.00    *                   vpermps	(%rax), %ymm1, %ymm2
 # CHECK-NEXT:  1      4     1.00                        vpermq	$1, %ymm0, %ymm2
 # CHECK-NEXT:  1      11    1.00    *                   vpermq	$1, (%rax), %ymm2
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherdd	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherdd	%ymm0, (%rax,%ymm1,2), %ymm2
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherdq	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherdq	%ymm0, (%rax,%xmm1,2), %ymm2
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherqd	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherqd	%xmm0, (%rax,%ymm1,2), %xmm2
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherqq	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherqq	%ymm0, (%rax,%ymm1,2), %ymm2
+# CHECK-NEXT:  24     15    5.00    *                   vpgatherdd	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT:  42     21    8.25    *                   vpgatherdd	%ymm0, (%rax,%ymm1,2), %ymm2
+# CHECK-NEXT:  17     12    4.00    *                   vpgatherdq	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT:  24     16    5.00    *                   vpgatherdq	%ymm0, (%rax,%xmm1,2), %ymm2
+# CHECK-NEXT:  18     13    4.00    *                   vpgatherqd	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT:  24     15    5.00    *                   vpgatherqd	%xmm0, (%rax,%ymm1,2), %xmm2
+# CHECK-NEXT:  17     12    4.00    *                   vpgatherqq	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT:  24     16    5.00    *                   vpgatherqq	%ymm0, (%rax,%ymm1,2), %ymm2
 # CHECK-NEXT:  3      3     3.00                        vphaddd	%ymm0, %ymm1, %ymm2
 # CHECK-NEXT:  4      10    3.00    *                   vphaddd	(%rax), %ymm1, %ymm2
 # CHECK-NEXT:  3      3     3.00                        vphaddsw	%ymm0, %ymm1, %ymm2
@@ -789,7 +789,7 @@ vpxor           (%rax), %ymm1, %ymm2
 
 # CHECK:      Resource pressure per iteration:
 # CHECK-NEXT: [0]    [1]    [2]    [3]    [4]    [5]    [6]    [7]    [8]    [9]    [10]   [11]   [12.0] [12.1] [13]   [14.0] [14.1] [14.2] [15.0] [15.1] [15.2] [16.0] [16.1]
-# CHECK-NEXT: 6.67   6.67   6.67    -      -      -      -      -     93.75  128.75 92.25  36.25  80.50  80.50  29.00  52.33  52.33  52.33  50.67  50.67  50.67  2.50   2.50
+# CHECK-NEXT: 21.33  21.33  21.33   -      -      -      -      -     174.25 209.25 172.75 116.75 110.50 110.50 29.00  67.00  67.00  67.00  65.33  65.33  65.33  2.50   2.50
 
 # CHECK:      Resource pressure by instruction:
 # CHECK-NEXT: [0]    [1]    [2]    [3]    [4]    [5]    [6]    [7]    [8]    [9]    [10]   [11]   [12.0] [12.1] [13]   [14.0] [14.1] [14.2] [15.0] [15.1] [15.2] [16.0] [16.1] Instructions:
@@ -798,14 +798,14 @@ vpxor           (%rax), %ymm1, %ymm2
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00    -      -      -      -      -      -      -      -      -      -      -      -      -     vbroadcastss	%xmm0, %ymm0
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     1.00    -      -      -      -      -      -      -      -      -      -      -      -      -      -     vextracti128	$1, %ymm0, %xmm2
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     1.00    -      -      -     0.50   0.50   1.00   0.33   0.33   0.33    -      -      -     0.50   0.50   vextracti128	$1, %ymm0, (%rax)
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherdpd	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherdpd	%ymm0, (%rax,%xmm1,2), %ymm2
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherdps	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherdps	%ymm0, (%rax,%ymm1,2), %ymm2
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherqpd	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherqpd	%ymm0, (%rax,%ymm1,2), %ymm2
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherqps	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherqps	%xmm0, (%rax,%ymm1,2), %xmm2
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00    -     0.67   0.67   0.67   0.67   0.67   0.67    -      -     vgatherdpd	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vgatherdpd	%ymm0, (%rax,%xmm1,2), %ymm2
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vgatherdps	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     8.25   8.25   8.25   8.25   4.00   4.00    -     2.67   2.67   2.67   2.67   2.67   2.67    -      -     vgatherdps	%ymm0, (%rax,%ymm1,2), %ymm2
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00    -     0.67   0.67   0.67   0.67   0.67   0.67    -      -     vgatherqpd	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vgatherqpd	%ymm0, (%rax,%ymm1,2), %ymm2
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00    -     0.67   0.67   0.67   0.67   0.67   0.67    -      -     vgatherqps	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vgatherqps	%xmm0, (%rax,%ymm1,2), %xmm2
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00    -      -      -      -      -      -      -      -      -      -      -      -      -     vinserti128	$1, %xmm0, %ymm1, %ymm2
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00    -      -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vinserti128	$1, (%rax), %ymm1, %ymm2
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -      -      -      -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vmovntdqa	(%rax), %ymm0
@@ -901,14 +901,14 @@ vpxor           (%rax), %ymm1, %ymm2
 # CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -     1.00    -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpermps	(%rax), %ymm1, %ymm2
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00    -      -      -      -      -      -      -      -      -      -      -      -      -     vpermq	$1, %ymm0, %ymm2
 # CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -     1.00    -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpermq	$1, (%rax), %ymm2
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherdd	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherdd	%ymm0, (%rax,%ymm1,2), %ymm2
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherdq	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherdq	%ymm0, (%rax,%xmm1,2), %ymm2
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherqd	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherqd	%xmm0, (%rax,%ymm1,2), %xmm2
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherqq	%xmm0, (%rax,%xmm1,2), %xmm2
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherqq	%ymm0, (%rax,%ymm1,2), %ymm2
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vpgatherdd	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     8.25   8.25   8.25   8.25   4.00   4.00    -     2.67   2.67   2.67   2.67   2.67   2.67    -      -     vpgatherdd	%ymm0, (%rax,%ymm1,2), %ymm2
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00    -     0.67   0.67   0.67   0.67   0.67   0.67    -      -     vpgatherdq	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vpgatherdq	%ymm0, (%rax,%xmm1,2), %ymm2
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00    -     0.67   0.67   0.67   0.67   0.67   0.67    -      -     vpgatherqd	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vpgatherqd	%xmm0, (%rax,%ymm1,2), %xmm2
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00    -     0.67   0.67   0.67   0.67   0.67   0.67    -      -     vpgatherqq	%xmm0, (%rax,%xmm1,2), %xmm2
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vpgatherqq	%ymm0, (%rax,%ymm1,2), %ymm2
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     3.00    -      -      -      -      -      -      -      -      -      -      -      -      -      -     vphaddd	%ymm0, %ymm1, %ymm2
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     3.00    -      -      -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vphaddd	(%rax), %ymm1, %ymm2
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     3.00    -      -      -      -      -      -      -      -      -      -      -      -      -      -     vphaddsw	%ymm0, %ymm1, %ymm2
diff --git a/llvm/test/tools/llvm-mca/X86/Znver4/resources-avx512.s b/llvm/test/tools/llvm-mca/X86/Znver4/resources-avx512.s
index f1712aca08311..566a1a9e48126 100644
--- a/llvm/test/tools/llvm-mca/X86/Znver4/resources-avx512.s
+++ b/llvm/test/tools/llvm-mca/X86/Znver4/resources-avx512.s
@@ -1495,10 +1495,10 @@ vunpcklps         (%rax){1to16}, %zmm17, %zmm19 {z}{k1}
 # CHECK-NEXT:  1      4     1.00                        vfmadd231ps	%zmm16, %zmm17, %zmm19 {%k1} {z}
 # CHECK-NEXT:  1      11    1.00    *                   vfmadd231ps	(%rax), %zmm17, %zmm19 {%k1} {z}
 # CHECK-NEXT:  1      11    1.00    *                   vfmadd231ps	(%rax){1to16}, %zmm17, %zmm19 {%k1} {z}
-# CHECK-NEXT:  1      5     0.33    *                   vgatherdpd	(%rax,%ymm1,2), %zmm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vgatherdps	(%rax,%zmm1,2), %zmm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vgatherqpd	(%rax,%zmm1,2), %zmm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vgatherqps	(%rax,%zmm1,2), %ymm2 {%k1}
+# CHECK-NEXT:  46     22    9.50    *                   vgatherdpd	(%rax,%ymm1,2), %zmm2 {%k1}
+# CHECK-NEXT:  81     32    17.50   *                   vgatherdps	(%rax,%zmm1,2), %zmm2 {%k1}
+# CHECK-NEXT:  48     25    9.75    *                   vgatherqpd	(%rax,%zmm1,2), %zmm2 {%k1}
+# CHECK-NEXT:  46     24    9.75    *                   vgatherqps	(%rax,%zmm1,2), %ymm2 {%k1}
 # CHECK-NEXT:  1      2     1.00                        vmaxpd	%zmm16, %zmm17, %zmm19
 # CHECK-NEXT:  1      9     1.00    *                   vmaxpd	(%rax), %zmm17, %zmm19
 # CHECK-NEXT:  1      9     1.00    *                   vmaxpd	(%rax){1to8}, %zmm17, %zmm19
@@ -1732,10 +1732,10 @@ vunpcklps         (%rax){1to16}, %zmm17, %zmm19 {z}{k1}
 # CHECK-NEXT:  1      1     0.50                        vpcmpequq	%zmm0, %zmm1, %k2 {%k3}
 # CHECK-NEXT:  1      8     0.50    *                   vpcmpequq	(%rax), %zmm1, %k2 {%k3}
 # CHECK-NEXT:  1      8     0.50    *                   vpcmpequq	(%rax){1to8}, %zmm1, %k2 {%k3}
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherdq	(%rax,%ymm1,2), %zmm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherdd	(%rax,%zmm1,2), %zmm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherqq	(%rax,%zmm1,2), %zmm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherqd	(%rax,%zmm1,2), %ymm2 {%k1}
+# CHECK-NEXT:  46     22    9.50    *                   vpgatherdq	(%rax,%ymm1,2), %zmm2 {%k1}
+# CHECK-NEXT:  81     32    17.50   *                   vpgatherdd	(%rax,%zmm1,2), %zmm2 {%k1}
+# CHECK-NEXT:  48     25    9.75    *                   vpgatherqq	(%rax,%zmm1,2), %zmm2 {%k1}
+# CHECK-NEXT:  46     24    9.75    *                   vpgatherqd	(%rax,%zmm1,2), %ymm2 {%k1}
 # CHECK-NEXT:  1      5     1.00                        vpmovdb	%zmm19, %xmm16
 # CHECK-NEXT:  1      11    1.50           *            vpmovdb	%zmm19, (%rax)
 # CHECK-NEXT:  1      5     1.00                        vpmovdb	%zmm19, %xmm16 {%k1}
@@ -1970,10 +1970,10 @@ vunpcklps         (%rax){1to16}, %zmm17, %zmm19 {z}{k1}
 # CHECK-NEXT:  2      1     0.50                        vpermq	%zmm16, %zmm17, %zmm19 {%k1} {z}
 # CHECK-NEXT:  2      8     0.50    *                   vpermq	(%rax), %zmm17, %zmm19 {%k1} {z}
 # CHECK-NEXT:  2      8     0.50    *                   vpermq	(%rax){1to8}, %zmm17, %zmm19 {%k1} {z}
-# CHECK-NEXT:  1      1     1.00           *            vpscatterdd	%zmm1, (%rdx,%zmm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vpscatterdq	%zmm1, (%rdx,%ymm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vpscatterqd	%ymm1, (%rdx,%zmm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vpscatterqq	%zmm1, (%rdx,%zmm0,4) {%k1}
+# CHECK-NEXT:  89     24    25.75          *            vpscatterdd	%zmm1, (%rdx,%zmm0,4) {%k1}
+# CHECK-NEXT:  48     14    11.00          *            vpscatterdq	%zmm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT:  48     14    11.50          *            vpscatterqd	%ymm1, (%rdx,%zmm0,4) {%k1}
+# CHECK-NEXT:  48     14    14.00          *            vpscatterqq	%zmm1, (%rdx,%zmm0,4) {%k1}
 # CHECK-NEXT:  1      1     1.00                        vpshufd	$0, %zmm16, %zmm19
 # CHECK-NEXT:  1      8     1.00    *                   vpshufd	$0, (%rax), %zmm19
 # CHECK-NEXT:  1      8     1.00    *                   vpshufd	$0, (%rax){1to16}, %zmm19
@@ -2037,10 +2037,10 @@ vunpcklps         (%rax){1to16}, %zmm17, %zmm19 {z}{k1}
 # CHECK-NEXT:  1      1     1.00                        vpunpcklqdq	%zmm16, %zmm17, %zmm19 {%k1} {z}
 # CHECK-NEXT:  1      8     1.00    *                   vpunpcklqdq	(%rax), %zmm17, %zmm19 {%k1} {z}
 # CHECK-NEXT:  1      8     1.00    *                   vpunpcklqdq	(%rax){1to8}, %zmm17, %zmm19 {%k1} {z}
-# CHECK-NEXT:  1      1     1.00           *            vscatterdps	%zmm1, (%rdx,%zmm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vscatterdpd	%zmm1, (%rdx,%ymm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vscatterqps	%ymm1, (%rdx,%zmm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vscatterqpd	%zmm1, (%rdx,%zmm0,4) {%k1}
+# CHECK-NEXT:  89     24    25.75          *            vscatterdps	%zmm1, (%rdx,%zmm0,4) {%k1}
+# CHECK-NEXT:  48     14    11.00          *            vscatterdpd	%zmm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT:  48     14    11.50          *            vscatterqps	%ymm1, (%rdx,%zmm0,4) {%k1}
+# CHECK-NEXT:  48     14    14.00          *            vscatterqpd	%zmm1, (%rdx,%zmm0,4) {%k1}
 # CHECK-NEXT:  1      2     1.00                        vshuff32x4	$0, %zmm16, %zmm17, %zmm19
 # CHECK-NEXT:  3      9     1.00    *                   vshuff32x4	$0, (%rax), %zmm17, %zmm19
 # CHECK-NEXT:  3      9     1.00    *                   vshuff32x4	$0, (%rax){1to16}, %zmm17, %zmm19
@@ -2233,7 +2233,7 @@ vunpcklps         (%rax){1to16}, %zmm17, %zmm19 {z}{k1}
 
 # CHECK:      Resource pressure per iteration:
 # CHECK-NEXT: [0]    [1]    [2]    [3]    [4]    [5]    [6]    [7]    [8]    [9]    [10]   [11]   [12.0] [12.1] [13]   [14.0] [14.1] [14.2] [15.0] [15.1] [15.2] [16.0] [16.1]
-# CHECK-NEXT: 5.33   5.33   5.33    -      -      -      -      -     221.00 1120.50 678.00 352.50 312.50 312.50 17.00 215.67 215.67 215.67 204.67 204.67 204.67 16.50  16.50
+# CHECK-NEXT: 53.33  53.33  53.33   -      -      -      -      -     438.50 1338.00 895.50 570.00 392.50 392.50 97.00 261.00 261.00 261.00 228.67 228.67 228.67 48.50  48.50
 
 # CHECK:      Resource pressure by instruction:
 # CHECK-NEXT: [0]    [1]    [2]    [3]    [4]    [5]    [6]    [7]    [8]    [9]    [10]   [11]   [12.0] [12.1] [13]   [14.0] [14.1] [14.2] [15.0] [15.1] [15.2] [16.0] [16.1] Instructions:
@@ -2552,10 +2552,10 @@ vunpcklps         (%rax){1to16}, %zmm17, %zmm19 {z}{k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     1.00   1.00    -      -      -      -      -      -      -      -      -      -      -      -      -     vfmadd231ps	%zmm16, %zmm17, %zmm19 {%k1} {z}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     1.00   1.00    -      -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vfmadd231ps	(%rax), %zmm17, %zmm19 {%k1} {z}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     1.00   1.00    -      -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vfmadd231ps	(%rax){1to16}, %zmm17, %zmm19 {%k1} {z}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherdpd	(%rax,%ymm1,2), %zmm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherdps	(%rax,%zmm1,2), %zmm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherqpd	(%rax,%zmm1,2), %zmm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherqps	(%rax,%zmm1,2), %ymm2 {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     9.50   9.50   9.50   9.50   4.00   4.00    -     2.67   2.67   2.67   2.67   2.67   2.67    -      -     vgatherdpd	(%rax,%ymm1,2), %zmm2 {%k1}
+# CHECK-NEXT: 5.33   5.33   5.33    -      -      -      -      -     17.50  17.50  17.50  17.50  8.00   8.00    -     5.33   5.33   5.33   5.33   5.33   5.33    -      -     vgatherdps	(%rax,%zmm1,2), %zmm2 {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     9.75   9.75   9.75   9.75   4.00   4.00    -     2.67   2.67   2.67   2.67   2.67   2.67    -      -     vgatherqpd	(%rax,%zmm1,2), %zmm2 {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     9.75   9.75   9.75   9.75   4.00   4.00    -     2.67   2.67   2.67   2.67   2.67   2.67    -      -     vgatherqps	(%rax,%zmm1,2), %ymm2 {%k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     1.00   1.00    -      -      -      -      -      -      -      -      -      -      -      -      -     vmaxpd	%zmm16, %zmm17, %zmm19
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     1.00   1.00    -      -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vmaxpd	(%rax), %zmm17, %zmm19
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     1.00   1.00    -      -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vmaxpd	(%rax){1to8}, %zmm17, %zmm19
@@ -2789,10 +2789,10 @@ vunpcklps         (%rax){1to16}, %zmm17, %zmm19 {z}{k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     0.50   0.50   0.50   0.50    -      -      -      -      -      -      -      -      -      -      -     vpcmpequq	%zmm0, %zmm1, %k2 {%k3}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     0.50   0.50   0.50   0.50   0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpcmpequq	(%rax), %zmm1, %k2 {%k3}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     0.50   0.50   0.50   0.50   0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpcmpequq	(%rax){1to8}, %zmm1, %k2 {%k3}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherdq	(%rax,%ymm1,2), %zmm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherdd	(%rax,%zmm1,2), %zmm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherqq	(%rax,%zmm1,2), %zmm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherqd	(%rax,%zmm1,2), %ymm2 {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     9.50   9.50   9.50   9.50   4.00   4.00    -     2.67   2.67   2.67   2.67   2.67   2.67    -      -     vpgatherdq	(%rax,%ymm1,2), %zmm2 {%k1}
+# CHECK-NEXT: 5.33   5.33   5.33    -      -      -      -      -     17.50  17.50  17.50  17.50  8.00   8.00    -     5.33   5.33   5.33   5.33   5.33   5.33    -      -     vpgatherdd	(%rax,%zmm1,2), %zmm2 {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     9.75   9.75   9.75   9.75   4.00   4.00    -     2.67   2.67   2.67   2.67   2.67   2.67    -      -     vpgatherqq	(%rax,%zmm1,2), %zmm2 {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     9.75   9.75   9.75   9.75   4.00   4.00    -     2.67   2.67   2.67   2.67   2.67   2.67    -      -     vpgatherqd	(%rax,%zmm1,2), %ymm2 {%k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00   1.00    -      -      -      -      -      -      -      -      -      -      -      -     vpmovdb	%zmm19, %xmm16
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.50   1.50    -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpmovdb	%zmm19, (%rax)
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00   1.00    -      -      -      -      -      -      -      -      -      -      -      -     vpmovdb	%zmm19, %xmm16 {%k1}
@@ -3027,10 +3027,10 @@ vunpcklps         (%rax){1to16}, %zmm17, %zmm19 {z}{k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     0.50   0.50    -      -      -      -      -      -      -      -      -      -      -      -     vpermq	%zmm16, %zmm17, %zmm19 {%k1} {z}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     0.50   0.50    -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpermq	(%rax), %zmm17, %zmm19 {%k1} {z}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     0.50   0.50    -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpermq	(%rax){1to8}, %zmm17, %zmm19 {%k1} {z}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vpscatterdd	%zmm1, (%rdx,%zmm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vpscatterdq	%zmm1, (%rdx,%ymm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vpscatterqd	%ymm1, (%rdx,%zmm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vpscatterqq	%zmm1, (%rdx,%zmm0,4) {%k1}
+# CHECK-NEXT: 5.33   5.33   5.33    -      -      -      -      -     25.75  25.75  25.75  25.75  8.00   8.00   16.00  5.33   5.33   5.33    -      -      -     8.00   8.00   vpscatterdd	%zmm1, (%rdx,%zmm0,4) {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     11.00  11.00  11.00  11.00  4.00   4.00   8.00   2.67   2.67   2.67    -      -      -     4.00   4.00   vpscatterdq	%zmm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     11.50  11.50  11.50  11.50  4.00   4.00   8.00   2.67   2.67   2.67    -      -      -     4.00   4.00   vpscatterqd	%ymm1, (%rdx,%zmm0,4) {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     14.00  14.00  14.00  14.00  4.00   4.00   8.00   2.67   2.67   2.67    -      -      -     4.00   4.00   vpscatterqq	%zmm1, (%rdx,%zmm0,4) {%k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00   1.00    -      -      -      -      -      -      -      -      -      -      -      -     vpshufd	$0, %zmm16, %zmm19
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00   1.00    -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpshufd	$0, (%rax), %zmm19
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00   1.00    -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpshufd	$0, (%rax){1to16}, %zmm19
@@ -3094,10 +3094,10 @@ vunpcklps         (%rax){1to16}, %zmm17, %zmm19 {z}{k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00   1.00    -      -      -      -      -      -      -      -      -      -      -      -     vpunpcklqdq	%zmm16, %zmm17, %zmm19 {%k1} {z}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00   1.00    -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpunpcklqdq	(%rax), %zmm17, %zmm19 {%k1} {z}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00   1.00    -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpunpcklqdq	(%rax){1to8}, %zmm17, %zmm19 {%k1} {z}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vscatterdps	%zmm1, (%rdx,%zmm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vscatterdpd	%zmm1, (%rdx,%ymm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vscatterqps	%ymm1, (%rdx,%zmm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vscatterqpd	%zmm1, (%rdx,%zmm0,4) {%k1}
+# CHECK-NEXT: 5.33   5.33   5.33    -      -      -      -      -     25.75  25.75  25.75  25.75  8.00   8.00   16.00  5.33   5.33   5.33    -      -      -     8.00   8.00   vscatterdps	%zmm1, (%rdx,%zmm0,4) {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     11.00  11.00  11.00  11.00  4.00   4.00   8.00   2.67   2.67   2.67    -      -      -     4.00   4.00   vscatterdpd	%zmm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     11.50  11.50  11.50  11.50  4.00   4.00   8.00   2.67   2.67   2.67    -      -      -     4.00   4.00   vscatterqps	%ymm1, (%rdx,%zmm0,4) {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     14.00  14.00  14.00  14.00  4.00   4.00   8.00   2.67   2.67   2.67    -      -      -     4.00   4.00   vscatterqpd	%zmm1, (%rdx,%zmm0,4) {%k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00    -      -      -      -      -      -      -      -      -      -      -      -      -     vshuff32x4	$0, %zmm16, %zmm17, %zmm19
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00    -      -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vshuff32x4	$0, (%rax), %zmm17, %zmm19
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00    -      -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vshuff32x4	$0, (%rax){1to16}, %zmm17, %zmm19
diff --git a/llvm/test/tools/llvm-mca/X86/Znver4/resources-avx512vl.s b/llvm/test/tools/llvm-mca/X86/Znver4/resources-avx512vl.s
index ead609e33da4d..68623e7e27ecc 100644
--- a/llvm/test/tools/llvm-mca/X86/Znver4/resources-avx512vl.s
+++ b/llvm/test/tools/llvm-mca/X86/Znver4/resources-avx512vl.s
@@ -2392,14 +2392,14 @@ vunpcklps         (%rax){1to8}, %ymm17, %ymm19 {z}{k1}
 # CHECK-NEXT:  1      4     0.50                        vfmadd231ps	%ymm16, %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  1      11    0.50    *                   vfmadd231ps	(%rax), %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  1      11    0.50    *                   vfmadd231ps	(%rax){1to8}, %ymm17, %ymm19 {%k1} {z}
-# CHECK-NEXT:  1      5     0.33    *                   vgatherdpd	(%rax,%xmm1,2), %ymm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vgatherdps	(%rax,%ymm1,2), %ymm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vgatherqpd	(%rax,%ymm1,2), %ymm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vgatherqps	(%rax,%ymm1,2), %xmm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vgatherdpd	(%rax,%xmm1,2), %xmm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vgatherdps	(%rax,%xmm1,2), %xmm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vgatherqpd	(%rax,%xmm1,2), %xmm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vgatherqps	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT:  24     16    5.00    *                   vgatherdpd	(%rax,%xmm1,2), %ymm2 {%k1}
+# CHECK-NEXT:  41     20    9.25    *                   vgatherdps	(%rax,%ymm1,2), %ymm2 {%k1}
+# CHECK-NEXT:  24     16    5.00    *                   vgatherqpd	(%rax,%ymm1,2), %ymm2 {%k1}
+# CHECK-NEXT:  24     15    5.00    *                   vgatherqps	(%rax,%ymm1,2), %xmm2 {%k1}
+# CHECK-NEXT:  17     12    4.00    *                   vgatherdpd	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT:  24     15    5.00    *                   vgatherdps	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT:  17     12    4.00    *                   vgatherqpd	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT:  18     13    4.00    *                   vgatherqps	(%rax,%xmm1,2), %xmm2 {%k1}
 # CHECK-NEXT:  1      2     1.00                        vmaxpd	%xmm16, %xmm17, %xmm19
 # CHECK-NEXT:  1      9     0.50    *                   vmaxpd	(%rax), %xmm17, %xmm19
 # CHECK-NEXT:  1      9     0.50    *                   vmaxpd	(%rax){1to2}, %xmm17, %xmm19
@@ -2956,14 +2956,14 @@ vunpcklps         (%rax){1to8}, %ymm17, %ymm19 {z}{k1}
 # CHECK-NEXT:  1      4     1.00                        vpermq	%ymm16, %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  2      8     0.50    *                   vpermq	(%rax), %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  2      8     0.50    *                   vpermq	(%rax){1to4}, %ymm17, %ymm19 {%k1} {z}
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherdq	(%rax,%xmm1,2), %ymm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherdd	(%rax,%ymm1,2), %ymm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherqq	(%rax,%ymm1,2), %ymm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherqd	(%rax,%ymm1,2), %xmm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherdq	(%rax,%xmm1,2), %xmm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherdd	(%rax,%xmm1,2), %xmm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherqq	(%rax,%xmm1,2), %xmm2 {%k1}
-# CHECK-NEXT:  1      5     0.33    *                   vpgatherqd	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT:  24     16    5.00    *                   vpgatherdq	(%rax,%xmm1,2), %ymm2 {%k1}
+# CHECK-NEXT:  41     20    9.25    *                   vpgatherdd	(%rax,%ymm1,2), %ymm2 {%k1}
+# CHECK-NEXT:  24     16    5.00    *                   vpgatherqq	(%rax,%ymm1,2), %ymm2 {%k1}
+# CHECK-NEXT:  24     15    5.00    *                   vpgatherqd	(%rax,%ymm1,2), %xmm2 {%k1}
+# CHECK-NEXT:  17     12    4.00    *                   vpgatherdq	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT:  24     15    5.00    *                   vpgatherdd	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT:  17     12    4.00    *                   vpgatherqq	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT:  18     13    4.00    *                   vpgatherqd	(%rax,%xmm1,2), %xmm2 {%k1}
 # CHECK-NEXT:  1      2     0.50                        vpmovdb	%xmm19, %xmm16
 # CHECK-NEXT:  1      11    1.50           *            vpmovdb	%xmm19, (%rax)
 # CHECK-NEXT:  1      2     0.50                        vpmovdb	%xmm19, %xmm16 {%k1}
@@ -3252,14 +3252,14 @@ vunpcklps         (%rax){1to8}, %ymm17, %ymm19 {z}{k1}
 # CHECK-NEXT:  1      3     0.50                        vpmulld	%ymm16, %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  1      10    0.50    *                   vpmulld	(%rax), %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  1      10    0.50    *                   vpmulld	(%rax){1to8}, %ymm17, %ymm19 {%k1} {z}
-# CHECK-NEXT:  1      1     1.00           *            vpscatterdd	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vpscatterdq	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vpscatterqd	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vpscatterqq	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vpscatterdd	%ymm1, (%rdx,%ymm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vpscatterdq	%ymm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vpscatterqd	%xmm1, (%rdx,%ymm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vpscatterqq	%ymm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT:  27     9     8.25           *            vpscatterdd	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT:  17     6     4.00           *            vpscatterdq	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT:  17     6     4.00           *            vpscatterqd	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT:  17     6     4.00           *            vpscatterqq	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT:  48     14    12.25          *            vpscatterdd	%ymm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT:  28     9     6.00           *            vpscatterdq	%ymm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT:  27     9     6.50           *            vpscatterqd	%xmm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT:  27     9     6.50           *            vpscatterqq	%ymm1, (%rdx,%ymm0,4) {%k1}
 # CHECK-NEXT:  1      1     0.50                        vpshufd	$0, %xmm16, %xmm19
 # CHECK-NEXT:  1      8     0.50    *                   vpshufd	$0, (%rax), %xmm19
 # CHECK-NEXT:  1      8     0.50    *                   vpshufd	$0, (%rax){1to4}, %xmm19
@@ -3398,14 +3398,14 @@ vunpcklps         (%rax){1to8}, %ymm17, %ymm19 {z}{k1}
 # CHECK-NEXT:  1      1     0.50                        vpunpckldq	%ymm16, %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  1      8     0.50    *                   vpunpckldq	(%rax), %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  1      8     0.50    *                   vpunpckldq	(%rax){1to8}, %ymm17, %ymm19 {%k1} {z}
-# CHECK-NEXT:  1      1     1.00           *            vscatterdps	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vscatterdpd	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vscatterqps	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vscatterqpd	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vscatterdps	%ymm1, (%rdx,%ymm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vscatterdpd	%ymm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vscatterqps	%xmm1, (%rdx,%ymm0,4) {%k1}
-# CHECK-NEXT:  1      1     1.00           *            vscatterqpd	%ymm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT:  27     9     8.25           *            vscatterdps	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT:  17     6     4.00           *            vscatterdpd	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT:  17     6     4.00           *            vscatterqps	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT:  17     6     4.00           *            vscatterqpd	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT:  48     14    12.25          *            vscatterdps	%ymm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT:  28     9     6.00           *            vscatterdpd	%ymm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT:  27     9     6.50           *            vscatterqps	%xmm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT:  27     9     6.50           *            vscatterqpd	%ymm1, (%rdx,%ymm0,4) {%k1}
 # CHECK-NEXT:  1      2     1.00                        vshuff32x4	$0, %ymm16, %ymm17, %ymm19
 # CHECK-NEXT:  3      9     1.00    *                   vshuff32x4	$0, (%rax), %ymm17, %ymm19
 # CHECK-NEXT:  3      9     1.00    *                   vshuff32x4	$0, (%rax){1to8}, %ymm17, %ymm19
@@ -3614,7 +3614,7 @@ vunpcklps         (%rax){1to8}, %ymm17, %ymm19 {z}{k1}
 
 # CHECK:      Resource pressure per iteration:
 # CHECK-NEXT: [0]    [1]    [2]    [3]    [4]    [5]    [6]    [7]    [8]    [9]    [10]   [11]   [12.0] [12.1] [13]   [14.0] [14.1] [14.2] [15.0] [15.1] [15.2] [16.0] [16.1]
-# CHECK-NEXT: 10.67  10.67  10.67   -      -      -      -      -     208.00 1084.00 637.50 261.50 509.50 509.50 32.00 355.67 355.67 355.67 334.33 334.33 334.33 32.00  32.00
+# CHECK-NEXT: 40.00  40.00  40.00   -      -      -      -      -     393.50 1269.50 823.00 447.00 569.50 569.50 92.00 379.67 379.67 379.67 349.00 349.00 349.00 46.00  46.00
 
 # CHECK:      Resource pressure by instruction:
 # CHECK-NEXT: [0]    [1]    [2]    [3]    [4]    [5]    [6]    [7]    [8]    [9]    [10]   [11]   [12.0] [12.1] [13]   [14.0] [14.1] [14.2] [15.0] [15.1] [15.2] [16.0] [16.1] Instructions:
@@ -4098,14 +4098,14 @@ vunpcklps         (%rax){1to8}, %ymm17, %ymm19 {z}{k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     0.50   0.50    -      -      -      -      -      -      -      -      -      -      -      -      -     vfmadd231ps	%ymm16, %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     0.50   0.50    -      -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vfmadd231ps	(%rax), %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     0.50   0.50    -      -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vfmadd231ps	(%rax){1to8}, %ymm17, %ymm19 {%k1} {z}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherdpd	(%rax,%xmm1,2), %ymm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherdps	(%rax,%ymm1,2), %ymm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherqpd	(%rax,%ymm1,2), %ymm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherqps	(%rax,%ymm1,2), %xmm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherdpd	(%rax,%xmm1,2), %xmm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherdps	(%rax,%xmm1,2), %xmm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherqpd	(%rax,%xmm1,2), %xmm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vgatherqps	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vgatherdpd	(%rax,%xmm1,2), %ymm2 {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     9.25   9.25   9.25   9.25   4.00   4.00    -     2.67   2.67   2.67   2.67   2.67   2.67    -      -     vgatherdps	(%rax,%ymm1,2), %ymm2 {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vgatherqpd	(%rax,%ymm1,2), %ymm2 {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vgatherqps	(%rax,%ymm1,2), %xmm2 {%k1}
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00    -     0.67   0.67   0.67   0.67   0.67   0.67    -      -     vgatherdpd	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vgatherdps	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00    -     0.67   0.67   0.67   0.67   0.67   0.67    -      -     vgatherqpd	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00    -     0.67   0.67   0.67   0.67   0.67   0.67    -      -     vgatherqps	(%rax,%xmm1,2), %xmm2 {%k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     1.00   1.00    -      -      -      -      -      -      -      -      -      -      -      -      -     vmaxpd	%xmm16, %xmm17, %xmm19
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     0.50   0.50    -      -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vmaxpd	(%rax), %xmm17, %xmm19
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     0.50   0.50    -      -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vmaxpd	(%rax){1to2}, %xmm17, %xmm19
@@ -4662,14 +4662,14 @@ vunpcklps         (%rax){1to8}, %ymm17, %ymm19 {z}{k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00   1.00    -      -      -      -      -      -      -      -      -      -      -      -     vpermq	%ymm16, %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     0.50   0.50    -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpermq	(%rax), %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     0.50   0.50    -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpermq	(%rax){1to4}, %ymm17, %ymm19 {%k1} {z}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherdq	(%rax,%xmm1,2), %ymm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherdd	(%rax,%ymm1,2), %ymm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherqq	(%rax,%ymm1,2), %ymm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherqd	(%rax,%ymm1,2), %xmm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherdq	(%rax,%xmm1,2), %xmm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherdd	(%rax,%xmm1,2), %xmm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherqq	(%rax,%xmm1,2), %xmm2 {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpgatherqd	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vpgatherdq	(%rax,%xmm1,2), %ymm2 {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     9.25   9.25   9.25   9.25   4.00   4.00    -     2.67   2.67   2.67   2.67   2.67   2.67    -      -     vpgatherdd	(%rax,%ymm1,2), %ymm2 {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vpgatherqq	(%rax,%ymm1,2), %ymm2 {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vpgatherqd	(%rax,%ymm1,2), %xmm2 {%k1}
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00    -     0.67   0.67   0.67   0.67   0.67   0.67    -      -     vpgatherdq	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     5.00   5.00   5.00   5.00   2.00   2.00    -     1.33   1.33   1.33   1.33   1.33   1.33    -      -     vpgatherdd	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00    -     0.67   0.67   0.67   0.67   0.67   0.67    -      -     vpgatherqq	(%rax,%xmm1,2), %xmm2 {%k1}
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00    -     0.67   0.67   0.67   0.67   0.67   0.67    -      -     vpgatherqd	(%rax,%xmm1,2), %xmm2 {%k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     0.50   0.50    -      -      -      -      -      -      -      -      -      -      -      -     vpmovdb	%xmm19, %xmm16
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.50   1.50    -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpmovdb	%xmm19, (%rax)
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     0.50   0.50    -      -      -      -      -      -      -      -      -      -      -      -     vpmovdb	%xmm19, %xmm16 {%k1}
@@ -4958,14 +4958,14 @@ vunpcklps         (%rax){1to8}, %ymm17, %ymm19 {z}{k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     0.50    -      -     0.50    -      -      -      -      -      -      -      -      -      -      -     vpmulld	%ymm16, %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     0.50    -      -     0.50   0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpmulld	(%rax), %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -     0.50    -      -     0.50   0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpmulld	(%rax){1to8}, %ymm17, %ymm19 {%k1} {z}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vpscatterdd	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vpscatterdq	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vpscatterqd	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vpscatterqq	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vpscatterdd	%ymm1, (%rdx,%ymm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vpscatterdq	%ymm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vpscatterqd	%xmm1, (%rdx,%ymm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vpscatterqq	%ymm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     8.25   8.25   8.25   8.25   2.00   2.00   4.00   1.33   1.33   1.33    -      -      -     2.00   2.00   vpscatterdd	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00   2.00   0.67   0.67   0.67    -      -      -     1.00   1.00   vpscatterdq	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00   2.00   0.67   0.67   0.67    -      -      -     1.00   1.00   vpscatterqd	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00   2.00   0.67   0.67   0.67    -      -      -     1.00   1.00   vpscatterqq	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     12.25  12.25  12.25  12.25  4.00   4.00   8.00   2.67   2.67   2.67    -      -      -     4.00   4.00   vpscatterdd	%ymm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     6.00   6.00   6.00   6.00   2.00   2.00   4.00   1.33   1.33   1.33    -      -      -     2.00   2.00   vpscatterdq	%ymm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     6.50   6.50   6.50   6.50   2.00   2.00   4.00   1.33   1.33   1.33    -      -      -     2.00   2.00   vpscatterqd	%xmm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     6.50   6.50   6.50   6.50   2.00   2.00   4.00   1.33   1.33   1.33    -      -      -     2.00   2.00   vpscatterqq	%ymm1, (%rdx,%ymm0,4) {%k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     0.50   0.50    -      -      -      -      -      -      -      -      -      -      -      -     vpshufd	$0, %xmm16, %xmm19
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     0.50   0.50    -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpshufd	$0, (%rax), %xmm19
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     0.50   0.50    -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpshufd	$0, (%rax){1to4}, %xmm19
@@ -5104,14 +5104,14 @@ vunpcklps         (%rax){1to8}, %ymm17, %ymm19 {z}{k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     0.50   0.50    -      -      -      -      -      -      -      -      -      -      -      -     vpunpckldq	%ymm16, %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     0.50   0.50    -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpunpckldq	(%rax), %ymm17, %ymm19 {%k1} {z}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     0.50   0.50    -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vpunpckldq	(%rax){1to8}, %ymm17, %ymm19 {%k1} {z}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vscatterdps	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vscatterdpd	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vscatterqps	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vscatterqpd	%xmm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vscatterdps	%ymm1, (%rdx,%ymm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vscatterdpd	%ymm1, (%rdx,%xmm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vscatterqps	%xmm1, (%rdx,%ymm0,4) {%k1}
-# CHECK-NEXT: 0.33   0.33   0.33    -      -      -      -      -      -      -      -      -      -      -      -     0.67   0.67   0.67    -      -      -     1.00   1.00   vscatterqpd	%ymm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     8.25   8.25   8.25   8.25   2.00   2.00   4.00   1.33   1.33   1.33    -      -      -     2.00   2.00   vscatterdps	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00   2.00   0.67   0.67   0.67    -      -      -     1.00   1.00   vscatterdpd	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00   2.00   0.67   0.67   0.67    -      -      -     1.00   1.00   vscatterqps	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT: 0.67   0.67   0.67    -      -      -      -      -     4.00   4.00   4.00   4.00   1.00   1.00   2.00   0.67   0.67   0.67    -      -      -     1.00   1.00   vscatterqpd	%xmm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT: 2.67   2.67   2.67    -      -      -      -      -     12.25  12.25  12.25  12.25  4.00   4.00   8.00   2.67   2.67   2.67    -      -      -     4.00   4.00   vscatterdps	%ymm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     6.00   6.00   6.00   6.00   2.00   2.00   4.00   1.33   1.33   1.33    -      -      -     2.00   2.00   vscatterdpd	%ymm1, (%rdx,%xmm0,4) {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     6.50   6.50   6.50   6.50   2.00   2.00   4.00   1.33   1.33   1.33    -      -      -     2.00   2.00   vscatterqps	%xmm1, (%rdx,%ymm0,4) {%k1}
+# CHECK-NEXT: 1.33   1.33   1.33    -      -      -      -      -     6.50   6.50   6.50   6.50   2.00   2.00   4.00   1.33   1.33   1.33    -      -      -     2.00   2.00   vscatterqpd	%ymm1, (%rdx,%ymm0,4) {%k1}
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00    -      -      -      -      -      -      -      -      -      -      -      -      -     vshuff32x4	$0, %ymm16, %ymm17, %ymm19
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00    -      -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vshuff32x4	$0, (%rax), %ymm17, %ymm19
 # CHECK-NEXT:  -      -      -      -      -      -      -      -      -     1.00    -      -     0.50   0.50    -     0.33   0.33   0.33   0.33   0.33   0.33    -      -     vshuff32x4	$0, (%rax){1to8}, %ymm17, %ymm19

>From 2c48f092c9b03b8d607ba2ee1f5df705e9065e28 Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Wed, 2 Sep 2026 01:18:37 +0530
Subject: [PATCH 2/5] [X86][CostModel] Pay masked gather/scatter overhead once
 per operation

getGSVectorCost handles a wide masked gather/scatter that legalises into
SplitFactor sub-ops by recursing on the narrower shape, which charged the
flat vectorize-vs-scalarize break-even overhead (getGatherOverhead /
getScatterOverhead) once per sub-op. That overhead is a property of the
original, pre-split operation and should be paid once; multiplying it by the
split factor over-counted every shape wide enough to split and made those
gathers/scatters look artificially expensive.

Charge the overhead once for the original shape and scale only the
instruction body (VF * scalar-memory-op cost) by the split factor, and fold
the TCK_CodeSize path into the same structure (the sub-op count). The body is
computed through a single helper so the alignment/address-space inputs cannot
diverge between the split and unsplit paths.

This is independent of any target-specific tuning; it slightly lowers the cost
of split AVX-512/AVX2 gathers and scatters and, in a few cases, lets the loop
vectorizer choose a wider VF. The autogenerated cost-model and LoopVectorize
expectations are regenerated accordingly.
---
 .../lib/Target/X86/X86TargetTransformInfo.cpp | 45 ++++++++----
 .../X86/masked-intrinsic-cost-inseltpoison.ll | 38 +++++------
 .../CostModel/X86/masked-intrinsic-cost.ll    | 38 +++++------
 .../X86/CostModel/gather-i32-with-i8-index.ll | 10 +--
 .../X86/CostModel/gather-i64-with-i8-index.ll | 12 ++--
 .../interleaved-load-f32-stride-2.ll          |  3 -
 .../interleaved-load-f32-stride-4.ll          |  5 --
 .../interleaved-load-f32-stride-8.ll          |  9 ---
 .../interleaved-load-f64-stride-5.ll          | 30 ++++----
 .../interleaved-load-f64-stride-6.ll          | 36 +++++-----
 .../interleaved-load-f64-stride-7.ll          | 42 ++++++------
 .../interleaved-load-f64-stride-8.ll          | 48 ++++++-------
 ...nterleaved-load-i32-stride-2-indices-0u.ll |  2 -
 .../interleaved-load-i32-stride-2.ll          |  3 -
 ...erleaved-load-i32-stride-4-indices-012u.ll |  4 --
 ...erleaved-load-i32-stride-4-indices-01uu.ll |  3 -
 ...erleaved-load-i32-stride-4-indices-0uuu.ll |  4 +-
 .../interleaved-load-i32-stride-4.ll          |  5 --
 .../interleaved-load-i32-stride-8.ll          |  9 ---
 .../interleaved-load-i64-stride-2.ll          |  8 +--
 .../interleaved-load-i64-stride-4.ll          | 24 +++----
 .../interleaved-load-i64-stride-5.ll          | 30 ++++----
 .../interleaved-load-i64-stride-6.ll          | 36 +++++-----
 .../interleaved-load-i64-stride-7.ll          | 42 ++++++------
 .../interleaved-load-i64-stride-8.ll          | 48 ++++++-------
 .../interleaved-store-f32-stride-8.ll         | 27 --------
 .../interleaved-store-f64-stride-4.ll         |  5 --
 .../interleaved-store-f64-stride-8.ll         | 48 ++++++-------
 .../interleaved-store-i32-stride-8.ll         | 27 --------
 .../interleaved-store-i64-stride-4.ll         |  5 --
 .../interleaved-store-i64-stride-8.ll         | 48 ++++++-------
 .../masked-gather-i32-with-i8-index.ll        | 10 +--
 .../masked-gather-i64-with-i8-index.ll        | 12 ++--
 .../masked-scatter-i32-with-i8-index.ll       |  4 +-
 .../masked-scatter-i64-with-i8-index.ll       |  6 +-
 .../CostModel/scatter-i32-with-i8-index.ll    |  4 +-
 .../CostModel/scatter-i64-with-i8-index.ll    |  6 +-
 .../LoopVectorize/X86/cast-costs.ll           | 24 +++----
 .../LoopVectorize/X86/interleave-cost.ll      | 68 +++++++++----------
 39 files changed, 369 insertions(+), 459 deletions(-)

diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
index 8e0cf1fc5a153..4acd2d898e503 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
@@ -6659,24 +6659,41 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
   std::pair<InstructionCost, MVT> SrcLT = getTypeLegalizationCost(SrcVTy);
   InstructionCost::CostType SplitFactor =
       std::max(IdxsLT.first, SrcLT.first).getValue();
+  const bool IsLoad = Opcode == Instruction::Load;
+
+  // If a wide gather/scatter splits (e.g. a 16-wide op whose 64-bit indices
+  // don't fit in one zmm) the hardware issues SplitFactor sub-ops; for code
+  // size that is simply the sub-op count.
+  if (CostKind == TTI::TCK_CodeSize)
+    return SplitFactor;
+
+  // A masked gather/scatter cost is a flat vectorize-vs-scalarize break-even
+  // overhead plus the cost of the instruction body:
+  //
+  //   cost = break-even overhead + instruction body cost
+  //
+  // The flat overhead is a property of the original, pre-split shape and is
+  // paid once regardless of how many sub-ops the op splits into (previously it
+  // was charged once per sub-op via the recursive split, over-counting wide
+  // shapes). Only the body scales with the split factor. The body is computed
+  // once here so the alignment/address-space inputs cannot diverge.
+  const int GSOverhead = IsLoad ? getGatherOverhead() : getScatterOverhead();
+
+  auto BodyCostOf = [&](Type *VTy) -> InstructionCost {
+    unsigned Lanes = cast<FixedVectorType>(VTy)->getNumElements();
+    return Lanes * getMemoryOpCost(Opcode, VTy->getScalarType(), Alignment,
+                                   AddressSpace, CostKind);
+  };
+
+  InstructionCost BodyCost;
   if (SplitFactor > 1) {
-    // Handle splitting of vector of pointers
     auto *SplitSrcTy =
         FixedVectorType::get(SrcVTy->getScalarType(), VF / SplitFactor);
-    return SplitFactor * getGSVectorCost(Opcode, CostKind, SplitSrcTy, Ptr,
-                                         Alignment, AddressSpace);
+    BodyCost = SplitFactor * BodyCostOf(SplitSrcTy);
+  } else {
+    BodyCost = BodyCostOf(SrcVTy);
   }
-
-  // If we didn't split, this will be a single gather/scatter instruction.
-  if (CostKind == TTI::TCK_CodeSize)
-    return 1;
-
-  // The gather / scatter cost is given by Intel architects. It is a rough
-  // number since we are looking at one instruction in a time.
-  const int GSOverhead = (Opcode == Instruction::Load) ? getGatherOverhead()
-                                                       : getScatterOverhead();
-  return GSOverhead + VF * getMemoryOpCost(Opcode, SrcVTy->getScalarType(),
-                                           Alignment, AddressSpace, CostKind);
+  return GSOverhead + BodyCost;
 }
 
 /// Calculate the cost of Gather / Scatter operation
diff --git a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll
index f835a8aedbbcc..5bf852c7ebfdd 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll
@@ -840,20 +840,20 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret i32 0
 ;
 ; SKL-LABEL: 'masked_gather'
-; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8F64 = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x double> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8F64 = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x double> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F64 = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x double> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2F64 = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x double> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1F64 = call <1 x double> @llvm.masked.gather.v1f64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x double> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F32 = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2F32 = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x float> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8I64 = call <8 x i64> @llvm.masked.gather.v8i64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i64> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8I64 = call <8 x i64> @llvm.masked.gather.v8i64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I64 = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:107 CodeSize:139 Lat:235 SizeLat:139 for: %V32I16 = call <32 x i16> @llvm.masked.gather.v32i16.v32p0(<32 x ptr> align 1 undef, <32 x i1> %m32, <32 x i16> undef)
@@ -871,7 +871,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:32 SizeLat:20 for: %V4F64 = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x double> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F64 = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x double> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1F64 = call <1 x double> @llvm.masked.gather.v1f64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x double> undef)
-; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
+; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:32 SizeLat:20 for: %V4F32 = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x float> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F32 = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x float> undef)
@@ -879,7 +879,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:22 Lat:34 SizeLat:22 for: %V4I64 = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i64> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
-; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:17 CodeSize:21 Lat:33 SizeLat:21 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -898,7 +898,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F64 = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x double> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F64 = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x double> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1F64 = call <1 x double> @llvm.masked.gather.v1f64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x double> undef)
-; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
+; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F32 = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x float> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F32 = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x float> undef)
@@ -906,7 +906,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I64 = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i64> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
-; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -1094,7 +1094,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f64.v2p0(<2 x double> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1f64.v1p0(<1 x double> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8f32.v8p0(<8 x float> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f32.v2p0(<2 x float> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1102,7 +1102,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:22 Lat:22 SizeLat:22 for: call void @llvm.masked.scatter.v4i64.v4p0(<4 x i64> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:17 CodeSize:21 Lat:21 SizeLat:21 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1121,7 +1121,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f64.v2p0(<2 x double> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1f64.v1p0(<1 x double> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8f32.v8p0(<8 x float> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f32.v2p0(<2 x float> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1129,7 +1129,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4i64.v4p0(<4 x i64> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1909,7 +1909,7 @@ define <16 x float> @test_gather_16f32_const_mask(ptr %base, <16 x i32> %ind) {
 ; SKL-LABEL: 'test_gather_16f32_const_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_const_mask'
@@ -1953,7 +1953,7 @@ define <16 x float> @test_gather_16f32_var_mask(ptr %base, <16 x i32> %ind, <16
 ; SKL-LABEL: 'test_gather_16f32_var_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_var_mask'
@@ -1997,13 +1997,13 @@ define <16 x float> @test_gather_16f32_ra_var_mask(<16 x ptr> %ptrs, <16 x i32>
 ; SKL-LABEL: 'test_gather_16f32_ra_var_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, <16 x ptr> %ptrs, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_ra_var_mask'
 ; AVX512-NEXT:  Cost Model: Found costs of RThru:3 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; AVX512-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, <16 x ptr> %ptrs, <16 x i64> %sext_ind
-; AVX512-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; AVX512-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
 ; AVX512-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
   %sext_ind = sext <16 x i32> %ind to <16 x i64>
@@ -2051,7 +2051,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
 ; SKL-NEXT:  Cost Model: Found costs of RThru:1 CodeSize:1 Lat:3 SizeLat:2 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> poison, <16 x i32> zeroinitializer
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_const_mask2'
diff --git a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
index ed1b534fac8f8..32dd349bca659 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
@@ -840,20 +840,20 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret i32 0
 ;
 ; SKL-LABEL: 'masked_gather'
-; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8F64 = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x double> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8F64 = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x double> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F64 = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x double> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2F64 = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x double> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1F64 = call <1 x double> @llvm.masked.gather.v1f64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x double> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F32 = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2F32 = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x float> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8I64 = call <8 x i64> @llvm.masked.gather.v8i64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i64> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8I64 = call <8 x i64> @llvm.masked.gather.v8i64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I64 = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:107 CodeSize:139 Lat:235 SizeLat:139 for: %V32I16 = call <32 x i16> @llvm.masked.gather.v32i16.v32p0(<32 x ptr> align 1 undef, <32 x i1> %m32, <32 x i16> undef)
@@ -871,7 +871,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:32 SizeLat:20 for: %V4F64 = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x double> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F64 = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x double> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1F64 = call <1 x double> @llvm.masked.gather.v1f64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x double> undef)
-; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
+; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:32 SizeLat:20 for: %V4F32 = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x float> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F32 = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x float> undef)
@@ -879,7 +879,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:22 Lat:34 SizeLat:22 for: %V4I64 = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i64> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
-; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:17 CodeSize:21 Lat:33 SizeLat:21 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -898,7 +898,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F64 = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x double> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F64 = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x double> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1F64 = call <1 x double> @llvm.masked.gather.v1f64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x double> undef)
-; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
+; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F32 = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x float> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F32 = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x float> undef)
@@ -906,7 +906,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I64 = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i64> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
-; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -1094,7 +1094,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f64.v2p0(<2 x double> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1f64.v1p0(<1 x double> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8f32.v8p0(<8 x float> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f32.v2p0(<2 x float> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1102,7 +1102,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:22 Lat:22 SizeLat:22 for: call void @llvm.masked.scatter.v4i64.v4p0(<4 x i64> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:17 CodeSize:21 Lat:21 SizeLat:21 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1121,7 +1121,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f64.v2p0(<2 x double> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1f64.v1p0(<1 x double> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8f32.v8p0(<8 x float> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f32.v2p0(<2 x float> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1129,7 +1129,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4i64.v4p0(<4 x i64> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1909,7 +1909,7 @@ define <16 x float> @test_gather_16f32_const_mask(ptr %base, <16 x i32> %ind) {
 ; SKL-LABEL: 'test_gather_16f32_const_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_const_mask'
@@ -1953,7 +1953,7 @@ define <16 x float> @test_gather_16f32_var_mask(ptr %base, <16 x i32> %ind, <16
 ; SKL-LABEL: 'test_gather_16f32_var_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_var_mask'
@@ -1997,13 +1997,13 @@ define <16 x float> @test_gather_16f32_ra_var_mask(<16 x ptr> %ptrs, <16 x i32>
 ; SKL-LABEL: 'test_gather_16f32_ra_var_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, <16 x ptr> %ptrs, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_ra_var_mask'
 ; AVX512-NEXT:  Cost Model: Found costs of RThru:3 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; AVX512-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, <16 x ptr> %ptrs, <16 x i64> %sext_ind
-; AVX512-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; AVX512-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
 ; AVX512-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
   %sext_ind = sext <16 x i32> %ind to <16 x i64>
@@ -2051,7 +2051,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
 ; SKL-NEXT:  Cost Model: Found costs of RThru:1 CodeSize:1 Lat:3 SizeLat:2 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> undef, <16 x i32> zeroinitializer
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_const_mask2'
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll
index 2d5a30019bacd..9191e95c0236d 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll
@@ -50,9 +50,9 @@ define void @test() {
 ; AVX2-FASTGATHER:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB = load i32, ptr %inB, align 4
 ; AVX2-FASTGATHER:  Cost of 4 for VF 2: WIDEN ir<%valB> = load ir<%inB>
 ; AVX2-FASTGATHER:  Cost of 6 for VF 4: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER:  Cost of 12 for VF 8: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER:  Cost of 24 for VF 16: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER:  Cost of 48 for VF 32: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER:  Cost of 10 for VF 8: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER:  Cost of 18 for VF 16: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER:  Cost of 34 for VF 32: WIDEN ir<%valB> = load ir<%inB>
 ;
 ; AVX512-LABEL: 'test'
 ; AVX512:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB = load i32, ptr %inB, align 4
@@ -60,8 +60,8 @@ define void @test() {
 ; AVX512:  Cost of 13 for VF 4: REPLICATE ir<%valB> = load ir<%inB>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%valB> = load ir<%inB>
 ; AVX512:  Cost of 18 for VF 16: WIDEN ir<%valB> = load ir<%inB>
-; AVX512:  Cost of 36 for VF 32: WIDEN ir<%valB> = load ir<%inB>
-; AVX512:  Cost of 72 for VF 64: WIDEN ir<%valB> = load ir<%inB>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%valB> = load ir<%inB>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%valB> = load ir<%inB>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i64-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i64-with-i8-index.ll
index ce5828a46eea7..189d5c837501d 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i64-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i64-with-i8-index.ll
@@ -50,18 +50,18 @@ define void @test() {
 ; AVX2-FASTGATHER:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB = load i64, ptr %inB, align 8
 ; AVX2-FASTGATHER:  Cost of 4 for VF 2: WIDEN ir<%valB> = load ir<%inB>
 ; AVX2-FASTGATHER:  Cost of 6 for VF 4: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER:  Cost of 12 for VF 8: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER:  Cost of 24 for VF 16: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER:  Cost of 48 for VF 32: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER:  Cost of 10 for VF 8: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER:  Cost of 18 for VF 16: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER:  Cost of 34 for VF 32: WIDEN ir<%valB> = load ir<%inB>
 ;
 ; AVX512-LABEL: 'test'
 ; AVX512:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB = load i64, ptr %inB, align 8
 ; AVX512:  Cost of 6 for VF 2: REPLICATE ir<%valB> = load ir<%inB>
 ; AVX512:  Cost of 14 for VF 4: REPLICATE ir<%valB> = load ir<%inB>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%valB> = load ir<%inB>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%valB> = load ir<%inB>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%valB> = load ir<%inB>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%valB> = load ir<%inB>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%valB> = load ir<%inB>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%valB> = load ir<%inB>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%valB> = load ir<%inB>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-2.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-2.ll
index 3b3c3dcb78041..6232ac0e913ca 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-2.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-2.ll
@@ -70,9 +70,6 @@ define void @test() {
 ; AVX512:  Cost of 34 for VF 32: INTERLEAVE-GROUP with factor 2, ir<%in0>
 ; AVX512:    ir<%v0> = load from index 0
 ; AVX512:    ir<%v1> = load from index 1
-; AVX512:  Cost of 148 for VF 64: INTERLEAVE-GROUP with factor 2, ir<%in0>
-; AVX512:    ir<%v0> = load from index 0
-; AVX512:    ir<%v1> = load from index 1
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-4.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-4.ll
index 679ff5ecec952..c2dfbe4f26870 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-4.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-4.ll
@@ -77,11 +77,6 @@ define void @test() {
 ; AVX512:    ir<%v1> = load from index 1
 ; AVX512:    ir<%v2> = load from index 2
 ; AVX512:    ir<%v3> = load from index 3
-; AVX512:  Cost of 148 for VF 32: INTERLEAVE-GROUP with factor 4, ir<%in0>
-; AVX512:    ir<%v0> = load from index 0
-; AVX512:    ir<%v1> = load from index 1
-; AVX512:    ir<%v2> = load from index 2
-; AVX512:    ir<%v3> = load from index 3
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-8.ll
index eafe9cfd912c4..a760b1ca0d1e7 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-8.ll
@@ -72,15 +72,6 @@ define void @test() {
 ; AVX512:    ir<%v5> = load from index 5
 ; AVX512:    ir<%v6> = load from index 6
 ; AVX512:    ir<%v7> = load from index 7
-; AVX512:  Cost of 148 for VF 16: INTERLEAVE-GROUP with factor 8, ir<%in0>
-; AVX512:    ir<%v0> = load from index 0
-; AVX512:    ir<%v1> = load from index 1
-; AVX512:    ir<%v2> = load from index 2
-; AVX512:    ir<%v3> = load from index 3
-; AVX512:    ir<%v4> = load from index 4
-; AVX512:    ir<%v5> = load from index 5
-; AVX512:    ir<%v6> = load from index 6
-; AVX512:    ir<%v7> = load from index 7
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-5.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-5.ll
index 7d18d05a344a8..a3cf9835db917 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-5.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-5.ll
@@ -54,21 +54,21 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v2> = load ir<%in2>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v3> = load ir<%in3>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-6.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-6.ll
index 909e0a9c8ed18..7ca2a371b05fa 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-6.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-6.ll
@@ -75,24 +75,24 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v3> = load ir<%in3>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v4> = load ir<%in4>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v5> = load ir<%in5>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-7.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-7.ll
index a9cef6bd260a1..b807cf344cbe5 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-7.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-7.ll
@@ -59,27 +59,27 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v4> = load ir<%in4>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v5> = load ir<%in5>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v6> = load ir<%in6>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-8.ll
index ba511ad467b1a..ddbb47e9e7f55 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-8.ll
@@ -62,30 +62,30 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v5> = load ir<%in5>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v6> = load ir<%in6>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v7> = load ir<%in7>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v7> = load ir<%in7>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v7> = load ir<%in7>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v7> = load ir<%in7>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v7> = load ir<%in7>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v7> = load ir<%in7>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v7> = load ir<%in7>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2-indices-0u.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2-indices-0u.ll
index 3d0e795d653a4..3c8ae5ec46962 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2-indices-0u.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2-indices-0u.ll
@@ -56,8 +56,6 @@ define void @test() {
 ; AVX512:    ir<%v0> = load from index 0
 ; AVX512:  Cost of 19 for VF 32: INTERLEAVE-GROUP with factor 2, ir<%in0>
 ; AVX512:    ir<%v0> = load from index 0
-; AVX512:  Cost of 78 for VF 64: INTERLEAVE-GROUP with factor 2, ir<%in0>
-; AVX512:    ir<%v0> = load from index 0
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2.ll
index 302e0cb447d65..34e448379a105 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2.ll
@@ -70,9 +70,6 @@ define void @test() {
 ; AVX512:  Cost of 34 for VF 32: INTERLEAVE-GROUP with factor 2, ir<%in0>
 ; AVX512:    ir<%v0> = load from index 0
 ; AVX512:    ir<%v1> = load from index 1
-; AVX512:  Cost of 148 for VF 64: INTERLEAVE-GROUP with factor 2, ir<%in0>
-; AVX512:    ir<%v0> = load from index 0
-; AVX512:    ir<%v1> = load from index 1
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-012u.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-012u.ll
index 41bd92713c386..c21f1e670528b 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-012u.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-012u.ll
@@ -68,10 +68,6 @@ define void @test() {
 ; AVX512:    ir<%v0> = load from index 0
 ; AVX512:    ir<%v1> = load from index 1
 ; AVX512:    ir<%v2> = load from index 2
-; AVX512:  Cost of 113 for VF 32: INTERLEAVE-GROUP with factor 4, ir<%in0>
-; AVX512:    ir<%v0> = load from index 0
-; AVX512:    ir<%v1> = load from index 1
-; AVX512:    ir<%v2> = load from index 2
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-01uu.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-01uu.ll
index 9a8ae47d51907..6c43b50da956c 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-01uu.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-01uu.ll
@@ -59,9 +59,6 @@ define void @test() {
 ; AVX512:  Cost of 19 for VF 16: INTERLEAVE-GROUP with factor 4, ir<%in0>
 ; AVX512:    ir<%v0> = load from index 0
 ; AVX512:    ir<%v1> = load from index 1
-; AVX512:  Cost of 78 for VF 32: INTERLEAVE-GROUP with factor 4, ir<%in0>
-; AVX512:    ir<%v0> = load from index 0
-; AVX512:    ir<%v1> = load from index 1
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-0uuu.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-0uuu.ll
index d2e518a981de5..939e58c9381cf 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-0uuu.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-0uuu.ll
@@ -50,8 +50,8 @@ define void @test() {
 ; AVX512:    ir<%v0> = load from index 0
 ; AVX512:  Cost of 8 for VF 16: INTERLEAVE-GROUP with factor 4, ir<%in0>
 ; AVX512:    ir<%v0> = load from index 0
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4.ll
index bd7a63609acde..2d56cdf49e0a1 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4.ll
@@ -77,11 +77,6 @@ define void @test() {
 ; AVX512:    ir<%v1> = load from index 1
 ; AVX512:    ir<%v2> = load from index 2
 ; AVX512:    ir<%v3> = load from index 3
-; AVX512:  Cost of 148 for VF 32: INTERLEAVE-GROUP with factor 4, ir<%in0>
-; AVX512:    ir<%v0> = load from index 0
-; AVX512:    ir<%v1> = load from index 1
-; AVX512:    ir<%v2> = load from index 2
-; AVX512:    ir<%v3> = load from index 3
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-8.ll
index c64076fca0d62..a276757bce5f9 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-8.ll
@@ -72,15 +72,6 @@ define void @test() {
 ; AVX512:    ir<%v5> = load from index 5
 ; AVX512:    ir<%v6> = load from index 6
 ; AVX512:    ir<%v7> = load from index 7
-; AVX512:  Cost of 148 for VF 16: INTERLEAVE-GROUP with factor 8, ir<%in0>
-; AVX512:    ir<%v0> = load from index 0
-; AVX512:    ir<%v1> = load from index 1
-; AVX512:    ir<%v2> = load from index 2
-; AVX512:    ir<%v3> = load from index 3
-; AVX512:    ir<%v4> = load from index 4
-; AVX512:    ir<%v5> = load from index 5
-; AVX512:    ir<%v6> = load from index 6
-; AVX512:    ir<%v7> = load from index 7
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-2.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-2.ll
index d0a99efab706d..46836b631c9e0 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-2.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-2.ll
@@ -63,10 +63,10 @@ define void @test() {
 ; AVX512:  Cost of 34 for VF 16: INTERLEAVE-GROUP with factor 2, ir<%in0>
 ; AVX512:    ir<%v0> = load from index 0
 ; AVX512:    ir<%v1> = load from index 1
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-4.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-4.ll
index 180ef142675f5..60b2cdcba7626 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-4.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-4.ll
@@ -68,18 +68,18 @@ define void @test() {
 ; AVX512:    ir<%v1> = load from index 1
 ; AVX512:    ir<%v2> = load from index 2
 ; AVX512:    ir<%v3> = load from index 3
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-5.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-5.ll
index 65a27fc96f223..3c5ddeab269cf 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-5.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-5.ll
@@ -54,21 +54,21 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v2> = load ir<%in2>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v3> = load ir<%in3>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-6.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-6.ll
index d261e5242f8d8..611c1e28dbc42 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-6.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-6.ll
@@ -75,24 +75,24 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v3> = load ir<%in3>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v4> = load ir<%in4>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v5> = load ir<%in5>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-7.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-7.ll
index 04c8db7a58357..3e9886ee10ac3 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-7.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-7.ll
@@ -59,27 +59,27 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v4> = load ir<%in4>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v5> = load ir<%in5>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v6> = load ir<%in6>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-8.ll
index a2b5d1c4df2f8..6e0c384d73e7a 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-8.ll
@@ -62,30 +62,30 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v5> = load ir<%in5>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v6> = load ir<%in6>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v7> = load ir<%in7>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v7> = load ir<%in7>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v7> = load ir<%in7>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v7> = load ir<%in7>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v7> = load ir<%in7>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v7> = load ir<%in7>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v7> = load ir<%in7>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f32-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f32-stride-8.ll
index 8d813d73120cb..4f3aa27971213 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f32-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f32-stride-8.ll
@@ -172,33 +172,6 @@ define void @test() {
 ; AVX512:    store ir<%v5> to index 5
 ; AVX512:    store ir<%v6> to index 6
 ; AVX512:    store ir<%v7> to index 7
-; AVX512:  Cost of 148 for VF 16: INTERLEAVE-GROUP with factor 8, ir<%out0>
-; AVX512:    store ir<%v0> to index 0
-; AVX512:    store ir<%v1> to index 1
-; AVX512:    store ir<%v2> to index 2
-; AVX512:    store ir<%v3> to index 3
-; AVX512:    store ir<%v4> to index 4
-; AVX512:    store ir<%v5> to index 5
-; AVX512:    store ir<%v6> to index 6
-; AVX512:    store ir<%v7> to index 7
-; AVX512:  Cost of 296 for VF 32: INTERLEAVE-GROUP with factor 8, ir<%out0>
-; AVX512:    store ir<%v0> to index 0
-; AVX512:    store ir<%v1> to index 1
-; AVX512:    store ir<%v2> to index 2
-; AVX512:    store ir<%v3> to index 3
-; AVX512:    store ir<%v4> to index 4
-; AVX512:    store ir<%v5> to index 5
-; AVX512:    store ir<%v6> to index 6
-; AVX512:    store ir<%v7> to index 7
-; AVX512:  Cost of 592 for VF 64: INTERLEAVE-GROUP with factor 8, ir<%out0>
-; AVX512:    store ir<%v0> to index 0
-; AVX512:    store ir<%v1> to index 1
-; AVX512:    store ir<%v2> to index 2
-; AVX512:    store ir<%v3> to index 3
-; AVX512:    store ir<%v4> to index 4
-; AVX512:    store ir<%v5> to index 5
-; AVX512:    store ir<%v6> to index 6
-; AVX512:    store ir<%v7> to index 7
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-4.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-4.ll
index 9799b132c3697..b636530a650ec 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-4.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-4.ll
@@ -114,11 +114,6 @@ define void @test() {
 ; AVX512:    store ir<%v1> to index 1
 ; AVX512:    store ir<%v2> to index 2
 ; AVX512:    store ir<%v3> to index 3
-; AVX512:  Cost of 272 for VF 64: INTERLEAVE-GROUP with factor 4, ir<%out0>
-; AVX512:    store ir<%v0> to index 0
-; AVX512:    store ir<%v1> to index 1
-; AVX512:    store ir<%v2> to index 2
-; AVX512:    store ir<%v3> to index 3
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-8.ll
index 22d395647482d..11635959289f9 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-8.ll
@@ -32,30 +32,30 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out5>, ir<%v5>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out6>, ir<%v6>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out7>, ir<%v7>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out0>, ir<%v0>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out1>, ir<%v1>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out2>, ir<%v2>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out3>, ir<%v3>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out4>, ir<%v4>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out5>, ir<%v5>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out6>, ir<%v6>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out7>, ir<%v7>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out0>, ir<%v0>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out1>, ir<%v1>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out2>, ir<%v2>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out3>, ir<%v3>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out4>, ir<%v4>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out5>, ir<%v5>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out6>, ir<%v6>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out7>, ir<%v7>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out0>, ir<%v0>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out1>, ir<%v1>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out2>, ir<%v2>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out3>, ir<%v3>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out4>, ir<%v4>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out5>, ir<%v5>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out6>, ir<%v6>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out7>, ir<%v7>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out0>, ir<%v0>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out1>, ir<%v1>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out2>, ir<%v2>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out3>, ir<%v3>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out4>, ir<%v4>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out5>, ir<%v5>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out6>, ir<%v6>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out7>, ir<%v7>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out0>, ir<%v0>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out1>, ir<%v1>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out2>, ir<%v2>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out3>, ir<%v3>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out4>, ir<%v4>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out5>, ir<%v5>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out6>, ir<%v6>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out7>, ir<%v7>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out0>, ir<%v0>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out1>, ir<%v1>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out2>, ir<%v2>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out3>, ir<%v3>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out4>, ir<%v4>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out5>, ir<%v5>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out6>, ir<%v6>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out7>, ir<%v7>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i32-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i32-stride-8.ll
index 506aaab0951e4..50888201df108 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i32-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i32-stride-8.ll
@@ -172,33 +172,6 @@ define void @test() {
 ; AVX512:    store ir<%v5> to index 5
 ; AVX512:    store ir<%v6> to index 6
 ; AVX512:    store ir<%v7> to index 7
-; AVX512:  Cost of 148 for VF 16: INTERLEAVE-GROUP with factor 8, ir<%out0>
-; AVX512:    store ir<%v> to index 0
-; AVX512:    store ir<%v1> to index 1
-; AVX512:    store ir<%v2> to index 2
-; AVX512:    store ir<%v3> to index 3
-; AVX512:    store ir<%v4> to index 4
-; AVX512:    store ir<%v5> to index 5
-; AVX512:    store ir<%v6> to index 6
-; AVX512:    store ir<%v7> to index 7
-; AVX512:  Cost of 296 for VF 32: INTERLEAVE-GROUP with factor 8, ir<%out0>
-; AVX512:    store ir<%v> to index 0
-; AVX512:    store ir<%v1> to index 1
-; AVX512:    store ir<%v2> to index 2
-; AVX512:    store ir<%v3> to index 3
-; AVX512:    store ir<%v4> to index 4
-; AVX512:    store ir<%v5> to index 5
-; AVX512:    store ir<%v6> to index 6
-; AVX512:    store ir<%v7> to index 7
-; AVX512:  Cost of 592 for VF 64: INTERLEAVE-GROUP with factor 8, ir<%out0>
-; AVX512:    store ir<%v> to index 0
-; AVX512:    store ir<%v1> to index 1
-; AVX512:    store ir<%v2> to index 2
-; AVX512:    store ir<%v3> to index 3
-; AVX512:    store ir<%v4> to index 4
-; AVX512:    store ir<%v5> to index 5
-; AVX512:    store ir<%v6> to index 6
-; AVX512:    store ir<%v7> to index 7
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-4.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-4.ll
index 805a50527749a..2b2e22eb74630 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-4.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-4.ll
@@ -114,11 +114,6 @@ define void @test() {
 ; AVX512:    store ir<%v1> to index 1
 ; AVX512:    store ir<%v2> to index 2
 ; AVX512:    store ir<%v3> to index 3
-; AVX512:  Cost of 272 for VF 64: INTERLEAVE-GROUP with factor 4, ir<%out0>
-; AVX512:    store ir<%v> to index 0
-; AVX512:    store ir<%v1> to index 1
-; AVX512:    store ir<%v2> to index 2
-; AVX512:    store ir<%v3> to index 3
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-8.ll
index 6e808e99d351f..9150ee0e631bb 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-8.ll
@@ -170,30 +170,30 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out5>, ir<%v5>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out6>, ir<%v6>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out7>, ir<%v7>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out0>, ir<%v>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out1>, ir<%v1>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out2>, ir<%v2>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out3>, ir<%v3>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out4>, ir<%v4>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out5>, ir<%v5>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out6>, ir<%v6>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out7>, ir<%v7>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out0>, ir<%v>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out1>, ir<%v1>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out2>, ir<%v2>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out3>, ir<%v3>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out4>, ir<%v4>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out5>, ir<%v5>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out6>, ir<%v6>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out7>, ir<%v7>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out0>, ir<%v>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out1>, ir<%v1>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out2>, ir<%v2>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out3>, ir<%v3>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out4>, ir<%v4>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out5>, ir<%v5>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out6>, ir<%v6>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out7>, ir<%v7>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out0>, ir<%v>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out1>, ir<%v1>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out2>, ir<%v2>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out3>, ir<%v3>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out4>, ir<%v4>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out5>, ir<%v5>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out6>, ir<%v6>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out7>, ir<%v7>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out0>, ir<%v>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out1>, ir<%v1>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out2>, ir<%v2>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out3>, ir<%v3>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out4>, ir<%v4>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out5>, ir<%v5>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out6>, ir<%v6>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out7>, ir<%v7>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out0>, ir<%v>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out1>, ir<%v1>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out2>, ir<%v2>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out3>, ir<%v3>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out4>, ir<%v4>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out5>, ir<%v5>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out6>, ir<%v6>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out7>, ir<%v7>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll
index 8f5da77027970..59714eb573bba 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll
@@ -43,9 +43,9 @@ define void @test() {
 ; AVX2-FASTGATHER:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB.loaded = load i32, ptr %inB, align 4
 ; AVX2-FASTGATHER:  Cost of 4 for VF 2: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ; AVX2-FASTGATHER:  Cost of 6 for VF 4: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER:  Cost of 12 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER:  Cost of 24 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER:  Cost of 48 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER:  Cost of 10 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER:  Cost of 18 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER:  Cost of 34 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ;
 ; AVX512-LABEL: 'test'
 ; AVX512:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB.loaded = load i32, ptr %inB, align 4
@@ -53,8 +53,8 @@ define void @test() {
 ; AVX512:  Cost of 17 for VF 4: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ; AVX512:  Cost of 18 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX512:  Cost of 36 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX512:  Cost of 72 for VF 64: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i64-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i64-with-i8-index.ll
index 6801a549e19c4..094829fc8aa22 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i64-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i64-with-i8-index.ll
@@ -43,18 +43,18 @@ define void @test() {
 ; AVX2-FASTGATHER:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB.loaded = load i64, ptr %inB, align 8
 ; AVX2-FASTGATHER:  Cost of 4 for VF 2: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ; AVX2-FASTGATHER:  Cost of 6 for VF 4: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER:  Cost of 12 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER:  Cost of 24 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER:  Cost of 48 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER:  Cost of 10 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER:  Cost of 18 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER:  Cost of 34 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ;
 ; AVX512-LABEL: 'test'
 ; AVX512:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB.loaded = load i64, ptr %inB, align 8
 ; AVX512:  Cost of 8 for VF 2: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ; AVX512:  Cost of 18 for VF 4: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX512:  Cost of 20 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX512:  Cost of 40 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX512:  Cost of 80 for VF 64: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX512:  Cost of 18 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX512:  Cost of 34 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX512:  Cost of 66 for VF 64: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i32-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i32-with-i8-index.ll
index 6959fea2d512b..b9b07ae9d2711 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i32-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i32-with-i8-index.ll
@@ -52,8 +52,8 @@ define void @test() {
 ; AVX512:  Cost of 10.5 for VF 4: REPLICATE store ir<%valB>, ir<%out>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
 ; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
-; AVX512:  Cost of 36 for VF 32: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
-; AVX512:  Cost of 72 for VF 64: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i64-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i64-with-i8-index.ll
index 41ae89933204e..2e01201fef3ed 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i64-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i64-with-i8-index.ll
@@ -51,9 +51,9 @@ define void @test() {
 ; AVX512:  Cost of 5 for VF 2: REPLICATE store ir<%valB>, ir<%out>
 ; AVX512:  Cost of 11 for VF 4: REPLICATE store ir<%valB>, ir<%out>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i32-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i32-with-i8-index.ll
index 5e0b3277dd5e9..043c14231846f 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i32-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i32-with-i8-index.ll
@@ -52,8 +52,8 @@ define void @test() {
 ; AVX512:  Cost of 13 for VF 4: REPLICATE store ir<%valB>, ir<%out>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out>, ir<%valB>
 ; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out>, ir<%valB>
-; AVX512:  Cost of 36 for VF 32: WIDEN store ir<%out>, ir<%valB>
-; AVX512:  Cost of 72 for VF 64: WIDEN store ir<%out>, ir<%valB>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out>, ir<%valB>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out>, ir<%valB>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i64-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i64-with-i8-index.ll
index fcbf6042dec14..fe934b4b62dfa 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i64-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i64-with-i8-index.ll
@@ -51,9 +51,9 @@ define void @test() {
 ; AVX512:  Cost of 6 for VF 2: REPLICATE store ir<%valB>, ir<%out>
 ; AVX512:  Cost of 14 for VF 4: REPLICATE store ir<%valB>, ir<%out>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out>, ir<%valB>
-; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out>, ir<%valB>
-; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out>, ir<%valB>
-; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out>, ir<%valB>
+; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out>, ir<%valB>
+; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out>, ir<%valB>
+; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out>, ir<%valB>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll b/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll
index 7f0ba19d3e7a8..bc2bea9708227 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll
@@ -123,28 +123,28 @@ define void @replicate_sext(i32 %N, ptr %dst, ptr %src) #0 {
 ; CHECK-NEXT:    [[FOUND_CONFLICT:%.*]] = and i1 [[BOUND0]], [[BOUND1]]
 ; CHECK-NEXT:    br i1 [[FOUND_CONFLICT]], label %[[SCALAR_PH]], label %[[VECTOR_PH:.*]]
 ; CHECK:       [[VECTOR_PH]]:
-; CHECK-NEXT:    [[N_MOD_VF:%.*]] = and i32 [[TMP0]], 3
+; CHECK-NEXT:    [[N_MOD_VF:%.*]] = and i32 [[TMP0]], 7
 ; CHECK-NEXT:    [[TMP18:%.*]] = icmp eq i32 [[N_MOD_VF]], 0
-; CHECK-NEXT:    [[TMP19:%.*]] = select i1 [[TMP18]], i32 4, i32 [[N_MOD_VF]]
+; CHECK-NEXT:    [[TMP19:%.*]] = select i1 [[TMP18]], i32 8, i32 [[N_MOD_VF]]
 ; CHECK-NEXT:    [[N_VEC:%.*]] = sub i32 [[TMP0]], [[TMP19]]
 ; CHECK-NEXT:    [[TMP20:%.*]] = shl i32 [[N_VEC]], 2
 ; CHECK-NEXT:    [[TMP21:%.*]] = mul i32 [[N_VEC]], 3
 ; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
 ; CHECK:       [[VECTOR_BODY]]:
 ; CHECK-NEXT:    [[INDEX:%.*]] = phi i32 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT:    [[VEC_IND:%.*]] = phi <4 x i32> [ <i32 0, i32 3, i32 6, i32 9>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[VEC_IND:%.*]] = phi <8 x i32> [ <i32 0, i32 3, i32 6, i32 9, i32 12, i32 15, i32 18, i32 21>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
 ; CHECK-NEXT:    [[OFFSET_IDX:%.*]] = shl i32 [[INDEX]], 2
 ; CHECK-NEXT:    [[TMP22:%.*]] = sext i32 [[OFFSET_IDX]] to i64
 ; CHECK-NEXT:    [[TMP23:%.*]] = getelementptr nusw i32, ptr [[SRC]], i64 [[TMP22]]
-; CHECK-NEXT:    [[WIDE_VEC:%.*]] = load <16 x i32>, ptr [[TMP23]], align 4, !alias.scope [[META4:![0-9]+]]
-; CHECK-NEXT:    [[STRIDED_VEC:%.*]] = shufflevector <16 x i32> [[WIDE_VEC]], <16 x i32> poison, <4 x i32> <i32 0, i32 4, i32 8, i32 12>
-; CHECK-NEXT:    [[STRIDED_VEC9:%.*]] = shufflevector <16 x i32> [[WIDE_VEC]], <16 x i32> poison, <4 x i32> <i32 1, i32 5, i32 9, i32 13>
-; CHECK-NEXT:    [[TMP25:%.*]] = sext <4 x i32> [[VEC_IND]] to <4 x i64>
-; CHECK-NEXT:    [[TMP27:%.*]] = getelementptr i32, ptr [[DST]], <4 x i64> [[TMP25]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[STRIDED_VEC]], <4 x ptr> align 4 [[TMP27]], <4 x i1> splat (i1 true)), !alias.scope [[META7:![0-9]+]], !noalias [[META4]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[STRIDED_VEC9]], <4 x ptr> align 4 [[TMP27]], <4 x i1> splat (i1 true)), !alias.scope [[META7]], !noalias [[META4]]
-; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 4
-; CHECK-NEXT:    [[VEC_IND_NEXT]] = add <4 x i32> [[VEC_IND]], splat (i32 12)
+; CHECK-NEXT:    [[WIDE_VEC:%.*]] = load <32 x i32>, ptr [[TMP23]], align 4, !alias.scope [[META4:![0-9]+]]
+; CHECK-NEXT:    [[STRIDED_VEC:%.*]] = shufflevector <32 x i32> [[WIDE_VEC]], <32 x i32> poison, <8 x i32> <i32 0, i32 4, i32 8, i32 12, i32 16, i32 20, i32 24, i32 28>
+; CHECK-NEXT:    [[STRIDED_VEC2:%.*]] = shufflevector <32 x i32> [[WIDE_VEC]], <32 x i32> poison, <8 x i32> <i32 1, i32 5, i32 9, i32 13, i32 17, i32 21, i32 25, i32 29>
+; CHECK-NEXT:    [[TMP24:%.*]] = sext <8 x i32> [[VEC_IND]] to <8 x i64>
+; CHECK-NEXT:    [[WIDE_GEP:%.*]] = getelementptr i32, ptr [[DST]], <8 x i64> [[TMP24]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[STRIDED_VEC]], <8 x ptr> align 4 [[WIDE_GEP]], <8 x i1> splat (i1 true)), !alias.scope [[META7:![0-9]+]], !noalias [[META4]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[STRIDED_VEC2]], <8 x ptr> align 4 [[WIDE_GEP]], <8 x i1> splat (i1 true)), !alias.scope [[META7]], !noalias [[META4]]
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 8
+; CHECK-NEXT:    [[VEC_IND_NEXT]] = add <8 x i32> [[VEC_IND]], splat (i32 24)
 ; CHECK-NEXT:    [[TMP26:%.*]] = icmp eq i32 [[INDEX_NEXT]], [[N_VEC]]
 ; CHECK-NEXT:    br i1 [[TMP26]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP9:![0-9]+]]
 ; CHECK:       [[MIDDLE_BLOCK]]:
diff --git a/llvm/test/Transforms/LoopVectorize/X86/interleave-cost.ll b/llvm/test/Transforms/LoopVectorize/X86/interleave-cost.ll
index aa5a0f5be24be..13692e2c41a84 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/interleave-cost.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/interleave-cost.ll
@@ -179,7 +179,7 @@ define void @geps_feeding_interleave_groups_with_reuse2(ptr %A, ptr %B, i64 %N)
 ; CHECK-NEXT:  [[ENTRY:.*]]:
 ; CHECK-NEXT:    [[TMP0:%.*]] = lshr i64 [[N]], 3
 ; CHECK-NEXT:    [[TMP1:%.*]] = add nuw nsw i64 [[TMP0]], 1
-; CHECK-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ule i64 [[TMP1]], 28
+; CHECK-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ule i64 [[TMP1]], 32
 ; CHECK-NEXT:    br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_SCEVCHECK:.*]]
 ; CHECK:       [[VECTOR_SCEVCHECK]]:
 ; CHECK-NEXT:    [[TMP2:%.*]] = lshr i64 [[N]], 3
@@ -256,48 +256,48 @@ define void @geps_feeding_interleave_groups_with_reuse2(ptr %A, ptr %B, i64 %N)
 ; CHECK-NEXT:    [[CONFLICT_RDX:%.*]] = or i1 [[FOUND_CONFLICT]], [[FOUND_CONFLICT40]]
 ; CHECK-NEXT:    br i1 [[CONFLICT_RDX]], label %[[SCALAR_PH]], label %[[VECTOR_PH:.*]]
 ; CHECK:       [[VECTOR_PH]]:
-; CHECK-NEXT:    [[N_MOD_VF:%.*]] = and i64 [[TMP1]], 3
+; CHECK-NEXT:    [[N_MOD_VF:%.*]] = and i64 [[TMP1]], 7
 ; CHECK-NEXT:    [[TMP48:%.*]] = icmp eq i64 [[N_MOD_VF]], 0
-; CHECK-NEXT:    [[TMP49:%.*]] = select i1 [[TMP48]], i64 4, i64 [[N_MOD_VF]]
+; CHECK-NEXT:    [[TMP49:%.*]] = select i1 [[TMP48]], i64 8, i64 [[N_MOD_VF]]
 ; CHECK-NEXT:    [[N_VEC:%.*]] = sub i64 [[TMP1]], [[TMP49]]
 ; CHECK-NEXT:    [[TMP50:%.*]] = shl i64 [[N_VEC]], 3
 ; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
 ; CHECK:       [[VECTOR_BODY]]:
 ; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT:    [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 8, i64 16, i64 24>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[VEC_IND:%.*]] = phi <8 x i64> [ <i64 0, i64 8, i64 16, i64 24, i64 32, i64 40, i64 48, i64 56>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
 ; CHECK-NEXT:    [[TMP51:%.*]] = shl nuw i64 [[INDEX]], 3
 ; CHECK-NEXT:    [[TMP52:%.*]] = lshr exact i64 [[TMP51]], 1
 ; CHECK-NEXT:    [[TMP53:%.*]] = getelementptr nusw i32, ptr [[B]], i64 [[TMP52]]
-; CHECK-NEXT:    [[WIDE_VEC:%.*]] = load <16 x i32>, ptr [[TMP53]], align 4, !alias.scope [[META3:![0-9]+]], !noalias [[META6:![0-9]+]]
-; CHECK-NEXT:    [[STRIDED_VEC:%.*]] = shufflevector <16 x i32> [[WIDE_VEC]], <16 x i32> poison, <4 x i32> <i32 0, i32 4, i32 8, i32 12>
-; CHECK-NEXT:    [[STRIDED_VEC41:%.*]] = shufflevector <16 x i32> [[WIDE_VEC]], <16 x i32> poison, <4 x i32> <i32 1, i32 5, i32 9, i32 13>
-; CHECK-NEXT:    [[WIDE_GEP:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[VEC_IND]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[STRIDED_VEC]], <4 x ptr> align 4 [[WIDE_GEP]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[TMP54:%.*]] = or disjoint <4 x i64> [[VEC_IND]], splat (i64 1)
-; CHECK-NEXT:    [[WIDE_GEP42:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[TMP54]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> zeroinitializer, <4 x ptr> align 4 [[WIDE_GEP42]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[TMP55:%.*]] = or disjoint <4 x i64> [[VEC_IND]], splat (i64 2)
-; CHECK-NEXT:    [[WIDE_GEP43:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[TMP55]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[STRIDED_VEC41]], <4 x ptr> align 4 [[WIDE_GEP43]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[TMP56:%.*]] = or disjoint <4 x i64> [[VEC_IND]], splat (i64 3)
-; CHECK-NEXT:    [[WIDE_GEP44:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[TMP56]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> zeroinitializer, <4 x ptr> align 4 [[WIDE_GEP44]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[TMP57:%.*]] = or disjoint <4 x i64> [[VEC_IND]], splat (i64 4)
-; CHECK-NEXT:    [[WIDE_GEP45:%.*]] = getelementptr i32, ptr [[B]], <4 x i64> [[VEC_IND]]
-; CHECK-NEXT:    [[WIDE_MASKED_GATHER:%.*]] = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 4 [[WIDE_GEP45]], <4 x i1> splat (i1 true), <4 x i32> poison), !alias.scope [[META8:![0-9]+]], !noalias [[META6]]
-; CHECK-NEXT:    [[WIDE_GEP46:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[TMP57]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[WIDE_MASKED_GATHER]], <4 x ptr> align 4 [[WIDE_GEP46]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[TMP58:%.*]] = or disjoint <4 x i64> [[VEC_IND]], splat (i64 5)
-; CHECK-NEXT:    [[WIDE_GEP47:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[TMP58]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> zeroinitializer, <4 x ptr> align 4 [[WIDE_GEP47]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[TMP59:%.*]] = or disjoint <4 x i64> [[VEC_IND]], splat (i64 6)
-; CHECK-NEXT:    [[WIDE_GEP48:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[TMP59]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> zeroinitializer, <4 x ptr> align 4 [[WIDE_GEP48]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[TMP60:%.*]] = or disjoint <4 x i64> [[VEC_IND]], splat (i64 7)
-; CHECK-NEXT:    [[WIDE_GEP49:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[TMP60]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> zeroinitializer, <4 x ptr> align 4 [[WIDE_GEP49]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; CHECK-NEXT:    [[VEC_IND_NEXT]] = add nuw nsw <4 x i64> [[VEC_IND]], splat (i64 32)
+; CHECK-NEXT:    [[WIDE_VEC:%.*]] = load <32 x i32>, ptr [[TMP53]], align 4, !alias.scope [[META3:![0-9]+]], !noalias [[META6:![0-9]+]]
+; CHECK-NEXT:    [[STRIDED_VEC:%.*]] = shufflevector <32 x i32> [[WIDE_VEC]], <32 x i32> poison, <8 x i32> <i32 0, i32 4, i32 8, i32 12, i32 16, i32 20, i32 24, i32 28>
+; CHECK-NEXT:    [[STRIDED_VEC17:%.*]] = shufflevector <32 x i32> [[WIDE_VEC]], <32 x i32> poison, <8 x i32> <i32 1, i32 5, i32 9, i32 13, i32 17, i32 21, i32 25, i32 29>
+; CHECK-NEXT:    [[WIDE_GEP:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[VEC_IND]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[STRIDED_VEC]], <8 x ptr> align 4 [[WIDE_GEP]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[TMP55:%.*]] = or disjoint <8 x i64> [[VEC_IND]], splat (i64 1)
+; CHECK-NEXT:    [[WIDE_GEP18:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[TMP55]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> zeroinitializer, <8 x ptr> align 4 [[WIDE_GEP18]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[TMP56:%.*]] = or disjoint <8 x i64> [[VEC_IND]], splat (i64 2)
+; CHECK-NEXT:    [[WIDE_GEP19:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[TMP56]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[STRIDED_VEC17]], <8 x ptr> align 4 [[WIDE_GEP19]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[TMP57:%.*]] = or disjoint <8 x i64> [[VEC_IND]], splat (i64 3)
+; CHECK-NEXT:    [[WIDE_GEP20:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[TMP57]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> zeroinitializer, <8 x ptr> align 4 [[WIDE_GEP20]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[TMP58:%.*]] = or disjoint <8 x i64> [[VEC_IND]], splat (i64 4)
+; CHECK-NEXT:    [[WIDE_GEP21:%.*]] = getelementptr i32, ptr [[B]], <8 x i64> [[VEC_IND]]
+; CHECK-NEXT:    [[WIDE_MASKED_GATHER:%.*]] = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 [[WIDE_GEP21]], <8 x i1> splat (i1 true), <8 x i32> poison), !alias.scope [[META8:![0-9]+]], !noalias [[META6]]
+; CHECK-NEXT:    [[WIDE_GEP22:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[TMP58]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[WIDE_MASKED_GATHER]], <8 x ptr> align 4 [[WIDE_GEP22]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[TMP59:%.*]] = or disjoint <8 x i64> [[VEC_IND]], splat (i64 5)
+; CHECK-NEXT:    [[WIDE_GEP23:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[TMP59]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> zeroinitializer, <8 x ptr> align 4 [[WIDE_GEP23]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[TMP60:%.*]] = or disjoint <8 x i64> [[VEC_IND]], splat (i64 6)
+; CHECK-NEXT:    [[WIDE_GEP24:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[TMP60]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> zeroinitializer, <8 x ptr> align 4 [[WIDE_GEP24]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[TMP64:%.*]] = or disjoint <8 x i64> [[VEC_IND]], splat (i64 7)
+; CHECK-NEXT:    [[WIDE_GEP25:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[TMP64]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> zeroinitializer, <8 x ptr> align 4 [[WIDE_GEP25]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 8
+; CHECK-NEXT:    [[VEC_IND_NEXT]] = add nuw nsw <8 x i64> [[VEC_IND]], splat (i64 64)
 ; CHECK-NEXT:    [[TMP61:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
 ; CHECK-NEXT:    br i1 [[TMP61]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP10:![0-9]+]]
 ; CHECK:       [[MIDDLE_BLOCK]]:

>From 98a1b24709873adba07eaed14f0ec4c993eb0663 Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Wed, 2 Sep 2026 01:18:57 +0530
Subject: [PATCH 3/5] [X86][CostModel] Add per-shape gather/scatter cost tables
 for AMD znver4+

The X86 cost model prices every masked gather/scatter through a single
flat overhead from getGatherOverhead / getScatterOverhead (2 on AVX-512)
plus VF * scalar-memory-op for the body. On Zen the real cost of these
instructions varies by more than 3x across shapes, so one flat number
forces the LoopVectorizer to mis-price loops that need indirect memory
access.

On znver4+ price them per shape as

  total = getZenGSOverhead(shape) + getModeledGSInstrCost(shape)

where the body is read live from the schedule model and the overhead is
the calibrated profitability premium stored here. Splitting it this way
keeps the hardware half in X86ScheduleZnver4.td as the single source of
truth, so a schedule-model retune needs no matching edit here.

Gate it on TuningPreferGSCostTable, inherited via ZN4Tuning. znver1..3
take the scalarise path for masked gather and never reach the code, so
the bit is deliberately not in ZNTuning. The feature carries InlineIgnore,
so it does not make callers and callees inline-incompatible.

The intent of a row depends on whether the masked instruction beats its
scalarised alternative for that shape. Where the masked op is faster
(most i32 / f32 / f64 shapes) the total sits at or below the cost at
which the LoopVectorizer would stop emitting it; where it is slower (all
i64 shapes, and v4i32 / v4f32 scatter) the total sits at or above that
flip, so the scalarised lowering is selected. The i64 rows are priced
above their f64 counterparts because the i64 scalar fallback runs on the
cheaper integer pipeline and is harder to beat.

Break-even totals were calibrated by sweeping a forced overhead over a
micro-benchmark with one indirect access per inner iteration, finding the
cost at which the vectorizer switches lowering, and timing the binary on
each side of the flip. Because the body term now comes from the schedule
model, retargeting X86ScheduleZnver4.td to Znver4 grew every body by 1 to
9 points; each row was re-checked against its flip and the six with no
headroom (v4f32, v4f64, v8f32, v8i32, v8f64 gather and v8f64 scatter)
were reduced by their body delta to restore the calibrated total.

Shapes with no table row -- VF < 4, which a 3-element gather can reach --
return no entry rather than asserting, and take the flat-overhead path
like any other unlisted shape.

Tests cover every live shape on znver4/znver5/znver6 plus the unchanged
behaviour for znver3, skx and x86-64-v4; pin the vectoriser decisions
end-to-end, including the guard that a unit-stride load must not become a
gather regardless of table values (issue #91370); and add SLP coverage.
---
 llvm/docs/ReleaseNotes.md                     |    5 +
 llvm/lib/Target/X86/X86.td                    |   15 +-
 .../lib/Target/X86/X86TargetTransformInfo.cpp |  210 +++-
 llvm/lib/Target/X86/X86TargetTransformInfo.h  |    9 +
 .../X86/masked-gather-scatter-cost-table.ll   | 1078 +++++++++++++++++
 .../gather-scatter-cost-table-decisions.ll    |   99 ++
 .../SLPVectorizer/X86/gather-cost-znver4.ll   |   86 ++
 7 files changed, 1484 insertions(+), 18 deletions(-)
 create mode 100644 llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-table.ll
 create mode 100644 llvm/test/Transforms/LoopVectorize/X86/gather-scatter-cost-table-decisions.ll
 create mode 100644 llvm/test/Transforms/SLPVectorizer/X86/gather-cost-znver4.ll

diff --git a/llvm/docs/ReleaseNotes.md b/llvm/docs/ReleaseNotes.md
index ef90a1f1f1c41..049db72db6e10 100644
--- a/llvm/docs/ReleaseNotes.md
+++ b/llvm/docs/ReleaseNotes.md
@@ -227,6 +227,11 @@ Makes programs 10x faster by doing Special New Thing.
 
 ### Changes to the X86 Backend
 
+* The cost model now uses per-shape masked gather/scatter costs measured on
+  AMD Zen hardware for `znver4`, `znver5`, and `znver6` (enabled via the
+  `prefer-gs-cost-table` tuning feature), replacing the flat AVX-512 overhead
+  for those targets.
+
 ### Changes to the OCaml bindings
 
 ### Changes to the Python bindings
diff --git a/llvm/lib/Target/X86/X86.td b/llvm/lib/Target/X86/X86.td
index 5786c659b0f98..cd043a3f5758e 100644
--- a/llvm/lib/Target/X86/X86.td
+++ b/llvm/lib/Target/X86/X86.td
@@ -804,6 +804,18 @@ def TuningFastGather
                        "Indicates if gather is reasonably fast (this is true for Skylake client and all AVX-512 CPUs)",
                        [], InlineIgnore>;
 
+// Use AMD Zen-tuned cost tables for masked gather/scatter intrinsics in the
+// X86 TargetTransformInfo cost model. Refines the flat overhead used by other
+// AVX-512 targets with per-element-type/per-VL costs measured on znver4 and
+// znver5. Inherited automatically by every znver4+ CPU via ZN4Tuning; not
+// applied to pre-AVX-512 Zen parts (znver1..3), which take the scalarise
+// path for masked gather anyway.
+def TuningPreferGSCostTable
+    : SubtargetFeature<"prefer-gs-cost-table",
+                       "HasPreferGSCostTable", "true",
+                       "Use per-shape gather/scatter cost tables in the cost model",
+                       [], InlineIgnore>;
+
 // Generate vpdpwssd instead of vpmaddwd+vpaddd sequence.
 def TuningFastDPWSSD
     : SubtargetFeature<
@@ -1737,7 +1749,8 @@ def ProcessorFeatures {
 
   list<SubtargetFeature> ZN4AdditionalTuning = [TuningFastDPWSSD,
                                                 TuningCOMPRESSFalseDeps,
-                                                TuningEXPANDFalseDeps];
+                                                TuningEXPANDFalseDeps,
+                                                TuningPreferGSCostTable];
   list<SubtargetFeature> ZN4Tuning =
     !listremove(!listconcat(ZN3Tuning, ZN4AdditionalTuning),
                 [TuningSlowVecMaskStore]);
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
index 4acd2d898e503..d6f68c9aa3993 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
@@ -57,6 +57,8 @@
 #include "llvm/CodeGen/TargetLowering.h"
 #include "llvm/IR/InstIterator.h"
 #include "llvm/IR/IntrinsicInst.h"
+#include "llvm/MC/MCSchedule.h"
+#include <cmath>
 #include <optional>
 
 using namespace llvm;
@@ -6588,11 +6590,163 @@ InstructionCost X86TTIImpl::getCFInstrCost(unsigned Opcode,
   return TTI::TCC_Free;
 }
 
+// Pick a representative masked gather/scatter opcode for a data shape, used to
+// read the body cost from the schedule model. Keyed on (NumElts, EltBits): the
+// integer and FP variants of a shape share a scheduling class, so the integer
+// form is returned. Index width is not part of the key; the matched-width form
+// is always used (dword index for 32-bit data, qword for 64-bit). Where a shape
+// has a separate class per index width the two agree to within a cycle, with
+// one exception: a v8i64 scatter is 14 on the qword-index class against 11 on
+// the dword-index one, so one reached through 32-bit indices is charged 3 too
+// many. The bare-pointer form, which is what the vectorizers produce, is the
+// qword-index one. Returns 0 for unsupported shapes (VF 5/6/7, or VF not in
+// {4,8,16}), whereupon the caller takes the generic fallback; the caller also
+// checks VLX, needed for the sub-512-bit forms.
+static unsigned getAVX512GSRepresentativeOpcode(bool IsLoad, MVT VT) {
+  if (!VT.isVector())
+    return 0;
+  unsigned NumElts = VT.getVectorNumElements();
+  unsigned EltBits = VT.getScalarSizeInBits();
+  if (IsLoad) {
+    if (EltBits == 32)
+      return NumElts == 4    ? X86::VPGATHERDDZ128rm
+             : NumElts == 8  ? X86::VPGATHERDDZ256rm
+             : NumElts == 16 ? X86::VPGATHERDDZrm
+                             : 0;
+    if (EltBits == 64)
+      return NumElts == 4   ? X86::VPGATHERQQZ256rm
+             : NumElts == 8 ? X86::VPGATHERQQZrm
+                            : 0;
+    return 0;
+  }
+  if (EltBits == 32)
+    return NumElts == 4    ? X86::VPSCATTERDDZ128mr
+           : NumElts == 8  ? X86::VPSCATTERDDZ256mr
+           : NumElts == 16 ? X86::VPSCATTERDDZmr
+                           : 0;
+  if (EltBits == 64)
+    return NumElts == 4   ? X86::VPSCATTERQQZ256mr
+           : NumElts == 8 ? X86::VPSCATTERQQZmr
+                          : 0;
+  return 0;
+}
+
+// Read an opcode's reciprocal-throughput body cost from this subtarget's
+// schedule model, or nullopt (caller falls back) when there is no
+// per-instruction model, the class is invalid or a variant (variants need a
+// real MachineInstr), or there is no real per-shape override.
+static std::optional<unsigned> getSchedModelGSBody(unsigned Opc,
+                                                   const X86Subtarget *ST) {
+  const MCSchedModel &SM = ST->getSchedModel();
+  if (!SM.hasInstrSchedModel())
+    return std::nullopt;
+
+  const auto *TII = ST->getInstrInfo();
+  auto ReciprocalThroughputOf = [&](unsigned Opcode) -> std::optional<double> {
+    unsigned SClassID = TII->get(Opcode).getSchedClass();
+    const MCSchedClassDesc *SCDesc = SM.getSchedClassDesc(SClassID);
+    if (!SCDesc || !SCDesc->isValid() || SCDesc->isVariant())
+      return std::nullopt;
+    return MCSchedModel::getReciprocalThroughput(*ST, *SCDesc);
+  };
+
+  std::optional<double> RThru = ReciprocalThroughputOf(Opc);
+  if (!RThru)
+    return std::nullopt;
+
+  // A valid class alone is not enough: an unmodelled op also has one, and
+  // would round to 0. The cheapest a real gather/scatter can be is a plain
+  // vector load, so anything at or below the load's throughput is the generic
+  // default in disguise. Reading the baseline from the model avoids a magic
+  // floor.
+  std::optional<double> LoadRThru = ReciprocalThroughputOf(X86::VMOVUPSZrm);
+  if (!LoadRThru || *RThru <= *LoadRThru)
+    return std::nullopt;
+
+  return static_cast<unsigned>(std::lround(*RThru));
+}
+
+// Hardware body cost of a single native-width masked gather/scatter, read live
+// from the schedule model; getZenGSOverhead adds the profitability premium.
+std::optional<unsigned>
+X86TTIImpl::getModeledGSInstrCost(bool IsLoad, Type *SrcVTy,
+                                  TTI::TargetCostKind CostKind) const {
+  if (CostKind != TTI::TCK_RecipThroughput || !ST->hasAVX512() ||
+      !ST->hasPreferGSCostTable() || !SrcVTy)
+    return std::nullopt;
+  EVT VT = TLI->getValueType(DL, SrcVTy);
+  if (!VT.isSimple())
+    return std::nullopt;
+  // The sub-512-bit (Z128/Z256) forms require AVX512VL; without it the opcodes
+  // do not exist, so fall back rather than read a meaningless class.
+  if (VT.getSizeInBits() < 512 && !ST->hasVLX())
+    return std::nullopt;
+  unsigned Opc = getAVX512GSRepresentativeOpcode(IsLoad, VT.getSimpleVT());
+  if (!Opc)
+    return std::nullopt;
+  return getSchedModelGSBody(Opc, ST);
+}
+
+// Per-shape vectorize-vs-scalarize OVERHEAD for AMD znver4+ gather/scatter
+// (TuningPreferGSCostTable): the calibrated break-even total, at which the
+// LoopVectorizer stops preferring the vectorised op, minus the schedule-model
+// body from getModeledGSInstrCost. Keeping the hardware half in
+// X86ScheduleZnver4.td leaves one source of truth and only the profitability
+// premium here. The premium belongs to the whole logical op and is paid once
+// however legalisation splits it; the caller scales by the split factor only
+// for genuinely wider-than-native shapes. Totals were calibrated on a Zen5
+// 9950X against a body term now measured on Znver4 (Ryzen 5 8645HS).
+//
+// Keyed by native shape (VF <= 16 for 32-bit, <= 8 for 64-bit).
+std::optional<unsigned> X86TTIImpl::getZenGSOverhead(bool IsLoad,
+                                                     Type *SrcVTy) const {
+  if (!ST->hasPreferGSCostTable() || !ST->hasAVX512() || !SrcVTy)
+    return std::nullopt;
+  // Each value places (overhead + body) on the intended side of the measured
+  // flip: below it where the masked op wins, above it where the scalarised
+  // lowering does (all i64 shapes, and v4i32/v4f32 scatter). Retargeting
+  // X86ScheduleZnver4.td to Znver4 grew the body, so every row was re-checked;
+  // the six with no headroom (v4f32/v4f64/v8f32/v8i32/v8f64 gather, v8f64
+  // scatter) were reduced by their body delta. The rest keep their value:
+  // moving them would assert precision the flip measurement does not have.
+  //
+  // i64 sits far above f64 (22 vs 15 at VF 8) by design: the scalarised i64
+  // alternative runs on the integer pipes, a much faster baseline than f64 on
+  // the FP pipes, so those rows are set above their flip to keep the
+  // vectoriser scalar (cf. llvm/llvm-project#198850).
+  static const CostTblEntry ZenGatherOverheadTable[] = {
+      {ISD::LOAD, MVT::v4i32, 7},   {ISD::LOAD, MVT::v8i32, 16},
+      {ISD::LOAD, MVT::v16i32, 17}, {ISD::LOAD, MVT::v4f32, 6},
+      {ISD::LOAD, MVT::v8f32, 16},  {ISD::LOAD, MVT::v16f32, 17},
+      {ISD::LOAD, MVT::v4f64, 6},   {ISD::LOAD, MVT::v8f64, 15},
+      {ISD::LOAD, MVT::v4i64, 10},  {ISD::LOAD, MVT::v8i64, 22},
+  };
+  static const CostTblEntry ZenScatterOverheadTable[] = {
+      {ISD::STORE, MVT::v4i32, 11}, {ISD::STORE, MVT::v8i32, 13},
+      {ISD::STORE, MVT::v16i32, 5}, {ISD::STORE, MVT::v4f32, 11},
+      {ISD::STORE, MVT::v8f32, 13}, {ISD::STORE, MVT::v16f32, 5},
+      {ISD::STORE, MVT::v4f64, 4},  {ISD::STORE, MVT::v8f64, 9},
+      {ISD::STORE, MVT::v4i64, 9},  {ISD::STORE, MVT::v8i64, 21},
+  };
+  // Any shape not in the table (e.g. the VF<4 forms the auto-vectoriser
+  // force-scalarises) returns nullopt for the flat-overhead path, not an error.
+  EVT VT = TLI->getValueType(DL, SrcVTy);
+  if (!VT.isSimple())
+    return std::nullopt;
+  const int ISDOpc = IsLoad ? ISD::LOAD : ISD::STORE;
+  ArrayRef<CostTblEntry> Table =
+      IsLoad ? ZenGatherOverheadTable : ZenScatterOverheadTable;
+  if (const auto *E = CostTableLookup(Table, ISDOpc, VT.getSimpleVT()))
+    return E->Cost;
+  return std::nullopt;
+}
+
 int X86TTIImpl::getGatherOverhead() const {
-  // Some CPUs have more overhead for gather. The specified overhead is relative
-  // to the Load operation. "2" is the number provided by Intel architects. This
-  // parameter is used for cost estimation of Gather Op and comparison with
-  // other alternatives.
+  // Flat, shape-independent gather break-even overhead relative to a plain
+  // Load. "2" is the number provided by Intel architects; the per-shape AMD
+  // znver4+ break-even lives in getZenGSOverhead + getModeledGSInstrCost
+  // instead. The hasAVX512() guard is symmetric with getScatterOverhead:
+  // without it a VEX gather would be charged an AVX-512-measured overhead.
   // TODO: Remove the explicit hasAVX512()?, That would mean we would only
   // enable gather with a -march.
   if (ST->hasAVX512() || (ST->hasAVX2() && ST->hasFastGather()))
@@ -6602,6 +6756,7 @@ int X86TTIImpl::getGatherOverhead() const {
 }
 
 int X86TTIImpl::getScatterOverhead() const {
+  // Flat, shape-independent scatter break-even overhead; see getGatherOverhead.
   if (ST->hasAVX512())
     return 2;
 
@@ -6661,22 +6816,43 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
       std::max(IdxsLT.first, SrcLT.first).getValue();
   const bool IsLoad = Opcode == Instruction::Load;
 
-  // If a wide gather/scatter splits (e.g. a 16-wide op whose 64-bit indices
-  // don't fit in one zmm) the hardware issues SplitFactor sub-ops; for code
-  // size that is simply the sub-op count.
+  // A split op (e.g. 16-wide with 64-bit indices that don't fit one zmm)
+  // issues SplitFactor sub-ops; for code size that is just the sub-op count.
   if (CostKind == TTI::TCK_CodeSize)
     return SplitFactor;
 
-  // A masked gather/scatter cost is a flat vectorize-vs-scalarize break-even
-  // overhead plus the cost of the instruction body:
-  //
-  //   cost = break-even overhead + instruction body cost
-  //
-  // The flat overhead is a property of the original, pre-split shape and is
-  // paid once regardless of how many sub-ops the op splits into (previously it
-  // was charged once per sub-op via the recursive split, over-counting wide
-  // shapes). Only the body scales with the split factor. The body is computed
-  // once here so the alignment/address-space inputs cannot diverge.
+  // AMD znver4+ (TuningPreferGSCostTable): overhead + schedule-model body, both
+  // keyed on the *native* shape, so a shape fitting one register is charged
+  // once however legalisation splits it (narrow- and wide-index forms then
+  // report alike). Wider-than-native shapes tile into GSMul native ops.
+  if (CostKind == TTI::TCK_RecipThroughput && ST->hasPreferGSCostTable() &&
+      ST->hasAVX512()) {
+    Type *EltTy = SrcVTy->getScalarType();
+    unsigned EltBits = EltTy->getPrimitiveSizeInBits();
+    if (EltBits == 32 || EltBits == 64) {
+      unsigned NativeMaxElts = 512 / EltBits;
+      // Split only for genuinely-wider-than-native shapes, and only when the
+      // width is an exact multiple so the native op tiles cleanly.
+      unsigned NativeVF = std::min(VF, NativeMaxElts);
+      if (VF % NativeVF == 0) {
+        unsigned GSMul = VF / NativeVF;
+        auto *NativeVTy = FixedVectorType::get(EltTy, NativeVF);
+        std::optional<unsigned> Overhead = getZenGSOverhead(IsLoad, NativeVTy);
+        std::optional<unsigned> Body =
+            getModeledGSInstrCost(IsLoad, NativeVTy, CostKind);
+        if (Overhead && Body)
+          return GSMul * (*Overhead + *Body);
+      }
+    }
+  }
+
+  // Otherwise (a non-enumerated VF, or a target without
+  // TuningPreferGSCostTable): flat break-even overhead plus body. The overhead
+  // belongs to the pre-split shape and is paid once; only the body
+  // (VF * scalar-memory-op) scales with the split factor. Not guaranteed
+  // monotonic against the calibrated table - v6i32 can read cheaper than
+  // v4i32's measured total - but such VFs are rare and the table above remains
+  // the source of truth.
   const int GSOverhead = IsLoad ? getGatherOverhead() : getScatterOverhead();
 
   auto BodyCostOf = [&](Type *VTy) -> InstructionCost {
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.h b/llvm/lib/Target/X86/X86TargetTransformInfo.h
index 422e4aad2316c..e97ca07a437c9 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.h
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.h
@@ -275,6 +275,15 @@ class X86TTIImpl final : public BasicTTIImplBase<X86TTIImpl> {
 
   int getGatherOverhead() const;
   int getScatterOverhead() const;
+  // Per-shape vectorize-vs-scalarize overhead for AMD znver4+ gather/scatter:
+  // the calibrated break-even TOTAL minus the schedule-model body, so that
+  // getZenGSOverhead + getModeledGSInstrCost reproduces the measured total.
+  std::optional<unsigned> getZenGSOverhead(bool IsLoad, Type *SrcVTy) const;
+  // Hardware body cost of a single (non-split) masked gather/scatter, read
+  // live from this subtarget's schedule model, or nullopt to fall back.
+  std::optional<unsigned>
+  getModeledGSInstrCost(bool IsLoad, Type *SrcVTy,
+                        TTI::TargetCostKind CostKind) const;
 
   /// @}
 };
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-table.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-table.ll
new file mode 100644
index 0000000000000..3f638adaaf76f
--- /dev/null
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-table.ll
@@ -0,0 +1,1078 @@
+; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py
+; Cost-model coverage for the per-shape masked gather/scatter overheads.
+;
+; ZNVER4 / ZNVER5 carry TuningPreferGSCostTable and have AVX-512, so
+; getGSVectorCost reads the measured hardware cost from the schedule model
+; (which stays honest for llvm-mca) and adds the per-shape break-even overhead
+; from the TTI tables on top.
+; ZNVER3 lacks AVX-512, so masked gather is scalarised; its numbers are the
+; unchanged scalar fallback, kept only to show no regression on pre-AVX-512 Zen.
+; SKX is a non-Zen AVX-512 baseline: not gated into the schedule-model path, so
+; it keeps the generic flat overhead and its numbers are unchanged.
+;
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=throughput -mcpu=znver4 | FileCheck %s --check-prefix=ZNVER4
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=throughput -mcpu=znver5 | FileCheck %s --check-prefix=ZNVER5
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=throughput -mcpu=znver6 | FileCheck %s --check-prefix=ZNVER6
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=throughput -mcpu=znver3 | FileCheck %s --check-prefix=ZNVER3
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=throughput -mcpu=skx        | FileCheck %s --check-prefix=SKX
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=throughput -mcpu=x86-64-v4 | FileCheck %s --check-prefix=X8664V4
+
+;------------------------------------------------------------------------------
+; Masked gather - i32 element type
+;------------------------------------------------------------------------------
+
+define <2 x i32> @gather_v2i32(<2 x ptr> %ptrs, <2 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v2i32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 8 for instruction: %v = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 4 %ptrs, <2 x i1> %mask, <2 x i32> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x i32> %v
+;
+; ZNVER5-LABEL: 'gather_v2i32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 8 for instruction: %v = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 4 %ptrs, <2 x i1> %mask, <2 x i32> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x i32> %v
+;
+; ZNVER6-LABEL: 'gather_v2i32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 8 for instruction: %v = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 4 %ptrs, <2 x i1> %mask, <2 x i32> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x i32> %v
+;
+; ZNVER3-LABEL: 'gather_v2i32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 7 for instruction: %v = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 4 %ptrs, <2 x i1> %mask, <2 x i32> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x i32> %v
+;
+; SKX-LABEL: 'gather_v2i32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 8 for instruction: %v = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 4 %ptrs, <2 x i1> %mask, <2 x i32> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x i32> %v
+;
+; X8664V4-LABEL: 'gather_v2i32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 8 for instruction: %v = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 4 %ptrs, <2 x i1> %mask, <2 x i32> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x i32> %v
+;
+  %v = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> %ptrs, i32 4, <2 x i1> %mask, <2 x i32> poison)
+  ret <2 x i32> %v
+}
+
+define <4 x i32> @gather_v4i32(<4 x ptr> %ptrs, <4 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v4i32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 12 for instruction: %v = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 4 %ptrs, <4 x i1> %mask, <4 x i32> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x i32> %v
+;
+; ZNVER5-LABEL: 'gather_v4i32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 12 for instruction: %v = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 4 %ptrs, <4 x i1> %mask, <4 x i32> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x i32> %v
+;
+; ZNVER6-LABEL: 'gather_v4i32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 12 for instruction: %v = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 4 %ptrs, <4 x i1> %mask, <4 x i32> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x i32> %v
+;
+; ZNVER3-LABEL: 'gather_v4i32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 14 for instruction: %v = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 4 %ptrs, <4 x i1> %mask, <4 x i32> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x i32> %v
+;
+; SKX-LABEL: 'gather_v4i32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: %v = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 4 %ptrs, <4 x i1> %mask, <4 x i32> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x i32> %v
+;
+; X8664V4-LABEL: 'gather_v4i32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: %v = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 4 %ptrs, <4 x i1> %mask, <4 x i32> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x i32> %v
+;
+  %v = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> %ptrs, i32 4, <4 x i1> %mask, <4 x i32> poison)
+  ret <4 x i32> %v
+}
+
+define <8 x i32> @gather_v8i32(<8 x ptr> %ptrs, <8 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v8i32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x i32> %v
+;
+; ZNVER5-LABEL: 'gather_v8i32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x i32> %v
+;
+; ZNVER6-LABEL: 'gather_v8i32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x i32> %v
+;
+; ZNVER3-LABEL: 'gather_v8i32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 28 for instruction: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x i32> %v
+;
+; SKX-LABEL: 'gather_v8i32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x i32> %v
+;
+; X8664V4-LABEL: 'gather_v8i32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x i32> %v
+;
+  %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x i32> poison)
+  ret <8 x i32> %v
+}
+
+define <16 x i32> @gather_v16i32(<16 x ptr> %ptrs, <16 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v16i32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
+;
+; ZNVER5-LABEL: 'gather_v16i32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
+;
+; ZNVER6-LABEL: 'gather_v16i32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
+;
+; ZNVER3-LABEL: 'gather_v16i32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 55 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
+;
+; SKX-LABEL: 'gather_v16i32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
+;
+; X8664V4-LABEL: 'gather_v16i32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
+;
+  %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+  ret <16 x i32> %v
+}
+
+;------------------------------------------------------------------------------
+; Masked gather - i64 element type
+;------------------------------------------------------------------------------
+
+define <2 x i64> @gather_v2i64(<2 x ptr> %ptrs, <2 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v2i64'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 8 for instruction: %v = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 8 %ptrs, <2 x i1> %mask, <2 x i64> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x i64> %v
+;
+; ZNVER5-LABEL: 'gather_v2i64'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 8 for instruction: %v = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 8 %ptrs, <2 x i1> %mask, <2 x i64> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x i64> %v
+;
+; ZNVER6-LABEL: 'gather_v2i64'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 8 for instruction: %v = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 8 %ptrs, <2 x i1> %mask, <2 x i64> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x i64> %v
+;
+; ZNVER3-LABEL: 'gather_v2i64'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 7 for instruction: %v = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 8 %ptrs, <2 x i1> %mask, <2 x i64> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x i64> %v
+;
+; SKX-LABEL: 'gather_v2i64'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 8 for instruction: %v = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 8 %ptrs, <2 x i1> %mask, <2 x i64> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x i64> %v
+;
+; X8664V4-LABEL: 'gather_v2i64'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 8 for instruction: %v = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 8 %ptrs, <2 x i1> %mask, <2 x i64> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x i64> %v
+;
+  %v = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> %ptrs, i32 8, <2 x i1> %mask, <2 x i64> poison)
+  ret <2 x i64> %v
+}
+
+define <4 x i64> @gather_v4i64(<4 x ptr> %ptrs, <4 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v4i64'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 15 for instruction: %v = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x i64> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x i64> %v
+;
+; ZNVER5-LABEL: 'gather_v4i64'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 15 for instruction: %v = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x i64> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x i64> %v
+;
+; ZNVER6-LABEL: 'gather_v4i64'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 15 for instruction: %v = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x i64> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x i64> %v
+;
+; ZNVER3-LABEL: 'gather_v4i64'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 15 for instruction: %v = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x i64> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x i64> %v
+;
+; SKX-LABEL: 'gather_v4i64'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: %v = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x i64> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x i64> %v
+;
+; X8664V4-LABEL: 'gather_v4i64'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: %v = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x i64> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x i64> %v
+;
+  %v = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> %ptrs, i32 8, <4 x i1> %mask, <4 x i64> poison)
+  ret <4 x i64> %v
+}
+
+define <8 x i64> @gather_v8i64(<8 x ptr> %ptrs, <8 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v8i64'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 32 for instruction: %v = call <8 x i64> @llvm.masked.gather.v8i64.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x i64> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x i64> %v
+;
+; ZNVER5-LABEL: 'gather_v8i64'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 32 for instruction: %v = call <8 x i64> @llvm.masked.gather.v8i64.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x i64> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x i64> %v
+;
+; ZNVER6-LABEL: 'gather_v8i64'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 32 for instruction: %v = call <8 x i64> @llvm.masked.gather.v8i64.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x i64> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x i64> %v
+;
+; ZNVER3-LABEL: 'gather_v8i64'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 29 for instruction: %v = call <8 x i64> @llvm.masked.gather.v8i64.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x i64> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x i64> %v
+;
+; SKX-LABEL: 'gather_v8i64'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: %v = call <8 x i64> @llvm.masked.gather.v8i64.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x i64> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x i64> %v
+;
+; X8664V4-LABEL: 'gather_v8i64'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: %v = call <8 x i64> @llvm.masked.gather.v8i64.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x i64> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x i64> %v
+;
+  %v = call <8 x i64> @llvm.masked.gather.v8i64.v8p0(<8 x ptr> %ptrs, i32 8, <8 x i1> %mask, <8 x i64> poison)
+  ret <8 x i64> %v
+}
+
+;------------------------------------------------------------------------------
+; Masked gather - f32 element type
+;------------------------------------------------------------------------------
+
+define <2 x float> @gather_v2f32(<2 x ptr> %ptrs, <2 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v2f32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 7 for instruction: %v = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 4 %ptrs, <2 x i1> %mask, <2 x float> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x float> %v
+;
+; ZNVER5-LABEL: 'gather_v2f32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 7 for instruction: %v = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 4 %ptrs, <2 x i1> %mask, <2 x float> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x float> %v
+;
+; ZNVER6-LABEL: 'gather_v2f32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 7 for instruction: %v = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 4 %ptrs, <2 x i1> %mask, <2 x float> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x float> %v
+;
+; ZNVER3-LABEL: 'gather_v2f32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: %v = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 4 %ptrs, <2 x i1> %mask, <2 x float> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x float> %v
+;
+; SKX-LABEL: 'gather_v2f32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 7 for instruction: %v = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 4 %ptrs, <2 x i1> %mask, <2 x float> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x float> %v
+;
+; X8664V4-LABEL: 'gather_v2f32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 7 for instruction: %v = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 4 %ptrs, <2 x i1> %mask, <2 x float> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x float> %v
+;
+  %v = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> %ptrs, i32 4, <2 x i1> %mask, <2 x float> poison)
+  ret <2 x float> %v
+}
+
+define <4 x float> @gather_v4f32(<4 x ptr> %ptrs, <4 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v4f32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 11 for instruction: %v = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 4 %ptrs, <4 x i1> %mask, <4 x float> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x float> %v
+;
+; ZNVER5-LABEL: 'gather_v4f32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 11 for instruction: %v = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 4 %ptrs, <4 x i1> %mask, <4 x float> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x float> %v
+;
+; ZNVER6-LABEL: 'gather_v4f32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 11 for instruction: %v = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 4 %ptrs, <4 x i1> %mask, <4 x float> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x float> %v
+;
+; ZNVER3-LABEL: 'gather_v4f32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 13 for instruction: %v = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 4 %ptrs, <4 x i1> %mask, <4 x float> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x float> %v
+;
+; SKX-LABEL: 'gather_v4f32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: %v = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 4 %ptrs, <4 x i1> %mask, <4 x float> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x float> %v
+;
+; X8664V4-LABEL: 'gather_v4f32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: %v = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 4 %ptrs, <4 x i1> %mask, <4 x float> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x float> %v
+;
+  %v = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> %ptrs, i32 4, <4 x i1> %mask, <4 x float> poison)
+  ret <4 x float> %v
+}
+
+define <8 x float> @gather_v8f32(<8 x ptr> %ptrs, <8 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v8f32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x float> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x float> %v
+;
+; ZNVER5-LABEL: 'gather_v8f32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x float> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x float> %v
+;
+; ZNVER6-LABEL: 'gather_v8f32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x float> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x float> %v
+;
+; ZNVER3-LABEL: 'gather_v8f32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 26 for instruction: %v = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x float> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x float> %v
+;
+; SKX-LABEL: 'gather_v8f32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: %v = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x float> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x float> %v
+;
+; X8664V4-LABEL: 'gather_v8f32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: %v = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x float> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x float> %v
+;
+  %v = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x float> poison)
+  ret <8 x float> %v
+}
+
+define <16 x float> @gather_v16f32(<16 x ptr> %ptrs, <16 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v16f32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
+;
+; ZNVER5-LABEL: 'gather_v16f32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
+;
+; ZNVER6-LABEL: 'gather_v16f32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
+;
+; ZNVER3-LABEL: 'gather_v16f32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 51 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
+;
+; SKX-LABEL: 'gather_v16f32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
+;
+; X8664V4-LABEL: 'gather_v16f32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
+;
+  %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x float> poison)
+  ret <16 x float> %v
+}
+
+;------------------------------------------------------------------------------
+; Masked gather - f64 element type
+;------------------------------------------------------------------------------
+
+define <2 x double> @gather_v2f64(<2 x ptr> %ptrs, <2 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v2f64'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 7 for instruction: %v = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 8 %ptrs, <2 x i1> %mask, <2 x double> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x double> %v
+;
+; ZNVER5-LABEL: 'gather_v2f64'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 7 for instruction: %v = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 8 %ptrs, <2 x i1> %mask, <2 x double> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x double> %v
+;
+; ZNVER6-LABEL: 'gather_v2f64'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 7 for instruction: %v = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 8 %ptrs, <2 x i1> %mask, <2 x double> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x double> %v
+;
+; ZNVER3-LABEL: 'gather_v2f64'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: %v = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 8 %ptrs, <2 x i1> %mask, <2 x double> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x double> %v
+;
+; SKX-LABEL: 'gather_v2f64'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 7 for instruction: %v = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 8 %ptrs, <2 x i1> %mask, <2 x double> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x double> %v
+;
+; X8664V4-LABEL: 'gather_v2f64'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 7 for instruction: %v = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 8 %ptrs, <2 x i1> %mask, <2 x double> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <2 x double> %v
+;
+  %v = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> %ptrs, i32 8, <2 x i1> %mask, <2 x double> poison)
+  ret <2 x double> %v
+}
+
+define <4 x double> @gather_v4f64(<4 x ptr> %ptrs, <4 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v4f64'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 11 for instruction: %v = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x double> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x double> %v
+;
+; ZNVER5-LABEL: 'gather_v4f64'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 11 for instruction: %v = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x double> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x double> %v
+;
+; ZNVER6-LABEL: 'gather_v4f64'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 11 for instruction: %v = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x double> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x double> %v
+;
+; ZNVER3-LABEL: 'gather_v4f64'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 13 for instruction: %v = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x double> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x double> %v
+;
+; SKX-LABEL: 'gather_v4f64'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: %v = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x double> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x double> %v
+;
+; X8664V4-LABEL: 'gather_v4f64'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: %v = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x double> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x double> %v
+;
+  %v = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> %ptrs, i32 8, <4 x i1> %mask, <4 x double> poison)
+  ret <4 x double> %v
+}
+
+define <8 x double> @gather_v8f64(<8 x ptr> %ptrs, <8 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v8f64'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x double> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x double> %v
+;
+; ZNVER5-LABEL: 'gather_v8f64'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x double> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x double> %v
+;
+; ZNVER6-LABEL: 'gather_v8f64'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x double> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x double> %v
+;
+; ZNVER3-LABEL: 'gather_v8f64'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x double> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x double> %v
+;
+; SKX-LABEL: 'gather_v8f64'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: %v = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x double> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x double> %v
+;
+; X8664V4-LABEL: 'gather_v8f64'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: %v = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x double> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x double> %v
+;
+  %v = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> %ptrs, i32 8, <8 x i1> %mask, <8 x double> poison)
+  ret <8 x double> %v
+}
+
+;------------------------------------------------------------------------------
+; Masked scatter - i32 element type
+;------------------------------------------------------------------------------
+
+define void @scatter_v4i32(<4 x i32> %src, <4 x ptr> %ptrs, <4 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v4i32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 19 for instruction: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v4i32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 19 for instruction: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v4i32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 19 for instruction: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v4i32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 14 for instruction: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v4i32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v4i32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> %src, <4 x ptr> %ptrs, i32 4, <4 x i1> %mask)
+  ret void
+}
+
+define void @scatter_v8i32(<8 x i32> %src, <8 x ptr> %ptrs, <8 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v8i32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> %src, <8 x ptr> align 4 %ptrs, <8 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v8i32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> %src, <8 x ptr> align 4 %ptrs, <8 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v8i32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> %src, <8 x ptr> align 4 %ptrs, <8 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v8i32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 28 for instruction: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> %src, <8 x ptr> align 4 %ptrs, <8 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v8i32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> %src, <8 x ptr> align 4 %ptrs, <8 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v8i32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> %src, <8 x ptr> align 4 %ptrs, <8 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> %src, <8 x ptr> %ptrs, i32 4, <8 x i1> %mask)
+  ret void
+}
+
+define void @scatter_v16i32(<16 x i32> %src, <16 x ptr> %ptrs, <16 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v16i32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v16i32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v16i32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v16i32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 55 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v16i32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v16i32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> %ptrs, i32 4, <16 x i1> %mask)
+  ret void
+}
+
+;------------------------------------------------------------------------------
+; Masked scatter - i64 element type
+;------------------------------------------------------------------------------
+
+define void @scatter_v4i64(<4 x i64> %src, <4 x ptr> %ptrs, <4 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v4i64'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 16 for instruction: call void @llvm.masked.scatter.v4i64.v4p0(<4 x i64> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v4i64'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 16 for instruction: call void @llvm.masked.scatter.v4i64.v4p0(<4 x i64> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v4i64'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 16 for instruction: call void @llvm.masked.scatter.v4i64.v4p0(<4 x i64> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v4i64'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 15 for instruction: call void @llvm.masked.scatter.v4i64.v4p0(<4 x i64> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v4i64'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: call void @llvm.masked.scatter.v4i64.v4p0(<4 x i64> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v4i64'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: call void @llvm.masked.scatter.v4i64.v4p0(<4 x i64> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  call void @llvm.masked.scatter.v4i64.v4p0(<4 x i64> %src, <4 x ptr> %ptrs, i32 8, <4 x i1> %mask)
+  ret void
+}
+
+define void @scatter_v8i64(<8 x i64> %src, <8 x ptr> %ptrs, <8 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v8i64'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: call void @llvm.masked.scatter.v8i64.v8p0(<8 x i64> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v8i64'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: call void @llvm.masked.scatter.v8i64.v8p0(<8 x i64> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v8i64'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: call void @llvm.masked.scatter.v8i64.v8p0(<8 x i64> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v8i64'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 29 for instruction: call void @llvm.masked.scatter.v8i64.v8p0(<8 x i64> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v8i64'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: call void @llvm.masked.scatter.v8i64.v8p0(<8 x i64> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v8i64'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: call void @llvm.masked.scatter.v8i64.v8p0(<8 x i64> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  call void @llvm.masked.scatter.v8i64.v8p0(<8 x i64> %src, <8 x ptr> %ptrs, i32 8, <8 x i1> %mask)
+  ret void
+}
+
+;------------------------------------------------------------------------------
+; Masked scatter - f32 element type
+;------------------------------------------------------------------------------
+
+define void @scatter_v4f32(<4 x float> %src, <4 x ptr> %ptrs, <4 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v4f32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 19 for instruction: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v4f32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 19 for instruction: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v4f32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 19 for instruction: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v4f32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 13 for instruction: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v4f32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v4f32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> %src, <4 x ptr> %ptrs, i32 4, <4 x i1> %mask)
+  ret void
+}
+
+define void @scatter_v8f32(<8 x float> %src, <8 x ptr> %ptrs, <8 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v8f32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: call void @llvm.masked.scatter.v8f32.v8p0(<8 x float> %src, <8 x ptr> align 4 %ptrs, <8 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v8f32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: call void @llvm.masked.scatter.v8f32.v8p0(<8 x float> %src, <8 x ptr> align 4 %ptrs, <8 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v8f32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: call void @llvm.masked.scatter.v8f32.v8p0(<8 x float> %src, <8 x ptr> align 4 %ptrs, <8 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v8f32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 26 for instruction: call void @llvm.masked.scatter.v8f32.v8p0(<8 x float> %src, <8 x ptr> align 4 %ptrs, <8 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v8f32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: call void @llvm.masked.scatter.v8f32.v8p0(<8 x float> %src, <8 x ptr> align 4 %ptrs, <8 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v8f32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: call void @llvm.masked.scatter.v8f32.v8p0(<8 x float> %src, <8 x ptr> align 4 %ptrs, <8 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  call void @llvm.masked.scatter.v8f32.v8p0(<8 x float> %src, <8 x ptr> %ptrs, i32 4, <8 x i1> %mask)
+  ret void
+}
+
+define void @scatter_v16f32(<16 x float> %src, <16 x ptr> %ptrs, <16 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v16f32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v16f32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v16f32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v16f32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 51 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v16f32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v16f32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> %ptrs, i32 4, <16 x i1> %mask)
+  ret void
+}
+
+;------------------------------------------------------------------------------
+; Masked scatter - f64 element type
+;------------------------------------------------------------------------------
+
+define void @scatter_v4f64(<4 x double> %src, <4 x ptr> %ptrs, <4 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v4f64'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 11 for instruction: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v4f64'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 11 for instruction: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v4f64'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 11 for instruction: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v4f64'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 13 for instruction: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v4f64'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v4f64'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> %src, <4 x ptr> %ptrs, i32 8, <4 x i1> %mask)
+  ret void
+}
+
+define void @scatter_v8f64(<8 x double> %src, <8 x ptr> %ptrs, <8 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v8f64'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 23 for instruction: call void @llvm.masked.scatter.v8f64.v8p0(<8 x double> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v8f64'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 23 for instruction: call void @llvm.masked.scatter.v8f64.v8p0(<8 x double> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v8f64'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 23 for instruction: call void @llvm.masked.scatter.v8f64.v8p0(<8 x double> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v8f64'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: call void @llvm.masked.scatter.v8f64.v8p0(<8 x double> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v8f64'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: call void @llvm.masked.scatter.v8f64.v8p0(<8 x double> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v8f64'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: call void @llvm.masked.scatter.v8f64.v8p0(<8 x double> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  call void @llvm.masked.scatter.v8f64.v8p0(<8 x double> %src, <8 x ptr> %ptrs, i32 8, <8 x i1> %mask)
+  ret void
+}
+
+;------------------------------------------------------------------------------
+; 16-wide gather/scatter with a single base + 32-bit index GEP.
+;
+; The bare <16 x ptr> forms above use 64-bit indices that don't fit in one zmm,
+; so the cost model splits them into two 8-wide sub-ops. These GEP forms carry a
+; 32-bit index off a common base, which getGSVectorCost reduces to a 32-bit
+; index so the op stays 16-wide and exercises the v16i32/v16f32 table rows
+; directly (rather than the v8 rows twice).
+;------------------------------------------------------------------------------
+
+define <16 x i32> @gather_v16i32_gep(ptr %base, <16 x i32> %idx, <16 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v16i32_gep'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds i32, ptr %base, <16 x i32> %idx
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
+;
+; ZNVER5-LABEL: 'gather_v16i32_gep'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds i32, ptr %base, <16 x i32> %idx
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
+;
+; ZNVER6-LABEL: 'gather_v16i32_gep'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds i32, ptr %base, <16 x i32> %idx
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
+;
+; ZNVER3-LABEL: 'gather_v16i32_gep'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds i32, ptr %base, <16 x i32> %idx
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 55 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
+;
+; SKX-LABEL: 'gather_v16i32_gep'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds i32, ptr %base, <16 x i32> %idx
+; SKX-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
+;
+; X8664V4-LABEL: 'gather_v16i32_gep'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds i32, ptr %base, <16 x i32> %idx
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
+;
+  %ptrs = getelementptr inbounds i32, ptr %base, <16 x i32> %idx
+  %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+  ret <16 x i32> %v
+}
+
+define <16 x float> @gather_v16f32_gep(ptr %base, <16 x i32> %idx, <16 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v16f32_gep'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds float, ptr %base, <16 x i32> %idx
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
+;
+; ZNVER5-LABEL: 'gather_v16f32_gep'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds float, ptr %base, <16 x i32> %idx
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
+;
+; ZNVER6-LABEL: 'gather_v16f32_gep'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds float, ptr %base, <16 x i32> %idx
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
+;
+; ZNVER3-LABEL: 'gather_v16f32_gep'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds float, ptr %base, <16 x i32> %idx
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 51 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
+;
+; SKX-LABEL: 'gather_v16f32_gep'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds float, ptr %base, <16 x i32> %idx
+; SKX-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
+;
+; X8664V4-LABEL: 'gather_v16f32_gep'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds float, ptr %base, <16 x i32> %idx
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
+;
+  %ptrs = getelementptr inbounds float, ptr %base, <16 x i32> %idx
+  %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x float> poison)
+  ret <16 x float> %v
+}
+
+define void @scatter_v16i32_gep(<16 x i32> %src, ptr %base, <16 x i32> %idx, <16 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v16i32_gep'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds i32, ptr %base, <16 x i32> %idx
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v16i32_gep'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds i32, ptr %base, <16 x i32> %idx
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v16i32_gep'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds i32, ptr %base, <16 x i32> %idx
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v16i32_gep'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds i32, ptr %base, <16 x i32> %idx
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 55 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v16i32_gep'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds i32, ptr %base, <16 x i32> %idx
+; SKX-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v16i32_gep'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds i32, ptr %base, <16 x i32> %idx
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  %ptrs = getelementptr inbounds i32, ptr %base, <16 x i32> %idx
+  call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> %ptrs, i32 4, <16 x i1> %mask)
+  ret void
+}
+
+define void @scatter_v16f32_gep(<16 x float> %src, ptr %base, <16 x i32> %idx, <16 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v16f32_gep'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds float, ptr %base, <16 x i32> %idx
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v16f32_gep'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds float, ptr %base, <16 x i32> %idx
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v16f32_gep'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds float, ptr %base, <16 x i32> %idx
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v16f32_gep'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds float, ptr %base, <16 x i32> %idx
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 51 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v16f32_gep'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds float, ptr %base, <16 x i32> %idx
+; SKX-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v16f32_gep'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: %ptrs = getelementptr inbounds float, ptr %base, <16 x i32> %idx
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  %ptrs = getelementptr inbounds float, ptr %base, <16 x i32> %idx
+  call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> %ptrs, i32 4, <16 x i1> %mask)
+  ret void
+}
+
+; Wider-than-native shapes: the cost must stay monotonic in VF (a VF32 gather is
+; measured at exactly 2x a VF16 on Zen5, VF64 at 4x), so these must not dip
+; below the widest in-table shape. Guards against the VF32-cheaper-than-VF16
+; regression that a flat out-of-table overhead would reintroduce.
+define <32 x i32> @gather_v32i32(<32 x ptr> %ptrs, <32 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v32i32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 70 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i32> %v
+;
+; ZNVER5-LABEL: 'gather_v32i32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 70 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i32> %v
+;
+; ZNVER6-LABEL: 'gather_v32i32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 70 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i32> %v
+;
+; ZNVER3-LABEL: 'gather_v32i32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 109 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i32> %v
+;
+; SKX-LABEL: 'gather_v32i32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 34 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i32> %v
+;
+; X8664V4-LABEL: 'gather_v32i32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 34 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i32> %v
+;
+  %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> %ptrs, i32 4, <32 x i1> %mask, <32 x i32> poison)
+  ret <32 x i32> %v
+}
+
+define <64 x i32> @gather_v64i32(<64 x ptr> %ptrs, <64 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v64i32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 140 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <64 x i32> %v
+;
+; ZNVER5-LABEL: 'gather_v64i32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 140 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <64 x i32> %v
+;
+; ZNVER6-LABEL: 'gather_v64i32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 140 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <64 x i32> %v
+;
+; ZNVER3-LABEL: 'gather_v64i32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 218 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <64 x i32> %v
+;
+; SKX-LABEL: 'gather_v64i32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 66 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <64 x i32> %v
+;
+; X8664V4-LABEL: 'gather_v64i32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 66 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <64 x i32> %v
+;
+  %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> %ptrs, i32 4, <64 x i1> %mask, <64 x i32> poison)
+  ret <64 x i32> %v
+}
+
+define void @scatter_v32i32(<32 x i32> %src, <32 x ptr> %ptrs, <32 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v32i32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 62 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v32i32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 62 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v32i32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 62 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v32i32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 109 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v32i32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 34 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v32i32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 34 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> %ptrs, i32 4, <32 x i1> %mask)
+  ret void
+}
+
+define void @scatter_v64i32(<64 x i32> %src, <64 x ptr> %ptrs, <64 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v64i32'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 124 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v64i32'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 124 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v64i32'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 124 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v64i32'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 218 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v64i32'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 66 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v64i32'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 66 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> %ptrs, i32 4, <64 x i1> %mask)
+  ret void
+}
+
+define <16 x i64> @gather_v16i64(<16 x ptr> %ptrs, <16 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v16i64'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 64 for instruction: %v = call <16 x i64> @llvm.masked.gather.v16i64.v16p0(<16 x ptr> align 8 %ptrs, <16 x i1> %mask, <16 x i64> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i64> %v
+;
+; ZNVER5-LABEL: 'gather_v16i64'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 64 for instruction: %v = call <16 x i64> @llvm.masked.gather.v16i64.v16p0(<16 x ptr> align 8 %ptrs, <16 x i1> %mask, <16 x i64> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i64> %v
+;
+; ZNVER6-LABEL: 'gather_v16i64'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 64 for instruction: %v = call <16 x i64> @llvm.masked.gather.v16i64.v16p0(<16 x ptr> align 8 %ptrs, <16 x i1> %mask, <16 x i64> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i64> %v
+;
+; ZNVER3-LABEL: 'gather_v16i64'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 57 for instruction: %v = call <16 x i64> @llvm.masked.gather.v16i64.v16p0(<16 x ptr> align 8 %ptrs, <16 x i1> %mask, <16 x i64> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i64> %v
+;
+; SKX-LABEL: 'gather_v16i64'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x i64> @llvm.masked.gather.v16i64.v16p0(<16 x ptr> align 8 %ptrs, <16 x i1> %mask, <16 x i64> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i64> %v
+;
+; X8664V4-LABEL: 'gather_v16i64'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x i64> @llvm.masked.gather.v16i64.v16p0(<16 x ptr> align 8 %ptrs, <16 x i1> %mask, <16 x i64> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i64> %v
+;
+  %v = call <16 x i64> @llvm.masked.gather.v16i64.v16p0(<16 x ptr> %ptrs, i32 8, <16 x i1> %mask, <16 x i64> poison)
+  ret <16 x i64> %v
+}
+
+define <32 x i64> @gather_v32i64(<32 x ptr> %ptrs, <32 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v32i64'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 128 for instruction: %v = call <32 x i64> @llvm.masked.gather.v32i64.v32p0(<32 x ptr> align 8 %ptrs, <32 x i1> %mask, <32 x i64> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i64> %v
+;
+; ZNVER5-LABEL: 'gather_v32i64'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 128 for instruction: %v = call <32 x i64> @llvm.masked.gather.v32i64.v32p0(<32 x ptr> align 8 %ptrs, <32 x i1> %mask, <32 x i64> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i64> %v
+;
+; ZNVER6-LABEL: 'gather_v32i64'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 128 for instruction: %v = call <32 x i64> @llvm.masked.gather.v32i64.v32p0(<32 x ptr> align 8 %ptrs, <32 x i1> %mask, <32 x i64> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i64> %v
+;
+; ZNVER3-LABEL: 'gather_v32i64'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 113 for instruction: %v = call <32 x i64> @llvm.masked.gather.v32i64.v32p0(<32 x ptr> align 8 %ptrs, <32 x i1> %mask, <32 x i64> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i64> %v
+;
+; SKX-LABEL: 'gather_v32i64'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 34 for instruction: %v = call <32 x i64> @llvm.masked.gather.v32i64.v32p0(<32 x ptr> align 8 %ptrs, <32 x i1> %mask, <32 x i64> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i64> %v
+;
+; X8664V4-LABEL: 'gather_v32i64'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 34 for instruction: %v = call <32 x i64> @llvm.masked.gather.v32i64.v32p0(<32 x ptr> align 8 %ptrs, <32 x i1> %mask, <32 x i64> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i64> %v
+;
+  %v = call <32 x i64> @llvm.masked.gather.v32i64.v32p0(<32 x ptr> %ptrs, i32 8, <32 x i1> %mask, <32 x i64> poison)
+  ret <32 x i64> %v
+}
diff --git a/llvm/test/Transforms/LoopVectorize/X86/gather-scatter-cost-table-decisions.ll b/llvm/test/Transforms/LoopVectorize/X86/gather-scatter-cost-table-decisions.ll
new file mode 100644
index 0000000000000..2474aa4bcb755
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/X86/gather-scatter-cost-table-decisions.ll
@@ -0,0 +1,99 @@
+; End-to-end loop-vectorize decisions driven by the per-shape gather/scatter
+; cost tables (TuningPreferGSCostTable, set on znver4+). The companion
+; cost-model test masked-gather-scatter-cost-table.ll pins the individual cost
+; numbers; this test pins the resulting vectorizer decisions so cost-model
+; refactors that re-enable harmful gathers (or suppress profitable ones) are
+; caught here.
+;
+; The three cases below:
+;   1. f64 indirect-load reduction -- gather IS chosen on znver5.
+;   2. i64 indirect-load reduction -- gather is NOT chosen on znver5 (the i64
+;      entry is set above the break-even to suppress vpgatherqq for harmful
+;      patterns, cf. PR llvm#198850).
+;   3. Unit-stride load -- must stay a plain wide load, not a gather.
+;      Regression guard for issue llvm#91370.
+;
+; RUN: opt < %s -S -passes=loop-vectorize -mtriple=x86_64-unknown-linux-gnu -mcpu=znver5 | FileCheck %s
+
+target datalayout = "e-m:e-p270:32:32-p271:32:32-p272:64:64-i64:64-f80:128-n8:16:32:64-S128"
+
+; --- Case 1: f64 indirect-load gather IS chosen on znver5 ----------------
+; CHECK-LABEL: define double @f64_indirect_gather_chosen
+; CHECK:       call <{{[0-9]+}} x double> @llvm.masked.gather.v{{[0-9]+}}f64
+define double @f64_indirect_gather_chosen(ptr noundef readonly %data, ptr noundef readonly %idx, i32 noundef %n) {
+entry:
+  %cmp = icmp ugt i32 %n, 0
+  br i1 %cmp, label %loop, label %exit
+
+loop:
+  %i = phi i32 [ 0, %entry ], [ %inc, %loop ]
+  %acc = phi double [ 0.0, %entry ], [ %acc.next, %loop ]
+  %idx.gep = getelementptr inbounds i32, ptr %idx, i32 %i
+  %idx.val = load i32, ptr %idx.gep, align 4
+  %idx.sext = sext i32 %idx.val to i64
+  %data.gep = getelementptr inbounds double, ptr %data, i64 %idx.sext
+  %data.val = load double, ptr %data.gep, align 8
+  %acc.next = fadd fast double %acc, %data.val
+  %inc = add nuw nsw i32 %i, 1
+  %done = icmp eq i32 %inc, %n
+  br i1 %done, label %exit, label %loop
+
+exit:
+  %ret = phi double [ 0.0, %entry ], [ %acc.next, %loop ]
+  ret double %ret
+}
+
+; --- Case 2: i64 indirect-load gather is NOT chosen on znver5 ------------
+; The positive CHECK on vector.body distinguishes "vectorized without a
+; gather" from "did not vectorize at all" -- without it, a future regression
+; that fails to vectorize the loop entirely would pass CHECK-NOT vacuously.
+; CHECK-LABEL: define i64 @i64_indirect_gather_avoided
+; CHECK:       vector.body
+; CHECK-NOT:   call <{{[0-9]+}} x i64> @llvm.masked.gather.v{{[0-9]+}}i64
+define i64 @i64_indirect_gather_avoided(ptr noundef readonly %data, ptr noundef readonly %idx, i32 noundef %n) {
+entry:
+  %cmp = icmp ugt i32 %n, 0
+  br i1 %cmp, label %loop, label %exit
+
+loop:
+  %i = phi i32 [ 0, %entry ], [ %inc, %loop ]
+  %acc = phi i64 [ 0, %entry ], [ %acc.next, %loop ]
+  %idx.gep = getelementptr inbounds i64, ptr %idx, i32 %i
+  %idx.val = load i64, ptr %idx.gep, align 8
+  %data.gep = getelementptr inbounds i64, ptr %data, i64 %idx.val
+  %data.val = load i64, ptr %data.gep, align 8
+  %acc.next = add i64 %acc, %data.val
+  %inc = add nuw nsw i32 %i, 1
+  %done = icmp eq i32 %inc, %n
+  br i1 %done, label %exit, label %loop
+
+exit:
+  %ret = phi i64 [ 0, %entry ], [ %acc.next, %loop ]
+  ret i64 %ret
+}
+
+; --- Case 3: unit-stride load must NOT become a gather (#91370 guard) -----
+; Same vector.body anchor as Case 2: ensures the loop did vectorize (to a
+; wide load) rather than failing to vectorize entirely.
+; CHECK-LABEL: define void @unit_stride_no_gather
+; CHECK:       vector.body
+; CHECK-NOT:   call <{{[0-9]+}} x double> @llvm.masked.gather
+define void @unit_stride_no_gather(ptr noundef writeonly %out, ptr noundef readonly %in, i32 noundef %n) {
+entry:
+  %cmp = icmp ugt i32 %n, 0
+  br i1 %cmp, label %loop, label %exit
+
+loop:
+  %i = phi i32 [ 0, %entry ], [ %inc, %loop ]
+  %in.gep = getelementptr inbounds double, ptr %in, i32 %i
+  %in.val = load double, ptr %in.gep, align 8
+  %mul = fmul fast double %in.val, 2.000000e+00
+  %out.gep = getelementptr inbounds double, ptr %out, i32 %i
+  store double %mul, ptr %out.gep, align 8
+  %inc = add nuw nsw i32 %i, 1
+  %done = icmp eq i32 %inc, %n
+  br i1 %done, label %exit, label %loop
+
+exit:
+  ret void
+}
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/gather-cost-znver4.ll b/llvm/test/Transforms/SLPVectorizer/X86/gather-cost-znver4.ll
new file mode 100644
index 0000000000000..0d43c16e29055
--- /dev/null
+++ b/llvm/test/Transforms/SLPVectorizer/X86/gather-cost-znver4.ll
@@ -0,0 +1,86 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 2
+; RUN: opt -passes=slp-vectorizer -S < %s -mtriple=x86_64-unknown-linux-gnu -mcpu=znver4 -slp-threshold=-5 | FileCheck %s --check-prefix=ZNVER4
+; RUN: opt -passes=slp-vectorizer -S < %s -mtriple=x86_64-unknown-linux-gnu -mcpu=skx    -slp-threshold=-5 | FileCheck %s --check-prefix=SKX
+; RUN: opt -passes=slp-vectorizer -S < %s -mtriple=x86_64-unknown-linux-gnu -mcpu=znver4 -slp-threshold=-8 | FileCheck %s --check-prefix=ZNVER4-BUDGET
+
+; Four loads addressed by loaded indices off a common base, feeding four
+; contiguous stores: SLP vectorizes the stores and gathers the source loads via
+; ScatterVectorize, reaching X86TTIImpl::getGSVectorCost for a 4-wide gather.
+; On znver4 that gather is priced from the per-shape cost table
+; (TuningPreferGSCostTable = schedule-model body + break-even overhead) at 11,
+; against 6 from the generic flat overhead on skx.
+;
+; The first two runs pin that the difference changes the decision and not just
+; the number: at the same threshold skx forms the gather, while znver4 leaves
+; the loads scalar and vectorizes only the stores. The third run gives znver4
+; the budget to afford the gather, so the ZNVER4 checks are a costing decision
+; rather than SLP never reaching the gather at all.
+define void @gather4(ptr %base, ptr %idxp, ptr noalias %out) {
+; ZNVER4-LABEL: define void @gather4
+; ZNVER4-SAME: (ptr [[BASE:%.*]], ptr [[IDXP:%.*]], ptr noalias [[OUT:%.*]]) #[[ATTR0:[0-9]+]] {
+; ZNVER4-NEXT:    [[I0:%.*]] = load i32, ptr [[IDXP]], align 4
+; ZNVER4-NEXT:    [[G1:%.*]] = getelementptr inbounds i32, ptr [[IDXP]], i64 1
+; ZNVER4-NEXT:    [[I1:%.*]] = load i32, ptr [[G1]], align 4
+; ZNVER4-NEXT:    [[G2:%.*]] = getelementptr inbounds i32, ptr [[IDXP]], i64 2
+; ZNVER4-NEXT:    [[TMP1:%.*]] = load <2 x i32>, ptr [[G2]], align 4
+; ZNVER4-NEXT:    [[P0:%.*]] = getelementptr inbounds float, ptr [[BASE]], i32 [[I0]]
+; ZNVER4-NEXT:    [[V0:%.*]] = load float, ptr [[P0]], align 4
+; ZNVER4-NEXT:    [[P1:%.*]] = getelementptr inbounds float, ptr [[BASE]], i32 [[I1]]
+; ZNVER4-NEXT:    [[V1:%.*]] = load float, ptr [[P1]], align 4
+; ZNVER4-NEXT:    [[TMP2:%.*]] = extractelement <2 x i32> [[TMP1]], i64 0
+; ZNVER4-NEXT:    [[P2:%.*]] = getelementptr inbounds float, ptr [[BASE]], i32 [[TMP2]]
+; ZNVER4-NEXT:    [[V2:%.*]] = load float, ptr [[P2]], align 4
+; ZNVER4-NEXT:    [[TMP3:%.*]] = extractelement <2 x i32> [[TMP1]], i64 1
+; ZNVER4-NEXT:    [[P3:%.*]] = getelementptr inbounds float, ptr [[BASE]], i32 [[TMP3]]
+; ZNVER4-NEXT:    [[V3:%.*]] = load float, ptr [[P3]], align 4
+; ZNVER4-NEXT:    [[TMP4:%.*]] = insertelement <4 x float> poison, float [[V0]], i64 0
+; ZNVER4-NEXT:    [[TMP5:%.*]] = insertelement <4 x float> [[TMP4]], float [[V1]], i64 1
+; ZNVER4-NEXT:    [[TMP6:%.*]] = insertelement <4 x float> [[TMP5]], float [[V2]], i64 2
+; ZNVER4-NEXT:    [[TMP7:%.*]] = insertelement <4 x float> [[TMP6]], float [[V3]], i64 3
+; ZNVER4-NEXT:    store <4 x float> [[TMP7]], ptr [[OUT]], align 4
+; ZNVER4-NEXT:    ret void
+;
+; SKX-LABEL: define void @gather4
+; SKX-SAME: (ptr [[BASE:%.*]], ptr [[IDXP:%.*]], ptr noalias [[OUT:%.*]]) #[[ATTR0:[0-9]+]] {
+; SKX-NEXT:    [[TMP1:%.*]] = load <4 x i32>, ptr [[IDXP]], align 4
+; SKX-NEXT:    [[TMP2:%.*]] = insertelement <4 x ptr> poison, ptr [[BASE]], i64 0
+; SKX-NEXT:    [[TMP3:%.*]] = shufflevector <4 x ptr> [[TMP2]], <4 x ptr> poison, <4 x i32> zeroinitializer
+; SKX-NEXT:    [[TMP4:%.*]] = getelementptr inbounds float, <4 x ptr> [[TMP3]], <4 x i32> [[TMP1]]
+; SKX-NEXT:    [[TMP5:%.*]] = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 4 [[TMP4]], <4 x i1> splat (i1 true), <4 x float> poison)
+; SKX-NEXT:    store <4 x float> [[TMP5]], ptr [[OUT]], align 4
+; SKX-NEXT:    ret void
+;
+; ZNVER4-BUDGET-LABEL: define void @gather4
+; ZNVER4-BUDGET-SAME: (ptr [[BASE:%.*]], ptr [[IDXP:%.*]], ptr noalias [[OUT:%.*]]) #[[ATTR0:[0-9]+]] {
+; ZNVER4-BUDGET-NEXT:    [[TMP1:%.*]] = load <4 x i32>, ptr [[IDXP]], align 4
+; ZNVER4-BUDGET-NEXT:    [[TMP2:%.*]] = insertelement <4 x ptr> poison, ptr [[BASE]], i64 0
+; ZNVER4-BUDGET-NEXT:    [[TMP3:%.*]] = shufflevector <4 x ptr> [[TMP2]], <4 x ptr> poison, <4 x i32> zeroinitializer
+; ZNVER4-BUDGET-NEXT:    [[TMP4:%.*]] = getelementptr inbounds float, <4 x ptr> [[TMP3]], <4 x i32> [[TMP1]]
+; ZNVER4-BUDGET-NEXT:    [[TMP5:%.*]] = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 4 [[TMP4]], <4 x i1> splat (i1 true), <4 x float> poison)
+; ZNVER4-BUDGET-NEXT:    store <4 x float> [[TMP5]], ptr [[OUT]], align 4
+; ZNVER4-BUDGET-NEXT:    ret void
+;
+  %i0 = load i32, ptr %idxp, align 4
+  %g1 = getelementptr inbounds i32, ptr %idxp, i64 1
+  %i1 = load i32, ptr %g1, align 4
+  %g2 = getelementptr inbounds i32, ptr %idxp, i64 2
+  %i2 = load i32, ptr %g2, align 4
+  %g3 = getelementptr inbounds i32, ptr %idxp, i64 3
+  %i3 = load i32, ptr %g3, align 4
+  %p0 = getelementptr inbounds float, ptr %base, i32 %i0
+  %v0 = load float, ptr %p0, align 4
+  %p1 = getelementptr inbounds float, ptr %base, i32 %i1
+  %v1 = load float, ptr %p1, align 4
+  %p2 = getelementptr inbounds float, ptr %base, i32 %i2
+  %v2 = load float, ptr %p2, align 4
+  %p3 = getelementptr inbounds float, ptr %base, i32 %i3
+  %v3 = load float, ptr %p3, align 4
+  store float %v0, ptr %out, align 4
+  %o1 = getelementptr inbounds float, ptr %out, i64 1
+  store float %v1, ptr %o1, align 4
+  %o2 = getelementptr inbounds float, ptr %out, i64 2
+  store float %v2, ptr %o2, align 4
+  %o3 = getelementptr inbounds float, ptr %out, i64 3
+  store float %v3, ptr %o3, align 4
+  ret void
+}

>From eff5fdabc67e8d36c35711492ae41eff4603d6b0 Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Wed, 2 Sep 2026 11:31:55 +0530
Subject: [PATCH 4/5] [X86][CostModel] Price pointer-element gather/scatter
 from the Zen tables

getGSVectorCost gated the znver4+ cost tables on the IR scalar type's
primitive size, which is 0 for a pointer, so a gather or scatter of
pointers never reached the tables and took the flat AVX-512 overhead
instead. A gather of pointers is the same VPGATHERQQ as a gather of
i64, so the two forms of one instruction were priced more than 3x
apart: <8 x ptr> at 10 against <8 x i64> at 32.

The gap defeated the rows where it matters most. The i64 entries exist
to keep the vectorizers off 64-bit gather/scatter on Zen, and
SLPVectorizer's ScatterVectorize reaches getGSVectorCost with pointer
element types as a matter of course.

Take the element width from the legalized type instead. Both
getZenGSOverhead and getModeledGSInstrCost already legalize internally,
so nothing downstream needs to change: pointer shapes now price
identically to their i64 equivalents, and no other shape moves.
---
 .../lib/Target/X86/X86TargetTransformInfo.cpp |   6 +-
 .../X86/masked-gather-scatter-cost-table.ll   | 125 ++++++++++++++++++
 2 files changed, 130 insertions(+), 1 deletion(-)

diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
index d6f68c9aa3993..03f5611eb10b6 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
@@ -6828,7 +6828,11 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
   if (CostKind == TTI::TCK_RecipThroughput && ST->hasPreferGSCostTable() &&
       ST->hasAVX512()) {
     Type *EltTy = SrcVTy->getScalarType();
-    unsigned EltBits = EltTy->getPrimitiveSizeInBits();
+    // Take the width from the legalized type, not the IR one: a pointer
+    // element gathers as its integer equivalent (the same VPGATHERQQ), but
+    // reports a primitive size of 0.
+    EVT SrcEVT = TLI->getValueType(DL, SrcVTy);
+    unsigned EltBits = SrcEVT.isVector() ? SrcEVT.getScalarSizeInBits() : 0;
     if (EltBits == 32 || EltBits == 64) {
       unsigned NativeMaxElts = 512 / EltBits;
       // Split only for genuinely-wider-than-native shapes, and only when the
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-table.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-table.ll
index 3f638adaaf76f..bbc79850980f2 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-table.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-table.ll
@@ -745,6 +745,131 @@ define void @scatter_v8f64(<8 x double> %src, <8 x ptr> %ptrs, <8 x i1> %mask) {
   ret void
 }
 
+;------------------------------------------------------------------------------
+; Pointer element type.
+;
+; A gather of pointers is the same VPGATHERQQ as a gather of i64, so it must be
+; priced from the i64 rows. The table is keyed on the legalized type, not the IR
+; one, which matters here because a pointer reports a primitive size of 0.
+; SLPVectorizer's ScatterVectorize reaches getGSVectorCost with these shapes.
+;------------------------------------------------------------------------------
+
+define <4 x ptr> @gather_v4ptr(<4 x ptr> %ptrs, <4 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v4ptr'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 15 for instruction: %v = call <4 x ptr> @llvm.masked.gather.v4p0.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x ptr> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x ptr> %v
+;
+; ZNVER5-LABEL: 'gather_v4ptr'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 15 for instruction: %v = call <4 x ptr> @llvm.masked.gather.v4p0.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x ptr> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x ptr> %v
+;
+; ZNVER6-LABEL: 'gather_v4ptr'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 15 for instruction: %v = call <4 x ptr> @llvm.masked.gather.v4p0.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x ptr> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x ptr> %v
+;
+; ZNVER3-LABEL: 'gather_v4ptr'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 15 for instruction: %v = call <4 x ptr> @llvm.masked.gather.v4p0.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x ptr> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x ptr> %v
+;
+; SKX-LABEL: 'gather_v4ptr'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: %v = call <4 x ptr> @llvm.masked.gather.v4p0.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x ptr> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x ptr> %v
+;
+; X8664V4-LABEL: 'gather_v4ptr'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: %v = call <4 x ptr> @llvm.masked.gather.v4p0.v4p0(<4 x ptr> align 8 %ptrs, <4 x i1> %mask, <4 x ptr> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <4 x ptr> %v
+;
+  %v = call <4 x ptr> @llvm.masked.gather.v4p0.v4p0(<4 x ptr> %ptrs, i32 8, <4 x i1> %mask, <4 x ptr> poison)
+  ret <4 x ptr> %v
+}
+
+define <8 x ptr> @gather_v8ptr(<8 x ptr> %ptrs, <8 x i1> %mask) {
+; ZNVER4-LABEL: 'gather_v8ptr'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 32 for instruction: %v = call <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x ptr> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x ptr> %v
+;
+; ZNVER5-LABEL: 'gather_v8ptr'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 32 for instruction: %v = call <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x ptr> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x ptr> %v
+;
+; ZNVER6-LABEL: 'gather_v8ptr'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 32 for instruction: %v = call <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x ptr> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x ptr> %v
+;
+; ZNVER3-LABEL: 'gather_v8ptr'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 29 for instruction: %v = call <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x ptr> poison)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x ptr> %v
+;
+; SKX-LABEL: 'gather_v8ptr'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: %v = call <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x ptr> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x ptr> %v
+;
+; X8664V4-LABEL: 'gather_v8ptr'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: %v = call <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr> align 8 %ptrs, <8 x i1> %mask, <8 x ptr> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x ptr> %v
+;
+  %v = call <8 x ptr> @llvm.masked.gather.v8p0.v8p0(<8 x ptr> %ptrs, i32 8, <8 x i1> %mask, <8 x ptr> poison)
+  ret <8 x ptr> %v
+}
+
+define void @scatter_v4ptr(<4 x ptr> %src, <4 x ptr> %ptrs, <4 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v4ptr'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 16 for instruction: call void @llvm.masked.scatter.v4p0.v4p0(<4 x ptr> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v4ptr'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 16 for instruction: call void @llvm.masked.scatter.v4p0.v4p0(<4 x ptr> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v4ptr'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 16 for instruction: call void @llvm.masked.scatter.v4p0.v4p0(<4 x ptr> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v4ptr'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 15 for instruction: call void @llvm.masked.scatter.v4p0.v4p0(<4 x ptr> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v4ptr'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: call void @llvm.masked.scatter.v4p0.v4p0(<4 x ptr> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v4ptr'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 6 for instruction: call void @llvm.masked.scatter.v4p0.v4p0(<4 x ptr> %src, <4 x ptr> align 8 %ptrs, <4 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  call void @llvm.masked.scatter.v4p0.v4p0(<4 x ptr> %src, <4 x ptr> %ptrs, i32 8, <4 x i1> %mask)
+  ret void
+}
+
+define void @scatter_v8ptr(<8 x ptr> %src, <8 x ptr> %ptrs, <8 x i1> %mask) {
+; ZNVER4-LABEL: 'scatter_v8ptr'
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: call void @llvm.masked.scatter.v8p0.v8p0(<8 x ptr> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER5-LABEL: 'scatter_v8ptr'
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: call void @llvm.masked.scatter.v8p0.v8p0(<8 x ptr> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER6-LABEL: 'scatter_v8ptr'
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: call void @llvm.masked.scatter.v8p0.v8p0(<8 x ptr> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; ZNVER3-LABEL: 'scatter_v8ptr'
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 29 for instruction: call void @llvm.masked.scatter.v8p0.v8p0(<8 x ptr> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SKX-LABEL: 'scatter_v8ptr'
+; SKX-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: call void @llvm.masked.scatter.v8p0.v8p0(<8 x ptr> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; X8664V4-LABEL: 'scatter_v8ptr'
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 10 for instruction: call void @llvm.masked.scatter.v8p0.v8p0(<8 x ptr> %src, <8 x ptr> align 8 %ptrs, <8 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+  call void @llvm.masked.scatter.v8p0.v8p0(<8 x ptr> %src, <8 x ptr> %ptrs, i32 8, <8 x i1> %mask)
+  ret void
+}
+
 ;------------------------------------------------------------------------------
 ; 16-wide gather/scatter with a single base + 32-bit index GEP.
 ;

>From 3eb4ee5d573c06e7b6bd03934756a2fd1c9937b2 Mon Sep 17 00:00:00 2001
From: Sumukh Bharadwaj <Sumukh.Bharadwaj at amd.com>
Date: Wed, 2 Sep 2026 14:38:02 +0530
Subject: [PATCH 5/5] [X86][CostModel] Correct Zen gather/scatter legalization
 costs

Select the schedule-model body from the actual index width and tile
irregular vectors with the legal shapes CodeGen emits. Use the exact part
count for code size, and keep one-time split overhead local to the
calibrated Zen path so non-AMD costs and decisions remain unchanged.
---
 .../lib/Target/X86/X86TargetTransformInfo.cpp | 157 ++++++++++--------
 llvm/lib/Target/X86/X86TargetTransformInfo.h  |   2 +-
 .../masked-gather-scatter-cost-edge-cases.ll  | 152 +++++++++++++++++
 .../X86/masked-gather-scatter-cost-table.ll   | 112 ++++++-------
 .../X86/masked-intrinsic-cost-inseltpoison.ll |  38 ++---
 .../CostModel/X86/masked-intrinsic-cost.ll    |  38 ++---
 .../X86/CostModel/gather-i32-with-i8-index.ll |  10 +-
 .../X86/CostModel/gather-i64-with-i8-index.ll |  12 +-
 .../interleaved-load-f32-stride-2.ll          |   3 +
 .../interleaved-load-f32-stride-4.ll          |   5 +
 .../interleaved-load-f32-stride-8.ll          |   9 +
 .../interleaved-load-f64-stride-5.ll          |  30 ++--
 .../interleaved-load-f64-stride-6.ll          |  36 ++--
 .../interleaved-load-f64-stride-7.ll          |  42 ++---
 .../interleaved-load-f64-stride-8.ll          |  48 +++---
 ...nterleaved-load-i32-stride-2-indices-0u.ll |   2 +
 .../interleaved-load-i32-stride-2.ll          |   3 +
 ...erleaved-load-i32-stride-4-indices-012u.ll |   4 +
 ...erleaved-load-i32-stride-4-indices-01uu.ll |   3 +
 ...erleaved-load-i32-stride-4-indices-0uuu.ll |   4 +-
 .../interleaved-load-i32-stride-4.ll          |   5 +
 .../interleaved-load-i32-stride-8.ll          |   9 +
 .../interleaved-load-i64-stride-2.ll          |   8 +-
 .../interleaved-load-i64-stride-4.ll          |  24 +--
 .../interleaved-load-i64-stride-5.ll          |  30 ++--
 .../interleaved-load-i64-stride-6.ll          |  36 ++--
 .../interleaved-load-i64-stride-7.ll          |  42 ++---
 .../interleaved-load-i64-stride-8.ll          |  48 +++---
 .../interleaved-store-f32-stride-8.ll         |  27 +++
 .../interleaved-store-f64-stride-4.ll         |   5 +
 .../interleaved-store-f64-stride-8.ll         |  48 +++---
 .../interleaved-store-i32-stride-8.ll         |  27 +++
 .../interleaved-store-i64-stride-4.ll         |   5 +
 .../interleaved-store-i64-stride-8.ll         |  48 +++---
 .../masked-gather-i32-with-i8-index.ll        |  10 +-
 .../masked-gather-i64-with-i8-index.ll        |  12 +-
 .../masked-scatter-i32-with-i8-index.ll       |   4 +-
 .../masked-scatter-i64-with-i8-index.ll       |   6 +-
 .../CostModel/scatter-i32-with-i8-index.ll    |   4 +-
 .../CostModel/scatter-i64-with-i8-index.ll    |   6 +-
 .../LoopVectorize/X86/cast-costs.ll           |  24 +--
 .../LoopVectorize/X86/interleave-cost.ll      |  68 ++++----
 42 files changed, 745 insertions(+), 461 deletions(-)
 create mode 100644 llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-edge-cases.ll

diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
index 03f5611eb10b6..148c76a5e787b 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
@@ -6590,44 +6590,60 @@ InstructionCost X86TTIImpl::getCFInstrCost(unsigned Opcode,
   return TTI::TCC_Free;
 }
 
-// Pick a representative masked gather/scatter opcode for a data shape, used to
-// read the body cost from the schedule model. Keyed on (NumElts, EltBits): the
-// integer and FP variants of a shape share a scheduling class, so the integer
-// form is returned. Index width is not part of the key; the matched-width form
-// is always used (dword index for 32-bit data, qword for 64-bit). Where a shape
-// has a separate class per index width the two agree to within a cycle, with
-// one exception: a v8i64 scatter is 14 on the qword-index class against 11 on
-// the dword-index one, so one reached through 32-bit indices is charged 3 too
-// many. The bare-pointer form, which is what the vectorizers produce, is the
-// qword-index one. Returns 0 for unsupported shapes (VF 5/6/7, or VF not in
-// {4,8,16}), whereupon the caller takes the generic fallback; the caller also
-// checks VLX, needed for the sub-512-bit forms.
-static unsigned getAVX512GSRepresentativeOpcode(bool IsLoad, MVT VT) {
+// Pick a representative masked gather/scatter opcode for a data and index
+// shape, used to read the body cost from the schedule model. The integer and FP
+// variants of a shape share a scheduling class, so the integer form is
+// returned. Returns 0 for unsupported shapes.
+static unsigned getAVX512GSRepresentativeOpcode(bool IsLoad, MVT VT,
+                                                unsigned IndexBits) {
   if (!VT.isVector())
     return 0;
   unsigned NumElts = VT.getVectorNumElements();
   unsigned EltBits = VT.getScalarSizeInBits();
+  if (IndexBits != 32 && IndexBits != 64)
+    return 0;
+
   if (IsLoad) {
-    if (EltBits == 32)
-      return NumElts == 4    ? X86::VPGATHERDDZ128rm
-             : NumElts == 8  ? X86::VPGATHERDDZ256rm
-             : NumElts == 16 ? X86::VPGATHERDDZrm
-                             : 0;
-    if (EltBits == 64)
+    if (EltBits == 32) {
+      if (IndexBits == 32)
+        return NumElts == 4    ? X86::VPGATHERDDZ128rm
+               : NumElts == 8  ? X86::VPGATHERDDZ256rm
+               : NumElts == 16 ? X86::VPGATHERDDZrm
+                               : 0;
+      return NumElts == 4   ? X86::VPGATHERQDZ256rm
+             : NumElts == 8 ? X86::VPGATHERQDZrm
+                            : 0;
+    }
+    if (EltBits == 64) {
+      if (IndexBits == 32)
+        return NumElts == 4   ? X86::VPGATHERDQZ256rm
+               : NumElts == 8 ? X86::VPGATHERDQZrm
+                              : 0;
       return NumElts == 4   ? X86::VPGATHERQQZ256rm
              : NumElts == 8 ? X86::VPGATHERQQZrm
                             : 0;
+    }
     return 0;
   }
-  if (EltBits == 32)
-    return NumElts == 4    ? X86::VPSCATTERDDZ128mr
-           : NumElts == 8  ? X86::VPSCATTERDDZ256mr
-           : NumElts == 16 ? X86::VPSCATTERDDZmr
-                           : 0;
-  if (EltBits == 64)
+  if (EltBits == 32) {
+    if (IndexBits == 32)
+      return NumElts == 4    ? X86::VPSCATTERDDZ128mr
+             : NumElts == 8  ? X86::VPSCATTERDDZ256mr
+             : NumElts == 16 ? X86::VPSCATTERDDZmr
+                             : 0;
+    return NumElts == 4   ? X86::VPSCATTERQDZ256mr
+           : NumElts == 8 ? X86::VPSCATTERQDZmr
+                          : 0;
+  }
+  if (EltBits == 64) {
+    if (IndexBits == 32)
+      return NumElts == 4   ? X86::VPSCATTERDQZ256mr
+             : NumElts == 8 ? X86::VPSCATTERDQZmr
+                            : 0;
     return NumElts == 4   ? X86::VPSCATTERQQZ256mr
            : NumElts == 8 ? X86::VPSCATTERQQZmr
                           : 0;
+  }
   return 0;
 }
 
@@ -6670,6 +6686,7 @@ static std::optional<unsigned> getSchedModelGSBody(unsigned Opc,
 // from the schedule model; getZenGSOverhead adds the profitability premium.
 std::optional<unsigned>
 X86TTIImpl::getModeledGSInstrCost(bool IsLoad, Type *SrcVTy,
+                                  unsigned IndexSize,
                                   TTI::TargetCostKind CostKind) const {
   if (CostKind != TTI::TCK_RecipThroughput || !ST->hasAVX512() ||
       !ST->hasPreferGSCostTable() || !SrcVTy)
@@ -6677,11 +6694,14 @@ X86TTIImpl::getModeledGSInstrCost(bool IsLoad, Type *SrcVTy,
   EVT VT = TLI->getValueType(DL, SrcVTy);
   if (!VT.isSimple())
     return std::nullopt;
-  // The sub-512-bit (Z128/Z256) forms require AVX512VL; without it the opcodes
-  // do not exist, so fall back rather than read a meaningless class.
-  if (VT.getSizeInBits() < 512 && !ST->hasVLX())
+  // An encoding whose data and index operands are both sub-512-bit requires
+  // AVX512VL. A v8i32 operation with i64 indices still uses a full zmm index
+  // and therefore does not.
+  unsigned IndexVectorBits = IndexSize * VT.getVectorNumElements();
+  if (VT.getSizeInBits() < 512 && IndexVectorBits < 512 && !ST->hasVLX())
     return std::nullopt;
-  unsigned Opc = getAVX512GSRepresentativeOpcode(IsLoad, VT.getSimpleVT());
+  unsigned Opc =
+      getAVX512GSRepresentativeOpcode(IsLoad, VT.getSimpleVT(), IndexSize);
   if (!Opc)
     return std::nullopt;
   return getSchedModelGSBody(Opc, ST);
@@ -6693,9 +6713,9 @@ X86TTIImpl::getModeledGSInstrCost(bool IsLoad, Type *SrcVTy,
 // body from getModeledGSInstrCost. Keeping the hardware half in
 // X86ScheduleZnver4.td leaves one source of truth and only the profitability
 // premium here. The premium belongs to the whole logical op and is paid once
-// however legalisation splits it; the caller scales by the split factor only
-// for genuinely wider-than-native shapes. Totals were calibrated on a Zen5
-// 9950X against a body term now measured on Znver4 (Ryzen 5 8645HS).
+// when that op has a calibrated row. Irregular and wider shapes are tiled with
+// legal table shapes instead. Totals were calibrated on a Zen5 9950X against a
+// body term now measured on Znver4 (Ryzen 5 8645HS).
 //
 // Keyed by native shape (VF <= 16 for 32-bit, <= 8 for 64-bit).
 std::optional<unsigned> X86TTIImpl::getZenGSOverhead(bool IsLoad,
@@ -6802,11 +6822,14 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
     return (unsigned)32;
   };
 
-  // Trying to reduce IndexSize to 32 bits for vector 16.
-  // By default the IndexSize is equal to pointer size.
-  unsigned IndexSize = (ST->hasAVX512() && VF >= 16)
-                           ? getIndexSizeInBits(Ptr, DL)
-                           : DL.getPointerSizeInBits();
+  // By default the index is pointer-sized. The generic model has historically
+  // tried to narrow only VF 16+, where this avoids a legalization split. The
+  // Zen schedule-driven path also needs the real index width at smaller VFs to
+  // select the matching DQ/QQ (or DD/QD) instruction class.
+  unsigned IndexSize =
+      ST->hasAVX512() && (VF >= 16 || ST->hasPreferGSCostTable())
+          ? getIndexSizeInBits(Ptr, DL)
+          : DL.getPointerSizeInBits();
 
   auto *IndexVTy = FixedVectorType::get(
       IntegerType::get(SrcVTy->getContext(), IndexSize), VF);
@@ -6814,19 +6837,24 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
   std::pair<InstructionCost, MVT> SrcLT = getTypeLegalizationCost(SrcVTy);
   InstructionCost::CostType SplitFactor =
       std::max(IdxsLT.first, SrcLT.first).getValue();
+  unsigned NumParts =
+      std::max(getNumberOfParts(IndexVTy), getNumberOfParts(SrcVTy));
   const bool IsLoad = Opcode == Instruction::Load;
 
-  // A split op (e.g. 16-wide with 64-bit indices that don't fit one zmm)
-  // issues SplitFactor sub-ops; for code size that is just the sub-op count.
+  // The legalization multiplier is rounded to a power of two. Code size needs
+  // the actual number of emitted parts for irregular vectors such as v24.
   if (CostKind == TTI::TCK_CodeSize)
-    return SplitFactor;
-
-  // AMD znver4+ (TuningPreferGSCostTable): overhead + schedule-model body, both
-  // keyed on the *native* shape, so a shape fitting one register is charged
-  // once however legalisation splits it (narrow- and wide-index forms then
-  // report alike). Wider-than-native shapes tile into GSMul native ops.
+    return NumParts;
+
+  // AMD znver4+ (TuningPreferGSCostTable): decompose the operation into the
+  // legal power-of-two shape that CodeGen emits, then read that shape's body
+  // from the schedule model using the actual index width. For a calibrated
+  // native data shape, its profitability overhead belongs to the whole logical
+  // operation and is paid once even if wider indices split it. Irregular and
+  // wider-than-native shapes have no logical-operation row, so tile them with
+  // the modeled legal shape and pay that shape's complete cost per part.
   if (CostKind == TTI::TCK_RecipThroughput && ST->hasPreferGSCostTable() &&
-      ST->hasAVX512()) {
+      ST->hasAVX512() && NumParts) {
     Type *EltTy = SrcVTy->getScalarType();
     // Take the width from the legalized type, not the IR one: a pointer
     // element gathers as its integer equivalent (the same VPGATHERQQ), but
@@ -6834,29 +6862,24 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
     EVT SrcEVT = TLI->getValueType(DL, SrcVTy);
     unsigned EltBits = SrcEVT.isVector() ? SrcEVT.getScalarSizeInBits() : 0;
     if (EltBits == 32 || EltBits == 64) {
-      unsigned NativeMaxElts = 512 / EltBits;
-      // Split only for genuinely-wider-than-native shapes, and only when the
-      // width is an exact multiple so the native op tiles cleanly.
-      unsigned NativeVF = std::min(VF, NativeMaxElts);
-      if (VF % NativeVF == 0) {
-        unsigned GSMul = VF / NativeVF;
-        auto *NativeVTy = FixedVectorType::get(EltTy, NativeVF);
-        std::optional<unsigned> Overhead = getZenGSOverhead(IsLoad, NativeVTy);
-        std::optional<unsigned> Body =
-            getModeledGSInstrCost(IsLoad, NativeVTy, CostKind);
-        if (Overhead && Body)
-          return GSMul * (*Overhead + *Body);
+      unsigned PartVF = PowerOf2Ceil(divideCeil(VF, NumParts));
+      auto *PartVTy = FixedVectorType::get(EltTy, PartVF);
+      std::optional<unsigned> Body =
+          getModeledGSInstrCost(IsLoad, PartVTy, IndexSize, CostKind);
+      if (Body) {
+        if (std::optional<unsigned> OpOverhead =
+                getZenGSOverhead(IsLoad, SrcVTy))
+          return *OpOverhead + NumParts * *Body;
+        if (std::optional<unsigned> PartOverhead =
+                getZenGSOverhead(IsLoad, PartVTy))
+          return NumParts * (*PartOverhead + *Body);
       }
     }
   }
 
-  // Otherwise (a non-enumerated VF, or a target without
-  // TuningPreferGSCostTable): flat break-even overhead plus body. The overhead
-  // belongs to the pre-split shape and is paid once; only the body
-  // (VF * scalar-memory-op) scales with the split factor. Not guaranteed
-  // monotonic against the calibrated table - v6i32 can read cheaper than
-  // v4i32's measured total - but such VFs are rare and the table above remains
-  // the source of truth.
+  // Otherwise use the flat break-even overhead plus body. Preserve the
+  // historical per-part overhead for targets outside the Zen cost-table path;
+  // the one-per-logical-operation policy is intentionally local to that path.
   const int GSOverhead = IsLoad ? getGatherOverhead() : getScatterOverhead();
 
   auto BodyCostOf = [&](Type *VTy) -> InstructionCost {
@@ -6873,7 +6896,9 @@ InstructionCost X86TTIImpl::getGSVectorCost(unsigned Opcode,
   } else {
     BodyCost = BodyCostOf(SrcVTy);
   }
-  return GSOverhead + BodyCost;
+  InstructionCost::CostType OverheadMultiplier =
+      SplitFactor > 1 && !ST->hasPreferGSCostTable() ? SplitFactor : 1;
+  return OverheadMultiplier * GSOverhead + BodyCost;
 }
 
 /// Calculate the cost of Gather / Scatter operation
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.h b/llvm/lib/Target/X86/X86TargetTransformInfo.h
index e97ca07a437c9..b707aa2a01929 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.h
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.h
@@ -282,7 +282,7 @@ class X86TTIImpl final : public BasicTTIImplBase<X86TTIImpl> {
   // Hardware body cost of a single (non-split) masked gather/scatter, read
   // live from this subtarget's schedule model, or nullopt to fall back.
   std::optional<unsigned>
-  getModeledGSInstrCost(bool IsLoad, Type *SrcVTy,
+  getModeledGSInstrCost(bool IsLoad, Type *SrcVTy, unsigned IndexSize,
                         TTI::TargetCostKind CostKind) const;
 
   /// @}
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-edge-cases.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-edge-cases.ll
new file mode 100644
index 0000000000000..4898e67e3749a
--- /dev/null
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-edge-cases.ll
@@ -0,0 +1,152 @@
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu \
+; RUN:   -passes="print<cost-model>" -disable-output -cost-kind=throughput \
+; RUN:   -mcpu=znver4 2>&1 | FileCheck %s --check-prefix=THROUGHPUT
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu \
+; RUN:   -passes="print<cost-model>" -disable-output -cost-kind=code-size \
+; RUN:   -mcpu=znver4 2>&1 | FileCheck %s --check-prefix=SIZE
+
+; Irregular widths are padded or split into the power-of-two gather shapes
+; emitted by CodeGen instead of falling back to the flat generic cost.
+
+define <3 x i32> @gather_v3i32(<3 x ptr> %ptrs, <3 x i1> %mask) {
+; THROUGHPUT-LABEL: 'gather_v3i32'
+; THROUGHPUT: Cost Model: Found an estimated cost of 12 for instruction: %v = call
+; SIZE-LABEL: 'gather_v3i32'
+; SIZE: Cost Model: Found an estimated cost of 1 for instruction: %v = call
+  %v = call <3 x i32> @llvm.masked.gather.v3i32.v3p0(<3 x ptr> %ptrs, i32 4, <3 x i1> %mask, <3 x i32> poison)
+  ret <3 x i32> %v
+}
+
+define <5 x i32> @gather_v5i32(<5 x ptr> %ptrs, <5 x i1> %mask) {
+; THROUGHPUT-LABEL: 'gather_v5i32'
+; THROUGHPUT: Cost Model: Found an estimated cost of 26 for instruction: %v = call
+; SIZE-LABEL: 'gather_v5i32'
+; SIZE: Cost Model: Found an estimated cost of 1 for instruction: %v = call
+  %v = call <5 x i32> @llvm.masked.gather.v5i32.v5p0(<5 x ptr> %ptrs, i32 4, <5 x i1> %mask, <5 x i32> poison)
+  ret <5 x i32> %v
+}
+
+define <6 x i32> @gather_v6i32(<6 x ptr> %ptrs, <6 x i1> %mask) {
+; THROUGHPUT-LABEL: 'gather_v6i32'
+; THROUGHPUT: Cost Model: Found an estimated cost of 26 for instruction: %v = call
+; SIZE-LABEL: 'gather_v6i32'
+; SIZE: Cost Model: Found an estimated cost of 1 for instruction: %v = call
+  %v = call <6 x i32> @llvm.masked.gather.v6i32.v6p0(<6 x ptr> %ptrs, i32 4, <6 x i1> %mask, <6 x i32> poison)
+  ret <6 x i32> %v
+}
+
+define <7 x i32> @gather_v7i32(<7 x ptr> %ptrs, <7 x i1> %mask) {
+; THROUGHPUT-LABEL: 'gather_v7i32'
+; THROUGHPUT: Cost Model: Found an estimated cost of 26 for instruction: %v = call
+; SIZE-LABEL: 'gather_v7i32'
+; SIZE: Cost Model: Found an estimated cost of 1 for instruction: %v = call
+  %v = call <7 x i32> @llvm.masked.gather.v7i32.v7p0(<7 x ptr> %ptrs, i32 4, <7 x i1> %mask, <7 x i32> poison)
+  ret <7 x i32> %v
+}
+
+define <12 x i32> @gather_v12i32(<12 x ptr> %ptrs, <12 x i1> %mask) {
+; THROUGHPUT-LABEL: 'gather_v12i32'
+; THROUGHPUT: Cost Model: Found an estimated cost of 52 for instruction: %v = call
+; SIZE-LABEL: 'gather_v12i32'
+; SIZE: Cost Model: Found an estimated cost of 2 for instruction: %v = call
+  %v = call <12 x i32> @llvm.masked.gather.v12i32.v12p0(<12 x ptr> %ptrs, i32 4, <12 x i1> %mask, <12 x i32> poison)
+  ret <12 x i32> %v
+}
+
+define <24 x i32> @gather_v24i32(<24 x ptr> %ptrs, <24 x i1> %mask) {
+; THROUGHPUT-LABEL: 'gather_v24i32'
+; THROUGHPUT: Cost Model: Found an estimated cost of 78 for instruction: %v = call
+; SIZE-LABEL: 'gather_v24i32'
+; SIZE: Cost Model: Found an estimated cost of 3 for instruction: %v = call
+  %v = call <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr> %ptrs, i32 4, <24 x i1> %mask, <24 x i32> poison)
+  ret <24 x i32> %v
+}
+
+define void @scatter_v24i32(<24 x i32> %src, <24 x ptr> %ptrs, <24 x i1> %mask) {
+; THROUGHPUT-LABEL: 'scatter_v24i32'
+; THROUGHPUT: Cost Model: Found an estimated cost of 75 for instruction: call void
+; SIZE-LABEL: 'scatter_v24i32'
+; SIZE: Cost Model: Found an estimated cost of 3 for instruction: call void
+  call void @llvm.masked.scatter.v24i32.v24p0(<24 x i32> %src, <24 x ptr> %ptrs, i32 4, <24 x i1> %mask)
+  ret void
+}
+
+; The schedule model distinguishes dword and qword indices. Ensure TTI selects
+; the encoding that matches the GEP rather than one based only on data width.
+
+define <8 x i32> @gather_v8i32_i32_index(ptr %base, <8 x i32> %index, <8 x i1> %mask) {
+; THROUGHPUT-LABEL: 'gather_v8i32_i32_index'
+; THROUGHPUT: Cost Model: Found an estimated cost of 25 for instruction: %v = call
+; SIZE-LABEL: 'gather_v8i32_i32_index'
+; SIZE: Cost Model: Found an estimated cost of 1 for instruction: %v = call
+  %ptrs = getelementptr i32, ptr %base, <8 x i32> %index
+  %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x i32> poison)
+  ret <8 x i32> %v
+}
+
+define <8 x i32> @gather_v8i32_i64_index(ptr %base, <8 x i64> %index, <8 x i1> %mask) {
+; THROUGHPUT-LABEL: 'gather_v8i32_i64_index'
+; THROUGHPUT: Cost Model: Found an estimated cost of 26 for instruction: %v = call
+; SIZE-LABEL: 'gather_v8i32_i64_index'
+; SIZE: Cost Model: Found an estimated cost of 1 for instruction: %v = call
+  %ptrs = getelementptr i32, ptr %base, <8 x i64> %index
+  %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x i32> poison)
+  ret <8 x i32> %v
+}
+
+define void @scatter_v8i64_i32_index(<8 x i64> %src, ptr %base, <8 x i32> %index, <8 x i1> %mask) {
+; THROUGHPUT-LABEL: 'scatter_v8i64_i32_index'
+; THROUGHPUT: Cost Model: Found an estimated cost of 32 for instruction: call void
+; SIZE-LABEL: 'scatter_v8i64_i32_index'
+; SIZE: Cost Model: Found an estimated cost of 1 for instruction: call void
+  %ptrs = getelementptr i64, ptr %base, <8 x i32> %index
+  call void @llvm.masked.scatter.v8i64.v8p0(<8 x i64> %src, <8 x ptr> %ptrs, i32 8, <8 x i1> %mask)
+  ret void
+}
+
+define void @scatter_v8i64_i64_index(<8 x i64> %src, ptr %base, <8 x i64> %index, <8 x i1> %mask) {
+; THROUGHPUT-LABEL: 'scatter_v8i64_i64_index'
+; THROUGHPUT: Cost Model: Found an estimated cost of 35 for instruction: call void
+; SIZE-LABEL: 'scatter_v8i64_i64_index'
+; SIZE: Cost Model: Found an estimated cost of 1 for instruction: call void
+  %ptrs = getelementptr i64, ptr %base, <8 x i64> %index
+  call void @llvm.masked.scatter.v8i64.v8p0(<8 x i64> %src, <8 x ptr> %ptrs, i32 8, <8 x i1> %mask)
+  ret void
+}
+
+; Exercise function-level vector-width attributes while checking that a native
+; data shape pays its overhead once when qword indices require two sub-ops.
+
+define <16 x i32> @gather_v16i32_i32_index_prefer256(ptr %base, <16 x i32> %index, <16 x i1> %mask) #0 {
+; THROUGHPUT-LABEL: 'gather_v16i32_i32_index_prefer256'
+; THROUGHPUT: Cost Model: Found an estimated cost of 35 for instruction: %v = call
+; SIZE-LABEL: 'gather_v16i32_i32_index_prefer256'
+; SIZE: Cost Model: Found an estimated cost of 1 for instruction: %v = call
+  %ptrs = getelementptr i32, ptr %base, <16 x i32> %index
+  %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+  ret <16 x i32> %v
+}
+
+define <16 x i32> @gather_v16i32_i64_index_minlegal0(ptr %base, <16 x i64> %index, <16 x i1> %mask) #1 {
+; THROUGHPUT-LABEL: 'gather_v16i32_i64_index_minlegal0'
+; THROUGHPUT: Cost Model: Found an estimated cost of 37 for instruction: %v = call
+; SIZE-LABEL: 'gather_v16i32_i64_index_minlegal0'
+; SIZE: Cost Model: Found an estimated cost of 2 for instruction: %v = call
+  %ptrs = getelementptr i32, ptr %base, <16 x i64> %index
+  %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
+  ret <16 x i32> %v
+}
+
+declare <3 x i32> @llvm.masked.gather.v3i32.v3p0(<3 x ptr>, i32 immarg, <3 x i1>, <3 x i32>)
+declare <5 x i32> @llvm.masked.gather.v5i32.v5p0(<5 x ptr>, i32 immarg, <5 x i1>, <5 x i32>)
+declare <6 x i32> @llvm.masked.gather.v6i32.v6p0(<6 x ptr>, i32 immarg, <6 x i1>, <6 x i32>)
+declare <7 x i32> @llvm.masked.gather.v7i32.v7p0(<7 x ptr>, i32 immarg, <7 x i1>, <7 x i32>)
+declare <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr>, i32 immarg, <8 x i1>, <8 x i32>)
+declare <12 x i32> @llvm.masked.gather.v12i32.v12p0(<12 x ptr>, i32 immarg, <12 x i1>, <12 x i32>)
+declare <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr>, i32 immarg, <16 x i1>, <16 x i32>)
+declare <24 x i32> @llvm.masked.gather.v24i32.v24p0(<24 x ptr>, i32 immarg, <24 x i1>, <24 x i32>)
+declare void @llvm.masked.scatter.v24i32.v24p0(<24 x i32>, <24 x ptr>, i32 immarg, <24 x i1>)
+declare void @llvm.masked.scatter.v8i64.v8p0(<8 x i64>, <8 x ptr>, i32 immarg, <8 x i1>)
+
+attributes #0 = { "prefer-vector-width"="256" }
+attributes #1 = { "min-legal-vector-width"="0" }
diff --git a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-table.ll b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-table.ll
index bbc79850980f2..3e9a70f5ec13a 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-table.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-gather-scatter-cost-table.ll
@@ -81,15 +81,15 @@ define <4 x i32> @gather_v4i32(<4 x ptr> %ptrs, <4 x i1> %mask) {
 
 define <8 x i32> @gather_v8i32(<8 x ptr> %ptrs, <8 x i1> %mask) {
 ; ZNVER4-LABEL: 'gather_v8i32'
-; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 26 for instruction: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
 ; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x i32> %v
 ;
 ; ZNVER5-LABEL: 'gather_v8i32'
-; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 26 for instruction: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
 ; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x i32> %v
 ;
 ; ZNVER6-LABEL: 'gather_v8i32'
-; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 26 for instruction: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
 ; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x i32> %v
 ;
 ; ZNVER3-LABEL: 'gather_v8i32'
@@ -110,15 +110,15 @@ define <8 x i32> @gather_v8i32(<8 x ptr> %ptrs, <8 x i1> %mask) {
 
 define <16 x i32> @gather_v16i32(<16 x ptr> %ptrs, <16 x i1> %mask) {
 ; ZNVER4-LABEL: 'gather_v16i32'
-; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 37 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
 ; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
 ;
 ; ZNVER5-LABEL: 'gather_v16i32'
-; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 37 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
 ; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
 ;
 ; ZNVER6-LABEL: 'gather_v16i32'
-; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 37 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
 ; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
 ;
 ; ZNVER3-LABEL: 'gather_v16i32'
@@ -126,11 +126,11 @@ define <16 x i32> @gather_v16i32(<16 x ptr> %ptrs, <16 x i1> %mask) {
 ; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
 ;
 ; SKX-LABEL: 'gather_v16i32'
-; SKX-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 20 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
 ; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
 ;
 ; X8664V4-LABEL: 'gather_v16i32'
-; X8664V4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 20 for instruction: %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x i32> poison)
 ; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i32> %v
 ;
   %v = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x i32> poison)
@@ -292,15 +292,15 @@ define <4 x float> @gather_v4f32(<4 x ptr> %ptrs, <4 x i1> %mask) {
 
 define <8 x float> @gather_v8f32(<8 x ptr> %ptrs, <8 x i1> %mask) {
 ; ZNVER4-LABEL: 'gather_v8f32'
-; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x float> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 26 for instruction: %v = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x float> poison)
 ; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x float> %v
 ;
 ; ZNVER5-LABEL: 'gather_v8f32'
-; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x float> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 26 for instruction: %v = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x float> poison)
 ; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x float> %v
 ;
 ; ZNVER6-LABEL: 'gather_v8f32'
-; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 25 for instruction: %v = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x float> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 26 for instruction: %v = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x float> poison)
 ; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <8 x float> %v
 ;
 ; ZNVER3-LABEL: 'gather_v8f32'
@@ -321,15 +321,15 @@ define <8 x float> @gather_v8f32(<8 x ptr> %ptrs, <8 x i1> %mask) {
 
 define <16 x float> @gather_v16f32(<16 x ptr> %ptrs, <16 x i1> %mask) {
 ; ZNVER4-LABEL: 'gather_v16f32'
-; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 37 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
 ; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
 ;
 ; ZNVER5-LABEL: 'gather_v16f32'
-; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 37 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
 ; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
 ;
 ; ZNVER6-LABEL: 'gather_v16f32'
-; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 35 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 37 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
 ; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
 ;
 ; ZNVER3-LABEL: 'gather_v16f32'
@@ -337,11 +337,11 @@ define <16 x float> @gather_v16f32(<16 x ptr> %ptrs, <16 x i1> %mask) {
 ; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
 ;
 ; SKX-LABEL: 'gather_v16f32'
-; SKX-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 20 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
 ; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
 ;
 ; X8664V4-LABEL: 'gather_v16f32'
-; X8664V4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 20 for instruction: %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %ptrs, <16 x i1> %mask, <16 x float> poison)
 ; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x float> %v
 ;
   %v = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> %ptrs, i32 4, <16 x i1> %mask, <16 x float> poison)
@@ -445,15 +445,15 @@ define <8 x double> @gather_v8f64(<8 x ptr> %ptrs, <8 x i1> %mask) {
 
 define void @scatter_v4i32(<4 x i32> %src, <4 x ptr> %ptrs, <4 x i1> %mask) {
 ; ZNVER4-LABEL: 'scatter_v4i32'
-; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 19 for instruction: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
 ; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER5-LABEL: 'scatter_v4i32'
-; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 19 for instruction: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
 ; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER6-LABEL: 'scatter_v4i32'
-; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 19 for instruction: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
 ; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER3-LABEL: 'scatter_v4i32'
@@ -503,15 +503,15 @@ define void @scatter_v8i32(<8 x i32> %src, <8 x ptr> %ptrs, <8 x i1> %mask) {
 
 define void @scatter_v16i32(<16 x i32> %src, <16 x ptr> %ptrs, <16 x i1> %mask) {
 ; ZNVER4-LABEL: 'scatter_v16i32'
-; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 29 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
 ; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER5-LABEL: 'scatter_v16i32'
-; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 29 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
 ; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER6-LABEL: 'scatter_v16i32'
-; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 29 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
 ; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER3-LABEL: 'scatter_v16i32'
@@ -519,11 +519,11 @@ define void @scatter_v16i32(<16 x i32> %src, <16 x ptr> %ptrs, <16 x i1> %mask)
 ; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; SKX-LABEL: 'scatter_v16i32'
-; SKX-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 20 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
 ; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; X8664V4-LABEL: 'scatter_v16i32'
-; X8664V4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 20 for instruction: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
 ; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
   call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> %src, <16 x ptr> %ptrs, i32 4, <16 x i1> %mask)
@@ -598,15 +598,15 @@ define void @scatter_v8i64(<8 x i64> %src, <8 x ptr> %ptrs, <8 x i1> %mask) {
 
 define void @scatter_v4f32(<4 x float> %src, <4 x ptr> %ptrs, <4 x i1> %mask) {
 ; ZNVER4-LABEL: 'scatter_v4f32'
-; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 19 for instruction: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
 ; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER5-LABEL: 'scatter_v4f32'
-; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 19 for instruction: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
 ; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER6-LABEL: 'scatter_v4f32'
-; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 19 for instruction: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> %src, <4 x ptr> align 4 %ptrs, <4 x i1> %mask)
 ; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER3-LABEL: 'scatter_v4f32'
@@ -656,15 +656,15 @@ define void @scatter_v8f32(<8 x float> %src, <8 x ptr> %ptrs, <8 x i1> %mask) {
 
 define void @scatter_v16f32(<16 x float> %src, <16 x ptr> %ptrs, <16 x i1> %mask) {
 ; ZNVER4-LABEL: 'scatter_v16f32'
-; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 29 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
 ; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER5-LABEL: 'scatter_v16f32'
-; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 29 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
 ; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER6-LABEL: 'scatter_v16f32'
-; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 31 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 29 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
 ; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER3-LABEL: 'scatter_v16f32'
@@ -672,11 +672,11 @@ define void @scatter_v16f32(<16 x float> %src, <16 x ptr> %ptrs, <16 x i1> %mask
 ; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; SKX-LABEL: 'scatter_v16f32'
-; SKX-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 20 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
 ; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; X8664V4-LABEL: 'scatter_v16f32'
-; X8664V4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 20 for instruction: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> align 4 %ptrs, <16 x i1> %mask)
 ; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
   call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> %src, <16 x ptr> %ptrs, i32 4, <16 x i1> %mask)
@@ -1030,15 +1030,15 @@ define void @scatter_v16f32_gep(<16 x float> %src, ptr %base, <16 x i32> %idx, <
 ; regression that a flat out-of-table overhead would reintroduce.
 define <32 x i32> @gather_v32i32(<32 x ptr> %ptrs, <32 x i1> %mask) {
 ; ZNVER4-LABEL: 'gather_v32i32'
-; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 70 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 104 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
 ; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i32> %v
 ;
 ; ZNVER5-LABEL: 'gather_v32i32'
-; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 70 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 104 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
 ; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i32> %v
 ;
 ; ZNVER6-LABEL: 'gather_v32i32'
-; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 70 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 104 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
 ; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i32> %v
 ;
 ; ZNVER3-LABEL: 'gather_v32i32'
@@ -1046,11 +1046,11 @@ define <32 x i32> @gather_v32i32(<32 x ptr> %ptrs, <32 x i1> %mask) {
 ; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i32> %v
 ;
 ; SKX-LABEL: 'gather_v32i32'
-; SKX-NEXT:  Cost Model: Found an estimated cost of 34 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 40 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
 ; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i32> %v
 ;
 ; X8664V4-LABEL: 'gather_v32i32'
-; X8664V4-NEXT:  Cost Model: Found an estimated cost of 34 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 40 for instruction: %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> align 4 %ptrs, <32 x i1> %mask, <32 x i32> poison)
 ; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i32> %v
 ;
   %v = call <32 x i32> @llvm.masked.gather.v32i32.v32p0(<32 x ptr> %ptrs, i32 4, <32 x i1> %mask, <32 x i32> poison)
@@ -1059,15 +1059,15 @@ define <32 x i32> @gather_v32i32(<32 x ptr> %ptrs, <32 x i1> %mask) {
 
 define <64 x i32> @gather_v64i32(<64 x ptr> %ptrs, <64 x i1> %mask) {
 ; ZNVER4-LABEL: 'gather_v64i32'
-; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 140 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 208 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
 ; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <64 x i32> %v
 ;
 ; ZNVER5-LABEL: 'gather_v64i32'
-; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 140 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 208 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
 ; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <64 x i32> %v
 ;
 ; ZNVER6-LABEL: 'gather_v64i32'
-; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 140 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 208 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
 ; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <64 x i32> %v
 ;
 ; ZNVER3-LABEL: 'gather_v64i32'
@@ -1075,11 +1075,11 @@ define <64 x i32> @gather_v64i32(<64 x ptr> %ptrs, <64 x i1> %mask) {
 ; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <64 x i32> %v
 ;
 ; SKX-LABEL: 'gather_v64i32'
-; SKX-NEXT:  Cost Model: Found an estimated cost of 66 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 80 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
 ; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <64 x i32> %v
 ;
 ; X8664V4-LABEL: 'gather_v64i32'
-; X8664V4-NEXT:  Cost Model: Found an estimated cost of 66 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 80 for instruction: %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> align 4 %ptrs, <64 x i1> %mask, <64 x i32> poison)
 ; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <64 x i32> %v
 ;
   %v = call <64 x i32> @llvm.masked.gather.v64i32.v64p0(<64 x ptr> %ptrs, i32 4, <64 x i1> %mask, <64 x i32> poison)
@@ -1088,15 +1088,15 @@ define <64 x i32> @gather_v64i32(<64 x ptr> %ptrs, <64 x i1> %mask) {
 
 define void @scatter_v32i32(<32 x i32> %src, <32 x ptr> %ptrs, <32 x i1> %mask) {
 ; ZNVER4-LABEL: 'scatter_v32i32'
-; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 62 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 100 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
 ; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER5-LABEL: 'scatter_v32i32'
-; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 62 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 100 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
 ; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER6-LABEL: 'scatter_v32i32'
-; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 62 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 100 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
 ; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER3-LABEL: 'scatter_v32i32'
@@ -1104,11 +1104,11 @@ define void @scatter_v32i32(<32 x i32> %src, <32 x ptr> %ptrs, <32 x i1> %mask)
 ; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; SKX-LABEL: 'scatter_v32i32'
-; SKX-NEXT:  Cost Model: Found an estimated cost of 34 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 40 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
 ; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; X8664V4-LABEL: 'scatter_v32i32'
-; X8664V4-NEXT:  Cost Model: Found an estimated cost of 34 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 40 for instruction: call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> align 4 %ptrs, <32 x i1> %mask)
 ; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
   call void @llvm.masked.scatter.v32i32.v32p0(<32 x i32> %src, <32 x ptr> %ptrs, i32 4, <32 x i1> %mask)
@@ -1117,15 +1117,15 @@ define void @scatter_v32i32(<32 x i32> %src, <32 x ptr> %ptrs, <32 x i1> %mask)
 
 define void @scatter_v64i32(<64 x i32> %src, <64 x ptr> %ptrs, <64 x i1> %mask) {
 ; ZNVER4-LABEL: 'scatter_v64i32'
-; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 124 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
+; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 200 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
 ; ZNVER4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER5-LABEL: 'scatter_v64i32'
-; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 124 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
+; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 200 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
 ; ZNVER5-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER6-LABEL: 'scatter_v64i32'
-; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 124 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
+; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 200 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
 ; ZNVER6-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; ZNVER3-LABEL: 'scatter_v64i32'
@@ -1133,11 +1133,11 @@ define void @scatter_v64i32(<64 x i32> %src, <64 x ptr> %ptrs, <64 x i1> %mask)
 ; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; SKX-LABEL: 'scatter_v64i32'
-; SKX-NEXT:  Cost Model: Found an estimated cost of 66 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 80 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
 ; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
 ; X8664V4-LABEL: 'scatter_v64i32'
-; X8664V4-NEXT:  Cost Model: Found an estimated cost of 66 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 80 for instruction: call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> align 4 %ptrs, <64 x i1> %mask)
 ; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret void
 ;
   call void @llvm.masked.scatter.v64i32.v64p0(<64 x i32> %src, <64 x ptr> %ptrs, i32 4, <64 x i1> %mask)
@@ -1162,11 +1162,11 @@ define <16 x i64> @gather_v16i64(<16 x ptr> %ptrs, <16 x i1> %mask) {
 ; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i64> %v
 ;
 ; SKX-LABEL: 'gather_v16i64'
-; SKX-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x i64> @llvm.masked.gather.v16i64.v16p0(<16 x ptr> align 8 %ptrs, <16 x i1> %mask, <16 x i64> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 20 for instruction: %v = call <16 x i64> @llvm.masked.gather.v16i64.v16p0(<16 x ptr> align 8 %ptrs, <16 x i1> %mask, <16 x i64> poison)
 ; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i64> %v
 ;
 ; X8664V4-LABEL: 'gather_v16i64'
-; X8664V4-NEXT:  Cost Model: Found an estimated cost of 18 for instruction: %v = call <16 x i64> @llvm.masked.gather.v16i64.v16p0(<16 x ptr> align 8 %ptrs, <16 x i1> %mask, <16 x i64> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 20 for instruction: %v = call <16 x i64> @llvm.masked.gather.v16i64.v16p0(<16 x ptr> align 8 %ptrs, <16 x i1> %mask, <16 x i64> poison)
 ; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <16 x i64> %v
 ;
   %v = call <16 x i64> @llvm.masked.gather.v16i64.v16p0(<16 x ptr> %ptrs, i32 8, <16 x i1> %mask, <16 x i64> poison)
@@ -1191,11 +1191,11 @@ define <32 x i64> @gather_v32i64(<32 x ptr> %ptrs, <32 x i1> %mask) {
 ; ZNVER3-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i64> %v
 ;
 ; SKX-LABEL: 'gather_v32i64'
-; SKX-NEXT:  Cost Model: Found an estimated cost of 34 for instruction: %v = call <32 x i64> @llvm.masked.gather.v32i64.v32p0(<32 x ptr> align 8 %ptrs, <32 x i1> %mask, <32 x i64> poison)
+; SKX-NEXT:  Cost Model: Found an estimated cost of 40 for instruction: %v = call <32 x i64> @llvm.masked.gather.v32i64.v32p0(<32 x ptr> align 8 %ptrs, <32 x i1> %mask, <32 x i64> poison)
 ; SKX-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i64> %v
 ;
 ; X8664V4-LABEL: 'gather_v32i64'
-; X8664V4-NEXT:  Cost Model: Found an estimated cost of 34 for instruction: %v = call <32 x i64> @llvm.masked.gather.v32i64.v32p0(<32 x ptr> align 8 %ptrs, <32 x i1> %mask, <32 x i64> poison)
+; X8664V4-NEXT:  Cost Model: Found an estimated cost of 40 for instruction: %v = call <32 x i64> @llvm.masked.gather.v32i64.v32p0(<32 x ptr> align 8 %ptrs, <32 x i1> %mask, <32 x i64> poison)
 ; X8664V4-NEXT:  Cost Model: Found an estimated cost of 0 for instruction: ret <32 x i64> %v
 ;
   %v = call <32 x i64> @llvm.masked.gather.v32i64.v32p0(<32 x ptr> %ptrs, i32 8, <32 x i1> %mask, <32 x i64> poison)
diff --git a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll
index 5bf852c7ebfdd..f835a8aedbbcc 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost-inseltpoison.ll
@@ -840,20 +840,20 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret i32 0
 ;
 ; SKL-LABEL: 'masked_gather'
-; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8F64 = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x double> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8F64 = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x double> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F64 = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x double> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2F64 = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x double> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1F64 = call <1 x double> @llvm.masked.gather.v1f64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x double> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F32 = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2F32 = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x float> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8I64 = call <8 x i64> @llvm.masked.gather.v8i64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i64> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8I64 = call <8 x i64> @llvm.masked.gather.v8i64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I64 = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:107 CodeSize:139 Lat:235 SizeLat:139 for: %V32I16 = call <32 x i16> @llvm.masked.gather.v32i16.v32p0(<32 x ptr> align 1 undef, <32 x i1> %m32, <32 x i16> undef)
@@ -871,7 +871,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:32 SizeLat:20 for: %V4F64 = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x double> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F64 = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x double> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1F64 = call <1 x double> @llvm.masked.gather.v1f64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x double> undef)
-; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
+; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:32 SizeLat:20 for: %V4F32 = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x float> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F32 = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x float> undef)
@@ -879,7 +879,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:22 Lat:34 SizeLat:22 for: %V4I64 = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i64> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
-; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:17 CodeSize:21 Lat:33 SizeLat:21 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -898,7 +898,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F64 = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x double> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F64 = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x double> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1F64 = call <1 x double> @llvm.masked.gather.v1f64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x double> undef)
-; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
+; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F32 = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x float> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F32 = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x float> undef)
@@ -906,7 +906,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I64 = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i64> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
-; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -1094,7 +1094,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f64.v2p0(<2 x double> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1f64.v1p0(<1 x double> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8f32.v8p0(<8 x float> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f32.v2p0(<2 x float> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1102,7 +1102,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:22 Lat:22 SizeLat:22 for: call void @llvm.masked.scatter.v4i64.v4p0(<4 x i64> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:17 CodeSize:21 Lat:21 SizeLat:21 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1121,7 +1121,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f64.v2p0(<2 x double> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1f64.v1p0(<1 x double> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8f32.v8p0(<8 x float> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f32.v2p0(<2 x float> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1129,7 +1129,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4i64.v4p0(<4 x i64> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1909,7 +1909,7 @@ define <16 x float> @test_gather_16f32_const_mask(ptr %base, <16 x i32> %ind) {
 ; SKL-LABEL: 'test_gather_16f32_const_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_const_mask'
@@ -1953,7 +1953,7 @@ define <16 x float> @test_gather_16f32_var_mask(ptr %base, <16 x i32> %ind, <16
 ; SKL-LABEL: 'test_gather_16f32_var_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_var_mask'
@@ -1997,13 +1997,13 @@ define <16 x float> @test_gather_16f32_ra_var_mask(<16 x ptr> %ptrs, <16 x i32>
 ; SKL-LABEL: 'test_gather_16f32_ra_var_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, <16 x ptr> %ptrs, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_ra_var_mask'
 ; AVX512-NEXT:  Cost Model: Found costs of RThru:3 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; AVX512-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, <16 x ptr> %ptrs, <16 x i64> %sext_ind
-; AVX512-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; AVX512-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
 ; AVX512-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
   %sext_ind = sext <16 x i32> %ind to <16 x i64>
@@ -2051,7 +2051,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
 ; SKL-NEXT:  Cost Model: Found costs of RThru:1 CodeSize:1 Lat:3 SizeLat:2 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> poison, <16 x i32> zeroinitializer
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_const_mask2'
diff --git a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
index 32dd349bca659..ed1b534fac8f8 100644
--- a/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
+++ b/llvm/test/Analysis/CostModel/X86/masked-intrinsic-cost.ll
@@ -840,20 +840,20 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret i32 0
 ;
 ; SKL-LABEL: 'masked_gather'
-; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8F64 = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x double> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8F64 = call <8 x double> @llvm.masked.gather.v8f64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x double> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F64 = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x double> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2F64 = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x double> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1F64 = call <1 x double> @llvm.masked.gather.v1f64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x double> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F32 = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2F32 = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x float> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8I64 = call <8 x i64> @llvm.masked.gather.v8i64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i64> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8I64 = call <8 x i64> @llvm.masked.gather.v8i64.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I64 = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
-; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:2 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:12 CodeSize:2 Lat:36 SizeLat:12 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:4 CodeSize:1 Lat:10 SizeLat:4 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:107 CodeSize:139 Lat:235 SizeLat:139 for: %V32I16 = call <32 x i16> @llvm.masked.gather.v32i16.v32p0(<32 x ptr> align 1 undef, <32 x i1> %m32, <32 x i16> undef)
@@ -871,7 +871,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:32 SizeLat:20 for: %V4F64 = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x double> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F64 = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x double> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1F64 = call <1 x double> @llvm.masked.gather.v1f64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x double> undef)
-; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
+; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:32 SizeLat:20 for: %V4F32 = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x float> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F32 = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x float> undef)
@@ -879,7 +879,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:22 Lat:34 SizeLat:22 for: %V4I64 = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i64> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
-; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:17 CodeSize:21 Lat:33 SizeLat:21 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -898,7 +898,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F64 = call <4 x double> @llvm.masked.gather.v4f64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x double> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F64 = call <2 x double> @llvm.masked.gather.v2f64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x double> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1F64 = call <1 x double> @llvm.masked.gather.v1f64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x double> undef)
-; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
+; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16F32 = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x float> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8F32 = call <8 x float> @llvm.masked.gather.v8f32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x float> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4F32 = call <4 x float> @llvm.masked.gather.v4f32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x float> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:15 SizeLat:9 for: %V2F32 = call <2 x float> @llvm.masked.gather.v2f32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x float> undef)
@@ -906,7 +906,7 @@ define i32 @masked_gather(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m8
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I64 = call <4 x i64> @llvm.masked.gather.v4i64.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i64> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I64 = call <2 x i64> @llvm.masked.gather.v2i64.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i64> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:6 SizeLat:3 for: %V1I64 = call <1 x i64> @llvm.masked.gather.v1i64.v1p0(<1 x ptr> align 1 undef, <1 x i1> %m1, <1 x i64> undef)
-; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
+; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %V16I32 = call <16 x i32> @llvm.masked.gather.v16i32.v16p0(<16 x ptr> align 1 undef, <16 x i1> %m16, <16 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %V8I32 = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 1 undef, <8 x i1> %m8, <8 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:18 SizeLat:6 for: %V4I32 = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 1 undef, <4 x i1> %m4, <4 x i32> undef)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:16 SizeLat:10 for: %V2I32 = call <2 x i32> @llvm.masked.gather.v2i32.v2p0(<2 x ptr> align 1 undef, <2 x i1> %m2, <2 x i32> undef)
@@ -1094,7 +1094,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f64.v2p0(<2 x double> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1f64.v1p0(<1 x double> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8f32.v8p0(<8 x float> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:16 CodeSize:20 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f32.v2p0(<2 x float> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1102,7 +1102,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:22 Lat:22 SizeLat:22 for: call void @llvm.masked.scatter.v4i64.v4p0(<4 x i64> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; KNL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; KNL-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:17 CodeSize:21 Lat:21 SizeLat:21 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; KNL-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1121,7 +1121,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4f64.v4p0(<4 x double> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f64.v2p0(<2 x double> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1f64.v1p0(<1 x double> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16f32.v16p0(<16 x float> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8f32.v8p0(<8 x float> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4f32.v4p0(<4 x float> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:7 CodeSize:9 Lat:9 SizeLat:9 for: call void @llvm.masked.scatter.v2f32.v2p0(<2 x float> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1129,7 +1129,7 @@ define i32 @masked_scatter(<1 x i1> %m1, <2 x i1> %m2, <4 x i1> %m4, <8 x i1> %m
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4i64.v4p0(<4 x i64> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:2 CodeSize:3 Lat:3 SizeLat:3 for: call void @llvm.masked.scatter.v1i64.v1p0(<1 x i64> undef, <1 x ptr> align 1 undef, <1 x i1> %m1)
-; SKX-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:18 SizeLat:18 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
+; SKX-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:20 SizeLat:20 for: call void @llvm.masked.scatter.v16i32.v16p0(<16 x i32> undef, <16 x ptr> align 1 undef, <16 x i1> %m16)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> undef, <8 x ptr> align 1 undef, <8 x i1> %m8)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:6 CodeSize:1 Lat:6 SizeLat:6 for: call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> undef, <4 x ptr> align 1 undef, <4 x i1> %m4)
 ; SKX-NEXT:  Cost Model: Found costs of RThru:8 CodeSize:10 Lat:10 SizeLat:10 for: call void @llvm.masked.scatter.v2i32.v2p0(<2 x i32> undef, <2 x ptr> align 1 undef, <2 x i1> %m2)
@@ -1909,7 +1909,7 @@ define <16 x float> @test_gather_16f32_const_mask(ptr %base, <16 x i32> %ind) {
 ; SKL-LABEL: 'test_gather_16f32_const_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> splat (i1 true), <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_const_mask'
@@ -1953,7 +1953,7 @@ define <16 x float> @test_gather_16f32_var_mask(ptr %base, <16 x i32> %ind, <16
 ; SKL-LABEL: 'test_gather_16f32_var_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, ptr %base, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_var_mask'
@@ -1997,13 +1997,13 @@ define <16 x float> @test_gather_16f32_ra_var_mask(<16 x ptr> %ptrs, <16 x i32>
 ; SKL-LABEL: 'test_gather_16f32_ra_var_mask'
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, <16 x ptr> %ptrs, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_ra_var_mask'
 ; AVX512-NEXT:  Cost Model: Found costs of RThru:3 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; AVX512-NEXT:  Cost Model: Found costs of 0 for: %gep.v = getelementptr float, <16 x ptr> %ptrs, <16 x i64> %sext_ind
-; AVX512-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:2 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
+; AVX512-NEXT:  Cost Model: Found costs of RThru:20 CodeSize:2 Lat:68 SizeLat:20 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.v, <16 x i1> %mask, <16 x float> undef)
 ; AVX512-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
   %sext_ind = sext <16 x i32> %ind to <16 x i64>
@@ -2051,7 +2051,7 @@ define <16 x float> @test_gather_16f32_const_mask2(ptr %base, <16 x i32> %ind) {
 ; SKL-NEXT:  Cost Model: Found costs of RThru:1 CodeSize:1 Lat:3 SizeLat:2 for: %broadcast.splat = shufflevector <16 x ptr> %broadcast.splatinsert, <16 x ptr> undef, <16 x i32> zeroinitializer
 ; SKL-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:1 SizeLat:1 for: %sext_ind = sext <16 x i32> %ind to <16 x i64>
 ; SKL-NEXT:  Cost Model: Found costs of 0 for: %gep.random = getelementptr float, <16 x ptr> %broadcast.splat, <16 x i64> %sext_ind
-; SKL-NEXT:  Cost Model: Found costs of RThru:18 CodeSize:4 Lat:66 SizeLat:18 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
+; SKL-NEXT:  Cost Model: Found costs of RThru:24 CodeSize:4 Lat:72 SizeLat:24 for: %res = call <16 x float> @llvm.masked.gather.v16f32.v16p0(<16 x ptr> align 4 %gep.random, <16 x i1> splat (i1 true), <16 x float> undef)
 ; SKL-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <16 x float> %res
 ;
 ; AVX512-LABEL: 'test_gather_16f32_const_mask2'
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll
index 9191e95c0236d..2d5a30019bacd 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i32-with-i8-index.ll
@@ -50,9 +50,9 @@ define void @test() {
 ; AVX2-FASTGATHER:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB = load i32, ptr %inB, align 4
 ; AVX2-FASTGATHER:  Cost of 4 for VF 2: WIDEN ir<%valB> = load ir<%inB>
 ; AVX2-FASTGATHER:  Cost of 6 for VF 4: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER:  Cost of 10 for VF 8: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER:  Cost of 18 for VF 16: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER:  Cost of 34 for VF 32: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER:  Cost of 12 for VF 8: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER:  Cost of 24 for VF 16: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER:  Cost of 48 for VF 32: WIDEN ir<%valB> = load ir<%inB>
 ;
 ; AVX512-LABEL: 'test'
 ; AVX512:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB = load i32, ptr %inB, align 4
@@ -60,8 +60,8 @@ define void @test() {
 ; AVX512:  Cost of 13 for VF 4: REPLICATE ir<%valB> = load ir<%inB>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%valB> = load ir<%inB>
 ; AVX512:  Cost of 18 for VF 16: WIDEN ir<%valB> = load ir<%inB>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%valB> = load ir<%inB>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%valB> = load ir<%inB>
+; AVX512:  Cost of 36 for VF 32: WIDEN ir<%valB> = load ir<%inB>
+; AVX512:  Cost of 72 for VF 64: WIDEN ir<%valB> = load ir<%inB>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i64-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i64-with-i8-index.ll
index 189d5c837501d..ce5828a46eea7 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i64-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/gather-i64-with-i8-index.ll
@@ -50,18 +50,18 @@ define void @test() {
 ; AVX2-FASTGATHER:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB = load i64, ptr %inB, align 8
 ; AVX2-FASTGATHER:  Cost of 4 for VF 2: WIDEN ir<%valB> = load ir<%inB>
 ; AVX2-FASTGATHER:  Cost of 6 for VF 4: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER:  Cost of 10 for VF 8: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER:  Cost of 18 for VF 16: WIDEN ir<%valB> = load ir<%inB>
-; AVX2-FASTGATHER:  Cost of 34 for VF 32: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER:  Cost of 12 for VF 8: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER:  Cost of 24 for VF 16: WIDEN ir<%valB> = load ir<%inB>
+; AVX2-FASTGATHER:  Cost of 48 for VF 32: WIDEN ir<%valB> = load ir<%inB>
 ;
 ; AVX512-LABEL: 'test'
 ; AVX512:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB = load i64, ptr %inB, align 8
 ; AVX512:  Cost of 6 for VF 2: REPLICATE ir<%valB> = load ir<%inB>
 ; AVX512:  Cost of 14 for VF 4: REPLICATE ir<%valB> = load ir<%inB>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%valB> = load ir<%inB>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%valB> = load ir<%inB>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%valB> = load ir<%inB>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%valB> = load ir<%inB>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%valB> = load ir<%inB>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%valB> = load ir<%inB>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%valB> = load ir<%inB>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-2.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-2.ll
index 6232ac0e913ca..3b3c3dcb78041 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-2.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-2.ll
@@ -70,6 +70,9 @@ define void @test() {
 ; AVX512:  Cost of 34 for VF 32: INTERLEAVE-GROUP with factor 2, ir<%in0>
 ; AVX512:    ir<%v0> = load from index 0
 ; AVX512:    ir<%v1> = load from index 1
+; AVX512:  Cost of 148 for VF 64: INTERLEAVE-GROUP with factor 2, ir<%in0>
+; AVX512:    ir<%v0> = load from index 0
+; AVX512:    ir<%v1> = load from index 1
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-4.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-4.ll
index c2dfbe4f26870..679ff5ecec952 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-4.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-4.ll
@@ -77,6 +77,11 @@ define void @test() {
 ; AVX512:    ir<%v1> = load from index 1
 ; AVX512:    ir<%v2> = load from index 2
 ; AVX512:    ir<%v3> = load from index 3
+; AVX512:  Cost of 148 for VF 32: INTERLEAVE-GROUP with factor 4, ir<%in0>
+; AVX512:    ir<%v0> = load from index 0
+; AVX512:    ir<%v1> = load from index 1
+; AVX512:    ir<%v2> = load from index 2
+; AVX512:    ir<%v3> = load from index 3
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-8.ll
index a760b1ca0d1e7..eafe9cfd912c4 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f32-stride-8.ll
@@ -72,6 +72,15 @@ define void @test() {
 ; AVX512:    ir<%v5> = load from index 5
 ; AVX512:    ir<%v6> = load from index 6
 ; AVX512:    ir<%v7> = load from index 7
+; AVX512:  Cost of 148 for VF 16: INTERLEAVE-GROUP with factor 8, ir<%in0>
+; AVX512:    ir<%v0> = load from index 0
+; AVX512:    ir<%v1> = load from index 1
+; AVX512:    ir<%v2> = load from index 2
+; AVX512:    ir<%v3> = load from index 3
+; AVX512:    ir<%v4> = load from index 4
+; AVX512:    ir<%v5> = load from index 5
+; AVX512:    ir<%v6> = load from index 6
+; AVX512:    ir<%v7> = load from index 7
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-5.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-5.ll
index a3cf9835db917..7d18d05a344a8 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-5.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-5.ll
@@ -54,21 +54,21 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v2> = load ir<%in2>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v3> = load ir<%in3>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-6.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-6.ll
index 7ca2a371b05fa..909e0a9c8ed18 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-6.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-6.ll
@@ -75,24 +75,24 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v3> = load ir<%in3>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v4> = load ir<%in4>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v5> = load ir<%in5>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-7.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-7.ll
index b807cf344cbe5..a9cef6bd260a1 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-7.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-7.ll
@@ -59,27 +59,27 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v4> = load ir<%in4>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v5> = load ir<%in5>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v6> = load ir<%in6>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-8.ll
index ddbb47e9e7f55..ba511ad467b1a 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-f64-stride-8.ll
@@ -62,30 +62,30 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v5> = load ir<%in5>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v6> = load ir<%in6>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v7> = load ir<%in7>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v7> = load ir<%in7>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v7> = load ir<%in7>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v7> = load ir<%in7>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v7> = load ir<%in7>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v7> = load ir<%in7>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v7> = load ir<%in7>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2-indices-0u.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2-indices-0u.ll
index 3c8ae5ec46962..3d0e795d653a4 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2-indices-0u.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2-indices-0u.ll
@@ -56,6 +56,8 @@ define void @test() {
 ; AVX512:    ir<%v0> = load from index 0
 ; AVX512:  Cost of 19 for VF 32: INTERLEAVE-GROUP with factor 2, ir<%in0>
 ; AVX512:    ir<%v0> = load from index 0
+; AVX512:  Cost of 78 for VF 64: INTERLEAVE-GROUP with factor 2, ir<%in0>
+; AVX512:    ir<%v0> = load from index 0
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2.ll
index 34e448379a105..302e0cb447d65 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-2.ll
@@ -70,6 +70,9 @@ define void @test() {
 ; AVX512:  Cost of 34 for VF 32: INTERLEAVE-GROUP with factor 2, ir<%in0>
 ; AVX512:    ir<%v0> = load from index 0
 ; AVX512:    ir<%v1> = load from index 1
+; AVX512:  Cost of 148 for VF 64: INTERLEAVE-GROUP with factor 2, ir<%in0>
+; AVX512:    ir<%v0> = load from index 0
+; AVX512:    ir<%v1> = load from index 1
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-012u.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-012u.ll
index c21f1e670528b..41bd92713c386 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-012u.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-012u.ll
@@ -68,6 +68,10 @@ define void @test() {
 ; AVX512:    ir<%v0> = load from index 0
 ; AVX512:    ir<%v1> = load from index 1
 ; AVX512:    ir<%v2> = load from index 2
+; AVX512:  Cost of 113 for VF 32: INTERLEAVE-GROUP with factor 4, ir<%in0>
+; AVX512:    ir<%v0> = load from index 0
+; AVX512:    ir<%v1> = load from index 1
+; AVX512:    ir<%v2> = load from index 2
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-01uu.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-01uu.ll
index 6c43b50da956c..9a8ae47d51907 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-01uu.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-01uu.ll
@@ -59,6 +59,9 @@ define void @test() {
 ; AVX512:  Cost of 19 for VF 16: INTERLEAVE-GROUP with factor 4, ir<%in0>
 ; AVX512:    ir<%v0> = load from index 0
 ; AVX512:    ir<%v1> = load from index 1
+; AVX512:  Cost of 78 for VF 32: INTERLEAVE-GROUP with factor 4, ir<%in0>
+; AVX512:    ir<%v0> = load from index 0
+; AVX512:    ir<%v1> = load from index 1
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-0uuu.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-0uuu.ll
index 939e58c9381cf..d2e518a981de5 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-0uuu.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4-indices-0uuu.ll
@@ -50,8 +50,8 @@ define void @test() {
 ; AVX512:    ir<%v0> = load from index 0
 ; AVX512:  Cost of 8 for VF 16: INTERLEAVE-GROUP with factor 4, ir<%in0>
 ; AVX512:    ir<%v0> = load from index 0
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4.ll
index 2d56cdf49e0a1..bd7a63609acde 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-4.ll
@@ -77,6 +77,11 @@ define void @test() {
 ; AVX512:    ir<%v1> = load from index 1
 ; AVX512:    ir<%v2> = load from index 2
 ; AVX512:    ir<%v3> = load from index 3
+; AVX512:  Cost of 148 for VF 32: INTERLEAVE-GROUP with factor 4, ir<%in0>
+; AVX512:    ir<%v0> = load from index 0
+; AVX512:    ir<%v1> = load from index 1
+; AVX512:    ir<%v2> = load from index 2
+; AVX512:    ir<%v3> = load from index 3
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-8.ll
index a276757bce5f9..c64076fca0d62 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i32-stride-8.ll
@@ -72,6 +72,15 @@ define void @test() {
 ; AVX512:    ir<%v5> = load from index 5
 ; AVX512:    ir<%v6> = load from index 6
 ; AVX512:    ir<%v7> = load from index 7
+; AVX512:  Cost of 148 for VF 16: INTERLEAVE-GROUP with factor 8, ir<%in0>
+; AVX512:    ir<%v0> = load from index 0
+; AVX512:    ir<%v1> = load from index 1
+; AVX512:    ir<%v2> = load from index 2
+; AVX512:    ir<%v3> = load from index 3
+; AVX512:    ir<%v4> = load from index 4
+; AVX512:    ir<%v5> = load from index 5
+; AVX512:    ir<%v6> = load from index 6
+; AVX512:    ir<%v7> = load from index 7
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-2.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-2.ll
index 46836b631c9e0..d0a99efab706d 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-2.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-2.ll
@@ -63,10 +63,10 @@ define void @test() {
 ; AVX512:  Cost of 34 for VF 16: INTERLEAVE-GROUP with factor 2, ir<%in0>
 ; AVX512:    ir<%v0> = load from index 0
 ; AVX512:    ir<%v1> = load from index 1
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-4.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-4.ll
index 60b2cdcba7626..180ef142675f5 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-4.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-4.ll
@@ -68,18 +68,18 @@ define void @test() {
 ; AVX512:    ir<%v1> = load from index 1
 ; AVX512:    ir<%v2> = load from index 2
 ; AVX512:    ir<%v3> = load from index 3
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-5.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-5.ll
index 3c5ddeab269cf..65a27fc96f223 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-5.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-5.ll
@@ -54,21 +54,21 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v2> = load ir<%in2>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v3> = load ir<%in3>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-6.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-6.ll
index 611c1e28dbc42..d261e5242f8d8 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-6.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-6.ll
@@ -75,24 +75,24 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v3> = load ir<%in3>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v4> = load ir<%in4>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v5> = load ir<%in5>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-7.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-7.ll
index 3e9886ee10ac3..04c8db7a58357 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-7.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-7.ll
@@ -59,27 +59,27 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v4> = load ir<%in4>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v5> = load ir<%in5>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v6> = load ir<%in6>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-8.ll
index 6e0c384d73e7a..a2b5d1c4df2f8 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-load-i64-stride-8.ll
@@ -62,30 +62,30 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v5> = load ir<%in5>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v6> = load ir<%in6>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%v7> = load ir<%in7>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%v7> = load ir<%in7>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%v7> = load ir<%in7>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v0> = load ir<%in0>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v1> = load ir<%in1>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v2> = load ir<%in2>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v3> = load ir<%in3>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v4> = load ir<%in4>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v5> = load ir<%in5>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v6> = load ir<%in6>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%v7> = load ir<%in7>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%v7> = load ir<%in7>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%v7> = load ir<%in7>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v0> = load ir<%in0>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v1> = load ir<%in1>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v2> = load ir<%in2>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v3> = load ir<%in3>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v4> = load ir<%in4>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v5> = load ir<%in5>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v6> = load ir<%in6>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%v7> = load ir<%in7>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f32-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f32-stride-8.ll
index 4f3aa27971213..8d813d73120cb 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f32-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f32-stride-8.ll
@@ -172,6 +172,33 @@ define void @test() {
 ; AVX512:    store ir<%v5> to index 5
 ; AVX512:    store ir<%v6> to index 6
 ; AVX512:    store ir<%v7> to index 7
+; AVX512:  Cost of 148 for VF 16: INTERLEAVE-GROUP with factor 8, ir<%out0>
+; AVX512:    store ir<%v0> to index 0
+; AVX512:    store ir<%v1> to index 1
+; AVX512:    store ir<%v2> to index 2
+; AVX512:    store ir<%v3> to index 3
+; AVX512:    store ir<%v4> to index 4
+; AVX512:    store ir<%v5> to index 5
+; AVX512:    store ir<%v6> to index 6
+; AVX512:    store ir<%v7> to index 7
+; AVX512:  Cost of 296 for VF 32: INTERLEAVE-GROUP with factor 8, ir<%out0>
+; AVX512:    store ir<%v0> to index 0
+; AVX512:    store ir<%v1> to index 1
+; AVX512:    store ir<%v2> to index 2
+; AVX512:    store ir<%v3> to index 3
+; AVX512:    store ir<%v4> to index 4
+; AVX512:    store ir<%v5> to index 5
+; AVX512:    store ir<%v6> to index 6
+; AVX512:    store ir<%v7> to index 7
+; AVX512:  Cost of 592 for VF 64: INTERLEAVE-GROUP with factor 8, ir<%out0>
+; AVX512:    store ir<%v0> to index 0
+; AVX512:    store ir<%v1> to index 1
+; AVX512:    store ir<%v2> to index 2
+; AVX512:    store ir<%v3> to index 3
+; AVX512:    store ir<%v4> to index 4
+; AVX512:    store ir<%v5> to index 5
+; AVX512:    store ir<%v6> to index 6
+; AVX512:    store ir<%v7> to index 7
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-4.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-4.ll
index b636530a650ec..9799b132c3697 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-4.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-4.ll
@@ -114,6 +114,11 @@ define void @test() {
 ; AVX512:    store ir<%v1> to index 1
 ; AVX512:    store ir<%v2> to index 2
 ; AVX512:    store ir<%v3> to index 3
+; AVX512:  Cost of 272 for VF 64: INTERLEAVE-GROUP with factor 4, ir<%out0>
+; AVX512:    store ir<%v0> to index 0
+; AVX512:    store ir<%v1> to index 1
+; AVX512:    store ir<%v2> to index 2
+; AVX512:    store ir<%v3> to index 3
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-8.ll
index 11635959289f9..22d395647482d 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-f64-stride-8.ll
@@ -32,30 +32,30 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out5>, ir<%v5>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out6>, ir<%v6>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out7>, ir<%v7>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out0>, ir<%v0>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out1>, ir<%v1>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out2>, ir<%v2>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out3>, ir<%v3>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out4>, ir<%v4>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out5>, ir<%v5>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out6>, ir<%v6>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out7>, ir<%v7>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out0>, ir<%v0>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out1>, ir<%v1>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out2>, ir<%v2>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out3>, ir<%v3>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out4>, ir<%v4>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out5>, ir<%v5>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out6>, ir<%v6>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out7>, ir<%v7>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out0>, ir<%v0>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out1>, ir<%v1>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out2>, ir<%v2>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out3>, ir<%v3>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out4>, ir<%v4>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out5>, ir<%v5>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out6>, ir<%v6>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out7>, ir<%v7>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out0>, ir<%v0>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out1>, ir<%v1>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out2>, ir<%v2>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out3>, ir<%v3>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out4>, ir<%v4>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out5>, ir<%v5>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out6>, ir<%v6>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out7>, ir<%v7>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out0>, ir<%v0>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out1>, ir<%v1>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out2>, ir<%v2>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out3>, ir<%v3>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out4>, ir<%v4>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out5>, ir<%v5>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out6>, ir<%v6>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out7>, ir<%v7>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out0>, ir<%v0>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out1>, ir<%v1>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out2>, ir<%v2>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out3>, ir<%v3>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out4>, ir<%v4>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out5>, ir<%v5>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out6>, ir<%v6>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out7>, ir<%v7>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i32-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i32-stride-8.ll
index 50888201df108..506aaab0951e4 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i32-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i32-stride-8.ll
@@ -172,6 +172,33 @@ define void @test() {
 ; AVX512:    store ir<%v5> to index 5
 ; AVX512:    store ir<%v6> to index 6
 ; AVX512:    store ir<%v7> to index 7
+; AVX512:  Cost of 148 for VF 16: INTERLEAVE-GROUP with factor 8, ir<%out0>
+; AVX512:    store ir<%v> to index 0
+; AVX512:    store ir<%v1> to index 1
+; AVX512:    store ir<%v2> to index 2
+; AVX512:    store ir<%v3> to index 3
+; AVX512:    store ir<%v4> to index 4
+; AVX512:    store ir<%v5> to index 5
+; AVX512:    store ir<%v6> to index 6
+; AVX512:    store ir<%v7> to index 7
+; AVX512:  Cost of 296 for VF 32: INTERLEAVE-GROUP with factor 8, ir<%out0>
+; AVX512:    store ir<%v> to index 0
+; AVX512:    store ir<%v1> to index 1
+; AVX512:    store ir<%v2> to index 2
+; AVX512:    store ir<%v3> to index 3
+; AVX512:    store ir<%v4> to index 4
+; AVX512:    store ir<%v5> to index 5
+; AVX512:    store ir<%v6> to index 6
+; AVX512:    store ir<%v7> to index 7
+; AVX512:  Cost of 592 for VF 64: INTERLEAVE-GROUP with factor 8, ir<%out0>
+; AVX512:    store ir<%v> to index 0
+; AVX512:    store ir<%v1> to index 1
+; AVX512:    store ir<%v2> to index 2
+; AVX512:    store ir<%v3> to index 3
+; AVX512:    store ir<%v4> to index 4
+; AVX512:    store ir<%v5> to index 5
+; AVX512:    store ir<%v6> to index 6
+; AVX512:    store ir<%v7> to index 7
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-4.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-4.ll
index 2b2e22eb74630..805a50527749a 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-4.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-4.ll
@@ -114,6 +114,11 @@ define void @test() {
 ; AVX512:    store ir<%v1> to index 1
 ; AVX512:    store ir<%v2> to index 2
 ; AVX512:    store ir<%v3> to index 3
+; AVX512:  Cost of 272 for VF 64: INTERLEAVE-GROUP with factor 4, ir<%out0>
+; AVX512:    store ir<%v> to index 0
+; AVX512:    store ir<%v1> to index 1
+; AVX512:    store ir<%v2> to index 2
+; AVX512:    store ir<%v3> to index 3
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-8.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-8.ll
index 9150ee0e631bb..6e808e99d351f 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-8.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/interleaved-store-i64-stride-8.ll
@@ -170,30 +170,30 @@ define void @test() {
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out5>, ir<%v5>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out6>, ir<%v6>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out7>, ir<%v7>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out0>, ir<%v>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out1>, ir<%v1>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out2>, ir<%v2>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out3>, ir<%v3>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out4>, ir<%v4>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out5>, ir<%v5>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out6>, ir<%v6>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out7>, ir<%v7>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out0>, ir<%v>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out1>, ir<%v1>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out2>, ir<%v2>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out3>, ir<%v3>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out4>, ir<%v4>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out5>, ir<%v5>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out6>, ir<%v6>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out7>, ir<%v7>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out0>, ir<%v>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out1>, ir<%v1>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out2>, ir<%v2>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out3>, ir<%v3>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out4>, ir<%v4>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out5>, ir<%v5>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out6>, ir<%v6>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out7>, ir<%v7>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out0>, ir<%v>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out1>, ir<%v1>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out2>, ir<%v2>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out3>, ir<%v3>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out4>, ir<%v4>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out5>, ir<%v5>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out6>, ir<%v6>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out7>, ir<%v7>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out0>, ir<%v>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out1>, ir<%v1>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out2>, ir<%v2>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out3>, ir<%v3>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out4>, ir<%v4>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out5>, ir<%v5>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out6>, ir<%v6>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out7>, ir<%v7>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out0>, ir<%v>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out1>, ir<%v1>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out2>, ir<%v2>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out3>, ir<%v3>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out4>, ir<%v4>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out5>, ir<%v5>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out6>, ir<%v6>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out7>, ir<%v7>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll
index 59714eb573bba..8f5da77027970 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i32-with-i8-index.ll
@@ -43,9 +43,9 @@ define void @test() {
 ; AVX2-FASTGATHER:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB.loaded = load i32, ptr %inB, align 4
 ; AVX2-FASTGATHER:  Cost of 4 for VF 2: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ; AVX2-FASTGATHER:  Cost of 6 for VF 4: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER:  Cost of 10 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER:  Cost of 18 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER:  Cost of 34 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER:  Cost of 12 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER:  Cost of 24 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER:  Cost of 48 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ;
 ; AVX512-LABEL: 'test'
 ; AVX512:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB.loaded = load i32, ptr %inB, align 4
@@ -53,8 +53,8 @@ define void @test() {
 ; AVX512:  Cost of 17 for VF 4: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ; AVX512:  Cost of 18 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX512:  Cost of 36 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX512:  Cost of 72 for VF 64: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i64-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i64-with-i8-index.ll
index 094829fc8aa22..6801a549e19c4 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i64-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-gather-i64-with-i8-index.ll
@@ -43,18 +43,18 @@ define void @test() {
 ; AVX2-FASTGATHER:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB.loaded = load i64, ptr %inB, align 8
 ; AVX2-FASTGATHER:  Cost of 4 for VF 2: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ; AVX2-FASTGATHER:  Cost of 6 for VF 4: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER:  Cost of 10 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER:  Cost of 18 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX2-FASTGATHER:  Cost of 34 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER:  Cost of 12 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER:  Cost of 24 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX2-FASTGATHER:  Cost of 48 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ;
 ; AVX512-LABEL: 'test'
 ; AVX512:  LV: Found an estimated cost of 1 for VF 1 For instruction: %valB.loaded = load i64, ptr %inB, align 8
 ; AVX512:  Cost of 8 for VF 2: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ; AVX512:  Cost of 18 for VF 4: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ; AVX512:  Cost of 10 for VF 8: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX512:  Cost of 18 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX512:  Cost of 34 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
-; AVX512:  Cost of 66 for VF 64: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX512:  Cost of 20 for VF 16: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX512:  Cost of 40 for VF 32: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
+; AVX512:  Cost of 80 for VF 64: WIDEN ir<%valB.loaded> = load ir<%inB>, ir<%canLoad>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i32-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i32-with-i8-index.ll
index b9b07ae9d2711..6959fea2d512b 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i32-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i32-with-i8-index.ll
@@ -52,8 +52,8 @@ define void @test() {
 ; AVX512:  Cost of 10.5 for VF 4: REPLICATE store ir<%valB>, ir<%out>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
 ; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
+; AVX512:  Cost of 36 for VF 32: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
+; AVX512:  Cost of 72 for VF 64: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i64-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i64-with-i8-index.ll
index 2e01201fef3ed..41ae89933204e 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i64-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/masked-scatter-i64-with-i8-index.ll
@@ -51,9 +51,9 @@ define void @test() {
 ; AVX512:  Cost of 5 for VF 2: REPLICATE store ir<%valB>, ir<%out>
 ; AVX512:  Cost of 11 for VF 4: REPLICATE store ir<%valB>, ir<%out>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out>, ir<%valB>, ir<%canStore>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i32-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i32-with-i8-index.ll
index 043c14231846f..5e0b3277dd5e9 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i32-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i32-with-i8-index.ll
@@ -52,8 +52,8 @@ define void @test() {
 ; AVX512:  Cost of 13 for VF 4: REPLICATE store ir<%valB>, ir<%out>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out>, ir<%valB>
 ; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out>, ir<%valB>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out>, ir<%valB>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out>, ir<%valB>
+; AVX512:  Cost of 36 for VF 32: WIDEN store ir<%out>, ir<%valB>
+; AVX512:  Cost of 72 for VF 64: WIDEN store ir<%out>, ir<%valB>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i64-with-i8-index.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i64-with-i8-index.ll
index fe934b4b62dfa..fcbf6042dec14 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i64-with-i8-index.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/scatter-i64-with-i8-index.ll
@@ -51,9 +51,9 @@ define void @test() {
 ; AVX512:  Cost of 6 for VF 2: REPLICATE store ir<%valB>, ir<%out>
 ; AVX512:  Cost of 14 for VF 4: REPLICATE store ir<%valB>, ir<%out>
 ; AVX512:  Cost of 10 for VF 8: WIDEN store ir<%out>, ir<%valB>
-; AVX512:  Cost of 18 for VF 16: WIDEN store ir<%out>, ir<%valB>
-; AVX512:  Cost of 34 for VF 32: WIDEN store ir<%out>, ir<%valB>
-; AVX512:  Cost of 66 for VF 64: WIDEN store ir<%out>, ir<%valB>
+; AVX512:  Cost of 20 for VF 16: WIDEN store ir<%out>, ir<%valB>
+; AVX512:  Cost of 40 for VF 32: WIDEN store ir<%out>, ir<%valB>
+; AVX512:  Cost of 80 for VF 64: WIDEN store ir<%out>, ir<%valB>
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll b/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll
index bc2bea9708227..7f0ba19d3e7a8 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/cast-costs.ll
@@ -123,28 +123,28 @@ define void @replicate_sext(i32 %N, ptr %dst, ptr %src) #0 {
 ; CHECK-NEXT:    [[FOUND_CONFLICT:%.*]] = and i1 [[BOUND0]], [[BOUND1]]
 ; CHECK-NEXT:    br i1 [[FOUND_CONFLICT]], label %[[SCALAR_PH]], label %[[VECTOR_PH:.*]]
 ; CHECK:       [[VECTOR_PH]]:
-; CHECK-NEXT:    [[N_MOD_VF:%.*]] = and i32 [[TMP0]], 7
+; CHECK-NEXT:    [[N_MOD_VF:%.*]] = and i32 [[TMP0]], 3
 ; CHECK-NEXT:    [[TMP18:%.*]] = icmp eq i32 [[N_MOD_VF]], 0
-; CHECK-NEXT:    [[TMP19:%.*]] = select i1 [[TMP18]], i32 8, i32 [[N_MOD_VF]]
+; CHECK-NEXT:    [[TMP19:%.*]] = select i1 [[TMP18]], i32 4, i32 [[N_MOD_VF]]
 ; CHECK-NEXT:    [[N_VEC:%.*]] = sub i32 [[TMP0]], [[TMP19]]
 ; CHECK-NEXT:    [[TMP20:%.*]] = shl i32 [[N_VEC]], 2
 ; CHECK-NEXT:    [[TMP21:%.*]] = mul i32 [[N_VEC]], 3
 ; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
 ; CHECK:       [[VECTOR_BODY]]:
 ; CHECK-NEXT:    [[INDEX:%.*]] = phi i32 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT:    [[VEC_IND:%.*]] = phi <8 x i32> [ <i32 0, i32 3, i32 6, i32 9, i32 12, i32 15, i32 18, i32 21>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[VEC_IND:%.*]] = phi <4 x i32> [ <i32 0, i32 3, i32 6, i32 9>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
 ; CHECK-NEXT:    [[OFFSET_IDX:%.*]] = shl i32 [[INDEX]], 2
 ; CHECK-NEXT:    [[TMP22:%.*]] = sext i32 [[OFFSET_IDX]] to i64
 ; CHECK-NEXT:    [[TMP23:%.*]] = getelementptr nusw i32, ptr [[SRC]], i64 [[TMP22]]
-; CHECK-NEXT:    [[WIDE_VEC:%.*]] = load <32 x i32>, ptr [[TMP23]], align 4, !alias.scope [[META4:![0-9]+]]
-; CHECK-NEXT:    [[STRIDED_VEC:%.*]] = shufflevector <32 x i32> [[WIDE_VEC]], <32 x i32> poison, <8 x i32> <i32 0, i32 4, i32 8, i32 12, i32 16, i32 20, i32 24, i32 28>
-; CHECK-NEXT:    [[STRIDED_VEC2:%.*]] = shufflevector <32 x i32> [[WIDE_VEC]], <32 x i32> poison, <8 x i32> <i32 1, i32 5, i32 9, i32 13, i32 17, i32 21, i32 25, i32 29>
-; CHECK-NEXT:    [[TMP24:%.*]] = sext <8 x i32> [[VEC_IND]] to <8 x i64>
-; CHECK-NEXT:    [[WIDE_GEP:%.*]] = getelementptr i32, ptr [[DST]], <8 x i64> [[TMP24]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[STRIDED_VEC]], <8 x ptr> align 4 [[WIDE_GEP]], <8 x i1> splat (i1 true)), !alias.scope [[META7:![0-9]+]], !noalias [[META4]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[STRIDED_VEC2]], <8 x ptr> align 4 [[WIDE_GEP]], <8 x i1> splat (i1 true)), !alias.scope [[META7]], !noalias [[META4]]
-; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 8
-; CHECK-NEXT:    [[VEC_IND_NEXT]] = add <8 x i32> [[VEC_IND]], splat (i32 24)
+; CHECK-NEXT:    [[WIDE_VEC:%.*]] = load <16 x i32>, ptr [[TMP23]], align 4, !alias.scope [[META4:![0-9]+]]
+; CHECK-NEXT:    [[STRIDED_VEC:%.*]] = shufflevector <16 x i32> [[WIDE_VEC]], <16 x i32> poison, <4 x i32> <i32 0, i32 4, i32 8, i32 12>
+; CHECK-NEXT:    [[STRIDED_VEC9:%.*]] = shufflevector <16 x i32> [[WIDE_VEC]], <16 x i32> poison, <4 x i32> <i32 1, i32 5, i32 9, i32 13>
+; CHECK-NEXT:    [[TMP25:%.*]] = sext <4 x i32> [[VEC_IND]] to <4 x i64>
+; CHECK-NEXT:    [[TMP27:%.*]] = getelementptr i32, ptr [[DST]], <4 x i64> [[TMP25]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[STRIDED_VEC]], <4 x ptr> align 4 [[TMP27]], <4 x i1> splat (i1 true)), !alias.scope [[META7:![0-9]+]], !noalias [[META4]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[STRIDED_VEC9]], <4 x ptr> align 4 [[TMP27]], <4 x i1> splat (i1 true)), !alias.scope [[META7]], !noalias [[META4]]
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 4
+; CHECK-NEXT:    [[VEC_IND_NEXT]] = add <4 x i32> [[VEC_IND]], splat (i32 12)
 ; CHECK-NEXT:    [[TMP26:%.*]] = icmp eq i32 [[INDEX_NEXT]], [[N_VEC]]
 ; CHECK-NEXT:    br i1 [[TMP26]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP9:![0-9]+]]
 ; CHECK:       [[MIDDLE_BLOCK]]:
diff --git a/llvm/test/Transforms/LoopVectorize/X86/interleave-cost.ll b/llvm/test/Transforms/LoopVectorize/X86/interleave-cost.ll
index 13692e2c41a84..aa5a0f5be24be 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/interleave-cost.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/interleave-cost.ll
@@ -179,7 +179,7 @@ define void @geps_feeding_interleave_groups_with_reuse2(ptr %A, ptr %B, i64 %N)
 ; CHECK-NEXT:  [[ENTRY:.*]]:
 ; CHECK-NEXT:    [[TMP0:%.*]] = lshr i64 [[N]], 3
 ; CHECK-NEXT:    [[TMP1:%.*]] = add nuw nsw i64 [[TMP0]], 1
-; CHECK-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ule i64 [[TMP1]], 32
+; CHECK-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ule i64 [[TMP1]], 28
 ; CHECK-NEXT:    br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_SCEVCHECK:.*]]
 ; CHECK:       [[VECTOR_SCEVCHECK]]:
 ; CHECK-NEXT:    [[TMP2:%.*]] = lshr i64 [[N]], 3
@@ -256,48 +256,48 @@ define void @geps_feeding_interleave_groups_with_reuse2(ptr %A, ptr %B, i64 %N)
 ; CHECK-NEXT:    [[CONFLICT_RDX:%.*]] = or i1 [[FOUND_CONFLICT]], [[FOUND_CONFLICT40]]
 ; CHECK-NEXT:    br i1 [[CONFLICT_RDX]], label %[[SCALAR_PH]], label %[[VECTOR_PH:.*]]
 ; CHECK:       [[VECTOR_PH]]:
-; CHECK-NEXT:    [[N_MOD_VF:%.*]] = and i64 [[TMP1]], 7
+; CHECK-NEXT:    [[N_MOD_VF:%.*]] = and i64 [[TMP1]], 3
 ; CHECK-NEXT:    [[TMP48:%.*]] = icmp eq i64 [[N_MOD_VF]], 0
-; CHECK-NEXT:    [[TMP49:%.*]] = select i1 [[TMP48]], i64 8, i64 [[N_MOD_VF]]
+; CHECK-NEXT:    [[TMP49:%.*]] = select i1 [[TMP48]], i64 4, i64 [[N_MOD_VF]]
 ; CHECK-NEXT:    [[N_VEC:%.*]] = sub i64 [[TMP1]], [[TMP49]]
 ; CHECK-NEXT:    [[TMP50:%.*]] = shl i64 [[N_VEC]], 3
 ; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
 ; CHECK:       [[VECTOR_BODY]]:
 ; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT:    [[VEC_IND:%.*]] = phi <8 x i64> [ <i64 0, i64 8, i64 16, i64 24, i64 32, i64 40, i64 48, i64 56>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 8, i64 16, i64 24>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
 ; CHECK-NEXT:    [[TMP51:%.*]] = shl nuw i64 [[INDEX]], 3
 ; CHECK-NEXT:    [[TMP52:%.*]] = lshr exact i64 [[TMP51]], 1
 ; CHECK-NEXT:    [[TMP53:%.*]] = getelementptr nusw i32, ptr [[B]], i64 [[TMP52]]
-; CHECK-NEXT:    [[WIDE_VEC:%.*]] = load <32 x i32>, ptr [[TMP53]], align 4, !alias.scope [[META3:![0-9]+]], !noalias [[META6:![0-9]+]]
-; CHECK-NEXT:    [[STRIDED_VEC:%.*]] = shufflevector <32 x i32> [[WIDE_VEC]], <32 x i32> poison, <8 x i32> <i32 0, i32 4, i32 8, i32 12, i32 16, i32 20, i32 24, i32 28>
-; CHECK-NEXT:    [[STRIDED_VEC17:%.*]] = shufflevector <32 x i32> [[WIDE_VEC]], <32 x i32> poison, <8 x i32> <i32 1, i32 5, i32 9, i32 13, i32 17, i32 21, i32 25, i32 29>
-; CHECK-NEXT:    [[WIDE_GEP:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[VEC_IND]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[STRIDED_VEC]], <8 x ptr> align 4 [[WIDE_GEP]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[TMP55:%.*]] = or disjoint <8 x i64> [[VEC_IND]], splat (i64 1)
-; CHECK-NEXT:    [[WIDE_GEP18:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[TMP55]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> zeroinitializer, <8 x ptr> align 4 [[WIDE_GEP18]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[TMP56:%.*]] = or disjoint <8 x i64> [[VEC_IND]], splat (i64 2)
-; CHECK-NEXT:    [[WIDE_GEP19:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[TMP56]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[STRIDED_VEC17]], <8 x ptr> align 4 [[WIDE_GEP19]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[TMP57:%.*]] = or disjoint <8 x i64> [[VEC_IND]], splat (i64 3)
-; CHECK-NEXT:    [[WIDE_GEP20:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[TMP57]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> zeroinitializer, <8 x ptr> align 4 [[WIDE_GEP20]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[TMP58:%.*]] = or disjoint <8 x i64> [[VEC_IND]], splat (i64 4)
-; CHECK-NEXT:    [[WIDE_GEP21:%.*]] = getelementptr i32, ptr [[B]], <8 x i64> [[VEC_IND]]
-; CHECK-NEXT:    [[WIDE_MASKED_GATHER:%.*]] = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 [[WIDE_GEP21]], <8 x i1> splat (i1 true), <8 x i32> poison), !alias.scope [[META8:![0-9]+]], !noalias [[META6]]
-; CHECK-NEXT:    [[WIDE_GEP22:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[TMP58]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> [[WIDE_MASKED_GATHER]], <8 x ptr> align 4 [[WIDE_GEP22]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[TMP59:%.*]] = or disjoint <8 x i64> [[VEC_IND]], splat (i64 5)
-; CHECK-NEXT:    [[WIDE_GEP23:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[TMP59]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> zeroinitializer, <8 x ptr> align 4 [[WIDE_GEP23]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[TMP60:%.*]] = or disjoint <8 x i64> [[VEC_IND]], splat (i64 6)
-; CHECK-NEXT:    [[WIDE_GEP24:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[TMP60]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> zeroinitializer, <8 x ptr> align 4 [[WIDE_GEP24]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[TMP64:%.*]] = or disjoint <8 x i64> [[VEC_IND]], splat (i64 7)
-; CHECK-NEXT:    [[WIDE_GEP25:%.*]] = getelementptr i32, ptr [[A]], <8 x i64> [[TMP64]]
-; CHECK-NEXT:    call void @llvm.masked.scatter.v8i32.v8p0(<8 x i32> zeroinitializer, <8 x ptr> align 4 [[WIDE_GEP25]], <8 x i1> splat (i1 true)), !alias.scope [[META6]]
-; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 8
-; CHECK-NEXT:    [[VEC_IND_NEXT]] = add nuw nsw <8 x i64> [[VEC_IND]], splat (i64 64)
+; CHECK-NEXT:    [[WIDE_VEC:%.*]] = load <16 x i32>, ptr [[TMP53]], align 4, !alias.scope [[META3:![0-9]+]], !noalias [[META6:![0-9]+]]
+; CHECK-NEXT:    [[STRIDED_VEC:%.*]] = shufflevector <16 x i32> [[WIDE_VEC]], <16 x i32> poison, <4 x i32> <i32 0, i32 4, i32 8, i32 12>
+; CHECK-NEXT:    [[STRIDED_VEC41:%.*]] = shufflevector <16 x i32> [[WIDE_VEC]], <16 x i32> poison, <4 x i32> <i32 1, i32 5, i32 9, i32 13>
+; CHECK-NEXT:    [[WIDE_GEP:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[VEC_IND]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[STRIDED_VEC]], <4 x ptr> align 4 [[WIDE_GEP]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[TMP54:%.*]] = or disjoint <4 x i64> [[VEC_IND]], splat (i64 1)
+; CHECK-NEXT:    [[WIDE_GEP42:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[TMP54]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> zeroinitializer, <4 x ptr> align 4 [[WIDE_GEP42]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[TMP55:%.*]] = or disjoint <4 x i64> [[VEC_IND]], splat (i64 2)
+; CHECK-NEXT:    [[WIDE_GEP43:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[TMP55]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[STRIDED_VEC41]], <4 x ptr> align 4 [[WIDE_GEP43]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[TMP56:%.*]] = or disjoint <4 x i64> [[VEC_IND]], splat (i64 3)
+; CHECK-NEXT:    [[WIDE_GEP44:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[TMP56]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> zeroinitializer, <4 x ptr> align 4 [[WIDE_GEP44]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[TMP57:%.*]] = or disjoint <4 x i64> [[VEC_IND]], splat (i64 4)
+; CHECK-NEXT:    [[WIDE_GEP45:%.*]] = getelementptr i32, ptr [[B]], <4 x i64> [[VEC_IND]]
+; CHECK-NEXT:    [[WIDE_MASKED_GATHER:%.*]] = call <4 x i32> @llvm.masked.gather.v4i32.v4p0(<4 x ptr> align 4 [[WIDE_GEP45]], <4 x i1> splat (i1 true), <4 x i32> poison), !alias.scope [[META8:![0-9]+]], !noalias [[META6]]
+; CHECK-NEXT:    [[WIDE_GEP46:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[TMP57]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> [[WIDE_MASKED_GATHER]], <4 x ptr> align 4 [[WIDE_GEP46]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[TMP58:%.*]] = or disjoint <4 x i64> [[VEC_IND]], splat (i64 5)
+; CHECK-NEXT:    [[WIDE_GEP47:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[TMP58]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> zeroinitializer, <4 x ptr> align 4 [[WIDE_GEP47]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[TMP59:%.*]] = or disjoint <4 x i64> [[VEC_IND]], splat (i64 6)
+; CHECK-NEXT:    [[WIDE_GEP48:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[TMP59]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> zeroinitializer, <4 x ptr> align 4 [[WIDE_GEP48]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[TMP60:%.*]] = or disjoint <4 x i64> [[VEC_IND]], splat (i64 7)
+; CHECK-NEXT:    [[WIDE_GEP49:%.*]] = getelementptr i32, ptr [[A]], <4 x i64> [[TMP60]]
+; CHECK-NEXT:    call void @llvm.masked.scatter.v4i32.v4p0(<4 x i32> zeroinitializer, <4 x ptr> align 4 [[WIDE_GEP49]], <4 x i1> splat (i1 true)), !alias.scope [[META6]]
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT:    [[VEC_IND_NEXT]] = add nuw nsw <4 x i64> [[VEC_IND]], splat (i64 32)
 ; CHECK-NEXT:    [[TMP61:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
 ; CHECK-NEXT:    br i1 [[TMP61]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP10:![0-9]+]]
 ; CHECK:       [[MIDDLE_BLOCK]]:



More information about the llvm-commits mailing list