[llvm] [X86][CostModel] Cost gathers by the instructions CodeGen emits (PR #220565)

Sumukh J Bharadwaj via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 8 06:23:53 PDT 2026


================
@@ -0,0 +1,91 @@
+; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --version 6
+; Costs for masked gather/scatter operations that need more than one register's
+; worth of pointers, across every cost kind.
+;
+; Two properties are pinned here. A vector length that is not a multiple of its
+; split factor must still be charged for all of its lanes. And the index width
+; is chosen once for the whole operation, so a length wide enough to qualify for
+; narrowing keeps the narrow index in each of its parts, rather than reverting
+; to pointer width because an individual part is too short to qualify.
+
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=skylake-avx512 | FileCheck %s --check-prefix=SKX
+; RUN: opt < %s -S -mtriple=x86_64-unknown-linux-gnu -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mcpu=haswell | FileCheck %s --check-prefix=AVX2
+
+; A length that divides its split factor: unchanged, and the reference point for
+; the two cases below.
+define <8 x i32> @gather_v8i32(<8 x ptr> %ptrs, <8 x i1> %mask) {
+; SKX-LABEL: 'gather_v8i32'
+; SKX-NEXT:  Cost Model: Found costs of RThru:10 CodeSize:1 Lat:34 SizeLat:10 for: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <8 x i32> %v
+;
+; AVX2-LABEL: 'gather_v8i32'
+; AVX2-NEXT:  Cost Model: Found costs of RThru:28 CodeSize:36 Lat:60 SizeLat:36 for: %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> align 4 %ptrs, <8 x i1> %mask, <8 x i32> poison)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <8 x i32> %v
+;
+  %v = call <8 x i32> @llvm.masked.gather.v8i32.v8p0(<8 x ptr> %ptrs, i32 4, <8 x i1> %mask, <8 x i32> poison)
+  ret <8 x i32> %v
+}
+
+; Remainder lanes: v9 splits into two parts but is not a multiple of two, so the
+; ninth lane must not be dropped from the body term.
+define <9 x i32> @gather_v9i32(<9 x ptr> %ptrs, <9 x i1> %mask) {
+; SKX-LABEL: 'gather_v9i32'
+; SKX-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:40 SizeLat:13 for: %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 4 %ptrs, <9 x i1> %mask, <9 x i32> poison)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <9 x i32> %v
+;
+; AVX2-LABEL: 'gather_v9i32'
+; AVX2-NEXT:  Cost Model: Found costs of RThru:32 CodeSize:41 Lat:68 SizeLat:41 for: %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> align 4 %ptrs, <9 x i1> %mask, <9 x i32> poison)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret <9 x i32> %v
+;
+  %v = call <9 x i32> @llvm.masked.gather.v9i32.v9p0(<9 x ptr> %ptrs, i32 4, <9 x i1> %mask, <9 x i32> poison)
+  ret <9 x i32> %v
+}
+
+define void @scatter_v9i32(<9 x i32> %val, <9 x ptr> %ptrs, <9 x i1> %mask) {
+; SKX-LABEL: 'scatter_v9i32'
+; SKX-NEXT:  Cost Model: Found costs of RThru:13 CodeSize:2 Lat:13 SizeLat:13 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> align 4 %ptrs, <9 x i1> %mask)
+; SKX-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+; AVX2-LABEL: 'scatter_v9i32'
+; AVX2-NEXT:  Cost Model: Found costs of RThru:31 CodeSize:40 Lat:40 SizeLat:40 for: call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> align 4 %ptrs, <9 x i1> %mask)
+; AVX2-NEXT:  Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret void
+;
+  call void @llvm.masked.scatter.v9i32.v9p0(<9 x i32> %val, <9 x ptr> %ptrs, i32 4, <9 x i1> %mask)
+  ret void
+}
+
+; Index width across a split: the GEP indices are 32-bit, and v24 is wide enough
+; to qualify for narrowing, so all three parts are priced with a dword index
----------------
amd-subharad wrote:

This has been fixed as well

https://github.com/llvm/llvm-project/pull/220565


More information about the llvm-commits mailing list