[llvm] [X86][SchedModel] Add masked gather/scatter overrides to Znver4Model (PR #212997)
via llvm-commits
llvm-commits at lists.llvm.org
Fri Sep 18 10:55:54 PDT 2026
================
@@ -509,6 +509,130 @@ defm : Zn4WriteResInt<WriteLoad, [Zn4AGU012, Zn4Load], !add(Znver4Model.LoadLate
// Does not cost anything by itself, only has latency, matching that of the WriteLoad,
defm : Zn4WriteResInt<WriteVecMaskedGatherWriteback, [], !add(Znver4Model.LoadLatency, 1), [], 0>;
+// AVX-512 masked GATHER / SCATTER. Microcoded, so entries are keyed by
+// (#elements, element width, index width); the two index widths share an entry
+// only where they measure alike. Throughput and latency were measured on
+// Znver4 with every mask bit set, and are accurate to about a cycle; scatter
+// latency is estimated rather than timed. Znver4Model also backs znver5 and
+// znver6, which run 25-55% faster and are pessimistic here.
+//
+// The FP occupancies are the uops.info Zen4 port distributions for the same
+// encodings, with the FP0-3 subgroups (FP01/FP12/FP23/FP123/FP0123) summed
+// into Zn4FPU0123. A scatter's element stores are charged to Zn4FPSt, which
+// keeps them subject to the one-FP-store-per-cycle limit; Zn4FP45 then carries
+// only the remaining FP4/5 uops, so the two together come to the measured
+// FP4/5 total.
+//
+// Those occupancies alone fall short of the measured throughput -- the
+// microcode sequencer, not any one pipe, is the limit -- so Zn4UcodeGS carries
+// the fitted residual, and a retune edits only that column. Charging it there
+// rather than to the FP pipes keeps a gather/scatter from falsely serialising
+// independent FP work. Zn4UcodeGS is abstract, not four pipes; it is declared
+// 4-wide only so that round(tput * 4) keeps the measurement's resolution, and
+// named to sort last so that adding it renumbers no existing resource.
+def Zn4UcodeGS : ProcResource<4>;
+
+class Zn4GatherRes<list<int> cyc, int lat, int uops>
+ : SchedWriteRes<[Zn4AGU012, Zn4Load, Zn4FPLd01, Zn4FPU0123, Zn4UcodeGS]> {
+ let ReleaseAtCycles = cyc; let Latency = lat; let NumMicroOps = uops;
+}
+class Zn4ScatterRes<list<int> cyc, int lat, int uops>
+ : SchedWriteRes<[Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FP45, Zn4FPU0123,
+ Zn4UcodeGS]> {
+ let ReleaseAtCycles = cyc; let Latency = lat; let NumMicroOps = uops;
+}
+// ReleaseAtCycles is [Zn4AGU012, Zn4Load, Zn4FPLd01, Zn4FPU0123, Zn4UcodeGS].
+def Zn4WriteVPGATHERQDZ128 : Zn4GatherRes<[ 2, 2, 4, 4, 16], 13, 18>;
+def : InstRW<[Zn4WriteVPGATHERQDZ128, WriteVecMaskedGatherWriteback],
+ (instrs VPGATHERQDZ128rm, VGATHERQPSZ128rm)>;
+def Zn4WriteVPGATHERDDZ128 : Zn4GatherRes<[ 4, 4, 6, 5, 20], 15, 24>;
+def : InstRW<[Zn4WriteVPGATHERDDZ128, WriteVecMaskedGatherWriteback],
+ (instrs VPGATHERDDZ128rm, VGATHERDPSZ128rm,
+ VPGATHERQDZ256rm, VGATHERQPSZ256rm)>;
+def Zn4WriteVPGATHERDDZ256 : Zn4GatherRes<[ 8, 8, 10, 9, 37], 20, 41>;
+def : InstRW<[Zn4WriteVPGATHERDDZ256, WriteVecMaskedGatherWriteback],
+ (instrs VPGATHERDDZ256rm, VGATHERDPSZ256rm)>;
+def Zn4WriteVPGATHERQDZ : Zn4GatherRes<[ 8, 8, 10, 12, 39], 24, 46>;
+def : InstRW<[Zn4WriteVPGATHERQDZ, WriteVecMaskedGatherWriteback],
+ (instrs VPGATHERQDZrm, VGATHERQPSZrm)>;
+def Zn4WriteVPGATHERDDZ : Zn4GatherRes<[ 16, 16, 18, 22, 70], 32, 81>;
+def : InstRW<[Zn4WriteVPGATHERDDZ, WriteVecMaskedGatherWriteback],
+ (instrs VPGATHERDDZrm, VGATHERDPSZrm)>;
+def Zn4WriteVPGATHERQQZ128 : Zn4GatherRes<[ 2, 2, 4, 3, 16], 12, 17>;
+def : InstRW<[Zn4WriteVPGATHERQQZ128, WriteVecMaskedGatherWriteback],
+ (instrs VPGATHERQQZ128rm, VGATHERQPDZ128rm,
+ VPGATHERDQZ128rm, VGATHERDPDZ128rm)>;
+def Zn4WriteVPGATHERQQZ256 : Zn4GatherRes<[ 4, 4, 6, 5, 20], 16, 24>;
+def : InstRW<[Zn4WriteVPGATHERQQZ256, WriteVecMaskedGatherWriteback],
+ (instrs VPGATHERQQZ256rm, VGATHERQPDZ256rm,
+ VPGATHERDQZ256rm, VGATHERDPDZ256rm)>;
+def Zn4WriteVPGATHERQQZ : Zn4GatherRes<[ 8, 8, 10, 14, 39], 25, 48>;
+def : InstRW<[Zn4WriteVPGATHERQQZ, WriteVecMaskedGatherWriteback],
+ (instrs VPGATHERQQZrm, VGATHERQPDZrm)>;
+def Zn4WriteVPGATHERDQZ : Zn4GatherRes<[ 8, 8, 10, 13, 38], 22, 46>;
+def : InstRW<[Zn4WriteVPGATHERDQZ, WriteVecMaskedGatherWriteback],
+ (instrs VPGATHERDQZrm, VGATHERDPDZrm)>;
+
+// AVX2 (VEX) gathers. AVX2 has no scatter and no VEX form exceeds 8 x i32. All
+// but 8 x i32 measure like the EVEX entry for their shape and reuse it; 8 x i32
+// is genuinely cheaper (8.35 vs 9.19 cyc) and gets its own entry. An InstRW
+// replaces the whole Sched<> list, so each one must restate the mask-writeback
+// def that X86InstrSSE.td attaches. That def is encoding-neutral and models the
+// mask clobber at the generic load latency, 5, for VEX and EVEX alike, while
+// the data result these entries measure is 12 to 32; narrowing that gap is
+// future work for both encodings.
+def : InstRW<[Zn4WriteVPGATHERQDZ128, WriteVecMaskedGatherWriteback],
+ (instrs VPGATHERQDrm, VGATHERQPSrm)>;
+def : InstRW<[Zn4WriteVPGATHERDDZ128, WriteVecMaskedGatherWriteback],
+ (instrs VPGATHERDDrm, VGATHERDPSrm,
+ VPGATHERQDYrm, VGATHERQPSYrm)>;
+def Zn4WriteVPGATHERDDY : Zn4GatherRes<[ 8, 8, 9, 17, 33], 21, 42>;
+def : InstRW<[Zn4WriteVPGATHERDDY, WriteVecMaskedGatherWriteback],
+ (instrs VPGATHERDDYrm, VGATHERDPSYrm)>;
+def : InstRW<[Zn4WriteVPGATHERQQZ128, WriteVecMaskedGatherWriteback],
+ (instrs VPGATHERQQrm, VGATHERQPDrm,
+ VPGATHERDQrm, VGATHERDPDrm)>;
+def : InstRW<[Zn4WriteVPGATHERQQZ256, WriteVecMaskedGatherWriteback],
+ (instrs VPGATHERQQYrm, VGATHERQPDYrm,
+ VPGATHERDQYrm, VGATHERDPDYrm)>;
+
+// ReleaseAtCycles is [Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FP45, Zn4FPU0123,
+// Zn4UcodeGS].
+def Zn4WriteVPSCATTERQDZ128 : Zn4ScatterRes<[ 2, 2, 2, 4, 1, 16], 6, 17>;
----------------
MattPD wrote:
Confirmed: The description now states the tied mask dependence and the six-cycle delay.
https://github.com/llvm/llvm-project/pull/212997
More information about the llvm-commits
mailing list