[llvm] [X86][SchedModel] Add masked gather/scatter overrides to Znver4Model (PR #212997)

via llvm-commits llvm-commits at lists.llvm.org
Fri Sep 18 10:55:43 PDT 2026


================
@@ -509,6 +526,248 @@ defm : Zn4WriteResInt<WriteLoad, [Zn4AGU012, Zn4Load], !add(Znver4Model.LoadLate
 // Does not cost anything by itself, only has latency, matching that of the WriteLoad,
 defm : Zn4WriteResInt<WriteVecMaskedGatherWriteback, [], !add(Znver4Model.LoadLatency, 1), [], 0>;
 
+// AVX-512 masked GATHER / SCATTER. Microcoded, so entries are keyed by
+// (#elements, element width, index width); the two index widths share an entry
+// only where they measure alike. Throughput and latency were measured on
+// Znver4 with every mask bit set, and are accurate to about a cycle; scatter
+// latency is estimated rather than timed. A scatter's mask operand is tied, so
+// that estimate is what the model makes a mask consumer wait for -- a kmov
+// reading the mask after a VPSCATTERQQZ128 stalls behind its 6 cycles -- and
+// it is the one number here that no measurement backs. Znver4Model also backs
+// znver5 and znver6, which run 25-55% faster and are pessimistic here.
+//
+// The FP occupancies are the uops.info Zen4 port distributions for the same
+// encodings, kept at the eligibility they were reported with: a uop restricted
+// to FP1/FP2 is charged to Zn4FPU12 and not to all four pipes, so the model
+// cannot schedule it somewhere the hardware would not. Several of these
+// sequences are genuinely narrow -- a 512-bit VPSCATTERQQ has no FP0-eligible
+// uop at all -- and summing them into Zn4FPU0123 would invent that
+// eligibility. A scatter's element stores are charged to Zn4FPSt, which keeps
+// them subject to the one-FP-store-per-cycle limit; Zn4FP45 then carries only
+// the remaining FP4/5 uops, so the two together come to the measured FP4/5
+// total.
+//
+// Three entries widen their few FP1/FP2 uops to FP1-3 rather than name
+// Zn4FPU12, which works around a tool limitation and is not a claim about the
+// part. llvm-mca does not terminate on a write that holds both Zn4FPU01 and
+// Zn4FPU12: they share FP1 with neither containing the other, and both also
+// overlap Zn4FPU123. Its resource manager settles such an assignment with one
+// greedy pass rather than a matching, so an earlier group takes the only unit
+// a later one still needed and the instruction never becomes issuable. The
+// machine scheduler is unaffected -- llc emits identical code either way.
+// Widening by the one pipe avoids that shape and leaves every modeled
+// throughput here unchanged. Restore the measured Zn4FPU12 eligibility on
+// these three entries once llvm-mca can allocate these groups; see
+// https://github.com/llvm/llvm-project/issues/224248.
+//
+// Those occupancies alone fall short of the measured throughput, so Zn4UcodeGS
+// (declared with the other resources above) carries the fitted residual. It is
+// named to sort last, so adding it renumbers no existing resource.
+//
+// Treating that residual as a serialising limit is what Znver4 measures.
+// Independent gathers do not overlap: with two to five of them in flight, the
+// cost per operation is flat to within 3% -- 17.5 to 17.7 cycles for a 512-bit
+// dword gather, 9.6 to 9.9 for a 512-bit qword one -- where a unit that
+// overlapped them would show the per-operation cost falling as the count rose.
+// A lone gather measures above that plateau, 18.2 and 12.7 cycles, because it
+// is latency-exposed rather than throughput-limited. A gather issued alongside
+// a scatter costs 90-96% of the two added together rather than the larger of
+// the two, so the directions share the limit rather than each having their own.
+//
+// A retune edits only that column for as long as its pressure stays the
+// highest in the entry: throughput is the maximum over resources, so the
+// physical columns are a floor beneath which Zn4UcodeGS cannot reach. The
+// FP4/5 pressure of Zn4WriteVPSCATTERDDZ, for instance, is already 17 cycles,
+// so no value in its last column models that instruction faster than 17. A
+// measurement below an entry's floor needs the physical columns revisited too.
+
+// Each entry names its own resources, since the FP eligibilities differ from
+// one encoding to the next. The leading columns are the addressing, memory and
+// FP4/5 occupancies; the middle ones the FP0-3 subgroups; the last is always
+// Zn4UcodeGS.
+class Zn4GSRes<list<ProcResourceKind> res, list<int> cyc, int lat, int uops>
+    : SchedWriteRes<res> {
+  let ReleaseAtCycles = cyc; let Latency = lat; let NumMicroOps = uops;
+}
+def Zn4WriteVPGATHERQDZ128 : Zn4GSRes<[Zn4AGU012, Zn4Load, Zn4FPLd01,
+                                       Zn4FPU0123, Zn4FPU123, Zn4UcodeGS],
+                                      [2, 2, 4, 2, 2, 16], 13, 18>;
+def : InstRW<[Zn4WriteVPGATHERQDZ128, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQDZ128rm, VGATHERQPSZ128rm)>;
+def Zn4WriteVPGATHERDDZ128 : Zn4GSRes<[Zn4AGU012, Zn4Load, Zn4FPLd01,
+                                       Zn4FPU0123, Zn4FPU123, Zn4UcodeGS],
+                                      [4, 4, 6, 3, 2, 20], 15, 24>;
+def : InstRW<[Zn4WriteVPGATHERDDZ128, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERDDZ128rm, VGATHERDPSZ128rm,
+                     VPGATHERQDZ256rm, VGATHERQPSZ256rm)>;
+def Zn4WriteVPGATHERDDZ256 : Zn4GSRes<[Zn4AGU012, Zn4Load, Zn4FPLd01,
+                                       Zn4FPU0123, Zn4FPU123, Zn4UcodeGS],
+                                      [8, 8, 10, 5, 4, 37], 20, 41>;
+def : InstRW<[Zn4WriteVPGATHERDDZ256, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERDDZ256rm, VGATHERDPSZ256rm)>;
+def Zn4WriteVPGATHERQDZ    : Zn4GSRes<[Zn4AGU012, Zn4Load, Zn4FPLd01,
+                                       Zn4FPU0123, Zn4FPU12, Zn4FPU123,
+                                       Zn4UcodeGS],
+                                      [8, 8, 10, 5, 2, 5, 39], 24, 46>;
+def : InstRW<[Zn4WriteVPGATHERQDZ, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQDZrm, VGATHERQPSZrm)>;
+def Zn4WriteVPGATHERDDZ    : Zn4GSRes<[Zn4AGU012, Zn4Load, Zn4FPLd01,
+                                       Zn4FPU01, Zn4FPU0123, Zn4FPU123,
+                                       Zn4FPU23, Zn4UcodeGS],
+                                      [16, 16, 18, 1, 8, 11, 2, 70], 32, 81>;
+def : InstRW<[Zn4WriteVPGATHERDDZ, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERDDZrm, VGATHERDPSZrm)>;
+def Zn4WriteVPGATHERQQZ128 : Zn4GSRes<[Zn4AGU012, Zn4Load, Zn4FPLd01,
+                                       Zn4FPU0123, Zn4FPU123, Zn4UcodeGS],
+                                      [2, 2, 4, 2, 1, 16], 12, 17>;
+def : InstRW<[Zn4WriteVPGATHERQQZ128, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQQZ128rm, VGATHERQPDZ128rm,
+                     VPGATHERDQZ128rm, VGATHERDPDZ128rm)>;
+def Zn4WriteVPGATHERQQZ256 : Zn4GSRes<[Zn4AGU012, Zn4Load, Zn4FPLd01,
+                                       Zn4FPU0123, Zn4FPU123, Zn4UcodeGS],
+                                      [4, 4, 6, 3, 2, 20], 16, 24>;
+def : InstRW<[Zn4WriteVPGATHERQQZ256, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQQZ256rm, VGATHERQPDZ256rm,
+                     VPGATHERDQZ256rm, VGATHERDPDZ256rm)>;
+def Zn4WriteVPGATHERQQZ    : Zn4GSRes<[Zn4AGU012, Zn4Load, Zn4FPLd01,
+                                       Zn4FPU0123, Zn4FPU12, Zn4FPU123,
+                                       Zn4FPU23, Zn4UcodeGS],
+                                      [8, 8, 10, 4, 3, 6, 1, 39], 25, 48>;
+def : InstRW<[Zn4WriteVPGATHERQQZ, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQQZrm, VGATHERQPDZrm)>;
+def Zn4WriteVPGATHERDQZ    : Zn4GSRes<[Zn4AGU012, Zn4Load, Zn4FPLd01,
+                                       Zn4FPU0123, Zn4FPU12, Zn4FPU123,
+                                       Zn4FPU23, Zn4UcodeGS],
+                                      [8, 8, 10, 4, 2, 6, 1, 38], 22, 46>;
+def : InstRW<[Zn4WriteVPGATHERDQZ, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERDQZrm, VGATHERDPDZrm)>;
+
+// AVX2 (VEX) gathers. AVX2 has no scatter and no VEX form exceeds 8 x i32.
+// Each measures at the throughput of the EVEX entry for its shape, except
+// 8 x i32, which is genuinely cheaper at 8.35 against 9.19 cyc. They are not
+// the same sequences underneath, though: a VEX VPGATHERDD xmm issues eight
+// FP0-3 uops and five FP4/5 where its EVEX counterpart issues five and six. So
+// each VEX form keeps the microcode limit measured for it while carrying the
+// physical occupancy measured for its own encoding.
+//
+// An InstRW replaces the whole Sched<> list, so each one must restate the
+// mask-writeback def that X86InstrSSE.td attaches. That def is
+// encoding-neutral and models the mask clobber at the generic load latency, 5,
+// for VEX and EVEX alike, while the data result these entries measure is 12 to
+// 32; narrowing that gap is future work for both encodings.
+def Zn4WriteVPGATHERQDX    : Zn4GSRes<[Zn4AGU012, Zn4Load, Zn4FPLd01,
+                                       Zn4FPU01, Zn4FPU0123, Zn4FPU123,
+                                       Zn4FPU23, Zn4UcodeGS],
+                                      [2, 2, 3, 2, 1, 1, 1, 16], 13, 18>;
+def : InstRW<[Zn4WriteVPGATHERQDX, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQDrm, VGATHERQPSrm)>;
+def Zn4WriteVPGATHERDDX    : Zn4GSRes<[Zn4AGU012, Zn4Load, Zn4FPLd01,
+                                       Zn4FPU01, Zn4FPU0123, Zn4FPU123,
+                                       Zn4FPU23, Zn4UcodeGS],
+                                      [4, 4, 5, 4, 1, 2, 1, 20], 15, 24>;
+def : InstRW<[Zn4WriteVPGATHERDDX, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERDDrm, VGATHERDPSrm,
+                     VPGATHERQDYrm, VGATHERQPSYrm)>;
+def Zn4WriteVPGATHERDDY    : Zn4GSRes<[Zn4AGU012, Zn4Load, Zn4FPLd01,
+                                       Zn4FPU01, Zn4FPU123, Zn4FPU23,
+                                       Zn4UcodeGS],
+                                      [8, 8, 9, 8, 7, 2, 33], 21, 42>;
+def : InstRW<[Zn4WriteVPGATHERDDY, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERDDYrm, VGATHERDPSYrm)>;
+// The VEX form retires one micro-op more than the EVEX form of the same
+// gather: uops.info measures 18 for VPGATHERQQ (XMM, VSIB_XMM, XMM) on Zen 4
+// against 17 for VPGATHERQQ (XMM, K, VSIB_XMM), which Zn4WriteVPGATHERQQZ128
+// carries. The two encodings agree on latency and throughput, so only the
+// micro-op count distinguishes them.
+def Zn4WriteVPGATHERQQX    : Zn4GSRes<[Zn4AGU012, Zn4Load, Zn4FPLd01,
+                                       Zn4FPU01, Zn4FPU0123, Zn4FPU123,
+                                       Zn4UcodeGS],
+                                      [2, 2, 3, 2, 1, 1, 16], 12, 18>;
+def : InstRW<[Zn4WriteVPGATHERQQX, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQQrm, VGATHERQPDrm,
+                     VPGATHERDQrm, VGATHERDPDrm)>;
+def Zn4WriteVPGATHERQQY    : Zn4GSRes<[Zn4AGU012, Zn4Load, Zn4FPLd01,
+                                       Zn4FPU01, Zn4FPU123, Zn4FPU23,
+                                       Zn4UcodeGS],
+                                      [4, 4, 5, 4, 3, 1, 20], 16, 24>;
+def : InstRW<[Zn4WriteVPGATHERQQY, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERQQYrm, VGATHERQPDYrm)>;
+def Zn4WriteVPGATHERDQY    : Zn4GSRes<[Zn4AGU012, Zn4Load, Zn4FPLd01,
+                                       Zn4FPU01, Zn4FPU0123, Zn4FPU123,
+                                       Zn4FPU23, Zn4UcodeGS],
+                                      [4, 4, 5, 4, 1, 2, 1, 20], 16, 24>;
+def : InstRW<[Zn4WriteVPGATHERDQY, WriteVecMaskedGatherWriteback],
+             (instrs VPGATHERDQYrm, VGATHERDPDYrm)>;
+
+def Zn4WriteVPSCATTERQDZ128 : Zn4GSRes<[Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FP45,
+                                        Zn4FPU123, Zn4UcodeGS],
+                                       [2, 2, 2, 4, 1, 16], 6, 17>;
+def : InstRW<[Zn4WriteVPSCATTERQDZ128],
+             (instrs VPSCATTERQDZ128mr, VSCATTERQPSZ128mr)>;
+def Zn4WriteVPSCATTERDDZ128 : Zn4GSRes<[Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FP45,
+                                        Zn4FPU123, Zn4FPU23, Zn4UcodeGS],
+                                       [4, 4, 4, 6, 2, 1, 33], 9, 27>;
+def : InstRW<[Zn4WriteVPSCATTERDDZ128],
+             (instrs VPSCATTERDDZ128mr, VSCATTERDPSZ128mr)>;
+def Zn4WriteVPSCATTERQDZ256 : Zn4GSRes<[Zn4AGU012, Zn4Store, Zn4FPSt, Zn4FP45,
+                                        Zn4FPU123, Zn4FPU23, Zn4UcodeGS],
+                                       [4, 4, 4, 6, 2, 1, 26], 9, 27>;
----------------
MattPD wrote:

The last commit makes the measured order of two lowerings a requirement and checks it at sixteen lanes. At four lanes the model has them the other way round: `vpscatterqd %xmm, (%rdx,%ymm)` at 6.50 against `vpscatterdd %xmm, (%rdx,%xmm)` at 8.25, at equal uops and latency. Every other pair in the table puts the qword index at or above the dword one. Your table in the #199488 thread has Zen4 at 5.38 for both and Zen5 at 4.61 against 5.53, so this entry carries the Zen5 ordering under the Znver4 name, and both four-lane entries sit 21% and 53% above the Zen4 readings.

I have answered the 26-or-33 question in the #199488 thread. Whichever value stays, could both four-lane readings be measured under the condition that calibrated the rest of the table?

https://github.com/llvm/llvm-project/pull/212997


More information about the llvm-commits mailing list