[llvm] [X86][SchedModel] Add masked gather/scatter overrides to Znver4Model (PR #212997)
Sumukh J Bharadwaj via llvm-commits
llvm-commits at lists.llvm.org
Thu Sep 17 02:20:15 PDT 2026
================
@@ -509,6 +509,130 @@ defm : Zn4WriteResInt<WriteLoad, [Zn4AGU012, Zn4Load], !add(Znver4Model.LoadLate
// Does not cost anything by itself, only has latency, matching that of the WriteLoad,
defm : Zn4WriteResInt<WriteVecMaskedGatherWriteback, [], !add(Znver4Model.LoadLatency, 1), [], 0>;
+// AVX-512 masked GATHER / SCATTER. Microcoded, so entries are keyed by
+// (#elements, element width, index width); the two index widths share an entry
+// only where they measure alike. Throughput and latency were measured on
+// Znver4 with every mask bit set, and are accurate to about a cycle; scatter
+// latency is estimated rather than timed. Znver4Model also backs znver5 and
+// znver6, which run 25-55% faster and are pessimistic here.
+//
+// The FP occupancies are the uops.info Zen4 port distributions for the same
+// encodings, with the FP0-3 subgroups (FP01/FP12/FP23/FP123/FP0123) summed
+// into Zn4FPU0123. A scatter's element stores are charged to Zn4FPSt, which
+// keeps them subject to the one-FP-store-per-cycle limit; Zn4FP45 then carries
+// only the remaining FP4/5 uops, so the two together come to the measured
+// FP4/5 total.
+//
+// Those occupancies alone fall short of the measured throughput -- the
+// microcode sequencer, not any one pipe, is the limit -- so Zn4UcodeGS carries
----------------
amd-subharad wrote:
Added this
https://github.com/llvm/llvm-project/pull/212997
More information about the llvm-commits
mailing list