[llvm] [X86][SchedModel] Add masked gather/scatter overrides to Znver4Model (PR #212997)

Sumukh J Bharadwaj via llvm-commits llvm-commits at lists.llvm.org
Sun Sep 6 08:21:30 PDT 2026


================
@@ -509,6 +509,130 @@ defm : Zn4WriteResInt<WriteLoad, [Zn4AGU012, Zn4Load], !add(Znver4Model.LoadLate
 // Does not cost anything by itself, only has latency, matching that of the WriteLoad,
 defm : Zn4WriteResInt<WriteVecMaskedGatherWriteback, [], !add(Znver4Model.LoadLatency, 1), [], 0>;
 
+// AVX-512 masked GATHER / SCATTER. Microcoded, so entries are keyed by
+// (#elements, element width, index width); the two index widths share an entry
+// only where they measure alike. Throughput and latency were measured on
+// Znver4 with every mask bit set, and are accurate to about a cycle; scatter
+// latency is estimated rather than timed. Znver4Model also backs znver5 and
+// znver6, which run 25-55% faster and are pessimistic here.
+//
+// The FP occupancies are the uops.info Zen4 port distributions for the same
----------------
amd-subharad wrote:

Adopted. The sums are gone. Each uop now stays on the sub-group uops.info reports it eligible for, with Zn4FPU01, Zn4FPU12 and Zn4FPU23 added so those constraints can be expressed.

https://github.com/llvm/llvm-project/pull/212997


More information about the llvm-commits mailing list