[llvm] [X86][CostModel] Add per-shape gather/scatter cost tables for AMD znver4+ (PR #199488)

Sumukh J Bharadwaj via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 1 13:55:23 PDT 2026


amd-subharad wrote:

Thank you @MattPD 
I have tried to address all your queries, several by changing the design rather than patching the symptom.

**The 3-element gather assertion:** Removed, and the lookup returns std::nullopt for any unlisted shape so the caller takes the flat-overhead path, which is what you suggested and is already the defined behaviour everywhere else. Your <3 x i32> case no longer aborts; it reports 5, the same as skx.

**The missing hasAVX512() on the gather path:** Both the overhead lookup and the schedule-model read now require it, symmetric with getScatterOverhead. The leak is closed: -mcpu=skylake gives 10 and -mcpu=skylake -mattr=+prefer-gs-cost-table also gives 10, where you measured 12 against 22.

**Class validity versus presence of an override:** The guard now derives a floor from the active schedule model: the cheapest a real gather can be is a plain vector load, so anything at or below VMOVUPSZrm's reciprocal throughput is treated as the generic default in disguise and takes the fallback. Your synthetic probe which is an unmodelled op reporting 0.3 and rounding to 0 is rejected instead of accepted, and a future subtarget that enables the feature without its own overrides gets the fallback rather than a zero body.

**Overhead paid once per part:** Restructured along the lines you proposed. The premium belongs to the whole logical operation and is looked up once on the native shape; only genuinely wider-than-native shapes multiply by a split factor. The v16i32 case now reports the same cost either way — 35 for both the GEP form and the bare <16 x ptr>, against the 27 and 46 measured.

**TCK_SizeAndLatency reading instruction latency:** The schedule-model path is now restricted to TCK_RecipThroughput, so that cost kind no longer routes through it at all and the inliner is unaffected.

**Representative opcode selection:** Now keyed on (NumElts, EltBits) as you suggested. The hasVLX() check is in place for the Z128/Z256 forms, and VF 5/6/7 take the fallback.

https://github.com/llvm/llvm-project/pull/199488


More information about the llvm-commits mailing list