[llvm] [X86] Add ISD::MULHS/MULHU v4i64/v8i64 lowering (PR #169819)
Rito Takeuchi via llvm-commits
llvm-commits at lists.llvm.org
Tue Jun 30 08:15:08 PDT 2026
Licht-T wrote:
I also tested with the following EC2 metal instance types ([code](https://gist.github.com/Licht-T/1e9fc2719d10a192e16c2974034b3ecd)):
| microarch | CPU | EC2 instance type |
|---|---|---|
| Skylake/Cascade Lake | Xeon Platinum 8275CL | `c5.metal` |
| Ice Lake-SP | Xeon Platinum 8375C | `c6i.metal` |
| Sapphire Rapids (SPR) | Xeon Platinum 8488C | `c7i.metal-24xl` |
| Zen 4 / Genoa | EPYC 9R14 | `c7a.metal-48xl` |
## Summary
- **`MULHU` (unsigned div) is a clean win** across all four microarchs: at **1.6 - 2.5x**. It stays a net win even after the frequency drop (Original concern 1).
- **Register pressure (Original concern 2) is real but negligible in practice.** It only shows up in a synthetic worst case (14 live FP accumulators kept alive across the loop). The patch's actual targets are integer-divide-bound loops, which aren't FP-heavy, and AVX-512's 32 registers absorb most of the spill component anyway.
- **Two regressions are not negligible, but both are addressable:**
- **full-128 / `wyhash` (`lo ^ hi`)** regresses on every microarch, because scalar `mulx` returns both halves in one op. Solution: guard the `UMUL_LOHI` case (don't vectorize when both lo and hi are used).
- **signed div (`MULHS`)** regresses only on slow-`vpmullq` Intel parts (Skylake/Cascade/SPR); it's actually a win on Ice Lake (1.28x) and Zen 4 (1.83x). Solution: gate the low multiply via `vpmuludq` with a `TuningSlowVPMULLQ` flag.
- **So split the PR into two steps:**
- **Step 1: `MULHU` only + the `UMUL_LOHI` guard.**
- **Step 2: `MULHS` (AVX-512 only) + `TuningSlowVPMULLQ`** for microarch-dependent perf optimization.
## Results
**Result 1. MULHU**: `uint64 / constant`: Numbers are ns/elem, `stock → patched (ratio)`.
| build | Cascade | Ice Lake | SPR | Zen 4 |
|---|---|---|---|---|
| AVX2 | 1.108→0.687 (1.61x) | 0.955→0.642 (1.49x) | 0.945→0.499 (1.89x) | 0.648→0.263 (2.46x) |
| AVX-512 256b | 1.099→0.658 (1.67x) | 0.955→0.566 (1.69x) | 0.963→0.546 (1.76x) | 0.516→0.286 (1.80x) |
| AVX-512 512b | 1.099→0.637 (1.73x) | 0.955→0.565 (1.69x) | 0.962→0.546 (1.76x) | 0.563→0.287 (1.96x) |
**Result 2. MULHS**: `int64 / constant`: Numbers are ns/elem, `stock → patched (ratio)`.
| build | Cascade | Ice Lake | SPR | Zen 4 |
|---|---|---|---|---|
| AVX2 | 1.077→1.326 (**0.81x**) | 0.919→1.170 (**0.79x**) | 0.942→1.191 (**0.79x**) | 0.716→0.588 (1.22x) |
| AVX-512 256b | 1.053→0.983 (1.07x) | 0.958→0.751 (1.28x) | 0.943→1.984 (**0.48x**) | 0.612→0.334 (**1.83x**) |
| AVX-512 512b | 1.053→0.986 (1.07x) | 0.961→0.751 (1.28x) | 0.941→1.983 (**0.47x**) | 0.667→0.334 (**2.00x**) |
**Result 3. full-128**: `wyhash` `lo ^ hi`: Numbers are ns/elem, `stock → patched (ratio)`.
| build | Cascade | Ice Lake | SPR | Zen 4 |
|---|---|---|---|---|
| AVX2 | 1.599→2.142 (0.75x) | 1.253→1.837 (0.68x) | 1.145→1.818 (0.63x) | 0.806→1.062 (0.76x) |
| AVX-512 256b | 1.482→2.016 (0.74x) | 1.178→1.839 (0.64x) | 1.137→1.831 (0.62x) | 0.811→1.159 (0.70x) |
| AVX-512 512b | 1.789→2.268 (0.79x) | 1.677→1.928 (0.87x) | 1.963→2.167 (0.91x) | 0.772→1.117 (0.69x) |
**Result 4. Frequency (Original concern 1)**: delivered all-core freq via APERF/MPERF, sustained, turbo on (MHz).
| microarch | scalar | 256b | 512b | 256b/scl | 512b/256b |
|---|---|---|---|---|---|
| Cascade (8275CL) | 3568 | 3100 | 2700 | 0.87 | 0.87 |
| Ice Lake (8375C) | 3485 | 3439 | 3121 | 0.99 | 0.91 |
| SPR (8488C) | 3200 | 3107 | 2403 | 0.97 | 0.77 |
| Zen 4 (9R14) | 2611 | 2318 | 2588 | 0.89 | 1.12 |
**Result 5. FP register-pressure (Original concern 2)**: 20-run medians, turbo off 14 live `<4 x double>` accumulators + one integer dividing with constant. ns `stock → patched (ratio)`. c0 = no call; c1 = a register-clobbering call each iteration (caller-saved spill).
| µarch | build | c0 (no call) | c1 (with call) |
|---|---|---|---|
| SPR| AVX2 (16 vec regs) | 1.612→2.040 (0.79x) | 2.655→3.247 (0.82x) |
| SPR | AVX-512 256b (32 regs) | 1.097→1.184 (0.93x) | 1.991→2.129 (0.94x) |
| SPR | AVX-512 512b (32 regs) | 1.097→1.184 (0.93x) | 2.252→2.532 (0.89x) |
| Cascade Lake | AVX2 (16 vec regs) | 1.571→1.748 (0.90x) | 2.846→3.171 (0.90x) |
| Cascade Lake | AVX-512 256b (32 regs) | 1.030→1.238 (0.83x) | 3.038→3.037 (1.00x) |
| Cascade Lake | AVX-512 512b (32 regs) | 1.030→1.242 (0.83x) | 3.040→2.949 (1.03x) |
https://github.com/llvm/llvm-project/pull/169819
More information about the llvm-commits
mailing list