[llvm] [X86] Add ISD::MULHS/MULHU v4i64/v8i64 lowering (PR #169819)

Rito Takeuchi via llvm-commits llvm-commits at lists.llvm.org
Sun Jun 28 05:55:42 PDT 2026


Licht-T wrote:

I am also interested in this PR and collected the microbench results.

**Method:** Built `llc` from this PR. `i7-1260P` (Alder Lake, **AVX2**), single pinned core, 11 trials, values are median, code is [here](https://gist.github.com/Licht-T/1e9fc2719d10a192e16c2974034b3ecd):

| Kernel | stock ns / MHz | patched ns / MHz | speedup |
|---|---|---|---|
| `udiv_vec`: explicit `<4 x i64>` **unsigned** / constant | 1.054 / 2200 | 0.678 / 2198 | **1.55×** |
| `div_vec`: same kernel but **signed** | 1.031 / 2200 | 1.314 / 2198 | **0.78×** |
| `wyhash`: hash mixer using the **full 128-bit product** | 1.250 / 2200 | 2.035 / 2200 | **0.61×** |
| `fp_adv`: **14 live `<4 x double>` accumulators + one integer `/C`** | 1.994 / 2200 | 2.129 / 2200 | 0.94× |

Here are answers for @phoebewang's two concerns and two real regressions I found while testing:

- Concern 1: CPU frequency
    At each test, frequency is flat scalar-vs-patched in every row (~2200 MHz) at least AVX2. I haven't tested AVX-512.
- Concern 2: Register pressure with heavy FP/vector use + calls
    For `fp_adv`, patched vs stock:
    ```
                vpmuludq   frame    ymm spill/reload moves
    stock           0      488 B            16
    patched        16      544 B            24    (+56 B, +8)
    ```
    So the patch does increase pressure, but the runtime hit is negligible here (0.94×).
- Real regression 1: **`div_vec`: signed `MULHS`, 0.78×.**
    Sign-correction needs `v4i64` arithmetic-shift-right (`vpsraq`), which AVX2 lacks (AVX-512 only), so it's emulated and loses to scalarizing. **Suggested fix: Keep `MULHU` only.**
- Real regression 2: **`wyhash`: full 128-bit product, 0.61×.**
   The patch vectorizes the high half (`MULHU` to `VPMULUDQ`) but then must reassemble 128 bits from 32-bit partials. **Suggested fix: When a `v4i64 MULHU` has a sibling `MUL` of the same operands (`UMUL_LOHI` / full product consumed), keep it scalar.**

@RKSimon Hope this helps. If you are okay, I am willing to take over this PR. Just please let me know. Thanks.

https://github.com/llvm/llvm-project/pull/169819


More information about the llvm-commits mailing list