[llvm] [AVX-512] make vpternlogq more aggressive for longer chains of bitmanipulations (PR #189971)

via llvm-commits llvm-commits at lists.llvm.org
Sat Sep 26 14:52:58 PDT 2026


mightsleep wrote:

Hi,
I ran into this from the Rust side: Miri can't run AVX-512 code that  uses_mm512_ternarylogic_epi64, since stdarch lowers it to an LLVM intrinsic. stdarch could use generic bit ops instead, but only if LLVM folds them back into one vpternlog, which is how I ended up here.

I ran this PR against its merge base on the same IR and executed everything on Zen 5 (no miscompiles). Of the 256 immediates written as generic ops, base folds 79 into one instruction, this PR 242, gcc all of them; in the Rust stdarch-style form, 37 and 158. The X86 CodeGen tests get slightly better, none worse. On random logic trees, 50 of 600 get worse than base, and a deep 16k-op tree takes 1.5 s in llc instead of 0.25 s.

Where I think it comes from:
- tryMatchBitSelect runs before tryVPTERNLOG and leaves a logic op outside the select.
- A NOT with other uses stops the walk, though it is free in the immediate.
- A node shared inside the same tree (DAG CSE, e.g. a | c in both halves) is taken as a leaf.
- Collect gathers every candidate in the subtree, hence the compile time.

Fixes for these, on top of this branch: https://gist.github.com/mightsleep/aaa12fc894eaee59c4f6a77c2330df66 With them, 256 of 256 in both forms, and 0.36 s on the deep tree.

The worse trees are a separate matter: they come from the cut choice. The cascade takes the cut that folds the most leaves at the root, while some trees need two or three opaque subtrees; the smallest, (((b | d) ^ f) ^ c) & (e | (f & a)), goes from 3 vpternlog to 4. In my measurements, choosing the cut by the instruction count over the whole tree removes all of them, but that's a design choice, so I left it out.

Last, an idea for redoing it: carry per node the 3-input cuts with their 8-bit truth tables and compose them, i.e. LUT mapping with k = 3, computed lazily in a bounded neighbourhood of the node, so no state between calls. It makes NOT, bit select and sharing ordinary cases and composes chains of vpternlog too; my prototype does a bit better still. I don't want to clutter this PR with it, though. What do you think?

https://github.com/llvm/llvm-project/pull/189971


More information about the llvm-commits mailing list