[llvm] [AVX-512] make vpternlogq more aggressive for longer chains of bitmanipulations (PR #189971)
Julian Pokrovsky via llvm-commits
llvm-commits at lists.llvm.org
Tue Jul 7 00:31:24 PDT 2026
raventid wrote:
@Validark thank you for a great example again! I finally have more time to wrap up the patch. For the example above (I have tried locally) if you run it through the whole llvm pipeline you get an optimal scheduling.
```
vpor %xmm2, %xmm3, %xmm2 # c | d
vpternlogq $224, %xmm0, %xmm1, %xmm2 # acc = (c|d) & (a | b)
vpternlogq $224, %xmm4, %xmm5, %xmm2 # acc &= (e | f)
vpternlogq $224, %xmm6, %xmm7, %xmm2 # acc &= (g | h)
vpternlogq $224, 8(%rsp), %xmm11, %xmm2 # acc &= (j | i) i from memory
vpternlogq $224, 40(%rsp), %xmm10, %xmm2 # acc &= (l | k) k from memory
vpternlogq $224, 72(%rsp), %xmm9, %xmm2 # acc &= (n | m) m from memory
vpternlogq $224, 104(%rsp), %xmm8, %xmm2 # acc &= (p | o) o from memory
vpand 136(%rsp), %xmm2, %xmm0 # acc &= q q from memory
vpternlogq $2, 168(%rsp), %xmm1, %xmm0 # s & ~(acc | r) = ~U & ~r & s
# s from memory
```
But if you use my `ternlog eval` implementation on a raw tree with your structure it won't be optimal.
tldr; The patch above looks good and I don't see any regressions if you use it as a part of the full -O2 pipeline, if you throw complex trees (not optimized by the previous steps) at it though, it might regress on rare cases.
I will try to compare those cases to GCC and add as a test cases.
https://github.com/llvm/llvm-project/pull/189971
More information about the llvm-commits
mailing list