[llvm] [AArch64] Fold CSEL of a CSEL with the same flags (PR #224992)
Henry Jiang via llvm-commits
llvm-commits at lists.llvm.org
Fri Sep 25 11:52:12 PDT 2026
mustartt wrote:
> > > Is there any particular reason we want to do it in ISel as opposed to a fold in InstCombine? Folding select of select with the same (or negated) condition sounds like all targets can benefit.
> >
> >
> > @SavchenkoValeriy Currently `InstCombine` handles the single use case only. But for our use, we need to support multi use for intermediate ops. It can definitely be implemented in `AggressiveInstCombine` where the IR instruction count is not as strict. A case can be made for implementing it in `InstCombine`, since it is still profitable for downstream optimizations that results in less instructions for both x86 and AArch64.
> > But if we implement it during ISel, the fold is guaranteed to not increase instruction count or regress performance.
>
> Please correct me if I'm wrong, but isn't `select %cond (select %cond %a %b) %c` always foldable to `select %cond %a %c`? If the nested `select` has more uses - fine, we can keep it, but a. maintain the same number of IR instructions and b. shorten the critical path. If we block it on the nested `select` having a single use, it sounds like a perfect opportunity to relax it.
Yes that's correct and it's already implemented in `InstCombine`. What `InstCombine` currently does not support are the 2 following forms:
```llvm
%s = select i1 %c, i32 %a, i32 %b
%i = add i32 %s, 1
%r0 = select i1 %c, i32 %i, i32 %x
%r1 = select i1 %c, i32 %i, i32 %y
; becomes
%s = select i1 %c, i32 %a, i32 %b
%i = add i32 %a, 1
%r0 = select i1 %c, i32 %i, i32 %x
%r1 = select i1 %c, i32 %i, i32 %y
```
where the `add` is only used on the same side. This can be rewritten in place (This is the fold that benefits all platforms without increasing instruction count).
But for the pattern we care about,
```llvm
%s = select i1 %c, i32 %b, i32 %a
%inc = add i32 %s, 1
%an = select i1 %c, i32 %a, i32 %inc
%bn = select i1 %c, i32 %inc, i32 %b
; becomes
%s = select i1 %c, i32 %b, i32 %a
%inc.f = add i32 %a, 1
%inc.t = add i32 %b, 1
%an = select i1 %c, i32 %a, i32 %inc.f
%bn = select i1 %c, i32 %inc.t, i32 %b
```
where `%inc` is used on both sides, we would need to duplicate `%inc` which increases the instruction count.
I'm thinking I would put up another orthogonal patch in `InstCombine` for the first case, and let each target handle what's profitable for the second case, which still requires this patch.
https://github.com/llvm/llvm-project/pull/224992
More information about the llvm-commits
mailing list