[llvm] [RISCV] Use `experimental-p` extension for zero-extended narrow unsigned types (PR #213260)
Craig Topper via llvm-commits
llvm-commits at lists.llvm.org
Sat Sep 12 20:44:38 PDT 2026
topperc wrote:
> Let me clarify the concerns here:
>
> For any pattern or combination of patterns (ex. mul/sub, add/div, etc.) we would have to visit each node at least once to check if we are able to match it.
>
> Next, we only have two options here:
>
> * Standard Recursion
> * Iterative Recursion
>
> In this case iteration is obviously better and more performant, speaking of the Worklist, that's just a simple vector nothing fancy here and is the minimum requirement to keep track of the nodes.
>
> Also, this preserves the order naturally, if you look closely to how I push the nodes, the only change I made is add the second node and then the first, since we use a queue here.
The queue doesn't preserve the structure. You're always producing a linear sequence even if the original sequence was a tree. With add and sub mixed with other instructions we will need to keep the structure.
>
> When we start popping them, it comes out in the forward order as intended, right now, I haven't extended this to sub or other order-strict op's so it may not feel that it's extendable.
Do you have a patch for sub that you can share so we can see that it is extendable?
> Now that the logic and complexity is clear, generalisation has only one thing to handle, i.e, at each depth level d(n) we have to check m possible instructions. So, for example in the minimal case of only combination of add/sub, we need 2 checks at each level.
>
> Assuming we don't have millions of types of instructions (the 'm' part) this shouldn't hurt the time complexity much. If we still care about that small constant factor, we can use bitmasking for preserving that as well and remvoving that constant factor too.
>
> Next, GISel/MachineIR wouldn't make much difference in the implementation or performance, cause it's the exact same algorithm that has to follow, that decision can always be made later since SelDag is much stable and can handle such additions, if needed we can always support it in GISel later.
>
> Regarding the scope of instructions, we actually already discovered a few right, like add, sub, mul, div, shift, etc.
There is no div in P and mul is weird. There is no packed mul. There is only packed multiply high and packed multiplies that produce full products.
> Indeed, it's not a comprehensive list but that doesn't interfere with the p-extension pass in any manner.
>
> The work of stabilising p-ext can continue as it is, at the same time, as we keep discovering more cases where we can use packed instructions than scalar ones, we just add a new case for it, it would be quite smooth.
>
> I don't think putting this in a separate pass is worth it, cause packed would always be preferred for performance in any scenario unless the consumer really wants to squeeze things to reduce size.
A separate pass in MIR would allow GISel and SelectionDAG to share a single implementation, but we lose the AssertZExt and need to preserve that somehow.
We may need to support these in the MachineCombiner's schedule based reassociation algorithm if we're going to start replacing scalar add/sub with packed.
We should also track zero bits through intermediate nodes that propagate zeros like and/or/xor/select/etc. Ideally we'd also be able to handle loop cases through phi nodes which SelectionDAG cannot see.
https://github.com/llvm/llvm-project/pull/213260
More information about the llvm-commits
mailing list