[llvm] [RISCV] Use `experimental-p` extension for zero-extended narrow unsigned types (PR #213260)
Rajveer Singh Bharadwaj via llvm-commits
llvm-commits at lists.llvm.org
Sat Sep 12 04:20:40 PDT 2026
Rajveer100 wrote:
Let me clarify a the concerns here:
For any pattern or combination of patterns (ex. mul/sub, add/div, etc.) we would have to visit each node at least once to check if we are able to match it.
Next, we only have two options here:
- Standard Recursion
- Iterative Recursion
In this case iteration is obviously better and more performant, speaking of the Worklist, that's just a simple vector nothing fancy here and is the minimum requirement to keep track of the nodes.
Also, this preserves the order naturally, if you look closely to how I push the nodes, the only change I made is add the second node and then the first, since we use a queue here.
When we start popping them, it comes out in the forward order as intended, right now, I haven't extended this to sub or other order-strict op's so it may not feel that it's extendable.
Now that the logic and complexity is clear, generalisation has only one thing to handle, i.e, at each depth level d(n) we have to check m possible instructions. So, for example in the minimal case of only combination of add/sub, we need 2 checks at each level.
Assuming we don't have a million combinations this shouldn't hurt the time complexity much. If we still care about that small constant factor, we can use bitmasking for preserving that as well and remvoving that constant factor too.
Next, GISel/MachineIR wouldn't make much difference in the implementation or performance, cause it's the exact same algorithm that has to follow, that decision can always be made later since SelDag is much stable and can handle such additions, if needed we can always support it in GISel later.
Regarding the scope of instructions, we actually already discovered a few right, like add, sub, mul, div, shift, etc.
Indeed, it's not a comprehensive list but that doesn't interfere with the p-extension pass in any manner.
The work of stabilising p-ext can continue as it is, at the same time, as we keep discovering more cases where we can use packed instructions than scalar ones, we just add a new case for it, it would quite smooth.
I don't think putting this in a separate pass is worth, cause packed would always be preferred for performance in any scenario unless the consumer really wants to squeeze things to reduce size.
https://github.com/llvm/llvm-project/pull/213260
More information about the llvm-commits
mailing list