[llvm] [AMDGPU] Break v2i32 and/or/xor ops into i32 pairs when followed by t… (PR #191422)
Wooseok Lee via llvm-commits
llvm-commits at lists.llvm.org
Fri Apr 10 11:54:28 PDT 2026
================
@@ -13956,6 +13956,58 @@ static uint32_t getPermuteMask(SDValue V) {
return ~0;
}
+static SDValue performAndOrXorv2i32Combine(SDNode *N, SelectionDAG &DAG) {
+ EVT VT = N->getValueType(0);
+ const unsigned Opc = N->getOpcode();
+ assert(VT == MVT::v2i32);
+ if (!N->isDivergent())
+ return SDValue();
+
+ SDValue LHS = N->getOperand(0);
+ SDValue RHS = N->getOperand(1);
+
+ // Reject v2i32 ops that would be handled by v2i32-v2i32 patterns
+ auto matchV2I32Patterns = [](const unsigned Opc,
+ const unsigned userOpc) -> bool {
+ return (Opc == ISD::AND || Opc == ISD::OR) && userOpc == ISD::OR;
+ };
+ if (matchV2I32Patterns(LHS.getOpcode(), Opc) ||
+ matchV2I32Patterns(RHS.getOpcode(), Opc))
+ return SDValue();
+
+ for (SDUse &use : N->uses()) {
+ SDNode *user = use.getUser();
+ if (user->getValueType(0) == MVT::v2i32 &&
+ matchV2I32Patterns(Opc, user->getOpcode()))
+ return SDValue();
+ }
+
+ // If the v2i32 op has a following i32 operation, break v2i32 into two i32 ops
----------------
wooseoklee wrote:
A bit of background why I decide to do it in dag combiner. The approach that you described has been implemented, but either (1) it is really hard to merge with existing SelectBITOP3() or (2) there exists some limitation if we have separate pattern to match.
Here, what we are trying to achieve is that
// Handle specific v2i32-i32 pattern
// i32_op(extract_elt(v2i32_op(LHS, RHS), 0),
// extract_elt(v2i32_op(LHS, RHS), 1))
// map to
// bitop3 LHS_lo, (i32_op LHS_hi, RHS_hi), RHS_lo
(1) SelectBITOP3() recursively visits ops to merge as many logical ops as possible by handling constant arguments. However, our pattern is one time conversion, which couldn't really fit in the way the SelectBITOP3() traverses.
(2) Implementing separate pattern to handle the case, such as SelectBITOP3v2i32() for my implementation, can be a solution, but we may lose some chances to further optimize the generated i32 logical ops if we maintain v2i32 and/or/xor until instruction selection and generate them.
So, as an alternative solution, I would like to break the v2i32 logical op into two i32 ops if it make sense at the last stage of dag combiner such that at the instruction selection stage, the logical ops can be merged as much as possible through SelectBITOP3().
https://github.com/llvm/llvm-project/pull/191422
More information about the llvm-commits
mailing list