[llvm] AMDGPU: Perform zero/any extend combine into permute (PR #177370)
Matt Arsenault via llvm-commits
llvm-commits at lists.llvm.org
Thu Jan 22 08:32:31 PST 2026
================
@@ -14470,7 +14473,44 @@ SDValue SITargetLowering::performZeroExtendCombine(SDNode *N,
if (Src.getValueType() != MVT::i16)
return SDValue();
- return SDValue();
+ if (1 < std::distance(Src->user_begin(), Src->user_end()))
+ return SDValue();
+
+ // TODO: We bail out below if SrcOffset is not in the first dword (>= 4). It's
+ // possible we're missing out on some combine opportunities, but we'd need to
+ // weigh the cost of extracting the byte from the upper dwords.
+
+ std::optional<ByteProvider<SDValue>> BP0 =
+ calculateByteProvider(SDValue(N, 0), 0, 0, 0);
+ if (!BP0.has_value() || 4 <= BP0->SrcOffset)
+ return SDValue();
+ SDValue V0 = BP0->Src.value_or(SDValue());
+
+ std::optional<ByteProvider<SDValue>> BP1 =
+ calculateByteProvider(SDValue(N, 0), 1, 0, 1);
+ if (!BP1.has_value() || 4 <= BP1->SrcOffset)
+ return SDValue();
----------------
arsenm wrote:
```suggestion
if (!BP1 || BP1->SrcOffset > 4 || !BP1->Src)
return SDValue();
```
https://github.com/llvm/llvm-project/pull/177370
More information about the llvm-commits
mailing list