[llvm] AMDGPU: Perform zero/any extend combine into permute (PR #177370)

Matt Arsenault via llvm-commits llvm-commits at lists.llvm.org
Thu Jan 22 08:32:31 PST 2026


================
@@ -14470,7 +14473,44 @@ SDValue SITargetLowering::performZeroExtendCombine(SDNode *N,
   if (Src.getValueType() != MVT::i16)
     return SDValue();
 
-  return SDValue();
+  if (1 < std::distance(Src->user_begin(), Src->user_end()))
+    return SDValue();
+
+  // TODO: We bail out below if SrcOffset is not in the first dword (>= 4). It's
+  // possible we're missing out on some combine opportunities, but we'd need to
+  // weigh the cost of extracting the byte from the upper dwords.
+
+  std::optional<ByteProvider<SDValue>> BP0 =
+      calculateByteProvider(SDValue(N, 0), 0, 0, 0);
+  if (!BP0.has_value() || 4 <= BP0->SrcOffset)
+    return SDValue();
+  SDValue V0 = BP0->Src.value_or(SDValue());
----------------
arsenm wrote:

```suggestion
  SDValue V0 = *BP0->Src;
```

https://github.com/llvm/llvm-project/pull/177370


More information about the llvm-commits mailing list