[llvm] [NVPTX] Use zero for undef bytes in v4i8 construction (PR #224575)

via llvm-commits llvm-commits at lists.llvm.org
Mon Sep 21 06:15:59 PDT 2026


sinan-flash wrote:

> Interesting, can't this be done in 1 PRMT? Maybe we should specialize LowerBUILD_VECTOR for undef/poison operands?

@AlexMaclean Thanks for the feedback! A single PRMT would suffice for this specific case. And I think this exposes a broader issue with mixed variable/constant inputs in `LowerBUILD_VECTOR`.

There are two related problems/optimization chances:
- undef/poison mixed with non-constant operands
We can address this one in this PR. My current reproducer uses undef; I haven’t encountered a poison case yet, but I can add poison case if useful.

<br>

- LowerBUILD_VECTOR handles fully constant inputs well, but lowers other v4i8 inputs through a fixed three-PRMT sequence, which can be suboptimal.


For testcase in this PR, choosing zero for the undef elements effectively gives:
```
[x, undef, undef, undef] → [x, 0, 0, 0]

lo = PRMT(x, 0, 0x3340)
hi = PRMT(0, 0, 0x3340)
r  = PRMT(lo, hi, 0x5410)
```

then combine stage(`combinePRMT`) folds full constant prmt `hi` to zero, leaving:
```
t = PRMT(x, 0, 0x3340)
r = PRMT(t, 0, 0x5410)
```

But we could construct the result directly with `r = PRMT(x, 0, 0x4440)`.

<br>

This isn’t specific to undef. LowerBUILD_VECTOR also misses these mixed variable/constant cases. More generally, a single variable mixed with constants can use one PRMT, for example:

e.g. one scalar pattern can be done with only one prmt, but the actual results are both 2 PRMTs. https://godbolt.org/z/fqnfhMTs6


```llvm
; [x,3,5,9]
; PRMT(x, 3 | (5 << 8) | (9 << 16), 0x6540)
define i32 @pack_x_3_5_9(i32 %input) {
  %sum = add i32 %input, 1
  %x = trunc i32 %sum to i8
  %v = insertelement <4 x i8> <i8 0, i8 3, i8 5, i8 9>, i8 %x, i32 0
  %packed = bitcast <4 x i8> %v to i32
  %result = xor i32 %packed, 305419866
  ret i32 %result
}

; [3,x,5,9]
; PRMT(x, 3 | (5 << 8) | (9 << 16), 0x6504)
define i32 @pack_3_x_5_9(i32 %input) {
  %sum = add i32 %input, 1
  %x = trunc i32 %sum to i8
  %v = insertelement <4 x i8> <i8 3, i8 0, i8 5, i8 9>, i8 %x, i32 1
  %packed = bitcast <4 x i8> %v to i32
  %result = xor i32 %packed, 305419866
  ret i32 %result
}
```

There are also optimization opportunities for other mixed variable/constant patterns, for example:
Pattern | currently PRMTs | expected PRMTs
-- | -- | --
[x,3,5,9] | 2 | 1
[x,x,3,5] | 2 | 1
[x,0,x,0] | 2 | 1
[x,y,0,0] | 2 | 1*
[x,y,3,5] | 2 | 2
[x,y,z,0] | 3 | 2*
[x,y,z,7] | 3 | 3

\* required a known-zero byte

Would it make sense to keep the undef-to-zero change in this PR and address the single-PRMT lowering, including these mixed variable  & constant cases, in a follow-up PR? Any suggestions? Thanks!

https://github.com/llvm/llvm-project/pull/224575


More information about the llvm-commits mailing list