[llvm] [NVPTX] Use zero for undef bytes in v4i8 construction (PR #224575)
via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 21 06:15:59 PDT 2026
sinan-flash wrote:
> Interesting, can't this be done in 1 PRMT? Maybe we should specialize LowerBUILD_VECTOR for undef/poison operands?
@AlexMaclean Thanks for the feedback! A single PRMT would suffice for this specific case. And I think this exposes a broader issue with mixed variable/constant inputs in `LowerBUILD_VECTOR`.
There are two related problems/optimization chances:
- undef/poison mixed with non-constant operands
We can address this one in this PR. My current reproducer uses undef; I haven’t encountered a poison case yet, but I can add poison case if useful.
<br>
- LowerBUILD_VECTOR handles fully constant inputs well, but lowers other v4i8 inputs through a fixed three-PRMT sequence, which can be suboptimal.
For testcase in this PR, choosing zero for the undef elements effectively gives:
```
[x, undef, undef, undef] → [x, 0, 0, 0]
lo = PRMT(x, 0, 0x3340)
hi = PRMT(0, 0, 0x3340)
r = PRMT(lo, hi, 0x5410)
```
then combine stage(`combinePRMT`) folds full constant prmt `hi` to zero, leaving:
```
t = PRMT(x, 0, 0x3340)
r = PRMT(t, 0, 0x5410)
```
But we could construct the result directly with `r = PRMT(x, 0, 0x4440)`.
<br>
This isn’t specific to undef. LowerBUILD_VECTOR also misses these mixed variable/constant cases. More generally, a single variable mixed with constants can use one PRMT, for example:
e.g. one scalar pattern can be done with only one prmt, but the actual results are both 2 PRMTs. https://godbolt.org/z/fqnfhMTs6
```llvm
; [x,3,5,9]
; PRMT(x, 3 | (5 << 8) | (9 << 16), 0x6540)
define i32 @pack_x_3_5_9(i32 %input) {
%sum = add i32 %input, 1
%x = trunc i32 %sum to i8
%v = insertelement <4 x i8> <i8 0, i8 3, i8 5, i8 9>, i8 %x, i32 0
%packed = bitcast <4 x i8> %v to i32
%result = xor i32 %packed, 305419866
ret i32 %result
}
; [3,x,5,9]
; PRMT(x, 3 | (5 << 8) | (9 << 16), 0x6504)
define i32 @pack_3_x_5_9(i32 %input) {
%sum = add i32 %input, 1
%x = trunc i32 %sum to i8
%v = insertelement <4 x i8> <i8 3, i8 0, i8 5, i8 9>, i8 %x, i32 1
%packed = bitcast <4 x i8> %v to i32
%result = xor i32 %packed, 305419866
ret i32 %result
}
```
There are also optimization opportunities for other mixed variable/constant patterns, for example:
Pattern | currently PRMTs | expected PRMTs
-- | -- | --
[x,3,5,9] | 2 | 1
[x,x,3,5] | 2 | 1
[x,0,x,0] | 2 | 1
[x,y,0,0] | 2 | 1*
[x,y,3,5] | 2 | 2
[x,y,z,0] | 3 | 2*
[x,y,z,7] | 3 | 3
\* required a known-zero byte
Would it make sense to keep the undef-to-zero change in this PR and address the single-PRMT lowering, including these mixed variable & constant cases, in a follow-up PR? Any suggestions? Thanks!
https://github.com/llvm/llvm-project/pull/224575
More information about the llvm-commits
mailing list