[llvm] [X86] Always use 128-bit V_SET0/AVX512_128_SET0 patterns, along with SUBREG_TO_REG for extension to 256/512-bit vectors (PR #212950)
Alina Sbirlea via llvm-commits
llvm-commits at lists.llvm.org
Wed Aug 26 13:47:37 PDT 2026
alinas wrote:
I think you're right that the explanation regarding the hardware instruction definitions from Intel's manual was incorrect. I apologize for not noticing that.
For the test that started failing, there appears to be a real issue between +avx512f, and +avx512f,+avx512vl where garbage starts getting read from the wider register in the +avx512f case.
I'm trying to figure this out, but could really use some help from folks directly working on this part of codegen.
The above test case might be pointing in the right direction.
An instruction like`call void asm sideeffect "# use $0", "{zmm16}"(<16 x i32> zeroinitializer)` used to get lowered to
```
vpxord %zmm16, %zmm16, %zmm16
#APP
# use %zmm16
#NO_APP
```
and with this patch it gets lowered to
```
vxorps %xmm0, %xmm0, %xmm0
vmovaps %zmm0, %zmm16
#APP
# use %zmm16
#NO_APP
vzeroupper
```
So it looks plausible that this change caused the 512-bit register to not get fully zeroed out - it would at least match what I'm seeing in this specific test.
I could be wrong here though, so some feedback would be appreciated.
https://github.com/llvm/llvm-project/pull/212950
More information about the llvm-commits
mailing list