[llvm] [AArch64] Avoid scalar DUP for promoted narrow extracts (PR #221186)
Oscar Priego via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 8 00:10:29 PDT 2026
Opriego wrote:
Thanks, you're right. I traced this further and the original fix was addressing the problem at the wrong boundary.
The promoted `EXTRACT_VECTOR_ELT` itself is fine: its high bits are not supposed to matter. The actual problem was in AArch64 `LowerBUILD_VECTOR()`, where a promoted extract from an `i8` vector element could be used to form a scalar-fed `DUP` with `i16` lanes. That made bits above the original element width observable.
I've revised the patch to remove the `SIGN_EXTEND_INREG` workaround entirely. Instead, AArch64 now avoids that scalar `DUP` lowering when the destination vector element is wider than the element the promoted extract came from, and lets the existing vector lowering path preserve the original element granularity.
For the reported case, this now stays as a byte-lane `dup` instead of extracting to a GPR and duplicating as halfwords.
I also reduced the fix to the specific `LowerBUILD_VECTOR()` path that produces the invalid widening. The full AArch64 CodeGen suite passes with 4204 PASS, 4 expected failures, and 0 unexpected failures.
https://github.com/llvm/llvm-project/pull/221186
More information about the llvm-commits
mailing list