[llvm] [SelectionDAG] Freeze dynamic vector indices lowered through the stack (PR #225031)
Mahmoud Naderi via llvm-commits
llvm-commits at lists.llvm.org
Thu Sep 24 03:04:24 PDT 2026
================
@@ -818,74 +818,97 @@ define <16 x i8> @var_shuffle_v16i8(<16 x i8> %v, <16 x i8> %indices) nounwind {
define <16 x i8> @var_shuffle_zero_v16i8(<16 x i8> %v, <16 x i8> %indices) nounwind {
; SSE3-LABEL: var_shuffle_zero_v16i8:
; SSE3: # %bb.0:
+; SSE3-NEXT: pushq %rbp
+; SSE3-NEXT: pushq %r15
+; SSE3-NEXT: pushq %r14
+; SSE3-NEXT: pushq %r13
+; SSE3-NEXT: pushq %r12
+; SSE3-NEXT: pushq %rbx
; SSE3-NEXT: movaps %xmm0, %xmm2
; SSE3-NEXT: movdqa {{.*#+}} xmm0 = [16,16,16,16,16,16,16,16,16,16,16,16,16,16,16,16]
; SSE3-NEXT: pmaxub %xmm1, %xmm0
; SSE3-NEXT: pcmpeqb %xmm1, %xmm0
; SSE3-NEXT: por %xmm0, %xmm1
; SSE3-NEXT: movdqa %xmm1, -40(%rsp)
-; SSE3-NEXT: movaps %xmm2, -24(%rsp)
----------------
slant14 wrote:
On SSE3 the index is already a stack load when the freeze is added, so freeze(load) blocks the zextload fold and increases register pressure. The freeze is needed for correctness because the index can be poison. I can try freezing earlier so it becomes a single freeze of the source vector. (I haven't tried freezing earlier but if you want I can give it a try)
https://github.com/llvm/llvm-project/pull/225031
More information about the llvm-commits
mailing list