[llvm] [SelectionDAG] Freeze dynamic vector indices lowered through the stack (PR #225031)

Mahmoud Naderi via llvm-commits llvm-commits at lists.llvm.org
Thu Sep 24 03:04:24 PDT 2026


================
@@ -818,74 +818,97 @@ define <16 x i8> @var_shuffle_v16i8(<16 x i8> %v, <16 x i8> %indices) nounwind {
 define <16 x i8> @var_shuffle_zero_v16i8(<16 x i8> %v, <16 x i8> %indices) nounwind {
 ; SSE3-LABEL: var_shuffle_zero_v16i8:
 ; SSE3:       # %bb.0:
+; SSE3-NEXT:    pushq %rbp
+; SSE3-NEXT:    pushq %r15
+; SSE3-NEXT:    pushq %r14
+; SSE3-NEXT:    pushq %r13
+; SSE3-NEXT:    pushq %r12
+; SSE3-NEXT:    pushq %rbx
 ; SSE3-NEXT:    movaps %xmm0, %xmm2
 ; SSE3-NEXT:    movdqa {{.*#+}} xmm0 = [16,16,16,16,16,16,16,16,16,16,16,16,16,16,16,16]
 ; SSE3-NEXT:    pmaxub %xmm1, %xmm0
 ; SSE3-NEXT:    pcmpeqb %xmm1, %xmm0
 ; SSE3-NEXT:    por %xmm0, %xmm1
 ; SSE3-NEXT:    movdqa %xmm1, -40(%rsp)
-; SSE3-NEXT:    movaps %xmm2, -24(%rsp)
----------------
slant14 wrote:

On SSE3 the index is already a stack load when the freeze is added, so freeze(load) blocks the zextload fold and increases register pressure. The freeze is needed for correctness because the index can be poison. I can try freezing earlier so it becomes a single freeze of the source vector. (I haven't tried freezing earlier but if you want I can give it a try)

https://github.com/llvm/llvm-project/pull/225031


More information about the llvm-commits mailing list