[llvm] [DAG] Re-add transitive users when FREEZE folds away (PR #193945)

via llvm-commits llvm-commits at lists.llvm.org
Fri Apr 24 04:32:48 PDT 2026


llvmbot wrote:


<!--LLVM PR SUMMARY COMMENT-->

@llvm/pr-subscribers-llvm-selectiondag

Author: Paweł Bylica (chfast)

<details>
<summary>Changes</summary>

When DAGCombiner::visitFREEZE folds a FREEZE of a guaranteed-non-poison value away, it returns the underlying value and relies on the combine framework to re-add the FREEZE's direct users. That is enough when the chain cascades: each re-visited user either simplifies (triggering ReplaceAllUsesWith on its own, which re-adds _its_ users) or is the consumer that cares.

It is not enough when an intermediate node in the chain is "transparent" to a peephole combine that lives on its users. The driving case is X86's checkBoolTestSetCCCombine, which is invoked from combineCMov, combineBrCond and combineSetCC and looks through

  CMOV/BRCOND/SETCC(E, CMP(AND(FREEZE(X86ISD::SETCC(COND, EFLAGS)), 1), 0))

Short-circuit AND lowering turns `select a, b, false` into `select a, freeze(b), 0` (foldBoolSelectToLogic), and a multi-word usub-with-carry chain ends up in that exact shape. On the first CMOV visit the combine stops at the FREEZE. visitFREEZE later folds the FREEZE (X86ISD::SETCC is marked unconditionally non-poison), AND simplifies to SETCC, and CMP's operand is updated — but CMP itself does not change identity, so CMP's users (the CMOV) are never re-added. The redundant setb+testb+cmove triple after the final SBB survives to codegen.

Fix: when visitFREEZE returns N0, walk the use chain of the freeze for up to 32 nodes and re-add them to the combine worklist. Worklist deduplication makes the cost on nodes that have nothing new to do negligible.

---

Patch is 2.01 MiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/193945.diff


6 Files Affected:

- (modified) llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp (+17-1) 
- (modified) llvm/test/CodeGen/RISCV/clmul.ll (+5237-5251) 
- (modified) llvm/test/CodeGen/RISCV/clmulh.ll (+7553-7567) 
- (modified) llvm/test/CodeGen/RISCV/clmulr.ll (+7274-7318) 
- (added) llvm/test/CodeGen/X86/cmp-setcc-cmov-reworklist.ll (+75) 
- (modified) llvm/test/CodeGen/X86/freeze-binary.ll (+7-13) 


``````````diff
diff --git a/llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp b/llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
index 20ff48789152b..2c5a454364db5 100644
--- a/llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
@@ -17641,8 +17641,24 @@ SDValue DAGCombiner::visitBUILD_PAIR(SDNode *N) {
 SDValue DAGCombiner::visitFREEZE(SDNode *N) {
   SDValue N0 = N->getOperand(0);
 
-  if (DAG.isGuaranteedNotToBeUndefOrPoison(N0, /*PoisonOnly*/ false))
+  if (DAG.isGuaranteedNotToBeUndefOrPoison(N0, /*PoisonOnly*/ false)) {
+    // Eliminating a freeze can unlock combines on nodes farther up the use
+    // chain — e.g., peephole combines that peek through intermediate nodes
+    // (CMP/AND/ZEXT) to reach a SETCC under the freeze. Normal worklist
+    // propagation only re-adds the direct users of the replaced node, which
+    // is not enough when those intermediate nodes do not themselves
+    // simplify. Re-add a bounded transitive user set (the set is for BFS
+    // dedup, separate from the worklist's own dedup).
+    // TODO: Not needed under -combiner-topological-sorting, which already
+    // visits operands before users. Remove when topo order is the default.
+    constexpr unsigned Limit = 32;
+    SmallSetVector<SDNode *, Limit> Users(N->users().begin(), N->users().end());
+    for (unsigned I = 0; I != Users.size() && Users.size() < Limit; ++I)
+      Users.insert_range(Users[I]->users());
+    for (SDNode *U : Users)
+      AddToWorklist(U);
     return N0;
+  }
 
   // If we have frozen and unfrozen users of N0, update so everything uses N.
   if (!N0.isUndef() && !N0.hasOneUse()) {
diff --git a/llvm/test/CodeGen/RISCV/clmul.ll b/llvm/test/CodeGen/RISCV/clmul.ll
index d9bbe20b8730d..5545fff6c9e60 100644
--- a/llvm/test/CodeGen/RISCV/clmul.ll
+++ b/llvm/test/CodeGen/RISCV/clmul.ll
@@ -2749,7 +2749,7 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
 ; RV32I-NEXT:    xor s1, a0, s1
 ; RV32I-NEXT:    lw a0, 60(sp) # 4-byte Folded Reload
 ; RV32I-NEXT:    xor ra, a0, ra
-; RV32I-NEXT:    xor s8, s8, s6
+; RV32I-NEXT:    xor s6, s8, s6
 ; RV32I-NEXT:    lw a0, 200(sp) # 4-byte Folded Reload
 ; RV32I-NEXT:    xor t6, a2, a0
 ; RV32I-NEXT:    lw a0, 288(sp) # 4-byte Folded Reload
@@ -2760,10 +2760,10 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
 ; RV32I-NEXT:    xor a5, a5, a0
 ; RV32I-NEXT:    slli a2, a3, 11
 ; RV32I-NEXT:    slli s2, s2, 11
-; RV32I-NEXT:    li s6, 1
-; RV32I-NEXT:    slli s6, s6, 11
-; RV32I-NEXT:    and a0, t3, s6
-; RV32I-NEXT:    and a1, t0, s6
+; RV32I-NEXT:    li s8, 1
+; RV32I-NEXT:    slli s8, s8, 11
+; RV32I-NEXT:    and a0, t3, s8
+; RV32I-NEXT:    and a1, t0, s8
 ; RV32I-NEXT:    seqz a0, a0
 ; RV32I-NEXT:    seqz a1, a1
 ; RV32I-NEXT:    addi a0, a0, -1
@@ -2799,7 +2799,7 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
 ; RV32I-NEXT:    lw s0, 72(sp) # 4-byte Folded Reload
 ; RV32I-NEXT:    xor s0, ra, s0
 ; RV32I-NEXT:    lw s1, 12(sp) # 4-byte Folded Reload
-; RV32I-NEXT:    xor s8, s8, s1
+; RV32I-NEXT:    xor s6, s6, s1
 ; RV32I-NEXT:    xor a4, t6, a4
 ; RV32I-NEXT:    lw t6, 272(sp) # 4-byte Folded Reload
 ; RV32I-NEXT:    xor a5, a5, t6
@@ -2828,7 +2828,7 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
 ; RV32I-NEXT:    lw t4, 80(sp) # 4-byte Folded Reload
 ; RV32I-NEXT:    xor t4, s0, t4
 ; RV32I-NEXT:    lw t5, 20(sp) # 4-byte Folded Reload
-; RV32I-NEXT:    xor t5, s8, t5
+; RV32I-NEXT:    xor t5, s6, t5
 ; RV32I-NEXT:    xor a4, a4, a5
 ; RV32I-NEXT:    lw a5, 252(sp) # 4-byte Folded Reload
 ; RV32I-NEXT:    xor a5, a0, a5
@@ -2897,271 +2897,271 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
 ; RV32I-NEXT:    xor a0, a5, a0
 ; RV32I-NEXT:    sw a0, 304(sp) # 4-byte Folded Spill
 ; RV32I-NEXT:    lui a1, 209715
-; RV32I-NEXT:    addi t6, a1, 819
+; RV32I-NEXT:    addi t5, a1, 819
 ; RV32I-NEXT:    srli a3, s1, 2
-; RV32I-NEXT:    and s1, s1, t6
-; RV32I-NEXT:    and a3, a3, t6
+; RV32I-NEXT:    and s1, s1, t5
+; RV32I-NEXT:    and a3, a3, t5
 ; RV32I-NEXT:    slli s1, s1, 2
 ; RV32I-NEXT:    or a3, a3, s1
 ; RV32I-NEXT:    srli a4, s2, 2
-; RV32I-NEXT:    and a5, s2, t6
-; RV32I-NEXT:    and a4, a4, t6
+; RV32I-NEXT:    and a5, s2, t5
+; RV32I-NEXT:    and a4, a4, t5
 ; RV32I-NEXT:    slli a5, a5, 2
 ; RV32I-NEXT:    or a4, a4, a5
 ; RV32I-NEXT:    xor a2, a2, a6
 ; RV32I-NEXT:    xor a5, a7, t0
 ; RV32I-NEXT:    xor a2, a5, a2
 ; RV32I-NEXT:    sw a2, 300(sp) # 4-byte Folded Spill
-; RV32I-NEXT:    lui s2, 349525
-; RV32I-NEXT:    addi s2, s2, 1365
+; RV32I-NEXT:    lui t6, 349525
+; RV32I-NEXT:    addi t6, t6, 1365
 ; RV32I-NEXT:    srli a2, a3, 1
-; RV32I-NEXT:    and a3, a3, s2
-; RV32I-NEXT:    and a2, a2, s2
+; RV32I-NEXT:    and a3, a3, t6
+; RV32I-NEXT:    and a2, a2, t6
 ; RV32I-NEXT:    slli a3, a3, 1
-; RV32I-NEXT:    or t5, a2, a3
+; RV32I-NEXT:    or t4, a2, a3
 ; RV32I-NEXT:    srli a2, a4, 1
-; RV32I-NEXT:    and a3, a4, s2
-; RV32I-NEXT:    and a2, a2, s2
+; RV32I-NEXT:    and a3, a4, t6
+; RV32I-NEXT:    and a2, a2, t6
 ; RV32I-NEXT:    slli a3, a3, 1
 ; RV32I-NEXT:    or t2, a2, a3
 ; RV32I-NEXT:    srli a3, a3, 31
 ; RV32I-NEXT:    seqz a2, a3
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 31
+; RV32I-NEXT:    slli a3, t4, 31
 ; RV32I-NEXT:    and a2, a2, a3
 ; RV32I-NEXT:    sw a2, 296(sp) # 4-byte Folded Spill
 ; RV32I-NEXT:    andi a2, t2, 2
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 1
+; RV32I-NEXT:    slli a3, t4, 1
 ; RV32I-NEXT:    and a2, a2, a3
 ; RV32I-NEXT:    sw a2, 292(sp) # 4-byte Folded Spill
 ; RV32I-NEXT:    andi a2, t2, 4
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 2
+; RV32I-NEXT:    slli a3, t4, 2
 ; RV32I-NEXT:    and a2, a2, a3
 ; RV32I-NEXT:    sw a2, 288(sp) # 4-byte Folded Spill
 ; RV32I-NEXT:    andi a2, t2, 8
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 3
+; RV32I-NEXT:    slli a3, t4, 3
 ; RV32I-NEXT:    and a2, a2, a3
 ; RV32I-NEXT:    sw a2, 284(sp) # 4-byte Folded Spill
 ; RV32I-NEXT:    andi a2, t2, 16
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 4
-; RV32I-NEXT:    and a2, a2, a3
-; RV32I-NEXT:    sw a2, 280(sp) # 4-byte Folded Spill
+; RV32I-NEXT:    slli s0, t4, 4
+; RV32I-NEXT:    and a2, a2, s0
+; RV32I-NEXT:    sw a2, 276(sp) # 4-byte Folded Spill
 ; RV32I-NEXT:    andi a2, t2, 32
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli s1, t5, 5
-; RV32I-NEXT:    and a2, a2, s1
-; RV32I-NEXT:    sw a2, 272(sp) # 4-byte Folded Spill
+; RV32I-NEXT:    slli a3, t4, 5
+; RV32I-NEXT:    and a2, a2, a3
+; RV32I-NEXT:    sw a2, 268(sp) # 4-byte Folded Spill
 ; RV32I-NEXT:    andi a2, t2, 64
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli s0, t5, 6
-; RV32I-NEXT:    and a2, a2, s0
-; RV32I-NEXT:    sw a2, 276(sp) # 4-byte Folded Spill
+; RV32I-NEXT:    slli s1, t4, 6
+; RV32I-NEXT:    and a2, a2, s1
+; RV32I-NEXT:    sw a2, 272(sp) # 4-byte Folded Spill
 ; RV32I-NEXT:    andi a2, t2, 128
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 7
+; RV32I-NEXT:    slli a3, t4, 7
 ; RV32I-NEXT:    and a2, a2, a3
-; RV32I-NEXT:    sw a2, 268(sp) # 4-byte Folded Spill
+; RV32I-NEXT:    sw a2, 280(sp) # 4-byte Folded Spill
 ; RV32I-NEXT:    andi a2, t2, 256
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 8
-; RV32I-NEXT:    and s5, a2, a3
+; RV32I-NEXT:    slli a3, t4, 8
+; RV32I-NEXT:    and a2, a2, a3
+; RV32I-NEXT:    sw a2, 264(sp) # 4-byte Folded Spill
 ; RV32I-NEXT:    andi a2, t2, 512
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 9
-; RV32I-NEXT:    and a2, a2, a3
-; RV32I-NEXT:    sw a2, 260(sp) # 4-byte Folded Spill
+; RV32I-NEXT:    slli a3, t4, 9
+; RV32I-NEXT:    and s6, a2, a3
 ; RV32I-NEXT:    andi a2, t2, 1024
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 10
+; RV32I-NEXT:    slli a3, t4, 10
 ; RV32I-NEXT:    and a2, a2, a3
-; RV32I-NEXT:    sw a2, 264(sp) # 4-byte Folded Spill
-; RV32I-NEXT:    and a2, t2, s6
+; RV32I-NEXT:    sw a2, 260(sp) # 4-byte Folded Spill
+; RV32I-NEXT:    and a2, t2, s8
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 11
-; RV32I-NEXT:    and s1, a2, a3
+; RV32I-NEXT:    slli a3, t4, 11
+; RV32I-NEXT:    and s3, a2, a3
 ; RV32I-NEXT:    lui a0, 1
 ; RV32I-NEXT:    and a2, t2, a0
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 12
-; RV32I-NEXT:    and s9, a2, a3
+; RV32I-NEXT:    slli a3, t4, 12
+; RV32I-NEXT:    and s5, a2, a3
 ; RV32I-NEXT:    lui a0, 2
 ; RV32I-NEXT:    and a2, t2, a0
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 13
+; RV32I-NEXT:    slli a3, t4, 13
 ; RV32I-NEXT:    and s7, a2, a3
 ; RV32I-NEXT:    lui a0, 4
 ; RV32I-NEXT:    and a2, t2, a0
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 14
-; RV32I-NEXT:    and s6, a2, a3
+; RV32I-NEXT:    slli a3, t4, 14
+; RV32I-NEXT:    and s11, a2, a3
 ; RV32I-NEXT:    lui a0, 8
 ; RV32I-NEXT:    and a2, t2, a0
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 15
+; RV32I-NEXT:    slli a3, t4, 15
 ; RV32I-NEXT:    and s8, a2, a3
 ; RV32I-NEXT:    lui a0, 16
 ; RV32I-NEXT:    and a2, t2, a0
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 16
-; RV32I-NEXT:    and s10, a2, a3
+; RV32I-NEXT:    slli a3, t4, 16
+; RV32I-NEXT:    and s9, a2, a3
 ; RV32I-NEXT:    lui a0, 32
 ; RV32I-NEXT:    and a2, t2, a0
 ; RV32I-NEXT:    seqz a2, a2
 ; RV32I-NEXT:    addi a2, a2, -1
-; RV32I-NEXT:    slli a3, t5, 17
-; RV32I-NEXT:    and s0, a2, a3
+; RV32I-NEXT:    slli a3, t4, 17
+; RV32I-NEXT:    and s10, a2, a3
 ; RV32I-NEXT:    lui a0, 64
-; RV32I-NEXT:    and a3, t2, a0
-; RV32I-NEXT:    seqz a3, a3
-; RV32I-NEXT:    addi a3, a3, -1
-; RV32I-NEXT:    slli a4, t5, 18
-; RV32I-NEXT:    and ra, a3, a4
+; RV32I-NEXT:    and a2, t2, a0
+; RV32I-NEXT:    seqz a2, a2
+; RV32I-NEXT:    addi a2, a2, -1
+; RV32I-NEXT:    slli a3, t4, 18
+; RV32I-NEXT:    and ra, a2, a3
 ; RV32I-NEXT:    lui a0, 128
-; RV32I-NEXT:    and a3, t2, a0
-; RV32I-NEXT:    seqz a3, a3
-; RV32I-NEXT:    addi a3, a3, -1
-; RV32I-NEXT:    slli a4, t5, 19
-; RV32I-NEXT:    and s11, a3, a4
+; RV32I-NEXT:    and a2, t2, a0
+; RV32I-NEXT:    seqz a2, a2
+; RV32I-NEXT:    addi a2, a2, -1
+; RV32I-NEXT:    slli a3, t4, 19
+; RV32I-NEXT:    and s1, a2, a3
 ; RV32I-NEXT:    lui a0, 256
+; RV32I-NEXT:    and a2, t2, a0
+; RV32I-NEXT:    seqz a2, a2
+; RV32I-NEXT:    addi a2, a2, -1
+; RV32I-NEXT:    slli a3, t4, 20
+; RV32I-NEXT:    and s0, a2, a3
+; RV32I-NEXT:    lui a0, 512
 ; RV32I-NEXT:    and a3, t2, a0
 ; RV32I-NEXT:    seqz a3, a3
 ; RV32I-NEXT:    addi a3, a3, -1
-; RV32I-NEXT:    slli a4, t5, 20
-; RV32I-NEXT:    and a4, a3, a4
-; RV32I-NEXT:    lui a0, 512
+; RV32I-NEXT:    slli a5, t4, 21
+; RV32I-NEXT:    and a5, a3, a5
+; RV32I-NEXT:    lui a0, 1024
 ; RV32I-NEXT:    and a3, t2, a0
 ; RV32I-NEXT:    seqz a3, a3
-; RV32I-NEXT:    addi a0, a3, -1
-; RV32I-NEXT:    slli a3, t5, 21
-; RV32I-NEXT:    and a3, a0, a3
-; RV32I-NEXT:    lui a0, 1024
+; RV32I-NEXT:    addi a3, a3, -1
+; RV32I-NEXT:    slli a4, t4, 22
+; RV32I-NEXT:    and a3, a3, a4
+; RV32I-NEXT:    lui a0, 2048
+; RV32I-NEXT:    and a4, t2, a0
+; RV32I-NEXT:    seqz a4, a4
+; RV32I-NEXT:    addi a0, a4, -1
+; RV32I-NEXT:    slli a4, t4, 23
+; RV32I-NEXT:    and a4, a0, a4
+; RV32I-NEXT:    lui a0, 4096
 ; RV32I-NEXT:    and a0, t2, a0
 ; RV32I-NEXT:    seqz a0, a0
 ; RV32I-NEXT:    addi a0, a0, -1
-; RV32I-NEXT:    slli t0, t5, 22
+; RV32I-NEXT:    slli t0, t4, 24
 ; RV32I-NEXT:    and a2, a0, t0
-; RV32I-NEXT:    lui a0, 2048
+; RV32I-NEXT:    lui a0, 8192
 ; RV32I-NEXT:    and t0, t2, a0
 ; RV32I-NEXT:    seqz t0, t0
 ; RV32I-NEXT:    addi t0, t0, -1
-; RV32I-NEXT:    slli a1, t5, 23
+; RV32I-NEXT:    slli a1, t4, 25
 ; RV32I-NEXT:    and a1, t0, a1
-; RV32I-NEXT:    lui a0, 4096
-; RV32I-NEXT:    and t0, t2, a0
-; RV32I-NEXT:    seqz t0, t0
-; RV32I-NEXT:    addi t0, t0, -1
-; RV32I-NEXT:    slli a5, t5, 24
-; RV32I-NEXT:    and a5, t0, a5
-; RV32I-NEXT:    lui a0, 8192
+; RV32I-NEXT:    lui a0, 16384
 ; RV32I-NEXT:    and t0, t2, a0
 ; RV32I-NEXT:    seqz t0, t0
 ; RV32I-NEXT:    addi t0, t0, -1
-; RV32I-NEXT:    slli a6, t5, 25
+; RV32I-NEXT:    slli a6, t4, 26
 ; RV32I-NEXT:    and a6, t0, a6
-; RV32I-NEXT:    lui a0, 16384
+; RV32I-NEXT:    lui a0, 32768
 ; RV32I-NEXT:    and t0, t2, a0
 ; RV32I-NEXT:    seqz t0, t0
 ; RV32I-NEXT:    addi t0, t0, -1
-; RV32I-NEXT:    slli a7, t5, 26
+; RV32I-NEXT:    slli a7, t4, 27
 ; RV32I-NEXT:    and a7, t0, a7
-; RV32I-NEXT:    lui a0, 32768
+; RV32I-NEXT:    lui a0, 65536
 ; RV32I-NEXT:    and t0, t2, a0
 ; RV32I-NEXT:    seqz t0, t0
 ; RV32I-NEXT:    addi t0, t0, -1
-; RV32I-NEXT:    slli t1, t5, 27
+; RV32I-NEXT:    slli t1, t4, 28
 ; RV32I-NEXT:    and t0, t0, t1
-; RV32I-NEXT:    lui a0, 65536
+; RV32I-NEXT:    lui a0, 131072
 ; RV32I-NEXT:    and t1, t2, a0
 ; RV32I-NEXT:    seqz t1, t1
 ; RV32I-NEXT:    addi t1, t1, -1
-; RV32I-NEXT:    slli t3, t5, 28
+; RV32I-NEXT:    slli t3, t4, 29
 ; RV32I-NEXT:    and t1, t1, t3
-; RV32I-NEXT:    lui a0, 131072
-; RV32I-NEXT:    and t3, t2, a0
-; RV32I-NEXT:    seqz t3, t3
-; RV32I-NEXT:    addi t3, t3, -1
-; RV32I-NEXT:    slli t4, t5, 29
-; RV32I-NEXT:    and t3, t3, t4
 ; RV32I-NEXT:    lui a0, 262144
-; RV32I-NEXT:    and t4, t2, a0
+; RV32I-NEXT:    and t3, t2, a0
 ; RV32I-NEXT:    andi t2, t2, 1
 ; RV32I-NEXT:    seqz t2, t2
 ; RV32I-NEXT:    addi t2, t2, -1
-; RV32I-NEXT:    and t2, t2, t5
-; RV32I-NEXT:    slli t5, t5, 30
-; RV32I-NEXT:    seqz t4, t4
-; RV32I-NEXT:    addi t4, t4, -1
-; RV32I-NEXT:    and t4, t4, t5
+; RV32I-NEXT:    and t2, t2, t4
+; RV32I-NEXT:    slli t4, t4, 30
+; RV32I-NEXT:    seqz t3, t3
+; RV32I-NEXT:    addi t3, t3, -1
+; RV32I-NEXT:    and t3, t3, t4
 ; RV32I-NEXT:    lw a0, 292(sp) # 4-byte Folded Reload
 ; RV32I-NEXT:    xor t2, t2, a0
 ; RV32I-NEXT:    lw a0, 288(sp) # 4-byte Folded Reload
-; RV32I-NEXT:    lw t5, 284(sp) # 4-byte Folded Reload
-; RV32I-NEXT:    xor t5, a0, t5
-; RV32I-NEXT:    lw a0, 280(sp) # 4-byte Folded Reload
-; RV32I-NEXT:    lw s3, 272(sp) # 4-byte Folded Reload
-; RV32I-NEXT:    xor a0, a0, s3
-; RV32I-NEXT:    lw s3, 268(sp) # 4-byte Folded Reload
-; RV32I-NEXT:    xor s5, s3, s5
-; RV32I-NEXT:    xor s1, s1, s9
-; RV32I-NEXT:    xor s0, s10, s0
+; RV32I-NEXT:    lw t4, 284(sp) # 4-byte Folded Reload
+; RV32I-NEXT:    xor t4, a0, t4
+; RV32I-NEXT:    lw a0, 276(sp) # 4-byte Folded Reload
+; RV32I-NEXT:    lw s2, 268(sp) # 4-byte Folded Reload
+; RV32I-NEXT:    xor a0, a0, s2
+; RV32I-NEXT:    lw s2, 264(sp) # 4-byte Folded Reload
+; RV32I-NEXT:    xor s6, s2, s6
+; RV32I-NEXT:    xor s7, s7, s11
+; RV32I-NEXT:    xor s0, s1, s0
 ; RV32I-NEXT:    xor a1, a2, a1
-; RV32I-NEXT:    xor a2, t3, t4
-; RV32I-NEXT:    xor t2, t2, t5
-; RV32I-NEXT:    lw t3, 276(sp) # 4-byte Folded Reload
-; RV32I-NEXT:    xor a0, a0, t3
-; RV32I-NEXT:    lw t3, 260(sp) # 4-byte Folded Reload
-; RV32I-NEXT:    xor t3, s5, t3
-; RV32I-NEXT:    xor t4, s1, s7
-; RV32I-NEXT:    xor t5, s0, ra
-; RV32I-NEXT:    xor a1, a1, a5
-; RV32I-NEXT:    lw a5, 296(sp) # 4-byte Folded Reload
-; RV32I-NEXT:    xor a2, a2, a5
-; RV32I-NEXT:    xor a0, t2, a0
-; RV32I-NEXT:    lw a5, 264(sp) # 4-byte Folded Reload
-; RV32I-NEXT:    xor a5, t3, a5
-; RV32I-NEXT:    xor t2, t4, s6
-; RV32I-NEXT:    xor t3, t5, s11
+; RV32I-NEXT:    xor a2, t2, t4
+; RV32I-NEXT:    lw t2, 272(sp) # 4-byte Folded Reload
+; RV32I-NEXT:    xor a0, a0, t2
+; RV32I-NEXT:    lw t2, 260(sp) # 4-byte Folded Reload
+; RV32I-NEXT:    xor t2, s6, t2
+; RV32I-NEXT:    xor t4, s7, s8
+; RV32I-NEXT:    xor a5, s0, a5
 ; RV32I-NEXT:    xor a1, a1, a6
-; RV32I-NEXT:    xor a0, a0, a5
-; RV32I-NEXT:    xor a5, t2, s8
-; RV32I-NEXT:    xor a4, t3, a4
+; RV32I-NEXT:    xor a0, a2, a0
+; RV32I-NEXT:    xor a2, t2, s3
+; RV32I-NEXT:    xor a6, t4, s9
+; RV32I-NEXT:    xor a3, a5, a3
 ; RV32I-NEXT:    xor a1, a1, a7
+; RV32I-NEXT:    lw a5, 280(sp) # 4-byte Folded Reload
 ; RV32I-NEXT:    xor a0, a0, a5
-; RV32I-NEXT:    xor a3, a4, a3
+; RV32I-NEXT:    xor a2, a2, s5
+; RV32I-NEXT:    xor a5, a6, s10
+; RV32I-NEXT:    xor a3, a3, a4
 ; RV32I-NEXT:    xor a1, a1, t0
-; RV32I-NEXT:    xor a0, a0, a3
+; RV32I-NEXT:    xor a4, a5, ra
 ; RV32I-NEXT:    xor a1, a1, t1
-; RV32I-NEXT:    xor a0, a0, a1
-; RV32I-NEXT:    xor a0, a0, a2
-; RV32I-NEXT:    srli a1, a0, 8
-; RV32I-NEXT:    srli a2, a0, 24
-; RV32I-NEXT:    lw a3, 312(sp) # 4-byte Folded Reload
-; RV32I-NEXT:    and a1, a1, a3
-; RV32I-NEXT:    or a1, a1, a2
-; RV32I-NEXT:    and a2, a0, a3
+; RV32I-NEXT:    xor a2, a0, a2
+; RV32I-NEXT:    xor a2, a2, a4
+; RV32I-NEXT:    xor a1, a1, t3
+; RV32I-NEXT:    xor a2, a2, a3
+; RV32I-NEXT:    lw a3, 296(sp) # 4-byte Folded Reload
+; RV32I-NEXT:    xor a1, a1, a3
+; RV32I-NEXT:    lw a4, 312(sp) # 4-byte Folded Reload
+; RV32I-NEXT:    and a3, a2, a4
+; RV32I-NEXT:    xor a1, a2, a1
+; RV32I-NEXT:    srli a2, a2, 8
+; RV32I-NEXT:    and a2, a2, a4
 ; RV32I-NEXT:    slli a0, a0, 24
-; RV32I-NEXT:    slli a2, a2, 8
-; RV32I-NEXT:    or a0, a0, a2
+; RV32I-NEXT:    slli a3, a3, 8
+; RV32I-NEXT:    or a0, a0, a3
+; RV32I-NEXT:    srli a1, a1, 24
+; RV32I-NEXT:    or a1, a2, a1
 ; RV32I-NEXT:    or a0, a0, a1
 ; RV32I-NEXT:    srli a1, a0, 4
 ; RV32I-NEXT:    and a0, a0, s4
@@ -3169,13 +3169,13 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
 ; RV32I-NEXT:    slli a0, a0, 4
 ; RV32I-NEXT:    or a0, a1, a0
 ; RV32I-NEXT:    srli a1, a0, 2
-; RV32I-NEXT:    and a0, a0, t6
-; RV32I-NEXT:    and a1, a1, t6
+; RV32I-NEXT:    and a0, a0, t5
+; RV32I-NEXT:    and a1, a1, t5
 ; RV32I-NEXT:    slli a0, a0, 2
 ; RV32I-NEXT:    or a0, a1, a0
 ; RV32I-NEXT:    lui a1, 349525
 ; RV32I-NEXT:    addi a1, a1, 1364
-; RV32I-NEXT:    and a2, a0, s2
+; RV32I-NEXT:    and a2, a0, t6
 ; RV32I-NEXT:    srli a0, a0, 1
 ; RV32I-NEXT:    and a0, a0, a1
 ; RV32I-NEXT:    slli a2, a2, 1
@@ -3870,16 +3870,16 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
 ; RV32IM-NEXT:    mv s4, a2
 ; RV32IM-NEXT:    mv s7, a1
 ; RV32IM-NEXT:    mv a1, a0
-; RV32IM-NEXT:    srli t0, a0, 8
+; RV32IM-NEXT:    srli t4, a0, 8
 ; RV32IM-NEXT:    srli t1, a0, 24
 ; RV32IM-NEXT:    srli t3, a2, 8
 ; RV32IM-NEXT:    srli t2, a2, 24
 ; RV32IM-NEXT:    slli a4, s7, 1
 ; RV32IM-NEXT:    andi a5, a2, 2
-; RV32IM-NEXT:    slli a7, s7, 2
+; RV32IM-NEXT:    slli a6, s7, 2
 ; RV32IM-NEXT:    andi s0, a2, 4
-; RV32IM-NEXT:    slli a6, s7, 3
-; RV32IM-NEXT:    andi t4, a2, 8
+; RV32IM-NEXT:    slli t5, s7, 3
+; RV32IM-NEXT:    andi a7, a2, 8
 ; RV32IM-NEXT:    slli t6, a0, 1
 ; RV32IM-NEXT:    andi s1, a3, 2
 ; RV32IM-NEXT:    lui a0, 16
@@ -3889,44 +3889,44 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
 ; RV32IM-NEXT:    lui a0, 32
 ; RV32IM-NEXT:    and a2, a2, a0
 ; RV32IM-NEXT:    lui s8, 2048
-; RV32IM-NEXT:    and s2, s4, s8
+; RV32IM-NEXT:    and t0, s4, s8
 ; RV32IM-NEXT:    lui s9, 4096
-; RV32IM-NEXT:    and t5, s4, s9
+; RV32IM-NEXT:    and s2, s4, s9
 ; RV32IM-NEXT:    and s3, a3, s3
 ; RV32IM-NEXT:    and a0, a3, a0
 ; RV32IM-NEXT:    sw s5, 312(sp) # 4-byte Folded Spill
-; RV32IM-NEXT:    and t0, t0, s5
+; RV32IM-NEXT:    and t4, t4, s5
 ; RV32IM-NEXT:    and t3, t3, s5
 ; RV32IM-NEXT:    mul s5, s7, s6
-; RV32IM-NEXT:    or t0, t0, t1
-; RV32IM-NEXT:    sw t0, 308(sp) # 4-byte Folded Spill
-; RV32IM-NEXT:    mul t0, s7, a2
-; RV32IM-NEXT:    or t1, t3, t2
-; RV32IM-NEXT:    sw t1, 304(sp) # 4-byte Folded Spill
+; RV32IM-NEXT:    or t1, t4, t1
+; RV32IM-NEXT:    sw t1, 308(sp) # 4-byte Folded Spill
+; RV32IM-NEXT:    mul t1, s7, a2
+; RV32IM-NEXT:    or t2, t3, t2
+; RV32IM-NEXT:    sw t2, 304(sp) # 4-byte Folded Spill
+; RV32IM-NEXT:    mul t2, s7, t0
+; RV32IM-NEXT:    xor t1, s5, t1
+; RV32IM-NEXT:    sw t1, 300(sp) # 4-byte Folded Spill
 ; RV32IM-NEXT:    mul t1, s7, s2
-; RV32IM-NEXT:    xor t0, s5, t0
-; RV32IM-NEXT:    sw t0, 300(sp) # 4-byte Folded Spill
-; RV32IM-NEXT:    mul t0, s7, t5
-; RV32IM-NEXT:    xor t0, t1, t0
-; RV32IM-NEXT:    sw t0, 296(s...
[truncated]

``````````

</details>


https://github.com/llvm/llvm-project/pull/193945


More information about the llvm-commits mailing list