[llvm] [DAG] Re-add transitive users when FREEZE folds away (PR #193945)
via llvm-commits
llvm-commits at lists.llvm.org
Fri Apr 24 04:32:48 PDT 2026
llvmbot wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-llvm-selectiondag
Author: Paweł Bylica (chfast)
<details>
<summary>Changes</summary>
When DAGCombiner::visitFREEZE folds a FREEZE of a guaranteed-non-poison value away, it returns the underlying value and relies on the combine framework to re-add the FREEZE's direct users. That is enough when the chain cascades: each re-visited user either simplifies (triggering ReplaceAllUsesWith on its own, which re-adds _its_ users) or is the consumer that cares.
It is not enough when an intermediate node in the chain is "transparent" to a peephole combine that lives on its users. The driving case is X86's checkBoolTestSetCCCombine, which is invoked from combineCMov, combineBrCond and combineSetCC and looks through
CMOV/BRCOND/SETCC(E, CMP(AND(FREEZE(X86ISD::SETCC(COND, EFLAGS)), 1), 0))
Short-circuit AND lowering turns `select a, b, false` into `select a, freeze(b), 0` (foldBoolSelectToLogic), and a multi-word usub-with-carry chain ends up in that exact shape. On the first CMOV visit the combine stops at the FREEZE. visitFREEZE later folds the FREEZE (X86ISD::SETCC is marked unconditionally non-poison), AND simplifies to SETCC, and CMP's operand is updated — but CMP itself does not change identity, so CMP's users (the CMOV) are never re-added. The redundant setb+testb+cmove triple after the final SBB survives to codegen.
Fix: when visitFREEZE returns N0, walk the use chain of the freeze for up to 32 nodes and re-add them to the combine worklist. Worklist deduplication makes the cost on nodes that have nothing new to do negligible.
---
Patch is 2.01 MiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/193945.diff
6 Files Affected:
- (modified) llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp (+17-1)
- (modified) llvm/test/CodeGen/RISCV/clmul.ll (+5237-5251)
- (modified) llvm/test/CodeGen/RISCV/clmulh.ll (+7553-7567)
- (modified) llvm/test/CodeGen/RISCV/clmulr.ll (+7274-7318)
- (added) llvm/test/CodeGen/X86/cmp-setcc-cmov-reworklist.ll (+75)
- (modified) llvm/test/CodeGen/X86/freeze-binary.ll (+7-13)
``````````diff
diff --git a/llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp b/llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
index 20ff48789152b..2c5a454364db5 100644
--- a/llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
@@ -17641,8 +17641,24 @@ SDValue DAGCombiner::visitBUILD_PAIR(SDNode *N) {
SDValue DAGCombiner::visitFREEZE(SDNode *N) {
SDValue N0 = N->getOperand(0);
- if (DAG.isGuaranteedNotToBeUndefOrPoison(N0, /*PoisonOnly*/ false))
+ if (DAG.isGuaranteedNotToBeUndefOrPoison(N0, /*PoisonOnly*/ false)) {
+ // Eliminating a freeze can unlock combines on nodes farther up the use
+ // chain — e.g., peephole combines that peek through intermediate nodes
+ // (CMP/AND/ZEXT) to reach a SETCC under the freeze. Normal worklist
+ // propagation only re-adds the direct users of the replaced node, which
+ // is not enough when those intermediate nodes do not themselves
+ // simplify. Re-add a bounded transitive user set (the set is for BFS
+ // dedup, separate from the worklist's own dedup).
+ // TODO: Not needed under -combiner-topological-sorting, which already
+ // visits operands before users. Remove when topo order is the default.
+ constexpr unsigned Limit = 32;
+ SmallSetVector<SDNode *, Limit> Users(N->users().begin(), N->users().end());
+ for (unsigned I = 0; I != Users.size() && Users.size() < Limit; ++I)
+ Users.insert_range(Users[I]->users());
+ for (SDNode *U : Users)
+ AddToWorklist(U);
return N0;
+ }
// If we have frozen and unfrozen users of N0, update so everything uses N.
if (!N0.isUndef() && !N0.hasOneUse()) {
diff --git a/llvm/test/CodeGen/RISCV/clmul.ll b/llvm/test/CodeGen/RISCV/clmul.ll
index d9bbe20b8730d..5545fff6c9e60 100644
--- a/llvm/test/CodeGen/RISCV/clmul.ll
+++ b/llvm/test/CodeGen/RISCV/clmul.ll
@@ -2749,7 +2749,7 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
; RV32I-NEXT: xor s1, a0, s1
; RV32I-NEXT: lw a0, 60(sp) # 4-byte Folded Reload
; RV32I-NEXT: xor ra, a0, ra
-; RV32I-NEXT: xor s8, s8, s6
+; RV32I-NEXT: xor s6, s8, s6
; RV32I-NEXT: lw a0, 200(sp) # 4-byte Folded Reload
; RV32I-NEXT: xor t6, a2, a0
; RV32I-NEXT: lw a0, 288(sp) # 4-byte Folded Reload
@@ -2760,10 +2760,10 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
; RV32I-NEXT: xor a5, a5, a0
; RV32I-NEXT: slli a2, a3, 11
; RV32I-NEXT: slli s2, s2, 11
-; RV32I-NEXT: li s6, 1
-; RV32I-NEXT: slli s6, s6, 11
-; RV32I-NEXT: and a0, t3, s6
-; RV32I-NEXT: and a1, t0, s6
+; RV32I-NEXT: li s8, 1
+; RV32I-NEXT: slli s8, s8, 11
+; RV32I-NEXT: and a0, t3, s8
+; RV32I-NEXT: and a1, t0, s8
; RV32I-NEXT: seqz a0, a0
; RV32I-NEXT: seqz a1, a1
; RV32I-NEXT: addi a0, a0, -1
@@ -2799,7 +2799,7 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
; RV32I-NEXT: lw s0, 72(sp) # 4-byte Folded Reload
; RV32I-NEXT: xor s0, ra, s0
; RV32I-NEXT: lw s1, 12(sp) # 4-byte Folded Reload
-; RV32I-NEXT: xor s8, s8, s1
+; RV32I-NEXT: xor s6, s6, s1
; RV32I-NEXT: xor a4, t6, a4
; RV32I-NEXT: lw t6, 272(sp) # 4-byte Folded Reload
; RV32I-NEXT: xor a5, a5, t6
@@ -2828,7 +2828,7 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
; RV32I-NEXT: lw t4, 80(sp) # 4-byte Folded Reload
; RV32I-NEXT: xor t4, s0, t4
; RV32I-NEXT: lw t5, 20(sp) # 4-byte Folded Reload
-; RV32I-NEXT: xor t5, s8, t5
+; RV32I-NEXT: xor t5, s6, t5
; RV32I-NEXT: xor a4, a4, a5
; RV32I-NEXT: lw a5, 252(sp) # 4-byte Folded Reload
; RV32I-NEXT: xor a5, a0, a5
@@ -2897,271 +2897,271 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
; RV32I-NEXT: xor a0, a5, a0
; RV32I-NEXT: sw a0, 304(sp) # 4-byte Folded Spill
; RV32I-NEXT: lui a1, 209715
-; RV32I-NEXT: addi t6, a1, 819
+; RV32I-NEXT: addi t5, a1, 819
; RV32I-NEXT: srli a3, s1, 2
-; RV32I-NEXT: and s1, s1, t6
-; RV32I-NEXT: and a3, a3, t6
+; RV32I-NEXT: and s1, s1, t5
+; RV32I-NEXT: and a3, a3, t5
; RV32I-NEXT: slli s1, s1, 2
; RV32I-NEXT: or a3, a3, s1
; RV32I-NEXT: srli a4, s2, 2
-; RV32I-NEXT: and a5, s2, t6
-; RV32I-NEXT: and a4, a4, t6
+; RV32I-NEXT: and a5, s2, t5
+; RV32I-NEXT: and a4, a4, t5
; RV32I-NEXT: slli a5, a5, 2
; RV32I-NEXT: or a4, a4, a5
; RV32I-NEXT: xor a2, a2, a6
; RV32I-NEXT: xor a5, a7, t0
; RV32I-NEXT: xor a2, a5, a2
; RV32I-NEXT: sw a2, 300(sp) # 4-byte Folded Spill
-; RV32I-NEXT: lui s2, 349525
-; RV32I-NEXT: addi s2, s2, 1365
+; RV32I-NEXT: lui t6, 349525
+; RV32I-NEXT: addi t6, t6, 1365
; RV32I-NEXT: srli a2, a3, 1
-; RV32I-NEXT: and a3, a3, s2
-; RV32I-NEXT: and a2, a2, s2
+; RV32I-NEXT: and a3, a3, t6
+; RV32I-NEXT: and a2, a2, t6
; RV32I-NEXT: slli a3, a3, 1
-; RV32I-NEXT: or t5, a2, a3
+; RV32I-NEXT: or t4, a2, a3
; RV32I-NEXT: srli a2, a4, 1
-; RV32I-NEXT: and a3, a4, s2
-; RV32I-NEXT: and a2, a2, s2
+; RV32I-NEXT: and a3, a4, t6
+; RV32I-NEXT: and a2, a2, t6
; RV32I-NEXT: slli a3, a3, 1
; RV32I-NEXT: or t2, a2, a3
; RV32I-NEXT: srli a3, a3, 31
; RV32I-NEXT: seqz a2, a3
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 31
+; RV32I-NEXT: slli a3, t4, 31
; RV32I-NEXT: and a2, a2, a3
; RV32I-NEXT: sw a2, 296(sp) # 4-byte Folded Spill
; RV32I-NEXT: andi a2, t2, 2
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 1
+; RV32I-NEXT: slli a3, t4, 1
; RV32I-NEXT: and a2, a2, a3
; RV32I-NEXT: sw a2, 292(sp) # 4-byte Folded Spill
; RV32I-NEXT: andi a2, t2, 4
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 2
+; RV32I-NEXT: slli a3, t4, 2
; RV32I-NEXT: and a2, a2, a3
; RV32I-NEXT: sw a2, 288(sp) # 4-byte Folded Spill
; RV32I-NEXT: andi a2, t2, 8
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 3
+; RV32I-NEXT: slli a3, t4, 3
; RV32I-NEXT: and a2, a2, a3
; RV32I-NEXT: sw a2, 284(sp) # 4-byte Folded Spill
; RV32I-NEXT: andi a2, t2, 16
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 4
-; RV32I-NEXT: and a2, a2, a3
-; RV32I-NEXT: sw a2, 280(sp) # 4-byte Folded Spill
+; RV32I-NEXT: slli s0, t4, 4
+; RV32I-NEXT: and a2, a2, s0
+; RV32I-NEXT: sw a2, 276(sp) # 4-byte Folded Spill
; RV32I-NEXT: andi a2, t2, 32
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli s1, t5, 5
-; RV32I-NEXT: and a2, a2, s1
-; RV32I-NEXT: sw a2, 272(sp) # 4-byte Folded Spill
+; RV32I-NEXT: slli a3, t4, 5
+; RV32I-NEXT: and a2, a2, a3
+; RV32I-NEXT: sw a2, 268(sp) # 4-byte Folded Spill
; RV32I-NEXT: andi a2, t2, 64
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli s0, t5, 6
-; RV32I-NEXT: and a2, a2, s0
-; RV32I-NEXT: sw a2, 276(sp) # 4-byte Folded Spill
+; RV32I-NEXT: slli s1, t4, 6
+; RV32I-NEXT: and a2, a2, s1
+; RV32I-NEXT: sw a2, 272(sp) # 4-byte Folded Spill
; RV32I-NEXT: andi a2, t2, 128
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 7
+; RV32I-NEXT: slli a3, t4, 7
; RV32I-NEXT: and a2, a2, a3
-; RV32I-NEXT: sw a2, 268(sp) # 4-byte Folded Spill
+; RV32I-NEXT: sw a2, 280(sp) # 4-byte Folded Spill
; RV32I-NEXT: andi a2, t2, 256
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 8
-; RV32I-NEXT: and s5, a2, a3
+; RV32I-NEXT: slli a3, t4, 8
+; RV32I-NEXT: and a2, a2, a3
+; RV32I-NEXT: sw a2, 264(sp) # 4-byte Folded Spill
; RV32I-NEXT: andi a2, t2, 512
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 9
-; RV32I-NEXT: and a2, a2, a3
-; RV32I-NEXT: sw a2, 260(sp) # 4-byte Folded Spill
+; RV32I-NEXT: slli a3, t4, 9
+; RV32I-NEXT: and s6, a2, a3
; RV32I-NEXT: andi a2, t2, 1024
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 10
+; RV32I-NEXT: slli a3, t4, 10
; RV32I-NEXT: and a2, a2, a3
-; RV32I-NEXT: sw a2, 264(sp) # 4-byte Folded Spill
-; RV32I-NEXT: and a2, t2, s6
+; RV32I-NEXT: sw a2, 260(sp) # 4-byte Folded Spill
+; RV32I-NEXT: and a2, t2, s8
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 11
-; RV32I-NEXT: and s1, a2, a3
+; RV32I-NEXT: slli a3, t4, 11
+; RV32I-NEXT: and s3, a2, a3
; RV32I-NEXT: lui a0, 1
; RV32I-NEXT: and a2, t2, a0
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 12
-; RV32I-NEXT: and s9, a2, a3
+; RV32I-NEXT: slli a3, t4, 12
+; RV32I-NEXT: and s5, a2, a3
; RV32I-NEXT: lui a0, 2
; RV32I-NEXT: and a2, t2, a0
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 13
+; RV32I-NEXT: slli a3, t4, 13
; RV32I-NEXT: and s7, a2, a3
; RV32I-NEXT: lui a0, 4
; RV32I-NEXT: and a2, t2, a0
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 14
-; RV32I-NEXT: and s6, a2, a3
+; RV32I-NEXT: slli a3, t4, 14
+; RV32I-NEXT: and s11, a2, a3
; RV32I-NEXT: lui a0, 8
; RV32I-NEXT: and a2, t2, a0
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 15
+; RV32I-NEXT: slli a3, t4, 15
; RV32I-NEXT: and s8, a2, a3
; RV32I-NEXT: lui a0, 16
; RV32I-NEXT: and a2, t2, a0
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 16
-; RV32I-NEXT: and s10, a2, a3
+; RV32I-NEXT: slli a3, t4, 16
+; RV32I-NEXT: and s9, a2, a3
; RV32I-NEXT: lui a0, 32
; RV32I-NEXT: and a2, t2, a0
; RV32I-NEXT: seqz a2, a2
; RV32I-NEXT: addi a2, a2, -1
-; RV32I-NEXT: slli a3, t5, 17
-; RV32I-NEXT: and s0, a2, a3
+; RV32I-NEXT: slli a3, t4, 17
+; RV32I-NEXT: and s10, a2, a3
; RV32I-NEXT: lui a0, 64
-; RV32I-NEXT: and a3, t2, a0
-; RV32I-NEXT: seqz a3, a3
-; RV32I-NEXT: addi a3, a3, -1
-; RV32I-NEXT: slli a4, t5, 18
-; RV32I-NEXT: and ra, a3, a4
+; RV32I-NEXT: and a2, t2, a0
+; RV32I-NEXT: seqz a2, a2
+; RV32I-NEXT: addi a2, a2, -1
+; RV32I-NEXT: slli a3, t4, 18
+; RV32I-NEXT: and ra, a2, a3
; RV32I-NEXT: lui a0, 128
-; RV32I-NEXT: and a3, t2, a0
-; RV32I-NEXT: seqz a3, a3
-; RV32I-NEXT: addi a3, a3, -1
-; RV32I-NEXT: slli a4, t5, 19
-; RV32I-NEXT: and s11, a3, a4
+; RV32I-NEXT: and a2, t2, a0
+; RV32I-NEXT: seqz a2, a2
+; RV32I-NEXT: addi a2, a2, -1
+; RV32I-NEXT: slli a3, t4, 19
+; RV32I-NEXT: and s1, a2, a3
; RV32I-NEXT: lui a0, 256
+; RV32I-NEXT: and a2, t2, a0
+; RV32I-NEXT: seqz a2, a2
+; RV32I-NEXT: addi a2, a2, -1
+; RV32I-NEXT: slli a3, t4, 20
+; RV32I-NEXT: and s0, a2, a3
+; RV32I-NEXT: lui a0, 512
; RV32I-NEXT: and a3, t2, a0
; RV32I-NEXT: seqz a3, a3
; RV32I-NEXT: addi a3, a3, -1
-; RV32I-NEXT: slli a4, t5, 20
-; RV32I-NEXT: and a4, a3, a4
-; RV32I-NEXT: lui a0, 512
+; RV32I-NEXT: slli a5, t4, 21
+; RV32I-NEXT: and a5, a3, a5
+; RV32I-NEXT: lui a0, 1024
; RV32I-NEXT: and a3, t2, a0
; RV32I-NEXT: seqz a3, a3
-; RV32I-NEXT: addi a0, a3, -1
-; RV32I-NEXT: slli a3, t5, 21
-; RV32I-NEXT: and a3, a0, a3
-; RV32I-NEXT: lui a0, 1024
+; RV32I-NEXT: addi a3, a3, -1
+; RV32I-NEXT: slli a4, t4, 22
+; RV32I-NEXT: and a3, a3, a4
+; RV32I-NEXT: lui a0, 2048
+; RV32I-NEXT: and a4, t2, a0
+; RV32I-NEXT: seqz a4, a4
+; RV32I-NEXT: addi a0, a4, -1
+; RV32I-NEXT: slli a4, t4, 23
+; RV32I-NEXT: and a4, a0, a4
+; RV32I-NEXT: lui a0, 4096
; RV32I-NEXT: and a0, t2, a0
; RV32I-NEXT: seqz a0, a0
; RV32I-NEXT: addi a0, a0, -1
-; RV32I-NEXT: slli t0, t5, 22
+; RV32I-NEXT: slli t0, t4, 24
; RV32I-NEXT: and a2, a0, t0
-; RV32I-NEXT: lui a0, 2048
+; RV32I-NEXT: lui a0, 8192
; RV32I-NEXT: and t0, t2, a0
; RV32I-NEXT: seqz t0, t0
; RV32I-NEXT: addi t0, t0, -1
-; RV32I-NEXT: slli a1, t5, 23
+; RV32I-NEXT: slli a1, t4, 25
; RV32I-NEXT: and a1, t0, a1
-; RV32I-NEXT: lui a0, 4096
-; RV32I-NEXT: and t0, t2, a0
-; RV32I-NEXT: seqz t0, t0
-; RV32I-NEXT: addi t0, t0, -1
-; RV32I-NEXT: slli a5, t5, 24
-; RV32I-NEXT: and a5, t0, a5
-; RV32I-NEXT: lui a0, 8192
+; RV32I-NEXT: lui a0, 16384
; RV32I-NEXT: and t0, t2, a0
; RV32I-NEXT: seqz t0, t0
; RV32I-NEXT: addi t0, t0, -1
-; RV32I-NEXT: slli a6, t5, 25
+; RV32I-NEXT: slli a6, t4, 26
; RV32I-NEXT: and a6, t0, a6
-; RV32I-NEXT: lui a0, 16384
+; RV32I-NEXT: lui a0, 32768
; RV32I-NEXT: and t0, t2, a0
; RV32I-NEXT: seqz t0, t0
; RV32I-NEXT: addi t0, t0, -1
-; RV32I-NEXT: slli a7, t5, 26
+; RV32I-NEXT: slli a7, t4, 27
; RV32I-NEXT: and a7, t0, a7
-; RV32I-NEXT: lui a0, 32768
+; RV32I-NEXT: lui a0, 65536
; RV32I-NEXT: and t0, t2, a0
; RV32I-NEXT: seqz t0, t0
; RV32I-NEXT: addi t0, t0, -1
-; RV32I-NEXT: slli t1, t5, 27
+; RV32I-NEXT: slli t1, t4, 28
; RV32I-NEXT: and t0, t0, t1
-; RV32I-NEXT: lui a0, 65536
+; RV32I-NEXT: lui a0, 131072
; RV32I-NEXT: and t1, t2, a0
; RV32I-NEXT: seqz t1, t1
; RV32I-NEXT: addi t1, t1, -1
-; RV32I-NEXT: slli t3, t5, 28
+; RV32I-NEXT: slli t3, t4, 29
; RV32I-NEXT: and t1, t1, t3
-; RV32I-NEXT: lui a0, 131072
-; RV32I-NEXT: and t3, t2, a0
-; RV32I-NEXT: seqz t3, t3
-; RV32I-NEXT: addi t3, t3, -1
-; RV32I-NEXT: slli t4, t5, 29
-; RV32I-NEXT: and t3, t3, t4
; RV32I-NEXT: lui a0, 262144
-; RV32I-NEXT: and t4, t2, a0
+; RV32I-NEXT: and t3, t2, a0
; RV32I-NEXT: andi t2, t2, 1
; RV32I-NEXT: seqz t2, t2
; RV32I-NEXT: addi t2, t2, -1
-; RV32I-NEXT: and t2, t2, t5
-; RV32I-NEXT: slli t5, t5, 30
-; RV32I-NEXT: seqz t4, t4
-; RV32I-NEXT: addi t4, t4, -1
-; RV32I-NEXT: and t4, t4, t5
+; RV32I-NEXT: and t2, t2, t4
+; RV32I-NEXT: slli t4, t4, 30
+; RV32I-NEXT: seqz t3, t3
+; RV32I-NEXT: addi t3, t3, -1
+; RV32I-NEXT: and t3, t3, t4
; RV32I-NEXT: lw a0, 292(sp) # 4-byte Folded Reload
; RV32I-NEXT: xor t2, t2, a0
; RV32I-NEXT: lw a0, 288(sp) # 4-byte Folded Reload
-; RV32I-NEXT: lw t5, 284(sp) # 4-byte Folded Reload
-; RV32I-NEXT: xor t5, a0, t5
-; RV32I-NEXT: lw a0, 280(sp) # 4-byte Folded Reload
-; RV32I-NEXT: lw s3, 272(sp) # 4-byte Folded Reload
-; RV32I-NEXT: xor a0, a0, s3
-; RV32I-NEXT: lw s3, 268(sp) # 4-byte Folded Reload
-; RV32I-NEXT: xor s5, s3, s5
-; RV32I-NEXT: xor s1, s1, s9
-; RV32I-NEXT: xor s0, s10, s0
+; RV32I-NEXT: lw t4, 284(sp) # 4-byte Folded Reload
+; RV32I-NEXT: xor t4, a0, t4
+; RV32I-NEXT: lw a0, 276(sp) # 4-byte Folded Reload
+; RV32I-NEXT: lw s2, 268(sp) # 4-byte Folded Reload
+; RV32I-NEXT: xor a0, a0, s2
+; RV32I-NEXT: lw s2, 264(sp) # 4-byte Folded Reload
+; RV32I-NEXT: xor s6, s2, s6
+; RV32I-NEXT: xor s7, s7, s11
+; RV32I-NEXT: xor s0, s1, s0
; RV32I-NEXT: xor a1, a2, a1
-; RV32I-NEXT: xor a2, t3, t4
-; RV32I-NEXT: xor t2, t2, t5
-; RV32I-NEXT: lw t3, 276(sp) # 4-byte Folded Reload
-; RV32I-NEXT: xor a0, a0, t3
-; RV32I-NEXT: lw t3, 260(sp) # 4-byte Folded Reload
-; RV32I-NEXT: xor t3, s5, t3
-; RV32I-NEXT: xor t4, s1, s7
-; RV32I-NEXT: xor t5, s0, ra
-; RV32I-NEXT: xor a1, a1, a5
-; RV32I-NEXT: lw a5, 296(sp) # 4-byte Folded Reload
-; RV32I-NEXT: xor a2, a2, a5
-; RV32I-NEXT: xor a0, t2, a0
-; RV32I-NEXT: lw a5, 264(sp) # 4-byte Folded Reload
-; RV32I-NEXT: xor a5, t3, a5
-; RV32I-NEXT: xor t2, t4, s6
-; RV32I-NEXT: xor t3, t5, s11
+; RV32I-NEXT: xor a2, t2, t4
+; RV32I-NEXT: lw t2, 272(sp) # 4-byte Folded Reload
+; RV32I-NEXT: xor a0, a0, t2
+; RV32I-NEXT: lw t2, 260(sp) # 4-byte Folded Reload
+; RV32I-NEXT: xor t2, s6, t2
+; RV32I-NEXT: xor t4, s7, s8
+; RV32I-NEXT: xor a5, s0, a5
; RV32I-NEXT: xor a1, a1, a6
-; RV32I-NEXT: xor a0, a0, a5
-; RV32I-NEXT: xor a5, t2, s8
-; RV32I-NEXT: xor a4, t3, a4
+; RV32I-NEXT: xor a0, a2, a0
+; RV32I-NEXT: xor a2, t2, s3
+; RV32I-NEXT: xor a6, t4, s9
+; RV32I-NEXT: xor a3, a5, a3
; RV32I-NEXT: xor a1, a1, a7
+; RV32I-NEXT: lw a5, 280(sp) # 4-byte Folded Reload
; RV32I-NEXT: xor a0, a0, a5
-; RV32I-NEXT: xor a3, a4, a3
+; RV32I-NEXT: xor a2, a2, s5
+; RV32I-NEXT: xor a5, a6, s10
+; RV32I-NEXT: xor a3, a3, a4
; RV32I-NEXT: xor a1, a1, t0
-; RV32I-NEXT: xor a0, a0, a3
+; RV32I-NEXT: xor a4, a5, ra
; RV32I-NEXT: xor a1, a1, t1
-; RV32I-NEXT: xor a0, a0, a1
-; RV32I-NEXT: xor a0, a0, a2
-; RV32I-NEXT: srli a1, a0, 8
-; RV32I-NEXT: srli a2, a0, 24
-; RV32I-NEXT: lw a3, 312(sp) # 4-byte Folded Reload
-; RV32I-NEXT: and a1, a1, a3
-; RV32I-NEXT: or a1, a1, a2
-; RV32I-NEXT: and a2, a0, a3
+; RV32I-NEXT: xor a2, a0, a2
+; RV32I-NEXT: xor a2, a2, a4
+; RV32I-NEXT: xor a1, a1, t3
+; RV32I-NEXT: xor a2, a2, a3
+; RV32I-NEXT: lw a3, 296(sp) # 4-byte Folded Reload
+; RV32I-NEXT: xor a1, a1, a3
+; RV32I-NEXT: lw a4, 312(sp) # 4-byte Folded Reload
+; RV32I-NEXT: and a3, a2, a4
+; RV32I-NEXT: xor a1, a2, a1
+; RV32I-NEXT: srli a2, a2, 8
+; RV32I-NEXT: and a2, a2, a4
; RV32I-NEXT: slli a0, a0, 24
-; RV32I-NEXT: slli a2, a2, 8
-; RV32I-NEXT: or a0, a0, a2
+; RV32I-NEXT: slli a3, a3, 8
+; RV32I-NEXT: or a0, a0, a3
+; RV32I-NEXT: srli a1, a1, 24
+; RV32I-NEXT: or a1, a2, a1
; RV32I-NEXT: or a0, a0, a1
; RV32I-NEXT: srli a1, a0, 4
; RV32I-NEXT: and a0, a0, s4
@@ -3169,13 +3169,13 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
; RV32I-NEXT: slli a0, a0, 4
; RV32I-NEXT: or a0, a1, a0
; RV32I-NEXT: srli a1, a0, 2
-; RV32I-NEXT: and a0, a0, t6
-; RV32I-NEXT: and a1, a1, t6
+; RV32I-NEXT: and a0, a0, t5
+; RV32I-NEXT: and a1, a1, t5
; RV32I-NEXT: slli a0, a0, 2
; RV32I-NEXT: or a0, a1, a0
; RV32I-NEXT: lui a1, 349525
; RV32I-NEXT: addi a1, a1, 1364
-; RV32I-NEXT: and a2, a0, s2
+; RV32I-NEXT: and a2, a0, t6
; RV32I-NEXT: srli a0, a0, 1
; RV32I-NEXT: and a0, a0, a1
; RV32I-NEXT: slli a2, a2, 1
@@ -3870,16 +3870,16 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
; RV32IM-NEXT: mv s4, a2
; RV32IM-NEXT: mv s7, a1
; RV32IM-NEXT: mv a1, a0
-; RV32IM-NEXT: srli t0, a0, 8
+; RV32IM-NEXT: srli t4, a0, 8
; RV32IM-NEXT: srli t1, a0, 24
; RV32IM-NEXT: srli t3, a2, 8
; RV32IM-NEXT: srli t2, a2, 24
; RV32IM-NEXT: slli a4, s7, 1
; RV32IM-NEXT: andi a5, a2, 2
-; RV32IM-NEXT: slli a7, s7, 2
+; RV32IM-NEXT: slli a6, s7, 2
; RV32IM-NEXT: andi s0, a2, 4
-; RV32IM-NEXT: slli a6, s7, 3
-; RV32IM-NEXT: andi t4, a2, 8
+; RV32IM-NEXT: slli t5, s7, 3
+; RV32IM-NEXT: andi a7, a2, 8
; RV32IM-NEXT: slli t6, a0, 1
; RV32IM-NEXT: andi s1, a3, 2
; RV32IM-NEXT: lui a0, 16
@@ -3889,44 +3889,44 @@ define i64 @clmul_i64(i64 %a, i64 %b) nounwind {
; RV32IM-NEXT: lui a0, 32
; RV32IM-NEXT: and a2, a2, a0
; RV32IM-NEXT: lui s8, 2048
-; RV32IM-NEXT: and s2, s4, s8
+; RV32IM-NEXT: and t0, s4, s8
; RV32IM-NEXT: lui s9, 4096
-; RV32IM-NEXT: and t5, s4, s9
+; RV32IM-NEXT: and s2, s4, s9
; RV32IM-NEXT: and s3, a3, s3
; RV32IM-NEXT: and a0, a3, a0
; RV32IM-NEXT: sw s5, 312(sp) # 4-byte Folded Spill
-; RV32IM-NEXT: and t0, t0, s5
+; RV32IM-NEXT: and t4, t4, s5
; RV32IM-NEXT: and t3, t3, s5
; RV32IM-NEXT: mul s5, s7, s6
-; RV32IM-NEXT: or t0, t0, t1
-; RV32IM-NEXT: sw t0, 308(sp) # 4-byte Folded Spill
-; RV32IM-NEXT: mul t0, s7, a2
-; RV32IM-NEXT: or t1, t3, t2
-; RV32IM-NEXT: sw t1, 304(sp) # 4-byte Folded Spill
+; RV32IM-NEXT: or t1, t4, t1
+; RV32IM-NEXT: sw t1, 308(sp) # 4-byte Folded Spill
+; RV32IM-NEXT: mul t1, s7, a2
+; RV32IM-NEXT: or t2, t3, t2
+; RV32IM-NEXT: sw t2, 304(sp) # 4-byte Folded Spill
+; RV32IM-NEXT: mul t2, s7, t0
+; RV32IM-NEXT: xor t1, s5, t1
+; RV32IM-NEXT: sw t1, 300(sp) # 4-byte Folded Spill
; RV32IM-NEXT: mul t1, s7, s2
-; RV32IM-NEXT: xor t0, s5, t0
-; RV32IM-NEXT: sw t0, 300(sp) # 4-byte Folded Spill
-; RV32IM-NEXT: mul t0, s7, t5
-; RV32IM-NEXT: xor t0, t1, t0
-; RV32IM-NEXT: sw t0, 296(s...
[truncated]
``````````
</details>
https://github.com/llvm/llvm-project/pull/193945
More information about the llvm-commits
mailing list