[llvm] [RISCV] Eliminate redundant materializations after register comparison (PR #227673)
via llvm-commits
llvm-commits at lists.llvm.org
Wed Sep 30 04:56:43 PDT 2026
llvmorg-github-actions[bot] wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-backend-risc-v
Author: tinfengyu
<details>
<summary>Changes</summary>
`RISCVRedundantCopyElimination` can eliminate redundant immediate materializations when a branch establishes a known immediate value, but it does not handle the equivalent case where a non-zero immediate is first materialized into a register and then used by `BEQ`/`BNE`.
For a register-register comparison, the equality edge can establish that one operand is equal to the materialized constant. If that successor materializes the same constant into that register again, the second materialization is redundant.
This patch extends the pass to recognize this case by finding an `ADDI` from `X0` or `QC_LI` that defines one of the branch operands before the terminator. On the equality edge of `BEQ`/`BNE`, the known immediate value is then used by the existing redundant-materialization elimination logic.
`X0` is explicitly excluded as a target register because writes to it are discarded and therefore cannot establish a non-zero value.
I also evaluated the change with a controlled microbenchmark on a SpacemiT X60 (RV64). With the equality edge taken 50% of the time, median measurements showed approximately:
* 3.39% fewer dynamic instructions
* 2.7% fewer cycles
* 2.7% lower wall-clock time
Correctness and final linked machine code were checked before measurement. Repeated baseline/candidate runs were interleaved to reduce run-order effects.
As expected for this transformation, the dynamic instruction reduction increased as the equality edge was executed more frequently. These measurements are from a focused microbenchmark and are not intended to represent an application-level performance improvement.
Tests cover the new register-register branch cases, RV32/RV64 code generation, the relevant Xqci and XAndes branch variants, `X0` handling, and cases where the materialization must not be removed.
AI tool usage: OpenAI Codex was used to assist with investigation and implementation. I reviewed the resulting implementation, tests, and validation results and take responsibility for the contribution.
Assisted-by: OpenAI Codex
---
Full diff: https://github.com/llvm/llvm-project/pull/227673.diff
5 Files Affected:
- (modified) llvm/lib/Target/RISCV/RISCVRedundantCopyElimination.cpp (+61-6)
- (added) llvm/test/CodeGen/RISCV/redundant-copy-elim-reg-imm.mir (+230)
- (modified) llvm/test/CodeGen/RISCV/redundant-copy-elim.ll (+17-6)
- (modified) llvm/test/CodeGen/RISCV/xandesperf-redundant-copy-elim.ll (+3-6)
- (modified) llvm/test/CodeGen/RISCV/xqcibi-redundant-copy-elim.ll (+6-10)
``````````diff
diff --git a/llvm/lib/Target/RISCV/RISCVRedundantCopyElimination.cpp b/llvm/lib/Target/RISCV/RISCVRedundantCopyElimination.cpp
index 2e0e6ab2408d7..651044d6e49ef 100644
--- a/llvm/lib/Target/RISCV/RISCVRedundantCopyElimination.cpp
+++ b/llvm/lib/Target/RISCV/RISCVRedundantCopyElimination.cpp
@@ -20,8 +20,8 @@
// This pass should be run after register allocation and is based on the
// earliest versions of AArch64RedundantCopyElimination.
//
-// FIXME: Support compare with non-zero immediates where the immediate is stored
-// in a register.
+// The pass also handles register-register branches when one operand is
+// materialized as a non-zero immediate in the predecessor block.
//
//===----------------------------------------------------------------------===//
@@ -32,6 +32,7 @@
#include "llvm/CodeGen/MachineRegisterInfo.h"
#include "llvm/CodeGen/RegisterClassInfo.h"
#include "llvm/Support/Debug.h"
+#include <optional>
using namespace llvm;
@@ -108,6 +109,32 @@ guaranteesRegEqualsImmInBlock(MachineBasicBlock &MBB,
return false;
}
+static std::optional<int64_t>
+getRegImmediateBeforeTerminator(MachineBasicBlock &MBB, Register Reg,
+ const TargetRegisterInfo *TRI) {
+ // A write to X0 is discarded, so it cannot establish a nonzero value.
+ if (Reg == RISCV::X0)
+ return std::nullopt;
+
+ for (auto I = MBB.getFirstTerminator(); I != MBB.begin();) {
+ MachineInstr &MI = *--I;
+ if (!MI.modifiesRegister(Reg, TRI))
+ continue;
+ // The last modification must define Reg itself to a known immediate.
+ if (MI.getOpcode() == RISCV::ADDI && MI.getOperand(0).isReg() &&
+ MI.getOperand(0).getReg() == Reg &&
+ MI.getOperand(1).isReg() && MI.getOperand(1).getReg() == RISCV::X0 &&
+ MI.getOperand(2).isImm())
+ return MI.getOperand(2).getImm();
+ if (MI.getOpcode() == RISCV::QC_LI && MI.getOperand(0).isReg() &&
+ MI.getOperand(0).getReg() == Reg &&
+ MI.getOperand(1).isImm())
+ return MI.getOperand(1).getImm();
+ return std::nullopt;
+ }
+ return std::nullopt;
+}
+
bool RISCVRedundantCopyElimination::optimizeBlock(MachineBasicBlock &MBB) {
// Check if the current basic block has a single predecessor.
if (MBB.pred_size() != 1)
@@ -131,8 +158,34 @@ bool RISCVRedundantCopyElimination::optimizeBlock(MachineBasicBlock &MBB) {
return false;
bool IsZeroCopy = guaranteesZeroRegInBlock(MBB, Cond, TBB);
+ bool IsImmCopy = guaranteesRegEqualsImmInBlock(MBB, Cond, TBB);
+ int64_t CompareImm = IsImmCopy ? Cond[2].getImm() : 0;
+ if (!IsZeroCopy && !IsImmCopy && Cond.size() == 3 &&
+ (Cond[0].getImm() == RISCV::BEQ || Cond[0].getImm() == RISCV::BNE) &&
+ Cond[2].isReg()) {
+ // One branch operand may have been materialized with ADDI or QC_LI.
+ // The other operand is known to have the same value on the equality edge,
+ // irrespective of the operand order.
+ std::optional<int64_t> Imm =
+ getRegImmediateBeforeTerminator(*PredMBB, Cond[2].getReg(), TRI);
+ if (Imm && *Imm != 0) {
+ TargetReg = Cond[1].getReg();
+ CompareImm = *Imm;
+ IsImmCopy = true;
+ } else {
+ Imm = getRegImmediateBeforeTerminator(*PredMBB, Cond[1].getReg(), TRI);
+ if (Imm && *Imm != 0) {
+ TargetReg = Cond[2].getReg();
+ CompareImm = *Imm;
+ IsImmCopy = true;
+ }
+ }
+ // For BEQ, equality is guaranteed on the taken edge. For BNE, it is
+ // guaranteed on the fallthrough edge.
+ IsImmCopy &= (Cond[0].getImm() == RISCV::BEQ) == (TBB == &MBB);
+ }
- if (!IsZeroCopy && !guaranteesRegEqualsImmInBlock(MBB, Cond, TBB))
+ if (!IsZeroCopy && !IsImmCopy)
return false;
bool Changed = false;
@@ -161,14 +214,14 @@ bool RISCVRedundantCopyElimination::optimizeBlock(MachineBasicBlock &MBB) {
Register SrcReg = MI->getOperand(1).getReg();
int64_t Imm = MI->getOperand(2).getImm();
if (SrcReg == RISCV::X0 && !MRI->isReserved(DefReg) &&
- TargetReg == DefReg && Imm == Cond[2].getImm())
+ TargetReg == DefReg && Imm == CompareImm)
RemoveMI = true;
} else if (MI->getOpcode() == RISCV::QC_LI && MI->getOperand(0).isReg() &&
MI->getOperand(1).isImm()) {
Register DefReg = MI->getOperand(0).getReg();
int64_t Imm = MI->getOperand(1).getImm();
if (!MRI->isReserved(DefReg) && TargetReg == DefReg &&
- Imm == Cond[2].getImm())
+ Imm == CompareImm)
RemoveMI = true;
}
}
@@ -203,7 +256,9 @@ bool RISCVRedundantCopyElimination::optimizeBlock(MachineBasicBlock &MBB) {
CondBr->getOpcode() == RISCV::NDS_BEQC ||
CondBr->getOpcode() == RISCV::NDS_BNEC) &&
"Unexpected opcode");
- assert(CondBr->getOperand(0).getReg() == TargetReg && "Unexpected register");
+ assert((CondBr->getOperand(0).getReg() == TargetReg ||
+ CondBr->getOperand(1).getReg() == TargetReg) &&
+ "Unexpected register");
// Otherwise, we have to fixup the use-def chain, starting with the
// BEQ(I)/BNE(I). Conservatively mark as much as we can live.
diff --git a/llvm/test/CodeGen/RISCV/redundant-copy-elim-reg-imm.mir b/llvm/test/CodeGen/RISCV/redundant-copy-elim-reg-imm.mir
new file mode 100644
index 0000000000000..2a6c23bee323e
--- /dev/null
+++ b/llvm/test/CodeGen/RISCV/redundant-copy-elim-reg-imm.mir
@@ -0,0 +1,230 @@
+# RUN: llc -mtriple=riscv32 -mattr=+xqcili -run-pass=riscv-copyelim -verify-machineinstrs %s -o - | FileCheck %s
+
+--- |
+ define void @beq_reg_imm() { ret void }
+ define void @bne_reg_imm() { ret void }
+ define void @clobber_const() { ret void }
+ define void @clobber_target() { ret void }
+ define void @inequality_edge() { ret void }
+ define void @beq_x0_is_zero() { ret void }
+ define void @bne_x0_is_zero() { ret void }
+ define void @beq_reg_imm_swapped() { ret void }
+ define void @bne_reg_imm_swapped() { ret void }
+ declare void @callee()
+...
+---
+name: beq_reg_imm
+tracksRegLiveness: true
+body: |
+ bb.0:
+ successors: %bb.1, %bb.2
+ liveins: $x10
+ $x11 = ADDI $x0, 7
+ BEQ $x11, killed $x10, %bb.1
+ PseudoBR %bb.2
+ bb.1:
+ successors: %bb.2
+ $x10 = ADDI $x0, 7
+ PseudoRET implicit $x10
+ bb.2:
+ PseudoRET
+...
+
+# CHECK-LABEL: name: beq_reg_imm
+# CHECK: $x11 = ADDI $x0, 7
+# CHECK: BEQ $x11, $x10, %bb.1
+# CHECK-LABEL: bb.1:
+# CHECK-NOT: $x10 = ADDI $x0, 7
+
+---
+name: bne_reg_imm
+tracksRegLiveness: true
+body: |
+ bb.0:
+ successors: %bb.2, %bb.1
+ liveins: $x10
+ $x11 = QC_LI 9
+ BNE killed $x10, $x11, %bb.2
+ PseudoBR %bb.1
+ bb.1:
+ $x10 = ADDI $x0, 9
+ PseudoRET implicit $x10
+ bb.2:
+ PseudoRET
+...
+
+# CHECK-LABEL: name: bne_reg_imm
+# CHECK: $x11 = QC_LI 9
+# CHECK: BNE $x10, $x11, %bb.2
+# CHECK-LABEL: bb.1:
+# CHECK-NOT: $x10 = ADDI $x0, 9
+
+---
+name: clobber_const
+tracksRegLiveness: true
+body: |
+ bb.0:
+ successors: %bb.1, %bb.2
+ liveins: $x8
+ $x11 = ADDI $x0, 7
+ PseudoCALL target-flags(riscv-call) @callee, csr_ilp32_lp64, implicit-def dead $x1, implicit-def $x11
+ BEQ $x8, $x11, %bb.1
+ PseudoBR %bb.2
+ bb.1:
+ $x8 = ADDI $x0, 7
+ PseudoRET implicit $x8
+ bb.2:
+ liveins: $x8
+ PseudoRET implicit $x8
+...
+
+# CHECK-LABEL: name: clobber_const
+# CHECK: PseudoCALL {{.*}}csr_ilp32_lp64
+# CHECK: BEQ $x8, $x11, %bb.1
+# CHECK-LABEL: bb.1:
+# CHECK: $x8 = ADDI $x0, 7
+
+---
+name: clobber_target
+tracksRegLiveness: true
+body: |
+ bb.0:
+ successors: %bb.1, %bb.2
+ liveins: $x10
+ $x11 = ADDI $x0, 7
+ BEQ $x10, $x11, %bb.1
+ PseudoBR %bb.2
+ bb.1:
+ liveins: $x10
+ $x10 = ADDI $x10, 1
+ $x10 = ADDI $x0, 7
+ PseudoRET implicit $x10
+ bb.2:
+ liveins: $x10
+ PseudoRET implicit $x10
+...
+
+# CHECK-LABEL: name: clobber_target
+# CHECK-LABEL: bb.1:
+# CHECK: $x10 = ADDI $x10, 1
+# CHECK: $x10 = ADDI $x0, 7
+
+---
+name: inequality_edge
+tracksRegLiveness: true
+body: |
+ bb.0:
+ successors: %bb.2, %bb.1
+ liveins: $x10
+ $x11 = ADDI $x0, 7
+ BEQ $x10, $x11, %bb.2
+ PseudoBR %bb.1
+ bb.1:
+ $x10 = ADDI $x0, 7
+ PseudoRET implicit $x10
+ bb.2:
+ liveins: $x10
+ PseudoRET implicit $x10
+...
+
+# CHECK-LABEL: name: inequality_edge
+# CHECK-LABEL: bb.1:
+# CHECK: $x10 = ADDI $x0, 7
+
+---
+name: beq_x0_is_zero
+tracksRegLiveness: true
+body: |
+ bb.0:
+ successors: %bb.1, %bb.2
+ liveins: $x10
+ $x0 = ADDI $x0, 7
+ BEQ $x0, $x10, %bb.1
+ PseudoBR %bb.2
+ bb.1:
+ $x10 = ADDI $x0, 7
+ PseudoRET implicit $x10
+ bb.2:
+ PseudoRET
+...
+
+# The formal ADDI definition cannot change x0's architectural zero value.
+# CHECK-LABEL: name: beq_x0_is_zero
+# CHECK: $x0 = ADDI $x0, 7
+# CHECK: BEQ $x0, $x10, %bb.1
+# CHECK-LABEL: bb.1:
+# CHECK: $x10 = ADDI $x0, 7
+
+---
+name: bne_x0_is_zero
+tracksRegLiveness: true
+body: |
+ bb.0:
+ successors: %bb.2, %bb.1
+ liveins: $x10
+ $x0 = ADDI $x0, 9
+ BNE $x0, $x10, %bb.2
+ PseudoBR %bb.1
+ bb.1:
+ $x10 = ADDI $x0, 9
+ PseudoRET implicit $x10
+ bb.2:
+ PseudoRET
+...
+
+# CHECK-LABEL: name: bne_x0_is_zero
+# CHECK: $x0 = ADDI $x0, 9
+# CHECK: BNE $x0, $x10, %bb.2
+# CHECK-LABEL: bb.1:
+# CHECK: $x10 = ADDI $x0, 9
+
+---
+name: beq_reg_imm_swapped
+tracksRegLiveness: true
+body: |
+ bb.0:
+ successors: %bb.1, %bb.2
+ liveins: $x10
+ $x11 = ADDI $x0, 7
+ BEQ killed $x10, $x11, %bb.1
+ PseudoBR %bb.2
+ bb.1:
+ liveins: $x10
+ $x12 = ADDI killed $x10, 1
+ $x10 = ADDI $x0, 7
+ PseudoRET implicit $x10, implicit $x12
+ bb.2:
+ PseudoRET
+...
+
+# The compared value is the first branch operand. Its kill in the successor
+# must be cleared when the later materialization is removed.
+# CHECK-LABEL: name: beq_reg_imm_swapped
+# CHECK: BEQ $x10, $x11, %bb.1
+# CHECK-LABEL: bb.1:
+# CHECK: $x12 = ADDI $x10, 1
+# CHECK-NOT: $x10 = ADDI $x0, 7
+# CHECK: PseudoRET implicit $x10, implicit $x12
+
+---
+name: bne_reg_imm_swapped
+tracksRegLiveness: true
+body: |
+ bb.0:
+ successors: %bb.2, %bb.1
+ liveins: $x10
+ $x11 = ADDI $x0, 9
+ BNE $x11, killed $x10, %bb.2
+ PseudoBR %bb.1
+ bb.1:
+ $x10 = ADDI $x0, 9
+ PseudoRET implicit $x10
+ bb.2:
+ PseudoRET
+...
+
+# CHECK-LABEL: name: bne_reg_imm_swapped
+# CHECK: BNE $x11, $x10, %bb.2
+# CHECK-LABEL: bb.1:
+# CHECK-NOT: $x10 = ADDI $x0, 9
+# CHECK: PseudoRET implicit $x10
diff --git a/llvm/test/CodeGen/RISCV/redundant-copy-elim.ll b/llvm/test/CodeGen/RISCV/redundant-copy-elim.ll
index a93514dfaedc6..828f8ebc3a309 100644
--- a/llvm/test/CodeGen/RISCV/redundant-copy-elim.ll
+++ b/llvm/test/CodeGen/RISCV/redundant-copy-elim.ll
@@ -3,17 +3,17 @@
; RUN: | FileCheck %s -check-prefixes=RV32I
; RUN: llc -mtriple=riscv32 -mattr=+experimental-zibi -verify-machineinstrs < %s \
; RUN: | FileCheck %s -check-prefixes=RV32IZIBI
+; RUN: llc -mtriple=riscv64 -verify-machineinstrs < %s \
+; RUN: | FileCheck %s -check-prefixes=RV64I
define i32 @test_beqi(i32 %a) nounwind {
; RV32I-LABEL: test_beqi:
; RV32I: # %bb.0: # %entry
; RV32I-NEXT: li a1, 7
-; RV32I-NEXT: bne a0, a1, .LBB0_2
-; RV32I-NEXT: # %bb.1: # %if.end
-; RV32I-NEXT: li a0, 7
-; RV32I-NEXT: ret
-; RV32I-NEXT: .LBB0_2: # %if.then
+; RV32I-NEXT: beq a0, a1, .LBB0_2
+; RV32I-NEXT: # %bb.1: # %if.then
; RV32I-NEXT: li a0, 1
+; RV32I-NEXT: .LBB0_2: # %if.end
; RV32I-NEXT: ret
;
; RV32IZIBI-LABEL: test_beqi:
@@ -23,6 +23,12 @@ define i32 @test_beqi(i32 %a) nounwind {
; RV32IZIBI-NEXT: li a0, 1
; RV32IZIBI-NEXT: .LBB0_2: # %if.end
; RV32IZIBI-NEXT: ret
+; RV64I-LABEL: test_beqi:
+; RV64I: li a1, 7
+; RV64I-NEXT: sext.w a0, a0
+; RV64I-NEXT: beq a0, a1, .LBB0_2
+; RV64I-NOT: li a0, 7
+; RV64I: ret
entry:
%cmp = icmp eq i32 %a, 7
br i1 %cmp, label %if.end, label %if.then
@@ -41,7 +47,6 @@ define i32 @test_bnei(i32 %a) nounwind {
; RV32I-NEXT: li a0, 1
; RV32I-NEXT: ret
; RV32I-NEXT: .LBB1_2: # %if.then
-; RV32I-NEXT: li a0, 7
; RV32I-NEXT: ret
;
; RV32IZIBI-LABEL: test_bnei:
@@ -52,6 +57,12 @@ define i32 @test_bnei(i32 %a) nounwind {
; RV32IZIBI-NEXT: ret
; RV32IZIBI-NEXT: .LBB1_2: # %if.then
; RV32IZIBI-NEXT: ret
+; RV64I-LABEL: test_bnei:
+; RV64I: li a1, 7
+; RV64I-NEXT: sext.w a0, a0
+; RV64I-NEXT: beq a0, a1, .LBB1_2
+; RV64I: .LBB1_2:
+; RV64I-NEXT: ret
entry:
%cmp = icmp ne i32 %a, 7
br i1 %cmp, label %if.end, label %if.then
diff --git a/llvm/test/CodeGen/RISCV/xandesperf-redundant-copy-elim.ll b/llvm/test/CodeGen/RISCV/xandesperf-redundant-copy-elim.ll
index d01eb2082a6fc..12497baa67556 100644
--- a/llvm/test/CodeGen/RISCV/xandesperf-redundant-copy-elim.ll
+++ b/llvm/test/CodeGen/RISCV/xandesperf-redundant-copy-elim.ll
@@ -8,12 +8,10 @@ define i32 @test_beqc(i32 %a) nounwind {
; RV32I-LABEL: test_beqc:
; RV32I: # %bb.0: # %entry
; RV32I-NEXT: li a1, 7
-; RV32I-NEXT: bne a0, a1, .LBB0_2
-; RV32I-NEXT: # %bb.1: # %if.end
-; RV32I-NEXT: li a0, 7
-; RV32I-NEXT: ret
-; RV32I-NEXT: .LBB0_2: # %if.then
+; RV32I-NEXT: beq a0, a1, .LBB0_2
+; RV32I-NEXT: # %bb.1: # %if.then
; RV32I-NEXT: li a0, 1
+; RV32I-NEXT: .LBB0_2: # %if.end
; RV32I-NEXT: ret
;
; RV32IXANDESPERF-LABEL: test_beqc:
@@ -41,7 +39,6 @@ define i32 @test_bnec(i32 %a) nounwind {
; RV32I-NEXT: li a0, 1
; RV32I-NEXT: ret
; RV32I-NEXT: .LBB1_2: # %if.then
-; RV32I-NEXT: li a0, 7
; RV32I-NEXT: ret
;
; RV32IXANDESPERF-LABEL: test_bnec:
diff --git a/llvm/test/CodeGen/RISCV/xqcibi-redundant-copy-elim.ll b/llvm/test/CodeGen/RISCV/xqcibi-redundant-copy-elim.ll
index df5c6a21b0868..8bd370c48947f 100644
--- a/llvm/test/CodeGen/RISCV/xqcibi-redundant-copy-elim.ll
+++ b/llvm/test/CodeGen/RISCV/xqcibi-redundant-copy-elim.ll
@@ -10,12 +10,10 @@ define dso_local i32 @test_beqi(i32 %a) nounwind {
; RV32I-LABEL: test_beqi:
; RV32I: # %bb.0: # %entry
; RV32I-NEXT: li a1, 7
-; RV32I-NEXT: bne a0, a1, .LBB0_2
-; RV32I-NEXT: # %bb.1: # %if.end
-; RV32I-NEXT: li a0, 7
-; RV32I-NEXT: ret
-; RV32I-NEXT: .LBB0_2: # %if.then
+; RV32I-NEXT: beq a0, a1, .LBB0_2
+; RV32I-NEXT: # %bb.1: # %if.then
; RV32I-NEXT: li a0, 1
+; RV32I-NEXT: .LBB0_2: # %if.end
; RV32I-NEXT: ret
;
; RV32IXQCIBI-LABEL: test_beqi:
@@ -54,12 +52,10 @@ define dso_local i32 @test_e_beqi(i32 %a) nounwind {
; RV32I-LABEL: test_e_beqi:
; RV32I: # %bb.0: # %entry
; RV32I-NEXT: li a1, 40
-; RV32I-NEXT: bne a0, a1, .LBB1_2
-; RV32I-NEXT: # %bb.1: # %if.end
-; RV32I-NEXT: li a0, 40
-; RV32I-NEXT: ret
-; RV32I-NEXT: .LBB1_2: # %if.then
+; RV32I-NEXT: beq a0, a1, .LBB1_2
+; RV32I-NEXT: # %bb.1: # %if.then
; RV32I-NEXT: li a0, 1
+; RV32I-NEXT: .LBB1_2: # %if.end
; RV32I-NEXT: ret
;
; RV32IXQCIBI-LABEL: test_e_beqi:
``````````
</details>
https://github.com/llvm/llvm-project/pull/227673
More information about the llvm-commits
mailing list