[llvm] [LoongArch] Enable callee saved register optimization (PR #226809)

via llvm-commits llvm-commits at lists.llvm.org
Sun Sep 27 09:59:22 PDT 2026


https://github.com/lrzlin created https://github.com/llvm/llvm-project/pull/226809

Enable called saved register optimization implemented in `RAGreedy::tryAssignCSRFirstTime()` for LoongArch, which could replace callee saved regs store/load in prologue/epilogue by register spill/reload in cold blocks or register splits according to https://github.com/llvm/llvm-project/pull/220090. 

Keep the default value 30 for `CSRCostScale`, larger numbers like 80 or 120 though lead to better performance, but may also cause more regressions.

Here are the SPEC CPU2006 results running on a Loongson 3B6000 (2.1 Ghz)
# CINT2006
| Benchmark | Samples | Base time (s) | Scale 30 time (s) | Base ratio | 30 ratio | Score change | Notes |
|---|---|---|---|---|---|---|---|
| 400.perlbench | 3 | 364.0 | 372.1 | 26.85 | 26.26 | **−2.18%** | Above noise |
| 401.bzip2 | 3 | 510.7 | 511.3 | 18.89 | 18.87 | −0.10% | |
| 403.gcc | 3 | 317.6 | 320.8 | 25.35 | 25.09 | −1.00% | |
| 429.mcf | 1 | 255.2 | 254.7 | 35.74 | 35.81 | +0.20% | |
| 445.gobmk | 3 | 465.2 | 462.3 | 22.55 | 22.69 | +0.64% | |
| 456.hmmer | 1 | 455.9 | 455.8 | 20.47 | 20.47 | +0.01% | |
| 458.sjeng | 3 | 514.9 | 508.2 | 23.50 | 23.81 | **+1.33%** | Above noise |
| 462.libquantum | 3 | 467.5 | 521.7 | 44.32 | 39.71 | −10.40% | Result unstable |
| 464.h264ref | 1 | 597.9 | 597.6 | 37.01 | 37.03 | +0.06% | |
| 471.omnetpp | 3 | 176.6 | 181.6 | 35.39 | 34.41 | −2.75% | Base noise 3.8% |
| 473.astar | 3 | 367.0 | 356.0 | 19.13 | 19.72 | **+3.10%** | Above noise |
| 483.xalancbmk | 1 | 211.0 | 213.0 | 32.70 | 32.38 | −0.97% | |
| **Geomean** | | | | | | **−1.06%** | Excluding libquantum: **−0.16%** |

# CFP2006
| Benchmark | Samples | Base time (s) | Scale 30 time (s) | Base ratio | 30 ratio | Score change | Notes |
|---|---|---|---|---|---|---|---|
| 410.bwaves | — | — | — | — | 0.00% | Byte identical to base |
| 416.gamess | 1 | 605.2 | 607.0 | 32.35 | 32.25 | −0.30% | |
| 433.milc | 3 | 356.1 | 360.1 | 25.78 | 25.49 | −1.13% | Base noise 1.4%, within noise |
| 434.zeusmp | 3 | 256.7 | 256.9 | 35.45 | 35.43 | −0.04% | |
| 435.gromacs | 1 | 405.1 | 404.5 | 17.63 | 17.65 | +0.14% | |
| 436.cactusADM | 1 | 204.5 | 204.7 | 58.44 | 58.39 | −0.09% | |
| 437.leslie3d | 3 | 271.6 | 265.4 | 34.61 | 35.41 | +2.30% | Base noise 3.4%, within noise |
| 444.namd | 1 | 350.2 | 349.9 | 22.91 | 22.93 | +0.08% | |
| 447.dealII | 3 | 312.3 | 312.5 | 36.64 | 36.62 | −0.05% | |
| 450.soplex | 3 | 227.0 | 224.8 | 36.73 | 37.09 | +0.99% | |
| 453.povray | 3 | 153.7 | 144.7 | 34.60 | 36.76 | **+6.22%** | Above noise |
| 454.calculix | 1 | 906.5 | 905.8 | 9.10 | 9.11 | +0.08% | |
| 459.GemsFDTD | 3 | 308.7 | 308.6 | 34.38 | 34.38 | +0.01% | |
| 465.tonto | 1 | 311.3 | 311.9 | 31.61 | 31.55 | −0.18% | |
| 470.lbm | 1 | 298.0 | 297.4 | 46.11 | 46.20 | +0.21% | |
| 481.wrf | 3 | 234.7 | 236.8 | 47.59 | 47.17 | −0.89% | |
| 482.sphinx3 | 3 | 561.9 | 562.1 | 34.68 | 34.67 | −0.04% | |
| **Geomean** | | | | | | **+0.42%** | |

There are some real improvments on `473.astar` and `453.povray`, and a regression on `400.perlbench`. If we raise scale to 30, improments on `453.povray` will comes to 9%, however the regression on `400.perlbench` will also aggravate to 3%.

For the two improvment tests, there're also other optimization methods for them, such as the following two PRs working on better shrink-wrapping support (by a register basis instead of a block) or multiple save/restore point support, they're better solutions IMO, but consider their changes are much larger and widespread, it may takes a long time for them to be merged into main branch.

[[RISCV][WIP] Let RA do the CSR saves.](https://github.com/llvm/llvm-project/pull/90819)
[[Draft] Support save/restore point splitting in shrink-wrap](https://github.com/llvm/llvm-project/pull/119359)
[[Draft] Data flow based shrink wrapping](https://github.com/llvm/llvm-project/pull/191784)
[Shrink Wrap Save/Restore Points Splitting](https://discourse.llvm.org/t/shrink-wrap-save-restore-points-splitting/83581)

>From 3cbfe8be6364609bd28509589c9f2f6249d71ced Mon Sep 17 00:00:00 2001
From: Lin Runze <linrunze at loongson.cn>
Date: Fri, 25 Sep 2026 01:41:31 +0800
Subject: [PATCH] [LoongArch] Enable callee saved register optimization

---
 llvm/lib/Target/LoongArch/LoongArchRegisterInfo.h | 8 ++++++++
 1 file changed, 8 insertions(+)

diff --git a/llvm/lib/Target/LoongArch/LoongArchRegisterInfo.h b/llvm/lib/Target/LoongArch/LoongArchRegisterInfo.h
index 74861e0f87d6f..ab081168244a5 100644
--- a/llvm/lib/Target/LoongArch/LoongArchRegisterInfo.h
+++ b/llvm/lib/Target/LoongArch/LoongArchRegisterInfo.h
@@ -46,6 +46,14 @@ struct LoongArchRegisterInfo : public LoongArchGenRegisterInfo {
     return true;
   }
   bool canRealignStack(const MachineFunction &MF) const override;
+
+  unsigned getCSRFirstUseCost(const MachineFunction &MF) const override {
+    // The cost of 2 means push and pop for each CSR.
+    return 2;
+  }
+  unsigned getCSRCostScale(const MachineFunction &MF) const override {
+    return 30;
+  }
 };
 } // end namespace llvm
 



More information about the llvm-commits mailing list