[llvm] [AtomicExpand] Don't sink the release fence into a weak cmpxchg (PR #214867)

Josef Schlehofer via llvm-commits llvm-commits at lists.llvm.org
Sun Aug 9 01:56:03 PDT 2026


https://github.com/BKPepe updated https://github.com/llvm/llvm-project/pull/214867

>From 068eed4c7428e3f279a10e87705a33b0c8492ae9 Mon Sep 17 00:00:00 2001
From: Josef Schlehofer <pepe at bloodkings.eu>
Date: Sat, 8 Aug 2026 00:31:56 +0200
Subject: [PATCH] [AtomicExpand] Let targets keep the release fence out of the
 reservation

expandAtomicCmpXchg sinks the leading fence into the conditional
cmpxchg.fencedstore block, which places it between the load-linked and the
store-conditional. For the strong form that is harmless: if the
store-conditional fails, the retry re-reserves after the fence and the loop
still converges. The weak form has no retry, so where a fence can clear the
reservation it can never succeed.

Add fenceClearsLoadLinkedReservation(), defaulting to false, and hoist the
fence for a weak cmpxchg only on targets that say yes. PowerPC does: the
Power ISA lets an implementation drop a reservation for reasons of its own,
and e500v2 does so for a sync between the lwarx and the stwcx., which is
where this was found. Every compare_exchange_weak with a release ordering
spins there forever, and since fetch_update in Rust's core is a weak retry
loop, async runtimes hang on the platform.

Hoisting moves the fence ahead of the comparison, so it also executes on the
path where the comparison fails and no store follows. The fence is only
needed to provide the release ordering of a successful store; executing it
on an attempt that fails before the store is permitted and has no additional
ordering effect. It adds the cost of the fence on that path and nothing
else -- no extra synchronisation, store or loop.

For targets where the hook stays false the UseUnconditionalReleaseBarrier
condition is unchanged from the previous implementation, so their fence
placement and codegen are preserved. ARM is one: it emits a DMB in the sunk
position and its monitor survives one, so it has no reason to opt in.

Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
---
 llvm/include/llvm/CodeGen/TargetLowering.h    |   7 +
 llvm/lib/CodeGen/AtomicExpandPass.cpp         |  13 +-
 llvm/lib/Target/PowerPC/PPCISelLowering.h     |   7 +
 llvm/test/CodeGen/PowerPC/atomics.ll          |   2 +-
 .../PowerPC/weak-cmpxchg-fence-placement.ll   | 166 ++++++++++++++++++
 5 files changed, 193 insertions(+), 2 deletions(-)
 create mode 100644 llvm/test/Transforms/AtomicExpand/PowerPC/weak-cmpxchg-fence-placement.ll

diff --git a/llvm/include/llvm/CodeGen/TargetLowering.h b/llvm/include/llvm/CodeGen/TargetLowering.h
index 9a525d69b3ee8..c735926946908 100644
--- a/llvm/include/llvm/CodeGen/TargetLowering.h
+++ b/llvm/include/llvm/CodeGen/TargetLowering.h
@@ -2315,6 +2315,13 @@ class LLVM_ABI TargetLoweringBase {
     return false;
   }
 
+  /// Whether a fence placed between the load-linked and the store-conditional
+  /// can clear the reservation on this target. When it can, AtomicExpandPass
+  /// must not sink the leading fence of a cmpxchg into the reservation window:
+  /// the store-conditional would fail, and a weak cmpxchg has no retry to
+  /// re-reserve and recover with, so it could never succeed. Defaults to false.
+  virtual bool fenceClearsLoadLinkedReservation() const { return false; }
+
   /// Whether AtomicExpandPass should automatically insert a seq_cst trailing
   /// fence without reducing the ordering for this atomic store. Defaults to
   /// false.
diff --git a/llvm/lib/CodeGen/AtomicExpandPass.cpp b/llvm/lib/CodeGen/AtomicExpandPass.cpp
index 2b9c9ea8229d0..0471599e52fe0 100644
--- a/llvm/lib/CodeGen/AtomicExpandPass.cpp
+++ b/llvm/lib/CodeGen/AtomicExpandPass.cpp
@@ -1498,7 +1498,18 @@ bool AtomicExpandImpl::expandAtomicCmpXchg(AtomicCmpXchgInst *CI) {
 
   // There's no overhead for sinking the release barrier in a weak cmpxchg, so
   // do it even on minsize.
-  bool UseUnconditionalReleaseBarrier = F->hasMinSize() && !CI->isWeak();
+  //
+  // Except where a fence can clear the reservation. Sinking puts the fence
+  // between the load-linked and the store-conditional, and there the
+  // store-conditional fails; the strong form recovers because its retry
+  // re-reserves afterwards, but a weak cmpxchg has no retry and so could never
+  // succeed. Hoisting is valid because the release ordering belongs to the
+  // successful store, not to individual failed reservation attempts, so
+  // executing the fence once before the load-linked establishes the same
+  // ordering.
+  bool UseUnconditionalReleaseBarrier =
+      (F->hasMinSize() && !CI->isWeak()) ||
+      (CI->isWeak() && TLI->fenceClearsLoadLinkedReservation());
 
   // Given: cmpxchg some_op iN* %addr, iN %desired, iN %new success_ord fail_ord
   //
diff --git a/llvm/lib/Target/PowerPC/PPCISelLowering.h b/llvm/lib/Target/PowerPC/PPCISelLowering.h
index a39f7e0d3c35f..295b4dbb42dd2 100644
--- a/llvm/lib/Target/PowerPC/PPCISelLowering.h
+++ b/llvm/lib/Target/PowerPC/PPCISelLowering.h
@@ -338,6 +338,13 @@ namespace llvm {
       return true;
     }
 
+    /// The Power ISA lets an implementation clear a reservation for reasons of
+    /// its own, and e500v2 does so for a sync between the lwarx and the
+    /// stwcx., which leaves a weak cmpxchg unable to ever succeed. There is no
+    /// e500 subtarget to narrow this to, and the cost elsewhere is one fence on
+    /// the comparison-failed path.
+    bool fenceClearsLoadLinkedReservation() const override { return true; }
+
     Value *emitLoadLinked(IRBuilderBase &Builder, Type *ValueTy, Value *Addr,
                           AtomicOrdering Ord) const override;
 
diff --git a/llvm/test/CodeGen/PowerPC/atomics.ll b/llvm/test/CodeGen/PowerPC/atomics.ll
index af2b8646d18ae..30897e06cbea1 100644
--- a/llvm/test/CodeGen/PowerPC/atomics.ll
+++ b/llvm/test/CodeGen/PowerPC/atomics.ll
@@ -310,12 +310,12 @@ define i64 @cas_weak_i64_release_monotonic(ptr %mem) {
 ;
 ; PPC64-LABEL: cas_weak_i64_release_monotonic:
 ; PPC64:       # %bb.0: # %cmpxchg.start
+; PPC64-NEXT:    lwsync
 ; PPC64-NEXT:    mr r4, r3
 ; PPC64-NEXT:    ldarx r3, 0, r3
 ; PPC64-NEXT:    cmpldi r3, 0
 ; PPC64-NEXT:    bnelr- cr0
 ; PPC64-NEXT:  # %bb.1: # %cmpxchg.fencedstore
-; PPC64-NEXT:    lwsync
 ; PPC64-NEXT:    li r5, 1
 ; PPC64-NEXT:    stdcx. r5, 0, r4
 ; PPC64-NEXT:    blr
diff --git a/llvm/test/Transforms/AtomicExpand/PowerPC/weak-cmpxchg-fence-placement.ll b/llvm/test/Transforms/AtomicExpand/PowerPC/weak-cmpxchg-fence-placement.ll
new file mode 100644
index 0000000000000..1a0fd06eb03a3
--- /dev/null
+++ b/llvm/test/Transforms/AtomicExpand/PowerPC/weak-cmpxchg-fence-placement.ll
@@ -0,0 +1,166 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 5
+; RUN: opt -S -mtriple=powerpc-unknown-linux-musl \
+; RUN:   -passes='require<libcall-lowering-info>,atomic-expand' %s | FileCheck %s
+
+; The release barrier for a weak cmpxchg belongs before the load-linked, not
+; between it and the store-conditional. A barrier inside the reservation window
+; can clear the reservation, and a weak cmpxchg has no retry to recover with, so
+; sinking it there leaves an operation that can never succeed. The strong form
+; keeps the sunk barrier: its retry re-reserves afterwards, so it converges
+; either way.
+
+define i1 @weak_acq_rel(ptr %addr) {
+; CHECK-LABEL: define i1 @weak_acq_rel(
+; CHECK-SAME: ptr [[ADDR:%.*]]) {
+; CHECK-NEXT:    call void @llvm.ppc.lwsync()
+; CHECK-NEXT:    br label %[[CMPXCHG_START:.*]]
+; CHECK:       [[CMPXCHG_START]]:
+; CHECK-NEXT:    [[LARX:%.*]] = call i32 @llvm.ppc.lwarx(ptr [[ADDR]])
+; CHECK-NEXT:    [[SHOULD_STORE:%.*]] = icmp eq i32 [[LARX]], 0
+; CHECK-NEXT:    br i1 [[SHOULD_STORE]], label %[[CMPXCHG_FENCEDSTORE:.*]], label %[[CMPXCHG_NOSTORE:.*]], !prof [[PROF0:![0-9]+]]
+; CHECK:       [[CMPXCHG_FENCEDSTORE]]:
+; CHECK-NEXT:    br label %[[CMPXCHG_TRYSTORE:.*]]
+; CHECK:       [[CMPXCHG_TRYSTORE]]:
+; CHECK-NEXT:    [[LOADED_TRYSTORE:%.*]] = phi i32 [ [[LARX]], %[[CMPXCHG_FENCEDSTORE]] ]
+; CHECK-NEXT:    [[STCX:%.*]] = call i32 @llvm.ppc.stwcx(ptr [[ADDR]], i32 1)
+; CHECK-NEXT:    [[TMP1:%.*]] = xor i32 [[STCX]], 1
+; CHECK-NEXT:    [[SUCCESS:%.*]] = icmp eq i32 [[TMP1]], 0
+; CHECK-NEXT:    br i1 [[SUCCESS]], label %[[CMPXCHG_SUCCESS:.*]], label %[[CMPXCHG_FAILURE:.*]], !prof [[PROF0]]
+; CHECK:       [[CMPXCHG_RELEASEDLOAD:.*:]]
+; CHECK-NEXT:    unreachable
+; CHECK:       [[CMPXCHG_SUCCESS]]:
+; CHECK-NEXT:    call void @llvm.ppc.lwsync()
+; CHECK-NEXT:    br label %[[CMPXCHG_END:.*]]
+; CHECK:       [[CMPXCHG_NOSTORE]]:
+; CHECK-NEXT:    [[LOADED_NOSTORE:%.*]] = phi i32 [ [[LARX]], %[[CMPXCHG_START]] ]
+; CHECK-NEXT:    br label %[[CMPXCHG_FAILURE]]
+; CHECK:       [[CMPXCHG_FAILURE]]:
+; CHECK-NEXT:    [[LOADED_FAILURE:%.*]] = phi i32 [ [[LOADED_NOSTORE]], %[[CMPXCHG_NOSTORE]] ], [ [[LOADED_TRYSTORE]], %[[CMPXCHG_TRYSTORE]] ]
+; CHECK-NEXT:    call void @llvm.ppc.lwsync()
+; CHECK-NEXT:    br label %[[CMPXCHG_END]]
+; CHECK:       [[CMPXCHG_END]]:
+; CHECK-NEXT:    [[LOADED_EXIT:%.*]] = phi i32 [ [[LOADED_TRYSTORE]], %[[CMPXCHG_SUCCESS]] ], [ [[LOADED_FAILURE]], %[[CMPXCHG_FAILURE]] ]
+; CHECK-NEXT:    [[SUCCESS1:%.*]] = phi i1 [ true, %[[CMPXCHG_SUCCESS]] ], [ false, %[[CMPXCHG_FAILURE]] ]
+; CHECK-NEXT:    ret i1 [[SUCCESS1]]
+;
+  %pair = cmpxchg weak ptr %addr, i32 0, i32 1 acq_rel acquire
+  %ok = extractvalue { i32, i1 } %pair, 1
+  ret i1 %ok
+}
+
+define i1 @weak_release(ptr %addr) {
+; CHECK-LABEL: define i1 @weak_release(
+; CHECK-SAME: ptr [[ADDR:%.*]]) {
+; CHECK-NEXT:    call void @llvm.ppc.lwsync()
+; CHECK-NEXT:    br label %[[CMPXCHG_START:.*]]
+; CHECK:       [[CMPXCHG_START]]:
+; CHECK-NEXT:    [[LARX:%.*]] = call i32 @llvm.ppc.lwarx(ptr [[ADDR]])
+; CHECK-NEXT:    [[SHOULD_STORE:%.*]] = icmp eq i32 [[LARX]], 0
+; CHECK-NEXT:    br i1 [[SHOULD_STORE]], label %[[CMPXCHG_FENCEDSTORE:.*]], label %[[CMPXCHG_NOSTORE:.*]], !prof [[PROF0]]
+; CHECK:       [[CMPXCHG_FENCEDSTORE]]:
+; CHECK-NEXT:    br label %[[CMPXCHG_TRYSTORE:.*]]
+; CHECK:       [[CMPXCHG_TRYSTORE]]:
+; CHECK-NEXT:    [[LOADED_TRYSTORE:%.*]] = phi i32 [ [[LARX]], %[[CMPXCHG_FENCEDSTORE]] ]
+; CHECK-NEXT:    [[STCX:%.*]] = call i32 @llvm.ppc.stwcx(ptr [[ADDR]], i32 1)
+; CHECK-NEXT:    [[TMP1:%.*]] = xor i32 [[STCX]], 1
+; CHECK-NEXT:    [[SUCCESS:%.*]] = icmp eq i32 [[TMP1]], 0
+; CHECK-NEXT:    br i1 [[SUCCESS]], label %[[CMPXCHG_SUCCESS:.*]], label %[[CMPXCHG_FAILURE:.*]], !prof [[PROF0]]
+; CHECK:       [[CMPXCHG_RELEASEDLOAD:.*:]]
+; CHECK-NEXT:    unreachable
+; CHECK:       [[CMPXCHG_SUCCESS]]:
+; CHECK-NEXT:    br label %[[CMPXCHG_END:.*]]
+; CHECK:       [[CMPXCHG_NOSTORE]]:
+; CHECK-NEXT:    [[LOADED_NOSTORE:%.*]] = phi i32 [ [[LARX]], %[[CMPXCHG_START]] ]
+; CHECK-NEXT:    br label %[[CMPXCHG_FAILURE]]
+; CHECK:       [[CMPXCHG_FAILURE]]:
+; CHECK-NEXT:    [[LOADED_FAILURE:%.*]] = phi i32 [ [[LOADED_NOSTORE]], %[[CMPXCHG_NOSTORE]] ], [ [[LOADED_TRYSTORE]], %[[CMPXCHG_TRYSTORE]] ]
+; CHECK-NEXT:    br label %[[CMPXCHG_END]]
+; CHECK:       [[CMPXCHG_END]]:
+; CHECK-NEXT:    [[LOADED_EXIT:%.*]] = phi i32 [ [[LOADED_TRYSTORE]], %[[CMPXCHG_SUCCESS]] ], [ [[LOADED_FAILURE]], %[[CMPXCHG_FAILURE]] ]
+; CHECK-NEXT:    [[SUCCESS1:%.*]] = phi i1 [ true, %[[CMPXCHG_SUCCESS]] ], [ false, %[[CMPXCHG_FAILURE]] ]
+; CHECK-NEXT:    ret i1 [[SUCCESS1]]
+;
+  %pair = cmpxchg weak ptr %addr, i32 0, i32 1 release monotonic
+  %ok = extractvalue { i32, i1 } %pair, 1
+  ret i1 %ok
+}
+
+define i1 @strong_acq_rel(ptr %addr) {
+; CHECK-LABEL: define i1 @strong_acq_rel(
+; CHECK-SAME: ptr [[ADDR:%.*]]) {
+; CHECK-NEXT:    br label %[[CMPXCHG_START:.*]]
+; CHECK:       [[CMPXCHG_START]]:
+; CHECK-NEXT:    [[LARX:%.*]] = call i32 @llvm.ppc.lwarx(ptr [[ADDR]])
+; CHECK-NEXT:    [[SHOULD_STORE:%.*]] = icmp eq i32 [[LARX]], 0
+; CHECK-NEXT:    br i1 [[SHOULD_STORE]], label %[[CMPXCHG_FENCEDSTORE:.*]], label %[[CMPXCHG_NOSTORE:.*]], !prof [[PROF0]]
+; CHECK:       [[CMPXCHG_FENCEDSTORE]]:
+; CHECK-NEXT:    call void @llvm.ppc.lwsync()
+; CHECK-NEXT:    br label %[[CMPXCHG_TRYSTORE:.*]]
+; CHECK:       [[CMPXCHG_TRYSTORE]]:
+; CHECK-NEXT:    [[LOADED_TRYSTORE:%.*]] = phi i32 [ [[LARX]], %[[CMPXCHG_FENCEDSTORE]] ], [ [[LARX1:%.*]], %[[CMPXCHG_RELEASEDLOAD:.*]] ]
+; CHECK-NEXT:    [[STCX:%.*]] = call i32 @llvm.ppc.stwcx(ptr [[ADDR]], i32 1)
+; CHECK-NEXT:    [[TMP1:%.*]] = xor i32 [[STCX]], 1
+; CHECK-NEXT:    [[SUCCESS:%.*]] = icmp eq i32 [[TMP1]], 0
+; CHECK-NEXT:    br i1 [[SUCCESS]], label %[[CMPXCHG_SUCCESS:.*]], label %[[CMPXCHG_RELEASEDLOAD]], !prof [[PROF0]]
+; CHECK:       [[CMPXCHG_RELEASEDLOAD]]:
+; CHECK-NEXT:    [[LARX1]] = call i32 @llvm.ppc.lwarx(ptr [[ADDR]])
+; CHECK-NEXT:    [[SHOULD_STORE2:%.*]] = icmp eq i32 [[LARX1]], 0
+; CHECK-NEXT:    br i1 [[SHOULD_STORE2]], label %[[CMPXCHG_TRYSTORE]], label %[[CMPXCHG_NOSTORE]], !prof [[PROF0]]
+; CHECK:       [[CMPXCHG_SUCCESS]]:
+; CHECK-NEXT:    call void @llvm.ppc.lwsync()
+; CHECK-NEXT:    br label %[[CMPXCHG_END:.*]]
+; CHECK:       [[CMPXCHG_NOSTORE]]:
+; CHECK-NEXT:    [[LOADED_NOSTORE:%.*]] = phi i32 [ [[LARX]], %[[CMPXCHG_START]] ], [ [[LARX1]], %[[CMPXCHG_RELEASEDLOAD]] ]
+; CHECK-NEXT:    br label %[[CMPXCHG_FAILURE:.*]]
+; CHECK:       [[CMPXCHG_FAILURE]]:
+; CHECK-NEXT:    [[LOADED_FAILURE:%.*]] = phi i32 [ [[LOADED_NOSTORE]], %[[CMPXCHG_NOSTORE]] ]
+; CHECK-NEXT:    call void @llvm.ppc.lwsync()
+; CHECK-NEXT:    br label %[[CMPXCHG_END]]
+; CHECK:       [[CMPXCHG_END]]:
+; CHECK-NEXT:    [[LOADED_EXIT:%.*]] = phi i32 [ [[LOADED_TRYSTORE]], %[[CMPXCHG_SUCCESS]] ], [ [[LOADED_FAILURE]], %[[CMPXCHG_FAILURE]] ]
+; CHECK-NEXT:    [[SUCCESS3:%.*]] = phi i1 [ true, %[[CMPXCHG_SUCCESS]] ], [ false, %[[CMPXCHG_FAILURE]] ]
+; CHECK-NEXT:    ret i1 [[SUCCESS3]]
+;
+  %pair = cmpxchg ptr %addr, i32 0, i32 1 acq_rel acquire
+  %ok = extractvalue { i32, i1 } %pair, 1
+  ret i1 %ok
+}
+
+define i1 @weak_monotonic(ptr %addr) {
+; CHECK-LABEL: define i1 @weak_monotonic(
+; CHECK-SAME: ptr [[ADDR:%.*]]) {
+; CHECK-NEXT:    br label %[[CMPXCHG_START:.*]]
+; CHECK:       [[CMPXCHG_START]]:
+; CHECK-NEXT:    [[LARX:%.*]] = call i32 @llvm.ppc.lwarx(ptr [[ADDR]])
+; CHECK-NEXT:    [[SHOULD_STORE:%.*]] = icmp eq i32 [[LARX]], 0
+; CHECK-NEXT:    br i1 [[SHOULD_STORE]], label %[[CMPXCHG_FENCEDSTORE:.*]], label %[[CMPXCHG_NOSTORE:.*]], !prof [[PROF0]]
+; CHECK:       [[CMPXCHG_FENCEDSTORE]]:
+; CHECK-NEXT:    br label %[[CMPXCHG_TRYSTORE:.*]]
+; CHECK:       [[CMPXCHG_TRYSTORE]]:
+; CHECK-NEXT:    [[LOADED_TRYSTORE:%.*]] = phi i32 [ [[LARX]], %[[CMPXCHG_FENCEDSTORE]] ]
+; CHECK-NEXT:    [[STCX:%.*]] = call i32 @llvm.ppc.stwcx(ptr [[ADDR]], i32 1)
+; CHECK-NEXT:    [[TMP1:%.*]] = xor i32 [[STCX]], 1
+; CHECK-NEXT:    [[SUCCESS:%.*]] = icmp eq i32 [[TMP1]], 0
+; CHECK-NEXT:    br i1 [[SUCCESS]], label %[[CMPXCHG_SUCCESS:.*]], label %[[CMPXCHG_FAILURE:.*]], !prof [[PROF0]]
+; CHECK:       [[CMPXCHG_RELEASEDLOAD:.*:]]
+; CHECK-NEXT:    unreachable
+; CHECK:       [[CMPXCHG_SUCCESS]]:
+; CHECK-NEXT:    br label %[[CMPXCHG_END:.*]]
+; CHECK:       [[CMPXCHG_NOSTORE]]:
+; CHECK-NEXT:    [[LOADED_NOSTORE:%.*]] = phi i32 [ [[LARX]], %[[CMPXCHG_START]] ]
+; CHECK-NEXT:    br label %[[CMPXCHG_FAILURE]]
+; CHECK:       [[CMPXCHG_FAILURE]]:
+; CHECK-NEXT:    [[LOADED_FAILURE:%.*]] = phi i32 [ [[LOADED_NOSTORE]], %[[CMPXCHG_NOSTORE]] ], [ [[LOADED_TRYSTORE]], %[[CMPXCHG_TRYSTORE]] ]
+; CHECK-NEXT:    br label %[[CMPXCHG_END]]
+; CHECK:       [[CMPXCHG_END]]:
+; CHECK-NEXT:    [[LOADED_EXIT:%.*]] = phi i32 [ [[LOADED_TRYSTORE]], %[[CMPXCHG_SUCCESS]] ], [ [[LOADED_FAILURE]], %[[CMPXCHG_FAILURE]] ]
+; CHECK-NEXT:    [[SUCCESS1:%.*]] = phi i1 [ true, %[[CMPXCHG_SUCCESS]] ], [ false, %[[CMPXCHG_FAILURE]] ]
+; CHECK-NEXT:    ret i1 [[SUCCESS1]]
+;
+  %pair = cmpxchg weak ptr %addr, i32 0, i32 1 monotonic monotonic
+  %ok = extractvalue { i32, i1 } %pair, 1
+  ret i1 %ok
+}
+;.
+; CHECK: [[PROF0]] = !{!"branch_weights", i32 1048575, i32 1}
+;.



More information about the llvm-commits mailing list