[clang] [llvm] [OMPIRBuilder] Perform atomic compare on padded FP types at their allocation size (PR #226567)
Akash Manna via llvm-commits
llvm-commits at lists.llvm.org
Fri Sep 25 11:57:35 PDT 2026
https://github.com/akash-manna-sky updated https://github.com/llvm/llvm-project/pull/226567
>From c2ac33e6fa1237f7d7fd171f1832f16194014475 Mon Sep 17 00:00:00 2001
From: Akash Manna <akash.manna.mymail at gmail.com>
Date: Sat, 26 Sep 2026 00:22:57 +0530
Subject: [PATCH 1/2] [OMPIRBuilder] Perform atomic compare on padded FP types
at their allocation size
createAtomicCompare cast floating-point operands to an integer of the
type's scalar width, so long double became i80. IRBuilder then filled
the missing cmpxchg alignment from the store size of i80, which is 10,
and Align asserted. An i80 atomic would not have passed the verifier
either.
Derive the integer type from the DataLayout allocation size instead, so
x86_fp80 is handled as i128 with its 16-byte alignment, and zero-fill
the padding bits when widening. Types without padding are unaffected.
Fixes #140080
---
clang/docs/ReleaseNotes.md | 1 +
.../atomic_compare_long_double_codegen.c | 54 ++++++++++++++++++
llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp | 56 +++++++++++++++----
3 files changed, 99 insertions(+), 12 deletions(-)
create mode 100644 clang/test/OpenMP/atomic_compare_long_double_codegen.c
diff --git a/clang/docs/ReleaseNotes.md b/clang/docs/ReleaseNotes.md
index f4a34a37aff52..ae537970b7055 100644
--- a/clang/docs/ReleaseNotes.md
+++ b/clang/docs/ReleaseNotes.md
@@ -547,6 +547,7 @@ features cannot lower the translation-unit ABI level;
- Fixed a crash when an `asm` label names the register for a global variable of incomplete type. (#GH219746)
- Fixed an ICE hat occurred when using `__imag int/float` as lvalue in assignment. (#GH119498)
- Fixed an assertion failure in `-Wsign-compare` when a negated or complemented vector of unsigned integers was compared against a signed constant. (#GH203575)
+- Fixed an assertion failure when `#pragma omp atomic compare` was applied to a `long double` on x86-64, whose 80-bit value is stored in 16 bytes. (#GH140080)
#### Bug Fixes to Compiler Builtins
diff --git a/clang/test/OpenMP/atomic_compare_long_double_codegen.c b/clang/test/OpenMP/atomic_compare_long_double_codegen.c
new file mode 100644
index 0000000000000..410920f71e49d
--- /dev/null
+++ b/clang/test/OpenMP/atomic_compare_long_double_codegen.c
@@ -0,0 +1,54 @@
+// RUN: %clang_cc1 -verify -triple x86_64-unknown-linux-gnu -fopenmp -x c -emit-llvm %s -o - | FileCheck %s
+// RUN: %clang_cc1 -verify -triple x86_64-unknown-linux-gnu -fopenmp-simd -x c -emit-llvm %s -o - | FileCheck --check-prefix SIMD-ONLY0 %s
+// expected-no-diagnostics
+
+// GH140080: x86_fp80 occupies 16 bytes, so the cmpxchg must be on i128 with
+// align 16, not on i80 (store size 10 is not a power of two).
+
+// CHECK-LABEL: define {{.*}}void @f(
+// CHECK: [[LD_ADDR:%.*]] = alloca x86_fp80, align 16
+// CHECK: cmpxchg ptr [[LD_ADDR]], i128 0, i128 302222231531620438900736 monotonic monotonic, align 16
+// CHECK-NOT: i80
+// CHECK: ret void
+void f(long double ld) {
+#pragma omp atomic compare
+ ld = ld == 0.0L ? 1.0L : ld;
+}
+
+// CHECK-LABEL: define {{.*}}void @g(
+// CHECK: [[X:%.*]] = load ptr, ptr %x.addr, align 8
+// CHECK-NEXT: [[E:%.*]] = load x86_fp80, ptr %e.addr, align 16
+// CHECK-NEXT: [[D:%.*]] = load x86_fp80, ptr %d.addr, align 16
+// CHECK-NEXT: [[E_BITS:%.*]] = bitcast x86_fp80 [[E]] to i80
+// CHECK-NEXT: [[E_INT:%.*]] = zext i80 [[E_BITS]] to i128
+// CHECK-NEXT: [[D_BITS:%.*]] = bitcast x86_fp80 [[D]] to i80
+// CHECK-NEXT: [[D_INT:%.*]] = zext i80 [[D_BITS]] to i128
+// CHECK-NEXT: cmpxchg ptr [[X]], i128 [[E_INT]], i128 [[D_INT]] monotonic monotonic, align 16
+void g(long double *x, long double e, long double d) {
+#pragma omp atomic compare
+ *x = *x == e ? d : *x;
+}
+
+// CHECK-LABEL: define {{.*}}void @h(
+// CHECK: [[RES:%.*]] = cmpxchg ptr {{%.*}}, i128 {{%.*}}, i128 {{%.*}} monotonic monotonic, align 16
+// CHECK-NEXT: [[OLD_INT:%.*]] = extractvalue { i128, i1 } [[RES]], 0
+// CHECK-NEXT: [[OLD_BITS:%.*]] = trunc i128 [[OLD_INT]] to i80
+// CHECK-NEXT: [[OLD:%.*]] = bitcast i80 [[OLD_BITS]] to x86_fp80
+// CHECK-NEXT: store x86_fp80 [[OLD]], ptr {{%.*}}, align 16
+void h(long double *x, long double *v, long double e, long double d) {
+#pragma omp atomic compare capture
+ {
+ *v = *x;
+ *x = *x == e ? d : *x;
+ }
+}
+
+// CHECK-LABEL: define {{.*}}void @f_double(
+// CHECK: [[D_ADDR:%.*]] = alloca double, align 8
+// CHECK: cmpxchg ptr [[D_ADDR]], i64 0, i64 4607182418800017408 monotonic monotonic, align 8
+void f_double(double d) {
+#pragma omp atomic compare
+ d = d == 0.0 ? 1.0 : d;
+}
+
+// SIMD-ONLY0-NOT: {{__kmpc|__tgt}}
diff --git a/llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp b/llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp
index 8b196c461f195..4bf9007ac0ed3 100644
--- a/llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp
+++ b/llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp
@@ -11700,6 +11700,38 @@ OpenMPIRBuilder::InsertPointOrErrorTy OpenMPIRBuilder::createAtomicCapture(
return Builder.saveIP();
}
+/// The integer type used to perform an atomic operation on a value of type
+/// \p FPTy: its allocation size, not its scalar width (x86_fp80 is 80 bits in
+/// 16 bytes, and an i80 atomic is not a legal power-of-two access).
+static IntegerType *getAtomicIntTypeForFP(Type *FPTy, const DataLayout &DL) {
+ return IntegerType::get(FPTy->getContext(), DL.getTypeAllocSizeInBits(FPTy));
+}
+
+/// Reinterpret \p V as \p IntTy, zero-filling any padding bits.
+static Value *castFPToAtomicInt(IRBuilderBase &Builder, const DataLayout &DL,
+ Value *V, IntegerType *IntTy) {
+ unsigned ValueBits = V->getType()->getPrimitiveSizeInBits();
+ if (ValueBits == IntTy->getBitWidth())
+ return Builder.CreateBitCast(V, IntTy);
+ // zext/trunc match the memory layout only on little-endian targets, which
+ // are the only ones with padded FP types.
+ assert(DL.isLittleEndian() && "padded FP type on a big-endian target");
+ Value *Bits =
+ Builder.CreateBitCast(V, IntegerType::get(V->getContext(), ValueBits));
+ return Builder.CreateZExt(Bits, IntTy);
+}
+
+static Value *castAtomicIntToFP(IRBuilderBase &Builder, const DataLayout &DL,
+ Value *V, Type *FPTy, const Twine &Name = "") {
+ unsigned ValueBits = FPTy->getPrimitiveSizeInBits();
+ if (ValueBits == V->getType()->getIntegerBitWidth())
+ return Builder.CreateBitCast(V, FPTy, Name);
+ assert(DL.isLittleEndian() && "padded FP type on a big-endian target");
+ Value *Bits =
+ Builder.CreateTrunc(V, IntegerType::get(V->getContext(), ValueBits));
+ return Builder.CreateBitCast(Bits, FPTy, Name);
+}
+
OpenMPIRBuilder::InsertPointTy OpenMPIRBuilder::createAtomicCompare(
const LocationDescription &Loc, AtomicOpValue &X, AtomicOpValue &V,
AtomicOpValue &R, Value *E, Value *D, AtomicOrdering AO,
@@ -11765,16 +11797,16 @@ OpenMPIRBuilder::InsertPointTy OpenMPIRBuilder::createAtomicCompare(
// br ExitBB
// ExitBB:
// phi merge
- IntegerType *IntCastTy =
- IntegerType::get(M.getContext(), X.ElemTy->getScalarSizeInBits());
- Value *EBCast = Builder.CreateBitCast(E, IntCastTy);
- Value *DBCast = Builder.CreateBitCast(D, IntCastTy);
+ const DataLayout &DL = M.getDataLayout();
+ IntegerType *IntCastTy = getAtomicIntTypeForFP(X.ElemTy, DL);
+ Value *EBCast = castFPToAtomicInt(Builder, DL, E, IntCastTy);
+ Value *DBCast = castFPToAtomicInt(Builder, DL, D, IntCastTy);
// Load X atomically.
LoadInst *XCurr = Builder.CreateLoad(IntCastTy, X.Var,
X.Var->getName() + ".atomic.load");
XCurr->setAtomic(AtomicOrdering::Monotonic);
- Value *XFP = Builder.CreateBitCast(XCurr, X.ElemTy);
+ Value *XFP = castAtomicIntToFP(Builder, DL, XCurr, X.ElemTy);
// IEEE 754: NaN != NaN, but cmpxchg would succeed if E and X have
// the same NaN bit pattern. Skip cmpxchg when either is NaN.
@@ -11854,16 +11886,16 @@ OpenMPIRBuilder::InsertPointTy OpenMPIRBuilder::createAtomicCompare(
Builder.SetInsertPoint(&*ExitBB->getFirstNonPHIIt());
}
- OldValue = Builder.CreateBitCast(OldIntPHI, X.ElemTy,
- X.Var->getName() + ".atomic.old.fp");
+ OldValue = castAtomicIntToFP(Builder, DL, OldIntPHI, X.ElemTy,
+ X.Var->getName() + ".atomic.old.fp");
SuccessOrFail = SuccessPHI;
} else {
AtomicCmpXchgInst *Result = nullptr;
+ const DataLayout &DL = M.getDataLayout();
if (!IsInteger) {
- IntegerType *IntCastTy =
- IntegerType::get(M.getContext(), X.ElemTy->getScalarSizeInBits());
- Value *EBCast = Builder.CreateBitCast(E, IntCastTy);
- Value *DBCast = Builder.CreateBitCast(D, IntCastTy);
+ IntegerType *IntCastTy = getAtomicIntTypeForFP(X.ElemTy, DL);
+ Value *EBCast = castFPToAtomicInt(Builder, DL, E, IntCastTy);
+ Value *DBCast = castFPToAtomicInt(Builder, DL, D, IntCastTy);
Result = Builder.CreateAtomicCmpXchg(X.Var, EBCast, DBCast,
MaybeAlign(), AO, Failure);
} else {
@@ -11875,7 +11907,7 @@ OpenMPIRBuilder::InsertPointTy OpenMPIRBuilder::createAtomicCompare(
if (V.Var) {
OldValue = Builder.CreateExtractValue(Result, /*Idxs=*/0);
if (!IsInteger)
- OldValue = Builder.CreateBitCast(OldValue, X.ElemTy);
+ OldValue = castAtomicIntToFP(Builder, DL, OldValue, X.ElemTy);
assert(OldValue->getType() == V.ElemTy &&
"OldValue and V must be of same type");
if (IsPostfixUpdate) {
>From 0831cb73a0cd1f2c74b21b57e1a0675c391524ad Mon Sep 17 00:00:00 2001
From: Akash Manna <akash.manna.mymail at gmail.com>
Date: Sat, 26 Sep 2026 00:27:13 +0530
Subject: [PATCH 2/2] Reposition the release notes to avoid the conflicts
---
clang/docs/ReleaseNotes.md | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/clang/docs/ReleaseNotes.md b/clang/docs/ReleaseNotes.md
index ae537970b7055..fec29a1440581 100644
--- a/clang/docs/ReleaseNotes.md
+++ b/clang/docs/ReleaseNotes.md
@@ -533,6 +533,7 @@ features cannot lower the translation-unit ABI level;
- Fixed USR generation for declarations whose signature mentions a class-type
non-type template parameter. (#GH212351)
- Fixed an assertion caused by Microsoft integer literals exceeding the maximum value. (#GH212504)
+- Fixed an assertion failure when `#pragma omp atomic compare` was applied to a `long double` on x86-64, whose 80-bit value is stored in 16 bytes. (#GH140080)
- Fixed an assertion failure when a value of a Unicode character type (`char8_t`, `char16_t`, `char32_t`) was implicitly splatted to a vector of the same element type, e.g. when comparing an `ext_vector_type` of `char32_t` with one of its elements. (#GH202317)
- Fixed a crash when checking scalar type with excess braces. (#GH69213), (#GH137845), (#GH198767), (#GH207566), (#GH106180)
- Fixed an assertion crash when instantiating a nested requirement with an invalid constraint. (#GH213575)
@@ -547,7 +548,6 @@ features cannot lower the translation-unit ABI level;
- Fixed a crash when an `asm` label names the register for a global variable of incomplete type. (#GH219746)
- Fixed an ICE hat occurred when using `__imag int/float` as lvalue in assignment. (#GH119498)
- Fixed an assertion failure in `-Wsign-compare` when a negated or complemented vector of unsigned integers was compared against a signed constant. (#GH203575)
-- Fixed an assertion failure when `#pragma omp atomic compare` was applied to a `long double` on x86-64, whose 80-bit value is stored in 16 bytes. (#GH140080)
#### Bug Fixes to Compiler Builtins
More information about the llvm-commits
mailing list