[flang-commits] [flang] [llvm] [mlir] [OpenMP] Honor barrier and last iteration flags in loop lowering (PR #224040)
Nicole Aschenbrenner via flang-commits
flang-commits at lists.llvm.org
Mon Oct 5 04:50:20 PDT 2026
https://github.com/nicebert updated https://github.com/llvm/llvm-project/pull/224040
>From d110cc067ca2e919980a4f0a85dd275d2678c7cf Mon Sep 17 00:00:00 2001
From: Nicole Aschenbrenner <nicole.aschenbrenner at amd.com>
Date: Thu, 27 Aug 2026 06:33:46 -0500
Subject: [PATCH 1/3] [OpenMP] Honor barrier and last iteration flags in loop
lowering
Emitting loop-free kernels for no-loop target regions in Clang requires
the shared OpenMP lowering to honor flags that reach it today and are
then discarded.
applyWorkshareLoop takes a NeedsBarrier flag, but the device path drops
it, so a worksharing loop without nowait emits no barrier at its exit.
The device path also never sets the canonical loop's last iteration
variable, which the linear clause finalization reads. Forward the flag
to applyWorkshareLoopTarget and compute the last iteration in the loop
body, mirroring how the host runtime reports it.
Auditing the surrounding lowering for the same class of problem turned
up one more. The barrier that follows privatization is emitted as part
of the firstprivate copy region, so a construct with lastprivate and no
firstprivate never gets one. Emit it independently of the copy region.
Flang skips it for taskloop, where the write-back already happens after
the reads.
---
.../lib/Lower/OpenMP/DataSharingProcessor.cpp | 7 ++
.../OpenMP/privatization-barrier.f90 | 38 ++++++++
flang/test/Lower/OpenMP/taskloop.f90 | 15 +++
.../llvm/Frontend/OpenMP/OMPIRBuilder.h | 8 +-
llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp | 27 +++++-
.../OpenMP/OpenMPToLLVMIRTranslation.cpp | 96 ++++++++++---------
mlir/test/Target/LLVMIR/omptarget-wsloop.mlir | 21 ++++
7 files changed, 163 insertions(+), 49 deletions(-)
create mode 100644 flang/test/Integration/OpenMP/privatization-barrier.f90
diff --git a/flang/lib/Lower/OpenMP/DataSharingProcessor.cpp b/flang/lib/Lower/OpenMP/DataSharingProcessor.cpp
index 8d0d191058cfb8..3e63c068e1cdb9 100644
--- a/flang/lib/Lower/OpenMP/DataSharingProcessor.cpp
+++ b/flang/lib/Lower/OpenMP/DataSharingProcessor.cpp
@@ -372,6 +372,13 @@ bool DataSharingProcessor::needBarrier() {
// initialization of firstprivate variables and post-update of lastprivate
// variables.
// Emit implicit barrier for linear clause in the OpenMPIRBuilder.
+ // Skip for taskloop: the write-back only happens after the reads are done.
+
+ const auto *ompEval = eval.getIf<parser::OpenMPConstruct>();
+ if (ompEval && llvm::omp::allTaskloopSet.test(
+ parser::omp::GetOmpDirectiveName(*ompEval).v))
+ return false;
+
for (const semantics::Symbol *sym : allPrivatizedSymbols) {
if (sym->test(semantics::Symbol::Flag::OmpLastPrivate) &&
(sym->test(semantics::Symbol::Flag::OmpFirstPrivate) ||
diff --git a/flang/test/Integration/OpenMP/privatization-barrier.f90 b/flang/test/Integration/OpenMP/privatization-barrier.f90
new file mode 100644
index 00000000000000..06ab26df2521e7
--- /dev/null
+++ b/flang/test/Integration/OpenMP/privatization-barrier.f90
@@ -0,0 +1,38 @@
+!===----------------------------------------------------------------------===!
+! This directory can be used to add Integration tests involving multiple
+! stages of the compiler (for eg. from Fortran to LLVM IR). It should not
+! contain executable tests. We should only add tests here sparingly and only
+! if there is no other way to test. Repeat this message in each test that is
+! added to this directory and sub-directories.
+!===----------------------------------------------------------------------===!
+
+! RUN: %flang_fc1 -fopenmp -emit-llvm %s -o - | FileCheck %s
+! RUN: %if amdgpu-registered-target %{ %flang_fc1 -triple amdgcn-amd-amdhsa -emit-llvm -fopenmp -fopenmp-is-target-device %s -o - | FileCheck %s %}
+
+subroutine lastprivate_allocatable_barrier_host
+ integer, allocatable :: a
+ integer :: i
+ !$omp parallel do lastprivate(a)
+ do i = 1, 10
+ a = i
+ end do
+ !$omp end parallel do
+end subroutine
+
+subroutine lastprivate_allocatable_barrier_device
+ integer, allocatable :: a
+ integer :: i
+ allocate(a)
+ !$omp target parallel do lastprivate(a)
+ do i = 1, 10
+ a = i
+ end do
+ !$omp end target parallel do
+end subroutine
+
+! CHECK-LABEL: define internal void @{{.*}}lastprivate_allocatable_barrier_{{(host|device)}}
+! CHECK: call void @__kmpc_barrier
+! CHECK-NEXT: br label %omp.wsloop.region
+! CHECK: call void @__kmpc_barrier
+! CHECK-NEXT: br label %omp_loop.after
+! CHECK-LABEL: define{{.*}}void @{{.*}}lastprivate_allocatable_barrier_device
diff --git a/flang/test/Lower/OpenMP/taskloop.f90 b/flang/test/Lower/OpenMP/taskloop.f90
index 81c57a0139d358..049291161de46a 100644
--- a/flang/test/Lower/OpenMP/taskloop.f90
+++ b/flang/test/Lower/OpenMP/taskloop.f90
@@ -2,6 +2,9 @@
! RUN: bbc -emit-hlfir %openmp_flags -fopenmp-version=50 -o - %s 2>&1 | FileCheck %s
! RUN: %flang_fc1 -emit-hlfir %openmp_flags -fopenmp-version=50 -o - %s 2>&1 | FileCheck %s
+! CHECK-LABEL: omp.private {type = firstprivate}
+! CHECK-SAME: @[[FIRST_LAST_PRIVATE_X:.*]] : i32
+
! CHECK-LABEL: omp.private
! CHECK-SAME: {type = private} @[[LAST_PRIVATE_I:.*]] : i32
@@ -260,3 +263,15 @@ subroutine omp_taskloop_lastprivate()
! CHECK: omp.terminator
!$omp end taskloop
end subroutine omp_taskloop_lastprivate
+
+! CHECK-LABEL: func @_QPomp_taskloop_first_and_lastprivate()
+subroutine omp_taskloop_first_and_lastprivate()
+ integer x
+ x = 0
+ ! CHECK: omp.taskloop.context private(@[[FIRST_LAST_PRIVATE_X]] {{.*}}) {
+ !$omp taskloop firstprivate(x) lastprivate(x)
+ do i = 1, 100
+ x = x + 1
+ end do
+ !$omp end taskloop
+end subroutine omp_taskloop_first_and_lastprivate
diff --git a/llvm/include/llvm/Frontend/OpenMP/OMPIRBuilder.h b/llvm/include/llvm/Frontend/OpenMP/OMPIRBuilder.h
index 24b542bd7b3587..4c0aa22b67c64e 100644
--- a/llvm/include/llvm/Frontend/OpenMP/OMPIRBuilder.h
+++ b/llvm/include/llvm/Frontend/OpenMP/OMPIRBuilder.h
@@ -1179,11 +1179,13 @@ class OpenMPIRBuilder {
/// \param NeedsBarrier Indicates whether a barrier must be inserted after
/// the loop.
/// \param NoLoop If true, no-loop code is generated.
+ /// \param NeedsLastIter If true, the last iteration variable is emitted.
///
/// \returns Point where to insert code after the workshare construct.
InsertPointOrErrorTy applyWorkshareLoopTarget(
DebugLoc DL, CanonicalLoopInfo *CLI, InsertPointTy AllocaIP,
- omp::WorksharingLoopType LoopType, bool NeedsBarrier, bool NoLoop);
+ omp::WorksharingLoopType LoopType, bool NeedsBarrier, bool NoLoop,
+ bool NeedsLastIter);
/// Modifies the canonical loop to be a statically-scheduled workshare loop.
///
@@ -1341,8 +1343,8 @@ class OpenMPIRBuilder {
/// \param NoLoop If true, no-loop code is generated.
/// \param HasDistSchedule Defines if the clause being lowered is
/// dist_schedule as this is handled slightly differently
- ///
/// \param DistScheduleChunkSize The chunk size for dist_schedule loop
+ /// \param NeedsLastIter If true, the last iteration variable is emitted.
///
/// \returns Point where to insert code after the workshare construct.
LLVM_ABI InsertPointOrErrorTy applyWorkshareLoop(
@@ -1355,7 +1357,7 @@ class OpenMPIRBuilder {
omp::WorksharingLoopType LoopType =
omp::WorksharingLoopType::ForStaticLoop,
bool NoLoop = false, bool HasDistSchedule = false,
- Value *DistScheduleChunkSize = nullptr);
+ Value *DistScheduleChunkSize = nullptr, bool NeedsLastIter = false);
/// Tile a loop nest.
///
diff --git a/llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp b/llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp
index 0cd72fe348f416..862055b33bac73 100644
--- a/llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp
+++ b/llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp
@@ -6606,9 +6606,30 @@ static void workshareLoopTargetCallback(
OpenMPIRBuilder::InsertPointOrErrorTy OpenMPIRBuilder::applyWorkshareLoopTarget(
DebugLoc DL, CanonicalLoopInfo *CLI, InsertPointTy AllocaIP,
- WorksharingLoopType LoopType, bool NeedsBarrier, bool NoLoop) {
+ WorksharingLoopType LoopType, bool NeedsBarrier, bool NoLoop,
+ bool NeedsLastIter) {
uint32_t SrcLocStrSize;
Constant *SrcLocStr = getOrCreateSrcLocStr(DL, SrcLocStrSize);
+
+ // Mirrors host runtime reporting of last iteration by in-body computation.
+ if (NeedsLastIter) {
+ Type *I32Type = Type::getInt32Ty(M.getContext());
+ Builder.restoreIP(AllocaIP);
+ AllocaInst *PLastIter =
+ Builder.CreateAlloca(I32Type, nullptr, "p.lastiter");
+ Builder.CreateStore(ConstantInt::get(I32Type, 0), PLastIter);
+ CLI->setLastIter(PLastIter);
+
+ Builder.SetInsertPoint(CLI->getBody(),
+ CLI->getBody()->getFirstInsertionPt());
+ Value *TripCount = CLI->getTripCount();
+ Value *LastIter =
+ Builder.CreateSub(TripCount, ConstantInt::get(TripCount->getType(), 1));
+ Value *IsLast =
+ Builder.CreateICmpEQ(CLI->getIndVar(), LastIter, "omp.is_last_iter");
+ Builder.CreateStore(Builder.CreateZExt(IsLast, I32Type), PLastIter);
+ }
+
IdentFlag Flag = IdentFlag(0);
switch (LoopType) {
case WorksharingLoopType::ForStaticLoop:
@@ -6728,10 +6749,10 @@ OpenMPIRBuilder::InsertPointOrErrorTy OpenMPIRBuilder::applyWorkshareLoop(
bool HasSimdModifier, bool HasMonotonicModifier,
bool HasNonmonotonicModifier, bool HasOrderedClause,
WorksharingLoopType LoopType, bool NoLoop, bool HasDistSchedule,
- Value *DistScheduleChunkSize) {
+ Value *DistScheduleChunkSize, bool NeedsLastIter) {
if (Config.isTargetDevice())
return applyWorkshareLoopTarget(DL, CLI, AllocaIP, LoopType, NeedsBarrier,
- NoLoop);
+ NoLoop, NeedsLastIter);
OMPScheduleType EffectiveScheduleType = computeOpenMPScheduleType(
SchedKind, ChunkSize, HasSimdModifier, HasMonotonicModifier,
HasNonmonotonicModifier, HasOrderedClause, DistScheduleChunkSize);
diff --git a/mlir/lib/Target/LLVMIR/Dialect/OpenMP/OpenMPToLLVMIRTranslation.cpp b/mlir/lib/Target/LLVMIR/Dialect/OpenMP/OpenMPToLLVMIRTranslation.cpp
index ba9b30e69aacef..30f37d1ee86c29 100644
--- a/mlir/lib/Target/LLVMIR/Dialect/OpenMP/OpenMPToLLVMIRTranslation.cpp
+++ b/mlir/lib/Target/LLVMIR/Dialect/OpenMP/OpenMPToLLVMIRTranslation.cpp
@@ -2151,13 +2151,27 @@ static bool opIsInSingleThread(mlir::Operation *op) {
return false;
}
-static LogicalResult copyFirstPrivateVars(
- mlir::Operation *op, llvm::IRBuilderBase &builder,
- LLVM::ModuleTranslation &moduleTranslation,
- SmallVectorImpl<llvm::Value *> &moldVars,
- ArrayRef<llvm::Value *> llvmPrivateVars,
- SmallVectorImpl<omp::PrivateClauseOp> &privateDecls, bool insertBarrier,
- llvm::DenseMap<Value, Value> *mappedPrivateVars = nullptr) {
+static LogicalResult
+emitPrivatizationBarrier(mlir::Operation *op, llvm::IRBuilderBase &builder,
+ LLVM::ModuleTranslation &moduleTranslation,
+ bool insertBarrier) {
+ if (!insertBarrier || opIsInSingleThread(op))
+ return success();
+
+ llvm::OpenMPIRBuilder *ompBuilder = moduleTranslation.getOpenMPBuilder();
+ llvm::OpenMPIRBuilder::InsertPointOrErrorTy res =
+ ompBuilder->createBarrier(builder, llvm::omp::OMPD_barrier);
+ return handleError(res, *op);
+}
+
+static LogicalResult
+completePrivateVars(mlir::Operation *op, llvm::IRBuilderBase &builder,
+ LLVM::ModuleTranslation &moduleTranslation,
+ SmallVectorImpl<llvm::Value *> &moldVars,
+ ArrayRef<llvm::Value *> llvmPrivateVars,
+ SmallVectorImpl<omp::PrivateClauseOp> &privateDecls,
+ bool insertBarrier,
+ llvm::DenseMap<Value, Value> *mappedPrivateVars = nullptr) {
// Apply copy region for firstprivate.
bool needsFirstprivate =
llvm::any_of(privateDecls, [](omp::PrivateClauseOp &privOp) {
@@ -2166,7 +2180,8 @@ static LogicalResult copyFirstPrivateVars(
});
if (!needsFirstprivate)
- return success();
+ return emitPrivatizationBarrier(op, builder, moduleTranslation,
+ insertBarrier);
llvm::BasicBlock *copyBlock =
splitBB(builder, /*CreateBranch=*/true, "omp.private.copy");
@@ -2205,24 +2220,18 @@ static LogicalResult copyFirstPrivateVars(
moduleTranslation.forgetMapping(copyRegion);
}
- if (insertBarrier && !opIsInSingleThread(op)) {
- llvm::OpenMPIRBuilder *ompBuilder = moduleTranslation.getOpenMPBuilder();
- llvm::OpenMPIRBuilder::InsertPointOrErrorTy res =
- ompBuilder->createBarrier(builder, llvm::omp::OMPD_barrier);
- if (failed(handleError(res, *op)))
- return failure();
- }
-
- return success();
+ return emitPrivatizationBarrier(op, builder, moduleTranslation,
+ insertBarrier);
}
-static LogicalResult copyFirstPrivateVars(
- mlir::Operation *op, llvm::IRBuilderBase &builder,
- LLVM::ModuleTranslation &moduleTranslation,
- SmallVectorImpl<mlir::Value> &mlirPrivateVars,
- ArrayRef<llvm::Value *> llvmPrivateVars,
- SmallVectorImpl<omp::PrivateClauseOp> &privateDecls, bool insertBarrier,
- llvm::DenseMap<Value, Value> *mappedPrivateVars = nullptr) {
+static LogicalResult
+completePrivateVars(mlir::Operation *op, llvm::IRBuilderBase &builder,
+ LLVM::ModuleTranslation &moduleTranslation,
+ SmallVectorImpl<mlir::Value> &mlirPrivateVars,
+ ArrayRef<llvm::Value *> llvmPrivateVars,
+ SmallVectorImpl<omp::PrivateClauseOp> &privateDecls,
+ bool insertBarrier,
+ llvm::DenseMap<Value, Value> *mappedPrivateVars = nullptr) {
llvm::SmallVector<llvm::Value *> moldVars(mlirPrivateVars.size());
llvm::transform(mlirPrivateVars, moldVars.begin(), [&](mlir::Value mlirVar) {
// map copyRegion rhs arg
@@ -2231,9 +2240,9 @@ static LogicalResult copyFirstPrivateVars(
assert(moldVar);
return moldVar;
});
- return copyFirstPrivateVars(op, builder, moduleTranslation, moldVars,
- llvmPrivateVars, privateDecls, insertBarrier,
- mappedPrivateVars);
+ return completePrivateVars(op, builder, moduleTranslation, moldVars,
+ llvmPrivateVars, privateDecls, insertBarrier,
+ mappedPrivateVars);
}
template <typename T>
@@ -2503,7 +2512,7 @@ convertOmpScope(omp::ScopeOp &scopeOp, llvm::IRBuilderBase &builder,
.failed())
return llvm::make_error<PreviouslyReportedError>();
- if (failed(copyFirstPrivateVars(
+ if (failed(completePrivateVars(
scopeOp, builder, moduleTranslation, privateVarsInfo.mlirVars,
privateVarsInfo.llvmVars, privateVarsInfo.privatizers,
scopeOp.getPrivateNeedsBarrier())))
@@ -3424,7 +3433,7 @@ convertOmpTaskOp(omp::TaskOp taskOp, llvm::IRBuilderBase &builder,
// firstprivate copy region
setInsertPointForPossiblyEmptyBlock(builder, copyBlock);
- if (failed(copyFirstPrivateVars(
+ if (failed(completePrivateVars(
taskOp, builder, moduleTranslation, privateVarsInfo.mlirVars,
taskStructMgr.getLLVMPrivateVarGEPs(), privateVarsInfo.privatizers,
taskOp.getPrivateNeedsBarrier())))
@@ -3896,7 +3905,7 @@ convertOmpTaskloopContextOp(omp::TaskloopContextOp contextOp,
// firstprivate copy region
setInsertPointForPossiblyEmptyBlock(builder, copyBlock);
- if (failed(copyFirstPrivateVars(
+ if (failed(completePrivateVars(
contextOp, builder, moduleTranslation, privateVarsInfo.mlirVars,
taskStructMgr.getLLVMPrivateVarGEPs(), privateVarsInfo.privatizers,
contextOp.getPrivateNeedsBarrier())))
@@ -4195,10 +4204,10 @@ convertOmpTaskloopContextOp(omp::TaskloopContextOp contextOp,
// through a stack allocated structure.
}
- if (failed(copyFirstPrivateVars(contextOp.getOperation(), builder,
- moduleTranslation, srcGEPs, destGEPs,
- privateVarsInfo.privatizers,
- contextOp.getPrivateNeedsBarrier())))
+ if (failed(completePrivateVars(contextOp.getOperation(), builder,
+ moduleTranslation, srcGEPs, destGEPs,
+ privateVarsInfo.privatizers,
+ contextOp.getPrivateNeedsBarrier())))
return llvm::make_error<PreviouslyReportedError>();
return builder.saveIP();
@@ -4794,7 +4803,7 @@ convertOmpWsloop(Operation &opInst, llvm::IRBuilderBase &builder,
.failed())
return failure();
- if (failed(copyFirstPrivateVars(
+ if (failed(completePrivateVars(
wsloopOp, builder, moduleTranslation, privateVarsInfo.mlirVars,
privateVarsInfo.llvmVars, privateVarsInfo.privatizers,
wsloopOp.getPrivateNeedsBarrier())))
@@ -4907,7 +4916,8 @@ convertOmpWsloop(Operation &opInst, llvm::IRBuilderBase &builder,
convertToScheduleKind(schedule), chunk, isSimd,
scheduleMod == omp::ScheduleModifier::monotonic,
scheduleMod == omp::ScheduleModifier::nonmonotonic, isOrdered,
- workshareLoopType, noLoopMode, hasDistSchedule, distScheduleChunk);
+ workshareLoopType, noLoopMode, hasDistSchedule, distScheduleChunk,
+ !wsloopOp.getLinearVars().empty());
if (failed(handleError(wsloopIP, opInst)))
return failure();
@@ -5012,7 +5022,7 @@ convertOmpParallel(omp::ParallelOp opInst, llvm::IRBuilderBase &builder,
.failed())
return llvm::make_error<PreviouslyReportedError>();
- if (failed(copyFirstPrivateVars(
+ if (failed(completePrivateVars(
opInst, builder, moduleTranslation, privateVarsInfo.mlirVars,
privateVarsInfo.llvmVars, privateVarsInfo.privatizers,
opInst.getPrivateNeedsBarrier())))
@@ -5251,7 +5261,7 @@ convertOmpSimd(Operation &opInst, llvm::IRBuilderBase &builder,
.failed())
return failure();
- // No call to copyFirstPrivateVars because FIRSTPRIVATE is not allowed for
+ // No call to completePrivateVars because FIRSTPRIVATE is not allowed for
// SIMD.
assert(afterAllocas.get()->getSinglePredecessor());
@@ -8821,10 +8831,10 @@ convertOmpDistribute(Operation &opInst, llvm::IRBuilderBase &builder,
.failed())
return llvm::make_error<PreviouslyReportedError>();
- if (failed(copyFirstPrivateVars(
- distributeOp, builder, moduleTranslation, privVarsInfo.mlirVars,
- privVarsInfo.llvmVars, privVarsInfo.privatizers,
- distributeOp.getPrivateNeedsBarrier())))
+ if (failed(completePrivateVars(distributeOp, builder, moduleTranslation,
+ privVarsInfo.mlirVars, privVarsInfo.llvmVars,
+ privVarsInfo.privatizers,
+ distributeOp.getPrivateNeedsBarrier())))
return llvm::make_error<PreviouslyReportedError>();
llvm::OpenMPIRBuilder *ompBuilder = moduleTranslation.getOpenMPBuilder();
@@ -9705,7 +9715,7 @@ convertOmpTarget(Operation &opInst, llvm::IRBuilderBase &builder,
.failed())
return llvm::make_error<PreviouslyReportedError>();
- if (failed(copyFirstPrivateVars(
+ if (failed(completePrivateVars(
targetOp, builder, moduleTranslation, privateVarsInfo.mlirVars,
privateVarsInfo.llvmVars, privateVarsInfo.privatizers,
targetOp.getPrivateNeedsBarrier(), &mappedPrivateVars)))
diff --git a/mlir/test/Target/LLVMIR/omptarget-wsloop.mlir b/mlir/test/Target/LLVMIR/omptarget-wsloop.mlir
index 18357db4f6a9c7..86dc03e1818848 100644
--- a/mlir/test/Target/LLVMIR/omptarget-wsloop.mlir
+++ b/mlir/test/Target/LLVMIR/omptarget-wsloop.mlir
@@ -53,6 +53,18 @@ module attributes {dlti.dl_spec = #dlti.dl_spec<#dlti.dl_entry<"dlti.alloca_memo
}
llvm.return
}
+
+ llvm.func @target_wsloop_linear(%arg0: !llvm.ptr) attributes {omp.declare_target = #omp.declaretarget<device_type = any, capture_clause = to>} {
+ %loop_ub = llvm.mlir.constant(9 : i32) : i32
+ %loop_lb = llvm.mlir.constant(0 : i32) : i32
+ %loop_step = llvm.mlir.constant(1 : i32) : i32
+ omp.wsloop linear(%arg0 : !llvm.ptr = %loop_step : i32) linear_var_types([i32]) {
+ omp.loop_nest (%loop_cnt) : i32 = (%loop_lb) to (%loop_ub) inclusive step (%loop_step) {
+ omp.yield
+ }
+ }
+ llvm.return
+ }
}
// CHECK-LABEL: define hidden void @target_wsloop(
@@ -104,3 +116,12 @@ module attributes {dlti.dl_spec = #dlti.dl_spec<#dlti.dl_entry<"dlti.alloca_memo
// CHECK: define internal void @[[ZERO_TRIP_BODY]](
// CHECK-NOT: @__kmpc{{.*}}barrier
// CHECK: ret void
+
+// CHECK: define hidden void @target_wsloop_linear(ptr %{{.*}})
+// CHECK: store i32 0, ptr{{.*}} %p.lastiter
+// CHECK: call void @__kmpc_for_static_loop_4u
+// CHECK: call void @__kmpc_barrier(
+
+// CHECK: define internal void @target_wsloop_linear
+// CHECK: %omp.is_last_iter = icmp eq
+// CHECK: store{{.*}}_p.lastiter
>From d4aad02042efd8d312cd5b0f6d727ff71830bf4f Mon Sep 17 00:00:00 2001
From: Nicole Aschenbrenner <nicole.aschenbrenner at amd.com>
Date: Fri, 25 Sep 2026 03:19:48 -0500
Subject: [PATCH 2/3] Reset last iteration on every loop entry, simplify
barrier helper
---
llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp | 4 +++-
.../OpenMP/OpenMPToLLVMIRTranslation.cpp | 20 ++++++++++---------
2 files changed, 14 insertions(+), 10 deletions(-)
diff --git a/llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp b/llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp
index 862055b33bac73..37580d4da7fb09 100644
--- a/llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp
+++ b/llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp
@@ -6617,9 +6617,11 @@ OpenMPIRBuilder::InsertPointOrErrorTy OpenMPIRBuilder::applyWorkshareLoopTarget(
Builder.restoreIP(AllocaIP);
AllocaInst *PLastIter =
Builder.CreateAlloca(I32Type, nullptr, "p.lastiter");
- Builder.CreateStore(ConstantInt::get(I32Type, 0), PLastIter);
CLI->setLastIter(PLastIter);
+ Builder.SetInsertPoint(CLI->getPreheader()->getTerminator());
+ Builder.CreateStore(ConstantInt::get(I32Type, 0), PLastIter);
+
Builder.SetInsertPoint(CLI->getBody(),
CLI->getBody()->getFirstInsertionPt());
Value *TripCount = CLI->getTripCount();
diff --git a/mlir/lib/Target/LLVMIR/Dialect/OpenMP/OpenMPToLLVMIRTranslation.cpp b/mlir/lib/Target/LLVMIR/Dialect/OpenMP/OpenMPToLLVMIRTranslation.cpp
index 30f37d1ee86c29..f5dcf17e33124a 100644
--- a/mlir/lib/Target/LLVMIR/Dialect/OpenMP/OpenMPToLLVMIRTranslation.cpp
+++ b/mlir/lib/Target/LLVMIR/Dialect/OpenMP/OpenMPToLLVMIRTranslation.cpp
@@ -2153,9 +2153,8 @@ static bool opIsInSingleThread(mlir::Operation *op) {
static LogicalResult
emitPrivatizationBarrier(mlir::Operation *op, llvm::IRBuilderBase &builder,
- LLVM::ModuleTranslation &moduleTranslation,
- bool insertBarrier) {
- if (!insertBarrier || opIsInSingleThread(op))
+ LLVM::ModuleTranslation &moduleTranslation) {
+ if (opIsInSingleThread(op))
return success();
llvm::OpenMPIRBuilder *ompBuilder = moduleTranslation.getOpenMPBuilder();
@@ -2180,8 +2179,9 @@ completePrivateVars(mlir::Operation *op, llvm::IRBuilderBase &builder,
});
if (!needsFirstprivate)
- return emitPrivatizationBarrier(op, builder, moduleTranslation,
- insertBarrier);
+ return insertBarrier
+ ? emitPrivatizationBarrier(op, builder, moduleTranslation)
+ : success();
llvm::BasicBlock *copyBlock =
splitBB(builder, /*CreateBranch=*/true, "omp.private.copy");
@@ -2220,8 +2220,9 @@ completePrivateVars(mlir::Operation *op, llvm::IRBuilderBase &builder,
moduleTranslation.forgetMapping(copyRegion);
}
- return emitPrivatizationBarrier(op, builder, moduleTranslation,
- insertBarrier);
+ return insertBarrier
+ ? emitPrivatizationBarrier(op, builder, moduleTranslation)
+ : success();
}
static LogicalResult
@@ -5261,8 +5262,9 @@ convertOmpSimd(Operation &opInst, llvm::IRBuilderBase &builder,
.failed())
return failure();
- // No call to completePrivateVars because FIRSTPRIVATE is not allowed for
- // SIMD.
+ // No call to completePrivateVars because SIMD does not allow FIRSTPRIVATE and
+ // is not a worksharing construct, so a privatization barrier here could
+ // deadlock.
assert(afterAllocas.get()->getSinglePredecessor());
if (failed(initReductionVars(simdOp, reductionArgs, builder,
>From cbe9f4d49fb70f12aca618ced7cf01cf1175ac22 Mon Sep 17 00:00:00 2001
From: Nicole Aschenbrenner <nicole.aschenbrenner at amd.com>
Date: Mon, 5 Oct 2026 05:47:05 -0500
Subject: [PATCH 3/3] [flang][OpenMP][NFC] Make the taskloop private_barrier
check explicit
The existing check already failed if the attribute was printed, but a
CHECK-NOT makes it obvious the barrier must not be there.
---
flang/test/Lower/OpenMP/taskloop.f90 | 4 +++-
llvm/include/llvm/Frontend/OpenMP/OMPIRBuilder.h | 9 +++++----
2 files changed, 8 insertions(+), 5 deletions(-)
diff --git a/flang/test/Lower/OpenMP/taskloop.f90 b/flang/test/Lower/OpenMP/taskloop.f90
index 049291161de46a..1ed5b7257fb11c 100644
--- a/flang/test/Lower/OpenMP/taskloop.f90
+++ b/flang/test/Lower/OpenMP/taskloop.f90
@@ -268,7 +268,9 @@ end subroutine omp_taskloop_lastprivate
subroutine omp_taskloop_first_and_lastprivate()
integer x
x = 0
- ! CHECK: omp.taskloop.context private(@[[FIRST_LAST_PRIVATE_X]] {{.*}}) {
+ ! CHECK: omp.taskloop.context private(@[[FIRST_LAST_PRIVATE_X]] {{.*}})
+ ! CHECK-NOT: private_barrier
+ ! CHECK-SAME: {
!$omp taskloop firstprivate(x) lastprivate(x)
do i = 1, 100
x = x + 1
diff --git a/llvm/include/llvm/Frontend/OpenMP/OMPIRBuilder.h b/llvm/include/llvm/Frontend/OpenMP/OMPIRBuilder.h
index 4c0aa22b67c64e..5acbb692483093 100644
--- a/llvm/include/llvm/Frontend/OpenMP/OMPIRBuilder.h
+++ b/llvm/include/llvm/Frontend/OpenMP/OMPIRBuilder.h
@@ -1182,10 +1182,11 @@ class OpenMPIRBuilder {
/// \param NeedsLastIter If true, the last iteration variable is emitted.
///
/// \returns Point where to insert code after the workshare construct.
- InsertPointOrErrorTy applyWorkshareLoopTarget(
- DebugLoc DL, CanonicalLoopInfo *CLI, InsertPointTy AllocaIP,
- omp::WorksharingLoopType LoopType, bool NeedsBarrier, bool NoLoop,
- bool NeedsLastIter);
+ InsertPointOrErrorTy
+ applyWorkshareLoopTarget(DebugLoc DL, CanonicalLoopInfo *CLI,
+ InsertPointTy AllocaIP,
+ omp::WorksharingLoopType LoopType, bool NeedsBarrier,
+ bool NoLoop, bool NeedsLastIter);
/// Modifies the canonical loop to be a statically-scheduled workshare loop.
///
More information about the flang-commits
mailing list