[flang] [llvm] [flang][OpenMP] Sink intervening code unguarded into collapsed loop body (PR #225159)

Caroline Newcombe via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 29 10:53:18 PDT 2026


https://github.com/cenewcombe updated https://github.com/llvm/llvm-project/pull/225159

>From 952363ac41848bb70df3c51349a9bf4af458a2fe Mon Sep 17 00:00:00 2001
From: Caroline Newcombe <caroline.newcombe at hpe.com>
Date: Mon, 21 Sep 2026 10:27:00 -0500
Subject: [PATCH 1/3] [flang][OpenMP] Sink intervening code unguarded into
 collapsed loop body

---
 flang/docs/OpenMPSupport.md                   |   2 +-
 flang/lib/Lower/OpenMP/OpenMP.cpp             | 149 +-----
 .../Lower/OpenMP/collapse-imperfect-nest.f90  | 455 ++++++------------
 .../collapse-target-intervening-todo.f90      |  28 --
 ...-teams-distribute-collapse-intervening.f90 |  96 ++++
 5 files changed, 256 insertions(+), 474 deletions(-)
 delete mode 100644 flang/test/Lower/OpenMP/collapse-target-intervening-todo.f90
 create mode 100644 offload/test/offloading/fortran/target-teams-distribute-collapse-intervening.f90

diff --git a/flang/docs/OpenMPSupport.md b/flang/docs/OpenMPSupport.md
index a63ae9f5a2716..7f73dc2eabb97 100644
--- a/flang/docs/OpenMPSupport.md
+++ b/flang/docs/OpenMPSupport.md
@@ -165,7 +165,7 @@ Parser/Semantics, MLIR, Lowering, or the OpenMPIRBuilder.
 | requires directive | <span class="part">partial</span> | | Frontend and lowering coverage exists (`flang/test/Semantics/OpenMP/requires01.f90`-`requires10.f90`, `flang/test/Lower/OpenMP/requires.f90`), but some clauses are still flagged as unsupported (for example reverse_offload warning path). | [llvm/llvm-project#204647](https://github.com/llvm/llvm-project/pull/204647) |
 | teams construct on host | <span class="good">done</span> | | Teams support is established and exercised across semantics/lowering coverage in Flang OpenMP tests. | |
 | loop construct and order(concurrent) clause | <span class="part">partial</span> | | Loop and order-related coverage exists (`flang/test/Semantics/OpenMP/compiler-directives-loop.f90`, `flang/test/Semantics/OpenMP/order-clause01.f90`, `flang/test/Lower/OpenMP/loop-directive.f90`, `flang/test/Lower/OpenMP/order-clause.f90`), with some transformations still evolving. | [llvm/llvm-project#169346](https://github.com/llvm/llvm-project/pull/169346), [llvm/llvm-project#208315](https://github.com/llvm/llvm-project/pull/208315) |
-| collapsing imperfectly nested loops | <span class="none">unclaimed</span> | | Current checks primarily diagnose non-perfect nests (for example `flang/test/Semantics/OpenMP/do-collapse.f90`), and no dedicated support for imperfect-nest collapsing was identified. | [llvm/llvm-project#202435](https://github.com/llvm/llvm-project/pull/202435) |
+| collapsing imperfectly nested loops | <span class="good">done</span> | | Intervening code is emitted into the collapsed loop body, so it executes once per collapsed logical iteration. Semantics coverage in `flang/test/Semantics/OpenMP/do-collapse.f90`, lowering coverage in `flang/test/Lower/OpenMP/collapse-imperfect-nest.f90`. | [llvm/llvm-project#211000](https://github.com/llvm/llvm-project/pull/211000) |
 | if clause and nontemporal clause on simd | <span class="part">partial</span> | | SIMD nontemporal coverage exists (`flang/test/Semantics/OpenMP/nontemporal.f90`), but complete OpenMP 5.0-level coverage for all if(simd)/nontemporal combinations remains incomplete. | [llvm/llvm-project#110015](https://github.com/llvm/llvm-project/pull/110015) |
 | atomic in simd | <span class="none">unclaimed</span> | | No dedicated Flang OpenMP coverage for atomic-in-simd forms was identified in current parser/semantics/lowering tests. | |
 | detach clause on task and omp_fulfill_event routine | <span class="good">done</span> | | Flang semantics and lowering coverage exists (`flang/test/Semantics/OpenMP/detach01.f90`, `flang/test/Semantics/OpenMP/detach02.f90`, `flang/test/Lower/OpenMP/task_detach.f90`); runtime routine is available in OpenMP module/runtime. | [llvm/llvm-project#119172](https://github.com/llvm/llvm-project/pull/119172), [llvm/llvm-project#119128](https://github.com/llvm/llvm-project/pull/119128) |
diff --git a/flang/lib/Lower/OpenMP/OpenMP.cpp b/flang/lib/Lower/OpenMP/OpenMP.cpp
index df14d13b76a4b..3d47695f690e0 100644
--- a/flang/lib/Lower/OpenMP/OpenMP.cpp
+++ b/flang/lib/Lower/OpenMP/OpenMP.cpp
@@ -1022,19 +1022,13 @@ static void genNestedEvaluations(lower::AbstractConverter &converter,
     converter.genEval(e);
 }
 
-static mlir::Operation *setLoopVar(lower::AbstractConverter &converter,
-                                   mlir::Location loc, mlir::Value indexVal,
-                                   const semantics::Symbol *sym);
-
 /// Emit the body of a collapsed loop nest, including any intervening code
 /// from imperfect nesting at intermediate levels (CLN relaxation, applied
 /// retroactively for all OMP versions).
 ///
-/// Because omp.loop_nest places its entire body at the innermost nesting
-/// level, intervening code must be guarded so that it only executes on the
-/// iterations where the corresponding inner induction variables are at their
-/// initial (for intervening code before nested loop) or final (for intervening
-/// code after nested loop) values.
+/// Intervening code is emitted unguarded into the innermost body, so it runs
+/// once per collapsed logical iteration. OpenMP 6.0 leaves that count
+/// unspecified between that and once per enclosing iteration.
 ///
 /// \param [in] converter - PFT to MLIR conversion interface.
 /// \param [in] outerEval - the evaluation containing the outermost loop
@@ -1044,42 +1038,17 @@ static void genCollapsedLoopNestBody(lower::AbstractConverter &converter,
                                      lower::pft::Evaluation &outerEval,
                                      int collapseValue) {
   assert(collapseValue >= 1);
-  if (collapseValue == 1) {
-    genNestedEvaluations(converter, outerEval, /*collapseValue=*/1);
-    return;
-  }
-
-  fir::FirOpBuilder &firOpBuilder = converter.getFirOpBuilder();
-  const mlir::Location loc = converter.getCurrentLocation();
-
-  // Get the enclosing omp.loop_nest to access induction variables and bounds.
-  auto loopNestOp = mlir::dyn_cast<mlir::omp::LoopNestOp>(
-      firOpBuilder.getInsertionBlock()->getParentOp());
-  assert(loopNestOp && "expected to be inside omp.loop_nest");
 
   // Collect before/after evaluations at each intermediate level.
   struct LevelInfo {
     llvm::SmallVector<lower::pft::Evaluation *> before;
     llvm::SmallVector<lower::pft::Evaluation *> after;
-
-    // Whether this level carries intervening code (i.e. imperfect nesting).
-    bool hasInterveningCode() const {
-      return !before.empty() || !after.empty();
-    }
   };
   llvm::SmallVector<LevelInfo> levels;
 
-  // DO-variable symbol of each collapsed level (index 0 = outermost). Used to
-  // restore an inner loop's variable to its Fortran terminal value before
-  // emitting "after" intervening code (see below).
-  llvm::SmallVector<const semantics::Symbol *> ivSyms;
-
   lower::pft::Evaluation *curEval = &outerEval;
   for (int i = 0; i < collapseValue - 1; ++i) {
     lower::pft::Evaluation *doEval = getNestedDoConstruct(*curEval);
-    const semantics::Symbol *ivSym = getIterationVariableSymbol(*doEval);
-    assert(ivSym && "expected iteration variable on collapsed DO loop");
-    ivSyms.push_back(ivSym);
     LevelInfo level;
     bool pastDo = false;
     for (lower::pft::Evaluation &e : doEval->getNestedEvaluations()) {
@@ -1104,117 +1073,19 @@ static void genCollapsedLoopNestBody(lower::AbstractConverter &converter,
     levels.push_back(std::move(level));
     curEval = doEval;
   }
-  // DO-variable symbol of the innermost collapsed loop must be restored
-  // inside enclosing "after" regions.
-  const semantics::Symbol *innermostIvSym =
-      getIterationVariableSymbol(*getNestedDoConstruct(*curEval));
-  assert(innermostIvSym && "expected iteration variable on collapsed DO loop");
-  ivSyms.push_back(innermostIvSym);
-
-  // Build a guard condition: all induction variables from
-  // startLevel..endLevel-1 equal their respective bound values.
-  // For "before" guards (useLowerBound=true), compare iv == lb (first iter).
-  // For "after" guards (useLowerBound=false), compare iv == last_iv.
-  const auto lbs = loopNestOp.getLoopLowerBounds();
-  const auto ubs = loopNestOp.getLoopUpperBounds();
-  const auto steps = loopNestOp.getLoopSteps();
-
-  // The intervening-code guards and terminal-value restoration do arithmetic
-  // on the collapsed loop bounds. If those bounds are host_eval block arguments
-  // of an enclosing omp.target region, such uses are illegal, so diagnose
-  // instead of emitting IR the omp.target verifier rejects.
-  const bool hasInterveningCode = llvm::any_of(
-      levels, [](const LevelInfo &l) { return l.hasInterveningCode(); });
-  if (hasInterveningCode) {
-    auto isHostEvalValue = [](mlir::Value v) {
-      auto blockArg = mlir::dyn_cast<mlir::BlockArgument>(v);
-      if (!blockArg)
-        return false;
-      auto iface = mlir::dyn_cast<mlir::omp::BlockArgOpenMPOpInterface>(
-          blockArg.getOwner()->getParentOp());
-      return iface &&
-             llvm::is_contained(iface.getHostEvalBlockArgs(), blockArg);
-    };
-    if (llvm::any_of(lbs, isHostEvalValue) ||
-        llvm::any_of(ubs, isHostEvalValue) ||
-        llvm::any_of(steps, isHostEvalValue))
-      TODO(loc, "collapsed loop nest with intervening code whose loop bounds "
-                "are evaluated on the host for an enclosing 'target' region");
-  }
-
-  // Last value the induction variable at \p lvl actually takes:
-  // lb + ((ub - lb) / step) * step. For unit steps this is exactly ub.
-  auto computeLastIV = [&](const int lvl) -> mlir::Value {
-    const std::optional<llvm::APInt> constStep =
-        fir::getIntIfConstant(steps[lvl]);
-    if (constStep && (constStep->isOne() || constStep->isAllOnes()))
-      return ubs[lvl];
-    const mlir::Value lb = lbs[lvl];
-    const mlir::Value ub = ubs[lvl];
-    const mlir::Value step = steps[lvl];
-    const mlir::Value range =
-        mlir::arith::SubIOp::create(firOpBuilder, loc, ub, lb);
-    const mlir::Value tripMinus1 =
-        mlir::arith::DivSIOp::create(firOpBuilder, loc, range, step);
-    const mlir::Value lastOffset =
-        mlir::arith::MulIOp::create(firOpBuilder, loc, tripMinus1, step);
-    return mlir::arith::AddIOp::create(firOpBuilder, loc, lb, lastOffset);
-  };
-
-  auto buildGuard = [&](const int startLevel, const int endLevel,
-                        const bool useLowerBound) -> mlir::Value {
-    mlir::Value cond;
-    for (int lvl = startLevel; lvl < endLevel; ++lvl) {
-      const mlir::Value iv = loopNestOp.getRegion().getArgument(lvl);
-      const mlir::Value target = useLowerBound ? lbs[lvl] : computeLastIV(lvl);
-      const mlir::Value cmp = mlir::arith::CmpIOp::create(
-          firOpBuilder, loc, mlir::arith::CmpIPredicate::eq, iv, target);
-      if (!cond)
-        cond = cmp;
-      else
-        cond = mlir::arith::AndIOp::create(firOpBuilder, loc, cond, cmp);
-    }
-    return cond;
-  };
 
-  // Emit "before" code at each level, guarded by inner IVs == lower bounds.
-  for (int i = 0; i < static_cast<int>(levels.size()); ++i) {
-    if (levels[i].before.empty())
-      continue;
-    const mlir::Value guard =
-        buildGuard(i + 1, collapseValue, /*useLowerBound=*/true);
-    auto ifOp = fir::IfOp::create(firOpBuilder, loc, guard, /*else*/ false);
-    firOpBuilder.setInsertionPointToStart(&ifOp.getThenRegion().front());
-    for (auto *e : levels[i].before)
+  // Emit before code, outermost-to-innermost.
+  for (const LevelInfo &level : levels)
+    for (lower::pft::Evaluation *e : level.before)
       converter.genEval(*e);
-    firOpBuilder.setInsertionPointAfter(ifOp);
-  }
 
-  // Emit innermost loop body.
+  // Emit the innermost loop body.
   genNestedEvaluations(converter, *curEval, /*collapseValue=*/1);
 
-  // Emit "after" code at each level (innermost first), guarded by
-  // inner IVs == last iteration values (accounts for non-unit steps).
-  for (int i = static_cast<int>(levels.size()) - 1; i >= 0; --i) {
-    if (levels[i].after.empty())
-      continue;
-    const mlir::Value guard =
-        buildGuard(i + 1, collapseValue, /*useLowerBound=*/false);
-    auto ifOp = fir::IfOp::create(firOpBuilder, loc, guard, /*else*/ false);
-    firOpBuilder.setInsertionPointToStart(&ifOp.getThenRegion().front());
-    // A normally-terminated Fortran DO loop leaves its variable one step past
-    // the last executed value, but the flattened nest leaves each at its last
-    // executed value. Restore the terminal value before running "after" code
-    // that may read it.
-    for (int lvl = i + 1; lvl < collapseValue; ++lvl) {
-      const mlir::Value terminal = mlir::arith::AddIOp::create(
-          firOpBuilder, loc, computeLastIV(lvl), steps[lvl]);
-      setLoopVar(converter, loc, terminal, ivSyms[lvl]);
-    }
-    for (auto *e : levels[i].after)
+  // Emit after code, innermost-to-outermost.
+  for (const LevelInfo &level : llvm::reverse(levels))
+    for (lower::pft::Evaluation *e : level.after)
       converter.genEval(*e);
-    firOpBuilder.setInsertionPointAfter(ifOp);
-  }
 }
 
 static fir::GlobalOp globalInitialization(lower::AbstractConverter &converter,
diff --git a/flang/test/Lower/OpenMP/collapse-imperfect-nest.f90 b/flang/test/Lower/OpenMP/collapse-imperfect-nest.f90
index e963049090f49..087d7b9857aa7 100644
--- a/flang/test/Lower/OpenMP/collapse-imperfect-nest.f90
+++ b/flang/test/Lower/OpenMP/collapse-imperfect-nest.f90
@@ -1,6 +1,8 @@
 ! Test lowering of imperfectly nested collapse loops (CLN relaxation).
-! Intervening code is guarded by IV comparisons to restore correct
-! execution frequency and ordering within the flat omp.loop_nest body.
+! Intervening code is emitted unguarded into the flat omp.loop_nest body, so it
+! executes once per collapsed logical iteration. OpenMP 6.0 leaves the count
+! unspecified between once per enclosing iteration and once per logical
+! iteration.
 
 ! RUN: %flang_fc1 -fopenmp -emit-hlfir %s -o - | FileCheck %s
 
@@ -22,23 +24,18 @@ subroutine collapse2_imperfect(n, x)
 
 ! CHECK:       omp.wsloop
 ! CHECK-NEXT:    omp.loop_nest (%[[I:.*]], %[[J:.*]]) : i32 =
-! CHECK-SAME:      (%{{.*}}, %[[LB_J:.*]]) to
 ! CHECK:           hlfir.assign %[[I]]
 ! CHECK:           hlfir.assign %[[J]]
-! Guard: j == lower_bound (before code executes once per i)
-! CHECK:           %[[CMP:.*]] = arith.cmpi eq, %[[J]], %[[LB_J]] : i32
-! CHECK:           fir.if %[[CMP]] {
 ! Intervening code: x = x + 1
-! CHECK:             %[[X1:.*]] = fir.load %{{.*}} : !fir.ref<i32>
-! CHECK:             %[[C1:.*]] = arith.constant 1 : i32
-! CHECK:             %[[ADD1:.*]] = arith.addi %[[X1]], %[[C1]] : i32
-! CHECK:             hlfir.assign %[[ADD1]] to %{{.*}} : i32, !fir.ref<i32>
-! CHECK:           }
+! CHECK:           %[[X1:.*]] = fir.load %{{.*}} : !fir.ref<i32>
+! CHECK:           %[[C1:.*]] = arith.constant 1 : i32
+! CHECK:           %[[ADD1:.*]] = arith.addi %[[X1]], %[[C1]] : i32
+! CHECK:           hlfir.assign %[[ADD1]] to %[[XADDR:.*]] : i32, !fir.ref<i32>
 ! Innermost body: x = x + j
-! CHECK:           %[[X2:.*]] = fir.load %{{.*}} : !fir.ref<i32>
+! CHECK:           %[[X2:.*]] = fir.load %[[XADDR]] : !fir.ref<i32>
 ! CHECK:           %[[JVAL:.*]] = fir.load %{{.*}} : !fir.ref<i32>
 ! CHECK:           %[[ADD2:.*]] = arith.addi %[[X2]], %[[JVAL]] : i32
-! CHECK:           hlfir.assign %[[ADD2]] to %{{.*}} : i32, !fir.ref<i32>
+! CHECK:           hlfir.assign %[[ADD2]] to %[[XADDR]] : i32, !fir.ref<i32>
 ! CHECK:           omp.yield
 
 ! CHECK-LABEL: func.func @_QPcollapse3_imperfect
@@ -62,37 +59,21 @@ subroutine collapse3_imperfect(n, x)
 
 ! CHECK:       omp.wsloop
 ! CHECK-NEXT:    omp.loop_nest (%[[I3:.*]], %[[J3:.*]], %[[K3:.*]]) : i32 =
-! CHECK-SAME:      (%{{.*}}, %[[LB_J3:.*]], %[[LB_K3:.*]]) to
 ! CHECK:           hlfir.assign %[[I3]]
 ! CHECK:           hlfir.assign %[[J3]]
 ! CHECK:           hlfir.assign %[[K3]]
-! Guard: j == lb_j AND k == lb_k (level 0 before code, once per i)
-! CHECK:           %[[CMP_J:.*]] = arith.cmpi eq, %[[J3]], %[[LB_J3]] : i32
-! CHECK:           %[[CMP_K1:.*]] = arith.cmpi eq, %[[K3]], %[[LB_K3]] : i32
-! CHECK:           %[[AND1:.*]] = arith.andi %[[CMP_J]], %[[CMP_K1]] : i1
-! CHECK:           fir.if %[[AND1]] {
-! Intervening code at level 0: x = x + i
-! CHECK:             %[[XI:.*]] = fir.load %{{.*}} : !fir.ref<i32>
-! CHECK:             %[[IVAL:.*]] = fir.load %{{.*}} : !fir.ref<i32>
-! CHECK:             %[[ADDI:.*]] = arith.addi %[[XI]], %[[IVAL]] : i32
-! CHECK:             hlfir.assign %[[ADDI]] to %{{.*}} : i32, !fir.ref<i32>
-! CHECK:           }
-! Guard: k == lb_k (level 1 before code, once per (i,j))
-! CHECK:           %[[CMP_K2:.*]] = arith.cmpi eq, %[[K3]], %[[LB_K3]] : i32
-! CHECK:           fir.if %[[CMP_K2]] {
-! Intervening code at level 1: x = x + j
-! CHECK:             %[[XJ:.*]] = fir.load %{{.*}} : !fir.ref<i32>
-! CHECK:             %[[JVAL3:.*]] = fir.load %{{.*}} : !fir.ref<i32>
-! CHECK:             %[[ADDJ:.*]] = arith.addi %[[XJ]], %[[JVAL3]] : i32
-! CHECK:             hlfir.assign %[[ADDJ]] to %{{.*}} : i32, !fir.ref<i32>
-! CHECK:           }
+! Level 0 intervening code: x = x + i
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
+! Level 1 intervening code: x = x + j
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
 ! Innermost body: x = x + k
-! CHECK:           %[[XK:.*]] = fir.load %{{.*}} : !fir.ref<i32>
-! CHECK:           %[[KVAL:.*]] = fir.load %{{.*}} : !fir.ref<i32>
-! CHECK:           %[[ADDK:.*]] = arith.addi %[[XK]], %[[KVAL]] : i32
-! CHECK:           hlfir.assign %[[ADDK]] to %{{.*}} : i32, !fir.ref<i32>
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
 ! CHECK:           omp.yield
 
+! Intervening code on both sides of the inner loop, emitted in source order.
 ! CHECK-LABEL: func.func @_QPcollapse2_both_sides
 subroutine collapse2_both_sides(n, x)
   integer, intent(in) :: n
@@ -112,29 +93,16 @@ subroutine collapse2_both_sides(n, x)
 
 ! CHECK:       omp.simd
 ! CHECK-NEXT:    omp.loop_nest (%[[I4:.*]], %[[J4:.*]]) : i32 =
-! CHECK-SAME:      (%{{[^)]*}}, %[[LB_J4:[^)]*]]) to (%{{[^)]*}}, %[[UB_J4:[^)]*]])
 ! CHECK:           hlfir.assign %[[I4]]
 ! CHECK:           hlfir.assign %[[J4]]
-! Guard: j == lower_bound (before code)
-! CHECK:           %[[CMP_B:.*]] = arith.cmpi eq, %[[J4]], %[[LB_J4]] : i32
-! CHECK:           fir.if %[[CMP_B]] {
-! Intervening code before inner loop: x = x + 1
-! CHECK:             %[[XB:.*]] = fir.load %{{.*}} : !fir.ref<i32>
-! CHECK:             %[[CB:.*]] = arith.constant 1 : i32
-! CHECK:             %[[ADDB:.*]] = arith.addi %[[XB]], %[[CB]] : i32
-! CHECK:             hlfir.assign %[[ADDB]] to %{{.*}} : i32, !fir.ref<i32>
-! CHECK:           }
+! Before the inner loop: x = x + 1
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
 ! Innermost body: x = x + j
-! CHECK:           %[[XIN:.*]] = fir.load %{{.*}} : !fir.ref<i32>
-! CHECK:           %[[JIN:.*]] = fir.load %{{.*}} : !fir.ref<i32>
-! CHECK:           %[[ADDIN:.*]] = arith.addi %[[XIN]], %[[JIN]] : i32
-! CHECK:           hlfir.assign %[[ADDIN]] to %{{.*}} : i32, !fir.ref<i32>
-! Guard: j == upper_bound (after code)
-! CHECK:           %[[CMP_A:.*]] = arith.cmpi eq, %[[J4]], %[[UB_J4]] : i32
-! CHECK:           fir.if %[[CMP_A]] {
-! Intervening code after inner loop: call ext_sub(x)
-! CHECK:             fir.call @_QPext_sub
-! CHECK:           }
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
+! After the inner loop: call ext_sub(x)
+! CHECK:           fir.call @_QPext_sub
 ! CHECK:           omp.yield
 
 ! Test collapse(3) with both before and after code at multiple levels.
@@ -159,50 +127,28 @@ subroutine collapse3_both_sides(n, x)
   !$omp end do
 end subroutine
 
+! Emission order is: before code outermost-to-innermost, innermost body, then
+! after code innermost-to-outermost.
 ! CHECK:       omp.wsloop
 ! CHECK-NEXT:    omp.loop_nest (%[[I:.*]], %[[J:.*]], %[[K:.*]]) : i32 =
-! CHECK-SAME:      (%{{[^)]*}}, %[[LB_J:[^)]*]], %[[LB_K:[^)]*]]) to (%{{[^)]*}}, %[[UB_J:[^)]*]], %[[UB_K:[^)]*]])
 ! CHECK:           hlfir.assign %[[I]]
 ! CHECK:           hlfir.assign %[[J]]
 ! CHECK:           hlfir.assign %[[K]]
-!
-! --- "before" guards (outermost level first) ---
-!
-! Guard level 0 before: j == lb_j AND k == lb_k (once per i)
-! CHECK:           %[[CJ1:.*]] = arith.cmpi eq, %[[J]], %[[LB_J]] : i32
-! CHECK:           %[[CK1:.*]] = arith.cmpi eq, %[[K]], %[[LB_K]] : i32
-! CHECK:           %[[AND1:.*]] = arith.andi %[[CJ1]], %[[CK1]] : i1
-! CHECK:           fir.if %[[AND1]] {
-! CHECK:             arith.addi
-! CHECK:             hlfir.assign
-! CHECK:           }
-! Guard level 1 before: k == lb_k (once per (i,j))
-! CHECK:           %[[CK2:.*]] = arith.cmpi eq, %[[K]], %[[LB_K]] : i32
-! CHECK:           fir.if %[[CK2]] {
-! CHECK:             arith.addi
-! CHECK:             hlfir.assign
-! CHECK:           }
-!
-! --- innermost body: x = x + k ---
+! Level 0 before: x = x + i
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
+! Level 1 before: x = x + j
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
+! Innermost body: x = x + k
 ! CHECK:           arith.addi
 ! CHECK:           hlfir.assign
-!
-! --- "after" guards (innermost level first) ---
-!
-! Guard level 1 after: k == ub_k (once per (i,j))
-! CHECK:           %[[CK3:.*]] = arith.cmpi eq, %[[K]], %[[UB_K]] : i32
-! CHECK:           fir.if %[[CK3]] {
-! CHECK:             arith.subi
-! CHECK:             hlfir.assign
-! CHECK:           }
-! Guard level 0 after: j == ub_j AND k == ub_k (once per i)
-! CHECK:           %[[CJ2:.*]] = arith.cmpi eq, %[[J]], %[[UB_J]] : i32
-! CHECK:           %[[CK4:.*]] = arith.cmpi eq, %[[K]], %[[UB_K]] : i32
-! CHECK:           %[[AND2:.*]] = arith.andi %[[CJ2]], %[[CK4]] : i1
-! CHECK:           fir.if %[[AND2]] {
-! CHECK:             arith.subi
-! CHECK:             hlfir.assign
-! CHECK:           }
+! Level 1 after: x = x - j
+! CHECK:           arith.subi
+! CHECK:           hlfir.assign
+! Level 0 after: x = x - i
+! CHECK:           arith.subi
+! CHECK:           hlfir.assign
 ! CHECK:           omp.yield
 
 ! Test collapse(4) with imperfect nesting at some levels and perfectly nested
@@ -232,56 +178,26 @@ subroutine collapse4_mixed(n, x)
 
 ! CHECK:       omp.wsloop
 ! CHECK-NEXT:    omp.loop_nest (%[[I:.*]], %[[J:.*]], %[[K:.*]], %[[L:.*]]) : i32 =
-! CHECK-SAME:      (%{{[^)]*}}, %[[LB_J:[^)]*]], %[[LB_K:[^)]*]], %[[LB_L:[^)]*]]) to (%{{[^)]*}}, %[[UB_J:[^)]*]], %[[UB_K:[^)]*]], %[[UB_L:[^)]*]])
 ! CHECK:           hlfir.assign %[[I]]
 ! CHECK:           hlfir.assign %[[J]]
 ! CHECK:           hlfir.assign %[[K]]
 ! CHECK:           hlfir.assign %[[L]]
-!
-! --- "before" guards ---
-!
-! Guard level 0 before: j == lb_j AND k == lb_k AND l == lb_l (once per i)
-! CHECK:           %[[CJ1:.*]] = arith.cmpi eq, %[[J]], %[[LB_J]] : i32
-! CHECK:           %[[CK1:.*]] = arith.cmpi eq, %[[K]], %[[LB_K]] : i32
-! CHECK:           %[[A1:.*]] = arith.andi %[[CJ1]], %[[CK1]] : i1
-! CHECK:           %[[CL1:.*]] = arith.cmpi eq, %[[L]], %[[LB_L]] : i32
-! CHECK:           %[[A2:.*]] = arith.andi %[[A1]], %[[CL1]] : i1
-! CHECK:           fir.if %[[A2]] {
-! CHECK:             arith.addi
-! CHECK:             hlfir.assign
-! CHECK:           }
-! Guard level 1 before: k == lb_k AND l == lb_l (once per (i,j))
-! CHECK:           %[[CK2:.*]] = arith.cmpi eq, %[[K]], %[[LB_K]] : i32
-! CHECK:           %[[CL2:.*]] = arith.cmpi eq, %[[L]], %[[LB_L]] : i32
-! CHECK:           %[[A3:.*]] = arith.andi %[[CK2]], %[[CL2]] : i1
-! CHECK:           fir.if %[[A3]] {
-! CHECK:             arith.addi
-! CHECK:             hlfir.assign
-! CHECK:           }
-! Level 2 (k->l) is perfectly nested: no guard emitted.
-!
-! --- innermost body: x = x + l ---
+! Level 0 before: x = x + i
 ! CHECK:           arith.addi
 ! CHECK:           hlfir.assign
-!
-! --- "after" guards (innermost first) ---
-!
-! Level 2 after: empty (perfectly nested), no guard emitted.
-! Level 1 after: empty, no guard emitted.
-! Guard level 0 after: j == ub_j AND k == ub_k AND l == ub_l (once per i)
-! CHECK:           %[[CJ2:.*]] = arith.cmpi eq, %[[J]], %[[UB_J]] : i32
-! CHECK:           %[[CK3:.*]] = arith.cmpi eq, %[[K]], %[[UB_K]] : i32
-! CHECK:           %[[A4:.*]] = arith.andi %[[CJ2]], %[[CK3]] : i1
-! CHECK:           %[[CL3:.*]] = arith.cmpi eq, %[[L]], %[[UB_L]] : i32
-! CHECK:           %[[A5:.*]] = arith.andi %[[A4]], %[[CL3]] : i1
-! CHECK:           fir.if %[[A5]] {
-! CHECK:             arith.subi
-! CHECK:             hlfir.assign
-! CHECK:           }
+! Level 1 before: x = x + j
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
+! Innermost body: x = x + l
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
+! Level 0 after: x = x - i
+! CHECK:           arith.subi
+! CHECK:           hlfir.assign
 ! CHECK:           omp.yield
 
 ! Test collapse(2) with only after-code (no before-code). Exercises the path
-! where levels[i].before.empty() is true and the "before" loop is entirely skipped.
+! where levels[i].before is empty.
 ! CHECK-LABEL: func.func @_QPcollapse2_after_only
 subroutine collapse2_after_only(n, x)
   integer, intent(in) :: n
@@ -300,23 +216,17 @@ subroutine collapse2_after_only(n, x)
 
 ! CHECK:       omp.wsloop
 ! CHECK-NEXT:    omp.loop_nest (%[[I:.*]], %[[J:.*]]) : i32 =
-! CHECK-SAME:      (%{{[^)]*}}, %{{[^)]*}}) to (%{{[^)]*}}, %[[UB_J:[^)]*]])
 ! CHECK:           hlfir.assign %[[I]]
 ! CHECK:           hlfir.assign %[[J]]
-! No "before" guard emitted (level 0 before is empty).
 ! Innermost body: x = x + j
 ! CHECK:           arith.addi
 ! CHECK:           hlfir.assign
-! Guard: j == upper_bound (after code)
-! CHECK:           %[[CMP:.*]] = arith.cmpi eq, %[[J]], %[[UB_J]] : i32
-! CHECK:           fir.if %[[CMP]] {
-! CHECK:             arith.subi
-! CHECK:             hlfir.assign
-! CHECK:           }
+! After code: x = x - i
+! CHECK:           arith.subi
+! CHECK:           hlfir.assign
 ! CHECK:           omp.yield
 
-! Test collapse(2) with multiple statements inside a single guard. Verifies
-! that all evals in level.before land inside the same fir.if block.
+! Test collapse(2) with multiple intervening statements at one level.
 ! CHECK-LABEL: func.func @_QPcollapse2_multi_stmt
 subroutine collapse2_multi_stmt(n, x)
   integer, intent(in) :: n
@@ -336,60 +246,54 @@ subroutine collapse2_multi_stmt(n, x)
 
 ! CHECK:       omp.wsloop
 ! CHECK-NEXT:    omp.loop_nest (%[[I:.*]], %[[J:.*]]) : i32 =
-! CHECK-SAME:      (%{{[^)]*}}, %[[LB_J:[^)]*]]) to
 ! CHECK:           hlfir.assign %[[I]]
 ! CHECK:           hlfir.assign %[[J]]
-! Guard: j == lower_bound (before code, multiple statements in one guard)
-! CHECK:           %[[CMP:.*]] = arith.cmpi eq, %[[J]], %[[LB_J]] : i32
-! CHECK:           fir.if %[[CMP]] {
 ! First intervening statement: x = x + 1
-! CHECK:             arith.addi
-! CHECK:             hlfir.assign
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
 ! Second intervening statement: x = x + i
-! CHECK:             arith.addi
-! CHECK:             hlfir.assign
-! CHECK:           }
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
 ! Innermost body: x = x + j
 ! CHECK:           arith.addi
 ! CHECK:           hlfir.assign
 ! CHECK:           omp.yield
 
-! Test collapse(2) with non-unit lower bound on inner loop. Verifies the guard
-! compares against the actual loop lower bound operand (3, not 1).
-! CHECK-LABEL: func.func @_QPcollapse2_nonunit_lb
-subroutine collapse2_nonunit_lb(n, x)
-  integer, intent(in) :: n
+! Test collapse(2) with a non-unit lower bound and a runtime step on the inner
+! loop.
+! CHECK-LABEL: func.func @_QPcollapse2_runtime_step
+subroutine collapse2_runtime_step(n, s, x)
+  integer, intent(in) :: n, s
   integer, intent(inout) :: x
   integer :: i, j
 
   !$omp do collapse(2)
   do i = 1, n
     x = x + i
-    do j = 3, n
+    do j = 3, n, s
       x = x + j
     end do
+    x = x - i
   end do
   !$omp end do
 end subroutine
 
 ! CHECK:       omp.wsloop
 ! CHECK-NEXT:    omp.loop_nest (%[[I:.*]], %[[J:.*]]) : i32 =
-! CHECK-SAME:      (%{{[^)]*}}, %[[LB_J:[^)]*]]) to
 ! CHECK:           hlfir.assign %[[I]]
 ! CHECK:           hlfir.assign %[[J]]
-! Guard: j == lb_j (lb_j is 3, not 1)
-! CHECK:           %[[CMP:.*]] = arith.cmpi eq, %[[J]], %[[LB_J]] : i32
-! CHECK:           fir.if %[[CMP]] {
-! CHECK:             arith.addi
-! CHECK:             hlfir.assign
-! CHECK:           }
+! Before code: x = x + i
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
 ! Innermost body: x = x + j
 ! CHECK:           arith.addi
 ! CHECK:           hlfir.assign
+! After code: x = x - i
+! CHECK:           arith.subi
+! CHECK:           hlfir.assign
 ! CHECK:           omp.yield
 
 ! Test collapse(3) with after-only at level 0 and before-only at level 1.
-! Exercises the independent skip logic at each level in both emission loops.
 ! CHECK-LABEL: func.func @_QPcollapse3_mixed_sides
 subroutine collapse3_mixed_sides(n, x)
   integer, intent(in) :: n
@@ -411,238 +315,177 @@ subroutine collapse3_mixed_sides(n, x)
 
 ! CHECK:       omp.wsloop
 ! CHECK-NEXT:    omp.loop_nest (%[[I:.*]], %[[J:.*]], %[[K:.*]]) : i32 =
-! CHECK-SAME:      (%{{[^)]*}}, %{{[^)]*}}, %[[LB_K:[^)]*]]) to (%{{[^)]*}}, %[[UB_J:[^)]*]], %[[UB_K:[^)]*]])
 ! CHECK:           hlfir.assign %[[I]]
 ! CHECK:           hlfir.assign %[[J]]
 ! CHECK:           hlfir.assign %[[K]]
-! Level 0 before: empty (skipped).
-! Guard level 1 before: k == lb_k (once per (i,j))
-! CHECK:           %[[CK:.*]] = arith.cmpi eq, %[[K]], %[[LB_K]] : i32
-! CHECK:           fir.if %[[CK]] {
-! CHECK:             arith.addi
-! CHECK:             hlfir.assign
-! CHECK:           }
+! Level 1 before: x = x + j
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
 ! Innermost body: x = x + k
 ! CHECK:           arith.addi
 ! CHECK:           hlfir.assign
-! Level 1 after: empty (skipped).
-! Guard level 0 after: j == ub_j AND k == ub_k (once per i)
-! CHECK:           %[[CJ:.*]] = arith.cmpi eq, %[[J]], %[[UB_J]] : i32
-! CHECK:           %[[CK2:.*]] = arith.cmpi eq, %[[K]], %[[UB_K]] : i32
-! CHECK:           %[[AND:.*]] = arith.andi %[[CJ]], %[[CK2]] : i1
-! CHECK:           fir.if %[[AND]] {
-! CHECK:             arith.subi
-! CHECK:             hlfir.assign
-! CHECK:           }
+! Level 0 after: x = x - i
+! CHECK:           arith.subi
+! CHECK:           hlfir.assign
 ! CHECK:           omp.yield
 
-! Test collapse(2) with non-unit positive step and after-code.
-! The after guard must compare iv against the last executed value
-! (lb + ((ub - lb) / step) * step), not the upper bound directly.
-! For do j = 1, 10, 4: last_iv = 1 + ((10-1)/4)*4 = 1 + 8 = 9.
-! CHECK-LABEL: func.func @_QPcollapse2_nonunit_step_after
-subroutine collapse2_nonunit_step_after(n, x)
+! OpenMP 6.0 6.4.3 requires each collapsed loop's iteration variable to hold the
+! value it would have in the unassociated nest, so "after" code reads j from the
+! current logical iteration rather than a synthesized Fortran terminal value.
+! CHECK-LABEL: func.func @_QPcollapse2_after_reads_inner
+subroutine collapse2_after_reads_inner(n, x)
   integer, intent(in) :: n
   integer, intent(inout) :: x
   integer :: i, j
 
   !$omp do collapse(2)
   do i = 1, n
-    do j = 1, 10, 4
-      x = x + j
+    do j = 1, n
+      x = x + 1
     end do
-    x = x - i
+    x = x + j
   end do
   !$omp end do
 end subroutine
 
 ! CHECK:       omp.wsloop
 ! CHECK-NEXT:    omp.loop_nest (%[[I:.*]], %[[J:.*]]) : i32 =
-! CHECK-SAME:      (%{{[^)]*}}, %[[LB_J:[^)]*]]) to (%{{[^)]*}}, %[[UB_J:[^)]*]]) inclusive step (%{{[^)]*}}, %[[STEP_J:[^)]*]])
 ! CHECK:           hlfir.assign %[[I]]
-! CHECK:           hlfir.assign %[[J]]
-! Innermost body: x = x + j
+! CHECK:           hlfir.assign %[[J]] to %[[J_ADDR:.*]] : i32, !fir.ref<i32>
+! Innermost body: x = x + 1
 ! CHECK:           arith.addi
 ! CHECK:           hlfir.assign
-! After guard: compute last_iv = lb + ((ub - lb) / step) * step
-! CHECK:           %[[RANGE:.*]] = arith.subi %[[UB_J]], %[[LB_J]] : i32
-! CHECK:           %[[DIV:.*]] = arith.divsi %[[RANGE]], %[[STEP_J]] : i32
-! CHECK:           %[[MUL:.*]] = arith.muli %[[DIV]], %[[STEP_J]] : i32
-! CHECK:           %[[LAST:.*]] = arith.addi %[[LB_J]], %[[MUL]] : i32
-! CHECK:           %[[CMP:.*]] = arith.cmpi eq, %[[J]], %[[LAST]] : i32
-! CHECK:           fir.if %[[CMP]] {
-! CHECK:             arith.subi
-! CHECK:             hlfir.assign
-! CHECK:           }
+! After code reads j directly, with no restore of a terminal value.
+! CHECK:           %[[XLD:.*]] = fir.load %{{.*}} : !fir.ref<i32>
+! CHECK:           %[[JLD:.*]] = fir.load %[[J_ADDR]] : !fir.ref<i32>
+! CHECK:           arith.addi %[[XLD]], %[[JLD]] : i32
+! CHECK:           hlfir.assign
 ! CHECK:           omp.yield
 
-! Test collapse(2) with negative step and after-code.
-! For do j = 10, 1, -4: last_iv = 10 + ((1-10)/(-4))*(-4) = 10 + (2*-4) = 2.
-! CHECK-LABEL: func.func @_QPcollapse2_negative_step_after
-subroutine collapse2_negative_step_after(n, x)
+! Labeled DO form: the terminating CONTINUE survives canonicalization as a
+! sibling of the inner loop.
+! CHECK-LABEL: func.func @_QPcollapse2_labeled_do
+subroutine collapse2_labeled_do(n, x)
   integer, intent(in) :: n
   integer, intent(inout) :: x
   integer :: i, j
 
   !$omp do collapse(2)
-  do i = 1, n
-    do j = 10, 1, -4
+  do 10 i = 1, n
+    x = x + i
+    do 20 j = 1, n
       x = x + j
-    end do
+20  continue
     x = x - i
-  end do
+10 continue
   !$omp end do
 end subroutine
 
 ! CHECK:       omp.wsloop
 ! CHECK-NEXT:    omp.loop_nest (%[[I:.*]], %[[J:.*]]) : i32 =
-! CHECK-SAME:      (%{{[^)]*}}, %[[LB_J:[^)]*]]) to (%{{[^)]*}}, %[[UB_J:[^)]*]]) inclusive step (%{{[^)]*}}, %[[STEP_J:[^)]*]])
 ! CHECK:           hlfir.assign %[[I]]
 ! CHECK:           hlfir.assign %[[J]]
+! Before code: x = x + i
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
 ! Innermost body: x = x + j
 ! CHECK:           arith.addi
 ! CHECK:           hlfir.assign
-! After guard: compute last_iv for negative step
-! CHECK:           %[[RANGE:.*]] = arith.subi %[[UB_J]], %[[LB_J]] : i32
-! CHECK:           %[[DIV:.*]] = arith.divsi %[[RANGE]], %[[STEP_J]] : i32
-! CHECK:           %[[MUL:.*]] = arith.muli %[[DIV]], %[[STEP_J]] : i32
-! CHECK:           %[[LAST:.*]] = arith.addi %[[LB_J]], %[[MUL]] : i32
-! CHECK:           %[[CMP:.*]] = arith.cmpi eq, %[[J]], %[[LAST]] : i32
-! CHECK:           fir.if %[[CMP]] {
-! CHECK:             arith.subi
-! CHECK:             hlfir.assign
-! CHECK:           }
+! After code: x = x - i
+! CHECK:           arith.subi
+! CHECK:           hlfir.assign
 ! CHECK:           omp.yield
 
-! Test collapse(3) with non-unit step on the middle loop (not innermost).
-! For do j = 1, n, 3: last_iv = lb + ((ub - lb) / step) * step (runtime).
-! CHECK-LABEL: func.func @_QPcollapse3_nonunit_step_middle
-subroutine collapse3_nonunit_step_middle(n, x)
+! A compiler directive between the loops is transparent to perfect nesting.
+! CHECK-LABEL: func.func @_QPcollapse2_compiler_directive
+subroutine collapse2_compiler_directive(n, x)
   integer, intent(in) :: n
   integer, intent(inout) :: x
-  integer :: i, j, k
+  integer :: i, j
 
-  !$omp do collapse(3)
+  !$omp do collapse(2)
   do i = 1, n
-    do j = 1, n, 3
+    x = x + i
+    !dir$ vector always
+    do j = 1, n
       x = x + j
-      do k = 1, n
-        x = x + k
-      end do
     end do
-    x = x - i
   end do
   !$omp end do
 end subroutine
 
 ! CHECK:       omp.wsloop
-! CHECK-NEXT:    omp.loop_nest (%[[I:.*]], %[[J:.*]], %[[K:.*]]) : i32 =
-! CHECK-SAME:      (%{{[^)]*}}, %[[LB_J:[^)]*]], %[[LB_K:[^)]*]]) to (%{{[^)]*}}, %[[UB_J:[^)]*]], %[[UB_K:[^)]*]]) inclusive step (%{{[^)]*}}, %[[STEP_J:[^)]*]], %{{[^)]*}})
+! CHECK-NEXT:    omp.loop_nest (%[[I:.*]], %[[J:.*]]) : i32 =
 ! CHECK:           hlfir.assign %[[I]]
 ! CHECK:           hlfir.assign %[[J]]
-! CHECK:           hlfir.assign %[[K]]
-! Guard level 1 before: k == lb_k (once per (i,j))
-! CHECK:           %[[CK1:.*]] = arith.cmpi eq, %[[K]], %[[LB_K]] : i32
-! CHECK:           fir.if %[[CK1]] {
-! CHECK:             arith.addi
-! CHECK:             hlfir.assign
-! CHECK:           }
-! Innermost body: x = x + k
+! Before code: x = x + i
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
+! Innermost body: x = x + j
 ! CHECK:           arith.addi
 ! CHECK:           hlfir.assign
-! Guard level 0 after: must compute last_iv for j (non-unit step) AND k == ub_k
-! CHECK:           %[[RANGE:.*]] = arith.subi %[[UB_J]], %[[LB_J]] : i32
-! CHECK:           %[[DIV:.*]] = arith.divsi %[[RANGE]], %[[STEP_J]] : i32
-! CHECK:           %[[MUL:.*]] = arith.muli %[[DIV]], %[[STEP_J]] : i32
-! CHECK:           %[[LASTJ:.*]] = arith.addi %[[LB_J]], %[[MUL]] : i32
-! CHECK:           %[[CJ:.*]] = arith.cmpi eq, %[[J]], %[[LASTJ]] : i32
-! CHECK:           %[[CK2:.*]] = arith.cmpi eq, %[[K]], %[[UB_K]] : i32
-! CHECK:           %[[AND:.*]] = arith.andi %[[CJ]], %[[CK2]] : i1
-! CHECK:           fir.if %[[AND]] {
-! CHECK:             arith.subi
-! CHECK:             hlfir.assign
-! CHECK:           }
 ! CHECK:           omp.yield
 
-! Test collapse(2) with a dynamic (runtime) step value.
-! The step is not a compile-time constant, so the last_iv computation
-! cannot be folded away and must remain as arith ops in the IR.
-! CHECK-LABEL: func.func @_QPcollapse2_dynamic_step_after
-subroutine collapse2_dynamic_step_after(n, s, x)
-  integer, intent(in) :: n, s
+! Intervening code need not be straight-line: a block IF is valid intervening
+! code and brings its own control flow into the loop_nest body.
+! CHECK-LABEL: func.func @_QPcollapse2_if_construct
+subroutine collapse2_if_construct(n, x)
+  integer, intent(in) :: n
   integer, intent(inout) :: x
   integer :: i, j
 
   !$omp do collapse(2)
   do i = 1, n
-    do j = 1, n, s
+    if (i > 2) then
+      x = x + i
+    end if
+    do j = 1, n
       x = x + j
     end do
-    x = x - i
   end do
   !$omp end do
 end subroutine
 
 ! CHECK:       omp.wsloop
 ! CHECK-NEXT:    omp.loop_nest (%[[I:.*]], %[[J:.*]]) : i32 =
-! CHECK-SAME:      (%{{[^)]*}}, %[[LB_J:[^)]*]]) to (%{{[^)]*}}, %[[UB_J:[^)]*]]) inclusive step (%{{[^)]*}}, %[[STEP_J:[^)]*]])
 ! CHECK:           hlfir.assign %[[I]]
 ! CHECK:           hlfir.assign %[[J]]
+! Before code: if (i > 2) x = x + i
+! CHECK:           arith.cmpi sgt
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
 ! Innermost body: x = x + j
 ! CHECK:           arith.addi
 ! CHECK:           hlfir.assign
-! After guard: dynamic step forces last_iv computation to stay in IR
-! CHECK:           %[[RANGE:.*]] = arith.subi %[[UB_J]], %[[LB_J]] : i32
-! CHECK:           %[[DIV:.*]] = arith.divsi %[[RANGE]], %[[STEP_J]] : i32
-! CHECK:           %[[MUL:.*]] = arith.muli %[[DIV]], %[[STEP_J]] : i32
-! CHECK:           %[[LAST:.*]] = arith.addi %[[LB_J]], %[[MUL]] : i32
-! CHECK:           %[[CMP:.*]] = arith.cmpi eq, %[[J]], %[[LAST]] : i32
-! CHECK:           fir.if %[[CMP]] {
-! CHECK:             arith.subi
-! CHECK:             hlfir.assign
-! CHECK:           }
 ! CHECK:           omp.yield
 
-! Test that "after" code reading an inner DO variable sees the Fortran
-! terminal value (lb + tripcount*step, i.e. one past the last executed
-! value), not the last executed value the flattened nest leaves it at.
-! The after guard must restore j to ub + step before the after code runs.
-! CHECK-LABEL: func.func @_QPcollapse2_after_reads_inner
-subroutine collapse2_after_reads_inner(n, x)
-  integer, intent(in) :: n
+! Intervening code inside a target region, whose collapsed bounds are host_eval
+! block arguments of the enclosing omp.target.
+! CHECK-LABEL: func.func @_QPcollapse2_target_intervening
+subroutine collapse2_target_intervening(n, m, x)
+  integer, intent(in) :: n, m
   integer, intent(inout) :: x
   integer :: i, j
 
-  !$omp do collapse(2)
+  !$omp target teams distribute parallel do collapse(2) map(tofrom:x)
   do i = 1, n
-    do j = 1, n
+    do j = 1, m
       x = x + 1
     end do
     x = x + j
   end do
-  !$omp end do
 end subroutine
 
-! CHECK:       omp.wsloop
-! CHECK-NEXT:    omp.loop_nest (%[[I:.*]], %[[J:.*]]) : i32 =
-! CHECK-SAME:      (%{{[^)]*}}, %{{[^)]*}}) to (%{{[^)]*}}, %[[UB_J:[^)]*]]) inclusive step (%{{[^)]*}}, %[[STEP_J:[^)]*]])
+! CHECK:       omp.target
+! CHECK-SAME:    host_eval(
+! CHECK:         omp.loop_nest (%[[I:.*]], %[[J:.*]]) : i32 =
 ! CHECK:           hlfir.assign %[[I]]
-! Capture j's storage from the loop-variable store so the restore and the
-! after-code read can be tied to the same address.
 ! CHECK:           hlfir.assign %[[J]] to %[[J_ADDR:.*]] : i32, !fir.ref<i32>
 ! Innermost body: x = x + 1
 ! CHECK:           arith.addi
 ! CHECK:           hlfir.assign
-! Guard: j == ub_j (unit-step fast path compares against the upper bound)
-! CHECK:           %[[CMP:.*]] = arith.cmpi eq, %[[J]], %[[UB_J]] : i32
-! CHECK:           fir.if %[[CMP]] {
-! Restore j's storage to its Fortran terminal value (ub + step).
-! CHECK:             %[[TERM:.*]] = arith.addi %[[UB_J]], %[[STEP_J]] : i32
-! CHECK:             hlfir.assign %[[TERM]] to %[[J_ADDR]] : i32, !fir.ref<i32>
-! After code reads the restored j: x = x + j.
-! CHECK:             %[[XLD:.*]] = fir.load %{{.*}} : !fir.ref<i32>
-! CHECK:             %[[JLD:.*]] = fir.load %[[J_ADDR]] : !fir.ref<i32>
-! CHECK:             arith.addi %[[XLD]], %[[JLD]] : i32
-! CHECK:             hlfir.assign
-! CHECK:           }
+! Intervening code: x = x + j
+! CHECK:           %[[JLD:.*]] = fir.load %[[J_ADDR]] : !fir.ref<i32>
+! CHECK:           arith.addi
+! CHECK:           hlfir.assign
 ! CHECK:           omp.yield
diff --git a/flang/test/Lower/OpenMP/collapse-target-intervening-todo.f90 b/flang/test/Lower/OpenMP/collapse-target-intervening-todo.f90
deleted file mode 100644
index 48072802db9ff..0000000000000
--- a/flang/test/Lower/OpenMP/collapse-target-intervening-todo.f90
+++ /dev/null
@@ -1,28 +0,0 @@
-! Regression test: a collapsed imperfect loop nest with genuine intervening
-! code, whose loop bounds are host-evaluated for an enclosing omp.target SPMD
-! region, is not yet supported and must be diagnosed cleanly rather than
-! producing IR that fails the omp.target verifier.
-!
-! The intervening statement "x = x + j" runs at intermediate nest level. Its
-! guard and terminal-IV restoration would perform arith.cmpi/arith.addi on the
-! collapsed loop bounds, which are omp.target host_eval block arguments in this
-! context -- an illegal use. Until that path is implemented, emit a "not yet
-! implemented" message.
-
-! RUN: not %flang_fc1 -emit-hlfir -fopenmp %s 2>&1 | FileCheck %s
-
-! CHECK: not yet implemented: collapsed loop nest with intervening code whose loop bounds are evaluated on the host for an enclosing 'target' region
-subroutine repro(n, m, x)
-  implicit none
-  integer, intent(in) :: n, m
-  integer, intent(inout) :: x
-  integer :: i, j
-
-  !$omp target teams distribute parallel do collapse(2) map(tofrom:x)
-  do i = 1, n
-    do j = 1, m
-      x = x + 1
-    end do
-    x = x + j
-  end do
-end subroutine
diff --git a/offload/test/offloading/fortran/target-teams-distribute-collapse-intervening.f90 b/offload/test/offloading/fortran/target-teams-distribute-collapse-intervening.f90
new file mode 100644
index 0000000000000..7604d711505e0
--- /dev/null
+++ b/offload/test/offloading/fortran/target-teams-distribute-collapse-intervening.f90
@@ -0,0 +1,96 @@
+! Collapsed imperfect loop nests in a target region. OpenMP 6.0 leaves the
+! execution count of intervening code unspecified between once per iteration of
+! the enclosing loop and once per collapsed logical iteration, so the array
+! results here are idempotent and the after-code counter is only range-checked.
+
+! REQUIRES: flang, gpu
+! UNSUPPORTED: nvptx64-nvidia-cuda-LTO
+
+! RUN: %libomptarget-compile-fortran-generic
+! RUN: env LIBOMPTARGET_INFO=16 %libomptarget-run-generic 2>&1 | %fcheck-generic
+
+module collapse_intervening
+  implicit none
+contains
+  subroutine fill(n, m, a)
+    integer, intent(in) :: n, m
+    integer, intent(out) :: a(n * m)
+    integer :: i, j, offset
+
+    !$omp target teams distribute parallel do collapse(2) private(offset) &
+    !$omp   map(from: a)
+    do i = 1, n
+      offset = (i - 1) * m
+      do j = 1, m
+        a(offset + j) = i * 100 + j
+      end do
+    end do
+  end subroutine
+
+  ! Intervening code before the loop at two levels, and after the loop at the
+  ! outermost level. The after-code is a reduction so that counting it is not a
+  ! race.
+  subroutine fill3(n, m, p, b, total)
+    integer, intent(in) :: n, m, p
+    integer, intent(out) :: b(n * m * p)
+    integer, intent(out) :: total
+    integer :: i, j, k, base, row
+
+    total = 0
+    !$omp target teams distribute parallel do collapse(3) private(base, row) &
+    !$omp   reduction(+: total) map(from: b)
+    do i = 1, n
+      base = (i - 1) * m * p
+      do j = 1, m
+        row = base + (j - 1) * p
+        do k = 1, p
+          b(row + k) = i * 10000 + j * 100 + k
+        end do
+      end do
+      total = total + 1
+    end do
+  end subroutine
+end module
+
+program target_teams_distribute_collapse_intervening
+  use collapse_intervening
+  implicit none
+  integer :: n, m, p, i, j, k, errors, total
+  integer, allocatable :: a(:), b(:)
+
+  ! Runtime values, so the collapsed bounds are host-evaluated.
+  n = command_argument_count() + 10
+  m = command_argument_count() + 8
+  p = command_argument_count() + 6
+
+  allocate(a(n * m), b(n * m * p))
+  a = -1
+  b = -1
+
+  call fill(n, m, a)
+  call fill3(n, m, p, b, total)
+
+  errors = 0
+  do i = 1, n
+    do j = 1, m
+      if (a((i - 1) * m + j) /= i * 100 + j) errors = errors + 1
+    end do
+  end do
+
+  do i = 1, n
+    do j = 1, m
+      do k = 1, p
+        if (b((i - 1) * m * p + (j - 1) * p + k) /= i * 10000 + j * 100 + k) &
+          errors = errors + 1
+      end do
+    end do
+  end do
+
+  if (total < n .or. total > n * m * p) errors = errors + 1
+
+  print *, "number of errors: ", errors
+  deallocate(a, b)
+  if (errors /= 0) stop 1
+end program
+
+! CHECK: number of errors:  0

>From ce220ba39c9646a8fa2aa5870760fcdb1c76adf3 Mon Sep 17 00:00:00 2001
From: Caroline Newcombe <caroline.newcombe at hpe.com>
Date: Mon, 21 Sep 2026 13:33:33 -0500
Subject: [PATCH 2/3] Update OpenMP support documentation

---
 flang/docs/OpenMPSupport.md | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/flang/docs/OpenMPSupport.md b/flang/docs/OpenMPSupport.md
index 7f73dc2eabb97..d5b7ff14ba3c6 100644
--- a/flang/docs/OpenMPSupport.md
+++ b/flang/docs/OpenMPSupport.md
@@ -165,7 +165,7 @@ Parser/Semantics, MLIR, Lowering, or the OpenMPIRBuilder.
 | requires directive | <span class="part">partial</span> | | Frontend and lowering coverage exists (`flang/test/Semantics/OpenMP/requires01.f90`-`requires10.f90`, `flang/test/Lower/OpenMP/requires.f90`), but some clauses are still flagged as unsupported (for example reverse_offload warning path). | [llvm/llvm-project#204647](https://github.com/llvm/llvm-project/pull/204647) |
 | teams construct on host | <span class="good">done</span> | | Teams support is established and exercised across semantics/lowering coverage in Flang OpenMP tests. | |
 | loop construct and order(concurrent) clause | <span class="part">partial</span> | | Loop and order-related coverage exists (`flang/test/Semantics/OpenMP/compiler-directives-loop.f90`, `flang/test/Semantics/OpenMP/order-clause01.f90`, `flang/test/Lower/OpenMP/loop-directive.f90`, `flang/test/Lower/OpenMP/order-clause.f90`), with some transformations still evolving. | [llvm/llvm-project#169346](https://github.com/llvm/llvm-project/pull/169346), [llvm/llvm-project#208315](https://github.com/llvm/llvm-project/pull/208315) |
-| collapsing imperfectly nested loops | <span class="good">done</span> | | Intervening code is emitted into the collapsed loop body, so it executes once per collapsed logical iteration. Semantics coverage in `flang/test/Semantics/OpenMP/do-collapse.f90`, lowering coverage in `flang/test/Lower/OpenMP/collapse-imperfect-nest.f90`. | [llvm/llvm-project#211000](https://github.com/llvm/llvm-project/pull/211000) |
+| collapsing imperfectly nested loops | <span class="good">done</span> | | Intervening code is emitted into the collapsed loop body, so it executes once per collapsed logical iteration. Semantics coverage in `flang/test/Semantics/OpenMP/do-collapse.f90`, lowering coverage in `flang/test/Lower/OpenMP/collapse-imperfect-nest.f90`. | [llvm/llvm-project#211000](https://github.com/llvm/llvm-project/pull/211000), [llvm/llvm-project#225159](https://github.com/llvm/llvm-project/pull/225159) |
 | if clause and nontemporal clause on simd | <span class="part">partial</span> | | SIMD nontemporal coverage exists (`flang/test/Semantics/OpenMP/nontemporal.f90`), but complete OpenMP 5.0-level coverage for all if(simd)/nontemporal combinations remains incomplete. | [llvm/llvm-project#110015](https://github.com/llvm/llvm-project/pull/110015) |
 | atomic in simd | <span class="none">unclaimed</span> | | No dedicated Flang OpenMP coverage for atomic-in-simd forms was identified in current parser/semantics/lowering tests. | |
 | detach clause on task and omp_fulfill_event routine | <span class="good">done</span> | | Flang semantics and lowering coverage exists (`flang/test/Semantics/OpenMP/detach01.f90`, `flang/test/Semantics/OpenMP/detach02.f90`, `flang/test/Lower/OpenMP/task_detach.f90`); runtime routine is available in OpenMP module/runtime. | [llvm/llvm-project#119172](https://github.com/llvm/llvm-project/pull/119172), [llvm/llvm-project#119128](https://github.com/llvm/llvm-project/pull/119128) |

>From 173032a4cae7d3d860323d6282f92c8992588fb5 Mon Sep 17 00:00:00 2001
From: Caroline Newcombe <caroline.newcombe at hpe.com>
Date: Tue, 29 Sep 2026 12:52:56 -0500
Subject: [PATCH 3/3] Re-trigger CI




More information about the llvm-commits mailing list