[llvm] [CodeGen] Drop tail-local IR values from memoperands in tail duplication (PR #226768)

MMS IT GmbH via llvm-commits llvm-commits at lists.llvm.org
Sun Sep 27 00:52:18 PDT 2026


https://github.com/mms-it-ch created https://github.com/llvm/llvm-project/pull/226768

When `TailDuplicator` copies an instruction into a predecessor, the memory operands of the copy keep their IR values. A value that is defined in the IR block of the tail (e.g. an address computed from one of its PHIs) then denotes the value of the *next* execution of that block, while alias analysis relates it to the values of the predecessor as if both belonged to the same execution. For a computed-goto dispatch that stores the next target to `p + 16` and then (in the duplicated tail) loads the target from `p'` = `p + 16`, the load's memory operand still says `p`, and AA answers `NoAlias`.

On SystemZ, which uses AA in the machine schedulers, the post-RA scheduler moved the load of the branch target above the store of that target, so the program branched to a stale target. This was found with GnuCOBOL-generated code (PERFORM/GENERATE are translated into a frame stack with computed gotos) on z/OS: the NIST COBOL85 test RW103A looped back into an earlier paragraph. Details and the machine code before/after `postmisched` are in #226758.

This patch drops IR values that are instructions of the tail's IR block (and the AA metadata) from the memory operands of the copies, so that they are treated conservatively. Values defined elsewhere are kept.

New test `taildup-indirectbr-memoperand.ll` checks that the copied loads have no IR value after early tail duplication (without the patch they are `from %ir.scevgep4`, the address computed in the tail block).

Tests: `ninja check-llvm-codegen` on `main` (126dbd97f) with all targets passes (31892 passed, 62 expectedly failed, 242 unsupported, 0 failed).

Fixes #226758.

Assisted-by: Claude Code (Anthropic)

🤖 Generated with [Claude Code](https://claude.com/claude-code)


>From 72c3a64a5df910b25ffc4665aa771c0310eecca6 Mon Sep 17 00:00:00 2001
From: mms-it-ch <info at mms-it.ch>
Date: Sat, 26 Sep 2026 23:17:03 +0200
Subject: [PATCH] [CodeGen] Drop tail-local IR values from memoperands in tail
 duplication

When TailDuplicator copies an instruction into a predecessor, the IR values
in its memory operands still describe addresses as computed in the IR block
of the tail. A value defined there (e.g. derived from one of its PHIs) then
denotes the value of the next execution of the tail block, while alias
analysis relates it to values of the predecessor as if both belonged to the
same execution. For a computed-goto dispatch that loads the next target
from "p" after a store to "p + 40" (where p becomes p + 40 on the back edge)
AA answered NoAlias and the SystemZ post-RA scheduler moved the load of the
branch target above the store, so the program branched to a stale target.

Drop such IR values (and AA metadata) from the memory operands of the copy.

Found with GnuCOBOL-generated code on z/OS (NIST COBOL85 test RW103A).

Assisted-by: Claude Code (Anthropic)
---
 llvm/lib/CodeGen/TailDuplicator.cpp           | 38 ++++++++++++++
 .../SystemZ/taildup-indirectbr-memoperand.ll  | 50 +++++++++++++++++++
 2 files changed, 88 insertions(+)
 create mode 100644 llvm/test/CodeGen/SystemZ/taildup-indirectbr-memoperand.ll

diff --git a/llvm/lib/CodeGen/TailDuplicator.cpp b/llvm/lib/CodeGen/TailDuplicator.cpp
index 3572cd8a3ac17..8ecd2df5f7c88 100644
--- a/llvm/lib/CodeGen/TailDuplicator.cpp
+++ b/llvm/lib/CodeGen/TailDuplicator.cpp
@@ -33,6 +33,7 @@
 #include "llvm/CodeGen/TargetSubtargetInfo.h"
 #include "llvm/IR/DebugLoc.h"
 #include "llvm/IR/Function.h"
+#include "llvm/IR/Instruction.h"
 #include "llvm/Support/CommandLine.h"
 #include "llvm/Support/Debug.h"
 #include "llvm/Support/ErrorHandling.h"
@@ -385,6 +386,42 @@ void TailDuplicator::processPHI(
 
 /// Duplicate a TailBB instruction to PredBB and update
 /// the source operands due to earlier PHI translation.
+/// The IR values in the memory operands of an instruction in \p TailBB
+/// describe addresses as seen in the IR block of \p TailBB. When the
+/// instruction is duplicated into a predecessor, a value defined in that IR
+/// block (e.g. an address computed from one of its PHIs) denotes the value
+/// the block would compute next, not a value of the predecessor. Alias
+/// analysis would relate it to the values used in the predecessor as if both
+/// belonged to the same execution of the block and could wrongly answer
+/// NoAlias (e.g. a store to "p + 40" and a load from "p" where, after the
+/// back edge, p is the old p + 40). Drop such IR values from the memory
+/// operands of the copy, so that it is treated conservatively.
+static void dropTailLocalMemOperandValues(MachineInstr &MI,
+                                          const MachineBasicBlock *TailBB) {
+  const BasicBlock *BB = TailBB->getBasicBlock();
+  if (!BB || MI.memoperands_empty())
+    return;
+  MachineFunction &MF = *MI.getMF();
+  SmallVector<MachineMemOperand *, 2> NewMMOs;
+  bool Changed = false;
+  for (MachineMemOperand *MMO : MI.memoperands()) {
+    const auto *I = dyn_cast_or_null<Instruction>(MMO->getValue());
+    if (I && I->getParent() == BB) {
+      NewMMOs.push_back(MF.getMachineMemOperand(
+          MachinePointerInfo(MMO->getAddrSpace()), MMO->getFlags(),
+          MMO->getSize(), MMO->getBaseAlign(),
+          MMOMetadata(AAMDNodes(), MMO->getRanges(), MMO->getMemCacheHint()),
+          MMO->getSyncScopeID(), MMO->getSuccessOrdering(),
+          MMO->getFailureOrdering()));
+      Changed = true;
+    } else {
+      NewMMOs.push_back(MMO);
+    }
+  }
+  if (Changed)
+    MI.setMemRefs(MF, NewMMOs);
+}
+
 void TailDuplicator::duplicateInstruction(
     MachineInstr *MI, MachineBasicBlock *TailBB, MachineBasicBlock *PredBB,
     DenseMap<Register, RegSubRegPair> &LocalVRMap,
@@ -398,6 +435,7 @@ void TailDuplicator::duplicateInstruction(
     return;
   }
   MachineInstr &NewMI = TII->duplicate(*PredBB, PredBB->end(), *MI);
+  dropTailLocalMemOperandValues(NewMI, TailBB);
   if (!PreRegAlloc)
     return;
   for (unsigned i = 0, e = NewMI.getNumOperands(); i != e; ++i) {
diff --git a/llvm/test/CodeGen/SystemZ/taildup-indirectbr-memoperand.ll b/llvm/test/CodeGen/SystemZ/taildup-indirectbr-memoperand.ll
new file mode 100644
index 0000000000000..c4895135e043c
--- /dev/null
+++ b/llvm/test/CodeGen/SystemZ/taildup-indirectbr-memoperand.ll
@@ -0,0 +1,50 @@
+; RUN: llc < %s -mtriple=s390x-linux-gnu -O2 -stop-after=early-tailduplication \
+; RUN:   | FileCheck %s
+;
+; Computed-goto dispatch: %ig loads the next target from a slot whose address
+; depends on a PHI of %ig; l1 and l2 store the next target one slot further
+; and branch back. Early tail duplication copies %ig into its predecessors.
+; In a copy, the IR value of the load address describes the slot as computed
+; by the *next* execution of %ig, so it must not be kept in the memory
+; operand: alias analysis would relate it to the store in the same block as
+; if both belonged to one execution of %ig and could answer NoAlias although
+; both access the same slot. The scheduler could then move the load of the
+; branch target above the store (seen on z/OS with GnuCOBOL-generated code).
+
+define i64 @f(ptr %base, ptr %ext) {
+; CHECK-LABEL: name: f
+; CHECK:       bb.{{[0-9]+}}.l1
+; CHECK:         STG {{.*}} :: (store (s64) into %ir.
+; CHECK-NEXT:    {{.*}} = LA
+; CHECK-NEXT:    {{.*}} = LG {{.*}} :: (load (s64))
+; CHECK:       bb.{{[0-9]+}}.l2
+; CHECK:         STG {{.*}} :: (store (s64) into %ir.
+; CHECK-NEXT:    {{.*}} = LA
+; CHECK-NEXT:    {{.*}} = LG {{.*}} :: (load (s64))
+entry:
+  store ptr blockaddress(@f, %l1), ptr %base
+  br label %ig
+
+ig:
+  %idx = phi i64 [ 0, %entry ], [ %idx1, %l1 ], [ %idx2, %l2 ]
+  %p = getelementptr i8, ptr %base, i64 %idx
+  %t = load ptr, ptr %p, align 8
+  indirectbr ptr %t, [label %l1, label %l2, label %done]
+
+l1:
+  store volatile i64 1, ptr %ext
+  %idx1 = add nsw i64 %idx, 16
+  %q1 = getelementptr inbounds i8, ptr %base, i64 %idx1
+  store ptr blockaddress(@f, %l2), ptr %q1, align 8
+  br label %ig
+
+l2:
+  store volatile i64 2, ptr %ext
+  %idx2 = add nsw i64 %idx, 16
+  %q2 = getelementptr inbounds i8, ptr %base, i64 %idx2
+  store ptr blockaddress(@f, %done), ptr %q2, align 8
+  br label %ig
+
+done:
+  ret i64 %idx
+}



More information about the llvm-commits mailing list