[llvm] [AMDGPU][NewInsertWaitcnt] Add Event Tracker (PR #226970)

Pierre van Houtryve via llvm-commits llvm-commits at lists.llvm.org
Fri Oct 2 02:59:46 PDT 2026


https://github.com/Pierre-vh updated https://github.com/llvm/llvm-project/pull/226970

>From f8c431c6a6db376431c2f2cea3072be15840baa6 Mon Sep 17 00:00:00 2001
From: pvanhout <pierre.vanhoutryve at amd.com>
Date: Mon, 28 Sep 2026 12:44:27 +0200
Subject: [PATCH 1/5] [AMDGPU][NewInsertWaitcnt] Add Event Tracker

See #226335

Introduction

Adds the basic infrastructure needed to track in-flight records for each `InstCounterType`.
This is not a full replacement of `SIInsertWaitcnts::WaitcntBrackets` and it does not
aim to be one.

One goal with the new implementation is to avoid having a single "tracker" class
that does everything. Instead, the `EventTracker` aims to do one thing and do it right:
preserve the history of a counter (instructions + events issued) and the value of the counter.
Any specific queries, such as "what STORE_CNT do I need to use this RU" is something that belongs
to helper methods in the client of the class or a utils file.

Class Design

This class is designed to use a single "source of truth" for all information, which is a compact
timeline of events. The intent is that we'll just iterate the timeline to answer certain queries
instead of baking-in a bunch of `DenseMaps` for every query dimension we need.
This considerably simplifies the implementation. Most of the code in the files here are
boilerplate, assertions, debug dumps, verification methods, and so on. There is very little "critical" logic.

The design/API presented here is *not* final - it'll evolve as we aim for feature parity with
`SIInsertWaitcnts` in the `NewInsertWaitcnts` pass.

Performance

As this design leans on iterating the timeline possibly multiple times per instruction, I made sure it'd be fast
by keeping the record type small (16B right now, may grow to 32) and stored contiguously.
I also did some profiling on very big test cases by shadowing `WaitcntBrackets` and doing many (dozens) of timeline
iteration each instruction for every counter, and there is no visible spike in the flame graph of the profiler.

Testing

I tested this implementation by shadowing the `WaitcntBrackets` class in a downstream branch and checking
my results against it, until I was confident that the values given by the `EventTracker` were as precise or
more precise than the values given by `WaitcntBrackets` for the entire test suite.

Use of AI

AI was exclusively used for 2 things:

- As a design companion, helping me flesh out architecture ideas in "plan mode" so I could iterate fast.
  I also used it as a research tool to find prior art, gather data for me to verify, etc.
- As a refactoring/boilerplate generation tool, such as creating files with basic boilerplate,
  generating scaffolding like methods/class declarations/etc.

AI was not used to write any serious code/algorithms in any function of either the unit tests or the implementation files.

Assisted-by: Claude Opus (4.8/5.5) and Sonnet (5)
---
 .../lib/Target/AMDGPU/AMDGPUEventTracking.cpp | 477 ++++++++++++++
 llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h  | 354 ++++++++++
 llvm/lib/Target/AMDGPU/AMDGPUHWEvents.h       |  27 +
 llvm/lib/Target/AMDGPU/CMakeLists.txt         |   1 +
 .../Target/AMDGPU/AMDGPUEventTrackingTest.cpp | 618 ++++++++++++++++++
 llvm/unittests/Target/AMDGPU/CMakeLists.txt   |   1 +
 6 files changed, 1478 insertions(+)
 create mode 100644 llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
 create mode 100644 llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
 create mode 100644 llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp

diff --git a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
new file mode 100644
index 0000000000000..8b383b45de67a
--- /dev/null
+++ b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
@@ -0,0 +1,477 @@
+//===- AMDGPUEventTracking.cpp ----------------------------------*- C++ -*-===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+#include "AMDGPUEventTracking.h"
+#include "AMDGPUHWEvents.h"
+#include "AMDGPUWaitcntUtils.h"
+#include "llvm/CodeGen/MachineBasicBlock.h"
+#include "llvm/CodeGen/MachineFunction.h"
+#include "llvm/CodeGen/MachineInstr.h"
+#include "llvm/Support/CommandLine.h"
+#include "llvm/Support/Debug.h"
+#include <algorithm>
+#include <optional>
+
+#define DEBUG_TYPE "amdgpu-event-tracking"
+
+namespace llvm {
+
+/// Mimic legacy (coarse) tracking of counter state.
+static cl::opt<bool> MimicLegacyTracking("amdgpu-event-legacy-tracking",
+                                         cl::init(false));
+
+#ifndef NDEBUG
+static cl::opt<bool> EventTrackerPrintAll(
+    "amdgpu-event-tracker-print-all", cl::init(false),
+    cl::desc("When using -debug, print the full set of live events every time "
+             "an event is added or removed"));
+#endif
+
+namespace AMDGPU {
+namespace eventtracking {
+
+namespace {
+static bool greaterThan(const EventTrackerRecord &A,
+                        const EventTrackerRecord &B) {
+  return A.getScore() > B.getScore();
+}
+} // namespace
+
+EventTrackerRecord::EventTrackerRecord(EventTrackingContext &Ctx,
+                                       MachineInstr *MI, SingleHWEvent Kind,
+                                       uint32_t Score)
+    : EventTrackerRecord(Ctx.nextDynamicInstanceID(), MI, Kind, Score) {}
+
+void EventTrackerRecord::print(raw_ostream &OS, bool PrintMI,
+                               unsigned Indent) const {
+  OS.indent(Indent) << "#" << ID << " " << Kind << " (Score=" << Score << "): ";
+  if (PrintMI && MI)
+    OS << *MI;
+  else
+    OS << "MI@" << (void *)MI << "\n";
+}
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+LLVM_DUMP_METHOD void EventTrackerRecord::dump() const {
+  dbgs() << "\n";
+  print(dbgs(), /*PrintMI=*/true);
+  dbgs() << "\n";
+}
+#endif
+
+EventTracker::EventTracker(MachineBasicBlock &MBB, EventTrackingContext &ETC)
+    : MBB(&MBB), Ctx(&ETC) {
+  Counters.resize(ETC.counters().size());
+  for (const CounterInfo &Info : ETC.counters())
+    Counters[Info.CounterT].CI = &Info;
+}
+
+void EventTracker::enterBlock() {
+  LLVM_DEBUG(dbgs() << "\n[EventTracker] Entering ";
+             MBB->printAsOperand(dbgs()); dbgs() << "\n");
+
+  // FIXME: This is a bit hacky, but we need to save the old state in case the
+  // MBB is also its own predecessor. Revisit when the design and clients of
+  // this class are set in stone.
+
+  SmallVector<EventTracker *> Preds;
+  bool IsSelfPred = false;
+  if (Preds.empty() && !MBB->pred_empty()) {
+    for (MachineBasicBlock *Pred : MBB->predecessors()) {
+      if (Pred == MBB)
+        IsSelfPred = true;
+      else
+        Preds.push_back(&(*Ctx)[Pred]);
+    }
+  }
+
+  if (IsSelfPred) {
+    EventTracker SelfCopy = *this;
+    Preds.push_back(&SelfCopy);
+    clear();
+    recordIncomings(*Ctx, Preds);
+  } else {
+    clear();
+    recordIncomings(*Ctx, Preds);
+  }
+}
+
+void EventTracker::leaveBlock() {
+  LLVM_DEBUG(dbgs() << "[EventTracker] Leaving "; MBB->printAsOperand(dbgs());
+             dbgs() << "\n");
+}
+
+void EventTracker::record(MachineInstr &MI, SingleHWEvent Event) {
+  LLVM_DEBUG(dbgs() << "[EventTracker] Recording " << Event << ": " << MI);
+
+  EventTrackerRecord Rec = EventTrackerRecord(*Ctx, &MI, Event);
+  [[maybe_unused]] bool FoundMatch = false;
+  for (auto &CD : Counters) {
+    if (!CD.CI->Events.contains(Event))
+      continue;
+
+    FoundMatch = true;
+    ++CD.Count;
+    CD.LegacyPendingEvents |= Event;
+
+    // Do not age records if we are out-of-order.
+    if (!CD.IsOutOfOrder) {
+      // NB: There is an intentional tradeoff here. We could avoid this loop by
+      // instead storing a timestamp in each record, and having a
+      // constantly-increasing clock to infer the score (clock-timestamp is
+      // score). However, it'd:
+      //  - Complexify fetching the score (`EventTrackerRecord` cannot answer it
+      //    on its own anymore and we need a separate query/wrapper).
+      //  - Make merge of incoming records a bit more annoying (we'd need to
+      //    rebase the `clock`).
+      //  - Potentially demand (much) more space in EventTrackerRecord to store
+      //    bigger numbers.
+      //
+      // All in all, I think this small loop is fine for now, but we can still
+      // change the system if we have data backed up by profiling to
+      // justify the change.
+      for (auto &Live : CD.LiveRecords)
+        Live.setScore(Live.getScore() + 1); // Age all existing events.
+    }
+
+    CD.LiveRecords.push_back(Rec);
+
+#ifndef NDEBUG
+    LLVM_DEBUG(if (EventTrackerPrintAll) {
+      dbgs().indent(2) << "Updated Timeline:\n";
+      print(dbgs(), CD, /*Indent=*/4);
+    });
+#endif
+  }
+
+  assert(FoundMatch && "Event has no matching InstCounterType!");
+}
+
+void EventTracker::wait(InstCounterType T, unsigned N) {
+  auto &CD = get(T);
+  LLVM_DEBUG(dbgs() << "[EventTracker] Wait on " << getInstCounterName(T)
+                    << " for " << N << "\n");
+
+  // Fast path for clearing the counter
+  if (N == 0) {
+    CD.LiveRecords.clear();
+    CD.Count = 0;
+    CD.IsIndeterminate = false;
+    CD.IsOutOfOrder = false;
+    CD.LegacyPendingEvents = HWEvents();
+    return;
+  }
+
+  CD.Count = std::min(CD.Count, N);
+
+  // Don't bother erasing stuff if we are out-of-order. All records have a score
+  // of zero in such cases.
+  if (!CD.IsOutOfOrder) {
+    auto *RmIt = remove_if(CD.LiveRecords, [&](EventTrackerRecord &E) {
+      if (E.getScore() < N)
+        return false;
+      LLVM_DEBUG(dbgs() << "  | Removing "; E.print(dbgs()));
+      return true;
+    });
+    CD.LiveRecords.erase(RmIt, CD.LiveRecords.end());
+  }
+
+  LLVM_DEBUG(dbgs() << "  | => Updated Count:" << CD.Count << "\n");
+
+#ifndef NDEBUG
+  LLVM_DEBUG(if (EventTrackerPrintAll) {
+    dbgs().indent(2) << "Updated Timeline:\n";
+    print(dbgs(), CD, /*Indent=*/4);
+  });
+#endif
+}
+
+void EventTracker::markIndeterminate(InstCounterType T) {
+  LLVM_DEBUG(dbgs() << "[EventTracker] Marking " << getInstCounterName(T)
+                    << " as indeterminate!\n");
+  auto &CD = get(T);
+  CD.IsIndeterminate = true;
+  markOutOfOrder(T);
+}
+
+void EventTracker::markOutOfOrder(InstCounterType T) {
+  LLVM_DEBUG(dbgs() << "[EventTracker] Marking " << getInstCounterName(T)
+                    << " as out-of-order!\n");
+  auto &CD = get(T);
+  CD.IsOutOfOrder = true;
+  for (auto &Rec : CD.LiveRecords)
+    Rec.setScore(0);
+}
+
+std::optional<unsigned> EventTracker::count(InstCounterType T) const {
+  auto &CD = get(T);
+  if (CD.IsIndeterminate)
+    return std::nullopt;
+  return get(T).Count;
+}
+
+bool EventTracker::isIndeterminate(InstCounterType T) const {
+  return get(T).IsIndeterminate;
+}
+
+bool EventTracker::isOutOfOrder(InstCounterType T) const {
+  return get(T).IsOutOfOrder;
+}
+
+HWEvents EventTracker::getPendingEvents(InstCounterType T) const {
+  const auto &CD = get(T);
+  if (CD.IsIndeterminate)
+    return CD.CI->Events; // return all events
+
+  if (MimicLegacyTracking)
+    return CD.LegacyPendingEvents;
+
+  HWEvents Res;
+  for (const auto &E : get(T).LiveRecords)
+    Res |= E.getKind();
+  return Res;
+}
+
+ArrayRef<EventTrackerRecord>
+EventTracker::getLiveRecords(InstCounterType T) const {
+  return get(T).LiveRecords;
+}
+
+void EventTracker::print(raw_ostream &OS, InstCounterType T) const {
+  print(OS, get(T));
+}
+
+bool EventTracker::mimicsLegacyTracking() { return MimicLegacyTracking; }
+
+#if !defined(NDEBUG) || defined(EXPENSIVE_CHECKS)
+void EventTracker::verify() const {
+  assert(MBB && Ctx && "Invalid internal state!");
+
+  for (auto &C : Counters) {
+    const auto OnError = [&]() {
+      dbgs() << "EventTracker verification error\n";
+      print(dbgs(), C);
+    };
+
+    if (C.IsIndeterminate) {
+      if (!C.IsOutOfOrder) {
+        OnError();
+        assert(false && "IsIndeterminate but not IsOutOfOrder");
+      }
+      continue;
+    }
+
+    if (C.IsOutOfOrder) {
+      if (!all_of(C.LiveRecords, [](auto &R) { return R.getScore() == 0; })) {
+        OnError();
+        assert(false &&
+               "IsOutOfOrder but some records do not have a score of 0!");
+      }
+    }
+
+    if (C.Count > C.LiveRecords.size() && !mimicsLegacyTracking()) {
+      OnError();
+      assert(false &&
+             "'Count' is inconsistent with the number of live records");
+    }
+
+    for (auto &E : C.LiveRecords) {
+      if (E.getScore() > C.Count) {
+        OnError();
+        dbgs() << "Concerning Record:";
+        E.print(dbgs());
+        assert(false && "record score is out of range");
+      }
+    }
+
+    // Check live records are sorted
+    if (!is_sorted(C.LiveRecords, greaterThan)) {
+      OnError();
+      assert(false && "live records are not sorted!");
+    }
+  }
+}
+#endif
+
+void EventTracker::print(raw_ostream &OS, bool IgnoreEmpty,
+                         unsigned Indent) const {
+  OS.indent(Indent) << "EventTracker for ";
+  MBB->printAsOperand(OS);
+  OS << ":";
+  if (IgnoreEmpty) {
+    if (all_of(Counters, [](const auto &CD) { return CD.Count == 0; })) {
+      OS << " (empty)\n";
+      return;
+    }
+  }
+
+  OS << "\n";
+  for (auto &C : Counters) {
+    if (IgnoreEmpty && C.Count == 0)
+      continue;
+    print(dbgs(), C, Indent + 2);
+  }
+}
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+LLVM_DUMP_METHOD void EventTracker::dump() const {
+  dbgs() << "\n";
+  print(dbgs());
+  dbgs() << "\n";
+}
+#endif
+
+void EventTracker::clear() {
+  for (auto &C : Counters) {
+    C.LiveRecords.clear();
+    C.Count = 0;
+    C.IsIndeterminate = false;
+    C.IsOutOfOrder = false;
+  }
+}
+
+void EventTracker::recordIncomings(EventTrackingContext &ETC,
+                                   ArrayRef<EventTracker *> Preds) {
+  LLVM_DEBUG(if (!Preds.empty()) {
+    dbgs() << "[EventTracker] Recording incoming events (merge) from "
+              "predecessors:\n";
+    for (EventTracker *Pred : Preds) {
+      Pred->print(dbgs(), /*IgnoreEmpty=*/true, /*Indent=*/2);
+    }
+  });
+
+  /// Iterate over all counters that are available to us.
+  for (auto &CI : ETC.counters()) {
+    auto &CData = Counters[CI.CounterT];
+    assert(CData.LiveRecords.empty());
+
+    DenseMap<EventTrackerRecord::DynamicInstanceID, EventTrackerRecord> Acc;
+
+    for (EventTracker *Pred : Preds) {
+      auto &PredCData = Pred->Counters[CI.CounterT];
+
+      // Merge domain for the count value:
+      CData.Count = std::max(CData.Count, PredCData.Count);
+      // Merge domain for the legacy pending events.
+      CData.LegacyPendingEvents |= PredCData.LegacyPendingEvents;
+      // Merge domain for the indeterminate state.
+      CData.IsIndeterminate |= PredCData.IsIndeterminate;
+      // Merge domain for the out-of-order state.
+      CData.IsOutOfOrder |= PredCData.IsOutOfOrder;
+
+      for (EventTrackerRecord &PredEntry : PredCData.LiveRecords) {
+        EventTrackerRecord::DynamicInstanceID ID = PredEntry.getID();
+        auto It = Acc.find(ID);
+        if (It != Acc.end()) {
+          auto &AccVal = It->second;
+          AccVal.setScore(std::min(AccVal.getScore(), PredEntry.getScore()));
+          assert(
+              PredEntry.getMI() == AccVal.getMI() &&
+              PredEntry.getKind() == AccVal.getKind() &&
+              "EventTrackerRecord have same DynamicInstanceID, but different "
+              "MachineInstr/HWEvent kind, which should not be possible");
+        } else
+          Acc.insert({ID, PredEntry});
+      }
+    }
+
+    if (MimicLegacyTracking) {
+      CData.PersistentUpperBound =
+          std::max(CData.PersistentUpperBound, CData.Count);
+      CData.Count = CData.PersistentUpperBound;
+    }
+
+    auto AccVals = Acc.values();
+    CData.LiveRecords.append(AccVals.begin(), AccVals.end());
+
+    // Sort records by Score (descending) for consistent iteration.
+    stable_sort(CData.LiveRecords, greaterThan);
+  }
+
+  LLVM_DEBUG(if (!Preds.empty()) {
+    dbgs() << "[EventTracker] Timeline after recording incomings:\n";
+    print(dbgs(), /*IgnoreEmpty=*/true, /*Indent=*/2);
+  });
+}
+
+void EventTracker::print(raw_ostream &OS, const CounterData &CD,
+                         unsigned Indent) {
+  OS.indent(Indent) << getInstCounterName(CD.CI->CounterT)
+                    << " (Count=" << CD.Count
+                    << ", PersistentUpperBound=" << CD.PersistentUpperBound
+                    << ", LiveRecords=" << CD.LiveRecords.size()
+                    << ", IsOutOfOrder=" << CD.IsOutOfOrder
+                    << ", IsIndeterminate=" << CD.IsIndeterminate << ")\n";
+  for (const auto &E : CD.LiveRecords) {
+    OS.indent(Indent + 2);
+    E.print(OS);
+  }
+}
+
+EventTrackingContext::EventTrackingContext(MachineFunction &MF,
+                                           ArrayRef<CounterInfo> Counters)
+    : CounterInfos(Counters) {
+  LLVM_DEBUG(dbgs() << "\n[EventTrackingContext] CounterInfos for "
+                    << MF.getName() << "\n";
+             for (const auto &CI
+                  : CounterInfos) {
+               dbgs().indent(2)
+                   << AMDGPU::getInstCounterName(CI.CounterT) << " ";
+               if (CI.Events.none()) {
+                 dbgs() << " (unused - no HWEvents assigned)\n";
+               } else {
+                 dbgs() << "(Limit=" << CI.Limit << ") " << CI.Events << "\n";
+               }
+             });
+
+  Trackers.reserve(MF.size());
+  for (MachineBasicBlock &MBB : MF)
+    Trackers[&MBB] = std::make_unique<EventTracker>(MBB, *this);
+}
+
+EventTracker &EventTrackingContext::operator[](MachineBasicBlock *MBB) {
+  assert(MBB);
+  return *Trackers.at(MBB);
+}
+
+#if !defined(NDEBUG) || defined(EXPENSIVE_CHECKS)
+void EventTrackingContext::verify() const {
+  for (const auto &[MBB, Tracker] : Trackers) {
+    assert(MBB && "Unexpected nullptr entry!");
+    Tracker->verify();
+  }
+
+  // Check CounterInfos is sane.
+  for (auto [Idx, CI] : enumerate(CounterInfos)) {
+    assert(Idx == CI.CounterT && "CounterInfo is in wrong position!");
+    assert(CI.Events.any() &&
+           "InstCounterType has no event associated with it!");
+  }
+}
+#endif
+
+void EventTrackingContext::print(raw_ostream &OS) const {
+  for (const auto &[MBB, Tracker] : Trackers) {
+    MBB->printAsOperand(OS);
+    OS << ":\n";
+    Tracker->print(OS, /*Indent=*/2);
+  }
+}
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+LLVM_DUMP_METHOD void EventTrackingContext::dump() const {
+  dbgs() << "\n";
+  print(dbgs());
+  dbgs() << "\n";
+}
+#endif
+
+} // namespace eventtracking
+} // namespace AMDGPU
+
+} // namespace llvm
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
new file mode 100644
index 0000000000000..c1e56fc184569
--- /dev/null
+++ b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
@@ -0,0 +1,354 @@
+//===- AMDGPUEventTracking.h ------------------------------------*- C++ -*-===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+//
+/// \file Tracks a per-InstCounterType event timeline which preserves
+/// information about previously encountered MachineInstr and HWEvents.
+//
+//===----------------------------------------------------------------------===//
+
+#ifndef LLVM_LIB_TARGET_AMDGPU_UTILS_AMDGPUEVENTTRACKING_H
+#define LLVM_LIB_TARGET_AMDGPU_UTILS_AMDGPUEVENTTRACKING_H
+
+#include "AMDGPUHWEvents.h"
+#include "AMDGPUWaitcntUtils.h"
+#include "llvm/ADT/ArrayRef.h"
+#include "llvm/ADT/DenseMap.h"
+#include "llvm/ADT/SmallVector.h"
+#include "llvm/ADT/Twine.h"
+#include <memory>
+
+namespace llvm {
+class raw_ostream;
+class MachineBasicBlock;
+class MachineOperand;
+class MachineFunction;
+
+namespace AMDGPU {
+namespace eventtracking {
+
+class EventTrackingContext;
+class EventTracker;
+
+/// FIXME: Make this a generic util?
+struct CounterInfo {
+  constexpr CounterInfo(InstCounterType T, HWEvents Events, unsigned Limit)
+      : CounterT(T), Events(Events), Limit(Limit) {}
+
+  /// Type of counter this is.
+  /// This must match the index in the container, e.g. CounterT=2 must be at
+  /// index 2.
+  InstCounterType CounterT;
+  /// HWEvents for this counter.
+  HWEvents Events;
+  /// Hardware limit for this counter.
+  unsigned Limit;
+};
+
+/// Represents entries in the \ref EventTracker.
+///
+/// Records are all uniquely identified by a \ref DynamicInstanceID. For each
+/// unique \ref DynamicInstanceID value, all \ref EventTrackerRecord that use
+/// that ID should have the same MI and Kind. This is enforced by exposing these
+/// as read-only, and making the constructor assign a new \ref DynamicInstanceID
+/// every time.
+///
+/// Only the score can change as it may be unique to each instance
+/// of \ref EventTracker that carry it.
+class EventTrackerRecord {
+public:
+  /// An always-increasing counter used to represent a dynamic instance of a
+  /// record.
+  ///
+  /// Whenever we add a new \ref EventTrackerRecord, even if it's one we already
+  /// have seen in a previous dataflow iteration, this counter is increased so
+  /// that the new record has a unique `DynamicInstanceID`.
+  using DynamicInstanceID = uint32_t;
+
+  EventTrackerRecord(EventTrackingContext &Ctx, MachineInstr *MI,
+                     SingleHWEvent Kind, uint32_t Score = 0);
+
+  /// \returns the ID uniquely identifying this record across an entire
+  /// EventTrackingContext. Whenever we revisit an instruction (when iterating
+  /// until a fixpoint is reached), we give it a new ID. This is used to
+  /// represent records carried over from previous iterations of the same basic
+  /// block.
+  DynamicInstanceID getID() const { return ID; }
+
+  /// \returns the MachineInstr that originated this record.
+  MachineInstr *getMI() const { return MI; }
+
+  /// \returns the kind of record this is, as a \ref SingleHWEvent.
+  SingleHWEvent getKind() const { return Kind; }
+
+  /// \returns the score of this record.
+  uint32_t getScore() const { return Score; }
+
+  /// Sets the score of this record to \p NewScore.
+  void setScore(uint32_t NewScore) {
+    Score = NewScore;
+    assert(Score == NewScore && "Score overflow!");
+  }
+
+  void print(raw_ostream &OS, bool PrintMI = true, unsigned Indent = 0) const;
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+  LLVM_DUMP_METHOD void dump() const;
+#endif
+
+private:
+  EventTrackerRecord(DynamicInstanceID ID, MachineInstr *MI, SingleHWEvent Kind,
+                     uint32_t Score = 0)
+      : MI(MI), ID(ID), Kind(Kind) {
+    setScore(Score);
+  }
+
+  MachineInstr *MI;
+  DynamicInstanceID ID;
+  SingleHWEvent Kind;
+  // Score should already never exceed uint8_t limit in normal circumstances as
+  // most counter types only use up to 6 bits encoding for the waitcnts. 16 bit
+  // is a very generous limit, we can probably shrink that at some point.
+  uint16_t Score;
+};
+
+/// This assert serves as a reminder to be mindful of the size of the object.
+static_assert(sizeof(EventTrackerRecord) == 16,
+              "EventTrackerRecord should remain small to optimize its layout "
+              "within cache lines, for maximum iteration speed");
+
+/// Per-MBB tracking context.
+///
+/// Tracks data across the following domains:
+///   - Current value (count) of each instruction counter.
+///   - In-flight (alive) \ref EventTrackerRecord of each instruction counter.
+///
+/// This class is only responsible for tracking records for every
+/// InstCounterType. It does not deal with calculating the waitcnts needed, or
+/// doing more advance reasoning over the timeline for specific queries (e.g.
+/// finding an aliasing store). These responsibilities are for
+/// utils/wrappers/users of the class.
+///
+/// The API should be kept as simple and clear as possible.
+class EventTracker {
+public:
+  EventTracker(MachineBasicBlock &MBB, EventTrackingContext &ET);
+
+  /// \defgroup MachineBasicBlock entry and exit
+  /// \{
+
+  /// Notify this EventTracker that we are going to begin recording events.
+  /// In case this is not the first time we are going through this block, this
+  /// clears the internal state of the tracker and re-imports all incoming
+  /// tracking state from the predecessors.
+  void enterBlock();
+
+  /// Notify this EventTracker that we are done recording events.
+  void leaveBlock();
+
+  /// \}
+
+  /// \defgroup InstCounters Tracking Entrypoints
+  /// Methods update the state of the InstCounters by adding/removing events
+  /// or signaling certain special conditions.
+  /// \{
+
+  /// Record an event of type \p Event at a MachineInstr \p MI, which will
+  /// affect all counters that have \p Event in their event set.
+  void record(MachineInstr &MI, SingleHWEvent Event);
+
+  /// Notify that we waited until the counter \p T reached the value \p N before
+  /// continuing execution of the program (and recording more events).
+  ///
+  /// This affects the count of \p T, an removes all records that have a score
+  /// greater than or equal to \p N.
+  ///
+  /// If \p N is zero, then \p T will no longer be in an indeterminate or
+  /// out-of-order state afterwards if it previously was in such a state.
+  void wait(InstCounterType T, unsigned N = 0);
+
+  /// Mark the counter \p T as being in an indeterminate state. This means that
+  /// we no longer accurately track \p T because there may be more records we do
+  /// not know about. This primarily affects \ref getPendingEvents and
+  /// \ref count.
+  ///
+  /// Implies \ref markOutOfOrder for \p T as well.
+  void markIndeterminate(InstCounterType T);
+
+  /// Mark the counter \p T as being "out-of-order", meaning records may retire
+  /// in any order. This sets the score of all records to zero.
+  void markOutOfOrder(InstCounterType T);
+
+  /// \}
+
+  /// \defgroup InstCounters Tracking Queries
+  /// Query the current state of each InstCounter without modifying it.
+  /// \{
+
+  /// \returns the current value of the counter \p T at this point in time, or
+  /// std::nullopt if \p T is in the indeterminate state.
+  std::optional<unsigned> count(InstCounterType T) const;
+
+  /// \returns true if the counter \p T is in an indeterminate state.
+  bool isIndeterminate(InstCounterType T) const;
+
+  /// \returns true if the counter \p T is out-of-order
+  bool isOutOfOrder(InstCounterType T) const;
+
+  /// \returns the set of pending HWEvents for \p T. If \p T is in an
+  /// indeterminate state, returns a conservative set of pending events instead.
+  HWEvents getPendingEvents(InstCounterType T) const;
+
+  /// \returns the set of live records recorded for \p T. This is the list of
+  /// all instructions in-flight for that counter.
+  /// Note that if \p T is indeterminate, then this set is non-exhaustive. It
+  /// only contains the records this class knows about.
+  ArrayRef<EventTrackerRecord> getLiveRecords(InstCounterType T) const;
+
+  /// \}
+
+  /// \defgroup Miscellaneous helpers
+  /// \{
+
+  /// Prints a dump of the internal tracking state of this class for \p T to the
+  /// stream \p OS.
+  void print(raw_ostream &OS, InstCounterType T) const;
+
+  /// Prints a dump of all internal tracking state of this class to the stream
+  /// \p OS. If \p IgnoreEmpty is true, do not print counters with a count of 0.
+  void print(raw_ostream &OS, bool IgnoreEmpty = false,
+             unsigned Indent = 0) const;
+
+  /// \returns true if the option to mimic legacy (SIInsertWaitcnts
+  ///          scoreboard-style) tracking of counters and pending events.
+  /// TODO: Remove in the future when legacy tracking is no longer needed.
+  static bool mimicsLegacyTracking();
+
+#if !defined(NDEBUG) || defined(EXPENSIVE_CHECKS)
+  /// Verifies invariants of this class are respected.
+  void verify() const;
+#endif
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+  LLVM_DUMP_METHOD void dump() const;
+#endif
+
+  /// \}
+
+private:
+  struct CounterData {
+    const CounterInfo *CI = nullptr;
+
+    /// Set of live records (events) that make up the `Count`.
+    SmallVector<EventTrackerRecord, 16> LiveRecords;
+    /// The current value of the counter. This is a max (upper bound) across
+    /// all possible execution paths at runtime. It cannot be inferred from the
+    /// LiveRecords alone and is thus a separate tracking domain.
+    uint32_t Count = 0;
+    /// An upper bound that persists across fixpoint iterations. This is only
+    /// used when \ref mimicsLegacyTracking returns true.
+    uint32_t PersistentUpperBound = 0;
+    /// Whether this counter is in an indeterminate state, which means that both
+    /// the set of LiveRecords and the Count are imprecise. This implies that
+    /// the counter is out-of-order as well.
+    bool IsIndeterminate = false;
+    /// Whether this counter is out-of-order, meaning records may retire in any
+    /// order and they all exist at a score of zero.
+    bool IsOutOfOrder = false;
+    /// Legacy-style tracking of pending events that is coarse and does not
+    /// leverage the live set of records. Only used when
+    /// \ref mimicsLegacyTracking returns true and not cleared between
+    /// iterations.
+    HWEvents LegacyPendingEvents;
+
+    // TODO: We could imagine storing the per-predecessor score for incoming
+    // events. We could achieve that by storing that as a map of ((ID, Pred),
+    // Score). This would allow identifying events that are "deep" in one branch
+    // but "shallow" in another, e.g. an event needing a waitcnt 1 for one pred,
+    // but a waitcnt 8 for another. Not sure if we can exploit that though?
+  };
+
+  CounterData &get(InstCounterType T) {
+    assert(Counters.size() > T && "T is out of range!");
+    return Counters[T];
+  }
+
+  const CounterData &get(InstCounterType T) const {
+    assert(Counters.size() > T && "T is out of range!");
+    return Counters[T];
+  }
+
+  /// Clears the tracked data, used when entering a block.
+  void clear();
+
+  /// Import all events from the incoming basic blocks in \p Preds and reconcile
+  /// divergence at joints.
+  void recordIncomings(EventTrackingContext &ETC,
+                       ArrayRef<EventTracker *> Preds);
+
+  static void print(raw_ostream &OS, const CounterData &CD,
+                    unsigned Indent = 0);
+
+  MachineBasicBlock *MBB;
+  EventTrackingContext *Ctx;
+
+  // NB: This, combined with the inline storage of LiveRecords, can lead to this
+  // class becoming quite big - verify the size of this object whenever a change
+  // is made.
+  SmallVector<CounterData, InstCounterType::NUM_INST_CNTS> Counters;
+};
+
+/// Per-MF Tracking Context.
+///
+/// This owns all \ref EventTrackers and keeps track of state that persists
+/// across dataflow analysis iterations, such as the current value of
+/// \ref DynamicInstanceID.
+class EventTrackingContext {
+public:
+  /// \param MF Machine Function
+  /// \param Counters The counters available to \p MF on this target.
+  EventTrackingContext(MachineFunction &MF, ArrayRef<CounterInfo> Counters);
+
+  /// Fetch the \ref MBBEventTracker of \p MBB.
+  EventTracker &operator[](MachineBasicBlock *MBB);
+
+  /// \returns the list of counters available to the current target.
+  ArrayRef<CounterInfo> counters() const { return CounterInfos; }
+
+#if !defined(NDEBUG) || defined(EXPENSIVE_CHECKS)
+  /// Verifies invariants of this class are respected.
+  void verify() const;
+#endif
+
+  void print(raw_ostream &OS) const;
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+  LLVM_DUMP_METHOD void dump() const;
+#endif
+
+private:
+  friend class EventTrackerRecord;
+
+  /// \returns a new, unique \ref DynamicInstanceID - only for use by
+  /// \ref EventTrackerRecord.
+  EventTrackerRecord::DynamicInstanceID nextDynamicInstanceID() {
+    assert(NextDynID + 1 > NextDynID && "DynamicInstanceIDs overflow!");
+    return ++NextDynID;
+  }
+
+  SmallVector<CounterInfo> CounterInfos;
+
+  EventTrackerRecord::DynamicInstanceID NextDynID = 0;
+  DenseMap<MachineBasicBlock *, std::unique_ptr<EventTracker>> Trackers;
+};
+
+} // namespace eventtracking
+} // namespace AMDGPU
+
+} // namespace llvm
+
+#endif
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUHWEvents.h b/llvm/lib/Target/AMDGPU/AMDGPUHWEvents.h
index eb25206b5ee04..b4ce2fa93eb16 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUHWEvents.h
+++ b/llvm/lib/Target/AMDGPU/AMDGPUHWEvents.h
@@ -165,6 +165,33 @@ class HWEvents {
   value_type Data = NONE;
 };
 
+/// Class to store singular HW events in the most compact way possible (u8).
+class SingleHWEvent {
+public:
+  using value_type = uint8_t;
+
+  static constexpr SingleHWEvent encode(HWEvents E) {
+    assert(E.size() == 1 && "expected exactly one event!");
+    return SingleHWEvent(countr_zero_constexpr(E.value()));
+  }
+
+  constexpr operator HWEvents() const { return HWEvents(1 << Data); }
+
+  constexpr value_type rawValue() const { return Data; }
+
+  constexpr bool operator==(const SingleHWEvent &Other) const {
+    return Data == Other.Data;
+  }
+  constexpr bool operator!=(const SingleHWEvent &Other) const {
+    return Data != Other.Data;
+  }
+
+private:
+  constexpr SingleHWEvent(value_type Data) : Data(Data) {}
+
+  value_type Data;
+};
+
 /// \param Inst A VMEM instruction (as per `SIInstrInfo::isVMEM`).
 /// \returns the simplified set of events triggered by the VMEM instruction \p
 /// Inst. The returned mask is not exhaustive, but is guaranteed to be a subset
diff --git a/llvm/lib/Target/AMDGPU/CMakeLists.txt b/llvm/lib/Target/AMDGPU/CMakeLists.txt
index b7e679a69a80d..07aa3aa33bc1e 100644
--- a/llvm/lib/Target/AMDGPU/CMakeLists.txt
+++ b/llvm/lib/Target/AMDGPU/CMakeLists.txt
@@ -54,6 +54,7 @@ add_llvm_target(AMDGPUCodeGen
   AMDGPUCodeGenPrepare.cpp
   AMDGPUCombinerHelper.cpp
   AMDGPUCtorDtorLowering.cpp
+  AMDGPUEventTracking.cpp
   AMDGPUExportClustering.cpp
   AMDGPUExportKernelRuntimeHandles.cpp
   AMDGPUFrameLowering.cpp
diff --git a/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp b/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
new file mode 100644
index 0000000000000..e00788f2386bf
--- /dev/null
+++ b/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
@@ -0,0 +1,618 @@
+//===- AMDGPUEventTrackingTest.cpp ------------------------------*- C++ -*-===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+#include "AMDGPUEventTracking.h"
+#include "AMDGPUUnitTests.h"
+#include "AMDGPUWaitcntUtils.h"
+#include "MCTargetDesc/AMDGPUMCTargetDesc.h"
+#include "llvm/CodeGen/MachineFunction.h"
+#include "gtest/gtest.h"
+
+using namespace llvm;
+using namespace llvm::AMDGPU;
+using namespace llvm::AMDGPU::eventtracking;
+
+namespace {
+
+static constexpr unsigned CounterLimit = 12;
+
+// These are not accurate, they are simply for testing purposes.
+// We do not need to test every single counter accurately, that is the job
+// of the IR/MIR tests in tests/CodeGen/AMDGPU. We just need enough here
+// to validate that the EventTracker works.
+std::array<CounterInfo, 4> GFX12CounterInfos = {{
+    {LOAD_CNT, HWEvents::VMEM_READ_ACCESS, CounterLimit},
+    {DS_CNT, HWEvents::LDS_ACCESS, CounterLimit},
+    {EXP_CNT, HWEvents::EXP_GPR_LOCK, CounterLimit},
+    {STORE_CNT, HWEvents::VMEM_WRITE_ACCESS | HWEvents::SCRATCH_WRITE_ACCESS,
+     CounterLimit},
+}};
+
+class AMDGPUGFX12EventTrackingTest : public AMDGPUCodeGenTestBase {
+public:
+  void SetUp() override { setUpImpl("amdgpu12.00-amd-amdhsa", "gfx1200", ""); }
+};
+
+namespace {
+static EventTracker &
+visitAll(EventTrackingContext &Ctx, MachineBasicBlock &MBB,
+         function_ref<void(EventTracker &ET)> AfterVisit = nullptr) {
+  const GCNSubtarget &ST = MBB.getParent()->getSubtarget<GCNSubtarget>();
+
+  EventTracker &ET = Ctx[&MBB];
+  ET.enterBlock();
+  for (MachineInstr &MI : MBB) {
+    HWEvents Events =
+        getEventsFor(MI, ST, /*IsExpertMode=*/false, /*TgSplit=*/false);
+    for (HWEvents SingleEv : Events) {
+      ET.record(MI, SingleHWEvent::encode(SingleEv));
+    }
+  }
+  if (AfterVisit)
+    AfterVisit(ET);
+  ET.leaveBlock();
+  return ET;
+}
+
+/// Provides some helpers to declaratively check the state of the EventTracker
+/// for one counter. This provides helpers to check the general counter state
+/// (value, etc) but also allows iterating over the timeline from the oldest to
+/// the earliest element.
+///
+/// This makes the actual test cases clearer.
+struct TrackerRecordsChecker {
+  TrackerRecordsChecker(EventTracker &ET, InstCounterType T)
+      : ET(ET), T(T), Records(ET.getLiveRecords(T)) {}
+
+  bool hasCount() { return ET.count(T).has_value(); }
+
+  unsigned getCount() { return *ET.count(T); }
+
+  bool empty() { return Records.empty(); }
+
+  // Has no records and score is zero (if there is one)
+  bool unused() { return empty() && (!hasCount() || !getCount()); }
+
+  const EventTrackerRecord &cur() { return Records[CurElt]; }
+
+  /// Move to the next record.
+  /// \returns true on success, false if the end has been reached.
+  bool next() {
+    ++CurElt;
+    if (CurElt >= Records.size())
+      return false;
+    return true;
+  }
+
+  EventTracker &ET;
+  InstCounterType T;
+  ArrayRef<EventTrackerRecord> Records;
+  unsigned CurElt = 0;
+};
+
+} // namespace
+
+/// Check trivial straight line counting.
+TEST_F(AMDGPUGFX12EventTrackingTest, BasicTimeline) {
+  StringRef MIR = R"(
+name:            BasicTimeline
+body:             |
+  bb.0:
+
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    $vgpr4 = GLOBAL_LOAD_DWORD $vgpr4_vgpr5, 0, 0, implicit $exec
+    $vgpr0 = DS_READ_B32_gfx9 $vgpr1, 0, 0, implicit $exec
+    GLOBAL_STORE_DWORDX2 $vgpr1_vgpr2, $vgpr3_vgpr4, 0, 0, implicit $exec
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("BasicTimeline");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  auto &ET = visitAll(Ctx, BB0);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.hasCount());
+  EXPECT_EQ(LoadCnt.getCount(), 1u);
+  EXPECT_FALSE(LoadCnt.empty());
+  EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
+  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(LoadCnt.next());
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.hasCount());
+  EXPECT_EQ(DsCnt.getCount(), 1u);
+  EXPECT_FALSE(DsCnt.empty());
+  EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
+  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(DsCnt.next());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+  EXPECT_TRUE(StoreCnt.hasCount());
+  EXPECT_EQ(StoreCnt.getCount(), 2u);
+  EXPECT_FALSE(StoreCnt.empty());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 1u);
+  EXPECT_TRUE(StoreCnt.next());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(StoreCnt.next());
+}
+
+/// Check cases where a block loops back on itself.
+TEST_F(AMDGPUGFX12EventTrackingTest, SelfPredecessor) {
+  StringRef MIR = R"(
+name:            BasicTimeline
+body:             |
+  bb.0:
+
+    $vgpr4 = GLOBAL_LOAD_DWORD $vgpr4_vgpr5, 0, 0, implicit $exec
+    $vgpr0 = DS_READ_B32_gfx9 $vgpr1, 0, 0, implicit $exec
+    GLOBAL_STORE_DWORDX2 $vgpr1_vgpr2, $vgpr3_vgpr4, 0, 0, implicit $exec
+    S_CBRANCH_SCC1 %bb.0, implicit $scc
+    S_BRANCH %bb.1
+
+  bb.1:
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("BasicTimeline");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  // Iterate twice
+  visitAll(Ctx, BB0);
+  auto &ET = visitAll(Ctx, BB0);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.hasCount());
+  EXPECT_EQ(LoadCnt.getCount(), 2u);
+  EXPECT_FALSE(LoadCnt.empty());
+  EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
+  EXPECT_EQ(LoadCnt.cur().getScore(), 1u);
+  EXPECT_TRUE(LoadCnt.next());
+  EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
+  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(LoadCnt.next());
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.hasCount());
+  EXPECT_EQ(DsCnt.getCount(), 2u);
+  EXPECT_FALSE(DsCnt.empty());
+  EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
+  EXPECT_EQ(DsCnt.cur().getScore(), 1u);
+  EXPECT_TRUE(DsCnt.next());
+  EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
+  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(DsCnt.next());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+  EXPECT_TRUE(StoreCnt.hasCount());
+  EXPECT_EQ(StoreCnt.getCount(), 2u);
+  EXPECT_FALSE(StoreCnt.empty());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 1u);
+  EXPECT_TRUE(StoreCnt.next());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(StoreCnt.next());
+}
+
+/// Check all events carry into the next incoming block.
+TEST_F(AMDGPUGFX12EventTrackingTest, SingleIncomingBlock) {
+  StringRef MIR = R"(
+name:            SingleIncomingBlock
+body:             |
+  bb.0:
+    successors: %bb.1
+
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    $vgpr4 = GLOBAL_LOAD_DWORD $vgpr4_vgpr5, 0, 0, implicit $exec
+    $vgpr0 = DS_READ_B32_gfx9 $vgpr1, 0, 0, implicit $exec
+    GLOBAL_STORE_DWORDX2 $vgpr1_vgpr2, $vgpr3_vgpr4, 0, 0, implicit $exec
+
+  bb.1:
+
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("SingleIncomingBlock");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+  MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  visitAll(Ctx, BB0);
+  auto &ET = visitAll(Ctx, BB1);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.hasCount());
+  EXPECT_EQ(LoadCnt.getCount(), 1u);
+  EXPECT_FALSE(LoadCnt.empty());
+  EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
+  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(LoadCnt.next());
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.hasCount());
+  EXPECT_EQ(DsCnt.getCount(), 1u);
+  EXPECT_FALSE(DsCnt.empty());
+  EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
+  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(DsCnt.next());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+  EXPECT_TRUE(StoreCnt.hasCount());
+  EXPECT_EQ(StoreCnt.getCount(), 2u);
+  EXPECT_FALSE(StoreCnt.empty());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 1u);
+  EXPECT_TRUE(StoreCnt.next());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(StoreCnt.next());
+}
+
+/// Check a simple merge where all events differ in each incoming block.
+TEST_F(AMDGPUGFX12EventTrackingTest, SimpleDisjointMerge) {
+  StringRef MIR = R"(
+name:            SimpleDisjointMerge
+body:             |
+  bb.0:
+    successors: %bb.2
+
+    $vgpr0 = DS_READ_B32_gfx9 $vgpr1, 0, 0, implicit $exec
+    $vgpr4 = GLOBAL_LOAD_DWORD $vgpr4_vgpr5, 0, 0, implicit $exec
+    S_BRANCH %bb.2
+
+  bb.1:
+    successors: %bb.2
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    GLOBAL_STORE_DWORDX2 $vgpr1_vgpr2, $vgpr3_vgpr4, 0, 0, implicit $exec
+    S_BRANCH %bb.2
+
+  bb.2:
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("SimpleDisjointMerge");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+  MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
+  MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  visitAll(Ctx, BB0);
+  visitAll(Ctx, BB1);
+  auto &ET = visitAll(Ctx, BB2);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.hasCount());
+  EXPECT_EQ(LoadCnt.getCount(), 1u);
+  EXPECT_FALSE(LoadCnt.empty());
+  EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
+  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(LoadCnt.next());
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.hasCount());
+  EXPECT_EQ(DsCnt.getCount(), 1u);
+  EXPECT_FALSE(DsCnt.empty());
+  EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
+  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(DsCnt.next());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+  EXPECT_TRUE(StoreCnt.hasCount());
+  EXPECT_EQ(StoreCnt.getCount(), 2u);
+  EXPECT_FALSE(StoreCnt.empty());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 1u);
+  EXPECT_TRUE(StoreCnt.next());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(StoreCnt.next());
+}
+
+/// Merge with divergent events in a counter.
+TEST_F(AMDGPUGFX12EventTrackingTest, DivergenceMerge) {
+  StringRef MIR = R"(
+name:            DivergenceMerge
+body:             |
+  bb.0:
+    successors: %bb.2
+
+    $vgpr0 = DS_READ_B32_gfx9 $vgpr1, 0, 0, implicit $exec
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    S_BRANCH %bb.2
+
+  bb.1:
+    successors: %bb.2
+    $vgpr4 = GLOBAL_LOAD_DWORD $vgpr4_vgpr5, 0, 0, implicit $exec
+    GLOBAL_STORE_DWORDX2 $vgpr1_vgpr2, $vgpr3_vgpr4, 0, 0, implicit $exec
+    S_BRANCH %bb.2
+
+  bb.2:
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("DivergenceMerge");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+  MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
+  MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  visitAll(Ctx, BB0);
+  visitAll(Ctx, BB1);
+  auto &ET = visitAll(Ctx, BB2);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.hasCount());
+  EXPECT_EQ(LoadCnt.getCount(), 1u);
+  EXPECT_FALSE(LoadCnt.empty());
+  EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
+  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(LoadCnt.next());
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.hasCount());
+  EXPECT_EQ(DsCnt.getCount(), 1u);
+  EXPECT_FALSE(DsCnt.empty());
+  EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
+  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(DsCnt.next());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+  EXPECT_TRUE(StoreCnt.hasCount());
+  // Count is 1 because we have 1 event max across all predecessors.
+  EXPECT_EQ(StoreCnt.getCount(), 1u);
+  EXPECT_FALSE(StoreCnt.empty());
+  // First successor has GLOBAL_STORE_DWORD at height 0
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_TRUE(StoreCnt.next());
+  // Second successor has GLOBAL_STORE_DWORDX2 at height 0 too
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(StoreCnt.next());
+}
+
+/// Basic diamond CFG, the store is carried all the way into bb3 and uniqued
+/// again so only 1 instance of the record is present in bb3.
+TEST_F(AMDGPUGFX12EventTrackingTest, BasicDiamond) {
+  StringRef MIR = R"(
+name:            BasicDiamond
+body:             |
+  bb.0:
+    successors: %bb.1, %bb.2
+
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    S_CBRANCH_SCC1 %bb.1, implicit $scc
+    S_BRANCH %bb.2
+
+  bb.1:
+    successors: %bb.3
+    $vgpr0 = DS_READ_B32_gfx9 $vgpr1, 0, 0, implicit $exec
+    S_BRANCH %bb.3
+
+  bb.2:
+    successors: %bb.3
+    $vgpr4 = GLOBAL_LOAD_DWORD $vgpr4_vgpr5, 0, 0, implicit $exec
+    S_BRANCH %bb.3
+
+  bb.3:
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("BasicDiamond");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+  MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
+  MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
+  MachineBasicBlock &BB3 = *MF.getBlockNumbered(3);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  visitAll(Ctx, BB0);
+  visitAll(Ctx, BB1);
+  visitAll(Ctx, BB2);
+  auto &ET = visitAll(Ctx, BB3);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.hasCount());
+  EXPECT_EQ(LoadCnt.getCount(), 1u);
+  EXPECT_FALSE(LoadCnt.empty());
+  EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
+  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(LoadCnt.next());
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.hasCount());
+  EXPECT_EQ(DsCnt.getCount(), 1u);
+  EXPECT_FALSE(DsCnt.empty());
+  EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
+  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(DsCnt.next());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+  EXPECT_TRUE(StoreCnt.hasCount());
+  EXPECT_EQ(StoreCnt.getCount(), 1u);
+  EXPECT_FALSE(StoreCnt.empty());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(StoreCnt.next());
+}
+
+/// Basic diamond CFG, the store is carried all the way into bb3 and uniqued
+/// again so only 1 instance of the record is present in bb3.
+TEST_F(AMDGPUGFX12EventTrackingTest, BasicDiamondFlags) {
+  StringRef MIR = R"(
+name:            BasicDiamondFlags
+body:             |
+  bb.0:
+    successors: %bb.1, %bb.2
+
+    S_CBRANCH_SCC1 %bb.1, implicit $scc
+    S_BRANCH %bb.2
+
+  bb.1:
+    successors: %bb.3
+    S_BRANCH %bb.3
+
+  bb.2:
+    successors: %bb.3
+    S_BRANCH %bb.3
+
+  bb.3:
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("BasicDiamondFlags");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+  MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
+  MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
+  MachineBasicBlock &BB3 = *MF.getBlockNumbered(3);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  visitAll(Ctx, BB0);
+  visitAll(Ctx, BB1, /*AfterVisit=*/[&](EventTracker &ET) {
+    ET.markIndeterminate(STORE_CNT);
+  });
+  visitAll(Ctx, BB2, [&](EventTracker &ET) { ET.markOutOfOrder(LOAD_CNT); });
+  auto &ET = visitAll(Ctx, BB3);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.unused());
+  // We inherit the out of order flag is either predecessor has it.
+  EXPECT_TRUE(LoadCnt.ET.isOutOfOrder(LOAD_CNT));
+  EXPECT_FALSE(LoadCnt.ET.isIndeterminate(LOAD_CNT));
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.unused());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+  EXPECT_TRUE(StoreCnt.unused());
+  // We inherit the indeterminate flag is either predecessor has it.
+  EXPECT_TRUE(StoreCnt.ET.isIndeterminate(STORE_CNT));
+}
+
+/// Assymetrical diamond
+///   - bb0 has two stores
+///   - bb1 adds another store without any waits.
+///   - bb2 adds a store and waits on the two stores from bb0 afterwards
+///
+/// The timeline will be:
+///   - Stores from bb0 exist at 1/2
+///   - The added store from bb1 exist at score 0
+///   - The added store from bb2 exist at score 0
+TEST_F(AMDGPUGFX12EventTrackingTest, AssymetricalDiamond) {
+  StringRef MIR = R"(
+name:            AssymetricalDiamond
+body:             |
+  bb.0:
+    successors: %bb.1, %bb.2
+
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    GLOBAL_STORE_DWORDX2 $vgpr1_vgpr2, $vgpr3_vgpr4, 0, 0, implicit $exec
+    S_CBRANCH_SCC1 %bb.1, implicit $scc
+    S_BRANCH %bb.2
+
+  bb.1:
+    successors: %bb.3
+    GLOBAL_STORE_DWORD $vgpr1_vgpr2, $vgpr3, 0, 0, implicit $exec
+    S_BRANCH %bb.3
+
+  bb.2:
+    successors: %bb.3
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    S_BRANCH %bb.3
+
+  bb.3:
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("AssymetricalDiamond");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+  MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
+  MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
+  MachineBasicBlock &BB3 = *MF.getBlockNumbered(3);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  visitAll(Ctx, BB0);
+  visitAll(Ctx, BB1);
+  visitAll(Ctx, BB2,
+           /*AfterVisit=*/[&](EventTracker &ET) { ET.wait(STORE_CNT, 1); });
+  auto &ET = visitAll(Ctx, BB3);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.unused());
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.unused());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+
+  EXPECT_TRUE(StoreCnt.hasCount());
+  EXPECT_EQ(StoreCnt.getCount(), 3u);
+  EXPECT_FALSE(StoreCnt.empty());
+
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB0);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 2u);
+  EXPECT_TRUE(StoreCnt.next());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 1u);
+  EXPECT_TRUE(StoreCnt.next());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_TRUE(StoreCnt.next());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB1);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(StoreCnt.next());
+}
+
+} // namespace
diff --git a/llvm/unittests/Target/AMDGPU/CMakeLists.txt b/llvm/unittests/Target/AMDGPU/CMakeLists.txt
index 2bb9b6dedba13..882ceef9f0310 100644
--- a/llvm/unittests/Target/AMDGPU/CMakeLists.txt
+++ b/llvm/unittests/Target/AMDGPU/CMakeLists.txt
@@ -23,6 +23,7 @@ set(LLVM_LINK_COMPONENTS
   )
 
 add_llvm_target_unittest(AMDGPUTests
+  AMDGPUEventTrackingTest.cpp
   AMDGPUMCExprTest.cpp
   AMDGPUUnitTests.cpp
   CSETest.cpp

>From 74d8ea44faa105097a35439a4f33d38fab4d693e Mon Sep 17 00:00:00 2001
From: pvanhout <pierre.vanhoutryve at amd.com>
Date: Mon, 28 Sep 2026 15:46:04 +0200
Subject: [PATCH 2/5] Comments

---
 .../lib/Target/AMDGPU/AMDGPUEventTracking.cpp | 82 +++++++++----------
 llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h  | 10 ---
 .../Target/AMDGPU/AMDGPUEventTrackingTest.cpp |  4 +-
 3 files changed, 43 insertions(+), 53 deletions(-)

diff --git a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
index 8b383b45de67a..51fd04df503ce 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
@@ -53,14 +53,14 @@ void EventTrackerRecord::print(raw_ostream &OS, bool PrintMI,
   if (PrintMI && MI)
     OS << *MI;
   else
-    OS << "MI@" << (void *)MI << "\n";
+    OS << "MI@" << (void *)MI << '\n';
 }
 
 #if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
 LLVM_DUMP_METHOD void EventTrackerRecord::dump() const {
-  dbgs() << "\n";
+  dbgs() << '\n';
   print(dbgs(), /*PrintMI=*/true);
-  dbgs() << "\n";
+  dbgs() << '\n';
 }
 #endif
 
@@ -73,7 +73,7 @@ EventTracker::EventTracker(MachineBasicBlock &MBB, EventTrackingContext &ETC)
 
 void EventTracker::enterBlock() {
   LLVM_DEBUG(dbgs() << "\n[EventTracker] Entering ";
-             MBB->printAsOperand(dbgs()); dbgs() << "\n");
+             MBB->printAsOperand(dbgs()); dbgs() << '\n');
 
   // FIXME: This is a bit hacky, but we need to save the old state in case the
   // MBB is also its own predecessor. Revisit when the design and clients of
@@ -103,7 +103,7 @@ void EventTracker::enterBlock() {
 
 void EventTracker::leaveBlock() {
   LLVM_DEBUG(dbgs() << "[EventTracker] Leaving "; MBB->printAsOperand(dbgs());
-             dbgs() << "\n");
+             dbgs() << '\n');
 }
 
 void EventTracker::record(MachineInstr &MI, SingleHWEvent Event) {
@@ -111,7 +111,7 @@ void EventTracker::record(MachineInstr &MI, SingleHWEvent Event) {
 
   EventTrackerRecord Rec = EventTrackerRecord(*Ctx, &MI, Event);
   [[maybe_unused]] bool FoundMatch = false;
-  for (auto &CD : Counters) {
+  for (CounterData &CD : Counters) {
     if (!CD.CI->Events.contains(Event))
       continue;
 
@@ -153,9 +153,9 @@ void EventTracker::record(MachineInstr &MI, SingleHWEvent Event) {
 }
 
 void EventTracker::wait(InstCounterType T, unsigned N) {
-  auto &CD = get(T);
+  CounterData &CD = Counters[T];
   LLVM_DEBUG(dbgs() << "[EventTracker] Wait on " << getInstCounterName(T)
-                    << " for " << N << "\n");
+                    << " for " << N << '\n');
 
   // Fast path for clearing the counter
   if (N == 0) {
@@ -181,7 +181,7 @@ void EventTracker::wait(InstCounterType T, unsigned N) {
     CD.LiveRecords.erase(RmIt, CD.LiveRecords.end());
   }
 
-  LLVM_DEBUG(dbgs() << "  | => Updated Count:" << CD.Count << "\n");
+  LLVM_DEBUG(dbgs() << "  | => Updated Count:" << CD.Count << '\n');
 
 #ifndef NDEBUG
   LLVM_DEBUG(if (EventTrackerPrintAll) {
@@ -194,7 +194,7 @@ void EventTracker::wait(InstCounterType T, unsigned N) {
 void EventTracker::markIndeterminate(InstCounterType T) {
   LLVM_DEBUG(dbgs() << "[EventTracker] Marking " << getInstCounterName(T)
                     << " as indeterminate!\n");
-  auto &CD = get(T);
+  CounterData &CD = Counters[T];
   CD.IsIndeterminate = true;
   markOutOfOrder(T);
 }
@@ -202,29 +202,29 @@ void EventTracker::markIndeterminate(InstCounterType T) {
 void EventTracker::markOutOfOrder(InstCounterType T) {
   LLVM_DEBUG(dbgs() << "[EventTracker] Marking " << getInstCounterName(T)
                     << " as out-of-order!\n");
-  auto &CD = get(T);
+  CounterData &CD = Counters[T];
   CD.IsOutOfOrder = true;
   for (auto &Rec : CD.LiveRecords)
     Rec.setScore(0);
 }
 
 std::optional<unsigned> EventTracker::count(InstCounterType T) const {
-  auto &CD = get(T);
+  const CounterData &CD = Counters[T];
   if (CD.IsIndeterminate)
     return std::nullopt;
-  return get(T).Count;
+  return Counters[T].Count;
 }
 
 bool EventTracker::isIndeterminate(InstCounterType T) const {
-  return get(T).IsIndeterminate;
+  return Counters[T].IsIndeterminate;
 }
 
 bool EventTracker::isOutOfOrder(InstCounterType T) const {
-  return get(T).IsOutOfOrder;
+  return Counters[T].IsOutOfOrder;
 }
 
 HWEvents EventTracker::getPendingEvents(InstCounterType T) const {
-  const auto &CD = get(T);
+  const CounterData &CD = Counters[T];
   if (CD.IsIndeterminate)
     return CD.CI->Events; // return all events
 
@@ -232,18 +232,18 @@ HWEvents EventTracker::getPendingEvents(InstCounterType T) const {
     return CD.LegacyPendingEvents;
 
   HWEvents Res;
-  for (const auto &E : get(T).LiveRecords)
+  for (const auto &E : Counters[T].LiveRecords)
     Res |= E.getKind();
   return Res;
 }
 
 ArrayRef<EventTrackerRecord>
 EventTracker::getLiveRecords(InstCounterType T) const {
-  return get(T).LiveRecords;
+  return Counters[T].LiveRecords;
 }
 
 void EventTracker::print(raw_ostream &OS, InstCounterType T) const {
-  print(OS, get(T));
+  print(OS, Counters[T]);
 }
 
 bool EventTracker::mimicsLegacyTracking() { return MimicLegacyTracking; }
@@ -252,7 +252,7 @@ bool EventTracker::mimicsLegacyTracking() { return MimicLegacyTracking; }
 void EventTracker::verify() const {
   assert(MBB && Ctx && "Invalid internal state!");
 
-  for (auto &C : Counters) {
+  for (const CounterData &C : Counters) {
     const auto OnError = [&]() {
       dbgs() << "EventTracker verification error\n";
       print(dbgs(), C);
@@ -261,7 +261,7 @@ void EventTracker::verify() const {
     if (C.IsIndeterminate) {
       if (!C.IsOutOfOrder) {
         OnError();
-        assert(false && "IsIndeterminate but not IsOutOfOrder");
+        llvm_unreachable("IsIndeterminate but not IsOutOfOrder");
       }
       continue;
     }
@@ -269,30 +269,30 @@ void EventTracker::verify() const {
     if (C.IsOutOfOrder) {
       if (!all_of(C.LiveRecords, [](auto &R) { return R.getScore() == 0; })) {
         OnError();
-        assert(false &&
-               "IsOutOfOrder but some records do not have a score of 0!");
+        llvm_unreachable(
+            "IsOutOfOrder but some records do not have a score of 0!");
       }
     }
 
     if (C.Count > C.LiveRecords.size() && !mimicsLegacyTracking()) {
       OnError();
-      assert(false &&
-             "'Count' is inconsistent with the number of live records");
+      llvm_unreachable(
+          "'Count' is inconsistent with the number of live records");
     }
 
-    for (auto &E : C.LiveRecords) {
+    for (const EventTrackerRecord &E : C.LiveRecords) {
       if (E.getScore() > C.Count) {
         OnError();
         dbgs() << "Concerning Record:";
         E.print(dbgs());
-        assert(false && "record score is out of range");
+        llvm_unreachable("record score is out of range");
       }
     }
 
     // Check live records are sorted
     if (!is_sorted(C.LiveRecords, greaterThan)) {
       OnError();
-      assert(false && "live records are not sorted!");
+      llvm_unreachable("live records are not sorted!");
     }
   }
 }
@@ -302,16 +302,16 @@ void EventTracker::print(raw_ostream &OS, bool IgnoreEmpty,
                          unsigned Indent) const {
   OS.indent(Indent) << "EventTracker for ";
   MBB->printAsOperand(OS);
-  OS << ":";
+  OS << ':';
   if (IgnoreEmpty) {
-    if (all_of(Counters, [](const auto &CD) { return CD.Count == 0; })) {
+    if (all_of(Counters, [](const CounterData &CD) { return CD.Count == 0; })) {
       OS << " (empty)\n";
       return;
     }
   }
 
-  OS << "\n";
-  for (auto &C : Counters) {
+  OS << '\n';
+  for (const CounterData &C : Counters) {
     if (IgnoreEmpty && C.Count == 0)
       continue;
     print(dbgs(), C, Indent + 2);
@@ -320,14 +320,14 @@ void EventTracker::print(raw_ostream &OS, bool IgnoreEmpty,
 
 #if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
 LLVM_DUMP_METHOD void EventTracker::dump() const {
-  dbgs() << "\n";
+  dbgs() << '\n';
   print(dbgs());
-  dbgs() << "\n";
+  dbgs() << '\n';
 }
 #endif
 
 void EventTracker::clear() {
-  for (auto &C : Counters) {
+  for (CounterData &C : Counters) {
     C.LiveRecords.clear();
     C.Count = 0;
     C.IsIndeterminate = false;
@@ -346,7 +346,7 @@ void EventTracker::recordIncomings(EventTrackingContext &ETC,
   });
 
   /// Iterate over all counters that are available to us.
-  for (auto &CI : ETC.counters()) {
+  for (const CounterInfo &CI : ETC.counters()) {
     auto &CData = Counters[CI.CounterT];
     assert(CData.LiveRecords.empty());
 
@@ -407,7 +407,7 @@ void EventTracker::print(raw_ostream &OS, const CounterData &CD,
                     << ", LiveRecords=" << CD.LiveRecords.size()
                     << ", IsOutOfOrder=" << CD.IsOutOfOrder
                     << ", IsIndeterminate=" << CD.IsIndeterminate << ")\n";
-  for (const auto &E : CD.LiveRecords) {
+  for (const EventTrackerRecord &E : CD.LiveRecords) {
     OS.indent(Indent + 2);
     E.print(OS);
   }
@@ -417,7 +417,7 @@ EventTrackingContext::EventTrackingContext(MachineFunction &MF,
                                            ArrayRef<CounterInfo> Counters)
     : CounterInfos(Counters) {
   LLVM_DEBUG(dbgs() << "\n[EventTrackingContext] CounterInfos for "
-                    << MF.getName() << "\n";
+                    << MF.getName() << '\n';
              for (const auto &CI
                   : CounterInfos) {
                dbgs().indent(2)
@@ -425,7 +425,7 @@ EventTrackingContext::EventTrackingContext(MachineFunction &MF,
                if (CI.Events.none()) {
                  dbgs() << " (unused - no HWEvents assigned)\n";
                } else {
-                 dbgs() << "(Limit=" << CI.Limit << ") " << CI.Events << "\n";
+                 dbgs() << "(Limit=" << CI.Limit << ") " << CI.Events << '\n';
                }
              });
 
@@ -465,9 +465,9 @@ void EventTrackingContext::print(raw_ostream &OS) const {
 
 #if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
 LLVM_DUMP_METHOD void EventTrackingContext::dump() const {
-  dbgs() << "\n";
+  dbgs() << '\n';
   print(dbgs());
-  dbgs() << "\n";
+  dbgs() << '\n';
 }
 #endif
 
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
index c1e56fc184569..4a717f794c7a7 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
+++ b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
@@ -272,16 +272,6 @@ class EventTracker {
     // but a waitcnt 8 for another. Not sure if we can exploit that though?
   };
 
-  CounterData &get(InstCounterType T) {
-    assert(Counters.size() > T && "T is out of range!");
-    return Counters[T];
-  }
-
-  const CounterData &get(InstCounterType T) const {
-    assert(Counters.size() > T && "T is out of range!");
-    return Counters[T];
-  }
-
   /// Clears the tracked data, used when entering a block.
   void clear();
 
diff --git a/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp b/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
index e00788f2386bf..ac168c9465a8e 100644
--- a/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
+++ b/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
@@ -25,7 +25,7 @@ static constexpr unsigned CounterLimit = 12;
 // We do not need to test every single counter accurately, that is the job
 // of the IR/MIR tests in tests/CodeGen/AMDGPU. We just need enough here
 // to validate that the EventTracker works.
-std::array<CounterInfo, 4> GFX12CounterInfos = {{
+static constexpr std::array<CounterInfo, 4> GFX12CounterInfos = {{
     {LOAD_CNT, HWEvents::VMEM_READ_ACCESS, CounterLimit},
     {DS_CNT, HWEvents::LDS_ACCESS, CounterLimit},
     {EXP_CNT, HWEvents::EXP_GPR_LOCK, CounterLimit},
@@ -35,7 +35,7 @@ std::array<CounterInfo, 4> GFX12CounterInfos = {{
 
 class AMDGPUGFX12EventTrackingTest : public AMDGPUCodeGenTestBase {
 public:
-  void SetUp() override { setUpImpl("amdgpu12.00-amd-amdhsa", "gfx1200", ""); }
+  void SetUp() override { setUpImpl("amdgpu12.00-amd-amdhsa", "", ""); }
 };
 
 namespace {

>From 10c8693f2fc6bfedadbe58537bb2a3c4295ec2e2 Mon Sep 17 00:00:00 2001
From: pvanhout <pierre.vanhoutryve at amd.com>
Date: Tue, 29 Sep 2026 11:28:36 +0200
Subject: [PATCH 3/5] Address some comments

---
 .../lib/Target/AMDGPU/AMDGPUEventTracking.cpp | 11 +--
 llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h  | 66 ++++-------------
 .../Target/AMDGPU/AMDGPUEventTrackingTest.cpp | 73 +++++++++++++++++++
 3 files changed, 91 insertions(+), 59 deletions(-)

diff --git a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
index 51fd04df503ce..bd491bb74a326 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
@@ -35,13 +35,6 @@ static cl::opt<bool> EventTrackerPrintAll(
 namespace AMDGPU {
 namespace eventtracking {
 
-namespace {
-static bool greaterThan(const EventTrackerRecord &A,
-                        const EventTrackerRecord &B) {
-  return A.getScore() > B.getScore();
-}
-} // namespace
-
 EventTrackerRecord::EventTrackerRecord(EventTrackingContext &Ctx,
                                        MachineInstr *MI, SingleHWEvent Kind,
                                        uint32_t Score)
@@ -290,7 +283,7 @@ void EventTracker::verify() const {
     }
 
     // Check live records are sorted
-    if (!is_sorted(C.LiveRecords, greaterThan)) {
+    if (!is_sorted(C.LiveRecords, EventTrackerRecord::isScoreGreaterThan)) {
       OnError();
       llvm_unreachable("live records are not sorted!");
     }
@@ -390,7 +383,7 @@ void EventTracker::recordIncomings(EventTrackingContext &ETC,
     CData.LiveRecords.append(AccVals.begin(), AccVals.end());
 
     // Sort records by Score (descending) for consistent iteration.
-    stable_sort(CData.LiveRecords, greaterThan);
+    stable_sort(CData.LiveRecords, EventTrackerRecord::isScoreGreaterThan);
   }
 
   LLVM_DEBUG(if (!Preds.empty()) {
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
index 4a717f794c7a7..627b29bae06e3 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
+++ b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
@@ -49,24 +49,18 @@ struct CounterInfo {
   unsigned Limit;
 };
 
-/// Represents entries in the \ref EventTracker.
-///
 /// Records are all uniquely identified by a \ref DynamicInstanceID. For each
 /// unique \ref DynamicInstanceID value, all \ref EventTrackerRecord that use
 /// that ID should have the same MI and Kind. This is enforced by exposing these
 /// as read-only, and making the constructor assign a new \ref DynamicInstanceID
-/// every time.
-///
-/// Only the score can change as it may be unique to each instance
+/// every time. Only the score can change as it may be unique to each instance
 /// of \ref EventTracker that carry it.
 class EventTrackerRecord {
 public:
   /// An always-increasing counter used to represent a dynamic instance of a
-  /// record.
-  ///
-  /// Whenever we add a new \ref EventTrackerRecord, even if it's one we already
-  /// have seen in a previous dataflow iteration, this counter is increased so
-  /// that the new record has a unique `DynamicInstanceID`.
+  /// record. Whenever we add a new \ref EventTrackerRecord, even if it's one we
+  /// already have seen in a previous dataflow iteration, this counter is
+  /// increased so that the new record has a unique `DynamicInstanceID`.
   using DynamicInstanceID = uint32_t;
 
   EventTrackerRecord(EventTrackingContext &Ctx, MachineInstr *MI,
@@ -100,6 +94,11 @@ class EventTrackerRecord {
   LLVM_DUMP_METHOD void dump() const;
 #endif
 
+  static bool isScoreGreaterThan(const EventTrackerRecord &A,
+                                 const EventTrackerRecord &B) {
+    return A.getScore() > B.getScore();
+  }
+
 private:
   EventTrackerRecord(DynamicInstanceID ID, MachineInstr *MI, SingleHWEvent Kind,
                      uint32_t Score = 0)
@@ -121,26 +120,17 @@ static_assert(sizeof(EventTrackerRecord) == 16,
               "EventTrackerRecord should remain small to optimize its layout "
               "within cache lines, for maximum iteration speed");
 
-/// Per-MBB tracking context.
-///
-/// Tracks data across the following domains:
-///   - Current value (count) of each instruction counter.
-///   - In-flight (alive) \ref EventTrackerRecord of each instruction counter.
-///
-/// This class is only responsible for tracking records for every
-/// InstCounterType. It does not deal with calculating the waitcnts needed, or
-/// doing more advance reasoning over the timeline for specific queries (e.g.
-/// finding an aliasing store). These responsibilities are for
-/// utils/wrappers/users of the class.
-///
-/// The API should be kept as simple and clear as possible.
+/// Per-MBB tracking context which tracks the current value (count) of each
+/// instruction counter and the set of in-flight (alive) \ref EventTrackerRecord
+/// of each instruction counter. This class is only responsible for tracking
+/// records for every InstCounterType. It does not deal with calculating the
+/// waitcnts needed, or doing more advance reasoning over the timeline for
+/// specific queries (e.g. finding an aliasing store). These responsibilities
+/// are for utils/wrappers/users of the class.
 class EventTracker {
 public:
   EventTracker(MachineBasicBlock &MBB, EventTrackingContext &ET);
 
-  /// \defgroup MachineBasicBlock entry and exit
-  /// \{
-
   /// Notify this EventTracker that we are going to begin recording events.
   /// In case this is not the first time we are going through this block, this
   /// clears the internal state of the tracker and re-imports all incoming
@@ -150,23 +140,14 @@ class EventTracker {
   /// Notify this EventTracker that we are done recording events.
   void leaveBlock();
 
-  /// \}
-
-  /// \defgroup InstCounters Tracking Entrypoints
-  /// Methods update the state of the InstCounters by adding/removing events
-  /// or signaling certain special conditions.
-  /// \{
-
   /// Record an event of type \p Event at a MachineInstr \p MI, which will
   /// affect all counters that have \p Event in their event set.
   void record(MachineInstr &MI, SingleHWEvent Event);
 
   /// Notify that we waited until the counter \p T reached the value \p N before
   /// continuing execution of the program (and recording more events).
-  ///
   /// This affects the count of \p T, an removes all records that have a score
   /// greater than or equal to \p N.
-  ///
   /// If \p N is zero, then \p T will no longer be in an indeterminate or
   /// out-of-order state afterwards if it previously was in such a state.
   void wait(InstCounterType T, unsigned N = 0);
@@ -175,7 +156,6 @@ class EventTracker {
   /// we no longer accurately track \p T because there may be more records we do
   /// not know about. This primarily affects \ref getPendingEvents and
   /// \ref count.
-  ///
   /// Implies \ref markOutOfOrder for \p T as well.
   void markIndeterminate(InstCounterType T);
 
@@ -183,12 +163,6 @@ class EventTracker {
   /// in any order. This sets the score of all records to zero.
   void markOutOfOrder(InstCounterType T);
 
-  /// \}
-
-  /// \defgroup InstCounters Tracking Queries
-  /// Query the current state of each InstCounter without modifying it.
-  /// \{
-
   /// \returns the current value of the counter \p T at this point in time, or
   /// std::nullopt if \p T is in the indeterminate state.
   std::optional<unsigned> count(InstCounterType T) const;
@@ -209,11 +183,6 @@ class EventTracker {
   /// only contains the records this class knows about.
   ArrayRef<EventTrackerRecord> getLiveRecords(InstCounterType T) const;
 
-  /// \}
-
-  /// \defgroup Miscellaneous helpers
-  /// \{
-
   /// Prints a dump of the internal tracking state of this class for \p T to the
   /// stream \p OS.
   void print(raw_ostream &OS, InstCounterType T) const;
@@ -237,8 +206,6 @@ class EventTracker {
   LLVM_DUMP_METHOD void dump() const;
 #endif
 
-  /// \}
-
 private:
   struct CounterData {
     const CounterInfo *CI = nullptr;
@@ -293,7 +260,6 @@ class EventTracker {
 };
 
 /// Per-MF Tracking Context.
-///
 /// This owns all \ref EventTrackers and keeps track of state that persists
 /// across dataflow analysis iterations, such as the current value of
 /// \ref DynamicInstanceID.
diff --git a/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp b/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
index ac168c9465a8e..565a54ffdd99e 100644
--- a/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
+++ b/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
@@ -615,4 +615,77 @@ body:             |
   EXPECT_FALSE(StoreCnt.next());
 }
 
+/// Assymetrical diamond
+///   - bb0 has one store
+///   - bb1 just falls through
+///   - bb2 adds 1 more store.
+///
+/// The timeline will showcase how the score of a counter can be different from
+/// the content of the timeline.
+///   - The timeline will have all records at a height of 0.
+///   - The counter will still be at 2.
+TEST_F(AMDGPUGFX12EventTrackingTest, AssymetricalDiamond2) {
+  StringRef MIR = R"(
+name:            AssymetricalDiamond2
+body:             |
+  bb.0:
+    successors: %bb.1, %bb.2
+
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    S_CBRANCH_SCC1 %bb.1, implicit $scc
+    S_BRANCH %bb.2
+
+  bb.1:
+    successors: %bb.3
+    S_BRANCH %bb.3
+
+  bb.2:
+    successors: %bb.3
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    S_BRANCH %bb.3
+
+  bb.3:
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("AssymetricalDiamond2");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+  MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
+  MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
+  MachineBasicBlock &BB3 = *MF.getBlockNumbered(3);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  visitAll(Ctx, BB0);
+  visitAll(Ctx, BB1);
+  visitAll(Ctx, BB2);
+
+  auto &ET = visitAll(Ctx, BB3);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.unused());
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.unused());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+
+  EXPECT_TRUE(StoreCnt.hasCount());
+  EXPECT_EQ(StoreCnt.getCount(), 2u);
+  EXPECT_FALSE(StoreCnt.empty());
+
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_TRUE(StoreCnt.next());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB0);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(StoreCnt.next());
+}
+
 } // namespace

>From 1899f5ab8b41a9535a85cbc52d916a493bd76213 Mon Sep 17 00:00:00 2001
From: pvanhout <pierre.vanhoutryve at amd.com>
Date: Fri, 2 Oct 2026 11:53:43 +0200
Subject: [PATCH 4/5] Simplify implementation; part 1

---
 .../lib/Target/AMDGPU/AMDGPUEventTracking.cpp | 247 +++++++++---------
 llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h  | 156 ++++-------
 .../Target/AMDGPU/AMDGPUEventTrackingTest.cpp | 203 +++++++-------
 3 files changed, 274 insertions(+), 332 deletions(-)

diff --git a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
index bd491bb74a326..50d77db364d0c 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
@@ -35,20 +35,19 @@ static cl::opt<bool> EventTrackerPrintAll(
 namespace AMDGPU {
 namespace eventtracking {
 
-EventTrackerRecord::EventTrackerRecord(EventTrackingContext &Ctx,
-                                       MachineInstr *MI, SingleHWEvent Kind,
-                                       uint32_t Score)
-    : EventTrackerRecord(Ctx.nextDynamicInstanceID(), MI, Kind, Score) {}
-
 void EventTrackerRecord::print(raw_ostream &OS, bool PrintMI,
                                unsigned Indent) const {
-  OS.indent(Indent) << "#" << ID << " " << Kind << " (Score=" << Score << "): ";
+  OS.indent(Indent) << "Record Kind=" << Kind << " Height=" << Height << ": ";
   if (PrintMI && MI)
     OS << *MI;
   else
     OS << "MI@" << (void *)MI << '\n';
 }
 
+hash_code EventTrackerRecord::getIdentity() const {
+  return hash_combine((void *)MI, Kind.rawValue());
+}
+
 #if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
 LLVM_DUMP_METHOD void EventTrackerRecord::dump() const {
   dbgs() << '\n';
@@ -57,14 +56,15 @@ LLVM_DUMP_METHOD void EventTrackerRecord::dump() const {
 }
 #endif
 
-EventTracker::EventTracker(MachineBasicBlock &MBB, EventTrackingContext &ETC)
-    : MBB(&MBB), Ctx(&ETC) {
-  Counters.resize(ETC.counters().size());
-  for (const CounterInfo &Info : ETC.counters())
+EventTracker::EventTracker(const MachineBasicBlock &MBB,
+                           ArrayRef<CounterInfo> CounterInfos)
+    : MBB(&MBB) {
+  Counters.resize(CounterInfos.size());
+  for (const CounterInfo &Info : CounterInfos)
     Counters[Info.CounterT].CI = &Info;
 }
 
-void EventTracker::enterBlock() {
+void EventTracker::enterBlock(GetEventTrackerFn EventTrackerGetter) {
   LLVM_DEBUG(dbgs() << "\n[EventTracker] Entering ";
              MBB->printAsOperand(dbgs()); dbgs() << '\n');
 
@@ -79,7 +79,7 @@ void EventTracker::enterBlock() {
       if (Pred == MBB)
         IsSelfPred = true;
       else
-        Preds.push_back(&(*Ctx)[Pred]);
+        Preds.push_back(&EventTrackerGetter(*Pred));
     }
   }
 
@@ -87,10 +87,10 @@ void EventTracker::enterBlock() {
     EventTracker SelfCopy = *this;
     Preds.push_back(&SelfCopy);
     clear();
-    recordIncomings(*Ctx, Preds);
+    recordIncomings(Preds);
   } else {
     clear();
-    recordIncomings(*Ctx, Preds);
+    recordIncomings(Preds);
   }
 }
 
@@ -102,7 +102,7 @@ void EventTracker::leaveBlock() {
 void EventTracker::record(MachineInstr &MI, SingleHWEvent Event) {
   LLVM_DEBUG(dbgs() << "[EventTracker] Recording " << Event << ": " << MI);
 
-  EventTrackerRecord Rec = EventTrackerRecord(*Ctx, &MI, Event);
+  EventTrackerRecord Rec = EventTrackerRecord(&MI, Event);
   [[maybe_unused]] bool FoundMatch = false;
   for (CounterData &CD : Counters) {
     if (!CD.CI->Events.contains(Event))
@@ -116,9 +116,10 @@ void EventTracker::record(MachineInstr &MI, SingleHWEvent Event) {
     if (!CD.IsOutOfOrder) {
       // NB: There is an intentional tradeoff here. We could avoid this loop by
       // instead storing a timestamp in each record, and having a
-      // constantly-increasing clock to infer the score (clock-timestamp is
-      // score). However, it'd:
-      //  - Complexify fetching the score (`EventTrackerRecord` cannot answer it
+      // constantly-increasing clock to infer the height (clock-timestamp is
+      // height). However, it'd:
+      //  - Complexify fetching the height (`EventTrackerRecord` cannot answer
+      //  it
       //    on its own anymore and we need a separate query/wrapper).
       //  - Make merge of incoming records a bit more annoying (we'd need to
       //    rebase the `clock`).
@@ -129,9 +130,13 @@ void EventTracker::record(MachineInstr &MI, SingleHWEvent Event) {
       // change the system if we have data backed up by profiling to
       // justify the change.
       for (auto &Live : CD.LiveRecords)
-        Live.setScore(Live.getScore() + 1); // Age all existing events.
+        Live.setHeight(Live.getHeight() + 1);
     }
 
+    // NOTE: We do not merge with a previous record that has the same (MI+Kind),
+    // unlike in recordIncomings. We are okay with having 2 separate records
+    // with the same identity (MI+Kind), if one is carried over from a backedge
+    // and one is from a more recent iteration.
     CD.LiveRecords.push_back(Rec);
 
 #ifndef NDEBUG
@@ -143,6 +148,10 @@ void EventTracker::record(MachineInstr &MI, SingleHWEvent Event) {
   }
 
   assert(FoundMatch && "Event has no matching InstCounterType!");
+
+#ifdef EXPENSIVE_CHECKS
+  verify();
+#endif
 }
 
 void EventTracker::wait(InstCounterType T, unsigned N) {
@@ -157,30 +166,33 @@ void EventTracker::wait(InstCounterType T, unsigned N) {
     CD.IsIndeterminate = false;
     CD.IsOutOfOrder = false;
     CD.LegacyPendingEvents = HWEvents();
-    return;
-  }
+  } else {
+    CD.Count = std::min(CD.Count, N);
 
-  CD.Count = std::min(CD.Count, N);
+    // Don't bother erasing stuff if we are out-of-order. All records have a
+    // height of zero in such cases.
+    if (!CD.IsOutOfOrder) {
+      auto *RmIt = remove_if(CD.LiveRecords, [&](EventTrackerRecord &E) {
+        if (E.getHeight() < N)
+          return false;
+        LLVM_DEBUG(dbgs() << "  | Removing "; E.print(dbgs()));
+        return true;
+      });
+      CD.LiveRecords.erase(RmIt, CD.LiveRecords.end());
+    }
+
+    LLVM_DEBUG(dbgs() << "  | => Updated Count:" << CD.Count << '\n');
 
-  // Don't bother erasing stuff if we are out-of-order. All records have a score
-  // of zero in such cases.
-  if (!CD.IsOutOfOrder) {
-    auto *RmIt = remove_if(CD.LiveRecords, [&](EventTrackerRecord &E) {
-      if (E.getScore() < N)
-        return false;
-      LLVM_DEBUG(dbgs() << "  | Removing "; E.print(dbgs()));
-      return true;
+#ifndef NDEBUG
+    LLVM_DEBUG(if (EventTrackerPrintAll) {
+      dbgs().indent(2) << "Updated Timeline:\n";
+      print(dbgs(), CD, /*Indent=*/4);
     });
-    CD.LiveRecords.erase(RmIt, CD.LiveRecords.end());
+#endif
   }
 
-  LLVM_DEBUG(dbgs() << "  | => Updated Count:" << CD.Count << '\n');
-
-#ifndef NDEBUG
-  LLVM_DEBUG(if (EventTrackerPrintAll) {
-    dbgs().indent(2) << "Updated Timeline:\n";
-    print(dbgs(), CD, /*Indent=*/4);
-  });
+#ifdef EXPENSIVE_CHECKS
+  verify();
 #endif
 }
 
@@ -198,7 +210,7 @@ void EventTracker::markOutOfOrder(InstCounterType T) {
   CounterData &CD = Counters[T];
   CD.IsOutOfOrder = true;
   for (auto &Rec : CD.LiveRecords)
-    Rec.setScore(0);
+    Rec.setHeight(0);
 }
 
 std::optional<unsigned> EventTracker::count(InstCounterType T) const {
@@ -243,7 +255,8 @@ bool EventTracker::mimicsLegacyTracking() { return MimicLegacyTracking; }
 
 #if !defined(NDEBUG) || defined(EXPENSIVE_CHECKS)
 void EventTracker::verify() const {
-  assert(MBB && Ctx && "Invalid internal state!");
+  if (!MBB)
+    llvm_unreachable("EventTracker has no MBB!");
 
   for (const CounterData &C : Counters) {
     const auto OnError = [&]() {
@@ -260,30 +273,42 @@ void EventTracker::verify() const {
     }
 
     if (C.IsOutOfOrder) {
-      if (!all_of(C.LiveRecords, [](auto &R) { return R.getScore() == 0; })) {
+      if (!all_of(C.LiveRecords, [](auto &R) { return R.getHeight() == 0; })) {
         OnError();
         llvm_unreachable(
-            "IsOutOfOrder but some records do not have a score of 0!");
+            "IsOutOfOrder but some records do not have a height of 0!");
       }
     }
 
-    if (C.Count > C.LiveRecords.size() && !mimicsLegacyTracking()) {
-      OnError();
-      llvm_unreachable(
-          "'Count' is inconsistent with the number of live records");
-    }
-
+    // Check some basic invariants
+    //  - Height of a record cannot exceed the value of the counter
+    //  - We cannot have two records with same "identity" at the same height.
+    //    e.g. We can't have 2 VMEM_READ_ACCESS at the same instruction at
+    //    the same height. This is because `recordIncomings` merges based on
+    //    identity, and unless there is a bug, `record` should increment all
+    //    pre-existing records when a new one is inserted.
+    DenseSet<std::pair<hash_code, unsigned>> RecIdentityCheck;
     for (const EventTrackerRecord &E : C.LiveRecords) {
-      if (E.getScore() > C.Count) {
+      if (E.getHeight() > C.Count) {
+        OnError();
+        dbgs() << "Concerning Record:";
+        E.print(dbgs());
+        llvm_unreachable("record height is out of range");
+      }
+
+      auto [It, Inserted] =
+          RecIdentityCheck.insert({E.getIdentity(), E.getHeight()});
+      if (!Inserted) {
         OnError();
         dbgs() << "Concerning Record:";
         E.print(dbgs());
-        llvm_unreachable("record score is out of range");
+        llvm_unreachable("The timeline cannot have two records with the same "
+                         "MachineInstr, Kind and Height at the same time!");
       }
     }
 
     // Check live records are sorted
-    if (!is_sorted(C.LiveRecords, EventTrackerRecord::isScoreGreaterThan)) {
+    if (!is_sorted(C.LiveRecords, compareRecords)) {
       OnError();
       llvm_unreachable("live records are not sorted!");
     }
@@ -319,6 +344,34 @@ LLVM_DUMP_METHOD void EventTracker::dump() const {
 }
 #endif
 
+bool EventTracker::compareRecords(const EventTrackerRecord &A,
+                                  const EventTrackerRecord &B) {
+  // Different height
+  if (A.getHeight() != B.getHeight())
+    return A.getHeight() > B.getHeight();
+
+  // Same height but different kind.
+  unsigned AKindV = A.getKind().rawValue();
+  unsigned BKindV = B.getKind().rawValue();
+  if (AKindV != BKindV)
+    return AKindV > BKindV;
+
+  const MachineInstr *AMI = A.getMI();
+  const MachineInstr *BMI = B.getMI();
+  if (AMI == BMI)
+    return false;
+
+  // Same height/kind but MIs are in different BBs:
+  // The one in the earlier BB comes first.
+  const MachineBasicBlock *Bbb = BMI->getParent();
+  const MachineBasicBlock *Abb = AMI->getParent();
+  if (Abb != Bbb)
+    return Abb->getNumber() < Bbb->getNumber();
+
+  llvm_unreachable("We cannot have two distinct instructions from the same "
+                   "MBB, at the same height and with the same event kind!");
+}
+
 void EventTracker::clear() {
   for (CounterData &C : Counters) {
     C.LiveRecords.clear();
@@ -328,8 +381,7 @@ void EventTracker::clear() {
   }
 }
 
-void EventTracker::recordIncomings(EventTrackingContext &ETC,
-                                   ArrayRef<EventTracker *> Preds) {
+void EventTracker::recordIncomings(ArrayRef<EventTracker *> Preds) {
   LLVM_DEBUG(if (!Preds.empty()) {
     dbgs() << "[EventTracker] Recording incoming events (merge) from "
               "predecessors:\n";
@@ -339,14 +391,13 @@ void EventTracker::recordIncomings(EventTrackingContext &ETC,
   });
 
   /// Iterate over all counters that are available to us.
-  for (const CounterInfo &CI : ETC.counters()) {
-    auto &CData = Counters[CI.CounterT];
+  for (CounterData &CData : Counters) {
     assert(CData.LiveRecords.empty());
 
-    DenseMap<EventTrackerRecord::DynamicInstanceID, EventTrackerRecord> Acc;
+    DenseMap<hash_code, EventTrackerRecord> Acc;
 
     for (EventTracker *Pred : Preds) {
-      auto &PredCData = Pred->Counters[CI.CounterT];
+      auto &PredCData = Pred->Counters[CData.CI->CounterT];
 
       // Merge domain for the count value:
       CData.Count = std::max(CData.Count, PredCData.Count);
@@ -358,18 +409,16 @@ void EventTracker::recordIncomings(EventTrackingContext &ETC,
       CData.IsOutOfOrder |= PredCData.IsOutOfOrder;
 
       for (EventTrackerRecord &PredEntry : PredCData.LiveRecords) {
-        EventTrackerRecord::DynamicInstanceID ID = PredEntry.getID();
-        auto It = Acc.find(ID);
+        // At a join, we collapse records from all predecessors with the same
+        // MI+Kind to a single entry with the Height of the entry being the
+        // minimum across predecessors.
+        hash_code Identity = PredEntry.getIdentity();
+        auto It = Acc.find(Identity);
         if (It != Acc.end()) {
           auto &AccVal = It->second;
-          AccVal.setScore(std::min(AccVal.getScore(), PredEntry.getScore()));
-          assert(
-              PredEntry.getMI() == AccVal.getMI() &&
-              PredEntry.getKind() == AccVal.getKind() &&
-              "EventTrackerRecord have same DynamicInstanceID, but different "
-              "MachineInstr/HWEvent kind, which should not be possible");
+          AccVal.setHeight(std::min(AccVal.getHeight(), PredEntry.getHeight()));
         } else
-          Acc.insert({ID, PredEntry});
+          Acc.insert({Identity, PredEntry});
       }
     }
 
@@ -382,10 +431,14 @@ void EventTracker::recordIncomings(EventTrackingContext &ETC,
     auto AccVals = Acc.values();
     CData.LiveRecords.append(AccVals.begin(), AccVals.end());
 
-    // Sort records by Score (descending) for consistent iteration.
-    stable_sort(CData.LiveRecords, EventTrackerRecord::isScoreGreaterThan);
+    // Sort records by Height (descending) for consistent iteration.
+    stable_sort(CData.LiveRecords, compareRecords);
   }
 
+#ifdef EXPENSIVE_CHECKS
+  verify();
+#endif
+
   LLVM_DEBUG(if (!Preds.empty()) {
     dbgs() << "[EventTracker] Timeline after recording incomings:\n";
     print(dbgs(), /*IgnoreEmpty=*/true, /*Indent=*/2);
@@ -406,64 +459,6 @@ void EventTracker::print(raw_ostream &OS, const CounterData &CD,
   }
 }
 
-EventTrackingContext::EventTrackingContext(MachineFunction &MF,
-                                           ArrayRef<CounterInfo> Counters)
-    : CounterInfos(Counters) {
-  LLVM_DEBUG(dbgs() << "\n[EventTrackingContext] CounterInfos for "
-                    << MF.getName() << '\n';
-             for (const auto &CI
-                  : CounterInfos) {
-               dbgs().indent(2)
-                   << AMDGPU::getInstCounterName(CI.CounterT) << " ";
-               if (CI.Events.none()) {
-                 dbgs() << " (unused - no HWEvents assigned)\n";
-               } else {
-                 dbgs() << "(Limit=" << CI.Limit << ") " << CI.Events << '\n';
-               }
-             });
-
-  Trackers.reserve(MF.size());
-  for (MachineBasicBlock &MBB : MF)
-    Trackers[&MBB] = std::make_unique<EventTracker>(MBB, *this);
-}
-
-EventTracker &EventTrackingContext::operator[](MachineBasicBlock *MBB) {
-  assert(MBB);
-  return *Trackers.at(MBB);
-}
-
-#if !defined(NDEBUG) || defined(EXPENSIVE_CHECKS)
-void EventTrackingContext::verify() const {
-  for (const auto &[MBB, Tracker] : Trackers) {
-    assert(MBB && "Unexpected nullptr entry!");
-    Tracker->verify();
-  }
-
-  // Check CounterInfos is sane.
-  for (auto [Idx, CI] : enumerate(CounterInfos)) {
-    assert(Idx == CI.CounterT && "CounterInfo is in wrong position!");
-    assert(CI.Events.any() &&
-           "InstCounterType has no event associated with it!");
-  }
-}
-#endif
-
-void EventTrackingContext::print(raw_ostream &OS) const {
-  for (const auto &[MBB, Tracker] : Trackers) {
-    MBB->printAsOperand(OS);
-    OS << ":\n";
-    Tracker->print(OS, /*Indent=*/2);
-  }
-}
-
-#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
-LLVM_DUMP_METHOD void EventTrackingContext::dump() const {
-  dbgs() << '\n';
-  print(dbgs());
-  dbgs() << '\n';
-}
-#endif
-
 } // namespace eventtracking
 } // namespace AMDGPU
 
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
index 627b29bae06e3..df7a6e76d6b93 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
+++ b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
@@ -36,6 +36,8 @@ class EventTracker;
 
 /// FIXME: Make this a generic util?
 struct CounterInfo {
+  CounterInfo() = default;
+
   constexpr CounterInfo(InstCounterType T, HWEvents Events, unsigned Limit)
       : CounterT(T), Events(Events), Limit(Limit) {}
 
@@ -49,29 +51,14 @@ struct CounterInfo {
   unsigned Limit;
 };
 
-/// Records are all uniquely identified by a \ref DynamicInstanceID. For each
-/// unique \ref DynamicInstanceID value, all \ref EventTrackerRecord that use
-/// that ID should have the same MI and Kind. This is enforced by exposing these
-/// as read-only, and making the constructor assign a new \ref DynamicInstanceID
-/// every time. Only the score can change as it may be unique to each instance
-/// of \ref EventTracker that carry it.
+/// Represent a record within the \ref EventTracker's timeline for one
+/// \ref InstCounterType.
 class EventTrackerRecord {
 public:
-  /// An always-increasing counter used to represent a dynamic instance of a
-  /// record. Whenever we add a new \ref EventTrackerRecord, even if it's one we
-  /// already have seen in a previous dataflow iteration, this counter is
-  /// increased so that the new record has a unique `DynamicInstanceID`.
-  using DynamicInstanceID = uint32_t;
-
-  EventTrackerRecord(EventTrackingContext &Ctx, MachineInstr *MI,
-                     SingleHWEvent Kind, uint32_t Score = 0);
-
-  /// \returns the ID uniquely identifying this record across an entire
-  /// EventTrackingContext. Whenever we revisit an instruction (when iterating
-  /// until a fixpoint is reached), we give it a new ID. This is used to
-  /// represent records carried over from previous iterations of the same basic
-  /// block.
-  DynamicInstanceID getID() const { return ID; }
+  EventTrackerRecord(MachineInstr *MI, SingleHWEvent Kind, uint32_t Height = 0)
+      : MI(MI), Kind(Kind) {
+    setHeight(Height);
+  }
 
   /// \returns the MachineInstr that originated this record.
   MachineInstr *getMI() const { return MI; }
@@ -79,13 +66,13 @@ class EventTrackerRecord {
   /// \returns the kind of record this is, as a \ref SingleHWEvent.
   SingleHWEvent getKind() const { return Kind; }
 
-  /// \returns the score of this record.
-  uint32_t getScore() const { return Score; }
+  /// \returns the height of this record.
+  uint32_t getHeight() const { return Height; }
 
-  /// Sets the score of this record to \p NewScore.
-  void setScore(uint32_t NewScore) {
-    Score = NewScore;
-    assert(Score == NewScore && "Score overflow!");
+  /// Sets the height of this record to \p NewHeight.
+  void setHeight(uint32_t NewHeight) {
+    Height = NewHeight;
+    assert(Height == NewHeight && "Height overflow!");
   }
 
   void print(raw_ostream &OS, bool PrintMI = true, unsigned Indent = 0) const;
@@ -94,25 +81,18 @@ class EventTrackerRecord {
   LLVM_DUMP_METHOD void dump() const;
 #endif
 
-  static bool isScoreGreaterThan(const EventTrackerRecord &A,
-                                 const EventTrackerRecord &B) {
-    return A.getScore() > B.getScore();
-  }
+  /// \returns the "identity" of a this record which is a hash of its kind and
+  /// MI.
+  hash_code getIdentity() const;
 
 private:
-  EventTrackerRecord(DynamicInstanceID ID, MachineInstr *MI, SingleHWEvent Kind,
-                     uint32_t Score = 0)
-      : MI(MI), ID(ID), Kind(Kind) {
-    setScore(Score);
-  }
-
   MachineInstr *MI;
-  DynamicInstanceID ID;
   SingleHWEvent Kind;
-  // Score should already never exceed uint8_t limit in normal circumstances as
+  // Height should never exceed uint8_t limit in normal circumstances as
   // most counter types only use up to 6 bits encoding for the waitcnts. 16 bit
-  // is a very generous limit, we can probably shrink that at some point.
-  uint16_t Score;
+  // is a very generous limit, we can probably shrink that at some point once we
+  // refine how we handle overflowing inst counters.
+  uint16_t Height;
 };
 
 /// This assert serves as a reminder to be mindful of the size of the object.
@@ -129,24 +109,34 @@ static_assert(sizeof(EventTrackerRecord) == 16,
 /// are for utils/wrappers/users of the class.
 class EventTracker {
 public:
-  EventTracker(MachineBasicBlock &MBB, EventTrackingContext &ET);
-
-  /// Notify this EventTracker that we are going to begin recording events.
-  /// In case this is not the first time we are going through this block, this
-  /// clears the internal state of the tracker and re-imports all incoming
-  /// tracking state from the predecessors.
-  void enterBlock();
-
-  /// Notify this EventTracker that we are done recording events.
+  using GetEventTrackerFn =
+      function_ref<EventTracker &(const MachineBasicBlock &)>;
+
+  EventTracker(const MachineBasicBlock &MBB,
+               ArrayRef<CounterInfo> CounterInfos);
+
+  /// Prepare this EventTracker for iteration through its basic block.
+  ///
+  /// When entering a block, we merge state from the EventTrackers of
+  /// incoming MBBs. Records from incoming MBBs are merged using the `identity`
+  /// of the \ref EventTrackerRecord and only the record with the lowest height
+  /// is kept.
+  ///
+  /// \param EventTrackerGetter Is a function that map a MachimeBasicBlock to
+  /// the EventTracker used for that MachineBasicBlock.
+  void enterBlock(GetEventTrackerFn EventTrackerGetter);
+
+  /// End iteration through the basic block.
+  /// FIXME: Currently does nothing, it's just for symmetry and debug logs.
   void leaveBlock();
 
   /// Record an event of type \p Event at a MachineInstr \p MI, which will
   /// affect all counters that have \p Event in their event set.
   void record(MachineInstr &MI, SingleHWEvent Event);
 
-  /// Notify that we waited until the counter \p T reached the value \p N before
+  /// Wait until the counter \p T reaches the value \p N before
   /// continuing execution of the program (and recording more events).
-  /// This affects the count of \p T, an removes all records that have a score
+  /// This affects the count of \p T, and removes all records that have a height
   /// greater than or equal to \p N.
   /// If \p N is zero, then \p T will no longer be in an indeterminate or
   /// out-of-order state afterwards if it previously was in such a state.
@@ -160,7 +150,7 @@ class EventTracker {
   void markIndeterminate(InstCounterType T);
 
   /// Mark the counter \p T as being "out-of-order", meaning records may retire
-  /// in any order. This sets the score of all records to zero.
+  /// in any order. This sets the height of all records to zero.
   void markOutOfOrder(InstCounterType T);
 
   /// \returns the current value of the counter \p T at this point in time, or
@@ -215,6 +205,9 @@ class EventTracker {
     /// The current value of the counter. This is a max (upper bound) across
     /// all possible execution paths at runtime. It cannot be inferred from the
     /// LiveRecords alone and is thus a separate tracking domain.
+    ///
+    /// FIXME: We shouldn't need it for correctness so perhaps it should be
+    /// removed entirely.
     uint32_t Count = 0;
     /// An upper bound that persists across fixpoint iterations. This is only
     /// used when \ref mimicsLegacyTracking returns true.
@@ -224,7 +217,7 @@ class EventTracker {
     /// the counter is out-of-order as well.
     bool IsIndeterminate = false;
     /// Whether this counter is out-of-order, meaning records may retire in any
-    /// order and they all exist at a score of zero.
+    /// order and they all exist at a height of zero.
     bool IsOutOfOrder = false;
     /// Legacy-style tracking of pending events that is coarse and does not
     /// leverage the live set of records. Only used when
@@ -232,26 +225,28 @@ class EventTracker {
     /// iterations.
     HWEvents LegacyPendingEvents;
 
-    // TODO: We could imagine storing the per-predecessor score for incoming
+    // TODO: We could imagine storing the per-predecessor height for incoming
     // events. We could achieve that by storing that as a map of ((ID, Pred),
     // Score). This would allow identifying events that are "deep" in one branch
     // but "shallow" in another, e.g. an event needing a waitcnt 1 for one pred,
     // but a waitcnt 8 for another. Not sure if we can exploit that though?
   };
 
+  /// Compare records in order to achieve stable sorting by descending height.
+  static bool compareRecords(const EventTrackerRecord &A,
+                             const EventTrackerRecord &B);
+
   /// Clears the tracked data, used when entering a block.
   void clear();
 
   /// Import all events from the incoming basic blocks in \p Preds and reconcile
   /// divergence at joints.
-  void recordIncomings(EventTrackingContext &ETC,
-                       ArrayRef<EventTracker *> Preds);
+  void recordIncomings(ArrayRef<EventTracker *> Preds);
 
   static void print(raw_ostream &OS, const CounterData &CD,
                     unsigned Indent = 0);
 
-  MachineBasicBlock *MBB;
-  EventTrackingContext *Ctx;
+  const MachineBasicBlock *MBB;
 
   // NB: This, combined with the inline storage of LiveRecords, can lead to this
   // class becoming quite big - verify the size of this object whenever a change
@@ -259,49 +254,6 @@ class EventTracker {
   SmallVector<CounterData, InstCounterType::NUM_INST_CNTS> Counters;
 };
 
-/// Per-MF Tracking Context.
-/// This owns all \ref EventTrackers and keeps track of state that persists
-/// across dataflow analysis iterations, such as the current value of
-/// \ref DynamicInstanceID.
-class EventTrackingContext {
-public:
-  /// \param MF Machine Function
-  /// \param Counters The counters available to \p MF on this target.
-  EventTrackingContext(MachineFunction &MF, ArrayRef<CounterInfo> Counters);
-
-  /// Fetch the \ref MBBEventTracker of \p MBB.
-  EventTracker &operator[](MachineBasicBlock *MBB);
-
-  /// \returns the list of counters available to the current target.
-  ArrayRef<CounterInfo> counters() const { return CounterInfos; }
-
-#if !defined(NDEBUG) || defined(EXPENSIVE_CHECKS)
-  /// Verifies invariants of this class are respected.
-  void verify() const;
-#endif
-
-  void print(raw_ostream &OS) const;
-
-#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
-  LLVM_DUMP_METHOD void dump() const;
-#endif
-
-private:
-  friend class EventTrackerRecord;
-
-  /// \returns a new, unique \ref DynamicInstanceID - only for use by
-  /// \ref EventTrackerRecord.
-  EventTrackerRecord::DynamicInstanceID nextDynamicInstanceID() {
-    assert(NextDynID + 1 > NextDynID && "DynamicInstanceIDs overflow!");
-    return ++NextDynID;
-  }
-
-  SmallVector<CounterInfo> CounterInfos;
-
-  EventTrackerRecord::DynamicInstanceID NextDynID = 0;
-  DenseMap<MachineBasicBlock *, std::unique_ptr<EventTracker>> Trackers;
-};
-
 } // namespace eventtracking
 } // namespace AMDGPU
 
diff --git a/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp b/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
index 565a54ffdd99e..06d95f089797c 100644
--- a/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
+++ b/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
@@ -35,29 +35,42 @@ static constexpr std::array<CounterInfo, 4> GFX12CounterInfos = {{
 
 class AMDGPUGFX12EventTrackingTest : public AMDGPUCodeGenTestBase {
 public:
-  void SetUp() override { setUpImpl("amdgpu12.00-amd-amdhsa", "", ""); }
-};
+  void SetUp() override {
+    setUpImpl("amdgpu12.00-amd-amdhsa", "", "");
+
+    TrackerGetter = [&](const MachineBasicBlock &MBB) -> EventTracker & {
+      auto &Entry = Trackers[&MBB];
+      if (!Entry)
+        Entry = std::make_unique<EventTracker>(MBB, GFX12CounterInfos);
+      return *Entry;
+    };
+  }
 
-namespace {
-static EventTracker &
-visitAll(EventTrackingContext &Ctx, MachineBasicBlock &MBB,
-         function_ref<void(EventTracker &ET)> AfterVisit = nullptr) {
-  const GCNSubtarget &ST = MBB.getParent()->getSubtarget<GCNSubtarget>();
-
-  EventTracker &ET = Ctx[&MBB];
-  ET.enterBlock();
-  for (MachineInstr &MI : MBB) {
-    HWEvents Events =
-        getEventsFor(MI, ST, /*IsExpertMode=*/false, /*TgSplit=*/false);
-    for (HWEvents SingleEv : Events) {
-      ET.record(MI, SingleHWEvent::encode(SingleEv));
+  EventTracker &
+  visitAll(MachineBasicBlock &MBB,
+           function_ref<void(EventTracker &ET)> AfterVisit = nullptr) {
+    const GCNSubtarget &ST = MBB.getParent()->getSubtarget<GCNSubtarget>();
+
+    EventTracker &ET = TrackerGetter(MBB);
+    ET.enterBlock(TrackerGetter);
+    for (MachineInstr &MI : MBB) {
+      HWEvents Events =
+          getEventsFor(MI, ST, /*IsExpertMode=*/false, /*TgSplit=*/false);
+      for (HWEvents SingleEv : Events) {
+        ET.record(MI, SingleHWEvent::encode(SingleEv));
+      }
     }
+    if (AfterVisit)
+      AfterVisit(ET);
+    ET.leaveBlock();
+    return ET;
   }
-  if (AfterVisit)
-    AfterVisit(ET);
-  ET.leaveBlock();
-  return ET;
-}
+
+  std::function<EventTracker &(const MachineBasicBlock &)> TrackerGetter;
+  DenseMap<const MachineBasicBlock *, std::unique_ptr<EventTracker>> Trackers;
+};
+
+namespace {
 
 /// Provides some helpers to declaratively check the state of the EventTracker
 /// for one counter. This provides helpers to check the general counter state
@@ -75,7 +88,7 @@ struct TrackerRecordsChecker {
 
   bool empty() { return Records.empty(); }
 
-  // Has no records and score is zero (if there is one)
+  // Has no records and count is zero (if there is one)
   bool unused() { return empty() && (!hasCount() || !getCount()); }
 
   const EventTrackerRecord &cur() { return Records[CurElt]; }
@@ -115,16 +128,14 @@ body:             |
   MachineFunction &MF = getMF("BasicTimeline");
   MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
 
-  EventTrackingContext Ctx(MF, GFX12CounterInfos);
-
-  auto &ET = visitAll(Ctx, BB0);
+  auto &ET = visitAll(BB0);
 
   auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
   EXPECT_TRUE(LoadCnt.hasCount());
   EXPECT_EQ(LoadCnt.getCount(), 1u);
   EXPECT_FALSE(LoadCnt.empty());
   EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
-  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_EQ(LoadCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(LoadCnt.next());
 
   auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
@@ -132,7 +143,7 @@ body:             |
   EXPECT_EQ(DsCnt.getCount(), 1u);
   EXPECT_FALSE(DsCnt.empty());
   EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
-  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_EQ(DsCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(DsCnt.next());
 
   auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
@@ -143,10 +154,10 @@ body:             |
   EXPECT_EQ(StoreCnt.getCount(), 2u);
   EXPECT_FALSE(StoreCnt.empty());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 1u);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 1u);
   EXPECT_TRUE(StoreCnt.next());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(StoreCnt.next());
 }
 
@@ -171,21 +182,19 @@ body:             |
   MachineFunction &MF = getMF("BasicTimeline");
   MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
 
-  EventTrackingContext Ctx(MF, GFX12CounterInfos);
-
   // Iterate twice
-  visitAll(Ctx, BB0);
-  auto &ET = visitAll(Ctx, BB0);
+  visitAll(BB0);
+  auto &ET = visitAll(BB0);
 
   auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
   EXPECT_TRUE(LoadCnt.hasCount());
   EXPECT_EQ(LoadCnt.getCount(), 2u);
   EXPECT_FALSE(LoadCnt.empty());
   EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
-  EXPECT_EQ(LoadCnt.cur().getScore(), 1u);
+  EXPECT_EQ(LoadCnt.cur().getHeight(), 1u);
   EXPECT_TRUE(LoadCnt.next());
   EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
-  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_EQ(LoadCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(LoadCnt.next());
 
   auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
@@ -193,10 +202,10 @@ body:             |
   EXPECT_EQ(DsCnt.getCount(), 2u);
   EXPECT_FALSE(DsCnt.empty());
   EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
-  EXPECT_EQ(DsCnt.cur().getScore(), 1u);
+  EXPECT_EQ(DsCnt.cur().getHeight(), 1u);
   EXPECT_TRUE(DsCnt.next());
   EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
-  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_EQ(DsCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(DsCnt.next());
 
   auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
@@ -207,10 +216,10 @@ body:             |
   EXPECT_EQ(StoreCnt.getCount(), 2u);
   EXPECT_FALSE(StoreCnt.empty());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 1u);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 1u);
   EXPECT_TRUE(StoreCnt.next());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(StoreCnt.next());
 }
 
@@ -237,17 +246,15 @@ body:             |
   MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
   MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
 
-  EventTrackingContext Ctx(MF, GFX12CounterInfos);
-
-  visitAll(Ctx, BB0);
-  auto &ET = visitAll(Ctx, BB1);
+  visitAll(BB0);
+  auto &ET = visitAll(BB1);
 
   auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
   EXPECT_TRUE(LoadCnt.hasCount());
   EXPECT_EQ(LoadCnt.getCount(), 1u);
   EXPECT_FALSE(LoadCnt.empty());
   EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
-  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_EQ(LoadCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(LoadCnt.next());
 
   auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
@@ -255,7 +262,7 @@ body:             |
   EXPECT_EQ(DsCnt.getCount(), 1u);
   EXPECT_FALSE(DsCnt.empty());
   EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
-  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_EQ(DsCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(DsCnt.next());
 
   auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
@@ -266,10 +273,10 @@ body:             |
   EXPECT_EQ(StoreCnt.getCount(), 2u);
   EXPECT_FALSE(StoreCnt.empty());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 1u);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 1u);
   EXPECT_TRUE(StoreCnt.next());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(StoreCnt.next());
 }
 
@@ -301,18 +308,16 @@ body:             |
   MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
   MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
 
-  EventTrackingContext Ctx(MF, GFX12CounterInfos);
-
-  visitAll(Ctx, BB0);
-  visitAll(Ctx, BB1);
-  auto &ET = visitAll(Ctx, BB2);
+  visitAll(BB0);
+  visitAll(BB1);
+  auto &ET = visitAll(BB2);
 
   auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
   EXPECT_TRUE(LoadCnt.hasCount());
   EXPECT_EQ(LoadCnt.getCount(), 1u);
   EXPECT_FALSE(LoadCnt.empty());
   EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
-  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_EQ(LoadCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(LoadCnt.next());
 
   auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
@@ -320,7 +325,7 @@ body:             |
   EXPECT_EQ(DsCnt.getCount(), 1u);
   EXPECT_FALSE(DsCnt.empty());
   EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
-  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_EQ(DsCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(DsCnt.next());
 
   auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
@@ -331,10 +336,10 @@ body:             |
   EXPECT_EQ(StoreCnt.getCount(), 2u);
   EXPECT_FALSE(StoreCnt.empty());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 1u);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 1u);
   EXPECT_TRUE(StoreCnt.next());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(StoreCnt.next());
 }
 
@@ -366,18 +371,16 @@ body:             |
   MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
   MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
 
-  EventTrackingContext Ctx(MF, GFX12CounterInfos);
-
-  visitAll(Ctx, BB0);
-  visitAll(Ctx, BB1);
-  auto &ET = visitAll(Ctx, BB2);
+  visitAll(BB0);
+  visitAll(BB1);
+  auto &ET = visitAll(BB2);
 
   auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
   EXPECT_TRUE(LoadCnt.hasCount());
   EXPECT_EQ(LoadCnt.getCount(), 1u);
   EXPECT_FALSE(LoadCnt.empty());
   EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
-  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_EQ(LoadCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(LoadCnt.next());
 
   auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
@@ -385,7 +388,7 @@ body:             |
   EXPECT_EQ(DsCnt.getCount(), 1u);
   EXPECT_FALSE(DsCnt.empty());
   EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
-  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_EQ(DsCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(DsCnt.next());
 
   auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
@@ -398,11 +401,11 @@ body:             |
   EXPECT_FALSE(StoreCnt.empty());
   // First successor has GLOBAL_STORE_DWORD at height 0
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 0u);
   EXPECT_TRUE(StoreCnt.next());
   // Second successor has GLOBAL_STORE_DWORDX2 at height 0 too
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(StoreCnt.next());
 }
 
@@ -440,19 +443,17 @@ body:             |
   MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
   MachineBasicBlock &BB3 = *MF.getBlockNumbered(3);
 
-  EventTrackingContext Ctx(MF, GFX12CounterInfos);
-
-  visitAll(Ctx, BB0);
-  visitAll(Ctx, BB1);
-  visitAll(Ctx, BB2);
-  auto &ET = visitAll(Ctx, BB3);
+  visitAll(BB0);
+  visitAll(BB1);
+  visitAll(BB2);
+  auto &ET = visitAll(BB3);
 
   auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
   EXPECT_TRUE(LoadCnt.hasCount());
   EXPECT_EQ(LoadCnt.getCount(), 1u);
   EXPECT_FALSE(LoadCnt.empty());
   EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
-  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_EQ(LoadCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(LoadCnt.next());
 
   auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
@@ -460,7 +461,7 @@ body:             |
   EXPECT_EQ(DsCnt.getCount(), 1u);
   EXPECT_FALSE(DsCnt.empty());
   EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
-  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_EQ(DsCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(DsCnt.next());
 
   auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
@@ -471,7 +472,7 @@ body:             |
   EXPECT_EQ(StoreCnt.getCount(), 1u);
   EXPECT_FALSE(StoreCnt.empty());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(StoreCnt.next());
 }
 
@@ -506,14 +507,12 @@ body:             |
   MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
   MachineBasicBlock &BB3 = *MF.getBlockNumbered(3);
 
-  EventTrackingContext Ctx(MF, GFX12CounterInfos);
-
-  visitAll(Ctx, BB0);
-  visitAll(Ctx, BB1, /*AfterVisit=*/[&](EventTracker &ET) {
+  visitAll(BB0);
+  visitAll(BB1, /*AfterVisit=*/[&](EventTracker &ET) {
     ET.markIndeterminate(STORE_CNT);
   });
-  visitAll(Ctx, BB2, [&](EventTracker &ET) { ET.markOutOfOrder(LOAD_CNT); });
-  auto &ET = visitAll(Ctx, BB3);
+  visitAll(BB2, [&](EventTracker &ET) { ET.markOutOfOrder(LOAD_CNT); });
+  auto &ET = visitAll(BB3);
 
   auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
   EXPECT_TRUE(LoadCnt.unused());
@@ -540,8 +539,8 @@ body:             |
 ///
 /// The timeline will be:
 ///   - Stores from bb0 exist at 1/2
-///   - The added store from bb1 exist at score 0
-///   - The added store from bb2 exist at score 0
+///   - The added store from bb1 exist at height 0
+///   - The added store from bb2 exist at height 0
 TEST_F(AMDGPUGFX12EventTrackingTest, AssymetricalDiamond) {
   StringRef MIR = R"(
 name:            AssymetricalDiamond
@@ -575,13 +574,11 @@ body:             |
   MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
   MachineBasicBlock &BB3 = *MF.getBlockNumbered(3);
 
-  EventTrackingContext Ctx(MF, GFX12CounterInfos);
-
-  visitAll(Ctx, BB0);
-  visitAll(Ctx, BB1);
-  visitAll(Ctx, BB2,
+  visitAll(BB0);
+  visitAll(BB1);
+  visitAll(BB2,
            /*AfterVisit=*/[&](EventTracker &ET) { ET.wait(STORE_CNT, 1); });
-  auto &ET = visitAll(Ctx, BB3);
+  auto &ET = visitAll(BB3);
 
   auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
   EXPECT_TRUE(LoadCnt.unused());
@@ -600,18 +597,18 @@ body:             |
 
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
   EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB0);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 2u);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 2u);
   EXPECT_TRUE(StoreCnt.next());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 1u);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 1u);
   EXPECT_TRUE(StoreCnt.next());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
-  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB2);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB1);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 0u);
   EXPECT_TRUE(StoreCnt.next());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
-  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB1);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB2);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(StoreCnt.next());
 }
 
@@ -620,7 +617,7 @@ body:             |
 ///   - bb1 just falls through
 ///   - bb2 adds 1 more store.
 ///
-/// The timeline will showcase how the score of a counter can be different from
+/// The timeline will showcase how the value of a counter can be different from
 /// the content of the timeline.
 ///   - The timeline will have all records at a height of 0.
 ///   - The counter will still be at 2.
@@ -655,13 +652,11 @@ body:             |
   MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
   MachineBasicBlock &BB3 = *MF.getBlockNumbered(3);
 
-  EventTrackingContext Ctx(MF, GFX12CounterInfos);
-
-  visitAll(Ctx, BB0);
-  visitAll(Ctx, BB1);
-  visitAll(Ctx, BB2);
+  visitAll(BB0);
+  visitAll(BB1);
+  visitAll(BB2);
 
-  auto &ET = visitAll(Ctx, BB3);
+  auto &ET = visitAll(BB3);
 
   auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
   EXPECT_TRUE(LoadCnt.unused());
@@ -679,12 +674,12 @@ body:             |
   EXPECT_FALSE(StoreCnt.empty());
 
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
-  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB2);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB0);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 0u);
   EXPECT_TRUE(StoreCnt.next());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
-  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB0);
-  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB2);
+  EXPECT_EQ(StoreCnt.cur().getHeight(), 0u);
   EXPECT_FALSE(StoreCnt.next());
 }
 

>From ddbd374929cfaa0b34abeb016c94669463761a9e Mon Sep 17 00:00:00 2001
From: pvanhout <pierre.vanhoutryve at amd.com>
Date: Fri, 2 Oct 2026 11:59:13 +0200
Subject: [PATCH 5/5] Simplifications, part 2

---
 .../lib/Target/AMDGPU/AMDGPUEventTracking.cpp | 140 ++++--------------
 llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h  |  51 +------
 .../Target/AMDGPU/AMDGPUEventTrackingTest.cpp |  84 +----------
 3 files changed, 33 insertions(+), 242 deletions(-)

diff --git a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
index 50d77db364d0c..bac203d74c91a 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
@@ -15,16 +15,11 @@
 #include "llvm/Support/CommandLine.h"
 #include "llvm/Support/Debug.h"
 #include <algorithm>
-#include <optional>
 
 #define DEBUG_TYPE "amdgpu-event-tracking"
 
 namespace llvm {
 
-/// Mimic legacy (coarse) tracking of counter state.
-static cl::opt<bool> MimicLegacyTracking("amdgpu-event-legacy-tracking",
-                                         cl::init(false));
-
 #ifndef NDEBUG
 static cl::opt<bool> EventTrackerPrintAll(
     "amdgpu-event-tracker-print-all", cl::init(false),
@@ -110,28 +105,24 @@ void EventTracker::record(MachineInstr &MI, SingleHWEvent Event) {
 
     FoundMatch = true;
     ++CD.Count;
-    CD.LegacyPendingEvents |= Event;
-
-    // Do not age records if we are out-of-order.
-    if (!CD.IsOutOfOrder) {
-      // NB: There is an intentional tradeoff here. We could avoid this loop by
-      // instead storing a timestamp in each record, and having a
-      // constantly-increasing clock to infer the height (clock-timestamp is
-      // height). However, it'd:
-      //  - Complexify fetching the height (`EventTrackerRecord` cannot answer
-      //  it
-      //    on its own anymore and we need a separate query/wrapper).
-      //  - Make merge of incoming records a bit more annoying (we'd need to
-      //    rebase the `clock`).
-      //  - Potentially demand (much) more space in EventTrackerRecord to store
-      //    bigger numbers.
-      //
-      // All in all, I think this small loop is fine for now, but we can still
-      // change the system if we have data backed up by profiling to
-      // justify the change.
-      for (auto &Live : CD.LiveRecords)
-        Live.setHeight(Live.getHeight() + 1);
-    }
+
+    // NB: There is an intentional tradeoff here. We could avoid this loop by
+    // instead storing a timestamp in each record, and having a
+    // constantly-increasing clock to infer the height (clock-timestamp is
+    // height). However, it'd:
+    //  - Complexify fetching the height (`EventTrackerRecord` cannot answer
+    //  it
+    //    on its own anymore and we need a separate query/wrapper).
+    //  - Make merge of incoming records a bit more annoying (we'd need to
+    //    rebase the `clock`).
+    //  - Potentially demand (much) more space in EventTrackerRecord to store
+    //    bigger numbers.
+    //
+    // All in all, I think this small loop is fine for now, but we can still
+    // change the system if we have data backed up by profiling to
+    // justify the change.
+    for (auto &Live : CD.LiveRecords)
+      Live.setHeight(Live.getHeight() + 1);
 
     // NOTE: We do not merge with a previous record that has the same (MI+Kind),
     // unlike in recordIncomings. We are okay with having 2 separate records
@@ -163,23 +154,16 @@ void EventTracker::wait(InstCounterType T, unsigned N) {
   if (N == 0) {
     CD.LiveRecords.clear();
     CD.Count = 0;
-    CD.IsIndeterminate = false;
-    CD.IsOutOfOrder = false;
-    CD.LegacyPendingEvents = HWEvents();
   } else {
     CD.Count = std::min(CD.Count, N);
 
-    // Don't bother erasing stuff if we are out-of-order. All records have a
-    // height of zero in such cases.
-    if (!CD.IsOutOfOrder) {
-      auto *RmIt = remove_if(CD.LiveRecords, [&](EventTrackerRecord &E) {
-        if (E.getHeight() < N)
-          return false;
-        LLVM_DEBUG(dbgs() << "  | Removing "; E.print(dbgs()));
-        return true;
-      });
-      CD.LiveRecords.erase(RmIt, CD.LiveRecords.end());
-    }
+    auto *RmIt = remove_if(CD.LiveRecords, [&](EventTrackerRecord &E) {
+      if (E.getHeight() < N)
+        return false;
+      LLVM_DEBUG(dbgs() << "  | Removing "; E.print(dbgs()));
+      return true;
+    });
+    CD.LiveRecords.erase(RmIt, CD.LiveRecords.end());
 
     LLVM_DEBUG(dbgs() << "  | => Updated Count:" << CD.Count << '\n');
 
@@ -196,46 +180,11 @@ void EventTracker::wait(InstCounterType T, unsigned N) {
 #endif
 }
 
-void EventTracker::markIndeterminate(InstCounterType T) {
-  LLVM_DEBUG(dbgs() << "[EventTracker] Marking " << getInstCounterName(T)
-                    << " as indeterminate!\n");
-  CounterData &CD = Counters[T];
-  CD.IsIndeterminate = true;
-  markOutOfOrder(T);
-}
-
-void EventTracker::markOutOfOrder(InstCounterType T) {
-  LLVM_DEBUG(dbgs() << "[EventTracker] Marking " << getInstCounterName(T)
-                    << " as out-of-order!\n");
-  CounterData &CD = Counters[T];
-  CD.IsOutOfOrder = true;
-  for (auto &Rec : CD.LiveRecords)
-    Rec.setHeight(0);
-}
-
-std::optional<unsigned> EventTracker::count(InstCounterType T) const {
-  const CounterData &CD = Counters[T];
-  if (CD.IsIndeterminate)
-    return std::nullopt;
+unsigned EventTracker::count(InstCounterType T) const {
   return Counters[T].Count;
 }
 
-bool EventTracker::isIndeterminate(InstCounterType T) const {
-  return Counters[T].IsIndeterminate;
-}
-
-bool EventTracker::isOutOfOrder(InstCounterType T) const {
-  return Counters[T].IsOutOfOrder;
-}
-
 HWEvents EventTracker::getPendingEvents(InstCounterType T) const {
-  const CounterData &CD = Counters[T];
-  if (CD.IsIndeterminate)
-    return CD.CI->Events; // return all events
-
-  if (MimicLegacyTracking)
-    return CD.LegacyPendingEvents;
-
   HWEvents Res;
   for (const auto &E : Counters[T].LiveRecords)
     Res |= E.getKind();
@@ -251,8 +200,6 @@ void EventTracker::print(raw_ostream &OS, InstCounterType T) const {
   print(OS, Counters[T]);
 }
 
-bool EventTracker::mimicsLegacyTracking() { return MimicLegacyTracking; }
-
 #if !defined(NDEBUG) || defined(EXPENSIVE_CHECKS)
 void EventTracker::verify() const {
   if (!MBB)
@@ -264,22 +211,6 @@ void EventTracker::verify() const {
       print(dbgs(), C);
     };
 
-    if (C.IsIndeterminate) {
-      if (!C.IsOutOfOrder) {
-        OnError();
-        llvm_unreachable("IsIndeterminate but not IsOutOfOrder");
-      }
-      continue;
-    }
-
-    if (C.IsOutOfOrder) {
-      if (!all_of(C.LiveRecords, [](auto &R) { return R.getHeight() == 0; })) {
-        OnError();
-        llvm_unreachable(
-            "IsOutOfOrder but some records do not have a height of 0!");
-      }
-    }
-
     // Check some basic invariants
     //  - Height of a record cannot exceed the value of the counter
     //  - We cannot have two records with same "identity" at the same height.
@@ -376,8 +307,6 @@ void EventTracker::clear() {
   for (CounterData &C : Counters) {
     C.LiveRecords.clear();
     C.Count = 0;
-    C.IsIndeterminate = false;
-    C.IsOutOfOrder = false;
   }
 }
 
@@ -401,12 +330,6 @@ void EventTracker::recordIncomings(ArrayRef<EventTracker *> Preds) {
 
       // Merge domain for the count value:
       CData.Count = std::max(CData.Count, PredCData.Count);
-      // Merge domain for the legacy pending events.
-      CData.LegacyPendingEvents |= PredCData.LegacyPendingEvents;
-      // Merge domain for the indeterminate state.
-      CData.IsIndeterminate |= PredCData.IsIndeterminate;
-      // Merge domain for the out-of-order state.
-      CData.IsOutOfOrder |= PredCData.IsOutOfOrder;
 
       for (EventTrackerRecord &PredEntry : PredCData.LiveRecords) {
         // At a join, we collapse records from all predecessors with the same
@@ -422,12 +345,6 @@ void EventTracker::recordIncomings(ArrayRef<EventTracker *> Preds) {
       }
     }
 
-    if (MimicLegacyTracking) {
-      CData.PersistentUpperBound =
-          std::max(CData.PersistentUpperBound, CData.Count);
-      CData.Count = CData.PersistentUpperBound;
-    }
-
     auto AccVals = Acc.values();
     CData.LiveRecords.append(AccVals.begin(), AccVals.end());
 
@@ -449,10 +366,7 @@ void EventTracker::print(raw_ostream &OS, const CounterData &CD,
                          unsigned Indent) {
   OS.indent(Indent) << getInstCounterName(CD.CI->CounterT)
                     << " (Count=" << CD.Count
-                    << ", PersistentUpperBound=" << CD.PersistentUpperBound
-                    << ", LiveRecords=" << CD.LiveRecords.size()
-                    << ", IsOutOfOrder=" << CD.IsOutOfOrder
-                    << ", IsIndeterminate=" << CD.IsIndeterminate << ")\n";
+                    << ", LiveRecords=" << CD.LiveRecords.size() << ")\n";
   for (const EventTrackerRecord &E : CD.LiveRecords) {
     OS.indent(Indent + 2);
     E.print(OS);
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
index df7a6e76d6b93..f22c61c839fa0 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
+++ b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
@@ -17,10 +17,8 @@
 #include "AMDGPUHWEvents.h"
 #include "AMDGPUWaitcntUtils.h"
 #include "llvm/ADT/ArrayRef.h"
-#include "llvm/ADT/DenseMap.h"
 #include "llvm/ADT/SmallVector.h"
 #include "llvm/ADT/Twine.h"
-#include <memory>
 
 namespace llvm {
 class raw_ostream;
@@ -138,39 +136,16 @@ class EventTracker {
   /// continuing execution of the program (and recording more events).
   /// This affects the count of \p T, and removes all records that have a height
   /// greater than or equal to \p N.
-  /// If \p N is zero, then \p T will no longer be in an indeterminate or
-  /// out-of-order state afterwards if it previously was in such a state.
   void wait(InstCounterType T, unsigned N = 0);
 
-  /// Mark the counter \p T as being in an indeterminate state. This means that
-  /// we no longer accurately track \p T because there may be more records we do
-  /// not know about. This primarily affects \ref getPendingEvents and
-  /// \ref count.
-  /// Implies \ref markOutOfOrder for \p T as well.
-  void markIndeterminate(InstCounterType T);
+  /// \returns the current value of the counter \p T at this point in time
+  unsigned count(InstCounterType T) const;
 
-  /// Mark the counter \p T as being "out-of-order", meaning records may retire
-  /// in any order. This sets the height of all records to zero.
-  void markOutOfOrder(InstCounterType T);
-
-  /// \returns the current value of the counter \p T at this point in time, or
-  /// std::nullopt if \p T is in the indeterminate state.
-  std::optional<unsigned> count(InstCounterType T) const;
-
-  /// \returns true if the counter \p T is in an indeterminate state.
-  bool isIndeterminate(InstCounterType T) const;
-
-  /// \returns true if the counter \p T is out-of-order
-  bool isOutOfOrder(InstCounterType T) const;
-
-  /// \returns the set of pending HWEvents for \p T. If \p T is in an
-  /// indeterminate state, returns a conservative set of pending events instead.
+  /// \returns the set of pending HWEvents for \p T
   HWEvents getPendingEvents(InstCounterType T) const;
 
   /// \returns the set of live records recorded for \p T. This is the list of
   /// all instructions in-flight for that counter.
-  /// Note that if \p T is indeterminate, then this set is non-exhaustive. It
-  /// only contains the records this class knows about.
   ArrayRef<EventTrackerRecord> getLiveRecords(InstCounterType T) const;
 
   /// Prints a dump of the internal tracking state of this class for \p T to the
@@ -182,11 +157,6 @@ class EventTracker {
   void print(raw_ostream &OS, bool IgnoreEmpty = false,
              unsigned Indent = 0) const;
 
-  /// \returns true if the option to mimic legacy (SIInsertWaitcnts
-  ///          scoreboard-style) tracking of counters and pending events.
-  /// TODO: Remove in the future when legacy tracking is no longer needed.
-  static bool mimicsLegacyTracking();
-
 #if !defined(NDEBUG) || defined(EXPENSIVE_CHECKS)
   /// Verifies invariants of this class are respected.
   void verify() const;
@@ -209,21 +179,6 @@ class EventTracker {
     /// FIXME: We shouldn't need it for correctness so perhaps it should be
     /// removed entirely.
     uint32_t Count = 0;
-    /// An upper bound that persists across fixpoint iterations. This is only
-    /// used when \ref mimicsLegacyTracking returns true.
-    uint32_t PersistentUpperBound = 0;
-    /// Whether this counter is in an indeterminate state, which means that both
-    /// the set of LiveRecords and the Count are imprecise. This implies that
-    /// the counter is out-of-order as well.
-    bool IsIndeterminate = false;
-    /// Whether this counter is out-of-order, meaning records may retire in any
-    /// order and they all exist at a height of zero.
-    bool IsOutOfOrder = false;
-    /// Legacy-style tracking of pending events that is coarse and does not
-    /// leverage the live set of records. Only used when
-    /// \ref mimicsLegacyTracking returns true and not cleared between
-    /// iterations.
-    HWEvents LegacyPendingEvents;
 
     // TODO: We could imagine storing the per-predecessor height for incoming
     // events. We could achieve that by storing that as a map of ((ID, Pred),
diff --git a/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp b/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
index 06d95f089797c..5b722b9f09fa5 100644
--- a/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
+++ b/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
@@ -82,14 +82,12 @@ struct TrackerRecordsChecker {
   TrackerRecordsChecker(EventTracker &ET, InstCounterType T)
       : ET(ET), T(T), Records(ET.getLiveRecords(T)) {}
 
-  bool hasCount() { return ET.count(T).has_value(); }
-
-  unsigned getCount() { return *ET.count(T); }
+  unsigned getCount() { return ET.count(T); }
 
   bool empty() { return Records.empty(); }
 
-  // Has no records and count is zero (if there is one)
-  bool unused() { return empty() && (!hasCount() || !getCount()); }
+  // Has no records and count is zero
+  bool unused() { return empty() && !getCount(); }
 
   const EventTrackerRecord &cur() { return Records[CurElt]; }
 
@@ -131,7 +129,6 @@ body:             |
   auto &ET = visitAll(BB0);
 
   auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
-  EXPECT_TRUE(LoadCnt.hasCount());
   EXPECT_EQ(LoadCnt.getCount(), 1u);
   EXPECT_FALSE(LoadCnt.empty());
   EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
@@ -139,7 +136,6 @@ body:             |
   EXPECT_FALSE(LoadCnt.next());
 
   auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
-  EXPECT_TRUE(DsCnt.hasCount());
   EXPECT_EQ(DsCnt.getCount(), 1u);
   EXPECT_FALSE(DsCnt.empty());
   EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
@@ -150,7 +146,6 @@ body:             |
   EXPECT_TRUE(ExpCnt.unused());
 
   auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
-  EXPECT_TRUE(StoreCnt.hasCount());
   EXPECT_EQ(StoreCnt.getCount(), 2u);
   EXPECT_FALSE(StoreCnt.empty());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
@@ -187,7 +182,6 @@ body:             |
   auto &ET = visitAll(BB0);
 
   auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
-  EXPECT_TRUE(LoadCnt.hasCount());
   EXPECT_EQ(LoadCnt.getCount(), 2u);
   EXPECT_FALSE(LoadCnt.empty());
   EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
@@ -198,7 +192,6 @@ body:             |
   EXPECT_FALSE(LoadCnt.next());
 
   auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
-  EXPECT_TRUE(DsCnt.hasCount());
   EXPECT_EQ(DsCnt.getCount(), 2u);
   EXPECT_FALSE(DsCnt.empty());
   EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
@@ -212,7 +205,6 @@ body:             |
   EXPECT_TRUE(ExpCnt.unused());
 
   auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
-  EXPECT_TRUE(StoreCnt.hasCount());
   EXPECT_EQ(StoreCnt.getCount(), 2u);
   EXPECT_FALSE(StoreCnt.empty());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
@@ -250,7 +242,6 @@ body:             |
   auto &ET = visitAll(BB1);
 
   auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
-  EXPECT_TRUE(LoadCnt.hasCount());
   EXPECT_EQ(LoadCnt.getCount(), 1u);
   EXPECT_FALSE(LoadCnt.empty());
   EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
@@ -258,7 +249,6 @@ body:             |
   EXPECT_FALSE(LoadCnt.next());
 
   auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
-  EXPECT_TRUE(DsCnt.hasCount());
   EXPECT_EQ(DsCnt.getCount(), 1u);
   EXPECT_FALSE(DsCnt.empty());
   EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
@@ -269,7 +259,6 @@ body:             |
   EXPECT_TRUE(ExpCnt.unused());
 
   auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
-  EXPECT_TRUE(StoreCnt.hasCount());
   EXPECT_EQ(StoreCnt.getCount(), 2u);
   EXPECT_FALSE(StoreCnt.empty());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
@@ -313,7 +302,6 @@ body:             |
   auto &ET = visitAll(BB2);
 
   auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
-  EXPECT_TRUE(LoadCnt.hasCount());
   EXPECT_EQ(LoadCnt.getCount(), 1u);
   EXPECT_FALSE(LoadCnt.empty());
   EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
@@ -321,7 +309,6 @@ body:             |
   EXPECT_FALSE(LoadCnt.next());
 
   auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
-  EXPECT_TRUE(DsCnt.hasCount());
   EXPECT_EQ(DsCnt.getCount(), 1u);
   EXPECT_FALSE(DsCnt.empty());
   EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
@@ -332,7 +319,6 @@ body:             |
   EXPECT_TRUE(ExpCnt.unused());
 
   auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
-  EXPECT_TRUE(StoreCnt.hasCount());
   EXPECT_EQ(StoreCnt.getCount(), 2u);
   EXPECT_FALSE(StoreCnt.empty());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
@@ -376,7 +362,6 @@ body:             |
   auto &ET = visitAll(BB2);
 
   auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
-  EXPECT_TRUE(LoadCnt.hasCount());
   EXPECT_EQ(LoadCnt.getCount(), 1u);
   EXPECT_FALSE(LoadCnt.empty());
   EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
@@ -384,7 +369,6 @@ body:             |
   EXPECT_FALSE(LoadCnt.next());
 
   auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
-  EXPECT_TRUE(DsCnt.hasCount());
   EXPECT_EQ(DsCnt.getCount(), 1u);
   EXPECT_FALSE(DsCnt.empty());
   EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
@@ -395,7 +379,6 @@ body:             |
   EXPECT_TRUE(ExpCnt.unused());
 
   auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
-  EXPECT_TRUE(StoreCnt.hasCount());
   // Count is 1 because we have 1 event max across all predecessors.
   EXPECT_EQ(StoreCnt.getCount(), 1u);
   EXPECT_FALSE(StoreCnt.empty());
@@ -449,7 +432,6 @@ body:             |
   auto &ET = visitAll(BB3);
 
   auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
-  EXPECT_TRUE(LoadCnt.hasCount());
   EXPECT_EQ(LoadCnt.getCount(), 1u);
   EXPECT_FALSE(LoadCnt.empty());
   EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
@@ -457,7 +439,6 @@ body:             |
   EXPECT_FALSE(LoadCnt.next());
 
   auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
-  EXPECT_TRUE(DsCnt.hasCount());
   EXPECT_EQ(DsCnt.getCount(), 1u);
   EXPECT_FALSE(DsCnt.empty());
   EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
@@ -468,7 +449,6 @@ body:             |
   EXPECT_TRUE(ExpCnt.unused());
 
   auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
-  EXPECT_TRUE(StoreCnt.hasCount());
   EXPECT_EQ(StoreCnt.getCount(), 1u);
   EXPECT_FALSE(StoreCnt.empty());
   EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
@@ -476,62 +456,6 @@ body:             |
   EXPECT_FALSE(StoreCnt.next());
 }
 
-/// Basic diamond CFG, the store is carried all the way into bb3 and uniqued
-/// again so only 1 instance of the record is present in bb3.
-TEST_F(AMDGPUGFX12EventTrackingTest, BasicDiamondFlags) {
-  StringRef MIR = R"(
-name:            BasicDiamondFlags
-body:             |
-  bb.0:
-    successors: %bb.1, %bb.2
-
-    S_CBRANCH_SCC1 %bb.1, implicit $scc
-    S_BRANCH %bb.2
-
-  bb.1:
-    successors: %bb.3
-    S_BRANCH %bb.3
-
-  bb.2:
-    successors: %bb.3
-    S_BRANCH %bb.3
-
-  bb.3:
-    S_ENDPGM 0
-...
-)";
-  ASSERT_TRUE(parseMIR(MIR));
-  MachineFunction &MF = getMF("BasicDiamondFlags");
-  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
-  MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
-  MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
-  MachineBasicBlock &BB3 = *MF.getBlockNumbered(3);
-
-  visitAll(BB0);
-  visitAll(BB1, /*AfterVisit=*/[&](EventTracker &ET) {
-    ET.markIndeterminate(STORE_CNT);
-  });
-  visitAll(BB2, [&](EventTracker &ET) { ET.markOutOfOrder(LOAD_CNT); });
-  auto &ET = visitAll(BB3);
-
-  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
-  EXPECT_TRUE(LoadCnt.unused());
-  // We inherit the out of order flag is either predecessor has it.
-  EXPECT_TRUE(LoadCnt.ET.isOutOfOrder(LOAD_CNT));
-  EXPECT_FALSE(LoadCnt.ET.isIndeterminate(LOAD_CNT));
-
-  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
-  EXPECT_TRUE(DsCnt.unused());
-
-  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
-  EXPECT_TRUE(ExpCnt.unused());
-
-  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
-  EXPECT_TRUE(StoreCnt.unused());
-  // We inherit the indeterminate flag is either predecessor has it.
-  EXPECT_TRUE(StoreCnt.ET.isIndeterminate(STORE_CNT));
-}
-
 /// Assymetrical diamond
 ///   - bb0 has two stores
 ///   - bb1 adds another store without any waits.
@@ -591,7 +515,6 @@ body:             |
 
   auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
 
-  EXPECT_TRUE(StoreCnt.hasCount());
   EXPECT_EQ(StoreCnt.getCount(), 3u);
   EXPECT_FALSE(StoreCnt.empty());
 
@@ -669,7 +592,6 @@ body:             |
 
   auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
 
-  EXPECT_TRUE(StoreCnt.hasCount());
   EXPECT_EQ(StoreCnt.getCount(), 2u);
   EXPECT_FALSE(StoreCnt.empty());
 



More information about the llvm-commits mailing list