[llvm] [AMDGPU][NewInsertWaitcnt] Add Event Tracker (PR #226970)

Pierre van Houtryve via llvm-commits llvm-commits at lists.llvm.org
Mon Sep 28 05:33:20 PDT 2026


https://github.com/Pierre-vh updated https://github.com/llvm/llvm-project/pull/226970

>From f8c431c6a6db376431c2f2cea3072be15840baa6 Mon Sep 17 00:00:00 2001
From: pvanhout <pierre.vanhoutryve at amd.com>
Date: Mon, 28 Sep 2026 12:44:27 +0200
Subject: [PATCH] [AMDGPU][NewInsertWaitcnt] Add Event Tracker

See #226335

Introduction

Adds the basic infrastructure needed to track in-flight records for each `InstCounterType`.
This is not a full replacement of `SIInsertWaitcnts::WaitcntBrackets` and it does not
aim to be one.

One goal with the new implementation is to avoid having a single "tracker" class
that does everything. Instead, the `EventTracker` aims to do one thing and do it right:
preserve the history of a counter (instructions + events issued) and the value of the counter.
Any specific queries, such as "what STORE_CNT do I need to use this RU" is something that belongs
to helper methods in the client of the class or a utils file.

Class Design

This class is designed to use a single "source of truth" for all information, which is a compact
timeline of events. The intent is that we'll just iterate the timeline to answer certain queries
instead of baking-in a bunch of `DenseMaps` for every query dimension we need.
This considerably simplifies the implementation. Most of the code in the files here are
boilerplate, assertions, debug dumps, verification methods, and so on. There is very little "critical" logic.

The design/API presented here is *not* final - it'll evolve as we aim for feature parity with
`SIInsertWaitcnts` in the `NewInsertWaitcnts` pass.

Performance

As this design leans on iterating the timeline possibly multiple times per instruction, I made sure it'd be fast
by keeping the record type small (16B right now, may grow to 32) and stored contiguously.
I also did some profiling on very big test cases by shadowing `WaitcntBrackets` and doing many (dozens) of timeline
iteration each instruction for every counter, and there is no visible spike in the flame graph of the profiler.

Testing

I tested this implementation by shadowing the `WaitcntBrackets` class in a downstream branch and checking
my results against it, until I was confident that the values given by the `EventTracker` were as precise or
more precise than the values given by `WaitcntBrackets` for the entire test suite.

Use of AI

AI was exclusively used for 2 things:

- As a design companion, helping me flesh out architecture ideas in "plan mode" so I could iterate fast.
  I also used it as a research tool to find prior art, gather data for me to verify, etc.
- As a refactoring/boilerplate generation tool, such as creating files with basic boilerplate,
  generating scaffolding like methods/class declarations/etc.

AI was not used to write any serious code/algorithms in any function of either the unit tests or the implementation files.

Assisted-by: Claude Opus (4.8/5.5) and Sonnet (5)
---
 .../lib/Target/AMDGPU/AMDGPUEventTracking.cpp | 477 ++++++++++++++
 llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h  | 354 ++++++++++
 llvm/lib/Target/AMDGPU/AMDGPUHWEvents.h       |  27 +
 llvm/lib/Target/AMDGPU/CMakeLists.txt         |   1 +
 .../Target/AMDGPU/AMDGPUEventTrackingTest.cpp | 618 ++++++++++++++++++
 llvm/unittests/Target/AMDGPU/CMakeLists.txt   |   1 +
 6 files changed, 1478 insertions(+)
 create mode 100644 llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
 create mode 100644 llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
 create mode 100644 llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp

diff --git a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
new file mode 100644
index 0000000000000..8b383b45de67a
--- /dev/null
+++ b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.cpp
@@ -0,0 +1,477 @@
+//===- AMDGPUEventTracking.cpp ----------------------------------*- C++ -*-===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+#include "AMDGPUEventTracking.h"
+#include "AMDGPUHWEvents.h"
+#include "AMDGPUWaitcntUtils.h"
+#include "llvm/CodeGen/MachineBasicBlock.h"
+#include "llvm/CodeGen/MachineFunction.h"
+#include "llvm/CodeGen/MachineInstr.h"
+#include "llvm/Support/CommandLine.h"
+#include "llvm/Support/Debug.h"
+#include <algorithm>
+#include <optional>
+
+#define DEBUG_TYPE "amdgpu-event-tracking"
+
+namespace llvm {
+
+/// Mimic legacy (coarse) tracking of counter state.
+static cl::opt<bool> MimicLegacyTracking("amdgpu-event-legacy-tracking",
+                                         cl::init(false));
+
+#ifndef NDEBUG
+static cl::opt<bool> EventTrackerPrintAll(
+    "amdgpu-event-tracker-print-all", cl::init(false),
+    cl::desc("When using -debug, print the full set of live events every time "
+             "an event is added or removed"));
+#endif
+
+namespace AMDGPU {
+namespace eventtracking {
+
+namespace {
+static bool greaterThan(const EventTrackerRecord &A,
+                        const EventTrackerRecord &B) {
+  return A.getScore() > B.getScore();
+}
+} // namespace
+
+EventTrackerRecord::EventTrackerRecord(EventTrackingContext &Ctx,
+                                       MachineInstr *MI, SingleHWEvent Kind,
+                                       uint32_t Score)
+    : EventTrackerRecord(Ctx.nextDynamicInstanceID(), MI, Kind, Score) {}
+
+void EventTrackerRecord::print(raw_ostream &OS, bool PrintMI,
+                               unsigned Indent) const {
+  OS.indent(Indent) << "#" << ID << " " << Kind << " (Score=" << Score << "): ";
+  if (PrintMI && MI)
+    OS << *MI;
+  else
+    OS << "MI@" << (void *)MI << "\n";
+}
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+LLVM_DUMP_METHOD void EventTrackerRecord::dump() const {
+  dbgs() << "\n";
+  print(dbgs(), /*PrintMI=*/true);
+  dbgs() << "\n";
+}
+#endif
+
+EventTracker::EventTracker(MachineBasicBlock &MBB, EventTrackingContext &ETC)
+    : MBB(&MBB), Ctx(&ETC) {
+  Counters.resize(ETC.counters().size());
+  for (const CounterInfo &Info : ETC.counters())
+    Counters[Info.CounterT].CI = &Info;
+}
+
+void EventTracker::enterBlock() {
+  LLVM_DEBUG(dbgs() << "\n[EventTracker] Entering ";
+             MBB->printAsOperand(dbgs()); dbgs() << "\n");
+
+  // FIXME: This is a bit hacky, but we need to save the old state in case the
+  // MBB is also its own predecessor. Revisit when the design and clients of
+  // this class are set in stone.
+
+  SmallVector<EventTracker *> Preds;
+  bool IsSelfPred = false;
+  if (Preds.empty() && !MBB->pred_empty()) {
+    for (MachineBasicBlock *Pred : MBB->predecessors()) {
+      if (Pred == MBB)
+        IsSelfPred = true;
+      else
+        Preds.push_back(&(*Ctx)[Pred]);
+    }
+  }
+
+  if (IsSelfPred) {
+    EventTracker SelfCopy = *this;
+    Preds.push_back(&SelfCopy);
+    clear();
+    recordIncomings(*Ctx, Preds);
+  } else {
+    clear();
+    recordIncomings(*Ctx, Preds);
+  }
+}
+
+void EventTracker::leaveBlock() {
+  LLVM_DEBUG(dbgs() << "[EventTracker] Leaving "; MBB->printAsOperand(dbgs());
+             dbgs() << "\n");
+}
+
+void EventTracker::record(MachineInstr &MI, SingleHWEvent Event) {
+  LLVM_DEBUG(dbgs() << "[EventTracker] Recording " << Event << ": " << MI);
+
+  EventTrackerRecord Rec = EventTrackerRecord(*Ctx, &MI, Event);
+  [[maybe_unused]] bool FoundMatch = false;
+  for (auto &CD : Counters) {
+    if (!CD.CI->Events.contains(Event))
+      continue;
+
+    FoundMatch = true;
+    ++CD.Count;
+    CD.LegacyPendingEvents |= Event;
+
+    // Do not age records if we are out-of-order.
+    if (!CD.IsOutOfOrder) {
+      // NB: There is an intentional tradeoff here. We could avoid this loop by
+      // instead storing a timestamp in each record, and having a
+      // constantly-increasing clock to infer the score (clock-timestamp is
+      // score). However, it'd:
+      //  - Complexify fetching the score (`EventTrackerRecord` cannot answer it
+      //    on its own anymore and we need a separate query/wrapper).
+      //  - Make merge of incoming records a bit more annoying (we'd need to
+      //    rebase the `clock`).
+      //  - Potentially demand (much) more space in EventTrackerRecord to store
+      //    bigger numbers.
+      //
+      // All in all, I think this small loop is fine for now, but we can still
+      // change the system if we have data backed up by profiling to
+      // justify the change.
+      for (auto &Live : CD.LiveRecords)
+        Live.setScore(Live.getScore() + 1); // Age all existing events.
+    }
+
+    CD.LiveRecords.push_back(Rec);
+
+#ifndef NDEBUG
+    LLVM_DEBUG(if (EventTrackerPrintAll) {
+      dbgs().indent(2) << "Updated Timeline:\n";
+      print(dbgs(), CD, /*Indent=*/4);
+    });
+#endif
+  }
+
+  assert(FoundMatch && "Event has no matching InstCounterType!");
+}
+
+void EventTracker::wait(InstCounterType T, unsigned N) {
+  auto &CD = get(T);
+  LLVM_DEBUG(dbgs() << "[EventTracker] Wait on " << getInstCounterName(T)
+                    << " for " << N << "\n");
+
+  // Fast path for clearing the counter
+  if (N == 0) {
+    CD.LiveRecords.clear();
+    CD.Count = 0;
+    CD.IsIndeterminate = false;
+    CD.IsOutOfOrder = false;
+    CD.LegacyPendingEvents = HWEvents();
+    return;
+  }
+
+  CD.Count = std::min(CD.Count, N);
+
+  // Don't bother erasing stuff if we are out-of-order. All records have a score
+  // of zero in such cases.
+  if (!CD.IsOutOfOrder) {
+    auto *RmIt = remove_if(CD.LiveRecords, [&](EventTrackerRecord &E) {
+      if (E.getScore() < N)
+        return false;
+      LLVM_DEBUG(dbgs() << "  | Removing "; E.print(dbgs()));
+      return true;
+    });
+    CD.LiveRecords.erase(RmIt, CD.LiveRecords.end());
+  }
+
+  LLVM_DEBUG(dbgs() << "  | => Updated Count:" << CD.Count << "\n");
+
+#ifndef NDEBUG
+  LLVM_DEBUG(if (EventTrackerPrintAll) {
+    dbgs().indent(2) << "Updated Timeline:\n";
+    print(dbgs(), CD, /*Indent=*/4);
+  });
+#endif
+}
+
+void EventTracker::markIndeterminate(InstCounterType T) {
+  LLVM_DEBUG(dbgs() << "[EventTracker] Marking " << getInstCounterName(T)
+                    << " as indeterminate!\n");
+  auto &CD = get(T);
+  CD.IsIndeterminate = true;
+  markOutOfOrder(T);
+}
+
+void EventTracker::markOutOfOrder(InstCounterType T) {
+  LLVM_DEBUG(dbgs() << "[EventTracker] Marking " << getInstCounterName(T)
+                    << " as out-of-order!\n");
+  auto &CD = get(T);
+  CD.IsOutOfOrder = true;
+  for (auto &Rec : CD.LiveRecords)
+    Rec.setScore(0);
+}
+
+std::optional<unsigned> EventTracker::count(InstCounterType T) const {
+  auto &CD = get(T);
+  if (CD.IsIndeterminate)
+    return std::nullopt;
+  return get(T).Count;
+}
+
+bool EventTracker::isIndeterminate(InstCounterType T) const {
+  return get(T).IsIndeterminate;
+}
+
+bool EventTracker::isOutOfOrder(InstCounterType T) const {
+  return get(T).IsOutOfOrder;
+}
+
+HWEvents EventTracker::getPendingEvents(InstCounterType T) const {
+  const auto &CD = get(T);
+  if (CD.IsIndeterminate)
+    return CD.CI->Events; // return all events
+
+  if (MimicLegacyTracking)
+    return CD.LegacyPendingEvents;
+
+  HWEvents Res;
+  for (const auto &E : get(T).LiveRecords)
+    Res |= E.getKind();
+  return Res;
+}
+
+ArrayRef<EventTrackerRecord>
+EventTracker::getLiveRecords(InstCounterType T) const {
+  return get(T).LiveRecords;
+}
+
+void EventTracker::print(raw_ostream &OS, InstCounterType T) const {
+  print(OS, get(T));
+}
+
+bool EventTracker::mimicsLegacyTracking() { return MimicLegacyTracking; }
+
+#if !defined(NDEBUG) || defined(EXPENSIVE_CHECKS)
+void EventTracker::verify() const {
+  assert(MBB && Ctx && "Invalid internal state!");
+
+  for (auto &C : Counters) {
+    const auto OnError = [&]() {
+      dbgs() << "EventTracker verification error\n";
+      print(dbgs(), C);
+    };
+
+    if (C.IsIndeterminate) {
+      if (!C.IsOutOfOrder) {
+        OnError();
+        assert(false && "IsIndeterminate but not IsOutOfOrder");
+      }
+      continue;
+    }
+
+    if (C.IsOutOfOrder) {
+      if (!all_of(C.LiveRecords, [](auto &R) { return R.getScore() == 0; })) {
+        OnError();
+        assert(false &&
+               "IsOutOfOrder but some records do not have a score of 0!");
+      }
+    }
+
+    if (C.Count > C.LiveRecords.size() && !mimicsLegacyTracking()) {
+      OnError();
+      assert(false &&
+             "'Count' is inconsistent with the number of live records");
+    }
+
+    for (auto &E : C.LiveRecords) {
+      if (E.getScore() > C.Count) {
+        OnError();
+        dbgs() << "Concerning Record:";
+        E.print(dbgs());
+        assert(false && "record score is out of range");
+      }
+    }
+
+    // Check live records are sorted
+    if (!is_sorted(C.LiveRecords, greaterThan)) {
+      OnError();
+      assert(false && "live records are not sorted!");
+    }
+  }
+}
+#endif
+
+void EventTracker::print(raw_ostream &OS, bool IgnoreEmpty,
+                         unsigned Indent) const {
+  OS.indent(Indent) << "EventTracker for ";
+  MBB->printAsOperand(OS);
+  OS << ":";
+  if (IgnoreEmpty) {
+    if (all_of(Counters, [](const auto &CD) { return CD.Count == 0; })) {
+      OS << " (empty)\n";
+      return;
+    }
+  }
+
+  OS << "\n";
+  for (auto &C : Counters) {
+    if (IgnoreEmpty && C.Count == 0)
+      continue;
+    print(dbgs(), C, Indent + 2);
+  }
+}
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+LLVM_DUMP_METHOD void EventTracker::dump() const {
+  dbgs() << "\n";
+  print(dbgs());
+  dbgs() << "\n";
+}
+#endif
+
+void EventTracker::clear() {
+  for (auto &C : Counters) {
+    C.LiveRecords.clear();
+    C.Count = 0;
+    C.IsIndeterminate = false;
+    C.IsOutOfOrder = false;
+  }
+}
+
+void EventTracker::recordIncomings(EventTrackingContext &ETC,
+                                   ArrayRef<EventTracker *> Preds) {
+  LLVM_DEBUG(if (!Preds.empty()) {
+    dbgs() << "[EventTracker] Recording incoming events (merge) from "
+              "predecessors:\n";
+    for (EventTracker *Pred : Preds) {
+      Pred->print(dbgs(), /*IgnoreEmpty=*/true, /*Indent=*/2);
+    }
+  });
+
+  /// Iterate over all counters that are available to us.
+  for (auto &CI : ETC.counters()) {
+    auto &CData = Counters[CI.CounterT];
+    assert(CData.LiveRecords.empty());
+
+    DenseMap<EventTrackerRecord::DynamicInstanceID, EventTrackerRecord> Acc;
+
+    for (EventTracker *Pred : Preds) {
+      auto &PredCData = Pred->Counters[CI.CounterT];
+
+      // Merge domain for the count value:
+      CData.Count = std::max(CData.Count, PredCData.Count);
+      // Merge domain for the legacy pending events.
+      CData.LegacyPendingEvents |= PredCData.LegacyPendingEvents;
+      // Merge domain for the indeterminate state.
+      CData.IsIndeterminate |= PredCData.IsIndeterminate;
+      // Merge domain for the out-of-order state.
+      CData.IsOutOfOrder |= PredCData.IsOutOfOrder;
+
+      for (EventTrackerRecord &PredEntry : PredCData.LiveRecords) {
+        EventTrackerRecord::DynamicInstanceID ID = PredEntry.getID();
+        auto It = Acc.find(ID);
+        if (It != Acc.end()) {
+          auto &AccVal = It->second;
+          AccVal.setScore(std::min(AccVal.getScore(), PredEntry.getScore()));
+          assert(
+              PredEntry.getMI() == AccVal.getMI() &&
+              PredEntry.getKind() == AccVal.getKind() &&
+              "EventTrackerRecord have same DynamicInstanceID, but different "
+              "MachineInstr/HWEvent kind, which should not be possible");
+        } else
+          Acc.insert({ID, PredEntry});
+      }
+    }
+
+    if (MimicLegacyTracking) {
+      CData.PersistentUpperBound =
+          std::max(CData.PersistentUpperBound, CData.Count);
+      CData.Count = CData.PersistentUpperBound;
+    }
+
+    auto AccVals = Acc.values();
+    CData.LiveRecords.append(AccVals.begin(), AccVals.end());
+
+    // Sort records by Score (descending) for consistent iteration.
+    stable_sort(CData.LiveRecords, greaterThan);
+  }
+
+  LLVM_DEBUG(if (!Preds.empty()) {
+    dbgs() << "[EventTracker] Timeline after recording incomings:\n";
+    print(dbgs(), /*IgnoreEmpty=*/true, /*Indent=*/2);
+  });
+}
+
+void EventTracker::print(raw_ostream &OS, const CounterData &CD,
+                         unsigned Indent) {
+  OS.indent(Indent) << getInstCounterName(CD.CI->CounterT)
+                    << " (Count=" << CD.Count
+                    << ", PersistentUpperBound=" << CD.PersistentUpperBound
+                    << ", LiveRecords=" << CD.LiveRecords.size()
+                    << ", IsOutOfOrder=" << CD.IsOutOfOrder
+                    << ", IsIndeterminate=" << CD.IsIndeterminate << ")\n";
+  for (const auto &E : CD.LiveRecords) {
+    OS.indent(Indent + 2);
+    E.print(OS);
+  }
+}
+
+EventTrackingContext::EventTrackingContext(MachineFunction &MF,
+                                           ArrayRef<CounterInfo> Counters)
+    : CounterInfos(Counters) {
+  LLVM_DEBUG(dbgs() << "\n[EventTrackingContext] CounterInfos for "
+                    << MF.getName() << "\n";
+             for (const auto &CI
+                  : CounterInfos) {
+               dbgs().indent(2)
+                   << AMDGPU::getInstCounterName(CI.CounterT) << " ";
+               if (CI.Events.none()) {
+                 dbgs() << " (unused - no HWEvents assigned)\n";
+               } else {
+                 dbgs() << "(Limit=" << CI.Limit << ") " << CI.Events << "\n";
+               }
+             });
+
+  Trackers.reserve(MF.size());
+  for (MachineBasicBlock &MBB : MF)
+    Trackers[&MBB] = std::make_unique<EventTracker>(MBB, *this);
+}
+
+EventTracker &EventTrackingContext::operator[](MachineBasicBlock *MBB) {
+  assert(MBB);
+  return *Trackers.at(MBB);
+}
+
+#if !defined(NDEBUG) || defined(EXPENSIVE_CHECKS)
+void EventTrackingContext::verify() const {
+  for (const auto &[MBB, Tracker] : Trackers) {
+    assert(MBB && "Unexpected nullptr entry!");
+    Tracker->verify();
+  }
+
+  // Check CounterInfos is sane.
+  for (auto [Idx, CI] : enumerate(CounterInfos)) {
+    assert(Idx == CI.CounterT && "CounterInfo is in wrong position!");
+    assert(CI.Events.any() &&
+           "InstCounterType has no event associated with it!");
+  }
+}
+#endif
+
+void EventTrackingContext::print(raw_ostream &OS) const {
+  for (const auto &[MBB, Tracker] : Trackers) {
+    MBB->printAsOperand(OS);
+    OS << ":\n";
+    Tracker->print(OS, /*Indent=*/2);
+  }
+}
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+LLVM_DUMP_METHOD void EventTrackingContext::dump() const {
+  dbgs() << "\n";
+  print(dbgs());
+  dbgs() << "\n";
+}
+#endif
+
+} // namespace eventtracking
+} // namespace AMDGPU
+
+} // namespace llvm
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
new file mode 100644
index 0000000000000..c1e56fc184569
--- /dev/null
+++ b/llvm/lib/Target/AMDGPU/AMDGPUEventTracking.h
@@ -0,0 +1,354 @@
+//===- AMDGPUEventTracking.h ------------------------------------*- C++ -*-===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+//
+/// \file Tracks a per-InstCounterType event timeline which preserves
+/// information about previously encountered MachineInstr and HWEvents.
+//
+//===----------------------------------------------------------------------===//
+
+#ifndef LLVM_LIB_TARGET_AMDGPU_UTILS_AMDGPUEVENTTRACKING_H
+#define LLVM_LIB_TARGET_AMDGPU_UTILS_AMDGPUEVENTTRACKING_H
+
+#include "AMDGPUHWEvents.h"
+#include "AMDGPUWaitcntUtils.h"
+#include "llvm/ADT/ArrayRef.h"
+#include "llvm/ADT/DenseMap.h"
+#include "llvm/ADT/SmallVector.h"
+#include "llvm/ADT/Twine.h"
+#include <memory>
+
+namespace llvm {
+class raw_ostream;
+class MachineBasicBlock;
+class MachineOperand;
+class MachineFunction;
+
+namespace AMDGPU {
+namespace eventtracking {
+
+class EventTrackingContext;
+class EventTracker;
+
+/// FIXME: Make this a generic util?
+struct CounterInfo {
+  constexpr CounterInfo(InstCounterType T, HWEvents Events, unsigned Limit)
+      : CounterT(T), Events(Events), Limit(Limit) {}
+
+  /// Type of counter this is.
+  /// This must match the index in the container, e.g. CounterT=2 must be at
+  /// index 2.
+  InstCounterType CounterT;
+  /// HWEvents for this counter.
+  HWEvents Events;
+  /// Hardware limit for this counter.
+  unsigned Limit;
+};
+
+/// Represents entries in the \ref EventTracker.
+///
+/// Records are all uniquely identified by a \ref DynamicInstanceID. For each
+/// unique \ref DynamicInstanceID value, all \ref EventTrackerRecord that use
+/// that ID should have the same MI and Kind. This is enforced by exposing these
+/// as read-only, and making the constructor assign a new \ref DynamicInstanceID
+/// every time.
+///
+/// Only the score can change as it may be unique to each instance
+/// of \ref EventTracker that carry it.
+class EventTrackerRecord {
+public:
+  /// An always-increasing counter used to represent a dynamic instance of a
+  /// record.
+  ///
+  /// Whenever we add a new \ref EventTrackerRecord, even if it's one we already
+  /// have seen in a previous dataflow iteration, this counter is increased so
+  /// that the new record has a unique `DynamicInstanceID`.
+  using DynamicInstanceID = uint32_t;
+
+  EventTrackerRecord(EventTrackingContext &Ctx, MachineInstr *MI,
+                     SingleHWEvent Kind, uint32_t Score = 0);
+
+  /// \returns the ID uniquely identifying this record across an entire
+  /// EventTrackingContext. Whenever we revisit an instruction (when iterating
+  /// until a fixpoint is reached), we give it a new ID. This is used to
+  /// represent records carried over from previous iterations of the same basic
+  /// block.
+  DynamicInstanceID getID() const { return ID; }
+
+  /// \returns the MachineInstr that originated this record.
+  MachineInstr *getMI() const { return MI; }
+
+  /// \returns the kind of record this is, as a \ref SingleHWEvent.
+  SingleHWEvent getKind() const { return Kind; }
+
+  /// \returns the score of this record.
+  uint32_t getScore() const { return Score; }
+
+  /// Sets the score of this record to \p NewScore.
+  void setScore(uint32_t NewScore) {
+    Score = NewScore;
+    assert(Score == NewScore && "Score overflow!");
+  }
+
+  void print(raw_ostream &OS, bool PrintMI = true, unsigned Indent = 0) const;
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+  LLVM_DUMP_METHOD void dump() const;
+#endif
+
+private:
+  EventTrackerRecord(DynamicInstanceID ID, MachineInstr *MI, SingleHWEvent Kind,
+                     uint32_t Score = 0)
+      : MI(MI), ID(ID), Kind(Kind) {
+    setScore(Score);
+  }
+
+  MachineInstr *MI;
+  DynamicInstanceID ID;
+  SingleHWEvent Kind;
+  // Score should already never exceed uint8_t limit in normal circumstances as
+  // most counter types only use up to 6 bits encoding for the waitcnts. 16 bit
+  // is a very generous limit, we can probably shrink that at some point.
+  uint16_t Score;
+};
+
+/// This assert serves as a reminder to be mindful of the size of the object.
+static_assert(sizeof(EventTrackerRecord) == 16,
+              "EventTrackerRecord should remain small to optimize its layout "
+              "within cache lines, for maximum iteration speed");
+
+/// Per-MBB tracking context.
+///
+/// Tracks data across the following domains:
+///   - Current value (count) of each instruction counter.
+///   - In-flight (alive) \ref EventTrackerRecord of each instruction counter.
+///
+/// This class is only responsible for tracking records for every
+/// InstCounterType. It does not deal with calculating the waitcnts needed, or
+/// doing more advance reasoning over the timeline for specific queries (e.g.
+/// finding an aliasing store). These responsibilities are for
+/// utils/wrappers/users of the class.
+///
+/// The API should be kept as simple and clear as possible.
+class EventTracker {
+public:
+  EventTracker(MachineBasicBlock &MBB, EventTrackingContext &ET);
+
+  /// \defgroup MachineBasicBlock entry and exit
+  /// \{
+
+  /// Notify this EventTracker that we are going to begin recording events.
+  /// In case this is not the first time we are going through this block, this
+  /// clears the internal state of the tracker and re-imports all incoming
+  /// tracking state from the predecessors.
+  void enterBlock();
+
+  /// Notify this EventTracker that we are done recording events.
+  void leaveBlock();
+
+  /// \}
+
+  /// \defgroup InstCounters Tracking Entrypoints
+  /// Methods update the state of the InstCounters by adding/removing events
+  /// or signaling certain special conditions.
+  /// \{
+
+  /// Record an event of type \p Event at a MachineInstr \p MI, which will
+  /// affect all counters that have \p Event in their event set.
+  void record(MachineInstr &MI, SingleHWEvent Event);
+
+  /// Notify that we waited until the counter \p T reached the value \p N before
+  /// continuing execution of the program (and recording more events).
+  ///
+  /// This affects the count of \p T, an removes all records that have a score
+  /// greater than or equal to \p N.
+  ///
+  /// If \p N is zero, then \p T will no longer be in an indeterminate or
+  /// out-of-order state afterwards if it previously was in such a state.
+  void wait(InstCounterType T, unsigned N = 0);
+
+  /// Mark the counter \p T as being in an indeterminate state. This means that
+  /// we no longer accurately track \p T because there may be more records we do
+  /// not know about. This primarily affects \ref getPendingEvents and
+  /// \ref count.
+  ///
+  /// Implies \ref markOutOfOrder for \p T as well.
+  void markIndeterminate(InstCounterType T);
+
+  /// Mark the counter \p T as being "out-of-order", meaning records may retire
+  /// in any order. This sets the score of all records to zero.
+  void markOutOfOrder(InstCounterType T);
+
+  /// \}
+
+  /// \defgroup InstCounters Tracking Queries
+  /// Query the current state of each InstCounter without modifying it.
+  /// \{
+
+  /// \returns the current value of the counter \p T at this point in time, or
+  /// std::nullopt if \p T is in the indeterminate state.
+  std::optional<unsigned> count(InstCounterType T) const;
+
+  /// \returns true if the counter \p T is in an indeterminate state.
+  bool isIndeterminate(InstCounterType T) const;
+
+  /// \returns true if the counter \p T is out-of-order
+  bool isOutOfOrder(InstCounterType T) const;
+
+  /// \returns the set of pending HWEvents for \p T. If \p T is in an
+  /// indeterminate state, returns a conservative set of pending events instead.
+  HWEvents getPendingEvents(InstCounterType T) const;
+
+  /// \returns the set of live records recorded for \p T. This is the list of
+  /// all instructions in-flight for that counter.
+  /// Note that if \p T is indeterminate, then this set is non-exhaustive. It
+  /// only contains the records this class knows about.
+  ArrayRef<EventTrackerRecord> getLiveRecords(InstCounterType T) const;
+
+  /// \}
+
+  /// \defgroup Miscellaneous helpers
+  /// \{
+
+  /// Prints a dump of the internal tracking state of this class for \p T to the
+  /// stream \p OS.
+  void print(raw_ostream &OS, InstCounterType T) const;
+
+  /// Prints a dump of all internal tracking state of this class to the stream
+  /// \p OS. If \p IgnoreEmpty is true, do not print counters with a count of 0.
+  void print(raw_ostream &OS, bool IgnoreEmpty = false,
+             unsigned Indent = 0) const;
+
+  /// \returns true if the option to mimic legacy (SIInsertWaitcnts
+  ///          scoreboard-style) tracking of counters and pending events.
+  /// TODO: Remove in the future when legacy tracking is no longer needed.
+  static bool mimicsLegacyTracking();
+
+#if !defined(NDEBUG) || defined(EXPENSIVE_CHECKS)
+  /// Verifies invariants of this class are respected.
+  void verify() const;
+#endif
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+  LLVM_DUMP_METHOD void dump() const;
+#endif
+
+  /// \}
+
+private:
+  struct CounterData {
+    const CounterInfo *CI = nullptr;
+
+    /// Set of live records (events) that make up the `Count`.
+    SmallVector<EventTrackerRecord, 16> LiveRecords;
+    /// The current value of the counter. This is a max (upper bound) across
+    /// all possible execution paths at runtime. It cannot be inferred from the
+    /// LiveRecords alone and is thus a separate tracking domain.
+    uint32_t Count = 0;
+    /// An upper bound that persists across fixpoint iterations. This is only
+    /// used when \ref mimicsLegacyTracking returns true.
+    uint32_t PersistentUpperBound = 0;
+    /// Whether this counter is in an indeterminate state, which means that both
+    /// the set of LiveRecords and the Count are imprecise. This implies that
+    /// the counter is out-of-order as well.
+    bool IsIndeterminate = false;
+    /// Whether this counter is out-of-order, meaning records may retire in any
+    /// order and they all exist at a score of zero.
+    bool IsOutOfOrder = false;
+    /// Legacy-style tracking of pending events that is coarse and does not
+    /// leverage the live set of records. Only used when
+    /// \ref mimicsLegacyTracking returns true and not cleared between
+    /// iterations.
+    HWEvents LegacyPendingEvents;
+
+    // TODO: We could imagine storing the per-predecessor score for incoming
+    // events. We could achieve that by storing that as a map of ((ID, Pred),
+    // Score). This would allow identifying events that are "deep" in one branch
+    // but "shallow" in another, e.g. an event needing a waitcnt 1 for one pred,
+    // but a waitcnt 8 for another. Not sure if we can exploit that though?
+  };
+
+  CounterData &get(InstCounterType T) {
+    assert(Counters.size() > T && "T is out of range!");
+    return Counters[T];
+  }
+
+  const CounterData &get(InstCounterType T) const {
+    assert(Counters.size() > T && "T is out of range!");
+    return Counters[T];
+  }
+
+  /// Clears the tracked data, used when entering a block.
+  void clear();
+
+  /// Import all events from the incoming basic blocks in \p Preds and reconcile
+  /// divergence at joints.
+  void recordIncomings(EventTrackingContext &ETC,
+                       ArrayRef<EventTracker *> Preds);
+
+  static void print(raw_ostream &OS, const CounterData &CD,
+                    unsigned Indent = 0);
+
+  MachineBasicBlock *MBB;
+  EventTrackingContext *Ctx;
+
+  // NB: This, combined with the inline storage of LiveRecords, can lead to this
+  // class becoming quite big - verify the size of this object whenever a change
+  // is made.
+  SmallVector<CounterData, InstCounterType::NUM_INST_CNTS> Counters;
+};
+
+/// Per-MF Tracking Context.
+///
+/// This owns all \ref EventTrackers and keeps track of state that persists
+/// across dataflow analysis iterations, such as the current value of
+/// \ref DynamicInstanceID.
+class EventTrackingContext {
+public:
+  /// \param MF Machine Function
+  /// \param Counters The counters available to \p MF on this target.
+  EventTrackingContext(MachineFunction &MF, ArrayRef<CounterInfo> Counters);
+
+  /// Fetch the \ref MBBEventTracker of \p MBB.
+  EventTracker &operator[](MachineBasicBlock *MBB);
+
+  /// \returns the list of counters available to the current target.
+  ArrayRef<CounterInfo> counters() const { return CounterInfos; }
+
+#if !defined(NDEBUG) || defined(EXPENSIVE_CHECKS)
+  /// Verifies invariants of this class are respected.
+  void verify() const;
+#endif
+
+  void print(raw_ostream &OS) const;
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+  LLVM_DUMP_METHOD void dump() const;
+#endif
+
+private:
+  friend class EventTrackerRecord;
+
+  /// \returns a new, unique \ref DynamicInstanceID - only for use by
+  /// \ref EventTrackerRecord.
+  EventTrackerRecord::DynamicInstanceID nextDynamicInstanceID() {
+    assert(NextDynID + 1 > NextDynID && "DynamicInstanceIDs overflow!");
+    return ++NextDynID;
+  }
+
+  SmallVector<CounterInfo> CounterInfos;
+
+  EventTrackerRecord::DynamicInstanceID NextDynID = 0;
+  DenseMap<MachineBasicBlock *, std::unique_ptr<EventTracker>> Trackers;
+};
+
+} // namespace eventtracking
+} // namespace AMDGPU
+
+} // namespace llvm
+
+#endif
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUHWEvents.h b/llvm/lib/Target/AMDGPU/AMDGPUHWEvents.h
index eb25206b5ee04..b4ce2fa93eb16 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUHWEvents.h
+++ b/llvm/lib/Target/AMDGPU/AMDGPUHWEvents.h
@@ -165,6 +165,33 @@ class HWEvents {
   value_type Data = NONE;
 };
 
+/// Class to store singular HW events in the most compact way possible (u8).
+class SingleHWEvent {
+public:
+  using value_type = uint8_t;
+
+  static constexpr SingleHWEvent encode(HWEvents E) {
+    assert(E.size() == 1 && "expected exactly one event!");
+    return SingleHWEvent(countr_zero_constexpr(E.value()));
+  }
+
+  constexpr operator HWEvents() const { return HWEvents(1 << Data); }
+
+  constexpr value_type rawValue() const { return Data; }
+
+  constexpr bool operator==(const SingleHWEvent &Other) const {
+    return Data == Other.Data;
+  }
+  constexpr bool operator!=(const SingleHWEvent &Other) const {
+    return Data != Other.Data;
+  }
+
+private:
+  constexpr SingleHWEvent(value_type Data) : Data(Data) {}
+
+  value_type Data;
+};
+
 /// \param Inst A VMEM instruction (as per `SIInstrInfo::isVMEM`).
 /// \returns the simplified set of events triggered by the VMEM instruction \p
 /// Inst. The returned mask is not exhaustive, but is guaranteed to be a subset
diff --git a/llvm/lib/Target/AMDGPU/CMakeLists.txt b/llvm/lib/Target/AMDGPU/CMakeLists.txt
index b7e679a69a80d..07aa3aa33bc1e 100644
--- a/llvm/lib/Target/AMDGPU/CMakeLists.txt
+++ b/llvm/lib/Target/AMDGPU/CMakeLists.txt
@@ -54,6 +54,7 @@ add_llvm_target(AMDGPUCodeGen
   AMDGPUCodeGenPrepare.cpp
   AMDGPUCombinerHelper.cpp
   AMDGPUCtorDtorLowering.cpp
+  AMDGPUEventTracking.cpp
   AMDGPUExportClustering.cpp
   AMDGPUExportKernelRuntimeHandles.cpp
   AMDGPUFrameLowering.cpp
diff --git a/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp b/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
new file mode 100644
index 0000000000000..e00788f2386bf
--- /dev/null
+++ b/llvm/unittests/Target/AMDGPU/AMDGPUEventTrackingTest.cpp
@@ -0,0 +1,618 @@
+//===- AMDGPUEventTrackingTest.cpp ------------------------------*- C++ -*-===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+#include "AMDGPUEventTracking.h"
+#include "AMDGPUUnitTests.h"
+#include "AMDGPUWaitcntUtils.h"
+#include "MCTargetDesc/AMDGPUMCTargetDesc.h"
+#include "llvm/CodeGen/MachineFunction.h"
+#include "gtest/gtest.h"
+
+using namespace llvm;
+using namespace llvm::AMDGPU;
+using namespace llvm::AMDGPU::eventtracking;
+
+namespace {
+
+static constexpr unsigned CounterLimit = 12;
+
+// These are not accurate, they are simply for testing purposes.
+// We do not need to test every single counter accurately, that is the job
+// of the IR/MIR tests in tests/CodeGen/AMDGPU. We just need enough here
+// to validate that the EventTracker works.
+std::array<CounterInfo, 4> GFX12CounterInfos = {{
+    {LOAD_CNT, HWEvents::VMEM_READ_ACCESS, CounterLimit},
+    {DS_CNT, HWEvents::LDS_ACCESS, CounterLimit},
+    {EXP_CNT, HWEvents::EXP_GPR_LOCK, CounterLimit},
+    {STORE_CNT, HWEvents::VMEM_WRITE_ACCESS | HWEvents::SCRATCH_WRITE_ACCESS,
+     CounterLimit},
+}};
+
+class AMDGPUGFX12EventTrackingTest : public AMDGPUCodeGenTestBase {
+public:
+  void SetUp() override { setUpImpl("amdgpu12.00-amd-amdhsa", "gfx1200", ""); }
+};
+
+namespace {
+static EventTracker &
+visitAll(EventTrackingContext &Ctx, MachineBasicBlock &MBB,
+         function_ref<void(EventTracker &ET)> AfterVisit = nullptr) {
+  const GCNSubtarget &ST = MBB.getParent()->getSubtarget<GCNSubtarget>();
+
+  EventTracker &ET = Ctx[&MBB];
+  ET.enterBlock();
+  for (MachineInstr &MI : MBB) {
+    HWEvents Events =
+        getEventsFor(MI, ST, /*IsExpertMode=*/false, /*TgSplit=*/false);
+    for (HWEvents SingleEv : Events) {
+      ET.record(MI, SingleHWEvent::encode(SingleEv));
+    }
+  }
+  if (AfterVisit)
+    AfterVisit(ET);
+  ET.leaveBlock();
+  return ET;
+}
+
+/// Provides some helpers to declaratively check the state of the EventTracker
+/// for one counter. This provides helpers to check the general counter state
+/// (value, etc) but also allows iterating over the timeline from the oldest to
+/// the earliest element.
+///
+/// This makes the actual test cases clearer.
+struct TrackerRecordsChecker {
+  TrackerRecordsChecker(EventTracker &ET, InstCounterType T)
+      : ET(ET), T(T), Records(ET.getLiveRecords(T)) {}
+
+  bool hasCount() { return ET.count(T).has_value(); }
+
+  unsigned getCount() { return *ET.count(T); }
+
+  bool empty() { return Records.empty(); }
+
+  // Has no records and score is zero (if there is one)
+  bool unused() { return empty() && (!hasCount() || !getCount()); }
+
+  const EventTrackerRecord &cur() { return Records[CurElt]; }
+
+  /// Move to the next record.
+  /// \returns true on success, false if the end has been reached.
+  bool next() {
+    ++CurElt;
+    if (CurElt >= Records.size())
+      return false;
+    return true;
+  }
+
+  EventTracker &ET;
+  InstCounterType T;
+  ArrayRef<EventTrackerRecord> Records;
+  unsigned CurElt = 0;
+};
+
+} // namespace
+
+/// Check trivial straight line counting.
+TEST_F(AMDGPUGFX12EventTrackingTest, BasicTimeline) {
+  StringRef MIR = R"(
+name:            BasicTimeline
+body:             |
+  bb.0:
+
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    $vgpr4 = GLOBAL_LOAD_DWORD $vgpr4_vgpr5, 0, 0, implicit $exec
+    $vgpr0 = DS_READ_B32_gfx9 $vgpr1, 0, 0, implicit $exec
+    GLOBAL_STORE_DWORDX2 $vgpr1_vgpr2, $vgpr3_vgpr4, 0, 0, implicit $exec
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("BasicTimeline");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  auto &ET = visitAll(Ctx, BB0);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.hasCount());
+  EXPECT_EQ(LoadCnt.getCount(), 1u);
+  EXPECT_FALSE(LoadCnt.empty());
+  EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
+  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(LoadCnt.next());
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.hasCount());
+  EXPECT_EQ(DsCnt.getCount(), 1u);
+  EXPECT_FALSE(DsCnt.empty());
+  EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
+  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(DsCnt.next());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+  EXPECT_TRUE(StoreCnt.hasCount());
+  EXPECT_EQ(StoreCnt.getCount(), 2u);
+  EXPECT_FALSE(StoreCnt.empty());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 1u);
+  EXPECT_TRUE(StoreCnt.next());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(StoreCnt.next());
+}
+
+/// Check cases where a block loops back on itself.
+TEST_F(AMDGPUGFX12EventTrackingTest, SelfPredecessor) {
+  StringRef MIR = R"(
+name:            BasicTimeline
+body:             |
+  bb.0:
+
+    $vgpr4 = GLOBAL_LOAD_DWORD $vgpr4_vgpr5, 0, 0, implicit $exec
+    $vgpr0 = DS_READ_B32_gfx9 $vgpr1, 0, 0, implicit $exec
+    GLOBAL_STORE_DWORDX2 $vgpr1_vgpr2, $vgpr3_vgpr4, 0, 0, implicit $exec
+    S_CBRANCH_SCC1 %bb.0, implicit $scc
+    S_BRANCH %bb.1
+
+  bb.1:
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("BasicTimeline");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  // Iterate twice
+  visitAll(Ctx, BB0);
+  auto &ET = visitAll(Ctx, BB0);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.hasCount());
+  EXPECT_EQ(LoadCnt.getCount(), 2u);
+  EXPECT_FALSE(LoadCnt.empty());
+  EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
+  EXPECT_EQ(LoadCnt.cur().getScore(), 1u);
+  EXPECT_TRUE(LoadCnt.next());
+  EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
+  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(LoadCnt.next());
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.hasCount());
+  EXPECT_EQ(DsCnt.getCount(), 2u);
+  EXPECT_FALSE(DsCnt.empty());
+  EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
+  EXPECT_EQ(DsCnt.cur().getScore(), 1u);
+  EXPECT_TRUE(DsCnt.next());
+  EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
+  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(DsCnt.next());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+  EXPECT_TRUE(StoreCnt.hasCount());
+  EXPECT_EQ(StoreCnt.getCount(), 2u);
+  EXPECT_FALSE(StoreCnt.empty());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 1u);
+  EXPECT_TRUE(StoreCnt.next());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(StoreCnt.next());
+}
+
+/// Check all events carry into the next incoming block.
+TEST_F(AMDGPUGFX12EventTrackingTest, SingleIncomingBlock) {
+  StringRef MIR = R"(
+name:            SingleIncomingBlock
+body:             |
+  bb.0:
+    successors: %bb.1
+
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    $vgpr4 = GLOBAL_LOAD_DWORD $vgpr4_vgpr5, 0, 0, implicit $exec
+    $vgpr0 = DS_READ_B32_gfx9 $vgpr1, 0, 0, implicit $exec
+    GLOBAL_STORE_DWORDX2 $vgpr1_vgpr2, $vgpr3_vgpr4, 0, 0, implicit $exec
+
+  bb.1:
+
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("SingleIncomingBlock");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+  MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  visitAll(Ctx, BB0);
+  auto &ET = visitAll(Ctx, BB1);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.hasCount());
+  EXPECT_EQ(LoadCnt.getCount(), 1u);
+  EXPECT_FALSE(LoadCnt.empty());
+  EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
+  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(LoadCnt.next());
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.hasCount());
+  EXPECT_EQ(DsCnt.getCount(), 1u);
+  EXPECT_FALSE(DsCnt.empty());
+  EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
+  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(DsCnt.next());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+  EXPECT_TRUE(StoreCnt.hasCount());
+  EXPECT_EQ(StoreCnt.getCount(), 2u);
+  EXPECT_FALSE(StoreCnt.empty());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 1u);
+  EXPECT_TRUE(StoreCnt.next());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(StoreCnt.next());
+}
+
+/// Check a simple merge where all events differ in each incoming block.
+TEST_F(AMDGPUGFX12EventTrackingTest, SimpleDisjointMerge) {
+  StringRef MIR = R"(
+name:            SimpleDisjointMerge
+body:             |
+  bb.0:
+    successors: %bb.2
+
+    $vgpr0 = DS_READ_B32_gfx9 $vgpr1, 0, 0, implicit $exec
+    $vgpr4 = GLOBAL_LOAD_DWORD $vgpr4_vgpr5, 0, 0, implicit $exec
+    S_BRANCH %bb.2
+
+  bb.1:
+    successors: %bb.2
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    GLOBAL_STORE_DWORDX2 $vgpr1_vgpr2, $vgpr3_vgpr4, 0, 0, implicit $exec
+    S_BRANCH %bb.2
+
+  bb.2:
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("SimpleDisjointMerge");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+  MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
+  MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  visitAll(Ctx, BB0);
+  visitAll(Ctx, BB1);
+  auto &ET = visitAll(Ctx, BB2);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.hasCount());
+  EXPECT_EQ(LoadCnt.getCount(), 1u);
+  EXPECT_FALSE(LoadCnt.empty());
+  EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
+  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(LoadCnt.next());
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.hasCount());
+  EXPECT_EQ(DsCnt.getCount(), 1u);
+  EXPECT_FALSE(DsCnt.empty());
+  EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
+  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(DsCnt.next());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+  EXPECT_TRUE(StoreCnt.hasCount());
+  EXPECT_EQ(StoreCnt.getCount(), 2u);
+  EXPECT_FALSE(StoreCnt.empty());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 1u);
+  EXPECT_TRUE(StoreCnt.next());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(StoreCnt.next());
+}
+
+/// Merge with divergent events in a counter.
+TEST_F(AMDGPUGFX12EventTrackingTest, DivergenceMerge) {
+  StringRef MIR = R"(
+name:            DivergenceMerge
+body:             |
+  bb.0:
+    successors: %bb.2
+
+    $vgpr0 = DS_READ_B32_gfx9 $vgpr1, 0, 0, implicit $exec
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    S_BRANCH %bb.2
+
+  bb.1:
+    successors: %bb.2
+    $vgpr4 = GLOBAL_LOAD_DWORD $vgpr4_vgpr5, 0, 0, implicit $exec
+    GLOBAL_STORE_DWORDX2 $vgpr1_vgpr2, $vgpr3_vgpr4, 0, 0, implicit $exec
+    S_BRANCH %bb.2
+
+  bb.2:
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("DivergenceMerge");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+  MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
+  MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  visitAll(Ctx, BB0);
+  visitAll(Ctx, BB1);
+  auto &ET = visitAll(Ctx, BB2);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.hasCount());
+  EXPECT_EQ(LoadCnt.getCount(), 1u);
+  EXPECT_FALSE(LoadCnt.empty());
+  EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
+  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(LoadCnt.next());
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.hasCount());
+  EXPECT_EQ(DsCnt.getCount(), 1u);
+  EXPECT_FALSE(DsCnt.empty());
+  EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
+  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(DsCnt.next());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+  EXPECT_TRUE(StoreCnt.hasCount());
+  // Count is 1 because we have 1 event max across all predecessors.
+  EXPECT_EQ(StoreCnt.getCount(), 1u);
+  EXPECT_FALSE(StoreCnt.empty());
+  // First successor has GLOBAL_STORE_DWORD at height 0
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_TRUE(StoreCnt.next());
+  // Second successor has GLOBAL_STORE_DWORDX2 at height 0 too
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(StoreCnt.next());
+}
+
+/// Basic diamond CFG, the store is carried all the way into bb3 and uniqued
+/// again so only 1 instance of the record is present in bb3.
+TEST_F(AMDGPUGFX12EventTrackingTest, BasicDiamond) {
+  StringRef MIR = R"(
+name:            BasicDiamond
+body:             |
+  bb.0:
+    successors: %bb.1, %bb.2
+
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    S_CBRANCH_SCC1 %bb.1, implicit $scc
+    S_BRANCH %bb.2
+
+  bb.1:
+    successors: %bb.3
+    $vgpr0 = DS_READ_B32_gfx9 $vgpr1, 0, 0, implicit $exec
+    S_BRANCH %bb.3
+
+  bb.2:
+    successors: %bb.3
+    $vgpr4 = GLOBAL_LOAD_DWORD $vgpr4_vgpr5, 0, 0, implicit $exec
+    S_BRANCH %bb.3
+
+  bb.3:
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("BasicDiamond");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+  MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
+  MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
+  MachineBasicBlock &BB3 = *MF.getBlockNumbered(3);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  visitAll(Ctx, BB0);
+  visitAll(Ctx, BB1);
+  visitAll(Ctx, BB2);
+  auto &ET = visitAll(Ctx, BB3);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.hasCount());
+  EXPECT_EQ(LoadCnt.getCount(), 1u);
+  EXPECT_FALSE(LoadCnt.empty());
+  EXPECT_EQ(LoadCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_LOAD_DWORD);
+  EXPECT_EQ(LoadCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(LoadCnt.next());
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.hasCount());
+  EXPECT_EQ(DsCnt.getCount(), 1u);
+  EXPECT_FALSE(DsCnt.empty());
+  EXPECT_EQ(DsCnt.cur().getMI()->getOpcode(), AMDGPU::DS_READ_B32_gfx9);
+  EXPECT_EQ(DsCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(DsCnt.next());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+  EXPECT_TRUE(StoreCnt.hasCount());
+  EXPECT_EQ(StoreCnt.getCount(), 1u);
+  EXPECT_FALSE(StoreCnt.empty());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(StoreCnt.next());
+}
+
+/// Basic diamond CFG, the store is carried all the way into bb3 and uniqued
+/// again so only 1 instance of the record is present in bb3.
+TEST_F(AMDGPUGFX12EventTrackingTest, BasicDiamondFlags) {
+  StringRef MIR = R"(
+name:            BasicDiamondFlags
+body:             |
+  bb.0:
+    successors: %bb.1, %bb.2
+
+    S_CBRANCH_SCC1 %bb.1, implicit $scc
+    S_BRANCH %bb.2
+
+  bb.1:
+    successors: %bb.3
+    S_BRANCH %bb.3
+
+  bb.2:
+    successors: %bb.3
+    S_BRANCH %bb.3
+
+  bb.3:
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("BasicDiamondFlags");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+  MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
+  MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
+  MachineBasicBlock &BB3 = *MF.getBlockNumbered(3);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  visitAll(Ctx, BB0);
+  visitAll(Ctx, BB1, /*AfterVisit=*/[&](EventTracker &ET) {
+    ET.markIndeterminate(STORE_CNT);
+  });
+  visitAll(Ctx, BB2, [&](EventTracker &ET) { ET.markOutOfOrder(LOAD_CNT); });
+  auto &ET = visitAll(Ctx, BB3);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.unused());
+  // We inherit the out of order flag is either predecessor has it.
+  EXPECT_TRUE(LoadCnt.ET.isOutOfOrder(LOAD_CNT));
+  EXPECT_FALSE(LoadCnt.ET.isIndeterminate(LOAD_CNT));
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.unused());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+  EXPECT_TRUE(StoreCnt.unused());
+  // We inherit the indeterminate flag is either predecessor has it.
+  EXPECT_TRUE(StoreCnt.ET.isIndeterminate(STORE_CNT));
+}
+
+/// Assymetrical diamond
+///   - bb0 has two stores
+///   - bb1 adds another store without any waits.
+///   - bb2 adds a store and waits on the two stores from bb0 afterwards
+///
+/// The timeline will be:
+///   - Stores from bb0 exist at 1/2
+///   - The added store from bb1 exist at score 0
+///   - The added store from bb2 exist at score 0
+TEST_F(AMDGPUGFX12EventTrackingTest, AssymetricalDiamond) {
+  StringRef MIR = R"(
+name:            AssymetricalDiamond
+body:             |
+  bb.0:
+    successors: %bb.1, %bb.2
+
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    GLOBAL_STORE_DWORDX2 $vgpr1_vgpr2, $vgpr3_vgpr4, 0, 0, implicit $exec
+    S_CBRANCH_SCC1 %bb.1, implicit $scc
+    S_BRANCH %bb.2
+
+  bb.1:
+    successors: %bb.3
+    GLOBAL_STORE_DWORD $vgpr1_vgpr2, $vgpr3, 0, 0, implicit $exec
+    S_BRANCH %bb.3
+
+  bb.2:
+    successors: %bb.3
+    GLOBAL_STORE_DWORD $vgpr0_vgpr1, $vgpr2, 0, 0, implicit $exec
+    S_BRANCH %bb.3
+
+  bb.3:
+    S_ENDPGM 0
+...
+)";
+  ASSERT_TRUE(parseMIR(MIR));
+  MachineFunction &MF = getMF("AssymetricalDiamond");
+  MachineBasicBlock &BB0 = *MF.getBlockNumbered(0);
+  MachineBasicBlock &BB1 = *MF.getBlockNumbered(1);
+  MachineBasicBlock &BB2 = *MF.getBlockNumbered(2);
+  MachineBasicBlock &BB3 = *MF.getBlockNumbered(3);
+
+  EventTrackingContext Ctx(MF, GFX12CounterInfos);
+
+  visitAll(Ctx, BB0);
+  visitAll(Ctx, BB1);
+  visitAll(Ctx, BB2,
+           /*AfterVisit=*/[&](EventTracker &ET) { ET.wait(STORE_CNT, 1); });
+  auto &ET = visitAll(Ctx, BB3);
+
+  auto LoadCnt = TrackerRecordsChecker(ET, AMDGPU::LOAD_CNT);
+  EXPECT_TRUE(LoadCnt.unused());
+
+  auto DsCnt = TrackerRecordsChecker(ET, AMDGPU::DS_CNT);
+  EXPECT_TRUE(DsCnt.unused());
+
+  auto ExpCnt = TrackerRecordsChecker(ET, AMDGPU::EXP_CNT);
+  EXPECT_TRUE(ExpCnt.unused());
+
+  auto StoreCnt = TrackerRecordsChecker(ET, AMDGPU::STORE_CNT);
+
+  EXPECT_TRUE(StoreCnt.hasCount());
+  EXPECT_EQ(StoreCnt.getCount(), 3u);
+  EXPECT_FALSE(StoreCnt.empty());
+
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB0);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 2u);
+  EXPECT_TRUE(StoreCnt.next());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORDX2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 1u);
+  EXPECT_TRUE(StoreCnt.next());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB2);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_TRUE(StoreCnt.next());
+  EXPECT_EQ(StoreCnt.cur().getMI()->getOpcode(), AMDGPU::GLOBAL_STORE_DWORD);
+  EXPECT_EQ(StoreCnt.cur().getMI()->getParent(), &BB1);
+  EXPECT_EQ(StoreCnt.cur().getScore(), 0u);
+  EXPECT_FALSE(StoreCnt.next());
+}
+
+} // namespace
diff --git a/llvm/unittests/Target/AMDGPU/CMakeLists.txt b/llvm/unittests/Target/AMDGPU/CMakeLists.txt
index 2bb9b6dedba13..882ceef9f0310 100644
--- a/llvm/unittests/Target/AMDGPU/CMakeLists.txt
+++ b/llvm/unittests/Target/AMDGPU/CMakeLists.txt
@@ -23,6 +23,7 @@ set(LLVM_LINK_COMPONENTS
   )
 
 add_llvm_target_unittest(AMDGPUTests
+  AMDGPUEventTrackingTest.cpp
   AMDGPUMCExprTest.cpp
   AMDGPUUnitTests.cpp
   CSETest.cpp



More information about the llvm-commits mailing list