[llvm] [AMDGPU][CodeGen] Post-RA VGPR MSB group optimization pass (PR #222666)

via llvm-commits llvm-commits at lists.llvm.org
Thu Sep 10 07:02:57 PDT 2026


llvmorg-github-actions[bot] wrote:


<!--LLVM PR SUMMARY COMMENT-->

@llvm/pr-subscribers-backend-amdgpu

Author: Lucas Ramirez (lucas-rami)

<details>
<summary>Changes</summary>

This introduces a new pass meant to run between register allocation and virtual-to-physical register rewriting. On subtargets where the MSBs of some VGPRs come from the processor's MODE, the pass tries to modify the existing virtual-to-physical register mappings to reduce the number of MODE-setting instructions that will need to be inserted in the program to honor MSB group differences between physical VGPRs. The pass cannot cause extra spilling to occur.

The pass is showing promising results on a corpus of kernels I have been experimenting on, yielding an average reduction of 7% (geomean) in the number of `S_SET_VGPR_MSB` instructions in the final assembly on gfx1250.

The pass is off by default and enabled with `--amdgpu-enable-vgpr-encoding-optimization`.

I am working on creating unit tests for the pass, and will add them to this PR in the coming days.

In the future, the intent is for this pass to also try to minimize VGPR bank conflicts on subtargets where it is relevant.

---

Patch is 77.28 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/222666.diff


6 Files Affected:

- (modified) llvm/lib/Target/AMDGPU/AMDGPU.h (+3) 
- (added) llvm/lib/Target/AMDGPU/AMDGPUOptimizeVGPREncoding.cpp (+1835) 
- (added) llvm/lib/Target/AMDGPU/AMDGPUOptimizeVGPREncoding.h (+23) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPUPassRegistry.def (+1) 
- (modified) llvm/lib/Target/AMDGPU/AMDGPUTargetMachine.cpp (+13) 
- (modified) llvm/lib/Target/AMDGPU/CMakeLists.txt (+1) 


``````````diff
diff --git a/llvm/lib/Target/AMDGPU/AMDGPU.h b/llvm/lib/Target/AMDGPU/AMDGPU.h
index 9fe4123a27bbd..0466f472d92b8 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPU.h
+++ b/llvm/lib/Target/AMDGPU/AMDGPU.h
@@ -547,6 +547,9 @@ extern char &AMDGPUInsertDelayAluID;
 void initializeAMDGPULowerVGPREncodingLegacyPass(PassRegistry &);
 extern char &AMDGPULowerVGPREncodingLegacyID;
 
+void initializeAMDGPUOptimizeVGPREncodingLegacyPass(PassRegistry &);
+extern char &SIAMDGPUOptimizeVGPREncodingLegacyID;
+
 void initializeSIInsertHardClausesLegacyPass(PassRegistry &);
 extern char &SIInsertHardClausesID;
 
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUOptimizeVGPREncoding.cpp b/llvm/lib/Target/AMDGPU/AMDGPUOptimizeVGPREncoding.cpp
new file mode 100644
index 0000000000000..f19911ce6b797
--- /dev/null
+++ b/llvm/lib/Target/AMDGPU/AMDGPUOptimizeVGPREncoding.cpp
@@ -0,0 +1,1835 @@
+//===-- AMDGPUOptimizeVGPREncoding.cpp --------------------------*- C++- *-===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+//
+/// \file
+/// This pass is meant to run between register allocation and
+/// virtual-to-physical register rewriting. On subtargets where the MSBs of some
+/// VGPRs come from the processor's MODE, the pass tries to modify the existing
+/// virtual-to-physical register mappings to reduce the number of MODE-setting
+/// instructions that will need to be inserted in the program to honor MSB group
+/// differences between physical VGPRs. The pass cannot cause extra spilling to
+/// occur.
+///
+/// In the future, the intent is for this pass to also try to minimize VGPR bank
+/// conflicts on subtarget where it is relevant.
+//
+//===----------------------------------------------------------------------===//
+
+#include "AMDGPUOptimizeVGPREncoding.h"
+#include "AMDGPU.h"
+#include "GCNSubtarget.h"
+#include "SIInstrInfo.h"
+#include "SIRegisterInfo.h"
+#include "Utils/AMDGPUBaseInfo.h"
+#include "llvm/ADT/ArrayRef.h"
+#include "llvm/ADT/STLExtras.h"
+#include "llvm/ADT/Sequence.h"
+#include "llvm/ADT/SmallBitVector.h"
+#include "llvm/CodeGen/LiveIntervals.h"
+#include "llvm/CodeGen/LiveRegMatrix.h"
+#include "llvm/CodeGen/MachineBasicBlock.h"
+#include "llvm/CodeGen/MachineFunctionPass.h"
+#include "llvm/CodeGen/MachineRegisterInfo.h"
+#include "llvm/CodeGen/RegisterClassInfo.h"
+#include "llvm/CodeGen/VirtRegMap.h"
+#include "llvm/InitializePasses.h"
+#include "llvm/MC/MCRegister.h"
+#include "llvm/Support/Debug.h"
+#include "llvm/Support/ErrorHandling.h"
+#include "llvm/Support/raw_ostream.h"
+#include <cstring>
+#include <functional>
+#include <string>
+
+using namespace llvm;
+
+#define DEBUG_TYPE "amdgpu-optimize-vgpr-encoding"
+
+namespace {
+
+/// MSB group, identified by an unsigned ID in [0, NumMSBGroups).
+using MSBGroup = unsigned;
+static constexpr unsigned NumMSBGroups = 4;
+static constexpr unsigned DefaultGroup = 0;
+
+/// Operand type where the MSB group is relevant, identified by an unsigned ID
+/// in [0, NumOprdTypes).
+using OprdType = unsigned;
+static constexpr unsigned NumOprdTypes = 4;
+
+/// An instruction that has at least one VGPR operand whose MSBs are provided by
+/// MODE. Instructions are part of a list and refer to other "neighbor"
+/// instructions through their respective index in this list. Two instructions
+/// are neighbors if they have at least one VGPR operand in the same operand
+/// type and no other instruction with such an operand in between them. Two
+/// neighbor instructions whose respective VGPR in a shared operand type differ
+/// in MSB group require at least one MODE-setting instruction to be placed
+/// somewhere in between them.
+struct ModeInstr {
+  /// Sentinel value in previous/next index arrays to indicate the non-existence
+  /// of a previous/next instruction.
+  static constexpr unsigned NoIdx = ~0U;
+
+  /// For each operand type, virtual of physical VGPR operand used by the
+  /// instruction. A null register indicates the instruction has no VGPR operand
+  /// of that type. For VOPD instructions, this holds VGPRs for the first of the
+  /// two instruction which define a VGPR of each operand type.
+  std::array<Register, NumOprdTypes> Oprds;
+  /// For each operand type, indices of previous/next neighbor instructions with
+  /// defined operands in the instruction list this instruction is a part of.
+  /// \ref NoIdx indicates that there is no such previous or next instruction.
+  std::array<unsigned, NumOprdTypes> Prev, Next;
+
+  ModeInstr() {
+    Oprds.fill(Register());
+    Prev.fill(NoIdx);
+    Next.fill(NoIdx);
+  }
+};
+
+/// Summarizes MSB bits usage over a machine basic block.
+class MBBModeUsage {
+public:
+  /// The machine basic block.
+  const MachineBasicBlock &MBB;
+
+  /// Iterates over \p MBB's instructions to find those for which MSB bits
+  /// provided by MODE are relevant. Indices of virtual registers used at least
+  /// once in an operand reading MSB bits are set in \p OptVirtRegs. Indices of
+  /// those that cannot change MSB group throughout optimization are set in \p
+  /// PinnedVirtRegs.
+  MBBModeUsage(const MachineBasicBlock &MBB, BitVector &OptVirtRegs,
+               BitVector &PinnedVirtRegs);
+
+  /// Returns the list of MODE-using instructions in the block.
+  ArrayRef<ModeInstr> getInstructions() const { return Instructions; }
+
+  /// Returns the MODE-using instruction in the block at index \p Idx.
+  const ModeInstr &getInstruction(unsigned Idx) const {
+    assert(Idx < Instructions.size() && "out of bounds");
+    return Instructions[Idx];
+  }
+
+  /// Returns the index of the first instruction with a MODE-reading operand of
+  /// type \p Oprd in the block, or \ref ModeInstr::NoIdx if none exists.
+  unsigned getFirstOprd(OprdType Oprd) const {
+    return Instructions.empty() ? ModeInstr::NoIdx : getFirstInstrFrom(0, Oprd);
+  }
+
+  /// Returns the index of the last instruction with a MODE-reading operand of
+  /// type \p Oprd in the block, or \ref ModeInstr::NoIdx if none exists.
+  unsigned getLastOprd(OprdType Oprd) const {
+    return Instructions.empty()
+               ? ModeInstr::NoIdx
+               : getLastInstrUntil(Instructions.size() - 1, Oprd);
+  }
+
+  /// Returns the index of the first instruction from \p InstrIdx (included)
+  /// with a MODE-reading operand of type \p Oprd in the block, or \ref
+  /// ModeInstr::NoIdx if none exists.
+  unsigned getFirstInstrFrom(unsigned InstrIdx, OprdType Oprd) const {
+    return getInstrIdxImpl<false, false>(InstrIdx, Oprd);
+  }
+
+  /// Returns the index of the first instruction after \p InstrIdx (excluded)
+  /// with a MODE-reading operand of type \p Oprd in the block, or \ref
+  /// ModeInstr::NoIdx if none exists.
+  unsigned getFirstInstrAfter(unsigned InstrIdx, OprdType Oprd) const {
+    return getInstrIdxImpl<false, true>(InstrIdx, Oprd);
+  }
+
+  /// Returns the index of the last instruction until \p InstrIdx (included)
+  /// with a MODE-reading operand of type \p Oprd in the block, or \ref
+  /// ModeInstr::NoIdx if none exists.
+  unsigned getLastInstrUntil(unsigned InstrIdx, OprdType Oprd) const {
+    return getInstrIdxImpl<true, false>(InstrIdx, Oprd);
+  }
+
+  /// Returns the index of the last instruction before \p InstrIdx (excluded)
+  /// with a MODE-reading operand of type \p Oprd in the block, or \ref
+  /// ModeInstr::NoIdx if none exists.
+  unsigned getLastInstrBefore(unsigned InstrIdx, OprdType Oprd) const {
+    return getInstrIdxImpl<true, true>(InstrIdx, Oprd);
+  }
+
+#ifndef NDEBUG
+  Printable print(const VirtRegMap &VRM) const;
+#endif
+
+private:
+  /// List of all instructions in the machine basic block which have at least
+  /// one operand for which MSB bits provided by MODE are relevant, in program
+  /// order.
+  SmallVector<ModeInstr> Instructions;
+
+  template <bool UsePrev, bool SkipCurrent>
+  unsigned getInstrIdxImpl(unsigned InstrIdx, OprdType Oprd) const;
+};
+
+/// A virtual register whose assigned physical register's MSB group can be
+/// optimized for. An optimizable register maintains a per-group score encoding
+/// its current level of "neighboring-ness" to other registers in each MSB
+/// group, with higher scores indicating a higher number of "neighbor" registers
+/// currently in the corresponding MSB group. Two registers (virtual or
+/// physical) are initially considered neighbors when they are used as operands
+/// of the same type in two neighboring instructions (c.f. \ref ModeInstr). A
+/// register's neighborhood---and thefore its score---changes throughout the
+/// pass's lifetime to reflect the simulataed placement of MODE-setting
+/// instructions.
+///
+/// An optimizable register is said to be "pinned" when the MSB group of its
+/// assigned physical register cannot change. Once a register is pinned it never
+/// becomes unpinned. Score contributions from pinned and unpinned neighbors are
+/// kept separate to enable identification of unoptimizable MSB group conflicts.
+class OptReg {
+public:
+  using WeightedNeighbors = SmallDenseMap<OptReg *, unsigned, 4>;
+
+  /// Abstract coordinates for an occurence of this register.
+  struct Coordinates {
+    /// This index of the MBB.
+    unsigned MBBIndex;
+    /// The index of the MODE-using instruction.
+    unsigned InstrIdx;
+    /// The operand type.
+    OprdType Oprd;
+  };
+
+  /// Creates a neighbor-less optimizable register for register \p VirtReg.
+  /// Neighboring relations with other optimizable registers are
+  /// created/destroyed through class methods.
+  OptReg(Register VirtReg, const VirtRegMap &VRM);
+
+  /// Returns the total number of occurences of pinned neighbor registers in \p
+  /// Group.
+  unsigned getGroupPinnedScore(MSBGroup Group) const {
+    return PinnedScore[Group];
+  }
+
+  /// Returns the total number of occurences of neighbor registers in \p Group.
+  unsigned getGroupScore(MSBGroup Group) const {
+    return PinnedScore[Group] + Score[Group];
+  }
+
+  /// Returns the total number of occurences of neighbor registers in this
+  /// register's current MSB group.
+  unsigned getCurrentGroupScore() const { return getGroupScore(MSB); }
+
+  /// Returns a bitvector the size of the number of MSB groups whose set bits
+  /// indicate the MSB groups in which this register currently has at least one
+  /// pinned neighbor.
+  SmallBitVector getPinGroups() const;
+
+  /// Returns the register's current neighbors.
+  const WeightedNeighbors &getNeighbors() const { return Neighbors; }
+
+  /// Returns the list of coordinates corresponding to this register's
+  /// occurences.
+  ArrayRef<Coordinates> getOccurences() const { return Occurences; }
+
+  /// Returns the underlying virtual register.
+  Register getVirt() const { return VirtReg; }
+
+  /// Returns the underlying virtual register's index.
+  unsigned getVirtIndex() const { return VirtReg.virtRegIndex(); }
+
+  /// Returns the MSB group of this register's currently assigned physical
+  /// register.
+  MSBGroup getMSB() const { return MSB; }
+
+  /// Returns whether the register is pinned.
+  bool isPinned() const { return IsPinned; }
+
+  // addOccurence and record* methods used by OptimizableRegs to initialize the
+  // occurences and neighborhood of all optimizable registers at the beginning.
+
+  /// Adds an occurence of this register in operand type \p Oprd of instruction
+  /// \p InstrIdx of MBB \p MBBIdx.
+  void addOccurence(unsigned MBBIndex, unsigned InstrIdx, OprdType Oprd) {
+    Occurences.push_back({MBBIndex, InstrIdx, Oprd});
+  }
+
+  /// Records an occurence of \p NeighborReg as a neighbor.
+  void recordNeighborOccurence(OptReg &NeighborReg);
+
+  /// Records an occurence of physical register \p PhysReg as a neighbor.
+  void recordPhysNeighborOccurence(Register PhysReg, const VirtRegMap &VRM);
+
+  /// Records an occurence of this register at a block boundary. This adds a
+  /// "pinned occurence" of the default MSB group in which all MBBs start and
+  /// end.
+  void recordBlockBoundaryPin() { ++PinnedScore[DefaultGroup]; }
+
+  // pinMSBGroup and remove* methods used by ModeSetOptimizer to progressively
+  // simplify/destroy the neighborhood of all optimizable registers as it
+  // simulates placement of MODE-setting instructions. remove* methods mirror
+  // record* methods 1-to-1.
+
+  /// Pins this register's to the MSB group of its currently assigned physical
+  /// register.
+  void pinMSBGroup();
+
+  /// Removes an occurence of \p NeighborReg as a neighbor.
+  void removeNeighborOccurence(OptReg &NeighborReg);
+
+  /// Removes an occurence of physical register \p PhysReg as a neighbor.
+  void removePhysNeighborOccurence(Register PhysReg, const VirtRegMap &VRM);
+
+  /// Removes an occurence of this register at a block boundary. This removes a
+  /// "pinned occurence" of the default MSB group in which all MBBs start and
+  /// end.
+  void removeBlockBoundaryPin() {
+    assert(PinnedScore[DefaultGroup] > 0 && "underflow");
+    --PinnedScore[DefaultGroup];
+  }
+
+  /// Notifies the optimizable register that its assigned physcial register has
+  /// changed and that its new assignment belongs to MSB group \p NewGroup. It
+  /// is illegal to change the MSB group of a pinned register.
+  void notifyPhysAssignmentChanged(MSBGroup NewGroup);
+
+#ifndef NDEBUG
+  Printable print(const VirtRegMap &VRM) const;
+#endif
+
+private:
+  /// The virtual register.
+  Register VirtReg;
+  /// MSB group of the virtual register's current physcial register assignment.
+  MSBGroup MSB;
+  /// Per-MSB group score, counting the number of occurences of neighbor
+  /// registers in each group, separated between occurences of unpinned
+  /// optimizable registers from the others (physical registers, pinned
+  /// optimizable registers, and block boundaries). Reflects the register's
+  /// current neighborhood.
+  std::array<unsigned, NumMSBGroups> Score, PinnedScore;
+  /// Maps neighboring optimizable registers to the number of times they occur
+  /// in the neighborhood of one of this register's occurences. Neighbors can be
+  /// added or removed at will after construction, impacting the score.
+  WeightedNeighbors Neighbors;
+  /// Occurences of this register in the function under consideration.
+  /// Occurences can be added after construction but cannot be removed.
+  SmallVector<Coordinates> Occurences;
+  /// Whether the register is pinned to \ref MSB.
+  bool IsPinned = false;
+};
+
+/// Manages all optimizable virtual registers for a function.
+class OptimizableRegs {
+public:
+  /// Creates an \ref OptReg for each virtual register whose index is set in \p
+  /// OptVirtRegs, immediately pinning those whose index is set in \p
+  /// PinnedVirtRegs. Then derives neighborhood of all optimizable registers
+  /// from MODE-using instructions in each block of \p ModeUsage.
+  OptimizableRegs(const BitVector &OptVirtRegs, const BitVector &PinnedVirtRegs,
+                  ArrayRef<MBBModeUsage> ModeUsage, const VirtRegMap &VRM);
+
+  OptReg *operator[](Register Reg) {
+    if (Reg.isPhysical())
+      return nullptr;
+    unsigned Idx = VirtRegToStorageIdx[Reg.virtRegIndex()];
+    return Idx == NoIdx ? nullptr : &Storage[Idx];
+  }
+  OptReg &operator[](unsigned VirtRegIdx) {
+    assert(OptVirtRegs.test(VirtRegIdx) && "invalid index");
+    return Storage[VirtRegToStorageIdx[VirtRegIdx]];
+  }
+
+  /// Returns a bitvector whose set bits indicate indices of virtual registers
+  /// which are optimizable i.e., for which (*this)[VirtReg] returns a valid
+  /// \ref OptReg.
+  const BitVector &getAllOptVirtRegs() const { return OptVirtRegs; }
+
+  /// Returns the number of virtual registers, as reported by the MRI.
+  unsigned getNumVirtRegs() const { return OptVirtRegs.size(); }
+
+  using iterator = SmallVector<OptReg>::iterator;
+  using const_iterator = SmallVector<OptReg>::const_iterator;
+  iterator begin() { return Storage.begin(); }
+  iterator end() { return Storage.end(); }
+  const_iterator begin() const { return Storage.begin(); }
+  const_iterator end() const { return Storage.end(); }
+
+private:
+  /// Sentinel value in \p VirtRegToStorageIdx to indicate the non-existence of
+  /// a corresponding \ref OptReg.
+  static constexpr unsigned NoIdx = ~0;
+
+  /// Set bits indicate indices of virtual registers which are optimizable.
+  BitVector OptVirtRegs;
+  /// Backing storage for optimizable registers.
+  SmallVector<OptReg, 0> Storage;
+  /// Works as a map from virtual register indices to the index of the
+  /// corresponding \ref OptReg in \ref Storage. A virtual register that is not
+  /// optimizable "maps to" \ref NoIdx.
+  SmallVector<unsigned, 0> VirtRegToStorageIdx;
+};
+
+/// Simulates placement of MODE-setting instructions as unoptimizable MSB groups
+/// conflicts are detected, driving optimization forward by progressively
+/// pruning register neighborhoods and pinning optimizable registers once they
+/// reach an "ideal" MSB group.
+///
+/// The detection and resolution of unoptimizable MSB group conflicts is this
+/// class's main purpose. An optimizable register with non-null score
+/// contributions from pinned neighbors in more that one MSB group will
+/// necessarily require MODE-setting instructions around its occurences that
+/// neighbor pinned registers in all but one of those MSB groups. This is a
+/// conflict in the sense that we would need the register to be in multiple MSB
+/// groups at the same time to not need MODE-setting instructions. It is
+/// unoptimizable by the pass because pinned neighbors are not allowed to change
+/// MSB group, so no amount of register re-assignment can resolve it. The
+/// objective is to detect those situations early so that no effort is made
+/// attempting to optimize MSB conflits at code locations where we are
+/// guaranteed to be unable to solve them. The class resolves such conflicts by
+/// simulating the placement of MODE-setting instructions around problematic
+/// registers, effectively "breaking" their relationships with some pinned
+/// neighbors until any remaining MSB group conflict becomes optimizable again,
+/// at the known cost of "placed" MODE-setting instructions.
+///
+/// Resolving conflicts strictly lowers the per-group score of optimizable
+/// registers that neighbor "placed" MODE-setting instructions. This ensures
+/// forward progress (scores are lower bounded at 0) and can uncover new
+/// optimization opportunities as register neighborhoods become smaller and some
+/// registers reach an "ideal" MSB group that they can be pinned to.
+///
+/// FIXME: The current approach to determine where we place MODE-setting
+/// instructions to resolve conflicts is correct however when there are multiple
+/// possible locations to choose from it does not attempt to analyze the
+/// expected benefit of each. Picking the best location in such cases should
+/// improve overall pass performance.
+class ModeSetOptimizer {
+public:
+  /// Initializes the optimizer with all optimizable registers in \p OptRegs and
+  /// all blocks in \p ModeUsage. Performs a first round of register pinning and
+  /// conflict resolution on all relevant registers.
+  ModeSetOptimizer(OptimizableRegs &OptRegs, ArrayRef<MBBModeUsage> ModeUsage,
+                   const VirtRegMap &VRM);
+
+  /// Iteratively resolves MSB group conflicts by simulating placement of
+  /// MODE-setting instructions and pins newly eligible registers until reaching
+  /// a fixed-point. By the end there are no unoptimizable MSB group conflicts
+  /// and all registers that would be eligible for pinning are pinned. Returns
+  /// whether any register changed state.
+  bool resolveConflictsAndPinRegs();
+
+  /// Notifies the optimizer that \p Reg changed MSB group. Sets bits in \p
+  /// ScoreChanged for all virtual register indices whose score was affected by
+  /// the move. Tracks which registers can become eligible for pinning or may
+  /// exhibit a conflict as a result of the change.
+  void regChangedGroup(OptReg &Reg, BitVector &ScoreChanged);
+
+private:
+  /// Result of BitVector::find* methods when no bit was found.
+  static constexpr int NoBit = -1;
+
+  /// Optimizable registers under consideration.
+  OptimizableRegs &OptRegs;
+...
[truncated]

``````````

</details>


https://github.com/llvm/llvm-project/pull/222666


More information about the llvm-commits mailing list