[llvm] [AMDGPU] Select 64-bit VGPR copy opcode from the operand register class in copyPhysReg. (PR #219562)
Valery Pykhtin via llvm-commits
llvm-commits at lists.llvm.org
Fri Aug 28 11:54:57 PDT 2026
https://github.com/vpykhtin created https://github.com/llvm/llvm-project/pull/219562
copyPhysReg lowers a 64-bit VGPR copy to a wide move (V_MOV_B64_e32 / V_PK_MOV_B32) when one is available. Gate that choice on whether the destination is a member of the move's HwMode-resolved destination operand class, queried via getOpRegClassID(), instead of the register class's isProperlyAlignedRC() property.
isProperlyAlignedRC() answers a different question than the one that matters here: it inspects a property of the destination's own register class rather than whether the copy instruction can encode that destination. It is a separate, hand-maintained proxy for a constraint the operand register class already expresses, so it duplicates that knowledge and can diverge from it. Querying the operand register class - the authoritative description of what the instruction can encode - makes it the single source of truth: the wide move is selected only when the destination is a member of it, otherwise the copy falls through to the element-wise V_MOV_B32 expansion. This drops copyPhysReg's dependency on isProperlyAlignedRC(), consistent with the RegClassByHwMode operand handling.
The old and new criteria select the same opcode for the copies current targets emit, so there is no test change.
>From b9f902a83ba84f7046fcb7b9c4a2f888a0ff03b8 Mon Sep 17 00:00:00 2001
From: Valery Pykhtin <valery.pykhtin at amd.com>
Date: Fri, 28 Aug 2026 13:43:21 +0000
Subject: [PATCH] [AMDGPU] Select 64-bit VGPR copy opcode from the operand
register class
copyPhysReg lowers a 64-bit VGPR copy to a wide move (V_MOV_B64_e32 /
V_PK_MOV_B32) when one is available. Gate that choice on whether the destination
is a member of the move's HwMode-resolved destination operand class, queried via
getOpRegClassID(), instead of the register class's isProperlyAlignedRC()
property.
isProperlyAlignedRC() answers a different question than the one that matters
here: it inspects a property of the destination's own register class rather than
whether the copy instruction can encode that destination. It is a separate,
hand-maintained proxy for a constraint the operand register class already
expresses, so it duplicates that knowledge and can diverge from it. Querying the
operand register class - the authoritative description of what the instruction
can encode - makes it the single source of truth: the wide move is selected only
when the destination is a member of it, otherwise the copy falls through to the
element-wise V_MOV_B32 expansion. This drops copyPhysReg's dependency on
isProperlyAlignedRC(), consistent with the RegClassByHwMode operand handling.
The old and new criteria select the same opcode for the copies current targets
emit, so there is no test change.
Co-Authored-By: Claude <noreply at anthropic.com>
---
llvm/lib/Target/AMDGPU/SIInstrInfo.cpp | 45 ++++++++++++++++++++------
1 file changed, 35 insertions(+), 10 deletions(-)
diff --git a/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp b/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
index bc419d813f171..5f466faf30848 100644
--- a/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
+++ b/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
@@ -1119,13 +1119,24 @@ void SIInstrInfo::copyPhysReg(MachineBasicBlock &MBB,
return;
}
+ // Returns true if \p Opc can encode \p Dst as its destination, i.e. \p Dst is
+ // a member of \p Opc's HwMode-resolved destination operand class. A wide VALU
+ // copy is used only when this holds; otherwise the copy is expanded
+ // element-wise with V_MOV_B32 below.
+ auto IsDstInCopyInstrClass = [&](unsigned Opc, MCRegister Dst) {
+ int16_t RCID = getOpRegClassID(get(Opc).operands()[0]);
+ return RCID >= 0 && RI.getRegClass(RCID)->contains(Dst);
+ };
+
if (RC == RI.getVGPR64Class() && (SrcRC == RC || RI.isSGPRClass(SrcRC))) {
- if (ST.hasVMovB64Inst()) {
+ if (ST.hasVMovB64Inst() &&
+ IsDstInCopyInstrClass(AMDGPU::V_MOV_B64_e32, DestReg)) {
BuildMI(MBB, MI, DL, get(AMDGPU::V_MOV_B64_e32), DestReg)
.addReg(SrcReg, getKillRegState(KillSrc));
return;
}
- if (ST.hasPkMovB32()) {
+ if (ST.hasPkMovB32() &&
+ IsDstInCopyInstrClass(AMDGPU::V_PK_MOV_B32, DestReg)) {
BuildMI(MBB, MI, DL, get(AMDGPU::V_PK_MOV_B32), DestReg)
.addImm(SISrcMods::OP_SEL_1)
.addReg(SrcReg)
@@ -1166,15 +1177,29 @@ void SIInstrInfo::copyPhysReg(MachineBasicBlock &MBB,
} else if (RI.hasVGPRs(RC) && RI.isAGPRClass(SrcRC)) {
Opcode = AMDGPU::V_ACCVGPR_READ_B32_e64;
} else if ((Size % 64 == 0) && RI.hasVGPRs(RC) &&
- (RI.isProperlyAlignedRC(*RC) &&
- (SrcRC == RC || RI.isSGPRClass(SrcRC)))) {
+ (SrcRC == RC || RI.isSGPRClass(SrcRC))) {
// TODO: In 96-bit case, could do a 64-bit mov and then a 32-bit mov.
- if (ST.hasVMovB64Inst()) {
- Opcode = AMDGPU::V_MOV_B64_e32;
- EltSize = 8;
- } else if (ST.hasPkMovB32()) {
- Opcode = AMDGPU::V_PK_MOV_B32;
- EltSize = 8;
+ unsigned NewOpcode = AMDGPU::INSTRUCTION_LIST_END;
+ if (ST.hasVMovB64Inst())
+ NewOpcode = AMDGPU::V_MOV_B64_e32;
+ else if (ST.hasPkMovB32())
+ NewOpcode = AMDGPU::V_PK_MOV_B32;
+
+ if (NewOpcode != AMDGPU::INSTRUCTION_LIST_END) {
+ // The copy is expanded into 64-bit component moves below, so the wide
+ // move is usable only if it can encode each component. Components of a
+ // contiguous tuple share the base register's alignment, so it is enough
+ // to test the low component (sub0_sub1); for a 64-bit destination that
+ // component is the whole register. Otherwise fall through to the
+ // element-wise V_MOV_B32 expansion below (a Size == 64 copy reaches here
+ // when the wide moves above rejected its destination).
+ MCRegister WideDst = Size == 64
+ ? DestReg.asMCReg()
+ : RI.getSubReg(DestReg, AMDGPU::sub0_sub1);
+ if (IsDstInCopyInstrClass(NewOpcode, WideDst)) {
+ Opcode = NewOpcode;
+ EltSize = 8;
+ }
}
}
More information about the llvm-commits
mailing list