[llvm] [AMDGPU] Fix DPP combine when a mov feeds several REG_SEQUENCE lanes (PR #217902)
Matt Arsenault via llvm-commits
llvm-commits at lists.llvm.org
Sun Sep 13 09:33:59 PDT 2026
================
@@ -684,30 +686,51 @@ bool GCNDPPCombine::combineDPPMov(MachineInstr &MovMI) const {
assert((TII->get(OrigOp).getSize() != 4 || !AMDGPU::isTrue16Inst(OrigOp)) &&
"There should not be e32 True16 instructions pre-RA");
if (OrigOp == AMDGPU::REG_SEQUENCE) {
+ // A different reg is a forwarded lane of a nested REG_SEQUENCE.
+ if (Use->getReg() != DPPMovReg) {
+ LLVM_DEBUG(dbgs() << " failed: nested REG_SEQUENCE\n");
+ break;
+ }
+
Register FwdReg = OrigMI.getOperand(0).getReg();
- unsigned FwdSubReg = 0;
if (execMayBeModifiedBeforeAnyUse(*MRI, FwdReg, OrigMI)) {
LLVM_DEBUG(dbgs() << " failed: EXEC mask should remain the same"
" for all uses\n");
break;
}
- unsigned OpNo, E = OrigMI.getNumOperands();
- for (OpNo = 1; OpNo < E; OpNo += 2) {
- if (OrigMI.getOperand(OpNo).getReg() == DPPMovReg) {
- FwdSubReg = OrigMI.getOperand(OpNo + 1).getImm();
- break;
- }
+ // The DPP mov reg can appear in several operands.
+ unsigned OpNo = OrigMI.getOperandNo(Use);
+ unsigned FwdSubReg = OrigMI.getOperand(OpNo + 1).getImm();
+ auto It = RegSeqWithOpNos.find(&OrigMI);
+
+ // Duplicate subreg indices pass the verifier. Lane already queued.
----------------
arsenm wrote:
I think I opened an issue about this. The verifier really ought to not allow input subreg indices that overlap. The lowering is ambiguous otherwise.
I would just assume every subreg input cannot overlap others
https://github.com/llvm/llvm-project/pull/217902
More information about the llvm-commits
mailing list