[llvm] [AMDGPU][True16] Retain `hi16` subregisters through v2s copies (PR #211848)

Zach Goldthorpe via llvm-commits llvm-commits at lists.llvm.org
Mon Jul 27 06:54:19 PDT 2026


================
@@ -7894,18 +7898,26 @@ void SIInstrInfo::legalizeOperandsVALUt16(MachineInstr &MI, unsigned OpIdx,
   int16_t RCID = getOpRegClassID(get(Opcode).operands()[OpIdx]);
   const TargetRegisterClass *ExpectedRC = RI.getRegClass(RCID);
   if (RI.getMatchingSuperRegClass(CurrRC, ExpectedRC, AMDGPU::lo16)) {
-    Op.setSubReg(AMDGPU::lo16);
-  } else if (RI.getMatchingSuperRegClass(ExpectedRC, CurrRC, AMDGPU::lo16)) {
+    // Default to the lo16 only if the subregister is not specified.
+    if (Op.getSubReg() == AMDGPU::NoSubRegister)
+      Op.setSubReg(AMDGPU::lo16);
+    return;
+  }
+
+  const TargetRegisterClass *CurrSRC =
+      RI.getSubRegisterClass(CurrRC, Op.getSubReg());
+  if (RI.getMatchingSuperRegClass(ExpectedRC, CurrSRC, AMDGPU::lo16)) {
----------------
zGoldthorpe wrote:

This is branch is taken by the `salu16_usedby_[sv]alu32` tests; it checks if `CurrSRC` is a 16-bit subregister of `ExpectedRC`, which is to say that the operand is supposed to be 32-bit, but it is being fed something 16-bit.

I've added a couple more tests, though, which take this branch with `CurrSRC != CurrRC`; that is, we feed to a 32-bit operand a 16-bit subregister (`lo16` or `hi16`).

This part of the fix was necessary because I was otherwise seeing crashes e.g. in `CodeGen/AMDGPU/min.ll`.

https://github.com/llvm/llvm-project/pull/211848


More information about the llvm-commits mailing list