[llvm] [AMDGPU][True16] Retain `hi16` subregisters through v2s copies (PR #211848)
Zach Goldthorpe via llvm-commits
llvm-commits at lists.llvm.org
Mon Jul 27 06:54:19 PDT 2026
================
@@ -7894,18 +7898,26 @@ void SIInstrInfo::legalizeOperandsVALUt16(MachineInstr &MI, unsigned OpIdx,
int16_t RCID = getOpRegClassID(get(Opcode).operands()[OpIdx]);
const TargetRegisterClass *ExpectedRC = RI.getRegClass(RCID);
if (RI.getMatchingSuperRegClass(CurrRC, ExpectedRC, AMDGPU::lo16)) {
- Op.setSubReg(AMDGPU::lo16);
- } else if (RI.getMatchingSuperRegClass(ExpectedRC, CurrRC, AMDGPU::lo16)) {
+ // Default to the lo16 only if the subregister is not specified.
+ if (Op.getSubReg() == AMDGPU::NoSubRegister)
+ Op.setSubReg(AMDGPU::lo16);
+ return;
+ }
+
+ const TargetRegisterClass *CurrSRC =
+ RI.getSubRegisterClass(CurrRC, Op.getSubReg());
+ if (RI.getMatchingSuperRegClass(ExpectedRC, CurrSRC, AMDGPU::lo16)) {
----------------
zGoldthorpe wrote:
This is branch is taken by the `salu16_usedby_[sv]alu32` tests; it checks if `CurrSRC` is a 16-bit subregister of `ExpectedRC`, which is to say that the operand is supposed to be 32-bit, but it is being fed something 16-bit.
I've added a couple more tests, though, which take this branch with `CurrSRC != CurrRC`; that is, we feed to a 32-bit operand a 16-bit subregister (`lo16` or `hi16`).
This part of the fix was necessary because I was otherwise seeing crashes e.g. in `CodeGen/AMDGPU/min.ll`.
https://github.com/llvm/llvm-project/pull/211848
More information about the llvm-commits
mailing list