[llvm] AMDGPU/GlobalISel: Implement RegBankLegalizeRules for amdgcn_log, amdgcn_rcp, and amdgcn_sqrt (PR #195099)
Petar Avramovic via llvm-commits
llvm-commits at lists.llvm.org
Wed May 6 08:23:00 PDT 2026
petar-avramovic wrote:
## RegBankLegalize coverage check
New opcodes: `amdgcn_log`, `amdgcn_rcp` (added to existing `amdgcn_sqrt` rule block)
| Opcode | Rule | Test file | Function | Status |
|--------|------|-----------|----------|--------|
| amdgcn_log/rcp/sqrt | Div S16: Vgpr16, IntrId Vgpr16 | 💡 pseudo-scalar-transcendental.ll | 💡 v_amdgcn_log_f16_div / v_rcp_f16_div | ❌ MISSING |
| amdgcn_log/rcp/sqrt | Uni S16 (hasPST): Sgpr16, IntrId Sgpr16 | pseudo-scalar-transcendental.ll | v_s_amdgcn_log_f16 | ✅ |
| amdgcn_log/rcp/sqrt | Uni S16 (!hasPST): UniInVgprS16, IntrId Vgpr16 | 💡 pseudo-scalar-transcendental.ll | 💡 (needs non-PST target) | ❌ MISSING |
| amdgcn_log/rcp/sqrt | Div S32: Vgpr32, IntrId Vgpr32 | 💡 pseudo-scalar-transcendental.ll | 💡 v_amdgcn_log_f32_div / v_rcp_f32_div | ❌ MISSING |
| amdgcn_log/rcp/sqrt | Uni S32 (hasPST): Sgpr32, IntrId Sgpr32 | pseudo-scalar-transcendental.ll | v_s_log_f32 | ✅ |
| amdgcn_log/rcp/sqrt | Uni S32 (!hasPST): UniInVgprS32, IntrId Vgpr32 | 💡 pseudo-scalar-transcendental.ll | 💡 (needs non-PST target) | ❌ MISSING |
| amdgcn_log/rcp/sqrt | Div S64: Vgpr64, IntrId Vgpr64 | 💡 pseudo-scalar-transcendental.ll | 💡 v_amdgcn_log_f64_div / v_rcp_f64_div | ❌ MISSING |
| amdgcn_log/rcp/sqrt | Uni S64: UniInVgprS64, IntrId Vgpr64 | 💡 pseudo-scalar-transcendental.ll | 💡 v_s_amdgcn_log_f64 / v_s_rcp_f64 | ❌ MISSING |
### Suggested fix
Add to `llvm/test/CodeGen/AMDGPU/pseudo-scalar-transcendental.ll`:
**Divergent f16 (Div S16):**
```llvm
define void @v_amdgcn_log_f16_div(half %src, ptr addrspace(1) %out) {
%result = call half @llvm.amdgcn.log.f16(half %src)
store half %result, ptr addrspace(1) %out
ret void
}
define void @v_rcp_f16_div(half %src, ptr addrspace(1) %out) {
%result = call half @llvm.amdgcn.rcp.f16(half %src)
store half %result, ptr addrspace(1) %out
ret void
}
```
**Divergent f32 (Div S32):**
```llvm
define void @v_amdgcn_log_f32_div(float %src, ptr addrspace(1) %out) {
%result = call float @llvm.amdgcn.log.f32(float %src)
store float %result, ptr addrspace(1) %out
ret void
}
define void @v_rcp_f32_div(float %src, ptr addrspace(1) %out) {
%result = call float @llvm.amdgcn.rcp.f32(float %src)
store float %result, ptr addrspace(1) %out
ret void
}
```
**Divergent f64 (Div S64):**
```llvm
define void @v_amdgcn_log_f64_div(double %src, ptr addrspace(1) %out) {
%result = call double @llvm.amdgcn.log.f64(double %src)
store double %result, ptr addrspace(1) %out
ret void
}
define void @v_rcp_f64_div(double %src, ptr addrspace(1) %out) {
%result = call double @llvm.amdgcn.rcp.f64(double %src)
store double %result, ptr addrspace(1) %out
ret void
}
```
**Uniform f64 (Uni S64):**
```llvm
define amdgpu_cs double @v_s_amdgcn_log_f64(double inreg %src) {
%result = call double @llvm.amdgcn.log.f64(double %src)
ret double %result
}
define amdgpu_cs double @v_s_rcp_f64(double inreg %src) {
%result = call double @llvm.amdgcn.rcp.f64(double %src)
ret double %result
}
```
**Note:** The `!hasPST` rules (Uni S16 and Uni S32 without PST) require a non-PST target with `-new-reg-bank-select`. Since `pseudo-scalar-transcendental.ll` only targets `gfx1200` (which has PST), testing these would require adding a RUN line with a non-PST subtarget.
Run `update_llc_test_checks.py` to generate the check lines.
> *Automated check — not a human review. Generated by Claude (claude.ai).*
https://github.com/llvm/llvm-project/pull/195099
More information about the llvm-commits
mailing list