[llvm] [AMDGPU] Fix expensive checks in fmaak/fmamk f16 folding (PR #176238)
Joe Nash via llvm-commits
llvm-commits at lists.llvm.org
Mon Jan 19 07:47:30 PST 2026
================
@@ -205,7 +205,9 @@ define half @test_fmaak(half %x, half %y, half %z) {
; GFX11-SDAG-TRUE16-LABEL: test_fmaak:
; GFX11-SDAG-TRUE16: ; %bb.0:
; GFX11-SDAG-TRUE16-NEXT: s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
-; GFX11-SDAG-TRUE16-NEXT: v_fmaak_f16 v0.l, v0.l, v1.l, 0x4200
+; GFX11-SDAG-TRUE16-NEXT: v_mov_b16_e32 v0.h, v1.l
----------------
Sisyph wrote:
The net effect of this series of patches with fmaak/make VGPR16_Lo128 allocatable seems to be adding more v_mov_b16, and in some cases swapping v_fmac_f16_e32 to v_fmaak_f16. Can you do anything to fix the regressions? There may be something in register coalescer that would be a nice fix, but a simple one might be revert fmac/fma_e64 pre-ra shrinking?
https://github.com/llvm/llvm-project/pull/176238
More information about the llvm-commits
mailing list