[llvm] [AMDGPU] merge 16bit mov pairs in post-RA peephole (PR #208625)

Jay Foad via llvm-commits llvm-commits at lists.llvm.org
Thu Aug 27 02:41:27 PDT 2026


jayfoad wrote:

> Yes the docs does not say explictly about denormals in v_pack_b32_f16 section which is a bit annoying, but it does states the DENORM_IN64 as supported mode bits like the other FP insts. Do you think this is good enough or maybe a more clear statement is required? If not I will try to follow up with this seperately

I had not seen the part that mentions DENORM_IN64. I still think the description needs to be improved. In the public RDNA4 doc (for example) I see:

> D0[31 : 16].f16 = S1.f16;
> D0[15 : 0].f16 = S0.f16

This gives no indication that the instruction canonicalizes these values. Generally only arithmetic operations would be expected to canonicalize; bit operations like copying and fneg would not.

https://github.com/llvm/llvm-project/pull/208625


More information about the llvm-commits mailing list