[llvm] [AMDGPU] merge 16bit mov pairs in post-RA peephole (PR #208625)
Jay Foad via llvm-commits
llvm-commits at lists.llvm.org
Thu Aug 27 02:41:27 PDT 2026
jayfoad wrote:
> Yes the docs does not say explictly about denormals in v_pack_b32_f16 section which is a bit annoying, but it does states the DENORM_IN64 as supported mode bits like the other FP insts. Do you think this is good enough or maybe a more clear statement is required? If not I will try to follow up with this seperately
I had not seen the part that mentions DENORM_IN64. I still think the description needs to be improved. In the public RDNA4 doc (for example) I see:
> D0[31 : 16].f16 = S1.f16;
> D0[15 : 0].f16 = S0.f16
This gives no indication that the instruction canonicalizes these values. Generally only arithmetic operations would be expected to canonicalize; bit operations like copying and fneg would not.
https://github.com/llvm/llvm-project/pull/208625
More information about the llvm-commits
mailing list