[llvm] [PreISelIntrinsicLowering] Lower bounded memcpy/memmove to masked load/store (PR #212710)
via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 8 04:21:36 PDT 2026
================
----------------
Harishankar14 wrote:
@RKSimon Tbh Not really,the integer form is what maps to a single kmovq into a mask register, so it's the cheapest way to materialise this on the targets that currently pass isLegalMaskedLoad for <N x i8> , all of which are little-endian.
Kind of confused which one to take it because the portable alternative is `lvm.get.active.lane.mask`, which is lane-indexed and has no endianness dependence, but I haven't measured what it costs on x86 versus the two-instruction integer path. Would you prefer I measure that and switch if it's close, or keep the integer form with an assert documenting the little-endian assumption? I think even asserting would be bad !!
https://github.com/llvm/llvm-project/pull/212710
More information about the llvm-commits
mailing list