[llvm] [PreISelIntrinsicLowering] Lower bounded memcpy/memmove to masked load/store (PR #212710)
via llvm-commits
llvm-commits at lists.llvm.org
Thu Jul 30 00:01:53 PDT 2026
================
----------------
Harishankar14 wrote:
@RKSimon I tried the -mcpu RUN lines but they don't exercise the pass, and I think the reason is worth flagging:
-mcpu=x86-64-v4 sets "target-cpu" on the functions but not "target-features". This pass queries TTI (isLegalMaskedLoad/getRegisterBitWidth), which reads target-features so with only target-cpu set, it sees no AVX512 and declines. I confirmed -mcpu=x86-64-v4 doesn't fire in opt, and a function with just "target-cpu"="x86-64-v4" as an attribute also declines. Interestingly, llc -mcpu=x86-64-v4 does enable AVX512 for codegen (a <64 x i8> add lowers to vpaddb %zmm), so the CPU features are materialized later at ISel but aren't visible to TTI at pre-isel-intrinsic-lowering time.
-mattr=+avx512bw,+avx512vl sets the features directly and works. Real clang -O2 -mcpu=... is unaffected since clang expands the CPU into an explicit target-features list.
https://github.com/llvm/llvm-project/pull/212710
More information about the llvm-commits
mailing list