[llvm] [ExpandMemCmp] Only narrow load sizes for targets that opt in (PR #215186)

Sam Elliott via llvm-commits llvm-commits at lists.llvm.org
Mon Aug 10 21:18:57 PDT 2026


lenary wrote:

> BPF is the concrete example (see https://github.com/llvm/llvm-project/pull/209738): it reports LoadSizes = {8, 4, 2, 1}, does not allow misaligned scalar access, uses the
default NumLoadsPerBlock of 1, and has a very large MaxLoadsPerMemcmp. For an
align-1 memcmp the filter drops 8/4/2 and keeps only i8, so the compare is
expanded byte-by-byte into one block per byte. A 32-byte compare becomes 33
nested branches; a real BPF program then exceeds the in-kernel verifier's 1M
instruction limit and is rejected.

Why this fix rather than adjusting the BPF values for `NumLoadsPerBlock` (up to 8, i guess?), and maybe also `MaxLoadsPerMemcmp` (it's less clear what this should be scaled by, but maybe it should be 1/8 of what it was?)

Is BPF correctly reporting how it supports (or not) misaligned accesses?

Do we maybe need different filtering/costing for trying to understand how many loads a misaligned load/store will be split into, and scaling things with that?

https://github.com/llvm/llvm-project/pull/215186


More information about the llvm-commits mailing list