[libc-commits] [libc] [libc][aarch64] Use inline_memcpy_aligned_access_64bit under -mstrict-align (PR #227428)
via libc-commits
libc-commits at lists.llvm.org
Tue Sep 29 14:55:20 PDT 2026
https://github.com/PiJoules updated https://github.com/llvm/llvm-project/pull/227428
>From 3ab35bd411f0a0bc11911b70708f15c3ceeb16c8 Mon Sep 17 00:00:00 2001
From: Leonard Chan <leonardchan at google.com>
Date: Tue, 29 Sep 2026 11:21:40 -0700
Subject: [PATCH] [libc][aarch64] Avoid scalar byte loop in memcpy under
-mstrict-align
In libc/src/string/memory_utils/aarch64/inline_memcpy.h, when compiling
for AArch64 with -mstrict-align (!defined(__ARM_FEATURE_UNALIGNED)),
inline_memcpy_aarch64 uses builtin::Memcpy<64>::loop_and_tail(dst, src, count).
Because dst and src are typed as Ptr/CPtr (alignment 1), Clang under
-mstrict-align cannot emit multi-byte loads/stores, decomposing the
64-byte block into 64 individual ldrb and 64 individual strb
instructions (128 scalar byte operations per 64-byte block).
In bare-metal, firmware, and early kernel boot environments (such as
Fuchsia's kernel.phys boot shim where the MMU is disabled and data
accesses hit uncached DRAM), this causes severe performance degradation,
taking ~10.0 seconds for a 51 MB copy and triggering hardware watchdog
resets.
This patch routes large copies under !defined(__ARM_FEATURE_UNALIGNED)
to inline_memcpy_aligned_access_64bit() from
src/string/memory_utils/generic/aligned_access.h. This aligns dst to
an 8-byte boundary and uses 64-bit word transfers (load64_aligned /
store64_aligned<uint64_t>), finishing trailing remainders byte-by-byte.
---
.../string/memory_utils/aarch64/inline_memcpy.h | 17 +++++++++++++++++
1 file changed, 17 insertions(+)
diff --git a/libc/src/string/memory_utils/aarch64/inline_memcpy.h b/libc/src/string/memory_utils/aarch64/inline_memcpy.h
index 0c9224010784f..bc7dead70d93b 100644
--- a/libc/src/string/memory_utils/aarch64/inline_memcpy.h
+++ b/libc/src/string/memory_utils/aarch64/inline_memcpy.h
@@ -10,6 +10,7 @@
#include "src/__support/macros/attributes.h" // LIBC_INLINE
#include "src/__support/macros/properties/cpu_features.h"
+#include "src/string/memory_utils/generic/aligned_access.h"
#include "src/string/memory_utils/op_builtin.h"
#include "src/string/memory_utils/utils.h"
@@ -62,9 +63,25 @@ inline_memcpy_aarch64(Ptr __restrict dst, CPtr __restrict src, size_t count) {
return builtin::Memcpy<32>::head_tail(dst, src, count);
if (count < 128)
return builtin::Memcpy<64>::head_tail(dst, src, count);
+#if !defined(__ARM_FEATURE_UNALIGNED)
+ // When unaligned access is disabled (-mstrict-align), align_to_next_boundary
+ // below only aligns `src`. Because `dst` has type `cpp::byte*`, LLVM cannot
+ // statically prove that `dst` is aligned to more than 1 byte. Under strict
+ // alignment, this forces the compiler to lower the bulk copy loop to single-byte
+ // load and store instructions to prevent unaligned access faults, causing severe
+ // performance degradation.
+ //
+ // Instead, use inline_memcpy_aligned_access_64bit which explicitly aligns `dst`
+ // to an 8-byte boundary so 64-bit word stores can be safely used.
+ return inline_memcpy_aligned_access_64bit(dst, src, count);
+#else
+ // When hardware unaligned access is supported, aligning only `src` maximizes
+ // streaming read throughput while hardware handles unaligned vector stores to
+ // `dst` at full speed.
builtin::Memcpy<16>::block(dst, src);
align_to_next_boundary<16, Arg::Src>(dst, src, count);
return builtin::Memcpy<64>::loop_and_tail(dst, src, count);
+#endif
}
} // namespace LIBC_NAMESPACE_DECL
More information about the libc-commits
mailing list