[llvm] [BOLT][hugify] Support "--hugify-all-text" to enable THP for all code segments (PR #220258)
Jinjie Huang via llvm-commits
llvm-commits at lists.llvm.org
Wed Sep 9 04:42:20 PDT 2026
================
@@ -391,6 +391,107 @@ int __prctl(int option, unsigned long arg2, unsigned long arg3,
return ret;
}
+#if defined(BOLT_RT_HUGIFY)
+static void syncInstructionCache(uint8_t *From, uint8_t *To) {
+ uint64_t CacheType;
+ __asm__ __volatile__("mrs %0, ctr_el0" : "=r"(CacheType));
+
+ const uint64_t DCacheLineSize = 4 << ((CacheType >> 16) & 15);
+ const uint64_t ICacheLineSize = 4 << (CacheType & 15);
+ uint64_t Address = reinterpret_cast<uint64_t>(From) & ~(DCacheLineSize - 1);
+ const uint64_t End = reinterpret_cast<uint64_t>(To);
+
+ for (; Address < End; Address += DCacheLineSize)
+ __asm__ __volatile__("dc cvau, %0" : : "r"(Address) : "memory");
+ __asm__ __volatile__("dsb ish" : : : "memory");
+
+ Address = reinterpret_cast<uint64_t>(From) & ~(ICacheLineSize - 1);
+ for (; Address < End; Address += ICacheLineSize)
+ __asm__ __volatile__("ic ivau, %0" : : "r"(Address) : "memory");
+ __asm__ __volatile__("dsb ish\nisb" : : : "memory");
+}
+
+// This stub is copied to an isolated executable mapping before use. Its
+// operation is equivalent to:
+// if (mmap(Target, Size, PROT_READ | PROT_WRITE,
+// MAP_FIXED | MAP_ANONYMOUS, -1, 0) == MAP_FAILED)
+// return false;
+// madvise(Target, Size, MADV_HUGEPAGE);
+// memcpy(Target, Copy, Size);
----------------
Jinjie-Huang wrote:
Thanks! Yes, your understanding is correct. `--hugify-all-text` applies madvise huge pages to all executable segments, including all DSOs. Since we currently cannot retrieve hot/cold information for DSOs, we apply this consistent behavior to the main binary as well (although, for the main binary itself, it is technically possible to enable huge pages only for the hot regions).
Regarding the memory overhead, I have added more details in the PR description. The RSS increase is generally under 1% (which is quite small compared to the massive total memory footprint of our services). As for the startup time, based on my experiment with a binary featuring a 1.5GB .text segment, the time it takes to execute up to the main function increased from 2.054s to 3.667s.
Given these trade-offs, this feature seems more suited for long-running server-side workloads rather than mobile or client-side applications, which is exactly why it is an opt-in feature.
https://github.com/llvm/llvm-project/pull/220258
More information about the llvm-commits
mailing list