[llvm] [SampleProfile] Filter string NameTable lookups by module (PR #227900)

via llvm-commits llvm-commits at lists.llvm.org
Thu Oct 1 13:57:01 PDT 2026


llvmorg-github-actions[bot] wrote:


<!--LLVM PR SUMMARY COMMENT-->

@llvm/pr-subscribers-pgo

Author: Kunal Pathak (kunalspathak)

<details>
<summary>Changes</summary>

## Problem

StringSampleProfileNameTable::contains() lazily constructs a `DenseSet<StringRef>` containing every entry in the profile NameTable. This is expensive for a whole-program profile when the caller can only query names from one much smaller module.

The extensible-binary reader already collects canonical names from the current module in FuncsToUse for on-demand profile loading.

## Change

Reuse FuncsToUse as a filter and lazily construct a set containing only the intersection of the string NameTable and current-module names.

The filtered path is selected only when:

- The queried name belongs to FuncsToUse.
- The NameTable has at least 16 times as many entries as FuncsToUse.

Other queries use the existing unrestricted lookup. MD5 profiles, including Eytzinger profiles, are unchanged.

## Memory reduction

This measures the additional hash-table storage used for NameTable membership. The existing FuncsToUse set is excluded because both versions already construct it.

| Profile names | Module names | Full-profile index | Filtered index | Memory avoided | Reduction |
|---------------|--------------|--------------------|----------------|----------------|-----------|
| 100,000       | 1,000        | 4.03 MiB           | 32.3 KiB       | 4.00 MiB       | 99.2%     |
| 1,000,000     | 10,000       | 32.25 MiB          | 258 KiB        | 32.00 MiB      | 99.2%     |
| 5,000,000     | 50,000       | 129.00 MiB         | 2.02 MiB       | 126.98 MiB     | 98.4%     |
| 10,000,000    | 100,000      | 258.00 MiB         | 4.03 MiB       | 253.97 MiB     | 98.4%     |

The allocation sizes change in steps because DenseSet rounds its bucket capacity to a power of two.

## CPU and lookup performance

The microbenchmark uses LLVM's `DenseSet`, deterministic 29-byte names, a 50% hit rate, and five runs per configuration on an AArch64 Neoverse-V2 system.

| Profile names | Module names | Baseline construction | Filtered construction | Improvement | Query improvement |
|---------------|--------------|-----------------------|-----------------------|-------------|-------------------|
| 100,000       | 1,000        | 15.88 ms              | 13.86 ms              | 12.7%       | 30.4%             |
| 1,000,000     | 10,000       | 201.11 ms             | 134.82 ms             | 33.0%       | 42.0%             |
| 5,000,000     | 50,000       | 1,342.04 ms           | 833.32 ms             | 37.9%       | 44.2%             |
| 10,000,000    | 100,000      | 2,932.97 ms           | 1,755.35 ms           | 40.2%       | 42.8%             |

For the five-million-name case:

| Counter      | Baseline    | Filtered  | Improvement |
|--------------|-------------|-----------|-------------|
| Task clock   | 1,472.98 ms | 872.20 ms | 40.8%       |
| CPU cycles   | 4.859 B     | 2.883 B   | 40.7%       |
| Instructions | 5.436 B     | 5.448 B   | -0.2%       |
| Cache misses | 19.10 M     | 11.44 M   | 40.1%       |
| Page faults  | 5,874       | 3,875     | 34.0%       |

The instruction count is essentially unchanged. Most of the speedup comes from the smaller writable working set and fewer cache misses.

## End-to-end opt results

This benchmark runs the actual sample-profile pass using a one-million-name extensible-binary profile and a 1,000-function module. Results are medians of seven warmed runs from a Debug build with assertions enabled.

| Metric       | Baseline    | Filtered    | Improvement   |
|--------------|-------------|-------------|---------------|
| Wall time    | 1.66 s      | 1.48 s      | 10.8%         |
| Task clock   | 1,650.77 ms | 1,487.15 ms | 9.9%          |
| CPU cycles   | 5.437 B     | 4.892 B     | 10.0%         |
| Instructions | 8.766 B     | 8.180 B     | 6.7%          |
| Cache misses | 8.294 M     | 5.802 M     | 30.0%         |
| Page faults  | 3,848       | 3,333       | 13.4%         |
| Peak RSS     | 207 MiB     | 189 MiB     | 18 MiB (8.7%) |

The baseline and filtered configurations emitted byte-identical LLVM IR. The input is synthetic and intended to isolate NameTable scaling; it does not represent a complete production ThinLTO build.

Assisted-By: gpt-5.6-sol

---
Full diff: https://github.com/llvm/llvm-project/pull/227900.diff


3 Files Affected:

- (modified) llvm/include/llvm/ProfileData/SampleProfReader.h (+7) 
- (modified) llvm/lib/ProfileData/SampleProfReader.cpp (+23) 
- (modified) llvm/unittests/ProfileData/SampleProfTest.cpp (+37) 


``````````diff
diff --git a/llvm/include/llvm/ProfileData/SampleProfReader.h b/llvm/include/llvm/ProfileData/SampleProfReader.h
index 2bb616bb94b1f..94d309a1d704b 100644
--- a/llvm/include/llvm/ProfileData/SampleProfReader.h
+++ b/llvm/include/llvm/ProfileData/SampleProfReader.h
@@ -1200,6 +1200,10 @@ class LLVM_ABI SampleProfileReaderExtBinaryBase
   /// The set containing the functions to use when compiling a module.
   DenseSet<StringRef> FuncsToUse;
 
+  /// Name table entries matching functions in the current module. This avoids
+  /// building a whole-profile string index for module-scoped queries.
+  mutable std::optional<DenseSet<StringRef>> ModuleNameTableEntries;
+
 public:
   SampleProfileReaderExtBinaryBase(std::unique_ptr<MemoryBuffer> B,
                                    LLVMContext &C, SampleProfileFormat Format)
@@ -1227,6 +1231,9 @@ class LLVM_ABI SampleProfileReaderExtBinaryBase
   /// the reader has been given a module.
   bool collectFuncsFromModule() override;
 
+  using SampleProfileReaderBinary::contains;
+  bool contains(StringRef Key) const override;
+
   std::unique_ptr<ProfileSymbolList> getProfileSymbolList() override {
     return std::move(ProfSymList);
   };
diff --git a/llvm/lib/ProfileData/SampleProfReader.cpp b/llvm/lib/ProfileData/SampleProfReader.cpp
index 86610d5b0445b..9a404821904e2 100644
--- a/llvm/lib/ProfileData/SampleProfReader.cpp
+++ b/llvm/lib/ProfileData/SampleProfReader.cpp
@@ -1111,11 +1111,34 @@ bool SampleProfileReaderExtBinaryBase::collectFuncsFromModule() {
   if (!M)
     return false;
   FuncsToUse.clear();
+  ModuleNameTableEntries.reset();
   for (auto &F : *M)
     FuncsToUse.insert(FunctionSamples::getCanonicalFnName(F));
   return true;
 }
 
+bool SampleProfileReaderExtBinaryBase::contains(StringRef Key) const {
+  constexpr size_t MinProfileToModuleSizeRatio = 16;
+
+  assert(NameTable && "NameTable should be populated before querying");
+  // Use the existing module filter for module-scoped queries when it is
+  // substantially smaller than the profile NameTable.
+  if (useMD5() || !FuncsToUse.contains(Key) ||
+      FuncsToUse.size() > NameTable->size() / MinProfileToModuleSizeRatio)
+    return SampleProfileReaderBinary::contains(Key);
+
+  if (!ModuleNameTableEntries) {
+    ModuleNameTableEntries.emplace();
+    ModuleNameTableEntries->reserve(FuncsToUse.size());
+    for (FunctionId FID : *NameTable) {
+      StringRef Name = FID.stringRef();
+      if (FuncsToUse.contains(Name))
+        ModuleNameTableEntries->insert(Name);
+    }
+  }
+  return ModuleNameTableEntries->contains(Key);
+}
+
 std::error_code
 SampleProfileReaderExtBinaryBase::readFuncOffsetTable(bool IsEytzinger,
                                                       bool IsNested) {
diff --git a/llvm/unittests/ProfileData/SampleProfTest.cpp b/llvm/unittests/ProfileData/SampleProfTest.cpp
index a991333ef2bcf..3f21f21584f10 100644
--- a/llvm/unittests/ProfileData/SampleProfTest.cpp
+++ b/llvm/unittests/ProfileData/SampleProfTest.cpp
@@ -10,6 +10,7 @@
 #include "llvm/ADT/ArrayRef.h"
 #include "llvm/ADT/StringMap.h"
 #include "llvm/ADT/StringRef.h"
+#include "llvm/ADT/Twine.h"
 #include "llvm/Config/llvm-config.h"
 #include "llvm/IR/DebugInfoMetadata.h"
 #include "llvm/IR/LLVMContext.h"
@@ -719,6 +720,42 @@ TEST_F(SampleProfTest, roundtrip_ext_binary_profile) {
   testRoundTrip(SampleProfileFormat::SPF_Ext_Binary, false, false);
 }
 
+TEST_F(SampleProfTest, ExtBinaryModuleFilteredNameTableLookup) {
+  TempFile ProfileFile("profile", "", "", /*Unique=*/true);
+  createWriter(SampleProfileFormat::SPF_Ext_Binary, ProfileFile.path());
+
+  std::vector<std::string> ProfileNames;
+  ProfileNames.reserve(64);
+  SampleProfileMap Profiles;
+  for (unsigned I = 0; I != 64; ++I) {
+    ProfileNames.push_back((Twine("profile_function_") + Twine(I)).str());
+    FunctionSamples Samples;
+    Samples.setFunction(FunctionId(ProfileNames.back()));
+    Samples.addTotalSamples(1);
+    Profiles[StringRef(ProfileNames.back())] = std::move(Samples);
+  }
+
+  ASSERT_TRUE(NoError(Writer->write(Profiles)));
+  Writer->getOutputStream().flush();
+
+  Module M("my_module", Context);
+  FunctionType *FnType = FunctionType::get(Type::getVoidTy(Context), {}, false);
+  M.getOrInsertFunction("profile_function_0.llvm.123", FnType);
+  M.getOrInsertFunction("module_only_function.llvm.456", FnType);
+
+  readProfile(M, ProfileFile.path());
+  ASSERT_TRUE(NoError(Reader->read()));
+
+  // Exercise the module-filtered lookup using the canonical names collected
+  // from a present and an absent module function.
+  EXPECT_TRUE(Reader->contains("profile_function_0"));
+  EXPECT_FALSE(Reader->contains("module_only_function"));
+
+  // Queries outside the module must retain the unrestricted reader semantics.
+  EXPECT_TRUE(Reader->contains("profile_function_1"));
+  EXPECT_FALSE(Reader->contains("not_in_profile_or_module"));
+}
+
 // Verify the full ExtBinary round trip through composite profile sections.
 TEST_F(SampleProfTest, roundtrip_composite_ext_binary_profile) {
   testRoundTrip(SampleProfileFormat::SPF_Ext_Binary, false, false,

``````````

</details>


https://github.com/llvm/llvm-project/pull/227900


More information about the llvm-commits mailing list