[llvm] [llvm-strings] Use small buffer instead of reading whole file (PR #163073)

James Henderson via llvm-commits llvm-commits at lists.llvm.org
Wed Aug 26 01:31:50 PDT 2026


================
@@ -113,12 +113,20 @@ static void strings(raw_ostream &OS, StringRef FileName,
     }
   };
 
-  // llvm-strings should be able to process a very large file on a
-  // memory-budgeted machine, so the file is read in chunks. To handle this, we
-  // read the file in chunk instead of copying the whole file into memory.
+  // To handle very large files without consuming excessive memory, we read the
+  // file in a little at a time and process it then rather than reading the
+  // entire file at once.
+  //
+  // A string is only buffered until it is known to be long enough to print;
+  // from then on it is streamed out directly, so an arbitrarily long string
+  // never needs an arbitrarily large buffer. Candidate therefore only ever
+  // holds a run that is shorter than MinLength and that was cut off by the end
+  // of a chunk.
+  const size_t Min = MinLength;
   SmallString<DefaultMinLength> Candidate;
   bool InString = false;
-  unsigned StringStart = 0, Offset = 0;
+  // Offset of the start of the current chunk within the file.
+  unsigned ChunkOffset = 0;
----------------
jh7370 wrote:

`unsigned` seems to be the wrong type here. It should probably be `size_t` since we're dealing with file and container sizes.

https://github.com/llvm/llvm-project/pull/163073


More information about the llvm-commits mailing list