[llvm] [llvm-strings] Use small buffer instead of reading whole file (PR #163073)
Harald van Dijk via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 22 04:15:42 PDT 2026
================
@@ -0,0 +1,107 @@
+## Show that strings interacting with the read-chunk boundary are reported
+## correctly. The input files are crafted assuming the native read chunk size
+## (sys::fs::DefaultReadChunkSize) of 16384 bytes.
+##
+## The inputs, and the expected output lines that are too long to write as
+## CHECK lines, are generated by the Python script at the end of this file,
+## which is run as "%python %s". Each argument describes one part of the
+## output: a bare number is that many zero bytes, TEXT at N is TEXT repeated N
+## times, and anything else is literal text. With --line the parts are written
+## as a single line of llvm-strings output instead, preceded by the offset
+## header if --offset is given.
+
+## Case 1: at least min string size appears before the boundary, unprintable
+## byte as first byte of the next chunk. The string is printed on its own,
+## with the offset of its start (16376 = 16384 - 8).
+# RUN: %python %s 16376 ENDCHUNK 4 > %t.1
+# RUN: llvm-strings --radix=d %t.1 | FileCheck %s --check-prefix=CASE1 --match-full-lines --strict-whitespace --implicit-check-not={{.}}
+
+# CASE1: 16376 ENDCHUNK
+
+## Case 2: at least min string size appears before the boundary, printable
+## byte as first byte of the next chunk. The prefix is printed together with
+## the following characters, as a single string.
+# RUN: %python %s 16378 BEFORE AFTER! 4 > %t.2
+# RUN: llvm-strings --radix=d %t.2 | FileCheck %s --check-prefix=CASE2 --match-full-lines --strict-whitespace --implicit-check-not={{.}}
+
+# CASE2: 16378 BEFOREAFTER!
+
+## Case 3: less than min string size appears before the boundary, unprintable
+## byte as the next byte. The prefix is not printed.
+# RUN: %python %s 16382 AB 4 > %t.3
+# RUN: llvm-strings %t.3 | count 0
+
+## Case 4: less than min string size appears before the boundary, printable
+## bytes as the next bytes, forming a min length string. The prefix is
+## printed together with the following characters, with the offset of its
+## true start (16381 = 16384 - 3).
+# RUN: %python %s 16381 ABC DEF 4 > %t.4
+# RUN: llvm-strings --radix=d %t.4 | FileCheck %s --check-prefix=CASE4 --match-full-lines --strict-whitespace --implicit-check-not={{.}}
+
+# CASE4: 16381 ABCDEF
+
+## Case 5: the prefix is empty at the start of a chunk that starts with a min
+## length string (16384).
+# RUN: %python %s 16384 FRESH 4 > %t.5
+# RUN: llvm-strings --radix=d %t.5 | FileCheck %s --check-prefix=CASE5 --match-full-lines --strict-whitespace --implicit-check-not={{.}}
+
+# CASE5: 16384 FRESH
+
+## Case 6: a string that spans the entirety of one chunk, with one character
+## before the chunk start and one after its end. The string is printed
+## intact, once. The expected output is generated rather than written as a
+## CHECK line because it is a chunk long.
+# RUN: %python %s 16383 S M at 16384 E 4 > %t.6
+# RUN: %python %s --line S M at 16384 E > %t.6.expected
+# RUN: llvm-strings %t.6 > %t.6.out
+# RUN: diff %t.6.expected %t.6.out
+
+## Case 7: the minimum length is greater than the chunk size, so a candidate
+## has to be buffered across more than one chunk before it is known to be long
+## enough. A run of exactly the minimum length is printed, with the offset of
+## its start, which lies in an earlier chunk than the one where it is printed.
+# RUN: %python %s 8 A at 20000 4 > %t.7
+# RUN: %python %s --line --offset=8 A at 20000 > %t.7.expected
+# RUN: llvm-strings --radix=d --bytes=20000 %t.7 > %t.7.out
+# RUN: diff %t.7.expected %t.7.out
+
+## Case 8: as case 7, but the run is one byte short of the minimum length, so
+## the buffered candidate is discarded and nothing is printed.
+# RUN: %python %s 8 A at 19999 4 > %t.8
+# RUN: llvm-strings --bytes=20000 %t.8 | count 0
+
+## Case 9: a string is terminated by the end of the file, with no trailing
+## unprintable byte. The string must still be printed.
+# RUN: %python %s 16380 TRAILING > %t.9
+# RUN: llvm-strings %t.9 | FileCheck %s --check-prefix=CASE9 --match-full-lines --strict-whitespace --implicit-check-not={{.}}
+
+# CASE9:TRAILING
----------------
hvdijk wrote:
Just thinking a bit more: there is another option that avoids the need for this script entirely. If we have a `--chunk-size` option (maybe undocumented, only intended for use in tests) and run the test with a chunk size that's small enough, we can just hardcode the test data without needing to run a script to construct it.
https://github.com/llvm/llvm-project/pull/163073
More information about the llvm-commits
mailing list