[llvm] [llvm-strings] Use small buffer instead of reading whole file (PR #163073)
James Henderson via llvm-commits
llvm-commits at lists.llvm.org
Wed Sep 16 01:01:31 PDT 2026
================
@@ -1,70 +1,107 @@
## Show that strings interacting with the read-chunk boundary are reported
## correctly. The input files are crafted assuming the native read chunk size
## (sys::fs::DefaultReadChunkSize) of 16384 bytes.
+##
+## The inputs, and the expected output lines that are too long to write as
+## CHECK lines, are generated by the Python script at the end of this file,
+## which is run as "%python %s". Each argument describes one part of the
+## output: a bare number is that many zero bytes, TEXT at N is TEXT repeated N
+## times, and anything else is literal text. With --line the parts are written
+## as a single line of llvm-strings output instead, preceded by the offset
+## header if --offset is given.
## Case 1: at least min string size appears before the boundary, unprintable
## byte as first byte of the next chunk. The string is printed on its own,
-## with the offset of its start (0x3ff8 = 16384 - 8).
-RUN: %python -c "import sys; sys.stdout.buffer.write(b'\0' * 16376 + b'ENDCHUNK' + b'\0' * 4)" > %t.1
-RUN: llvm-strings --radix=x %t.1 | FileCheck %s --check-prefix=CASE1 --strict-whitespace --implicit-check-not={{.}}
+## with the offset of its start (16376 = 16384 - 8).
+# RUN: %python %s 16376 ENDCHUNK 4 > %t.1
+# RUN: llvm-strings --radix=d %t.1 | FileCheck %s --check-prefix=CASE1 --match-full-lines --strict-whitespace --implicit-check-not={{.}}
-CASE1:{{^}} 3ff8 ENDCHUNK{{$}}
+# CASE1: 16376 ENDCHUNK
## Case 2: at least min string size appears before the boundary, printable
## byte as first byte of the next chunk. The prefix is printed together with
## the following characters, as a single string.
-RUN: %python -c "import sys; sys.stdout.buffer.write(b'\0' * 16378 + b'BEFORE' + b'AFTER!' + b'\0' * 4)" > %t.2
-RUN: llvm-strings --radix=x %t.2 | FileCheck %s --check-prefix=CASE2 --strict-whitespace --implicit-check-not={{.}}
+# RUN: %python %s 16378 BEFORE AFTER! 4 > %t.2
+# RUN: llvm-strings --radix=d %t.2 | FileCheck %s --check-prefix=CASE2 --match-full-lines --strict-whitespace --implicit-check-not={{.}}
-CASE2:{{^}} 3ffa BEFOREAFTER!{{$}}
+# CASE2: 16378 BEFOREAFTER!
## Case 3: less than min string size appears before the boundary, unprintable
## byte as the next byte. The prefix is not printed.
-RUN: %python -c "import sys; sys.stdout.buffer.write(b'\0' * 16382 + b'AB' + b'\0' * 4)" > %t.3
-RUN: llvm-strings %t.3 | count 0
+# RUN: %python %s 16382 AB 4 > %t.3
+# RUN: llvm-strings %t.3 | count 0
## Case 4: less than min string size appears before the boundary, printable
## bytes as the next bytes, forming a min length string. The prefix is
## printed together with the following characters, with the offset of its
-## true start (0x3ffd = 16384 - 3).
-RUN: %python -c "import sys; sys.stdout.buffer.write(b'\0' * 16381 + b'ABC' + b'DEF' + b'\0' * 4)" > %t.4
-RUN: llvm-strings --radix=x %t.4 | FileCheck %s --check-prefix=CASE4 --strict-whitespace --implicit-check-not={{.}}
+## true start (16381 = 16384 - 3).
+# RUN: %python %s 16381 ABC DEF 4 > %t.4
+# RUN: llvm-strings --radix=d %t.4 | FileCheck %s --check-prefix=CASE4 --match-full-lines --strict-whitespace --implicit-check-not={{.}}
-CASE4:{{^}} 3ffd ABCDEF{{$}}
+# CASE4: 16381 ABCDEF
## Case 5: the prefix is empty at the start of a chunk that starts with a min
-## length string (0x4000 = 16384).
-RUN: %python -c "import sys; sys.stdout.buffer.write(b'\0' * 16384 + b'FRESH' + b'\0' * 4)" > %t.5
-RUN: llvm-strings --radix=x %t.5 | FileCheck %s --check-prefix=CASE5 --strict-whitespace --implicit-check-not={{.}}
+## length string (16384).
+# RUN: %python %s 16384 FRESH 4 > %t.5
+# RUN: llvm-strings --radix=d %t.5 | FileCheck %s --check-prefix=CASE5 --match-full-lines --strict-whitespace --implicit-check-not={{.}}
-CASE5:{{^}} 4000 FRESH{{$}}
+# CASE5: 16384 FRESH
## Case 6: a string that spans the entirety of one chunk, with one character
## before the chunk start and one after its end. The string is printed
## intact, once. The expected output is generated rather than written as a
## CHECK line because it is a chunk long.
-RUN: %python -c "import sys; sys.stdout.buffer.write(b'\0' * 16383 + b'S' + b'M' * 16384 + b'E' + b'\0' * 4)" > %t.6
-RUN: %python -c "import sys; sys.stdout.buffer.write(b'S' + b'M' * 16384 + b'E\n')" > %t.6.expected
-RUN: llvm-strings %t.6 > %t.6.out
-RUN: diff %t.6.expected %t.6.out
+# RUN: %python %s 16383 S M at 16384 E 4 > %t.6
+# RUN: %python %s --line S M at 16384 E > %t.6.expected
+# RUN: llvm-strings %t.6 > %t.6.out
+# RUN: diff %t.6.expected %t.6.out
## Case 7: the minimum length is greater than the chunk size, so a candidate
## has to be buffered across more than one chunk before it is known to be long
## enough. A run of exactly the minimum length is printed, with the offset of
-## its start (0x8), which precedes the chunk in which the decision is made.
-RUN: %python -c "import sys; sys.stdout.buffer.write(b'\0' * 8 + b'A' * 20000 + b'\0' * 4)" > %t.7
-RUN: %python -c "import sys; sys.stdout.buffer.write(b'%7x ' % 8 + b'A' * 20000 + b'\n')" > %t.7.expected
-RUN: llvm-strings --radix=x --bytes=20000 %t.7 > %t.7.out
-RUN: diff %t.7.expected %t.7.out
+## its start, which lies in an earlier chunk than the one where it is printed.
+# RUN: %python %s 8 A at 20000 4 > %t.7
+# RUN: %python %s --line --offset=8 A at 20000 > %t.7.expected
+# RUN: llvm-strings --radix=d --bytes=20000 %t.7 > %t.7.out
+# RUN: diff %t.7.expected %t.7.out
## Case 8: as case 7, but the run is one byte short of the minimum length, so
## the buffered candidate is discarded and nothing is printed.
-RUN: %python -c "import sys; sys.stdout.buffer.write(b'\0' * 8 + b'A' * 19999 + b'\0' * 4)" > %t.8
-RUN: llvm-strings --bytes=20000 %t.8 | count 0
+# RUN: %python %s 8 A at 19999 4 > %t.8
+# RUN: llvm-strings --bytes=20000 %t.8 | count 0
-## A string terminated by the end of the file (no trailing unprintable byte)
-## must still be printed.
-RUN: %python -c "import sys; sys.stdout.buffer.write(b'\0' * 16380 + b'TRAILING')" > %t.9
-RUN: llvm-strings %t.9 | FileCheck %s --check-prefix=EOF --strict-whitespace --implicit-check-not={{.}}
+## Case 9: a string is terminated by the end of the file, with no trailing
+## unprintable byte. The string must still be printed.
+# RUN: %python %s 16380 TRAILING > %t.9
+# RUN: llvm-strings %t.9 | FileCheck %s --check-prefix=CASE9 --match-full-lines --strict-whitespace --implicit-check-not={{.}}
-EOF:{{^}}TRAILING{{$}}
+# CASE9:TRAILING
+
+import sys
+
+Args = sys.argv[1:]
+Line = False
+Offset = None
+while Args and Args[0].startswith("--"):
+ Opt = Args.pop(0)
+ if Opt == "--line":
+ Line = True
+ elif Opt.startswith("--offset="):
+ Offset = int(Opt[len("--offset=") :])
+ else:
+ sys.exit("unknown option: " + Opt)
+
+Parts = []
+for Arg in Args:
+ if Arg.isdigit():
+ Parts.append(b"\0" * int(Arg))
+ else:
+ Text, At, Count = Arg.partition("@")
----------------
jh7370 wrote:
>From a readability perspective, I'm not convinced the use of "@" in a string to indicate this is a special case is great. I'd be tempted by a separate command-line option called e.g. "--repeat" which is either specified once per positional argument, or which takes a comma-separated list of numbers, which then causes the corresponding positional argument to be emitted that many times. Example:
```
%python %s 8 A 4 --repeat 1,20000,1
```
It would be an error to have mismatching sizes of --repeat and positional arguments, unless --repeat is empty. If unspecified, we'd assume each arg is emitted only once.
https://github.com/llvm/llvm-project/pull/163073
More information about the llvm-commits
mailing list