[llvm] [llvm-strings] Add --encoding={s,S,utf8} option (PR #221794)
James Henderson via llvm-commits
llvm-commits at lists.llvm.org
Thu Oct 1 01:23:06 PDT 2026
================
@@ -0,0 +1,71 @@
+# Show that llvm-strings uses the specified encoding.
+RUN: echo a > %t
+RUN: echo ab >> %t
+RUN: echo abc >> %t
+RUN: echo abcd >> %t
+# UTF-8 encoding of U+200A HAIR SPACE (printable)
+RUN: printf 'abcd\342\200\212\n' >> %t
+# UTF-8 encoding of U+200B ZERO WIDTH SPACE (non-printable)
+RUN: printf 'abcd\342\200\212\342\200\213\n' >> %t
+
+# Check that long form options work.
+RUN: llvm-strings --encoding s %t | FileCheck --check-prefixes CHECK,CHECK-ASCII --strict-whitespace --match-full-lines %s
+RUN: llvm-strings --encoding S %t | FileCheck --check-prefixes CHECK,CHECK-LOCALE --strict-whitespace --match-full-lines %s
+RUN: llvm-strings --encoding utf8 %t | FileCheck --check-prefixes CHECK,CHECK-UTF8 --strict-whitespace --match-full-lines %s
+
+# Check that short form options work.
+RUN: llvm-strings -e s %t | FileCheck --check-prefixes CHECK,CHECK-ASCII --strict-whitespace --match-full-lines %s
+RUN: llvm-strings -e S %t | FileCheck --check-prefixes CHECK,CHECK-LOCALE --strict-whitespace --match-full-lines %s
+RUN: llvm-strings -e utf8 %t | FileCheck --check-prefixes CHECK,CHECK-UTF8 --strict-whitespace --match-full-lines %s
+
+# Check that the default output matches that of -e S.
+RUN: llvm-strings -e S %t > %t.1
+RUN: llvm-strings %t > %t.2
+RUN: cmp %t.1 %t.2
+
+# The first line containing just abcd should always be matched.
+CHECK:abcd
+# In -e s, the bytes of U+200A and U+200B should not be included.
+CHECK-ASCII:abcd
+CHECK-ASCII:abcd
+# In -e S, any bytes of U+200A and U+200B may or may not be included depending
+# on the locale in effect. The locale in effect may also be altered by llvm-lit.
+# Make sure if changing the test that the byte representations should not allow
+# this to be split into multiple strings in non-UTF-8 locales.
+CHECK-LOCALE:abcd{{(\xe2(\x80\x8a?)?)?}}
+CHECK-LOCALE:abcd{{(\xe2(\x80(\x8a(\xe2(\x80\x8b?)?)?)?)?)?}}
+# In -e utf8, all bytes of U+200A should be included, but none of U+200B.
----------------
jh7370 wrote:
Perhaps worth a comment somewhere around here about what the default character set is for llvm-strings' output. I.e. how are non-ASCII UTF-8 characters printed.
https://github.com/llvm/llvm-project/pull/221794
More information about the llvm-commits
mailing list