[clang] [clang] Improve diagnostics for invalid named-universal-characters (PR #206326)

Aaron Ballman via cfe-commits cfe-commits at lists.llvm.org
Tue Jun 30 04:45:23 PDT 2026


================
@@ -550,6 +556,44 @@ static void DiagnoseInvalidUnicodeCharacterName(
 
   namespace u = llvm::sys::unicode;
 
+  auto StringifyCodePoint = [](llvm::UTF32 CodePoint) -> llvm::SmallString<16> {
+    llvm::SmallString<16> Result;
+    if (u::isPrintable(CodePoint)) {
+      std::string CharUTF8;
+      llvm::convertUTF32ToUTF8String(llvm::ArrayRef<llvm::UTF32>(&CodePoint, 1),
+                                     CharUTF8);
+      Result.append("'");
+      Result.append(CharUTF8);
+      Result.append("' U+");
+    } else {
+      Result.append("U+");
+    }
+    llvm::raw_svector_ostream OS(Result);
+    llvm::write_hex(OS, CodePoint, llvm::HexPrintStyle::Upper, 4);
+    return Result;
+  };
----------------
AaronBallman wrote:

Drive-by comment: I tend to agree with @eisenwave that the output is worse; I slightly prefer the angle brackets (but can live with them being lost), but I prefer printing the code point (at least for confusable characters if not every time) because that can be crucial information sometimes. Printing the quotes is mostly a consistency thing for other diagnostics printing characters like `err_ucn_escape_basic_scs`, `trigraph_converted`, `err_invalid_char_raw_delim`, etc.

The only purpose to `EscapeStringForDiagnostic` is to escape strings for diagnostics; if there are output regressions when using it for more diagnostics, the function can be updated to accommodate (potentially with policy parameters). What am I missing?

(I don't insist on a change, @cor3ntin is the expert here, but his suggestion to use that function seems like something I would have also suggested to ensure diagnostic consistency, hence the question.)

https://github.com/llvm/llvm-project/pull/206326


More information about the cfe-commits mailing list