[llvm] [llvm] Fix string and Twine concatenation for C++26 P2591R5 (PR #227959)

via llvm-commits llvm-commits at lists.llvm.org
Fri Oct 2 06:19:35 PDT 2026


rdong8 wrote:

>I should add that there's no reason not to keep an explicit conversion operator to std::string_view, as I assume it's the implicit nature of the operator that is causing the issue (assuming it is a conversion operator at all - I haven't actually checked).

I gave it a try and it looks like it would require a 19 file diff for Support alone vs 15 files for LLVM right now

>I think this is worth exploring; one of the benefits of Twine is that you generally don't have to explicitly construct it when concatenating objects of different string types;

I know this is a common pattern but I feel that it's more of an antipattern, simply because:

- `(x1 + x2 + ... + xk).str()` relying on `Twine` conversion is likely to be outperformed by a `concat(x1, ..., xk)` implementation: you don't need to traverse a tree structure, and you can reserve the exact size upfront
- New people such as myself might expect `StringRef` to behave like `string_view` wrt `operator+`, but it does not.

Some quick AI generated microbenchmarks:

<details>
<summary>`Twine::str()` vs naive `concat()` implemented with `reserve() + append()` vs `concat` with `resize_and_overwrite` vs C++26 `std::operator+(const std::string&, std::string_view)`</summary>

```
-----------------------------------------------------------------
Benchmark                       Time             CPU   Iterations
-----------------------------------------------------------------
# 1. Small Strings (SSO: ~11 bytes, fits in std::string small-string buffer)
BM_Twine_Small_SSO           58.9 ns         57.8 ns     10971478
BM_Reserve_Small_SSO         31.9 ns         31.3 ns     22539894   (1.8x faster)
BM_Overwrite_Small_SSO       14.4 ns         14.1 ns     47044477   (4.1x faster)
BM_StdPlus_Small_SSO         2.42 ns         2.38 ns    305064339  (24.3x faster)

# 2. Medium Strings (Heap allocated std::string, fits in Twine's SmallString<256>)
BM_Twine_Medium              89.2 ns         87.7 ns      7811980
BM_Reserve_Medium            53.0 ns         51.8 ns     11481115   (1.7x faster)
BM_Overwrite_Medium          33.6 ns         33.0 ns     28604146   (2.7x faster)
BM_StdPlus_Medium            17.0 ns         16.7 ns     58022289   (5.2x faster)

# 3. Large Strings (> 256 bytes, overflows Twine's SmallString<256>)
BM_Twine_Large                138 ns          135 ns      5632274
BM_Reserve_Large             50.5 ns         49.7 ns     13364358   (2.7x faster)
BM_Overwrite_Large           44.9 ns         43.8 ns     20377840   (3.1x faster)
BM_StdPlus_Large             33.4 ns         32.7 ns     19924956   (4.1x faster)

# 4. Multi-part Concatenation (3 parts: Prefix + Name + Suffix)
BM_Twine_3Parts               114 ns          112 ns      5934011
BM_Reserve_3Parts            58.9 ns         57.7 ns     11229078   (1.9x faster)
BM_Overwrite_3Parts          15.8 ns         15.5 ns     39283234   (7.2x faster)
BM_StdPlus_3Parts            56.4 ns         55.2 ns     11258969   (2.0x faster)
```

</details>

<!-- 

<details>
<summary>For  `Twine::NodeKind`'s that aren't string-like:</summary>


```
---------------------------------------------------------------------------
Benchmark                                 Time             CPU   Iterations
---------------------------------------------------------------------------
# 1. Decimal Integers (DecIKind)
BM_Int_Twine_Str                       77.3 ns         76.3 ns      6715694
BM_Int_ToString_Plus                   36.1 ns         35.3 ns     19796926  (2.1x faster than Twine)
BM_Int_StdFormat                        120 ns          118 ns      5370487  (slower, format parsing)
BM_Int_ToChars                         13.9 ns         13.6 ns     47967967  (5.6x faster than Twine)

# 2. Hexadecimal (UHexKind / utohexstr)
BM_Hex_Twine_Str                       66.3 ns         65.5 ns     11673191  (2.3x faster than std::format)
BM_Hex_StdFormat                        150 ns          146 ns      7174959

# 3. Single Character (CharKind)
BM_Char_Twine_Str                      60.4 ns         60.1 ns     10852059
BM_Char_StdPlus                        25.1 ns         24.9 ns     27752130  (2.4x faster than Twine)

# 4. Rich Formatting (FormatvObjectKind)
BM_Formatv_Twine_Str                    466 ns          459 ns      1427125
BM_Formatv_Direct_Str                   504 ns          494 ns      1880906
BM_Formatv_StdFormat                    162 ns          160 ns      5566142  (2.9x faster than formatv)

# 5. Streaming to raw_ostream (No .str() call)
BM_Stream_Twine_ToOS                   84.2 ns         82.6 ns      7512987
BM_Stream_Materialized_StdString       66.9 ns         65.3 ns     13112903
BM_Stream_Individual_Pieces            40.2 ns         39.4 ns     17894097  (2.1x faster than Twine)
```

</details>

-->

https://github.com/llvm/llvm-project/pull/227959


More information about the llvm-commits mailing list