[llvm] [InstCombine] Fold strlen/wcslen over select/phi of constants (PR #223246)
Haram Jeong via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 14 21:15:03 PDT 2026
haramj wrote:
I ran a broader regression check and a before/after performance experiment for this change.
Follow-up fix: https://github.com/llvm/llvm-project/pull/223246/commits/ceb15887eaf9
### Fixpoint regression
The original patch passed the four tests listed in the description, but eight existing strlen/wcslen tests failed the InstCombine fixpoint check. In `prepareWorklist()`, constant-folding a call's result can leave the call itself alive when its attributes are insufficient for immediate DCE. The unconditional `continue` then skipped adding that remaining instruction to the worklist.
The follow-up commit only skips further processing when the instruction was actually erased. Otherwise, the remaining call reaches the worklist and can be simplified in the same pass invocation. The existing tests reproduce this regression; their expectations were not weakened.
Validation after the fix:
- All 19 strlen/wcslen and libcall optimization-remark tests pass.
- InstCombine, InstSimplify, SCCP, and the libcall optimization-remark test: 2,263 passed, 92 unsupported, no failures in this configuration.
- `-verify-each` passes for the benchmark pipelines.
### Performance experiment
Baseline: `d283a5551095` (the parent of the original patch).
Original patch: `ae605b914df994b7cfb0b1b98c5469bdc9093389`.
The follow-up fix produces byte-identical optimized benchmark IR to the original patch for C O2/O3 and C++ O2.
Environment: macOS arm64, host CPU reported as apple-m1; Release LLVM build with assertions off. The same Clang 22.1.8 frontend emitted unoptimized IR with `-O2 -Xclang -disable-llvm-passes`; each optimizer ran the full `default<O2>` or `default<O3>` pipeline. The same locally built LLVM backend (`llc -O2`) and linker were used for both variants. These are not measurements comparing two different Clang releases.
Final O2 measurements (median ns per invocation, including driver/function-call overhead):
| Synthetic workload | Before | After | Speedup |
|---|---:|---:|---:|
| Two-literal select, existing optimization/control | 1.233 | 1.230 | 1.00x |
| Nested three-literal strlen | 5.086 | 1.233 | 4.13x |
| Status switch/PHI | 7.306 | 1.233 | 5.92x |
| Nested wcslen | 9.681 | 1.233 | 7.85x |
| 64/128/256-byte literal selection | 14.047 | 1.229 | 11.43x |
| Selection including a dynamic pointer/control | 7.033 | 6.992 | 1.01x |
| Dynamic strlen/control | 5.336 | 5.332 | 1.00x |
| Log-size accounting with surrounding work | 17.976 | 14.391 | 1.25x |
| C++ std::string_view construction and size | 5.478 | 1.229 | 4.46x |
| C++ std::string construction, append, and full-content hash | 102.283 | 93.261 | 1.10x |
The last workload showed an **8.8% reduction in elapsed time**. It includes string construction, appending a runtime message, and hashing the resulting bytes; it does not include I/O. C O3 showed the same call-removal pattern and similar improvements.
Methodology:
- Each optimized kernel was isolated into a separate executable using internalize/globaldce after optimization, with the same procedure for both variants.
- Runtime inputs use a deterministic 4,096-element xorshift sequence. The driver is a separate translation unit without LTO; kernels are noinline.
- 100,000 warmup calls; 15 repetitions of 20 million calls (2 million for the C++ message workload), with randomized variant order.
- Every final checksum was checked against an independently computed Python reference.
- No builds or large test runs overlapped timed measurements. CPU affinity/frequency and other desktop applications were not controlled.
An initial combined-executable experiment showed a sizeable slowdown even for an unchanged control. Its kernel assembly was unchanged, and the effect largely disappeared with per-kernel executables. The table reports the isolated experiment; I do not attribute that initial difference to the transformation itself or claim to have identified its precise microarchitectural cause.
These are **synthetic workload results, not whole-application speedups**. The relevant factor is the runtime frequency of constant select/PHI patterns, not strlen call count alone. The small differences in unchanged controls should not be interpreted as improvements. Real-application throughput, compiler build-time effects, and other platforms remain unmeasured.
### Reproduction sources
The files below contain the actual kernels, driver, and scripts. They use the macOS Homebrew compiler path and the repository layout from this experiment; adjust compiler paths for another machine.
Place them under `build/benchmarks/pr223246/`. Build the baseline optimizer from `d283a5551095` in a separate checkout with the same Release configuration, and copy it to that directory as `opt-base`. Copy the updated PR optimizer as `opt-pr`. The updated PR build should also provide `build/bin/opt` and `build/bin/llc`.
Run `python3 build/benchmarks/pr223246/run.py` to generate common C IR and the exploratory measurements, then `run_isolated.py` for the final 15-repeat C measurements, and `run_cpp.py` for C++ measurements. Final results are in `isolated/raw.csv`, `isolated/summary.json`, `cpp/raw.csv`, and `cpp/summary.json`; generated IR and assembly are retained alongside them. `run_isolated.py` copies `build/bin/opt` to `opt-fixed` and verifies the optimized IR matches `opt-pr`.
<details>
<summary>kernels.c</summary>
```cpp
#include <stddef.h>
#include <string.h>
#include <wchar.h>
#define NI __attribute__((noinline))
NI size_t direct(unsigned x, const char *p) { return strlen(x & 1 ? "ok" : "missing"); }
NI size_t nested(unsigned x, const char *p) { return strlen(x & 1 ? "ok" : x & 2 ? "missing" : "cancelled"); }
NI size_t status(unsigned x, const char *p) {
const char *s;
switch(x & 3) {case 0:s="ok";break;case 1:s="missing";break;case 2:s="cancelled";break;default:s="timed out!";break;}
return strlen(s);
}
NI size_t wide(unsigned x, const char *p) {return wcslen(x & 1 ? L"ok" : x & 2 ? L"missing" : L"cancelled");}
#define S64 "0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
NI size_t long_nested(unsigned x, const char *p) {return strlen(x & 1 ? S64 : x & 2 ? S64 S64 : S64 S64 S64 S64);}
NI size_t dynamic(unsigned x, const char *p) {return strlen(x & 1 ? "ok" : x & 2 ? "missing" : p);}
NI size_t scan(unsigned x, const char *p) {return strlen(p);}
// Synthetic log-size accounting: ordinary work surrounds the literal length.
NI size_t log_size(unsigned x, const char *p) {
size_t n=strlen(x & 1 ? "INFO" : x & 2 ? "WARNING" : "ERROR");
for(unsigned v=x;v;v/=10) ++n;
return n+strlen(p)+4;
}
```
</details>
<details>
<summary>kernels.cpp</summary>
```cpp
#include <string>
#include <string_view>
#include <cstddef>
extern "C" __attribute__((noinline)) std::size_t cpp_view(unsigned x, const char *) {
const char *label = x & 1 ? "ok" : x & 2 ? "missing" : "cancelled";
return std::string_view(label).size();
}
// Synthetic formatted-message workload: construct, append a dynamic message,
// then consume every output byte. Allocations and copies remain observable
// through the hash; no I/O is included in the measured function.
extern "C" __attribute__((noinline)) std::size_t cpp_message(unsigned x, const char *message) {
const char *label = x & 1 ? "INFO" : x & 2 ? "WARNING" : "ERROR";
std::string output(label);
output += ": ";
output += message;
std::size_t hash = 0;
for (unsigned char c : output) hash = hash * 33 + c;
return hash;
}
```
</details>
<details>
<summary>driver.c</summary>
```cpp
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <time.h>
#include <stddef.h>
#define DECL(n) extern size_t n(unsigned,const char*);
DECL(direct) DECL(nested) DECL(status) DECL(wide) DECL(long_nested) DECL(dynamic) DECL(scan) DECL(log_size)
static double now(void){struct timespec t;clock_gettime(CLOCK_MONOTONIC,&t);return t.tv_sec+t.tv_nsec*1e-9;}
int main(int argc,char**argv){
if(argc!=4)return 2;
#ifdef BENCH_CASE
extern size_t BENCH_CASE(unsigned,const char*);
size_t (*fn)(unsigned,const char*)=BENCH_CASE;
#else
size_t (*fn)(unsigned,const char*)=0;
#define PICK(n) if(!strcmp(argv[1],#n))fn=n;
PICK(direct) PICK(nested) PICK(status) PICK(wide) PICK(long_nested) PICK(dynamic) PICK(scan) PICK(log_size)
#endif
if(!fn)return 3;
uint32_t inputs[4096],s=123456789;for(int i=0;i<4096;i++){s^=s<<13;s^=s>>17;s^=s<<5;inputs[i]=s;}
unsigned long count=strtoul(argv[2],0,10);uint64_t sum=0;
for(unsigned i=0;i<100000;i++)sum+=fn(inputs[i&4095],argv[3]);
double begin=now();for(unsigned long i=0;i<count;i++)sum+=fn(inputs[i&4095],argv[3]);double elapsed=now()-begin;
printf("%.9f,%llu\n",elapsed,(unsigned long long)sum);return 0;
}
```
</details>
<details>
<summary>run.py</summary>
```python
from pathlib import Path
import subprocess,re,csv,random,statistics,json,hashlib,platform
D=Path(__file__).resolve().parent
CLANG='/opt/homebrew/opt/llvm/bin/clang'
def run(args): return subprocess.check_output([str(x) for x in args],text=True)
run([CLANG,'-O2','-Xclang','-disable-llvm-passes','-S','-emit-llvm',D/'kernels.c','-o',D/'input.ll'])
run([CLANG,'-O2','-c',D/'driver.c','-o',D/'driver.o'])
static={}
for variant in ['base','pr']:
for level in ['O2','O3']:
ir=D/f'{variant}-{level}.ll'
run([D/f'opt-{variant}',D/'input.ll',f'-passes=default<{level}>','-verify-each','-S','-o',ir])
static[f'{variant}-{level}']={m.group(1):len(re.findall(r'\b(?:call|invoke)\b[^\n]*@(?:strlen|wcslen)\(',m.group(2))) for m in re.finditer(r'^define[^@]+@(\w+)\([^\n]*\).*?\{(.*?)^}',ir.read_text(),re.M|re.S)}
run([D.parents[1]/'bin/llc','-O2','-filetype=obj',ir,'-o',D/f'{variant}-{level}.o'])
run([D.parents[1]/'bin/llc','-O2',ir,'-o',D/f'{variant}-{level}.s'])
run([CLANG,D/'driver.o',D/f'{variant}-{level}.o','-o',D/f'bench-{variant}-{level}'])
(D/'static.json').write_text(json.dumps(static,indent=2))
print(json.dumps(static,indent=2),flush=True)
cases=['direct','nested','status','wide','long_nested','dynamic','scan','log_size']
N=20000000;rows=[];rng=random.Random(223246)
for repeat in range(11):
order=[(case,level) for case in cases for level in ['O2','O3']];rng.shuffle(order)
for case,level in order:
variants=['base','pr'];rng.shuffle(variants);checks=[]
for variant in variants:
sec,checksum=run([D/f'bench-{variant}-{level}',case,str(N),'runtime message used for checksum validation']).strip().split(',')
checks.append(checksum);rows.append(dict(repeat=repeat,case=case,level=level,variant=variant,seconds=float(sec),iterations=N,checksum=checksum))
assert len(set(checks))==1,(case,checks)
print(f'repeat {repeat+1}/11 complete',flush=True)
with (D/'raw.csv').open('w') as f:
w=csv.DictWriter(f,fieldnames=rows[0]);w.writeheader();w.writerows(rows)
summary=[]
for level in ['O2','O3']:
for case in cases:
a={v:[r['seconds']/N*1e9 for r in rows if r['case']==case and r['level']==level and r['variant']==v] for v in ['base','pr']}
b,p=statistics.median(a['base']),statistics.median(a['pr'])
summary.append(dict(level=level,case=case,base_ns=b,pr_ns=p,speedup=b/p,reduction_percent=100*(1-p/b),base_min=min(a['base']),base_max=max(a['base']),pr_min=min(a['pr']),pr_max=max(a['pr'])))
(D/'summary.json').write_text(json.dumps(summary,indent=2))
(D/'environment.json').write_text(json.dumps(dict(platform=platform.platform(),clang=run([CLANG,'--version']),opt=run([D/'opt-pr','--version']),hashes={f:hashlib.sha256((D/f).read_bytes()).hexdigest() for f in ['opt-base','opt-pr','input.ll']},iterations=N,repeats=11,seed=223246),indent=2))
print(json.dumps(summary,indent=2))
```
</details>
<details>
<summary>run_isolated.py</summary>
```python
from pathlib import Path
import subprocess,re,csv,random,statistics,json,shutil,hashlib
D=Path(__file__).resolve().parent
CLANG='/opt/homebrew/opt/llvm/bin/clang'; LLC=D.parents[1]/'bin/llc'; FIX=D.parents[1]/'bin/opt'
def run(args):return subprocess.check_output([str(x) for x in args],text=True)
cases=['direct','nested','status','wide','long_nested','dynamic','scan','log_size']
out=D/'isolated';out.mkdir(exist_ok=True)
shutil.copy2(FIX,D/'opt-fixed')
# Confirm the regression fix preserves each benchmark's optimized IR.
for level in ['O2','O3']:
run([FIX,D/'input.ll',f'-passes=default<{level}>','-verify-each','-S','-o',D/f'fixed-{level}.ll'])
assert (D/f'fixed-{level}.ll').read_bytes()==(D/f'pr-{level}.ll').read_bytes()
for case in cases:
run([CLANG,'-O2',f'-DBENCH_CASE={case}','-c',D/'driver.c','-o',out/f'driver-{case}.o'])
for level in ['O2','O3']:
for variant in ['base','pr']:
stem=out/f'{case}-{variant}-{level}'
run([FIX,D/f'{variant}-{level}.ll','-passes=internalize,globaldce',f'-internalize-public-api-list={case}','-S','-o',str(stem)+'.ll'])
run([LLC,'-O2','-filetype=obj',str(stem)+'.ll','-o',str(stem)+'.o'])
run([LLC,'-O2',str(stem)+'.ll','-o',str(stem)+'.s'])
run([CLANG,out/f'driver-{case}.o',str(stem)+'.o','-o',stem])
# Independent expected checksum for deterministic input stream, including warmup.
s=123456789;inputs=[]
for _ in range(4096):
s^=(s<<13)&0xffffffff;s^=s>>17;s^=(s<<5)&0xffffffff;inputs.append(s)
message='runtime message used for checksum validation'
def reference(case,x):
if case=='direct':return 2 if x&1 else 7
if case in ['nested','wide']:return 2 if x&1 else 7 if x&2 else 9
if case=='status':return [2,7,9,10][x&3]
if case=='long_nested':return 64 if x&1 else 128 if x&2 else 256
if case=='dynamic':return 2 if x&1 else 7 if x&2 else len(message)
if case=='scan':return len(message)
if case=='log_size':return (4 if x&1 else 7 if x&2 else 5)+(len(str(x)) if x else 0)+len(message)+4
N=20000000;rows=[];rng=random.Random(223246)
expected={}
for case in cases:
vals=[reference(case,x) for x in inputs]
expected[case]=sum(sum(vals)*(n//4096)+sum(vals[:n%4096]) for n in [N,100000])
for repeat in range(15):
order=[(c,l) for c in cases for l in ['O2','O3']];rng.shuffle(order)
for case,level in order:
variants=['base','pr'];rng.shuffle(variants)
for variant in variants:
sec,checksum=run([out/f'{case}-{variant}-{level}',case,str(N),message]).strip().split(',')
assert int(checksum)==expected[case],(case,checksum,expected[case])
rows.append(dict(repeat=repeat,case=case,level=level,variant=variant,seconds=float(sec),iterations=N,checksum=checksum))
print(f'repeat {repeat+1}/15 complete',flush=True)
with (out/'raw.csv').open('w') as f:
w=csv.DictWriter(f,fieldnames=rows[0]);w.writeheader();w.writerows(rows)
summary=[]
for level in ['O2','O3']:
for case in cases:
a={v:[r['seconds']/N*1e9 for r in rows if r['case']==case and r['level']==level and r['variant']==v] for v in ['base','pr']}
b,p=statistics.median(a['base']),statistics.median(a['pr'])
pairs=[x/y for x,y in zip(a['base'],a['pr'])]
boot=sorted(statistics.median(rng.choices(pairs,k=len(pairs))) for _ in range(10000))
summary.append(dict(level=level,case=case,base_ns=b,pr_ns=p,speedup=b/p,reduction_percent=100*(1-p/b),paired_speedup_ci95=[boot[250],boot[9749]],base_min=min(a['base']),base_max=max(a['base']),pr_min=min(a['pr']),pr_max=max(a['pr'])))
(out/'summary.json').write_text(json.dumps(summary,indent=2))
print(json.dumps(summary,indent=2))
```
</details>
<details>
<summary>run_cpp.py</summary>
```python
from pathlib import Path
import subprocess,re,json,csv,random,statistics
D=Path(__file__).resolve().parent;O=D/'cpp';O.mkdir(exist_ok=True)
C='/opt/homebrew/opt/llvm/bin/clang';CC='/opt/homebrew/opt/llvm/bin/clang++';LLC=D.parents[1]/'bin/llc'
def run(args):return subprocess.check_output([str(x) for x in args],text=True)
run([CC,'-std=c++17','-O2','-Xclang','-disable-llvm-passes','-S','-emit-llvm',D/'kernels.cpp','-o',O/'input.ll'])
static={}
for v in ['base','pr','fixed']:
run([D/f'opt-{v}',O/'input.ll','-passes=default<O2>','-verify-each','-S','-o',O/f'{v}.ll'])
static[v]={m.group(1):len(re.findall(r'\b(?:call|invoke)\b[^\n]*@strlen\(',m.group(2))) for m in re.finditer(r'^define[^@]+@(\w+)\([^\n]*\).*?\{(.*?)^}',(O/f'{v}.ll').read_text(),re.M|re.S) if m.group(1).startswith('cpp_')}
assert (O/'pr.ll').read_bytes()==(O/'fixed.ll').read_bytes()
(O/'static.json').write_text(json.dumps(static,indent=2));print(static,flush=True)
for case in ['cpp_view','cpp_message']:
run([C,'-O2',f'-DBENCH_CASE={case}','-c',D/'driver.c','-o',O/f'driver-{case}.o'])
for v in ['base','pr']:
stem=O/f'{case}-{v}'
run([D/'opt-fixed',O/f'{v}.ll','-passes=internalize,globaldce',f'-internalize-public-api-list={case}','-S','-o',str(stem)+'.ll'])
run([LLC,'-O2','-filetype=obj',str(stem)+'.ll','-o',str(stem)+'.o'])
run([LLC,'-O2',str(stem)+'.ll','-o',str(stem)+'.s'])
run([CC,O/f'driver-{case}.o',str(stem)+'.o','-o',stem])
message='runtime message used for checksum validation';inputs=[];s=123456789
for _ in range(4096):
s^=(s<<13)&0xffffffff;s^=s>>17;s^=(s<<5)&0xffffffff;inputs.append(s)
def ref(case,x):
if case=='cpp_view':return 2 if x&1 else 7 if x&2 else 9
st=('INFO' if x&1 else 'WARNING' if x&2 else 'ERROR')+': '+message
h=0
for c in st.encode():h=(h*33+c)&((1<<64)-1)
return h
rows=[];rng=random.Random(223246)
for repeat in range(15):
for case,N in [('cpp_view',20000000),('cpp_message',2000000)]:
vals=[ref(case,x) for x in inputs];expected=sum(sum(vals)*(n//4096)+sum(vals[:n%4096]) for n in [N,100000])&((1<<64)-1)
variants=['base','pr'];rng.shuffle(variants)
for v in variants:
sec,ch=run([O/f'{case}-{v}',case,str(N),message]).strip().split(',');assert int(ch)==expected
rows.append(dict(repeat=repeat,case=case,variant=v,seconds=float(sec),iterations=N,checksum=ch))
print(f'repeat {repeat+1}/15 complete',flush=True)
with (O/'raw.csv').open('w') as f:
w=csv.DictWriter(f,fieldnames=rows[0]);w.writeheader();w.writerows(rows)
summary=[]
for case in ['cpp_view','cpp_message']:
a={v:[r['seconds']/r['iterations']*1e9 for r in rows if r['case']==case and r['variant']==v] for v in ['base','pr']}
b,p=statistics.median(a['base']),statistics.median(a['pr']);pairs=[x/y for x,y in zip(a['base'],a['pr'])];boot=sorted(statistics.median(rng.choices(pairs,k=15)) for _ in range(10000))
summary.append(dict(case=case,base_ns=b,pr_ns=p,speedup=b/p,reduction_percent=100*(1-p/b),paired_speedup_ci95=[boot[250],boot[9749]]))
(O/'summary.json').write_text(json.dumps(summary,indent=2));print(summary)
```
</details>
https://github.com/llvm/llvm-project/pull/223246
More information about the llvm-commits
mailing list