[llvm] [llvm-profgen] Add Arm SPE branch profile support (PR #223237)
Sergey Shcherbinin via llvm-commits
llvm-commits at lists.llvm.org
Sun Sep 13 04:54:16 PDT 2026
https://github.com/SergeyShch01 created https://github.com/llvm/llvm-project/pull/223237
Add `--spe-branch-profile` for generating AArch64 sample profiles from Arm SPE branch records.
Input may be raw `perf.data` or `perf script --show-mmap-events --itrace=bl1 -F ip,brstack` output.
Recording requires `arm_spe/branch_filter=1,event_filter=2/`, perf 6.15+, FEAT_SPEv1p2, and event filtering.
ArmSPEReader derives from LBRPerfReader and reuses the existing LBR/BRBE aggregation and profile-generation path.
perf provides each SPE record as a branch stack containing one entry (the second address of its executed range is then inferred as the next branching point by llvm-profgen).
Existing LBR/BRBE profile generation is unchanged; mismatch diagnostics now use aggregated sample weights.
Tests cover parsing, perf-data conversion, mmap relocation, taken state, range inference, symbolization, and target validation.
Documentation and release notes describe the option and recording requirements.
Assisted by GPT-5
>From 0c880df627bcf5671ca0e5a6abb97642580fda52 Mon Sep 17 00:00:00 2001
From: Sergey Shcherbinin <sscherbinin at nvidia.com>
Date: Sun, 13 Sep 2026 14:48:00 +0400
Subject: [PATCH] [llvm-profgen] Add Arm SPE branch profile support
---
llvm/docs/CommandGuide/llvm-profgen.md | 15 +-
llvm/docs/ReleaseNotes.md | 3 +
.../AArch64/spe-branch-profile.test | 484 +++++++++++
.../AArch64/spe-not-taken-branch.test | 232 ++++++
.../llvm-profgen/AArch64/spe-perfdata.test | 135 +++
.../llvm-profgen/AArch64/spe-range-gap.test | 774 ++++++++++++++++++
.../AArch64/spe-symbolized-profile.test | 109 +++
.../llvm-profgen/spe-target-validation.test | 45 +
llvm/tools/llvm-profgen/Options.h | 1 +
llvm/tools/llvm-profgen/PerfReader.cpp | 373 ++++++++-
llvm/tools/llvm-profgen/PerfReader.h | 130 ++-
llvm/tools/llvm-profgen/ProfiledBinary.cpp | 106 +++
llvm/tools/llvm-profgen/ProfiledBinary.h | 30 +
llvm/tools/llvm-profgen/llvm-profgen.cpp | 43 +-
14 files changed, 2430 insertions(+), 50 deletions(-)
create mode 100644 llvm/test/tools/llvm-profgen/AArch64/spe-branch-profile.test
create mode 100644 llvm/test/tools/llvm-profgen/AArch64/spe-not-taken-branch.test
create mode 100644 llvm/test/tools/llvm-profgen/AArch64/spe-perfdata.test
create mode 100644 llvm/test/tools/llvm-profgen/AArch64/spe-range-gap.test
create mode 100644 llvm/test/tools/llvm-profgen/AArch64/spe-symbolized-profile.test
create mode 100644 llvm/test/tools/llvm-profgen/spe-target-validation.test
diff --git a/llvm/docs/CommandGuide/llvm-profgen.md b/llvm/docs/CommandGuide/llvm-profgen.md
index 631a3e00f3ce2..a1a49f2140314 100644
--- a/llvm/docs/CommandGuide/llvm-profgen.md
+++ b/llvm/docs/CommandGuide/llvm-profgen.md
@@ -20,7 +20,10 @@ At least one of the following commands are required:
:::{option} --perfscript=<string[,string,...]>
Path of a trace created by the Linux `perf script` command. For LBR or BRBE
input, the raw perf data must contain branch stacks, for example from recording
-with `-b`.
+with `-b`. With `--spe-branch-profile`, the trace must contain only Arm SPE
+branch samples, for example from recording with
+`arm_spe/branch_filter=1,event_filter=2/`, and must be generated with
+`--show-mmap-events --itrace=bl1 -F ip,brstack`.
:::
:::{option} --etm=<string>
@@ -30,7 +33,9 @@ Requires the OpenCSD library version 1.5.4 or higher to be enabled during the bu
:::{option} --perfdata=<perfdata>, --pd
Path of raw perf data created by the Linux perf tool. For LBR or BRBE input, it
-must contain branch stacks, for example from recording with `-b`.
+must contain branch stacks, for example from recording with `-b`. For
+`--spe-branch-profile` input, it must contain only Arm SPE branch samples, from
+recording with `arm_spe/branch_filter=1,event_filter=2/`.
:::
:::{option} --unsymbolized-profile=<unsymbolized profile>, --up
@@ -69,6 +74,12 @@ descriptions of the format.
Print mmap events.
:::
+:::{option} --spe-branch-profile
+Read the `--perfscript` or `--perfdata` input as an Arm SPE branch profile of
+an AArch64 binary. Requires perf 6.15 or later and hardware with FEAT_SPEv1p2
+and event filtering.
+:::
+
:::{option} --warn-not-symbolized
Warn when an address covered by a recorded mmap range cannot be symbolized.
:::
diff --git a/llvm/docs/ReleaseNotes.md b/llvm/docs/ReleaseNotes.md
index 1fd90f5bd3be0..62c2e4e61c65d 100644
--- a/llvm/docs/ReleaseNotes.md
+++ b/llvm/docs/ReleaseNotes.md
@@ -282,6 +282,9 @@ Makes programs 10x faster by doing Special New Thing.
* llvm-mca no longer defaults -mcpu to "native"
+* llvm-profgen can now build a sample profile from an Arm SPE branch profile of
+ an AArch64 binary, with the new `--spe-branch-profile` option
+
* llvm-rc now supports `/showIncludes` to report header and resource-file
dependencies in a format compatible with Ninja's `deps = msvc` mode.
diff --git a/llvm/test/tools/llvm-profgen/AArch64/spe-branch-profile.test b/llvm/test/tools/llvm-profgen/AArch64/spe-branch-profile.test
new file mode 100644
index 0000000000000..d7500b49f91d5
--- /dev/null
+++ b/llvm/test/tools/llvm-profgen/AArch64/spe-branch-profile.test
@@ -0,0 +1,484 @@
+# RUN: split-file %s %t
+# RUN: yaml2obj %t/binary.yaml -o %t/spe-branch-profile.exe
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/perfscript --spe-branch-profile --skip-symbolization \
+# RUN: --use-offset=0 --format=text --output=%t/profile 2>&1 \
+# RUN: | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --input-file=%t/profile
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/no-ip.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --use-offset=0 --format=text \
+# RUN: --output=%t/no-ip.profile 2>&1 \
+# RUN: | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --check-prefix=NO-IP --input-file=%t/no-ip.profile
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/malformed.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --show-detailed-warning --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=INVALID
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/malformed.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=INVALID-SUMMARY
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/bad-fallthrough.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --use-offset=0 --format=text \
+# RUN: --output=%t/bad-fallthrough.profile 2>&1 \
+# RUN: | FileCheck %s --check-prefix=BAD-FALLTHROUGH
+# RUN: FileCheck %s --check-prefix=LAST-INST \
+# RUN: --input-file=%t/bad-fallthrough.profile
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/weighted-mismatch.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=WEIGHTED-MISMATCH
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/inconsistent-not-taken.perfscript \
+# RUN: --spe-branch-profile --skip-symbolization --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=INCONSISTENT-NOT-TAKEN
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/empty.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=EMPTY
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/unusable.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=NO-USABLE
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/partial.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --use-offset=0 --format=text \
+# RUN: --output=%t/partial.profile 2>&1 \
+# RUN: | FileCheck %s --check-prefix=PARTIAL
+# RUN: FileCheck %s --check-prefix=PARTIAL-PROFILE \
+# RUN: --input-file=%t/partial.profile
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/perfscript --spe-branch-profile --pid=2 \
+# RUN: --skip-symbolization --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=OTHER-PID
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/external-source.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --use-offset=0 --format=text \
+# RUN: --output=%t/external-source.profile
+# RUN: FileCheck %s --check-prefix=EXTERNAL-SOURCE \
+# RUN: --input-file=%t/external-source.profile
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/external-source.perfscript --spe-branch-profile \
+# RUN: --use-offset=0 --format=text --output=%t/external-source.symbolized
+# RUN: FileCheck %s --check-prefix=SYMBOLIZED \
+# RUN: --input-file=%t/external-source.symbolized
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/no-mmap.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=NO-MMAP
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/previous-target.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=DEEP-STACK
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/lbr.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=DEEP-STACK
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/damaged-line.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --show-detailed-warning --use-offset=0 \
+# RUN: --format=text --output=%t/damaged-line.profile 2>&1 \
+# RUN: | FileCheck %s --check-prefix=DAMAGED
+# RUN: FileCheck %s --check-prefix=DAMAGED-PROFILE \
+# RUN: --input-file=%t/damaged-line.profile
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/hybrid.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --show-detailed-warning --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=HYBRID
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/ip-addr.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=IP-ADDR
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/rejected.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --use-offset=0 --format=text \
+# RUN: --output=%t/rejected.profile 2>&1 \
+# RUN: | FileCheck %s --check-prefix=REJECTED
+# RUN: FileCheck %s --check-prefix=REJECTED-PROFILE \
+# RUN: --input-file=%t/rejected.profile
+# RUN: yaml2obj -DTYPE=ET_EXEC %t/binary.yaml -o %t/non-pie.exe
+# RUN: llvm-profgen --binary=%t/non-pie.exe \
+# RUN: --perfscript=%t/non-pie.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --use-offset=0 --format=text \
+# RUN: --output=%t/non-pie.profile 2>&1 \
+# RUN: | FileCheck %s --check-prefix=NON-PIE
+# RUN: FileCheck %s --check-prefix=NON-PIE-PROFILE \
+# RUN: --input-file=%t/non-pie.profile
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/perfscript --spe-branch-profile --skip-symbolization \
+# RUN: --target-triple=x86_64-unknown-linux-gnu --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=WRONG-TRIPLE
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/perfscript --skip-symbolization --use-offset=0 \
+# RUN: --format=text --output=%t/lbr.profile
+# RUN: FileCheck %s --check-prefix=WITHOUT-OPTION --input-file=%t/lbr.profile
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --unsymbolized-profile=%t/empty.perfscript --spe-branch-profile \
+# RUN: --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=INCOMPATIBLE-INPUT
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN: --perfscript=%t/perfscript --spe-branch-profile --show-disassembly-only \
+# RUN: --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=DISASM-ONLY
+
+## perf prints an Arm SPE branch as a branch stack of one entry preceded by
+## the instruction pointer of the sample. Use different source and destination
+## addresses to verify their order, canonicalize them through the runtime
+## MMAP, and infer the executed range up to the nearest transfer. The first
+## record is given three times, because an Arm SPE trace carries no repeat
+## count. The last record is a not-taken conditional branch, which perf marks
+## with the N flag and reports with the fall-through instruction as its
+## destination.
+# NO-WARN-NOT: warning: {{.*}}branch targets do not match the binary
+# NO-WARN-NOT: warning: Invalid Arm SPE branch record
+## A trace whose records yield counters must not be reported as yielding none,
+## and the report of the branch-stack readers, which describes ranges bounded
+## by consecutive entries, must not run for records that hold one entry each.
+# NO-WARN-NOT: warning: No Arm SPE branch record yields a counter
+# NO-WARN-NOT: warning: No samples in perf script
+
+## The not-taken record contributes the range that follows it, because
+## execution did continue at the fall-through instruction, but no branch
+## counter, because no control was transferred.
+# CHECK: 2
+# CHECK-NEXT: 1000-1008:3
+# CHECK-NEXT: 100c-100c:1
+# CHECK-NEXT: 1
+# CHECK-NEXT: 1008->1000:3
+
+## The instruction pointer of a record repeats the source of its branch, so a
+## trace printed with `-F brstack` alone holds every address the record needs
+## and is read the same way.
+# NO-IP: 1
+# NO-IP-NEXT: 1000-1008:1
+# NO-IP-NEXT: 1
+# NO-IP-NEXT: 1008->1000:1
+
+## Name each malformed record once when asked for detailed warnings: a record
+## holding the instruction pointer alone, a record followed by a trailing
+## field, as a wrong -F field list such as ip,brstack,sym produces, a record
+## whose addresses carry an appended shared-object path, as adding dso to that
+## list produces, and a record whose addresses carry no 0x prefix. A trailing
+## field is read as a branch-stack entry before a record is called a deeper
+## stack, or a field holding the slash that separates the addresses of an entry
+## would abort the run and blame an --itrace option that is already right. The
+## external-to-external record is valid input but cannot contribute to this
+## binary's profile. The last record holds an instruction pointer of decimal
+## digits alone, which is a repeat count of the branch-stack grammar and has to
+## be read as the damaged record it is instead, or the record that follows it
+## would be weighed by an address.
+# INVALID: warning: Invalid Arm SPE branch record at line 2: 70000000100a
+# INVALID-NEXT: warning: Invalid Arm SPE branch record at line 3: 700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/- foo
+# INVALID-NEXT: warning: Invalid Arm SPE branch record at line 4: 700000001008 0x700000001008(/tmp/spe-branch-profile.exe)/0x700000001000(/tmp/spe-branch-profile.exe)/P/-/-/10/COND/-
+# INVALID-NEXT: warning: Invalid Arm SPE branch record at line 5: 700000001008 700000001008/700000001000/P/-/-/10/COND/-
+# INVALID-NEXT: warning: Invalid Arm SPE branch record at line 6: 700000001004
+# INVALID-NOT: warning: Invalid Arm SPE branch record
+# INVALID-NOT: error:
+
+## Without that request the lines are reported as a share instead, because a
+## trace printed with a wrong field list holds millions of such lines. Every
+## share is of the same base, the sample lines of the trace, so that the
+## reports can be read against one another. The one record that was read is
+## dropped, which leaves the profile empty and is reported as well.
+# INVALID-SUMMARY-NOT: warning: Invalid Arm SPE branch record
+# INVALID-SUMMARY: warning: 83.33%(5/6) of the sample lines of the trace are not usable Arm SPE branch records.
+# INVALID-SUMMARY-NEXT: warning: 16.67%(1/6) of the sample lines of the trace are Arm SPE branch records whose destination is not an instruction of this binary.
+# INVALID-SUMMARY-NEXT: warning: No Arm SPE branch record yields a counter for this binary!
+
+## A trace whose every line holds two bare addresses was printed with the
+## ip,addr field list of the perf versions that synthesize no branch stack for
+## Arm SPE. That is only known once the whole trace has been read. Name the
+## perf command to use, not the individual lines.
+# IP-ADDR-NOT: warning:
+# IP-ADDR: error: --spe-branch-profile requires Arm SPE branch records created with `perf script --show-mmap-events --itrace=bl1 -F ip,brstack`, but no line of the input holds a usable one
+
+## Each way a line can fail to hold a usable record leaves the records around
+## it in place: a prediction field holding the value of another field, as a
+## shifted -F list prints; an entry printed without the prediction field; an
+## entry cut short right after the prediction field, whose P may be all that is
+## left of a PN; a source address of zero, which names no instruction; an
+## instruction pointer that is not the source of the entry that follows it, so
+## that neither address can be trusted; and an instruction pointer carrying the
+## 0x prefix that perf prints on the addresses of an entry alone. The one valid
+## record still yields its counters.
+# REJECTED: warning: 85.71%(6/7) of the sample lines of the trace are not usable Arm SPE branch records.
+# REJECTED-NOT: error:
+# REJECTED-PROFILE: 1
+# REJECTED-PROFILE-NEXT: 1000-1008:1
+# REJECTED-PROFILE-NEXT: 1
+# REJECTED-PROFILE-NEXT: 1008->1000:1
+
+## The addresses of a binary that is not position independent are already the
+## ones of the trace, so they are used as they are and the missing mmap event
+## is only reported.
+# NON-PIE: warning: No relevant mmap event is matched for non-pie.exe
+# NON-PIE-PROFILE: 1
+# NON-PIE-PROFILE-NEXT: 1000-1008:1
+# NON-PIE-PROFILE-NEXT: 1
+# NON-PIE-PROFILE-NEXT: 1008->1000:1
+
+## A direct unconditional branch cannot fall through. Diagnose a recorded
+## destination at the next instruction when it is not the encoded target, and
+## keep its branch counter, because control was transferred there.
+# BAD-FALLTHROUGH: of branch targets do not match the binary.
+
+## Weigh both the mismatch and the denominator by each aggregated sample.
+# WEIGHTED-MISMATCH: warning: 66.67%(2/3) of branch targets do not match the binary.
+
+## A not-taken record must name a conditional branch and its fall-through.
+# INCONSISTENT-NOT-TAKEN: warning: 50.00%(1/2) of branch samples are marked not taken but do not name a conditional branch.
+# INCONSISTENT-NOT-TAKEN-NEXT: warning: 50.00%(1/2) of branch targets do not match the binary.
+
+## No transfer follows the destination of the same record, so its range runs to
+## the last instruction of the binary, where it ends for want of a following
+## instruction.
+# LAST-INST: 1
+# LAST-INST-NEXT: 1010-1014:1
+# LAST-INST-NEXT: 1
+# LAST-INST-NEXT: 100c->1010:1
+
+## A trace that holds no Arm SPE record at all is the wrong input file. Fail
+## instead of writing an empty profile and reporting success, and say which of
+## the two reasons applies: here the trace holds no line to read at all.
+# EMPTY: error: --spe-branch-profile requires Arm SPE branch records created with `perf script --show-mmap-events --itrace=bl1 -F ip,brstack`, but the trace holds no record at all
+
+## A branch from this binary to external code is valid input, but a single SPE
+## record with an external target cannot produce a branch or range counter.
+# NO-USABLE: warning: 100.00%(1/1) of the sample lines of the trace are Arm SPE branch records whose destination is not an instruction of this binary.
+# NO-USABLE-NEXT: warning: No Arm SPE branch record yields a counter for this binary!
+
+## A record whose destination is not an instruction of this binary is dropped:
+## it reports a branch leaving the binary, a trace of another process or of
+## another build of the program, or a record whose branch target address was
+## lost. Report the share they take of the sample lines of the trace, so that a
+## profile thinned by a trace that does not belong to this binary is not
+## mistaken for a complete one.
+# PARTIAL: warning: 66.67%(2/3) of the sample lines of the trace are Arm SPE branch records whose destination is not an instruction of this binary.
+# PARTIAL-PROFILE: 1
+# PARTIAL-PROFILE-NEXT: 1000-1008:1
+# PARTIAL-PROFILE-NEXT: 1
+# PARTIAL-PROFILE-NEXT: 1008->1000:1
+
+## Only an mmap event of a perf script trace carries a process id, so --pid
+## selects the mapping that addresses are canonicalized against and does not
+## filter the records themselves. Another id therefore leaves every record of
+## this trace pointing outside the binary.
+# OTHER-PID: warning: No relevant mmap event is matched for
+# OTHER-PID: warning: 100.00%(4/4) of the sample lines of the trace are Arm SPE branch records whose destination is not an instruction of this binary.
+# OTHER-PID: warning: No Arm SPE branch record yields a counter for this binary!
+
+## The records of an Arm SPE trace are the branch-stack lines of an LBR trace,
+## holding one entry each, so the same trace is also valid input without the
+## option. It is then read as an LBR trace, which yields the branch counters
+## alone: no range is inferred, because where the range that follows the most
+## recent branch ends is what only the Arm SPE reader recovers from the binary.
+## The N flag that marks a branch as not taken is read by the Arm SPE reader
+## alone, because a hardware branch stack records taken branches only, so the
+## LBR reader counts the not-taken record as an ordinary branch here.
+# WITHOUT-OPTION: 0
+# WITHOUT-OPTION-NEXT: 2
+# WITHOUT-OPTION-NEXT: 1008->1000:3
+# WITHOUT-OPTION-NEXT: 1008->100c:1
+
+## Preserve an external source for a branch entering this binary so downstream
+## symbolization can recognize it as a function-entry sample.
+# EXTERNAL-SOURCE: 1
+# EXTERNAL-SOURCE-NEXT: 1000-1008:1
+# EXTERNAL-SOURCE-NEXT: 1
+# EXTERNAL-SOURCE-NEXT: 1->1000:1
+
+## Symbolization converts the external-to-internal branch into a head sample.
+# SYMBOLIZED: foo:0:1{{$}}
+
+## Without a matching mmap event every address is interpreted against the
+## preferred base address, which leaves all records external and the profile
+## empty. Report the missing event as the cause instead of the empty profile
+## alone.
+# NO-MMAP: warning: No relevant mmap event is matched for
+# NO-MMAP: warning: No Arm SPE branch record yields a counter for this binary!
+
+## A record decoded with a deeper branch stack holds more than the sampled
+## branch: further sampled branches, or the entry perf synthesizes from the
+## target of the branch that preceded the sampled operation, which it prints
+## with the source address 0x0 and which, counted as a branch, fabricates one
+## from outside the binary. Name the perf command that decodes one entry per
+## record instead of accepting it.
+# DEEP-STACK-NOT: warning:
+# DEEP-STACK: error: --spe-branch-profile requires Arm SPE branch records created with `perf script --show-mmap-events --itrace=bl1 -F ip,brstack`, but the record at line 2 holds a branch stack of more than one entry
+
+## A single unreadable line inside a trace whose other records parse is damaged
+## input, not input of the wrong kind, whichever order the lines come in. Warn
+## about that line alone and keep the records around it.
+# DAMAGED: warning: Invalid Arm SPE branch record at line 3: 700000001008 0x700000001008/0x
+# DAMAGED-NOT: error:
+# DAMAGED-PROFILE: 2
+# DAMAGED-PROFILE-NEXT: 1000-1008:1
+# DAMAGED-PROFILE-NEXT: 100c-100c:1
+# DAMAGED-PROFILE-NEXT: 1
+# DAMAGED-PROFILE-NEXT: 1008->1000:1
+
+## The call stack frames of a hybrid trace hold no branch-stack entry, so they
+## are named as unreadable lines, and the deeper branch stack that follows them
+## is rejected the way any deeper stack is.
+# HYBRID: warning: Invalid Arm SPE branch record at line 2: 4006ac
+# HYBRID-NEXT: warning: Invalid Arm SPE branch record at line 3: 40064c
+# HYBRID: error: --spe-branch-profile requires Arm SPE branch records created with `perf script --show-mmap-events --itrace=bl1 -F ip,brstack`, but the record at line 4 holds a branch stack of more than one entry
+
+## The check uses the triple llvm-profgen was told to use, so an override that
+## names another target is rejected as well.
+# WRONG-TRIPLE: error: --spe-branch-profile requires an AArch64 binary.
+
+## Name the conflicting mode instead of the input the user already gave.
+# INCOMPATIBLE-INPUT: error: --spe-branch-profile requires --perfscript or --perfdata input.
+# DISASM-ONLY: error: --spe-branch-profile cannot be used with --show-disassembly-only.
+
+#--- binary.yaml
+--- !ELF
+FileHeader:
+ Class: ELFCLASS64
+ Data: ELFDATA2LSB
+ ## Built a second time as ET_EXEC to cover a binary that is not position
+ ## independent.
+ Type: [[TYPE=ET_DYN]]
+ Machine: EM_AARCH64
+ Entry: 0x1000
+Sections:
+ - Name: .text
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x1000
+ Offset: 0x1000
+ AddressAlign: 0x4
+ ## nop
+ ## nop
+ ## b.eq 0x1000
+ ## b 0x1000
+ ## nop
+ ## nop
+ Content: 1F2003D51F2003D5C0FFFF54FDFFFF171F2003D51F2003D5
+ProgramHeaders:
+ - Type: PT_LOAD
+ Flags: [ PF_X, PF_R ]
+ Offset: 0
+ VAddr: 0
+ FileSize: 0x1018
+ MemSize: 0x1018
+ Align: 0x10000
+Symbols:
+ - Name: foo
+ Type: STT_FUNC
+ Section: .text
+ Binding: STB_GLOBAL
+ Value: 0x1000
+ Size: 0x18
+DWARF:
+ debug_abbrev:
+ - ID: 0
+ Table:
+ - Code: 1
+ Tag: DW_TAG_compile_unit
+ Children: DW_CHILDREN_yes
+ Attributes:
+ - Attribute: DW_AT_name
+ Form: DW_FORM_string
+ - Code: 2
+ Tag: DW_TAG_subprogram
+ Children: DW_CHILDREN_no
+ Attributes:
+ - Attribute: DW_AT_name
+ Form: DW_FORM_string
+ - Attribute: DW_AT_low_pc
+ Form: DW_FORM_addr
+ - Attribute: DW_AT_high_pc
+ Form: DW_FORM_data8
+ debug_info:
+ - Version: 4
+ AbbrevTableID: 0
+ AddrSize: 8
+ Entries:
+ - AbbrCode: 1
+ Values:
+ - CStr: spe-branch-profile.c
+ - AbbrCode: 2
+ Values:
+ - CStr: foo
+ - Value: 0x1000
+ - Value: 0x18
+ - AbbrCode: 0
+
+#--- perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+ 700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+ 700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+ 700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+700000001008 0x700000001008/0x70000000100c/PN/-/-/10/COND/-
+#--- no-ip.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+0x700000001008/0x700000001000/P/-/-/10/COND/-
+#--- malformed.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+70000000100a
+700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/- foo
+700000001008 0x700000001008(/tmp/spe-branch-profile.exe)/0x700000001000(/tmp/spe-branch-profile.exe)/P/-/-/10/COND/-
+700000001008 700000001008/700000001000/P/-/-/10/COND/-
+700000001004
+700000003000 0x700000003000/0x700000003004/P/-/-/10/COND/-
+#--- bad-fallthrough.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+70000000100c 0x70000000100c/0x700000001010/P/-/-/10/UNCOND/-
+#--- weighted-mismatch.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+70000000100c 0x70000000100c/0x700000001000/P/-/-/10/UNCOND/-
+70000000100c 0x70000000100c/0x700000001010/P/-/-/10/UNCOND/-
+70000000100c 0x70000000100c/0x700000001010/P/-/-/10/UNCOND/-
+#--- inconsistent-not-taken.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+70000000100c 0x70000000100c/0x700000001010/PN/-/-/10/UNCOND/-
+700000001008 0x700000001008/0x700000001000/PN/-/-/10/COND/-
+#--- empty.perfscript
+#--- unusable.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000001008 0x700000001008/0x700000003000/P/-/-/10/COND/-
+#--- partial.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000001008 0x700000001008/0x700000003000/P/-/-/10/COND/-
+700000001008 0x700000001008/0x0/P/-/-/10/COND/-
+700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+#--- external-source.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000003000 0x700000003000/0x700000001000/P/-/-/10/CALL/-
+#--- no-mmap.perfscript
+700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+#--- previous-target.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/- 0x0/0x70000000100c/-/-/-/0//-
+#--- lbr.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000001008 0x700000001008/0x700000001000/P/-/-/0 0x70000000100c/0x700000001000/P/-/-/0 0x700000001008/0x70000000100c/P/-/-/0
+#--- damaged-line.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+700000001008 0x700000001008/0x
+700000001008 0x700000001008/0x70000000100c/PN/-/-/10/COND/-
+#--- hybrid.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+4006ac
+40064c
+70000000100c 0x70000000100c/0x700000001000/P/-/-/0 0x700000001008/0x70000000100c/P/-/-/0
+#--- ip-addr.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000001000 700000001008
+70000000100c 700000001008
+#--- rejected.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000001008 0x700000001008/0x700000001000/COND/-/-/10
+700000001008 0x700000001008/0x700000001000
+700000001008 0x700000001008/0x700000001000/P
+0 0x0/0x700000001000/P/-/-/10/COND/-
+700000001000 0x700000001008/0x700000001000/P/-/-/10/COND/-
+0x700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+#--- non-pie.perfscript
+1008 0x1008/0x1000/P/-/-/10/COND/-
diff --git a/llvm/test/tools/llvm-profgen/AArch64/spe-not-taken-branch.test b/llvm/test/tools/llvm-profgen/AArch64/spe-not-taken-branch.test
new file mode 100644
index 0000000000000..d6295707b0aad
--- /dev/null
+++ b/llvm/test/tools/llvm-profgen/AArch64/spe-not-taken-branch.test
@@ -0,0 +1,232 @@
+## Arm SPE records a conditional branch that was not taken with the
+## instruction that follows it as its destination, and perf marks such an entry
+## with the N flag. Such a record transferred no control, so it must not be
+## counted as a branch, while the range that follows it did execute and is
+## counted. The binary below places the entry of a function right behind a
+## conditional branch, where counting the branch would invent a call of that
+## function and a head sample for it. The same two addresses without the flag
+## report a branch that was taken to the instruction that follows it, which is
+## a transfer and is counted, so the flag alone tells the two apart.
+
+# RUN: split-file %s %t
+# RUN: yaml2obj %t/binary.yaml -o %t/spe-not-taken-branch.exe
+# RUN: llvm-profgen --binary=%t/spe-not-taken-branch.exe \
+# RUN: --perfscript=%t/fall-through.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --use-offset=0 --format=text \
+# RUN: --output=%t/fall-through.profile 2>&1 \
+# RUN: | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --check-prefix=FALL-THROUGH \
+# RUN: --input-file=%t/fall-through.profile
+# RUN: llvm-profgen --binary=%t/spe-not-taken-branch.exe \
+# RUN: --perfscript=%t/fall-through.perfscript --spe-branch-profile \
+# RUN: --use-offset=0 --format=text --output=%t/fall-through.symbolized
+# RUN: FileCheck %s --check-prefix=HEAD-SAMPLE \
+# RUN: --input-file=%t/fall-through.symbolized
+# RUN: llvm-profgen --binary=%t/spe-not-taken-branch.exe \
+# RUN: --perfscript=%t/indirect.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --use-offset=0 --format=text \
+# RUN: --output=%t/indirect.profile 2>&1 \
+# RUN: | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --check-prefix=INDIRECT --input-file=%t/indirect.profile
+# RUN: llvm-profgen --binary=%t/spe-not-taken-branch.exe \
+# RUN: --perfscript=%t/taken.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --use-offset=0 --format=text \
+# RUN: --output=%t/taken.profile 2>&1 \
+# RUN: | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --check-prefix=TAKEN --input-file=%t/taken.profile
+# RUN: llvm-profgen --binary=%t/spe-not-taken-branch.exe \
+# RUN: --perfscript=%t/taken.perfscript --spe-branch-profile \
+# RUN: --use-offset=0 --format=text --output=%t/taken.symbolized
+# RUN: FileCheck %s --check-prefix=TAKEN-HEAD --input-file=%t/taken.symbolized
+# RUN: llvm-profgen --binary=%t/spe-not-taken-branch.exe \
+# RUN: --perfscript=%t/both.perfscript --spe-branch-profile \
+# RUN: --skip-symbolization --use-offset=0 --format=text \
+# RUN: --output=%t/both.profile 2>&1 \
+# RUN: | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --check-prefix=BOTH --input-file=%t/both.profile
+
+## A destination that the instruction does not encode is expected of a branch
+## that was not taken, so it is not reported as a mismatch either.
+# NO-WARN-NOT: warning:
+
+## Both records start their range at the entry of bar, so the range is counted
+## twice, but only the call transferred control there.
+# FALL-THROUGH: 1
+# FALL-THROUGH-NEXT: 1008-100c:2
+# FALL-THROUGH-NEXT: 1
+# FALL-THROUGH-NEXT: 1000->1008:1
+
+## The head count of bar is the number of calls of it, one, and not the number
+## of records whose destination is its entry, two.
+# HEAD-SAMPLE: bar:0:1{{$}}
+
+## An indirect branch encodes no target at all, so its recorded destination is
+## never the one of a branch that was not taken, even where it is the
+## instruction that follows the branch. Count it as the transfer it is. The
+## branch-type field of the record reads RET because that is what the kernel
+## maps an indirect branch to; this reader takes the kind of the branch from
+## the binary and ignores that field. Its prediction field reads M, for a
+## branch perf found mispredicted, which is counted like one it found
+## predicted: only the N suffix of that field changes what a record means.
+# INDIRECT: 1
+# INDIRECT-NEXT: 1014-1014:1
+# INDIRECT-NEXT: 1
+# INDIRECT-NEXT: 1010->1014:1
+
+## A conditional branch whose encoded target is the instruction that follows it
+## does transfer control when it is taken, and the record carries no N flag.
+## Count the branch, so that the function starting at that instruction keeps
+## the head sample of the transfer.
+# TAKEN: 1
+# TAKEN-NEXT: 101c-101c:1
+# TAKEN-NEXT: 1
+# TAKEN-NEXT: 1018->101c:1
+# TAKEN-HEAD: quux:0:1{{$}}
+
+## The same two addresses stand for a taken and for a not-taken branch, which
+## only the flag tells apart, so aggregation must keep the two records separate.
+## Both ranges executed; only the taken record transferred control. The flag is
+## read in both of its spellings: PN above, for a branch perf found predicted,
+## and MN here, for one it found mispredicted.
+# BOTH: 1
+# BOTH-NEXT: 101c-101c:2
+# BOTH-NEXT: 1
+# BOTH-NEXT: 1018->101c:1
+
+#--- binary.yaml
+--- !ELF
+FileHeader:
+ Class: ELFCLASS64
+ Data: ELFDATA2LSB
+ Type: ET_DYN
+ Machine: EM_AARCH64
+ Entry: 0x1000
+Sections:
+ - Name: .text
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x1000
+ Offset: 0x1000
+ AddressAlign: 0x4
+ ## foo:
+ ## bl 0x1008
+ ## b.eq 0x1000
+ ## bar:
+ ## nop
+ ## ret
+ ## baz:
+ ## br x0
+ ## ret
+ ## qux:
+ ## b.eq 0x101c
+ ## quux:
+ ## ret
+ Content: 02000094E0FFFF541F2003D5C0035FD600001FD6C0035FD620000054C0035FD6
+ProgramHeaders:
+ - Type: PT_LOAD
+ Flags: [ PF_X, PF_R ]
+ Offset: 0
+ VAddr: 0
+ FileSize: 0x1020
+ MemSize: 0x1020
+ Align: 0x10000
+Symbols:
+ - Name: foo
+ Type: STT_FUNC
+ Section: .text
+ Binding: STB_GLOBAL
+ Value: 0x1000
+ Size: 0x8
+ - Name: bar
+ Type: STT_FUNC
+ Section: .text
+ Binding: STB_GLOBAL
+ Value: 0x1008
+ Size: 0x8
+ - Name: baz
+ Type: STT_FUNC
+ Section: .text
+ Binding: STB_GLOBAL
+ Value: 0x1010
+ Size: 0x8
+ - Name: qux
+ Type: STT_FUNC
+ Section: .text
+ Binding: STB_GLOBAL
+ Value: 0x1018
+ Size: 0x4
+ - Name: quux
+ Type: STT_FUNC
+ Section: .text
+ Binding: STB_GLOBAL
+ Value: 0x101c
+ Size: 0x4
+DWARF:
+ debug_abbrev:
+ - ID: 0
+ Table:
+ - Code: 1
+ Tag: DW_TAG_compile_unit
+ Children: DW_CHILDREN_yes
+ Attributes:
+ - Attribute: DW_AT_name
+ Form: DW_FORM_string
+ - Code: 2
+ Tag: DW_TAG_subprogram
+ Children: DW_CHILDREN_no
+ Attributes:
+ - Attribute: DW_AT_name
+ Form: DW_FORM_string
+ - Attribute: DW_AT_low_pc
+ Form: DW_FORM_addr
+ - Attribute: DW_AT_high_pc
+ Form: DW_FORM_data8
+ debug_info:
+ - Version: 4
+ AbbrevTableID: 0
+ AddrSize: 8
+ Entries:
+ - AbbrCode: 1
+ Values:
+ - CStr: spe-not-taken-branch.c
+ - AbbrCode: 2
+ Values:
+ - CStr: foo
+ - Value: 0x1000
+ - Value: 0x8
+ - AbbrCode: 2
+ Values:
+ - CStr: bar
+ - Value: 0x1008
+ - Value: 0x8
+ - AbbrCode: 2
+ Values:
+ - CStr: baz
+ - Value: 0x1010
+ - Value: 0x8
+ - AbbrCode: 2
+ Values:
+ - CStr: qux
+ - Value: 0x1018
+ - Value: 0x4
+ - AbbrCode: 2
+ Values:
+ - CStr: quux
+ - Value: 0x101c
+ - Value: 0x4
+ - AbbrCode: 0
+
+#--- fall-through.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-not-taken-branch.exe
+700000001004 0x700000001004/0x700000001008/PN/-/-/10/COND/-
+700000001000 0x700000001000/0x700000001008/P/-/-/10/CALL/-
+#--- indirect.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-not-taken-branch.exe
+700000001010 0x700000001010/0x700000001014/M/-/-/10/RET/-
+#--- taken.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-not-taken-branch.exe
+700000001018 0x700000001018/0x70000000101c/P/-/-/10/COND/-
+#--- both.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-not-taken-branch.exe
+700000001018 0x700000001018/0x70000000101c/P/-/-/10/COND/-
+700000001018 0x700000001018/0x70000000101c/MN/-/-/10/COND/-
diff --git a/llvm/test/tools/llvm-profgen/AArch64/spe-perfdata.test b/llvm/test/tools/llvm-profgen/AArch64/spe-perfdata.test
new file mode 100644
index 0000000000000..98f1544fc6a9b
--- /dev/null
+++ b/llvm/test/tools/llvm-profgen/AArch64/spe-perfdata.test
@@ -0,0 +1,135 @@
+# REQUIRES: system-linux
+
+# RUN: split-file %s %t
+# RUN: yaml2obj %t/binary.yaml -o %t/spe-perfdata.exe
+# RUN: touch %t/perf.data
+# RUN: chmod +x %t/success/perf
+# RUN: chmod +x %t/no-record/perf
+# RUN: env PATH="%t/success%{pathsep}%{PATH}" \
+# RUN: llvm-profgen --binary=%t/spe-perfdata.exe \
+# RUN: --perfdata=%t/perf.data --spe-branch-profile --skip-symbolization \
+# RUN: --use-offset=0 --format=text --output=%t/profile
+# RUN: FileCheck %s --input-file=%t/profile
+# RUN: env PATH="%t/no-record%{pathsep}%{PATH}" \
+# RUN: not llvm-profgen --binary=%t/spe-perfdata.exe \
+# RUN: --perfdata=%t/perf.data --spe-branch-profile --skip-symbolization \
+# RUN: --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s --check-prefix=NO-RECORD
+
+## The mock validates both perf invocations, including the decoding of branch
+## events into a branch stack of one entry and the field list required for Arm
+## SPE. The profile checks that the second invocation's output is consumed.
+# CHECK: 1
+# CHECK-NEXT: 1000-1008:1
+# CHECK-NEXT: 1
+# CHECK-NEXT: 1008->1000:1
+
+## A recording that holds no Arm SPE branch event yields a trace of mmap
+## events alone. The perf command that printed it is one this tool ran, so
+## name what has to be recorded rather than what has to be printed.
+# NO-RECORD: error: --spe-branch-profile found no Arm SPE branch record in the perf data; record one with arm_spe/branch_filter=1,event_filter=2/ using perf 6.15 or later
+
+#--- success/perf
+#!/bin/sh
+
+if [ "$1" != "script" ]; then
+ echo "unexpected perf arguments: $*" >&2
+ exit 2
+fi
+
+if [ "$2" = "--show-mmap-events" ]; then
+ if [ "$3" != "-F" ] || [ "$4" != "comm,pid" ] || [ "$5" != "-i" ] || \
+ [ ! -f "$6" ] || [ "$#" -ne 6 ]; then
+ echo "unexpected mmap arguments: $*" >&2
+ exit 2
+ fi
+ printf '%s\n' 'PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-perfdata.exe'
+elif [ "$2" = "--itrace=bl1" ]; then
+ if [ "$3" != "--show-mmap-events" ] || [ "$4" != "-F" ] || \
+ [ "$5" != "ip,brstack" ] || [ "$6" != "-i" ] || [ ! -f "$7" ] || \
+ [ "$8" != "--pid" ] || [ "$9" != "1" ] || [ "$#" -ne 9 ]; then
+ echo "unexpected sample arguments: $*" >&2
+ exit 2
+ fi
+ printf '%s\n' 'PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-perfdata.exe'
+ printf '%s\n' '700000001008 0x700000001008/0x700000001000/P/-/-/10/UNCOND/-'
+else
+ echo "unexpected perf arguments: $*" >&2
+ exit 2
+fi
+
+#--- no-record/perf
+#!/bin/sh
+
+## A recording without Arm SPE branch events: the mmap events are printed by
+## both invocations and no branch record by either.
+printf '%s\n' 'PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-perfdata.exe'
+
+#--- binary.yaml
+--- !ELF
+FileHeader:
+ Class: ELFCLASS64
+ Data: ELFDATA2LSB
+ Type: ET_DYN
+ Machine: EM_AARCH64
+ Entry: 0x1000
+Sections:
+ - Name: .text
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x1000
+ Offset: 0x1000
+ AddressAlign: 0x4
+ ## nop
+ ## nop
+ ## b 0x1000
+ Content: 1F2003D51F2003D5FEFFFF17
+ProgramHeaders:
+ - Type: PT_LOAD
+ Flags: [ PF_X, PF_R ]
+ Offset: 0
+ VAddr: 0
+ FileSize: 0x100C
+ MemSize: 0x100C
+ Align: 0x10000
+Symbols:
+ - Name: foo
+ Type: STT_FUNC
+ Section: .text
+ Binding: STB_GLOBAL
+ Value: 0x1000
+ Size: 0xC
+DWARF:
+ debug_abbrev:
+ - ID: 0
+ Table:
+ - Code: 1
+ Tag: DW_TAG_compile_unit
+ Children: DW_CHILDREN_yes
+ Attributes:
+ - Attribute: DW_AT_name
+ Form: DW_FORM_string
+ - Code: 2
+ Tag: DW_TAG_subprogram
+ Children: DW_CHILDREN_no
+ Attributes:
+ - Attribute: DW_AT_name
+ Form: DW_FORM_string
+ - Attribute: DW_AT_low_pc
+ Form: DW_FORM_addr
+ - Attribute: DW_AT_high_pc
+ Form: DW_FORM_data8
+ debug_info:
+ - Version: 4
+ AbbrevTableID: 0
+ AddrSize: 8
+ Entries:
+ - AbbrCode: 1
+ Values:
+ - CStr: spe-perfdata.c
+ - AbbrCode: 2
+ Values:
+ - CStr: foo
+ - Value: 0x1000
+ - Value: 0xC
+ - AbbrCode: 0
diff --git a/llvm/test/tools/llvm-profgen/AArch64/spe-range-gap.test b/llvm/test/tools/llvm-profgen/AArch64/spe-range-gap.test
new file mode 100644
index 0000000000000..fac0e4d8b10a6
--- /dev/null
+++ b/llvm/test/tools/llvm-profgen/AArch64/spe-range-gap.test
@@ -0,0 +1,774 @@
+# RUN: split-file %s %t
+# RUN: yaml2obj %t/binary.yaml -o %t/spe-range-gap.exe
+# RUN: llvm-profgen --binary=%t/spe-range-gap.exe \
+# RUN: --perfscript=%t/perfscript --spe-branch-profile --skip-symbolization \
+# RUN: --use-offset=0 --format=text --output=%t/profile 2>&1 \
+# RUN: | FileCheck %s --check-prefix=SHRUNK
+# RUN: FileCheck %s --input-file=%t/profile
+# RUN: yaml2obj %t/no-debug.yaml -o %t/no-debug.exe
+# RUN: llvm-profgen --binary=%t/no-debug.exe \
+# RUN: --perfscript=%t/no-debug.perfscript --spe-branch-profile --skip-symbolization \
+# RUN: --use-offset=0 --format=text --output=%t/no-debug-profile 2>&1 \
+# RUN: | FileCheck %s --check-prefix=ALL-SHRUNK
+# RUN: FileCheck %s --check-prefix=NO-DEBUG --input-file=%t/no-debug-profile
+# RUN: yaml2obj %t/unordered.yaml -o %t/unordered.exe
+# RUN: llvm-profgen --binary=%t/unordered.exe \
+# RUN: --perfscript=%t/unordered.perfscript --spe-branch-profile --skip-symbolization \
+# RUN: --use-offset=0 --format=text --output=%t/unordered-profile 2>&1 \
+# RUN: | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --check-prefix=UNORDERED --input-file=%t/unordered-profile
+
+## A blank line between records is not a record, so it is not diagnosed. The
+## one record whose destination lies in code that debug info does not describe
+## is reported as a share, because its range covers a single instruction rather
+## than the code that ran.
+# SHRUNK-NOT: warning: Invalid
+# SHRUNK: warning: 7.69%(1/13) of the sample lines of the trace are Arm SPE branch records whose destination cannot be associated with a function, so only that instruction is profiled.
+
+## A binary that carries no debug info at all has every range reduced that way,
+## which the share states outright instead of leaving it to be inferred from a
+## profile of one-instruction ranges. The trace holds the same record twice, so
+## the two lines aggregate into one sample of weight two: the share counts the
+## lines of the trace, not the samples that remained after aggregation.
+# ALL-SHRUNK: warning: 100.00%(2/2) of the sample lines of the trace are Arm SPE branch records whose destination cannot be associated with a function, so only that instruction is profiled.
+
+## Ordering the instruction addresses is what keeps the two functions apart, so
+## neither range is reduced and nothing is reported.
+# NO-WARN-NOT: warning:
+
+## Each record below is checked by one paragraph, in address order of the
+## inferred range.
+##
+## The first record starts a range in a function that is adjacent to the next
+## function. Stop at the function boundary rather than at the transfer that
+## follows it.
+##
+## The second record starts a range in a function that is separated from the
+## following transfer by an address gap. Stop at the gap.
+##
+## The third record starts a range in a function that debug info describes as
+## two adjacent ranges. No symbol starts the second range, so it continues the
+## same function and execution runs through the boundary, reaching the transfer
+## in the second range instead of stopping at the end of the first. A symbol
+## that starts a function at the later range, as a split-out part of a function
+## carries, ends the range there instead; the fifth record covers that.
+##
+## The fourth record starts a range in a function whose two debug info ranges
+## are separated by an address gap. Stop at the gap even though both sides
+## belong to one function.
+##
+## The fifth record starts a range in a function that is adjacent to a
+## different function of the same name, as produced by static functions of
+## different compilation units. Debug info groups both under one name, so stop
+## at the entry of the second function rather than run into it.
+##
+## The sixth record has an external source and a destination in code that no
+## debug info describes. Reduce the range to the destination instead of
+## extending it into the next described function, and still count the branch.
+##
+## The seventh and the eighth record start ranges that a call and a return end,
+## which transfer execution just as a branch does.
+##
+## The ninth record starts a range in a described function that is immediately
+## followed by code no debug info describes. Stop at the last instruction of
+## the function rather than continue into code of an unknown function.
+##
+## The tenth record starts a range in a function that is adjacent to a
+## differently named function whose first range carries no symbol, as a part
+## split out of a function does. The change of function is what ends the range
+## there, because a range that no symbol starts is not marked as an entry.
+##
+## The eleventh record starts a range in a function that holds a trap. The trap
+## transfers control nowhere, so it is no branch, call or return, but execution
+## never continues past it, and the range ends there instead of running on to
+## the branch that ends the function.
+##
+## The twelfth record is a branch to itself, so its range starts at the branch
+## and the branch ends it, leaving a range of that one instruction.
+##
+## The thirteenth record is the back edge of a loop whose body holds a branch
+## of its own. The range starts at the top of the body and ends at that inner
+## branch, because what follows it ran only if it was not taken, which the
+## record does not say. A back edge therefore does not attest that the whole
+## body ran.
+# CHECK: 13
+# CHECK-NEXT: 500-500:1
+# CHECK-NEXT: 1000-1004:1
+# CHECK-NEXT: 2000-2004:1
+# CHECK-NEXT: 4000-400c:1
+# CHECK-NEXT: 6000-6004:1
+# CHECK-NEXT: 8000-8004:1
+# CHECK-NEXT: 9000-9004:1
+# CHECK-NEXT: a000-a004:1
+# CHECK-NEXT: b000-b004:1
+# CHECK-NEXT: c000-c004:1
+# CHECK-NEXT: d800-d804:1
+# CHECK-NEXT: d900-d900:1
+# CHECK-NEXT: da00-da04:1
+# CHECK-NEXT: 13
+# CHECK-NEXT: 1->500:1
+# CHECK-NEXT: 100c->1000:1
+# CHECK-NEXT: 3000->2000:1
+# CHECK-NEXT: 400c->4000:1
+# CHECK-NEXT: 7004->6000:1
+# CHECK-NEXT: 800c->8000:1
+# CHECK-NEXT: 900c->9000:1
+# CHECK-NEXT: a00c->a000:1
+# CHECK-NEXT: b00c->b000:1
+# CHECK-NEXT: c00c->c000:1
+# CHECK-NEXT: d80c->d800:1
+# CHECK-NEXT: d900->d900:1
+# CHECK-NEXT: da0c->da00:1
+
+## Function ranges come from debug info, which profiled binaries are expected
+## to carry. Where they are missing, an inferred range must not run past an
+## unknown function boundary, so it is reduced to the sampled instruction while
+## the branch itself is still counted. Both counters weigh two, because the
+## trace holds the record twice.
+# NO-DEBUG: 1
+# NO-DEBUG-NEXT: 1000-1000:2
+# NO-DEBUG-NEXT: 1
+# NO-DEBUG-NEXT: 1008->1000:2
+
+## Instructions are collected in section-header order, which an object file
+## need not keep in address order. Range inference orders them itself, so each
+## sampled branch destination is bounded by the branch that ends the function it
+## belongs to. Both functions are sampled: without that ordering the backward
+## walk over instructions loses the function whose section is listed first, and
+## its range would shrink to the sampled instruction alone.
+# UNORDERED: 2
+# UNORDERED-NEXT: 1000-1004:1
+# UNORDERED-NEXT: 2000-2004:1
+# UNORDERED-NEXT: 2
+# UNORDERED-NEXT: 1004->1000:1
+# UNORDERED-NEXT: 2004->2000:1
+
+#--- binary.yaml
+--- !ELF
+FileHeader:
+ Class: ELFCLASS64
+ Data: ELFDATA2LSB
+ Type: ET_DYN
+ Machine: EM_AARCH64
+ Entry: 0x1000
+Sections:
+ ## Code that no debug info describes. The symbol makes it disassembled, and
+ ## the described functions that follow provide range ends that must not be
+ ## reached from here.
+ - Name: .text.undescribed
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x500
+ Offset: 0x500
+ AddressAlign: 0x4
+ ## nop
+ ## nop
+ Content: 1F2003D51F2003D5
+ - Name: .text.first
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x1000
+ Offset: 0x1000
+ AddressAlign: 0x4
+ ## nop
+ ## nop
+ Content: 1F2003D51F2003D5
+ - Name: .text.adjacent
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x1008
+ Offset: 0x1008
+ AddressAlign: 0x4
+ ## nop
+ ## b 0x1000
+ Content: 1F2003D5FDFFFF17
+ - Name: .text.gap
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x2000
+ Offset: 0x2000
+ AddressAlign: 0x4
+ ## nop
+ ## nop
+ Content: 1F2003D51F2003D5
+ - Name: .text.transfer
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x3000
+ Offset: 0x3000
+ AddressAlign: 0x4
+ ## b 0x2000
+ Content: 00FCFF17
+ - Name: .text.split
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x4000
+ Offset: 0x4000
+ AddressAlign: 0x4
+ ## nop
+ ## nop
+ ## nop
+ ## b 0x4000
+ Content: 1F2003D51F2003D51F2003D5FDFFFF17
+ - Name: .text.split-gap-head
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x6000
+ Offset: 0x6000
+ AddressAlign: 0x4
+ ## nop
+ ## nop
+ Content: 1F2003D51F2003D5
+ - Name: .text.split-gap-tail
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x7000
+ Offset: 0x7000
+ AddressAlign: 0x4
+ ## nop
+ ## b 0x6000
+ Content: 1F2003D5FFFBFF17
+ - Name: .text.duplicate
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x8000
+ Offset: 0x8000
+ AddressAlign: 0x4
+ ## nop
+ ## nop
+ ## nop
+ ## b 0x8000
+ Content: 1F2003D51F2003D51F2003D5FDFFFF17
+ - Name: .text.call
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x9000
+ Offset: 0x9000
+ AddressAlign: 0x4
+ ## nop
+ ## bl 0x9000
+ ## nop
+ ## b 0x9000
+ Content: 1F2003D5FFFFFF971F2003D5FDFFFF17
+ - Name: .text.ret
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0xA000
+ Offset: 0xA000
+ AddressAlign: 0x4
+ ## nop
+ ## ret
+ ## nop
+ ## b 0xa000
+ Content: 1F2003D5C0035FD61F2003D5FDFFFF17
+ - Name: .text.tail
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0xB000
+ Offset: 0xB000
+ AddressAlign: 0x4
+ ## nop
+ ## nop
+ Content: 1F2003D51F2003D5
+ ## Code that no debug info describes, adjacent to the described function
+ ## above, so that only the unknown function can end the inferred range.
+ - Name: .text.tail-undescribed
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0xB008
+ Offset: 0xB008
+ AddressAlign: 0x4
+ ## nop
+ ## b 0xb000
+ Content: 1F2003D5FDFFFF17
+ - Name: .text.func-change
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0xC000
+ Offset: 0xC000
+ AddressAlign: 0x4
+ ## nop
+ ## nop
+ ## nop
+ ## b 0xc000
+ Content: 1F2003D51F2003D51F2003D5FDFFFF17
+ - Name: .text.func-change-tail
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0xD000
+ Offset: 0xD000
+ AddressAlign: 0x4
+ ## nop
+ ## nop
+ Content: 1F2003D51F2003D5
+ ## A trap in the middle of a function. It transfers control nowhere, so it is
+ ## no branch, call or return, but execution never continues past it either.
+ - Name: .text.trap
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0xD800
+ Offset: 0xD800
+ AddressAlign: 0x4
+ ## nop
+ ## brk #0
+ ## nop
+ ## b 0xd800
+ Content: 1F2003D5000020D41F2003D5FDFFFF17
+ ## A branch to itself, which is both the source and the destination of the
+ ## record below.
+ - Name: .text.self
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0xD900
+ Offset: 0xD900
+ AddressAlign: 0x4
+ ## b 0xd900
+ Content: 00000014
+ ## A loop whose body holds a branch of its own, so that the range inferred
+ ## from a record of the loop back edge covers a part of the body only.
+ - Name: .text.inner
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0xDA00
+ Offset: 0xDA00
+ AddressAlign: 0x4
+ ## nop
+ ## b.eq 0xda00
+ ## nop
+ ## b 0xda00
+ Content: 1F2003D5E0FFFF541F2003D5FDFFFF17
+ProgramHeaders:
+ - Type: PT_LOAD
+ Flags: [ PF_X, PF_R ]
+ Offset: 0
+ VAddr: 0
+ FileSize: 0xDA10
+ MemSize: 0xDA10
+ Align: 0x10000
+Symbols:
+ - Name: undescribed
+ Type: STT_FUNC
+ Section: .text.undescribed
+ Binding: STB_GLOBAL
+ Value: 0x500
+ Size: 0x8
+ - Name: first
+ Type: STT_FUNC
+ Section: .text.first
+ Binding: STB_GLOBAL
+ Value: 0x1000
+ Size: 0x8
+ - Name: second
+ Type: STT_FUNC
+ Section: .text.adjacent
+ Binding: STB_GLOBAL
+ Value: 0x1008
+ Size: 0x8
+ - Name: third
+ Type: STT_FUNC
+ Section: .text.gap
+ Binding: STB_GLOBAL
+ Value: 0x2000
+ Size: 0x8
+ - Name: fourth
+ Type: STT_FUNC
+ Section: .text.transfer
+ Binding: STB_GLOBAL
+ Value: 0x3000
+ Size: 0x4
+ - Name: split
+ Type: STT_FUNC
+ Section: .text.split
+ Binding: STB_GLOBAL
+ Value: 0x4000
+ Size: 0x10
+ - Name: split_gap
+ Type: STT_FUNC
+ Section: .text.split-gap-head
+ Binding: STB_GLOBAL
+ Value: 0x6000
+ Size: 0x8
+ ## A name of its own, so that this range is not marked as a function entry
+ ## and only the address gap can end the inferred range.
+ - Name: split_gap_tail
+ Type: STT_FUNC
+ Section: .text.split-gap-tail
+ Binding: STB_GLOBAL
+ Value: 0x7000
+ Size: 0x8
+ - Name: duplicate
+ Type: STT_FUNC
+ Section: .text.duplicate
+ Binding: STB_GLOBAL
+ Value: 0x8000
+ Size: 0x8
+ ## Two functions cannot share a symbol name, so the second one carries the
+ ## internal suffix that link-time renaming appends. Debug info keeps the
+ ## source name for both, and the suffix is stripped before the symbol is
+ ## compared with it, so this range is still marked as a function entry.
+ - Name: duplicate.llvm.1
+ Type: STT_FUNC
+ Section: .text.duplicate
+ Binding: STB_GLOBAL
+ Value: 0x8008
+ Size: 0x8
+ - Name: call_stop
+ Type: STT_FUNC
+ Section: .text.call
+ Binding: STB_GLOBAL
+ Value: 0x9000
+ Size: 0x10
+ - Name: ret_stop
+ Type: STT_FUNC
+ Section: .text.ret
+ Binding: STB_GLOBAL
+ Value: 0xA000
+ Size: 0x10
+ - Name: tail
+ Type: STT_FUNC
+ Section: .text.tail
+ Binding: STB_GLOBAL
+ Value: 0xB000
+ Size: 0x8
+ - Name: tail_undescribed
+ Type: STT_FUNC
+ Section: .text.tail-undescribed
+ Binding: STB_GLOBAL
+ Value: 0xB008
+ Size: 0x8
+ - Name: preceding
+ Type: STT_FUNC
+ Section: .text.func-change
+ Binding: STB_GLOBAL
+ Value: 0xC000
+ Size: 0x8
+ ## The function that follows starts at 0xc008, where no symbol sits, so that
+ ## range is not marked as a function entry. Its symbol sits at its later
+ ## range, which keeps the function from having no entry at all.
+ - Name: following
+ Type: STT_FUNC
+ Section: .text.func-change-tail
+ Binding: STB_GLOBAL
+ Value: 0xD000
+ Size: 0x8
+ - Name: trap_stop
+ Type: STT_FUNC
+ Section: .text.trap
+ Binding: STB_GLOBAL
+ Value: 0xD800
+ Size: 0x10
+ - Name: self_loop
+ Type: STT_FUNC
+ Section: .text.self
+ Binding: STB_GLOBAL
+ Value: 0xD900
+ Size: 0x4
+ - Name: inner_branch
+ Type: STT_FUNC
+ Section: .text.inner
+ Binding: STB_GLOBAL
+ Value: 0xDA00
+ Size: 0x10
+DWARF:
+ debug_abbrev:
+ - ID: 0
+ Table:
+ - Code: 1
+ Tag: DW_TAG_compile_unit
+ Children: DW_CHILDREN_yes
+ Attributes:
+ - Attribute: DW_AT_name
+ Form: DW_FORM_string
+ - Code: 2
+ Tag: DW_TAG_subprogram
+ Children: DW_CHILDREN_no
+ Attributes:
+ - Attribute: DW_AT_name
+ Form: DW_FORM_string
+ - Attribute: DW_AT_low_pc
+ Form: DW_FORM_addr
+ - Attribute: DW_AT_high_pc
+ Form: DW_FORM_data8
+ ## A function whose code is described by a range list instead of a
+ ## single low_pc/high_pc pair.
+ - Code: 3
+ Tag: DW_TAG_subprogram
+ Children: DW_CHILDREN_no
+ Attributes:
+ - Attribute: DW_AT_name
+ Form: DW_FORM_string
+ - Attribute: DW_AT_ranges
+ Form: DW_FORM_sec_offset
+ debug_ranges:
+ ## Two adjacent ranges of one function.
+ - Offset: 0x0
+ Entries:
+ - LowOffset: 0x4000
+ HighOffset: 0x4008
+ - LowOffset: 0x4008
+ HighOffset: 0x4010
+ ## End of the range list.
+ - LowOffset: 0x0
+ HighOffset: 0x0
+ ## Two ranges of one function separated by an address gap.
+ - Offset: 0x40
+ Entries:
+ - LowOffset: 0x6000
+ HighOffset: 0x6008
+ - LowOffset: 0x7000
+ HighOffset: 0x7008
+ - LowOffset: 0x0
+ HighOffset: 0x0
+ ## Two ranges of one function, the first of them adjacent to the function
+ ## that precedes it and started by no symbol.
+ - Offset: 0x80
+ Entries:
+ - LowOffset: 0xC008
+ HighOffset: 0xC010
+ - LowOffset: 0xD000
+ HighOffset: 0xD008
+ - LowOffset: 0x0
+ HighOffset: 0x0
+ debug_info:
+ - Version: 4
+ AbbrevTableID: 0
+ AddrSize: 8
+ Entries:
+ - AbbrCode: 1
+ Values:
+ - CStr: spe-range-gap.c
+ - AbbrCode: 2
+ Values:
+ - CStr: first
+ - Value: 0x1000
+ - Value: 0x8
+ - AbbrCode: 2
+ Values:
+ - CStr: second
+ - Value: 0x1008
+ - Value: 0x8
+ - AbbrCode: 2
+ Values:
+ - CStr: third
+ - Value: 0x2000
+ - Value: 0x8
+ - AbbrCode: 2
+ Values:
+ - CStr: fourth
+ - Value: 0x3000
+ - Value: 0x4
+ - AbbrCode: 3
+ Values:
+ - CStr: split
+ - Value: 0x0
+ - AbbrCode: 3
+ Values:
+ - CStr: split_gap
+ - Value: 0x40
+ ## Two functions of the same name, each described by its own entry.
+ - AbbrCode: 2
+ Values:
+ - CStr: duplicate
+ - Value: 0x8000
+ - Value: 0x8
+ - AbbrCode: 2
+ Values:
+ - CStr: duplicate
+ - Value: 0x8008
+ - Value: 0x8
+ - AbbrCode: 2
+ Values:
+ - CStr: call_stop
+ - Value: 0x9000
+ - Value: 0x10
+ - AbbrCode: 2
+ Values:
+ - CStr: ret_stop
+ - Value: 0xA000
+ - Value: 0x10
+ ## The code that follows this function has no entry of its own.
+ - AbbrCode: 2
+ Values:
+ - CStr: tail
+ - Value: 0xB000
+ - Value: 0x8
+ - AbbrCode: 2
+ Values:
+ - CStr: preceding
+ - Value: 0xC000
+ - Value: 0x8
+ - AbbrCode: 3
+ Values:
+ - CStr: following
+ - Value: 0x80
+ - AbbrCode: 2
+ Values:
+ - CStr: trap_stop
+ - Value: 0xD800
+ - Value: 0x10
+ - AbbrCode: 2
+ Values:
+ - CStr: self_loop
+ - Value: 0xD900
+ - Value: 0x4
+ - AbbrCode: 2
+ Values:
+ - CStr: inner_branch
+ - Value: 0xDA00
+ - Value: 0x10
+ - AbbrCode: 0
+
+#--- no-debug.yaml
+## The symbol table describes the function, but the binary carries no debug
+## info. Function ranges must not be recovered from the symbol table here.
+--- !ELF
+FileHeader:
+ Class: ELFCLASS64
+ Data: ELFDATA2LSB
+ Type: ET_DYN
+ Machine: EM_AARCH64
+ Entry: 0x1000
+Sections:
+ - Name: .text
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x1000
+ Offset: 0x1000
+ AddressAlign: 0x4
+ ## nop
+ ## nop
+ ## b 0x1000
+ Content: 1F2003D51F2003D5FEFFFF17
+ProgramHeaders:
+ - Type: PT_LOAD
+ Flags: [ PF_X, PF_R ]
+ Offset: 0
+ VAddr: 0
+ FileSize: 0x100C
+ MemSize: 0x100C
+ Align: 0x10000
+Symbols:
+ - Name: foo
+ Type: STT_FUNC
+ Section: .text
+ Binding: STB_GLOBAL
+ Value: 0x1000
+ Size: 0xC
+
+#--- perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0xe000) @ 0 00:00 0 0]: r-xp /tmp/spe-range-gap.exe
+70000000100c 0x70000000100c/0x700000001000/P/-/-/10/UNCOND/-
+
+700000003000 0x700000003000/0x700000002000/P/-/-/10/UNCOND/-
+70000000400c 0x70000000400c/0x700000004000/P/-/-/10/UNCOND/-
+700000007004 0x700000007004/0x700000006000/P/-/-/10/UNCOND/-
+70000000800c 0x70000000800c/0x700000008000/P/-/-/10/UNCOND/-
+70000000f000 0x70000000f000/0x700000000500/P/-/-/10/CALL/-
+70000000900c 0x70000000900c/0x700000009000/P/-/-/10/UNCOND/-
+70000000a00c 0x70000000a00c/0x70000000a000/P/-/-/10/UNCOND/-
+70000000b00c 0x70000000b00c/0x70000000b000/P/-/-/10/UNCOND/-
+70000000c00c 0x70000000c00c/0x70000000c000/P/-/-/10/UNCOND/-
+70000000d80c 0x70000000d80c/0x70000000d800/P/-/-/10/UNCOND/-
+70000000d900 0x70000000d900/0x70000000d900/P/-/-/10/UNCOND/-
+70000000da0c 0x70000000da0c/0x70000000da00/P/-/-/10/UNCOND/-
+
+#--- no-debug.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/no-debug.exe
+700000001008 0x700000001008/0x700000001000/P/-/-/10/UNCOND/-
+700000001008 0x700000001008/0x700000001000/P/-/-/10/UNCOND/-
+
+#--- unordered.yaml
+## Two described functions whose sections are listed in decreasing address
+## order, which leaves the collected instruction addresses unordered. The file
+## offsets still increase, as an object file requires.
+--- !ELF
+FileHeader:
+ Class: ELFCLASS64
+ Data: ELFDATA2LSB
+ Type: ET_DYN
+ Machine: EM_AARCH64
+ Entry: 0x1000
+Sections:
+ - Name: .text.high
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x2000
+ Offset: 0x1000
+ AddressAlign: 0x4
+ ## nop
+ ## b 0x2000
+ Content: 1F2003D5FFFFFF17
+ - Name: .text.low
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x1000
+ Offset: 0x1008
+ AddressAlign: 0x4
+ ## nop
+ ## b 0x1000
+ Content: 1F2003D5FFFFFF17
+ProgramHeaders:
+ - Type: PT_LOAD
+ Flags: [ PF_X, PF_R ]
+ Offset: 0
+ VAddr: 0
+ FileSize: 0x1010
+ MemSize: 0x1010
+ Align: 0x10000
+Symbols:
+ - Name: high
+ Type: STT_FUNC
+ Section: .text.high
+ Binding: STB_GLOBAL
+ Value: 0x2000
+ Size: 0x8
+ - Name: low
+ Type: STT_FUNC
+ Section: .text.low
+ Binding: STB_GLOBAL
+ Value: 0x1000
+ Size: 0x8
+DWARF:
+ debug_abbrev:
+ - ID: 0
+ Table:
+ - Code: 1
+ Tag: DW_TAG_compile_unit
+ Children: DW_CHILDREN_yes
+ Attributes:
+ - Attribute: DW_AT_name
+ Form: DW_FORM_string
+ - Code: 2
+ Tag: DW_TAG_subprogram
+ Children: DW_CHILDREN_no
+ Attributes:
+ - Attribute: DW_AT_name
+ Form: DW_FORM_string
+ - Attribute: DW_AT_low_pc
+ Form: DW_FORM_addr
+ - Attribute: DW_AT_high_pc
+ Form: DW_FORM_data8
+ debug_info:
+ - Version: 4
+ AbbrevTableID: 0
+ AddrSize: 8
+ Entries:
+ - AbbrCode: 1
+ Values:
+ - CStr: unordered.c
+ - AbbrCode: 2
+ Values:
+ - CStr: low
+ - Value: 0x1000
+ - Value: 0x8
+ - AbbrCode: 2
+ Values:
+ - CStr: high
+ - Value: 0x2000
+ - Value: 0x8
+ - AbbrCode: 0
+
+#--- unordered.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x3000) @ 0 00:00 0 0]: r-xp /tmp/unordered.exe
+700000001004 0x700000001004/0x700000001000/P/-/-/10/UNCOND/-
+700000002004 0x700000002004/0x700000002000/P/-/-/10/UNCOND/-
diff --git a/llvm/test/tools/llvm-profgen/AArch64/spe-symbolized-profile.test b/llvm/test/tools/llvm-profgen/AArch64/spe-symbolized-profile.test
new file mode 100644
index 0000000000000..a2aef4c3e96cf
--- /dev/null
+++ b/llvm/test/tools/llvm-profgen/AArch64/spe-symbolized-profile.test
@@ -0,0 +1,109 @@
+## The other Arm SPE tests skip symbolization and check the inferred ranges
+## themselves. This one carries a line table, so that the range and branch
+## counters of a single-branch record are checked where they end up: as the
+## line samples and the head count of the enclosing function.
+
+# RUN: split-file %s %t
+# RUN: yaml2obj %t/binary.yaml -o %t/spe-symbolized-profile.exe
+# RUN: llvm-profgen --binary=%t/spe-symbolized-profile.exe \
+# RUN: --perfscript=%t/perfscript --spe-branch-profile --format=text \
+# RUN: --output=%t/profile 2>&1 \
+# RUN: | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --input-file=%t/profile
+
+## The record is a return into a function that carries debug info, so neither
+## its source nor its destination is reported as unexpected, and no record is
+## dropped.
+# NO-WARN-NOT: warning: {{.*}}do not match the binary
+# NO-WARN-NOT: warning: {{.*}}not an instruction of this binary
+# NO-WARN-NOT: warning: No samples in perf script
+
+## The range inferred for the destination covers both instructions of
+## with_debug, which the line table attributes to the two lines that follow its
+## declaration, and the branch to its entry becomes the head count. The total
+## weighs each line sample by the size of the instructions it covers, as it
+## does for a branch-stack profile.
+# CHECK: with_debug:8:1{{$}}
+# CHECK-NEXT: 1: 1
+# CHECK-NEXT: 2: 1
+
+#--- binary.yaml
+## with_debug is described by both the line table and a subprogram; the
+## instructions of without_debug are described by neither, which leaves the
+## source of the record unsymbolized without affecting the counters.
+--- !ELF
+FileHeader:
+ Class: ELFCLASS64
+ Data: ELFDATA2LSB
+ Type: ET_DYN
+ Machine: EM_AARCH64
+ProgramHeaders:
+ - Type: PT_LOAD
+ Flags: [ PF_X, PF_R ]
+ VAddr: 0
+ Offset: 0
+ Align: 0x1000
+ FirstSec: .text
+ LastSec: .text
+Sections:
+ - Name: .text
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x27C
+ Offset: 0x27C
+ AddressAlign: 0x4
+ ## with_debug: add w0, w0, #1; ret
+ ## without_debug: add w0, w0, #2; ret
+ Content: 00040011C0035FD600080011C0035FD6
+ - Name: .debug_abbrev
+ Type: SHT_PROGBITS
+ AddressAlign: 0x1
+ Content: 011101252513050325721710171B25111B120673170000022E00111B1206401803253A0B3B0B49133F19000003240003253E0B0B0B000000
+ - Name: .debug_info
+ Type: SHT_PROGBITS
+ AddressAlign: 0x1
+ Content: 33000000050001080000000001001D0001080000000000000002000800000008000000020008000000016F030001320000000304050400
+ - Name: .debug_str_offsets
+ Type: SHT_PROGBITS
+ AddressAlign: 0x1
+ Content: 18000000050000000200000007000000000000001D00000003000000
+ - Name: .debug_line
+ Type: SHT_PROGBITS
+ AddressAlign: 0x1
+ Content: 460000000500080025000000010101FB0E0D00010101010000000100000101011F010000000002011F020F0102000000000400050C0A0009027C020000000000001305034B0204000101
+ - Name: .debug_line_str
+ Type: SHT_PROGBITS
+ Flags: [ SHF_MERGE, SHF_STRINGS ]
+ AddressAlign: 0x1
+ EntSize: 0x1
+ Content: 2F007761726E2D6E6F742D73796D626F6C697A65642E6300
+Symbols:
+ - Name: with_debug
+ Type: STT_FUNC
+ Section: .text
+ Binding: STB_GLOBAL
+ Value: 0x27C
+ Size: 0x8
+ - Name: without_debug
+ Type: STT_FUNC
+ Section: .text
+ Binding: STB_GLOBAL
+ Value: 0x284
+ Size: 0x8
+DWARF:
+ debug_str:
+ - '/'
+ - ''
+ - int
+ - warn-not-symbolized.c
+ - with_debug
+ debug_addr:
+ - Length: 0xC
+ Version: 0x5
+ AddressSize: 0x8
+ Entries:
+ - Address: 0x27C
+
+#--- perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x1000) @ 0 00:00 0 0]: r-xp /tmp/spe-symbolized-profile.exe
+700000000288 0x700000000288/0x70000000027c/P/-/-/10/RET/-
diff --git a/llvm/test/tools/llvm-profgen/spe-target-validation.test b/llvm/test/tools/llvm-profgen/spe-target-validation.test
new file mode 100644
index 0000000000000..fbcd792cbbfaf
--- /dev/null
+++ b/llvm/test/tools/llvm-profgen/spe-target-validation.test
@@ -0,0 +1,45 @@
+# REQUIRES: arm-registered-target
+#
+# RUN: split-file %s %t
+# RUN: yaml2obj %t/binary.yaml -o %t/arm.exe
+# RUN: not llvm-profgen --binary=%t/arm.exe --perfscript=%t/perfscript \
+# RUN: --spe-branch-profile --skip-symbolization --output=/dev/null 2>&1 \
+# RUN: | FileCheck %s
+#
+# CHECK: error: --spe-branch-profile requires an AArch64 binary.
+
+#--- binary.yaml
+--- !ELF
+FileHeader:
+ Class: ELFCLASS32
+ Data: ELFDATA2LSB
+ Type: ET_DYN
+ Machine: EM_ARM
+Sections:
+ - Name: .text
+ Type: SHT_PROGBITS
+ Flags: [ SHF_ALLOC, SHF_EXECINSTR ]
+ Address: 0x1000
+ Offset: 0x1000
+ AddressAlign: 0x4
+ ## bx lr
+ Content: 1EFF2FE1
+ProgramHeaders:
+ - Type: PT_LOAD
+ Flags: [ PF_X, PF_R ]
+ Offset: 0
+ VAddr: 0
+ FileSize: 0x1004
+ MemSize: 0x1004
+ Align: 0x1000
+Symbols:
+ - Name: foo
+ Type: STT_FUNC
+ Section: .text
+ Binding: STB_GLOBAL
+ Value: 0x1000
+ Size: 0x4
+
+#--- perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/arm.exe
+700000001000 0x700000001000/0x700000001000/P/-/-/10/RET/-
diff --git a/llvm/tools/llvm-profgen/Options.h b/llvm/tools/llvm-profgen/Options.h
index 395a55726274d..4886962790bd9 100644
--- a/llvm/tools/llvm-profgen/Options.h
+++ b/llvm/tools/llvm-profgen/Options.h
@@ -24,6 +24,7 @@ extern cl::opt<bool> EnableCSPreInliner;
extern cl::opt<bool> UseContextCostForPreInliner;
extern cl::opt<bool> LoadFunctionFromSymbol;
extern cl::opt<bool> TimeProfGen;
+extern cl::opt<bool> ReadSPEBranchProfile;
} // end namespace llvm
diff --git a/llvm/tools/llvm-profgen/PerfReader.cpp b/llvm/tools/llvm-profgen/PerfReader.cpp
index 18297d5f7665b..2b79714ec7cf3 100644
--- a/llvm/tools/llvm-profgen/PerfReader.cpp
+++ b/llvm/tools/llvm-profgen/PerfReader.cpp
@@ -70,6 +70,12 @@ cl::opt<bool> TimeProfGen("time-profgen", cl::desc("Time llvm-profgen phases"),
static const char *TimerGroupName = "profgen";
static const char *TimerGroupDesc = "llvm-profgen";
+cl::opt<bool> ReadSPEBranchProfile(
+ "spe-branch-profile",
+ cl::desc("Read the input as an Arm SPE branch profile of an AArch64 "
+ "binary; requires FEAT_SPEv1p2 and event filtering."),
+ cl::cat(ProfGenCategory));
+
namespace sampleprof {
void VirtualUnwinder::unwindCall(UnwindState &State) {
@@ -364,8 +370,15 @@ PerfReaderBase::create(ProfiledBinary *Binary, InputFile &Input,
return PerfReader;
}
+ // An Arm SPE trace cannot be detected: its lines look like the branch-stack
+ // lines of an LBR trace, so detection would read it as one. The kind is also
+ // needed before a recording is printed, because the perf command differs.
+ if (ReadSPEBranchProfile)
+ Input.Content = PerfContent::ArmSPE;
+
// For perf data input, we need to convert them into perf script first.
// If this is a kernel perf file, there is no need for retrieving PIDs.
+ bool ConvertedFromPerfData = Input.Format == InputFormat::PerfData;
if (Input.Format == InputFormat::PerfData)
Input = PerfScriptReader::convertPerfDataToTrace(Binary, Binary->isKernel(),
Input, PIDFilter);
@@ -373,12 +386,16 @@ PerfReaderBase::create(ProfiledBinary *Binary, InputFile &Input,
assert((Input.Format == InputFormat::PerfScript) &&
"Should be a perfscript!");
- Input.Content = PerfScriptReader::checkPerfScriptType(Input.InputFilePath);
+ if (Input.Content == PerfContent::UnknownContent)
+ Input.Content = PerfScriptReader::checkPerfScriptType(Input.InputFilePath);
if (Input.Content == PerfContent::LBRStack) {
PerfReader.reset(
new HybridPerfReader(Binary, Input.InputFilePath, PIDFilter));
} else if (Input.Content == PerfContent::LBR) {
PerfReader.reset(new LBRPerfReader(Binary, Input.InputFilePath, PIDFilter));
+ } else if (Input.Content == PerfContent::ArmSPE) {
+ PerfReader.reset(new ArmSPEReader(Binary, Input.InputFilePath, PIDFilter,
+ ConvertedFromPerfData));
} else {
exitWithError("Unsupported perfscript!");
}
@@ -553,6 +570,10 @@ PerfScriptReader::convertPerfDataToTrace(ProfiledBinary *Binary, bool SkipPID,
SmallVector<StringRef, 8> ScriptSampleArgs;
ScriptSampleArgs.push_back(PerfExecutablePath);
ScriptSampleArgs.push_back("script");
+ // Decode Arm SPE branch events into a branch stack of one entry: a larger
+ // limit adds a synthesized entry with source 0x0, which the reader rejects.
+ if (File.Content == PerfContent::ArmSPE)
+ ScriptSampleArgs.push_back("--itrace=bl1");
ScriptSampleArgs.push_back("--show-mmap-events");
ScriptSampleArgs.push_back("-F");
ScriptSampleArgs.push_back("ip,brstack");
@@ -564,8 +585,7 @@ PerfScriptReader::convertPerfDataToTrace(ProfiledBinary *Binary, bool SkipPID,
}
RunPerfScript(ScriptSampleArgs);
- return {std::string(PerfTraceFile), InputFormat::PerfScript,
- PerfContent::UnknownContent};
+ return {std::string(PerfTraceFile), InputFormat::PerfScript, File.Content};
}
static StringRef filename(StringRef Path, bool UseBackSlash) {
@@ -731,8 +751,45 @@ static bool parseAddress(StringRef Str, uint64_t &Addr, bool HasPrefix) {
return Str.getAsInteger(16, Addr);
}
+/// Clear the bit of the branch-stack entry at \p Index in \p TakenMask. Only a
+/// record format that holds a branch that was not taken reaches this, and the
+/// deepest such record holds one branch, so the index stays inside the mask.
+static void clearTaken(uint64_t &TakenMask, size_t Index) {
+ assert(Index < TakenMaskWidth && "Branch-stack entry past the mask");
+ TakenMask &= ~(1ULL << Index);
+}
+
+bool PerfScriptReader::parseBranchEntry(StringRef Token, uint64_t &Source,
+ uint64_t &Target, bool &Taken) {
+ // Only the source, the destination and the prediction field are read, so stop
+ // at the third slash and leave the rest of the entry as one field. Every
+ // field is kept, so a field count still counts the slashes before it.
+ SmallVector<StringRef, 4> Fields;
+ Token.split(Fields, "/", /*MaxSplit=*/3);
+ if (Fields.size() < 2 || parseAddress(Fields[0], Source, true) ||
+ parseAddress(Fields[1], Target, true))
+ return false;
+
+ // The third field holds the prediction flags, whose alphabet depends on the
+ // record format, so whether they mark a branch that was taken is left to the
+ // reader. A field the reader cannot read makes the entry unreadable. perf
+ // prints fields of its own after this one, so a prediction field that ends
+ // the entry was cut short and is withheld: its P may be all of a PN.
+ std::optional<bool> Flag =
+ parseTakenFlag(Fields.size() > 3 ? Fields[2] : StringRef());
+ if (!Flag)
+ return false;
+ Taken = *Flag;
+
+ // Canonicalize to use preferred load address as base address.
+ Source = Binary->canonicalizeVirtualAddress(Source);
+ Target = Binary->canonicalizeVirtualAddress(Target);
+ return true;
+}
+
bool PerfScriptReader::extractLBRStack(TraceStream &TraceIt,
- SmallVectorImpl<LBREntry> &LBRStack) {
+ SmallVectorImpl<LBREntry> &LBRStack,
+ uint64_t &TakenMask) {
// The raw format of LBR stack is like:
// 0x4005c8/0x4005dc/P/-/-/0 0x40062f/0x4005b0/P/-/-/0 ...
// ... 0x4005c8/0x4005dc/P/-/-/0
@@ -765,21 +822,16 @@ bool PerfScriptReader::extractLBRStack(TraceStream &TraceIt,
if (Token.size() == 0)
continue;
- SmallVector<StringRef, 8> Addresses;
- Token.split(Addresses, "/");
uint64_t Src;
uint64_t Dst;
+ bool Taken = true;
// Stop at broken LBR records.
- if (Addresses.size() < 2 || parseAddress(Addresses[0], Src, true) ||
- parseAddress(Addresses[1], Dst, true)) {
+ if (!parseBranchEntry(Token, Src, Dst, Taken)) {
WarnInvalidLBR(TraceIt);
break;
}
- // Canonicalize to use preferred load address as base address.
- Src = Binary->canonicalizeVirtualAddress(Src);
- Dst = Binary->canonicalizeVirtualAddress(Dst);
bool SrcIsInternal = Binary->addressIsCode(Src);
bool DstIsInternal = Binary->addressIsCode(Dst);
if (!SrcIsInternal)
@@ -790,12 +842,245 @@ bool PerfScriptReader::extractLBRStack(TraceStream &TraceIt,
if (!SrcIsInternal && !DstIsInternal)
continue;
+ if (!Taken)
+ clearTaken(TakenMask, LBRStack.size());
LBRStack.emplace_back(LBREntry(Src, Dst));
}
TraceIt.advance();
return !LBRStack.empty();
}
+uint64_t ArmSPEReader::parseAggregatedCount(TraceStream &) {
+ // perf prints no repeat count for Arm SPE, so leave the line to
+ // extractLBRStack, which reports it when unreadable, instead of silently
+ // weighing the next record by an address read off it as a count.
+ return 1;
+}
+
+std::optional<bool> ArmSPEReader::parseTakenFlag(StringRef PredictionFlags) {
+ // perf appends N to the prediction flags of a branch it recorded as not
+ // taken, and prints `-`, `P`, `M`, `PN` or `MN` for the field. Anything else
+ // leaves the entry unreadable rather than silently counted as taken.
+ if (PredictionFlags == "-" || PredictionFlags == "P" ||
+ PredictionFlags == "M")
+ return true;
+ if (PredictionFlags == "PN" || PredictionFlags == "MN")
+ return false;
+ return std::nullopt;
+}
+
+bool ArmSPEReader::extractLBRStack(TraceStream &TraceIt,
+ SmallVectorImpl<LBREntry> &LBRStack,
+ uint64_t &TakenMask) {
+ // perf prints an Arm SPE branch record as a branch stack of one entry,
+ // preceded by the instruction pointer of the sample, which repeats the
+ // source of the branch:
+ // ffff800080387878 0xffff800080387878/0xffff80008039c524/P/-/-/19/RET/-
+ SmallVector<StringRef, 4> Fields;
+ TraceIt.getCurrentLine().rtrim().split(Fields, " ", -1, false);
+
+ // A blank line carries no record, so skip it without a diagnostic.
+ if (Fields.empty()) {
+ TraceIt.advance();
+ return false;
+ }
+
+ // Report a line that holds no readable record. A trace printed with another
+ // field list holds millions of them, so name a line only on request;
+ // reportTraceLosses reports the share once.
+ auto ReportInvalidRecord = [&]() {
+ NumInvalidRecords++;
+ if (ShowDetailedWarning)
+ WithColor::warning() << "Invalid Arm SPE branch record at line "
+ << TraceIt.getLineNumber() << ": "
+ << TraceIt.getCurrentLine() << "\n";
+ TraceIt.advance();
+ };
+
+ // The leading instruction pointer is a bare address: no slash, and none of
+ // the 0x prefix that every address of an entry carries, so a prefixed field
+ // is something else and leaves the record unreadable.
+ size_t Index = 0;
+ uint64_t LeadingAddr = 0;
+ bool HasLeadingAddr = !Fields[0].contains('/');
+ if (HasLeadingAddr) {
+ if (parseAddress(Fields[0], LeadingAddr, /*HasPrefix=*/false)) {
+ ReportInvalidRecord();
+ return false;
+ }
+ Index = 1;
+ }
+
+ // A record holds the entry of the sampled branch right after the instruction
+ // pointer.
+ uint64_t Src = 0;
+ uint64_t Dst = 0;
+ bool Taken = true;
+ if (Fields.size() - Index < 1 ||
+ !parseBranchEntry(Fields[Index], Src, Dst, Taken)) {
+ ReportInvalidRecord();
+ return false;
+ }
+
+ // --itrace=bl1 leaves a record one entry. A second entry that parses is a
+ // deeper stack, whose synthesized source address 0x0 would fabricate a branch
+ // from outside the binary, so stop instead of weighing such a trace. Parse
+ // the field to decide, because a trailing field a wider -F list adds, a
+ // shared-object path above all, holds a slash too and leaves the record
+ // unreadable rather than the stack deeper.
+ if (Fields.size() - Index > 1) {
+ uint64_t ExtraSrc = 0;
+ uint64_t ExtraDst = 0;
+ bool ExtraTaken = true;
+ if (parseBranchEntry(Fields[Index + 1], ExtraSrc, ExtraDst, ExtraTaken))
+ exitWithError(
+ Twine("--spe-branch-profile requires Arm SPE branch records "
+ "created with `perf script --show-mmap-events --itrace=bl1 -F "
+ "ip,brstack`, but the record at line ") +
+ Twine(TraceIt.getLineNumber()) +
+ " holds a branch stack of more than one entry");
+ ReportInvalidRecord();
+ return false;
+ }
+
+ // A source address of zero names no instruction: perf prints it where the
+ // address of the sampled branch was lost, and left in place it reads as a
+ // branch from outside the binary. Compare against the canonical form of zero,
+ // because parseBranchEntry has already shifted every address of the trace
+ // onto the preferred load address of the binary.
+ if (Src == Binary->canonicalizeVirtualAddress(0)) {
+ ReportInvalidRecord();
+ return false;
+ }
+
+ // The leading instruction pointer repeats the source of the branch. Where the
+ // two differ, another field was printed first and neither can be trusted.
+ if (HasLeadingAddr &&
+ Binary->canonicalizeVirtualAddress(LeadingAddr) != Src) {
+ ReportInvalidRecord();
+ return false;
+ }
+
+ NumRecords++;
+
+ // The inherited parseSample reports a missing mmap event only for a record
+ // that yields a sample, and without the event every record is usually
+ // external and dropped below. The repeated call is a no-op once it has
+ // warned.
+ warnIfMissingMMap();
+
+ TraceIt.advance();
+ // A record holds one branch and no preceding one to bound a range with, so a
+ // destination outside the binary yields no counter at all. Count the record
+ // so that reportTraceLosses reports the share.
+ if (!Binary->addressIsCode(Dst)) {
+ NumDroppedRecords++;
+ return false;
+ }
+ // A branch from outside the binary keeps an external source, which the
+ // profile turns into a head sample of the function holding the destination.
+ if (!Binary->addressIsCode(Src))
+ Src = ExternalAddr;
+
+ // A branch that was not taken reports the following instruction as its
+ // destination, so keep the entry for the range that really was executed
+ // there; computeCounterFromLBR leaves out its branch counter from the flag.
+ if (!Taken)
+ clearTaken(TakenMask, LBRStack.size());
+ LBRStack.emplace_back(Src, Dst);
+ return true;
+}
+
+uint64_t ArmSPEReader::inferFirstRangeEnd(const PerfSample *Sample,
+ uint64_t Repeat) {
+ assert(!Sample->LBRStack.empty() &&
+ "an aggregated Arm SPE sample holds one branch");
+ // Execution continues at the destination of the recorded branch until the
+ // next transfer, so that transfer ends the range.
+ bool Uncovered = false;
+ uint64_t End =
+ Binary->findRangeEnd(Sample->LBRStack.front().Target, &Uncovered);
+ // Where the binary describes no function around the destination, the range
+ // is reduced to the sampled instruction. Count that here, in the one pass
+ // that looks the range of every sample up, so that reportTraceLosses needs
+ // no lookup of its own.
+ if (Uncovered)
+ NumUncoveredSamples += Repeat;
+ return End;
+}
+
+void ArmSPEReader::validateParsedTrace() {
+ // An input without a single Arm SPE branch record is the wrong input file, so
+ // reject it instead of writing an empty profile. Deciding after the whole
+ // trace was read keeps a damaged first line from discarding what follows.
+ if (NumRecords)
+ return;
+
+ // A recording is turned into a trace by this tool, with a perf command the
+ // user never wrote, so name what to record instead of what to print.
+ if (ConvertedFromPerfData)
+ exitWithError("--spe-branch-profile found no Arm SPE branch record in "
+ "the perf data; record one with "
+ "arm_spe/branch_filter=1,event_filter=2/ using perf 6.15 or "
+ "later");
+
+ // For a trace the user printed, the field list of the perf command is what
+ // has to be corrected, so name it together with what the lines looked like.
+ StringRef Cause = NumInvalidRecords
+ ? "no line of the input holds a usable one"
+ : "the trace holds no record at all";
+ exitWithError(
+ Twine("--spe-branch-profile requires Arm SPE branch records created "
+ "with `perf script --show-mmap-events --itrace=bl1 -F "
+ "ip,brstack`, but ") +
+ Cause);
+}
+
+void ArmSPEReader::generateUnsymbolizedProfile() {
+ PerfScriptReader::generateUnsymbolizedProfile();
+ // Computing the counters looks the range of every sample up, which is what
+ // tells how many of them had to be reduced, so the trace is reported on once
+ // that has run rather than from warnInvalidRange.
+ reportTraceLosses();
+}
+
+void ArmSPEReader::reportTraceLosses() {
+ // Every share below is of the same base, the lines of the trace that were
+ // read as records, so that the reports can be read against one another.
+ // Blank lines are left out because they carry no record to lose.
+ uint64_t NumLines = NumInvalidRecords + NumRecords;
+
+ // Report the unreadable lines as a share, the way the other readers report
+ // their losses. Naming each line is left to --show-detailed-warning, because
+ // a wrong field list makes every line of the trace unreadable.
+ emitWarningSummary(NumInvalidRecords, NumLines,
+ "of the sample lines of the trace are not usable Arm SPE "
+ "branch records.");
+
+ // A record whose destination is not an instruction of this binary carries no
+ // counter, and a trace of another process or build consists of such records.
+ // Report their share, so a thinned profile is not taken for a complete one.
+ emitWarningSummary(NumDroppedRecords, NumLines,
+ "of the sample lines of the trace are Arm SPE branch "
+ "records whose destination is not an instruction of this "
+ "binary.");
+
+ // The range of a record whose destination no function range covers is reduced
+ // to that one instruction. Report the share reduced this way; the count is of
+ // lines again because inferFirstRangeEnd weighs each sample by its repeats.
+ emitWarningSummary(NumUncoveredSamples, NumLines,
+ "of the sample lines of the trace are Arm SPE branch "
+ "records whose destination cannot be associated with a "
+ "function, so only that instruction is profiled.");
+
+ // A trace whose every record was dropped yields an empty profile. The base
+ // class decides that from ranges formed out of consecutive branch-stack
+ // entries, which a single-branch record never forms, so name the records.
+ if (NumRecords == NumDroppedRecords)
+ WithColor::warning() << "No Arm SPE branch record yields a counter for "
+ "this binary!\n";
+}
+
bool PerfScriptReader::extractCallstack(TraceStream &TraceIt,
SmallVectorImpl<uint64_t> &CallStack) {
// The raw format of call stack is like:
@@ -907,7 +1192,7 @@ void HybridPerfReader::parseSample(TraceStream &TraceIt, uint64_t Count) {
if (!TraceIt.isAtEoF() && isLBRSample(TraceIt.getCurrentLine(), true)) {
// Parsing LBR stack and populate into PerfSample.LBRStack
- if (extractLBRStack(TraceIt, Sample->LBRStack)) {
+ if (extractLBRStack(TraceIt, Sample->LBRStack, Sample->TakenMask)) {
if (IgnoreStackSamples) {
Sample->CallStack.clear();
} else {
@@ -1072,20 +1357,26 @@ void UnsymbolizedProfileReader::parsePerfTraces() {
void PerfScriptReader::computeCounterFromLBR(const PerfSample *Sample,
uint64_t Repeat) {
SampleCounter &Counter = SampleCounters.begin()->second;
- uint64_t EndAddress = 0;
- for (const LBREntry &LBR : Sample->LBRStack) {
+ uint64_t EndAddress = inferFirstRangeEnd(Sample, Repeat);
+ for (size_t I = 0; I < Sample->LBRStack.size(); I++) {
+ const LBREntry &LBR = Sample->LBRStack[I];
uint64_t SourceAddress = LBR.Source;
uint64_t TargetAddress = LBR.Target;
// Record the branch if its SourceAddress is external. It can be the case an
// external source call an internal function, later this branch will be used
// to generate the function's head sample.
- if (Binary->addressIsCode(TargetAddress)) {
+ // A branch that was not taken transferred no control, and counting it would
+ // make a fall-through that reaches the entry of the next function a call of
+ // that function and a head sample of it. A hardware branch stack holds
+ // taken branches alone, so every entry of one is flagged as taken.
+ if (Binary->addressIsCode(TargetAddress) && Sample->taken(I))
Counter.recordBranchCount(SourceAddress, TargetAddress, Repeat);
- }
// If this not the first LBR, update the range count between TO of current
- // LBR and FROM of next LBR.
+ // LBR and FROM of next LBR. The range that follows the first LBR is bound
+ // by no other entry, so it is only recorded for an input whose reader can
+ // infer where it ends.
uint64_t StartAddress = TargetAddress;
if (Binary->addressIsCode(StartAddress) &&
Binary->addressIsCode(EndAddress) &&
@@ -1098,7 +1389,7 @@ void PerfScriptReader::computeCounterFromLBR(const PerfSample *Sample,
void LBRPerfReader::parseSample(TraceStream &TraceIt, uint64_t Count) {
std::shared_ptr<PerfSample> Sample = std::make_shared<PerfSample>();
// Parsing LBR stack and populate into PerfSample.LBRStack
- if (extractLBRStack(TraceIt, Sample->LBRStack)) {
+ if (extractLBRStack(TraceIt, Sample->LBRStack, Sample->TakenMask)) {
warnIfMissingMMap();
// Record LBR only samples by aggregation
AggregatedSamples[Hashable<PerfSample>(Sample)] += Count;
@@ -1427,41 +1718,56 @@ void PerfScriptReader::warnInvalidRange() {
}
void PerfScriptReader::warnIfBranchTargetMismatch() {
- // Collect unique branch source and target addresses from LBR samples,
- // then check what percentage don't match known instructions in the binary.
+ // Check what share of branch samples with both endpoints in the binary
+ // disagrees with its instructions, weighing each aggregated sample by how
+ // often it was recorded.
uint64_t MismatchedBranches = 0;
+ uint64_t NonConditionalNotTaken = 0;
uint64_t MismatchedIndirectTargets = 0;
uint64_t MismatchedTargets = 0;
uint64_t TotalSamples = 0;
for (const auto &Item : AggregatedSamples) {
const PerfSample *Sample = Item.first.getPtr();
- for (const LBREntry &LBR : Sample->LBRStack) {
- uint64_t Source = LBR.Source;
- uint64_t Target = LBR.Target;
+ uint64_t Repeat = Item.second;
+ for (size_t I = 0; I < Sample->LBRStack.size(); I++) {
+ uint64_t Source = Sample->LBRStack[I].Source;
+ uint64_t Target = Sample->LBRStack[I].Target;
if (Source == ExternalAddr || Target == ExternalAddr)
continue;
- TotalSamples++;
+ TotalSamples += Repeat;
// Validate Branch sources are Call/Branch/Indirect Branch
if (!Binary->addressIsTransfer(Source))
- MismatchedBranches++;
-
- // Validate Indirect Branch targets landed in code. This may over estimate
- // the vaid targets only because there's no good way to determine jump
- // table targets
- if (Binary->addressIsIndirectBranch(Source)) {
+ MismatchedBranches += Repeat;
+
+ // A not-taken branch must be conditional and target its following
+ // instruction, even when the recorded target is otherwise known.
+ if (!Sample->taken(I)) {
+ if (!Binary->addressIsConditionalBranch(Source))
+ NonConditionalNotTaken += Repeat;
+ else if (Target <= Source ||
+ Target - Source != Binary->getInstSize(Source))
+ MismatchedTargets += Repeat;
+ } else if (Binary->addressIsIndirectBranch(Source)) {
+ // Validate indirect targets as code because jump-table targets cannot
+ // be identified more precisely.
if (!Binary->addressIsCode(Target))
- MismatchedIndirectTargets++;
+ MismatchedIndirectTargets += Repeat;
} else if (!Binary->addressIsBranchTarget(Target) &&
- !Binary->findFuncRangeForStartAddr(Target))
- MismatchedTargets++;
+ !Binary->findFuncRangeForStartAddr(Target)) {
+ MismatchedTargets += Repeat;
+ }
}
}
emitWarningSummary(MismatchedBranches, TotalSamples,
"of branch samples do not match the binary.");
+ emitWarningSummary(
+ NonConditionalNotTaken, TotalSamples,
+ "of branch samples are marked not taken but do not name a conditional "
+ "branch.");
emitWarningSummary(MismatchedTargets, TotalSamples,
"of branch targets do not match the binary.");
emitWarningSummary(MismatchedIndirectTargets, TotalSamples,
@@ -1471,6 +1777,7 @@ void PerfScriptReader::warnIfBranchTargetMismatch() {
void PerfScriptReader::parsePerfTraces() {
// Parse perf traces and do aggregation.
parseAndAggregateTrace();
+ validateParsedTrace();
if (Binary->isKernel() && !Binary->getIsLoadedByMMap()) {
exitWithError(
"Kernel is requested, but no kernel is found in mmap events.");
diff --git a/llvm/tools/llvm-profgen/PerfReader.h b/llvm/tools/llvm-profgen/PerfReader.h
index b9af0f19cb5d3..7aa707924871b 100644
--- a/llvm/tools/llvm-profgen/PerfReader.h
+++ b/llvm/tools/llvm-profgen/PerfReader.h
@@ -73,6 +73,7 @@ enum PerfContent {
UnknownContent = 0,
LBR = 1, // Only LBR sample.
LBRStack = 2, // Hybrid sample including call stack and LBR stack.
+ ArmSPE = 3, // Arm SPE branch samples.
};
struct InputFile {
@@ -95,11 +96,18 @@ struct LBREntry {
#endif
};
+// Number of branch-stack entries PerfSample::TakenMask can flag, one bit each.
+// Arm BRBE, the deepest branch stack in hardware, holds 64 records.
+constexpr size_t TakenMaskWidth = 64;
+
#ifndef NDEBUG
-static inline void printLBRStack(const SmallVectorImpl<LBREntry> &LBRStack) {
+static inline void printLBRStack(const SmallVectorImpl<LBREntry> &LBRStack,
+ uint64_t TakenMask = ~uint64_t(0)) {
for (size_t I = 0; I < LBRStack.size(); I++) {
dbgs() << "[" << I << "] ";
LBRStack[I].print();
+ if (I < TakenMaskWidth && !((TakenMask >> I) & 1))
+ dbgs() << " (not taken)";
dbgs() << "\n";
}
}
@@ -154,6 +162,19 @@ struct PerfSample {
// Call stack recorded in FILO(leaf to root) order, it's used for CS-profile
// generation
SmallVector<uint64_t, 16> CallStack;
+ // Bit I is clear where LBRStack[I] is a branch that was not taken, which perf
+ // flags with the N suffix of the prediction field; such an entry names the
+ // instruction that follows the branch as its target. Only Arm SPE records
+ // such a branch. Kept beside the stack so that an LBREntry stays two
+ // addresses wide. Every bit starts set, so a reader of taken branches alone
+ // leaves the mask as it is.
+ uint64_t TakenMask = ~uint64_t(0);
+
+ // Whether the branch of LBRStack[Index] was taken. An entry past the width of
+ // the mask carries no flag and counts as taken.
+ bool taken(size_t Index) const {
+ return Index >= TakenMaskWidth || (TakenMask >> Index) & 1;
+ }
virtual ~PerfSample() = default;
uint64_t getHashCode() const {
@@ -169,6 +190,9 @@ struct PerfSample {
Hash = HashCombine(Hash, Entry.Source);
Hash = HashCombine(Hash, Entry.Target);
}
+ // TakenMask is left out, so a taken and a not-taken record holding the
+ // same two addresses hash alike, which a conditional branch to the
+ // instruction that follows it produces; isEqual tells the two apart.
return Hash;
}
@@ -176,8 +200,11 @@ struct PerfSample {
const SmallVector<uint64_t, 16> &OtherCallStack = Other->CallStack;
const SmallVector<LBREntry, 16> &OtherLBRStack = Other->LBRStack;
+ // A branch that was taken and one that was not are separate events even
+ // where they hold the same two addresses, so the masks take part.
if (CallStack.size() != OtherCallStack.size() ||
- LBRStack.size() != OtherLBRStack.size())
+ LBRStack.size() != OtherLBRStack.size() ||
+ TakenMask != Other->TakenMask)
return false;
if (!std::equal(CallStack.begin(), CallStack.end(), OtherCallStack.begin()))
@@ -197,7 +224,7 @@ struct PerfSample {
void print() const {
dbgs() << "Line " << Linenum << "\n";
dbgs() << "LBR stack\n";
- printLBRStack(LBRStack);
+ printLBRStack(LBRStack, TakenMask);
dbgs() << "Call stack\n";
printCallStack(CallStack);
}
@@ -646,25 +673,53 @@ class PerfScriptReader : public PerfReaderBase {
void parseEventOrSample(TraceStream &TraceIt);
// Warn if the relevant mmap event is missing.
void warnIfMissingMMap();
+ // Reject an input that holds no record this reader can use, once the whole
+ // trace has been read. A branch-stack reader recognizes the shape of its
+ // trace before reading it, so this does nothing for one.
+ virtual void validateParsedTrace() {}
// Emit accumulate warnings.
void warnTruncatedStack();
// Warn if range is invalid.
- void warnInvalidRange();
+ virtual void warnInvalidRange();
// Warn if sampled branch/target addresses don't match the binary.
void warnIfBranchTargetMismatch();
// Extract call stack from the perf trace lines
bool extractCallstack(TraceStream &TraceIt,
SmallVectorImpl<uint64_t> &CallStack);
- // Extract LBR stack from one perf trace line
- bool extractLBRStack(TraceStream &TraceIt,
- SmallVectorImpl<LBREntry> &LBRStack);
- uint64_t parseAggregatedCount(TraceStream &TraceIt);
+ // Parse one branch-stack entry "<source>/<destination>/<flags>...", the way
+ // `perf script -F brstack` prints it, into addresses canonicalized against
+ // the runtime mapping and the one flag the profile depends on. Return false
+ // for any other token, which each reader reports in its own way.
+ bool parseBranchEntry(StringRef Token, uint64_t &Source, uint64_t &Target,
+ bool &Taken);
+ // Read from the prediction field of a branch-stack entry whether the branch
+ // was taken, or std::nullopt for a field the reader cannot read, which makes
+ // the entry unreadable. A hardware branch stack holds taken branches alone,
+ // so the base implementation answers yes without reading the field.
+ virtual std::optional<bool> parseTakenFlag(StringRef) { return true; }
+ // Extract LBR stack from one perf trace line. Bit I of \p TakenMask flags
+ // LBRStack[I], as PerfSample::TakenMask describes.
+ virtual bool extractLBRStack(TraceStream &TraceIt,
+ SmallVectorImpl<LBREntry> &LBRStack,
+ uint64_t &TakenMask);
+ // Consume a line that holds a decimal repeat count alone and return it, or
+ // return 1 and leave the line for the record parser. Override this where a
+ // record cannot be preceded by such a line, so it is not read as a count.
+ virtual uint64_t parseAggregatedCount(TraceStream &TraceIt);
// Parse one sample from multiple perf lines, override this for different
// sample type
void parseSample(TraceStream &TraceIt);
// An aggregated count is given to indicate how many times the sample is
// repeated.
virtual void parseSample(TraceStream &TraceIt, uint64_t Count){};
+ // Return the end of the executed range that follows the most recent branch of
+ // the sample, or zero when the reader cannot infer it. Consecutive entries of
+ // a branch stack bound every other range of the sample, but not this one, so
+ // the LBR and BRBE readers omit it while a reader of single-branch records
+ // infers it from the binary. Repeat is how often the sample was recorded.
+ virtual uint64_t inferFirstRangeEnd(const PerfSample *, uint64_t) {
+ return 0;
+ }
void computeCounterFromLBR(const PerfSample *Sample, uint64_t Repeat);
// Post process the profile after trace aggregation, we will do simple range
// overlap computation for AutoFDO, or unwind for CSSPGO(hybrid sample).
@@ -695,6 +750,65 @@ class LBRPerfReader : public PerfScriptReader {
void parseSample(TraceStream &TraceIt, uint64_t Count) override;
};
+// The reader of an Arm SPE branch profile, whose record holds a single branch
+// printed as a branch stack of one entry and fed into the same aggregation.
+class ArmSPEReader : public LBRPerfReader {
+public:
+ ArmSPEReader(ProfiledBinary *Binary, StringRef PerfTrace,
+ std::optional<int32_t> PID, bool ConvertedFromPerfData)
+ : LBRPerfReader(Binary, PerfTrace, PID),
+ ConvertedFromPerfData(ConvertedFromPerfData) {}
+
+protected:
+ // Weigh every Arm SPE record by one, since perf prints no repeat count for it
+ // and a record damaged down to its instruction pointer would read as a count.
+ uint64_t parseAggregatedCount(TraceStream &TraceIt) override;
+ // Read the N flag perf appends to the prediction field of a branch that was
+ // not taken, which only Arm SPE records.
+ std::optional<bool> parseTakenFlag(StringRef PredictionFlags) override;
+ // Parse one Arm SPE branch record, and reject a record that holds more than
+ // the sampled branch, which a trace decoded without --itrace=bl1 carries.
+ bool extractLBRStack(TraceStream &TraceIt,
+ SmallVectorImpl<LBREntry> &LBRStack,
+ uint64_t &TakenMask) override;
+ // Infer from the binary the range executed after the recorded branch, and
+ // count the sample as reduced when no function range covers its destination.
+ uint64_t inferFirstRangeEnd(const PerfSample *Sample,
+ uint64_t Repeat) override;
+ // Reject a trace that holds no Arm SPE branch record at all, naming the
+ // record format the lines were read as.
+ void validateParsedTrace() override;
+ // Report nothing: the base report describes ranges bounded by consecutive
+ // branch-stack entries, which a single-branch record cannot form.
+ void warnInvalidRange() override {}
+ // Compute the counters, then report what the trace lost on the way to them.
+ void generateUnsymbolizedProfile() override;
+
+private:
+ // Report, as shares of the sample lines of the trace, the lines that hold no
+ // record, the records that carry no counter and the samples whose range was
+ // reduced, and report a trace that yields no counter at all.
+ void reportTraceLosses();
+
+ // Number of parsable Arm SPE branch records, whether or not they yield a
+ // counter for this binary.
+ uint64_t NumRecords = 0;
+ // Number of lines that hold no readable record, every line of a trace printed
+ // with another field list among them.
+ uint64_t NumInvalidRecords = 0;
+ // Number of records dropped because their destination is not an instruction
+ // of this binary: a branch leaving the binary, a trace of another process or
+ // build, or a record whose branch target address was lost.
+ uint64_t NumDroppedRecords = 0;
+ // Number of samples, weighed by how often each was recorded, whose range was
+ // reduced to the sampled instruction because no function range covers it.
+ uint64_t NumUncoveredSamples = 0;
+ // Whether this tool printed the trace from a --perfdata recording. Only the
+ // report of an input that holds no record depends on it: the perf command to
+ // correct is one the user wrote for a --perfscript input alone.
+ bool ConvertedFromPerfData = false;
+};
+
/*
Hybrid perf script includes a group of hybrid samples(LBRs + call stack),
which is used to generate CS profile. An example of hybrid sample:
diff --git a/llvm/tools/llvm-profgen/ProfiledBinary.cpp b/llvm/tools/llvm-profgen/ProfiledBinary.cpp
index 188fb2c20608c..f1bc278dca988 100644
--- a/llvm/tools/llvm-profgen/ProfiledBinary.cpp
+++ b/llvm/tools/llvm-profgen/ProfiledBinary.cpp
@@ -26,6 +26,7 @@
#include "llvm/Support/Format.h"
#include "llvm/Support/TargetSelect.h"
#include "llvm/TargetParser/Triple.h"
+#include <algorithm>
#include <optional>
#define DEBUG_TYPE "load-binary"
@@ -659,6 +660,12 @@ bool ProfiledBinary::dissassembleSymbol(std::size_t SI, ArrayRef<uint8_t> Bytes,
if (MCDesc.isUnconditionalBranch())
UncondBranchAddrSet.insert(Address);
BranchAddressSet.insert(Address);
+ } else if (MCDesc.isBarrier() || MCDesc.isTrap()) {
+ // Transfers of control are handled above, so what reaches here stops
+ // execution without naming where it continues, the AArch64 BRK and UDF
+ // above all. Recording it lets range inference end a range there
+ // instead of running it into code the trap never let execute.
+ BarrierAddressSet.insert(Address);
}
if (MCDesc.isIndirectBranch()) {
@@ -1299,6 +1306,105 @@ void ProfiledBinary::inferMissingFrames(
MissingContextInferrer->inferMissingFrames(Context, NewContext);
}
+void ProfiledBinary::buildRangeEnds() {
+ // The merge below and findRangeEnd's binary search both need increasing
+ // addresses, and instructions are collected in section-header order, which an
+ // object file need not keep sorted, so copy and sort when it is not.
+ //
+ // TODO: Sort CodeAddressVec in disassemble() and reduce this to an assertion.
+ // getIndexForAddr and InstructionPointer already depend on that order, so a
+ // binary listed out of it yields a wrong profile on those paths as well.
+ std::vector<uint64_t> SortedAddresses;
+ ArrayRef<uint64_t> Addresses = CodeAddressVec;
+ if (!llvm::is_sorted(CodeAddressVec)) {
+ SortedAddresses.assign(CodeAddressVec.begin(), CodeAddressVec.end());
+ llvm::sort(SortedAddresses);
+ Addresses = SortedAddresses;
+ }
+
+ // Walk instructions and function ranges backwards at the same time, so that
+ // classifying every instruction costs one merge step instead of a search.
+ auto FuncIt = StartAddrToFuncRangeMap.rbegin();
+ const FuncRange *NextRange = nullptr;
+ for (size_t I = Addresses.size(); I != 0; --I) {
+ uint64_t Current = Addresses[I - 1];
+
+ // Advance to the last function range that starts at or before Current, then
+ // keep it only if Current is inside that range.
+ while (FuncIt != StartAddrToFuncRangeMap.rend() && Current < FuncIt->first)
+ ++FuncIt;
+ const FuncRange *CurrentRange = nullptr;
+ if (FuncIt != StartAddrToFuncRangeMap.rend() &&
+ Current < FuncIt->second.EndAddress)
+ CurrentRange = &FuncIt->second;
+
+ // An instruction outside every function range is never a valid range end,
+ // because a range must not extend into code whose function is unknown, and
+ // findRangeEnd rejects such an address before the search.
+ if (CurrentRange) {
+ // The following instruction continues the range only when it is adjacent
+ // and belongs to the same function. NextRange is null for the highest
+ // address and for a following instruction no function range covers.
+ bool FollowsAdjacently = I < Addresses.size() &&
+ Addresses[I] - Current == getInstSize(Current);
+
+ // Comparing the owning function rather than the range keeps a function
+ // described by several adjacent ranges, as function splitting produces,
+ // in one inferred range. Ranges are owned by name, so two same-named
+ // functions laid out back to back are separated only by IsFuncEntry,
+ // which a symbol on the second one sets.
+ bool EndsRange = !NextRange || !FollowsAdjacently ||
+ NextRange->Func != CurrentRange->Func ||
+ (NextRange != CurrentRange && NextRange->IsFuncEntry);
+ if (EndsRange || addressIsTransfer(Current) ||
+ BarrierAddressSet.count(Current))
+ RangeEnds.push_back(Current);
+ }
+ NextRange = CurrentRange;
+ }
+ // The traversal is backwards, so restore increasing order for binary search.
+ std::reverse(RangeEnds.begin(), RangeEnds.end());
+ RangeEndsBuilt = true;
+}
+
+uint64_t ProfiledBinary::findRangeEnd(uint64_t Address, bool *Uncovered) {
+ // Assigned on every path, so that a caller never reads a value left over
+ // from an earlier query.
+ if (Uncovered)
+ *Uncovered = false;
+
+ if (!addressIsCode(Address))
+ return 0;
+
+ // Function ranges come from debug info, and from the symbol table only where
+ // the binary carries pseudo probes. Where none covers Address, the executed
+ // range cannot be proven, so fail closed and keep that instruction alone.
+ if (!findFuncRange(Address)) {
+ if (Uncovered)
+ *Uncovered = true;
+ return Address;
+ }
+
+ // Every instruction before a range end has that end as its own, so indexing
+ // the ends alone keeps the index proportional to their number.
+ if (!RangeEndsBuilt)
+ buildRangeEnds();
+
+ // The greatest covered address always ends a range, and every covered address
+ // is at or below it, so the search finds an end.
+ auto End = llvm::lower_bound(RangeEnds, Address);
+ assert(End != RangeEnds.end() &&
+ "an address a function range covers has a range end at or above it");
+ // Should it not, fail closed as an uncovered address does rather than read
+ // past the index or hide the reduction from the caller.
+ if (End == RangeEnds.end()) {
+ if (Uncovered)
+ *Uncovered = true;
+ return Address;
+ }
+ return *End;
+}
+
InstructionPointer::InstructionPointer(const ProfiledBinary *Binary,
uint64_t Address, bool RoundToNext)
: Binary(Binary), Address(Address) {
diff --git a/llvm/tools/llvm-profgen/ProfiledBinary.h b/llvm/tools/llvm-profgen/ProfiledBinary.h
index e4af7c0fdf7f4..36949c89c05e3 100644
--- a/llvm/tools/llvm-profgen/ProfiledBinary.h
+++ b/llvm/tools/llvm-profgen/ProfiledBinary.h
@@ -283,6 +283,13 @@ class ProfiledBinary {
// An array of Addresses of all instructions sorted in increasing order. The
// sorting is needed to fast advance to the next forward/backward instruction.
std::vector<uint64_t> CodeAddressVec;
+ // Addresses that end an inferred executable range: transfer instructions and
+ // the last instructions before a code gap or a function boundary. Sorted in
+ // increasing order and filled on the first findRangeEnd query, so an input
+ // that needs no range inference never pays for it.
+ std::vector<uint64_t> RangeEnds;
+ // Whether RangeEnds has been built, which its emptiness does not tell.
+ bool RangeEndsBuilt = false;
// A set of call instruction addresses. Used by virtual unwinding.
DenseSet<uint64_t> CallAddressSet;
// A set of return instruction addresses. Used by virtual unwinding.
@@ -295,6 +302,10 @@ class ProfiledBinary {
DenseSet<uint64_t> IndirectBranchAddressSet;
// A set of branch target addresses (destinations of branches/calls).
DenseSet<uint64_t> BranchTargetAddressSet;
+ // A set of the addresses of instructions that execution never continues past,
+ // other than the calls, returns and branches held above: a trap such as the
+ // AArch64 BRK and UDF above all. buildRangeEnds ends an inferred range there.
+ DenseSet<uint64_t> BarrierAddressSet;
// Estimate and track function prolog and epilog ranges.
PrologEpilogTracker ProEpilogTracker;
@@ -421,6 +432,9 @@ class ProfiledBinary {
bool dissassembleSymbol(std::size_t SI, ArrayRef<uint8_t> Bytes,
SectionSymbolsTy &Symbols,
const object::SectionRef &Section);
+ /// Collect the addresses that end an inferred executable range into
+ /// RangeEnds. Called on the first findRangeEnd query.
+ void buildRangeEnds();
/// Symbolize a given instruction pointer and return a full call context.
SampleContextFrameVector symbolize(const InstructionPointer &IP,
bool UseCanonicalFnName = false,
@@ -505,11 +519,27 @@ class ProfiledBinary {
bool addressIsIndirectBranch(uint64_t Address) const {
return IndirectBranchAddressSet.count(Address);
}
+ // Whether Address is a direct branch that may fall through.
+ bool addressIsConditionalBranch(uint64_t Address) const {
+ return BranchAddressSet.count(Address) &&
+ !UncondBranchAddrSet.count(Address) &&
+ !IndirectBranchAddressSet.count(Address);
+ }
bool addressIsTransfer(uint64_t Address) {
return BranchAddressSet.count(Address) || RetAddressSet.count(Address) ||
CallAddressSet.count(Address);
}
+ // Return the end of the executable range that starts at Address: the nearest
+ // instruction at or after it that is a transfer or a trap, without crossing a
+ // code gap or a function boundary, or the last instruction before either
+ // boundary. Return zero when Address is not an instruction, and Address
+ // itself when no function range covers it, which keeps an inferred range out
+ // of code that may belong to another function. Set *Uncovered, when given, to
+ // whether the range had to be reduced to Address; it is assigned on every
+ // path.
+ uint64_t findRangeEnd(uint64_t Address, bool *Uncovered = nullptr);
+
bool rangeCrossUncondBranch(uint64_t Start, uint64_t End) {
if (Start >= End)
return false;
diff --git a/llvm/tools/llvm-profgen/llvm-profgen.cpp b/llvm/tools/llvm-profgen/llvm-profgen.cpp
index 31fda19ee8456..ed4f2db8b8fd4 100644
--- a/llvm/tools/llvm-profgen/llvm-profgen.cpp
+++ b/llvm/tools/llvm-profgen/llvm-profgen.cpp
@@ -33,9 +33,13 @@ static cl::opt<std::string> PerfScriptFilename(
"perfscript", cl::value_desc("perfscript"),
cl::desc("Path of a trace created by the Linux `perf script` command. For "
"LBR or BRBE input, the raw perf data must contain branch "
- "stacks, for example from recording with -b. "
- "Cannot be used with --perfdata, --unsymbolized-profile, or "
- "--llvm-sample-profile."),
+ "stacks, for example from recording with -b. For "
+ "--spe-branch-profile input, the trace must contain only Arm SPE "
+ "branch samples, from recording with "
+ "arm_spe/branch_filter=1,event_filter=2/, and must be generated "
+ "with `--show-mmap-events --itrace=bl1 -F ip,brstack`. Cannot be "
+ "used with --perfdata, "
+ "--unsymbolized-profile, or --llvm-sample-profile."),
cl::cat(ProfGenCategory));
static cl::alias PSA("ps", cl::desc("Alias for --perfscript"),
cl::aliasopt(PerfScriptFilename));
@@ -44,8 +48,11 @@ static cl::opt<std::string> PerfDataFilename(
"perfdata", cl::value_desc("perfdata"),
cl::desc("Path of raw perf data created by the Linux perf tool. For LBR or "
"BRBE input, it must contain branch stacks, for example from "
- "recording with -b. Cannot be used with --perfscript, "
- "--unsymbolized-profile, or --llvm-sample-profile."),
+ "recording with -b. For --spe-branch-profile input, it must "
+ "contain only Arm SPE branch samples, from recording with "
+ "arm_spe/branch_filter=1,event_filter=2/. Cannot be used with "
+ "--perfscript, --unsymbolized-profile, or "
+ "--llvm-sample-profile."),
cl::cat(ProfGenCategory));
static cl::alias PDA("pd", cl::desc("Alias for --perfdata"),
cl::aliasopt(PerfDataFilename));
@@ -104,11 +111,24 @@ static cl::opt<std::string>
// Validate the command line input.
static void validateCommandLine() {
+ bool HasPerfData = PerfDataFilename.getNumOccurrences() > 0;
+ bool HasPerfScript = PerfScriptFilename.getNumOccurrences() > 0;
+
+ // Arm SPE branch profiling only describes how a perf data or perfscript
+ // input is parsed. Reject the option for every other mode, where it would
+ // otherwise be silently meaningless.
+ if (ReadSPEBranchProfile) {
+ if (ShowDisassemblyOnly)
+ exitWithError("--spe-branch-profile cannot be used with "
+ "--show-disassembly-only.");
+ if (!HasPerfData && !HasPerfScript)
+ exitWithError("--spe-branch-profile requires --perfscript or "
+ "--perfdata input.");
+ }
+
// Allow the missing perfscript if we only use to show binary disassembly.
if (!ShowDisassemblyOnly) {
// Validate input profile is provided only once
- bool HasPerfData = PerfDataFilename.getNumOccurrences() > 0;
- bool HasPerfScript = PerfScriptFilename.getNumOccurrences() > 0;
bool HasUnsymbolizedProfile =
UnsymbolizedProfFilename.getNumOccurrences() > 0;
bool HasSampleProfile = SampleProfFilename.getNumOccurrences() > 0;
@@ -192,6 +212,15 @@ int main(int argc, const char *argv[]) {
std::make_unique<ProfiledBinary>(BinaryPath, DebugBinPath);
Binary->load(TargetTriple);
+ // Arm SPE samples describe AArch64 branch execution, so reject another
+ // target.
+ //
+ // TODO: --target-triple overrides the triple checked here, while the
+ // disassembler is looked up from the object file itself, so an override makes
+ // the two disagree.
+ if (ReadSPEBranchProfile && !Binary->getTriple().isAArch64())
+ exitWithError("--spe-branch-profile requires an AArch64 binary.");
+
if (ShowDisassemblyOnly)
return EXIT_SUCCESS;
More information about the llvm-commits
mailing list