[llvm] [llvm-profgen] Add Arm SPE branch profile support (PR #223237)

Sergey Shcherbinin via llvm-commits llvm-commits at lists.llvm.org
Sun Sep 13 04:54:16 PDT 2026


https://github.com/SergeyShch01 created https://github.com/llvm/llvm-project/pull/223237

Add `--spe-branch-profile` for generating AArch64 sample profiles from Arm SPE branch records.
Input may be raw `perf.data` or `perf script --show-mmap-events --itrace=bl1 -F ip,brstack` output.
Recording requires `arm_spe/branch_filter=1,event_filter=2/`, perf 6.15+, FEAT_SPEv1p2, and event filtering.

ArmSPEReader derives from LBRPerfReader and reuses the existing LBR/BRBE aggregation and profile-generation path.
perf provides each SPE record as a branch stack containing one entry (the second address of its executed range is then inferred as the next branching point by llvm-profgen). 

Existing LBR/BRBE profile generation is unchanged; mismatch diagnostics now use aggregated sample weights.

Tests cover parsing, perf-data conversion, mmap relocation, taken state, range inference, symbolization, and target validation.
Documentation and release notes describe the option and recording requirements.

Assisted by GPT-5

>From 0c880df627bcf5671ca0e5a6abb97642580fda52 Mon Sep 17 00:00:00 2001
From: Sergey Shcherbinin <sscherbinin at nvidia.com>
Date: Sun, 13 Sep 2026 14:48:00 +0400
Subject: [PATCH] [llvm-profgen] Add Arm SPE branch profile support

---
 llvm/docs/CommandGuide/llvm-profgen.md        |  15 +-
 llvm/docs/ReleaseNotes.md                     |   3 +
 .../AArch64/spe-branch-profile.test           | 484 +++++++++++
 .../AArch64/spe-not-taken-branch.test         | 232 ++++++
 .../llvm-profgen/AArch64/spe-perfdata.test    | 135 +++
 .../llvm-profgen/AArch64/spe-range-gap.test   | 774 ++++++++++++++++++
 .../AArch64/spe-symbolized-profile.test       | 109 +++
 .../llvm-profgen/spe-target-validation.test   |  45 +
 llvm/tools/llvm-profgen/Options.h             |   1 +
 llvm/tools/llvm-profgen/PerfReader.cpp        | 373 ++++++++-
 llvm/tools/llvm-profgen/PerfReader.h          | 130 ++-
 llvm/tools/llvm-profgen/ProfiledBinary.cpp    | 106 +++
 llvm/tools/llvm-profgen/ProfiledBinary.h      |  30 +
 llvm/tools/llvm-profgen/llvm-profgen.cpp      |  43 +-
 14 files changed, 2430 insertions(+), 50 deletions(-)
 create mode 100644 llvm/test/tools/llvm-profgen/AArch64/spe-branch-profile.test
 create mode 100644 llvm/test/tools/llvm-profgen/AArch64/spe-not-taken-branch.test
 create mode 100644 llvm/test/tools/llvm-profgen/AArch64/spe-perfdata.test
 create mode 100644 llvm/test/tools/llvm-profgen/AArch64/spe-range-gap.test
 create mode 100644 llvm/test/tools/llvm-profgen/AArch64/spe-symbolized-profile.test
 create mode 100644 llvm/test/tools/llvm-profgen/spe-target-validation.test

diff --git a/llvm/docs/CommandGuide/llvm-profgen.md b/llvm/docs/CommandGuide/llvm-profgen.md
index 631a3e00f3ce2..a1a49f2140314 100644
--- a/llvm/docs/CommandGuide/llvm-profgen.md
+++ b/llvm/docs/CommandGuide/llvm-profgen.md
@@ -20,7 +20,10 @@ At least one of the following commands are required:
 :::{option} --perfscript=<string[,string,...]>
 Path of a trace created by the Linux `perf script` command. For LBR or BRBE
 input, the raw perf data must contain branch stacks, for example from recording
-with `-b`.
+with `-b`. With `--spe-branch-profile`, the trace must contain only Arm SPE
+branch samples, for example from recording with
+`arm_spe/branch_filter=1,event_filter=2/`, and must be generated with
+`--show-mmap-events --itrace=bl1 -F ip,brstack`.
 :::
 
 :::{option} --etm=<string>
@@ -30,7 +33,9 @@ Requires the OpenCSD library version 1.5.4 or higher to be enabled during the bu
 
 :::{option} --perfdata=<perfdata>, --pd
 Path of raw perf data created by the Linux perf tool. For LBR or BRBE input, it
-must contain branch stacks, for example from recording with `-b`.
+must contain branch stacks, for example from recording with `-b`. For
+`--spe-branch-profile` input, it must contain only Arm SPE branch samples, from
+recording with `arm_spe/branch_filter=1,event_filter=2/`.
 :::
 
 :::{option} --unsymbolized-profile=<unsymbolized profile>, --up
@@ -69,6 +74,12 @@ descriptions of the format.
 Print mmap events.
 :::
 
+:::{option} --spe-branch-profile
+Read the `--perfscript` or `--perfdata` input as an Arm SPE branch profile of
+an AArch64 binary. Requires perf 6.15 or later and hardware with FEAT_SPEv1p2
+and event filtering.
+:::
+
 :::{option} --warn-not-symbolized
 Warn when an address covered by a recorded mmap range cannot be symbolized.
 :::
diff --git a/llvm/docs/ReleaseNotes.md b/llvm/docs/ReleaseNotes.md
index 1fd90f5bd3be0..62c2e4e61c65d 100644
--- a/llvm/docs/ReleaseNotes.md
+++ b/llvm/docs/ReleaseNotes.md
@@ -282,6 +282,9 @@ Makes programs 10x faster by doing Special New Thing.
 
 * llvm-mca no longer defaults -mcpu to "native"
 
+* llvm-profgen can now build a sample profile from an Arm SPE branch profile of
+  an AArch64 binary, with the new `--spe-branch-profile` option
+
 * llvm-rc now supports `/showIncludes` to report header and resource-file
   dependencies in a format compatible with Ninja's `deps = msvc` mode.
 
diff --git a/llvm/test/tools/llvm-profgen/AArch64/spe-branch-profile.test b/llvm/test/tools/llvm-profgen/AArch64/spe-branch-profile.test
new file mode 100644
index 0000000000000..d7500b49f91d5
--- /dev/null
+++ b/llvm/test/tools/llvm-profgen/AArch64/spe-branch-profile.test
@@ -0,0 +1,484 @@
+# RUN: split-file %s %t
+# RUN: yaml2obj %t/binary.yaml -o %t/spe-branch-profile.exe
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/perfscript --spe-branch-profile --skip-symbolization \
+# RUN:   --use-offset=0 --format=text --output=%t/profile 2>&1 \
+# RUN:   | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --input-file=%t/profile
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/no-ip.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --use-offset=0 --format=text \
+# RUN:   --output=%t/no-ip.profile 2>&1 \
+# RUN:   | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --check-prefix=NO-IP --input-file=%t/no-ip.profile
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/malformed.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --show-detailed-warning --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=INVALID
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/malformed.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=INVALID-SUMMARY
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/bad-fallthrough.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --use-offset=0 --format=text \
+# RUN:   --output=%t/bad-fallthrough.profile 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=BAD-FALLTHROUGH
+# RUN: FileCheck %s --check-prefix=LAST-INST \
+# RUN:   --input-file=%t/bad-fallthrough.profile
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/weighted-mismatch.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=WEIGHTED-MISMATCH
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/inconsistent-not-taken.perfscript \
+# RUN:   --spe-branch-profile --skip-symbolization --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=INCONSISTENT-NOT-TAKEN
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/empty.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=EMPTY
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/unusable.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=NO-USABLE
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/partial.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --use-offset=0 --format=text \
+# RUN:   --output=%t/partial.profile 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=PARTIAL
+# RUN: FileCheck %s --check-prefix=PARTIAL-PROFILE \
+# RUN:   --input-file=%t/partial.profile
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/perfscript --spe-branch-profile --pid=2 \
+# RUN:   --skip-symbolization --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=OTHER-PID
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/external-source.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --use-offset=0 --format=text \
+# RUN:   --output=%t/external-source.profile
+# RUN: FileCheck %s --check-prefix=EXTERNAL-SOURCE \
+# RUN:   --input-file=%t/external-source.profile
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/external-source.perfscript --spe-branch-profile \
+# RUN:   --use-offset=0 --format=text --output=%t/external-source.symbolized
+# RUN: FileCheck %s --check-prefix=SYMBOLIZED \
+# RUN:   --input-file=%t/external-source.symbolized
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/no-mmap.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=NO-MMAP
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/previous-target.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=DEEP-STACK
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/lbr.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=DEEP-STACK
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/damaged-line.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --show-detailed-warning --use-offset=0 \
+# RUN:   --format=text --output=%t/damaged-line.profile 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=DAMAGED
+# RUN: FileCheck %s --check-prefix=DAMAGED-PROFILE \
+# RUN:   --input-file=%t/damaged-line.profile
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/hybrid.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --show-detailed-warning --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=HYBRID
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/ip-addr.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=IP-ADDR
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/rejected.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --use-offset=0 --format=text \
+# RUN:   --output=%t/rejected.profile 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=REJECTED
+# RUN: FileCheck %s --check-prefix=REJECTED-PROFILE \
+# RUN:   --input-file=%t/rejected.profile
+# RUN: yaml2obj -DTYPE=ET_EXEC %t/binary.yaml -o %t/non-pie.exe
+# RUN: llvm-profgen --binary=%t/non-pie.exe \
+# RUN:   --perfscript=%t/non-pie.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --use-offset=0 --format=text \
+# RUN:   --output=%t/non-pie.profile 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=NON-PIE
+# RUN: FileCheck %s --check-prefix=NON-PIE-PROFILE \
+# RUN:   --input-file=%t/non-pie.profile
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/perfscript --spe-branch-profile --skip-symbolization \
+# RUN:   --target-triple=x86_64-unknown-linux-gnu --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=WRONG-TRIPLE
+# RUN: llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/perfscript --skip-symbolization --use-offset=0 \
+# RUN:   --format=text --output=%t/lbr.profile
+# RUN: FileCheck %s --check-prefix=WITHOUT-OPTION --input-file=%t/lbr.profile
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --unsymbolized-profile=%t/empty.perfscript --spe-branch-profile \
+# RUN:   --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=INCOMPATIBLE-INPUT
+# RUN: not llvm-profgen --binary=%t/spe-branch-profile.exe \
+# RUN:   --perfscript=%t/perfscript --spe-branch-profile --show-disassembly-only \
+# RUN:   --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=DISASM-ONLY
+
+## perf prints an Arm SPE branch as a branch stack of one entry preceded by
+## the instruction pointer of the sample. Use different source and destination
+## addresses to verify their order, canonicalize them through the runtime
+## MMAP, and infer the executed range up to the nearest transfer. The first
+## record is given three times, because an Arm SPE trace carries no repeat
+## count. The last record is a not-taken conditional branch, which perf marks
+## with the N flag and reports with the fall-through instruction as its
+## destination.
+# NO-WARN-NOT: warning: {{.*}}branch targets do not match the binary
+# NO-WARN-NOT: warning: Invalid Arm SPE branch record
+## A trace whose records yield counters must not be reported as yielding none,
+## and the report of the branch-stack readers, which describes ranges bounded
+## by consecutive entries, must not run for records that hold one entry each.
+# NO-WARN-NOT: warning: No Arm SPE branch record yields a counter
+# NO-WARN-NOT: warning: No samples in perf script
+
+## The not-taken record contributes the range that follows it, because
+## execution did continue at the fall-through instruction, but no branch
+## counter, because no control was transferred.
+# CHECK:      2
+# CHECK-NEXT: 1000-1008:3
+# CHECK-NEXT: 100c-100c:1
+# CHECK-NEXT: 1
+# CHECK-NEXT: 1008->1000:3
+
+## The instruction pointer of a record repeats the source of its branch, so a
+## trace printed with `-F brstack` alone holds every address the record needs
+## and is read the same way.
+# NO-IP:      1
+# NO-IP-NEXT: 1000-1008:1
+# NO-IP-NEXT: 1
+# NO-IP-NEXT: 1008->1000:1
+
+## Name each malformed record once when asked for detailed warnings: a record
+## holding the instruction pointer alone, a record followed by a trailing
+## field, as a wrong -F field list such as ip,brstack,sym produces, a record
+## whose addresses carry an appended shared-object path, as adding dso to that
+## list produces, and a record whose addresses carry no 0x prefix. A trailing
+## field is read as a branch-stack entry before a record is called a deeper
+## stack, or a field holding the slash that separates the addresses of an entry
+## would abort the run and blame an --itrace option that is already right. The
+## external-to-external record is valid input but cannot contribute to this
+## binary's profile. The last record holds an instruction pointer of decimal
+## digits alone, which is a repeat count of the branch-stack grammar and has to
+## be read as the damaged record it is instead, or the record that follows it
+## would be weighed by an address.
+# INVALID:      warning: Invalid Arm SPE branch record at line 2: 70000000100a
+# INVALID-NEXT: warning: Invalid Arm SPE branch record at line 3: 700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/- foo
+# INVALID-NEXT: warning: Invalid Arm SPE branch record at line 4: 700000001008 0x700000001008(/tmp/spe-branch-profile.exe)/0x700000001000(/tmp/spe-branch-profile.exe)/P/-/-/10/COND/-
+# INVALID-NEXT: warning: Invalid Arm SPE branch record at line 5: 700000001008 700000001008/700000001000/P/-/-/10/COND/-
+# INVALID-NEXT: warning: Invalid Arm SPE branch record at line 6: 700000001004
+# INVALID-NOT:  warning: Invalid Arm SPE branch record
+# INVALID-NOT:  error:
+
+## Without that request the lines are reported as a share instead, because a
+## trace printed with a wrong field list holds millions of such lines. Every
+## share is of the same base, the sample lines of the trace, so that the
+## reports can be read against one another. The one record that was read is
+## dropped, which leaves the profile empty and is reported as well.
+# INVALID-SUMMARY-NOT:  warning: Invalid Arm SPE branch record
+# INVALID-SUMMARY:      warning: 83.33%(5/6) of the sample lines of the trace are not usable Arm SPE branch records.
+# INVALID-SUMMARY-NEXT: warning: 16.67%(1/6) of the sample lines of the trace are Arm SPE branch records whose destination is not an instruction of this binary.
+# INVALID-SUMMARY-NEXT: warning: No Arm SPE branch record yields a counter for this binary!
+
+## A trace whose every line holds two bare addresses was printed with the
+## ip,addr field list of the perf versions that synthesize no branch stack for
+## Arm SPE. That is only known once the whole trace has been read. Name the
+## perf command to use, not the individual lines.
+# IP-ADDR-NOT: warning:
+# IP-ADDR:     error: --spe-branch-profile requires Arm SPE branch records created with `perf script --show-mmap-events --itrace=bl1 -F ip,brstack`, but no line of the input holds a usable one
+
+## Each way a line can fail to hold a usable record leaves the records around
+## it in place: a prediction field holding the value of another field, as a
+## shifted -F list prints; an entry printed without the prediction field; an
+## entry cut short right after the prediction field, whose P may be all that is
+## left of a PN; a source address of zero, which names no instruction; an
+## instruction pointer that is not the source of the entry that follows it, so
+## that neither address can be trusted; and an instruction pointer carrying the
+## 0x prefix that perf prints on the addresses of an entry alone. The one valid
+## record still yields its counters.
+# REJECTED:     warning: 85.71%(6/7) of the sample lines of the trace are not usable Arm SPE branch records.
+# REJECTED-NOT: error:
+# REJECTED-PROFILE:      1
+# REJECTED-PROFILE-NEXT: 1000-1008:1
+# REJECTED-PROFILE-NEXT: 1
+# REJECTED-PROFILE-NEXT: 1008->1000:1
+
+## The addresses of a binary that is not position independent are already the
+## ones of the trace, so they are used as they are and the missing mmap event
+## is only reported.
+# NON-PIE: warning: No relevant mmap event is matched for non-pie.exe
+# NON-PIE-PROFILE:      1
+# NON-PIE-PROFILE-NEXT: 1000-1008:1
+# NON-PIE-PROFILE-NEXT: 1
+# NON-PIE-PROFILE-NEXT: 1008->1000:1
+
+## A direct unconditional branch cannot fall through. Diagnose a recorded
+## destination at the next instruction when it is not the encoded target, and
+## keep its branch counter, because control was transferred there.
+# BAD-FALLTHROUGH: of branch targets do not match the binary.
+
+## Weigh both the mismatch and the denominator by each aggregated sample.
+# WEIGHTED-MISMATCH: warning: 66.67%(2/3) of branch targets do not match the binary.
+
+## A not-taken record must name a conditional branch and its fall-through.
+# INCONSISTENT-NOT-TAKEN: warning: 50.00%(1/2) of branch samples are marked not taken but do not name a conditional branch.
+# INCONSISTENT-NOT-TAKEN-NEXT: warning: 50.00%(1/2) of branch targets do not match the binary.
+
+## No transfer follows the destination of the same record, so its range runs to
+## the last instruction of the binary, where it ends for want of a following
+## instruction.
+# LAST-INST:      1
+# LAST-INST-NEXT: 1010-1014:1
+# LAST-INST-NEXT: 1
+# LAST-INST-NEXT: 100c->1010:1
+
+## A trace that holds no Arm SPE record at all is the wrong input file. Fail
+## instead of writing an empty profile and reporting success, and say which of
+## the two reasons applies: here the trace holds no line to read at all.
+# EMPTY: error: --spe-branch-profile requires Arm SPE branch records created with `perf script --show-mmap-events --itrace=bl1 -F ip,brstack`, but the trace holds no record at all
+
+## A branch from this binary to external code is valid input, but a single SPE
+## record with an external target cannot produce a branch or range counter.
+# NO-USABLE:      warning: 100.00%(1/1) of the sample lines of the trace are Arm SPE branch records whose destination is not an instruction of this binary.
+# NO-USABLE-NEXT: warning: No Arm SPE branch record yields a counter for this binary!
+
+## A record whose destination is not an instruction of this binary is dropped:
+## it reports a branch leaving the binary, a trace of another process or of
+## another build of the program, or a record whose branch target address was
+## lost. Report the share they take of the sample lines of the trace, so that a
+## profile thinned by a trace that does not belong to this binary is not
+## mistaken for a complete one.
+# PARTIAL: warning: 66.67%(2/3) of the sample lines of the trace are Arm SPE branch records whose destination is not an instruction of this binary.
+# PARTIAL-PROFILE:      1
+# PARTIAL-PROFILE-NEXT: 1000-1008:1
+# PARTIAL-PROFILE-NEXT: 1
+# PARTIAL-PROFILE-NEXT: 1008->1000:1
+
+## Only an mmap event of a perf script trace carries a process id, so --pid
+## selects the mapping that addresses are canonicalized against and does not
+## filter the records themselves. Another id therefore leaves every record of
+## this trace pointing outside the binary.
+# OTHER-PID: warning: No relevant mmap event is matched for
+# OTHER-PID: warning: 100.00%(4/4) of the sample lines of the trace are Arm SPE branch records whose destination is not an instruction of this binary.
+# OTHER-PID: warning: No Arm SPE branch record yields a counter for this binary!
+
+## The records of an Arm SPE trace are the branch-stack lines of an LBR trace,
+## holding one entry each, so the same trace is also valid input without the
+## option. It is then read as an LBR trace, which yields the branch counters
+## alone: no range is inferred, because where the range that follows the most
+## recent branch ends is what only the Arm SPE reader recovers from the binary.
+## The N flag that marks a branch as not taken is read by the Arm SPE reader
+## alone, because a hardware branch stack records taken branches only, so the
+## LBR reader counts the not-taken record as an ordinary branch here.
+# WITHOUT-OPTION:      0
+# WITHOUT-OPTION-NEXT: 2
+# WITHOUT-OPTION-NEXT: 1008->1000:3
+# WITHOUT-OPTION-NEXT: 1008->100c:1
+
+## Preserve an external source for a branch entering this binary so downstream
+## symbolization can recognize it as a function-entry sample.
+# EXTERNAL-SOURCE:      1
+# EXTERNAL-SOURCE-NEXT: 1000-1008:1
+# EXTERNAL-SOURCE-NEXT: 1
+# EXTERNAL-SOURCE-NEXT: 1->1000:1
+
+## Symbolization converts the external-to-internal branch into a head sample.
+# SYMBOLIZED: foo:0:1{{$}}
+
+## Without a matching mmap event every address is interpreted against the
+## preferred base address, which leaves all records external and the profile
+## empty. Report the missing event as the cause instead of the empty profile
+## alone.
+# NO-MMAP: warning: No relevant mmap event is matched for
+# NO-MMAP: warning: No Arm SPE branch record yields a counter for this binary!
+
+## A record decoded with a deeper branch stack holds more than the sampled
+## branch: further sampled branches, or the entry perf synthesizes from the
+## target of the branch that preceded the sampled operation, which it prints
+## with the source address 0x0 and which, counted as a branch, fabricates one
+## from outside the binary. Name the perf command that decodes one entry per
+## record instead of accepting it.
+# DEEP-STACK-NOT: warning:
+# DEEP-STACK: error: --spe-branch-profile requires Arm SPE branch records created with `perf script --show-mmap-events --itrace=bl1 -F ip,brstack`, but the record at line 2 holds a branch stack of more than one entry
+
+## A single unreadable line inside a trace whose other records parse is damaged
+## input, not input of the wrong kind, whichever order the lines come in. Warn
+## about that line alone and keep the records around it.
+# DAMAGED:      warning: Invalid Arm SPE branch record at line 3: 700000001008 0x700000001008/0x
+# DAMAGED-NOT:  error:
+# DAMAGED-PROFILE:      2
+# DAMAGED-PROFILE-NEXT: 1000-1008:1
+# DAMAGED-PROFILE-NEXT: 100c-100c:1
+# DAMAGED-PROFILE-NEXT: 1
+# DAMAGED-PROFILE-NEXT: 1008->1000:1
+
+## The call stack frames of a hybrid trace hold no branch-stack entry, so they
+## are named as unreadable lines, and the deeper branch stack that follows them
+## is rejected the way any deeper stack is.
+# HYBRID:      warning: Invalid Arm SPE branch record at line 2: 4006ac
+# HYBRID-NEXT: warning: Invalid Arm SPE branch record at line 3: 40064c
+# HYBRID:      error: --spe-branch-profile requires Arm SPE branch records created with `perf script --show-mmap-events --itrace=bl1 -F ip,brstack`, but the record at line 4 holds a branch stack of more than one entry
+
+## The check uses the triple llvm-profgen was told to use, so an override that
+## names another target is rejected as well.
+# WRONG-TRIPLE: error: --spe-branch-profile requires an AArch64 binary.
+
+## Name the conflicting mode instead of the input the user already gave.
+# INCOMPATIBLE-INPUT: error: --spe-branch-profile requires --perfscript or --perfdata input.
+# DISASM-ONLY: error: --spe-branch-profile cannot be used with --show-disassembly-only.
+
+#--- binary.yaml
+--- !ELF
+FileHeader:
+  Class:   ELFCLASS64
+  Data:    ELFDATA2LSB
+  ## Built a second time as ET_EXEC to cover a binary that is not position
+  ## independent.
+  Type:    [[TYPE=ET_DYN]]
+  Machine: EM_AARCH64
+  Entry:   0x1000
+Sections:
+  - Name:         .text
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x1000
+    Offset:       0x1000
+    AddressAlign: 0x4
+    ## nop
+    ## nop
+    ## b.eq 0x1000
+    ## b 0x1000
+    ## nop
+    ## nop
+    Content:      1F2003D51F2003D5C0FFFF54FDFFFF171F2003D51F2003D5
+ProgramHeaders:
+  - Type:     PT_LOAD
+    Flags:    [ PF_X, PF_R ]
+    Offset:   0
+    VAddr:    0
+    FileSize: 0x1018
+    MemSize:  0x1018
+    Align:    0x10000
+Symbols:
+  - Name:    foo
+    Type:    STT_FUNC
+    Section: .text
+    Binding: STB_GLOBAL
+    Value:   0x1000
+    Size:    0x18
+DWARF:
+  debug_abbrev:
+    - ID: 0
+      Table:
+        - Code:     1
+          Tag:      DW_TAG_compile_unit
+          Children: DW_CHILDREN_yes
+          Attributes:
+            - Attribute: DW_AT_name
+              Form:      DW_FORM_string
+        - Code:     2
+          Tag:      DW_TAG_subprogram
+          Children: DW_CHILDREN_no
+          Attributes:
+            - Attribute: DW_AT_name
+              Form:      DW_FORM_string
+            - Attribute: DW_AT_low_pc
+              Form:      DW_FORM_addr
+            - Attribute: DW_AT_high_pc
+              Form:      DW_FORM_data8
+  debug_info:
+    - Version:       4
+      AbbrevTableID: 0
+      AddrSize:      8
+      Entries:
+        - AbbrCode: 1
+          Values:
+            - CStr: spe-branch-profile.c
+        - AbbrCode: 2
+          Values:
+            - CStr:  foo
+            - Value: 0x1000
+            - Value: 0x18
+        - AbbrCode: 0
+
+#--- perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+    700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+    700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+    700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+700000001008 0x700000001008/0x70000000100c/PN/-/-/10/COND/-
+#--- no-ip.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+0x700000001008/0x700000001000/P/-/-/10/COND/-
+#--- malformed.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+70000000100a
+700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/- foo
+700000001008 0x700000001008(/tmp/spe-branch-profile.exe)/0x700000001000(/tmp/spe-branch-profile.exe)/P/-/-/10/COND/-
+700000001008 700000001008/700000001000/P/-/-/10/COND/-
+700000001004
+700000003000 0x700000003000/0x700000003004/P/-/-/10/COND/-
+#--- bad-fallthrough.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+70000000100c 0x70000000100c/0x700000001010/P/-/-/10/UNCOND/-
+#--- weighted-mismatch.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+70000000100c 0x70000000100c/0x700000001000/P/-/-/10/UNCOND/-
+70000000100c 0x70000000100c/0x700000001010/P/-/-/10/UNCOND/-
+70000000100c 0x70000000100c/0x700000001010/P/-/-/10/UNCOND/-
+#--- inconsistent-not-taken.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+70000000100c 0x70000000100c/0x700000001010/PN/-/-/10/UNCOND/-
+700000001008 0x700000001008/0x700000001000/PN/-/-/10/COND/-
+#--- empty.perfscript
+#--- unusable.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000001008 0x700000001008/0x700000003000/P/-/-/10/COND/-
+#--- partial.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000001008 0x700000001008/0x700000003000/P/-/-/10/COND/-
+700000001008 0x700000001008/0x0/P/-/-/10/COND/-
+700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+#--- external-source.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000003000 0x700000003000/0x700000001000/P/-/-/10/CALL/-
+#--- no-mmap.perfscript
+700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+#--- previous-target.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/- 0x0/0x70000000100c/-/-/-/0//-
+#--- lbr.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000001008 0x700000001008/0x700000001000/P/-/-/0 0x70000000100c/0x700000001000/P/-/-/0 0x700000001008/0x70000000100c/P/-/-/0
+#--- damaged-line.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+700000001008 0x700000001008/0x
+700000001008 0x700000001008/0x70000000100c/PN/-/-/10/COND/-
+#--- hybrid.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+4006ac
+40064c
+70000000100c 0x70000000100c/0x700000001000/P/-/-/0 0x700000001008/0x70000000100c/P/-/-/0
+#--- ip-addr.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000001000 700000001008
+70000000100c 700000001008
+#--- rejected.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-branch-profile.exe
+700000001008 0x700000001008/0x700000001000/COND/-/-/10
+700000001008 0x700000001008/0x700000001000
+700000001008 0x700000001008/0x700000001000/P
+0 0x0/0x700000001000/P/-/-/10/COND/-
+700000001000 0x700000001008/0x700000001000/P/-/-/10/COND/-
+0x700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+700000001008 0x700000001008/0x700000001000/P/-/-/10/COND/-
+#--- non-pie.perfscript
+1008 0x1008/0x1000/P/-/-/10/COND/-
diff --git a/llvm/test/tools/llvm-profgen/AArch64/spe-not-taken-branch.test b/llvm/test/tools/llvm-profgen/AArch64/spe-not-taken-branch.test
new file mode 100644
index 0000000000000..d6295707b0aad
--- /dev/null
+++ b/llvm/test/tools/llvm-profgen/AArch64/spe-not-taken-branch.test
@@ -0,0 +1,232 @@
+## Arm SPE records a conditional branch that was not taken with the
+## instruction that follows it as its destination, and perf marks such an entry
+## with the N flag. Such a record transferred no control, so it must not be
+## counted as a branch, while the range that follows it did execute and is
+## counted. The binary below places the entry of a function right behind a
+## conditional branch, where counting the branch would invent a call of that
+## function and a head sample for it. The same two addresses without the flag
+## report a branch that was taken to the instruction that follows it, which is
+## a transfer and is counted, so the flag alone tells the two apart.
+
+# RUN: split-file %s %t
+# RUN: yaml2obj %t/binary.yaml -o %t/spe-not-taken-branch.exe
+# RUN: llvm-profgen --binary=%t/spe-not-taken-branch.exe \
+# RUN:   --perfscript=%t/fall-through.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --use-offset=0 --format=text \
+# RUN:   --output=%t/fall-through.profile 2>&1 \
+# RUN:   | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --check-prefix=FALL-THROUGH \
+# RUN:   --input-file=%t/fall-through.profile
+# RUN: llvm-profgen --binary=%t/spe-not-taken-branch.exe \
+# RUN:   --perfscript=%t/fall-through.perfscript --spe-branch-profile \
+# RUN:   --use-offset=0 --format=text --output=%t/fall-through.symbolized
+# RUN: FileCheck %s --check-prefix=HEAD-SAMPLE \
+# RUN:   --input-file=%t/fall-through.symbolized
+# RUN: llvm-profgen --binary=%t/spe-not-taken-branch.exe \
+# RUN:   --perfscript=%t/indirect.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --use-offset=0 --format=text \
+# RUN:   --output=%t/indirect.profile 2>&1 \
+# RUN:   | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --check-prefix=INDIRECT --input-file=%t/indirect.profile
+# RUN: llvm-profgen --binary=%t/spe-not-taken-branch.exe \
+# RUN:   --perfscript=%t/taken.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --use-offset=0 --format=text \
+# RUN:   --output=%t/taken.profile 2>&1 \
+# RUN:   | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --check-prefix=TAKEN --input-file=%t/taken.profile
+# RUN: llvm-profgen --binary=%t/spe-not-taken-branch.exe \
+# RUN:   --perfscript=%t/taken.perfscript --spe-branch-profile \
+# RUN:   --use-offset=0 --format=text --output=%t/taken.symbolized
+# RUN: FileCheck %s --check-prefix=TAKEN-HEAD --input-file=%t/taken.symbolized
+# RUN: llvm-profgen --binary=%t/spe-not-taken-branch.exe \
+# RUN:   --perfscript=%t/both.perfscript --spe-branch-profile \
+# RUN:   --skip-symbolization --use-offset=0 --format=text \
+# RUN:   --output=%t/both.profile 2>&1 \
+# RUN:   | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --check-prefix=BOTH --input-file=%t/both.profile
+
+## A destination that the instruction does not encode is expected of a branch
+## that was not taken, so it is not reported as a mismatch either.
+# NO-WARN-NOT: warning:
+
+## Both records start their range at the entry of bar, so the range is counted
+## twice, but only the call transferred control there.
+# FALL-THROUGH:      1
+# FALL-THROUGH-NEXT: 1008-100c:2
+# FALL-THROUGH-NEXT: 1
+# FALL-THROUGH-NEXT: 1000->1008:1
+
+## The head count of bar is the number of calls of it, one, and not the number
+## of records whose destination is its entry, two.
+# HEAD-SAMPLE: bar:0:1{{$}}
+
+## An indirect branch encodes no target at all, so its recorded destination is
+## never the one of a branch that was not taken, even where it is the
+## instruction that follows the branch. Count it as the transfer it is. The
+## branch-type field of the record reads RET because that is what the kernel
+## maps an indirect branch to; this reader takes the kind of the branch from
+## the binary and ignores that field. Its prediction field reads M, for a
+## branch perf found mispredicted, which is counted like one it found
+## predicted: only the N suffix of that field changes what a record means.
+# INDIRECT:      1
+# INDIRECT-NEXT: 1014-1014:1
+# INDIRECT-NEXT: 1
+# INDIRECT-NEXT: 1010->1014:1
+
+## A conditional branch whose encoded target is the instruction that follows it
+## does transfer control when it is taken, and the record carries no N flag.
+## Count the branch, so that the function starting at that instruction keeps
+## the head sample of the transfer.
+# TAKEN:      1
+# TAKEN-NEXT: 101c-101c:1
+# TAKEN-NEXT: 1
+# TAKEN-NEXT: 1018->101c:1
+# TAKEN-HEAD: quux:0:1{{$}}
+
+## The same two addresses stand for a taken and for a not-taken branch, which
+## only the flag tells apart, so aggregation must keep the two records separate.
+## Both ranges executed; only the taken record transferred control. The flag is
+## read in both of its spellings: PN above, for a branch perf found predicted,
+## and MN here, for one it found mispredicted.
+# BOTH:      1
+# BOTH-NEXT: 101c-101c:2
+# BOTH-NEXT: 1
+# BOTH-NEXT: 1018->101c:1
+
+#--- binary.yaml
+--- !ELF
+FileHeader:
+  Class:   ELFCLASS64
+  Data:    ELFDATA2LSB
+  Type:    ET_DYN
+  Machine: EM_AARCH64
+  Entry:   0x1000
+Sections:
+  - Name:         .text
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x1000
+    Offset:       0x1000
+    AddressAlign: 0x4
+    ## foo:
+    ##   bl 0x1008
+    ##   b.eq 0x1000
+    ## bar:
+    ##   nop
+    ##   ret
+    ## baz:
+    ##   br x0
+    ##   ret
+    ## qux:
+    ##   b.eq 0x101c
+    ## quux:
+    ##   ret
+    Content:      02000094E0FFFF541F2003D5C0035FD600001FD6C0035FD620000054C0035FD6
+ProgramHeaders:
+  - Type:     PT_LOAD
+    Flags:    [ PF_X, PF_R ]
+    Offset:   0
+    VAddr:    0
+    FileSize: 0x1020
+    MemSize:  0x1020
+    Align:    0x10000
+Symbols:
+  - Name:    foo
+    Type:    STT_FUNC
+    Section: .text
+    Binding: STB_GLOBAL
+    Value:   0x1000
+    Size:    0x8
+  - Name:    bar
+    Type:    STT_FUNC
+    Section: .text
+    Binding: STB_GLOBAL
+    Value:   0x1008
+    Size:    0x8
+  - Name:    baz
+    Type:    STT_FUNC
+    Section: .text
+    Binding: STB_GLOBAL
+    Value:   0x1010
+    Size:    0x8
+  - Name:    qux
+    Type:    STT_FUNC
+    Section: .text
+    Binding: STB_GLOBAL
+    Value:   0x1018
+    Size:    0x4
+  - Name:    quux
+    Type:    STT_FUNC
+    Section: .text
+    Binding: STB_GLOBAL
+    Value:   0x101c
+    Size:    0x4
+DWARF:
+  debug_abbrev:
+    - ID: 0
+      Table:
+        - Code:     1
+          Tag:      DW_TAG_compile_unit
+          Children: DW_CHILDREN_yes
+          Attributes:
+            - Attribute: DW_AT_name
+              Form:      DW_FORM_string
+        - Code:     2
+          Tag:      DW_TAG_subprogram
+          Children: DW_CHILDREN_no
+          Attributes:
+            - Attribute: DW_AT_name
+              Form:      DW_FORM_string
+            - Attribute: DW_AT_low_pc
+              Form:      DW_FORM_addr
+            - Attribute: DW_AT_high_pc
+              Form:      DW_FORM_data8
+  debug_info:
+    - Version:       4
+      AbbrevTableID: 0
+      AddrSize:      8
+      Entries:
+        - AbbrCode: 1
+          Values:
+            - CStr: spe-not-taken-branch.c
+        - AbbrCode: 2
+          Values:
+            - CStr:  foo
+            - Value: 0x1000
+            - Value: 0x8
+        - AbbrCode: 2
+          Values:
+            - CStr:  bar
+            - Value: 0x1008
+            - Value: 0x8
+        - AbbrCode: 2
+          Values:
+            - CStr:  baz
+            - Value: 0x1010
+            - Value: 0x8
+        - AbbrCode: 2
+          Values:
+            - CStr:  qux
+            - Value: 0x1018
+            - Value: 0x4
+        - AbbrCode: 2
+          Values:
+            - CStr:  quux
+            - Value: 0x101c
+            - Value: 0x4
+        - AbbrCode: 0
+
+#--- fall-through.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-not-taken-branch.exe
+700000001004 0x700000001004/0x700000001008/PN/-/-/10/COND/-
+700000001000 0x700000001000/0x700000001008/P/-/-/10/CALL/-
+#--- indirect.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-not-taken-branch.exe
+700000001010 0x700000001010/0x700000001014/M/-/-/10/RET/-
+#--- taken.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-not-taken-branch.exe
+700000001018 0x700000001018/0x70000000101c/P/-/-/10/COND/-
+#--- both.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-not-taken-branch.exe
+700000001018 0x700000001018/0x70000000101c/P/-/-/10/COND/-
+700000001018 0x700000001018/0x70000000101c/MN/-/-/10/COND/-
diff --git a/llvm/test/tools/llvm-profgen/AArch64/spe-perfdata.test b/llvm/test/tools/llvm-profgen/AArch64/spe-perfdata.test
new file mode 100644
index 0000000000000..98f1544fc6a9b
--- /dev/null
+++ b/llvm/test/tools/llvm-profgen/AArch64/spe-perfdata.test
@@ -0,0 +1,135 @@
+# REQUIRES: system-linux
+
+# RUN: split-file %s %t
+# RUN: yaml2obj %t/binary.yaml -o %t/spe-perfdata.exe
+# RUN: touch %t/perf.data
+# RUN: chmod +x %t/success/perf
+# RUN: chmod +x %t/no-record/perf
+# RUN: env PATH="%t/success%{pathsep}%{PATH}" \
+# RUN:   llvm-profgen --binary=%t/spe-perfdata.exe \
+# RUN:   --perfdata=%t/perf.data --spe-branch-profile --skip-symbolization \
+# RUN:   --use-offset=0 --format=text --output=%t/profile
+# RUN: FileCheck %s --input-file=%t/profile
+# RUN: env PATH="%t/no-record%{pathsep}%{PATH}" \
+# RUN:   not llvm-profgen --binary=%t/spe-perfdata.exe \
+# RUN:   --perfdata=%t/perf.data --spe-branch-profile --skip-symbolization \
+# RUN:   --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=NO-RECORD
+
+## The mock validates both perf invocations, including the decoding of branch
+## events into a branch stack of one entry and the field list required for Arm
+## SPE. The profile checks that the second invocation's output is consumed.
+# CHECK:      1
+# CHECK-NEXT: 1000-1008:1
+# CHECK-NEXT: 1
+# CHECK-NEXT: 1008->1000:1
+
+## A recording that holds no Arm SPE branch event yields a trace of mmap
+## events alone. The perf command that printed it is one this tool ran, so
+## name what has to be recorded rather than what has to be printed.
+# NO-RECORD: error: --spe-branch-profile found no Arm SPE branch record in the perf data; record one with arm_spe/branch_filter=1,event_filter=2/ using perf 6.15 or later
+
+#--- success/perf
+#!/bin/sh
+
+if [ "$1" != "script" ]; then
+  echo "unexpected perf arguments: $*" >&2
+  exit 2
+fi
+
+if [ "$2" = "--show-mmap-events" ]; then
+  if [ "$3" != "-F" ] || [ "$4" != "comm,pid" ] || [ "$5" != "-i" ] || \
+     [ ! -f "$6" ] || [ "$#" -ne 6 ]; then
+    echo "unexpected mmap arguments: $*" >&2
+    exit 2
+  fi
+  printf '%s\n' 'PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-perfdata.exe'
+elif [ "$2" = "--itrace=bl1" ]; then
+  if [ "$3" != "--show-mmap-events" ] || [ "$4" != "-F" ] || \
+     [ "$5" != "ip,brstack" ] || [ "$6" != "-i" ] || [ ! -f "$7" ] || \
+     [ "$8" != "--pid" ] || [ "$9" != "1" ] || [ "$#" -ne 9 ]; then
+    echo "unexpected sample arguments: $*" >&2
+    exit 2
+  fi
+  printf '%s\n' 'PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-perfdata.exe'
+  printf '%s\n' '700000001008 0x700000001008/0x700000001000/P/-/-/10/UNCOND/-'
+else
+  echo "unexpected perf arguments: $*" >&2
+  exit 2
+fi
+
+#--- no-record/perf
+#!/bin/sh
+
+## A recording without Arm SPE branch events: the mmap events are printed by
+## both invocations and no branch record by either.
+printf '%s\n' 'PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/spe-perfdata.exe'
+
+#--- binary.yaml
+--- !ELF
+FileHeader:
+  Class:   ELFCLASS64
+  Data:    ELFDATA2LSB
+  Type:    ET_DYN
+  Machine: EM_AARCH64
+  Entry:   0x1000
+Sections:
+  - Name:         .text
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x1000
+    Offset:       0x1000
+    AddressAlign: 0x4
+    ## nop
+    ## nop
+    ## b 0x1000
+    Content:      1F2003D51F2003D5FEFFFF17
+ProgramHeaders:
+  - Type:     PT_LOAD
+    Flags:    [ PF_X, PF_R ]
+    Offset:   0
+    VAddr:    0
+    FileSize: 0x100C
+    MemSize:  0x100C
+    Align:    0x10000
+Symbols:
+  - Name:    foo
+    Type:    STT_FUNC
+    Section: .text
+    Binding: STB_GLOBAL
+    Value:   0x1000
+    Size:    0xC
+DWARF:
+  debug_abbrev:
+    - ID: 0
+      Table:
+        - Code:     1
+          Tag:      DW_TAG_compile_unit
+          Children: DW_CHILDREN_yes
+          Attributes:
+            - Attribute: DW_AT_name
+              Form:      DW_FORM_string
+        - Code:     2
+          Tag:      DW_TAG_subprogram
+          Children: DW_CHILDREN_no
+          Attributes:
+            - Attribute: DW_AT_name
+              Form:      DW_FORM_string
+            - Attribute: DW_AT_low_pc
+              Form:      DW_FORM_addr
+            - Attribute: DW_AT_high_pc
+              Form:      DW_FORM_data8
+  debug_info:
+    - Version:       4
+      AbbrevTableID: 0
+      AddrSize:      8
+      Entries:
+        - AbbrCode: 1
+          Values:
+            - CStr: spe-perfdata.c
+        - AbbrCode: 2
+          Values:
+            - CStr:  foo
+            - Value: 0x1000
+            - Value: 0xC
+        - AbbrCode: 0
diff --git a/llvm/test/tools/llvm-profgen/AArch64/spe-range-gap.test b/llvm/test/tools/llvm-profgen/AArch64/spe-range-gap.test
new file mode 100644
index 0000000000000..fac0e4d8b10a6
--- /dev/null
+++ b/llvm/test/tools/llvm-profgen/AArch64/spe-range-gap.test
@@ -0,0 +1,774 @@
+# RUN: split-file %s %t
+# RUN: yaml2obj %t/binary.yaml -o %t/spe-range-gap.exe
+# RUN: llvm-profgen --binary=%t/spe-range-gap.exe \
+# RUN:   --perfscript=%t/perfscript --spe-branch-profile --skip-symbolization \
+# RUN:   --use-offset=0 --format=text --output=%t/profile 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=SHRUNK
+# RUN: FileCheck %s --input-file=%t/profile
+# RUN: yaml2obj %t/no-debug.yaml -o %t/no-debug.exe
+# RUN: llvm-profgen --binary=%t/no-debug.exe \
+# RUN:   --perfscript=%t/no-debug.perfscript --spe-branch-profile --skip-symbolization \
+# RUN:   --use-offset=0 --format=text --output=%t/no-debug-profile 2>&1 \
+# RUN:   | FileCheck %s --check-prefix=ALL-SHRUNK
+# RUN: FileCheck %s --check-prefix=NO-DEBUG --input-file=%t/no-debug-profile
+# RUN: yaml2obj %t/unordered.yaml -o %t/unordered.exe
+# RUN: llvm-profgen --binary=%t/unordered.exe \
+# RUN:   --perfscript=%t/unordered.perfscript --spe-branch-profile --skip-symbolization \
+# RUN:   --use-offset=0 --format=text --output=%t/unordered-profile 2>&1 \
+# RUN:   | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --check-prefix=UNORDERED --input-file=%t/unordered-profile
+
+## A blank line between records is not a record, so it is not diagnosed. The
+## one record whose destination lies in code that debug info does not describe
+## is reported as a share, because its range covers a single instruction rather
+## than the code that ran.
+# SHRUNK-NOT: warning: Invalid
+# SHRUNK:     warning: 7.69%(1/13) of the sample lines of the trace are Arm SPE branch records whose destination cannot be associated with a function, so only that instruction is profiled.
+
+## A binary that carries no debug info at all has every range reduced that way,
+## which the share states outright instead of leaving it to be inferred from a
+## profile of one-instruction ranges. The trace holds the same record twice, so
+## the two lines aggregate into one sample of weight two: the share counts the
+## lines of the trace, not the samples that remained after aggregation.
+# ALL-SHRUNK: warning: 100.00%(2/2) of the sample lines of the trace are Arm SPE branch records whose destination cannot be associated with a function, so only that instruction is profiled.
+
+## Ordering the instruction addresses is what keeps the two functions apart, so
+## neither range is reduced and nothing is reported.
+# NO-WARN-NOT: warning:
+
+## Each record below is checked by one paragraph, in address order of the
+## inferred range.
+##
+## The first record starts a range in a function that is adjacent to the next
+## function. Stop at the function boundary rather than at the transfer that
+## follows it.
+##
+## The second record starts a range in a function that is separated from the
+## following transfer by an address gap. Stop at the gap.
+##
+## The third record starts a range in a function that debug info describes as
+## two adjacent ranges. No symbol starts the second range, so it continues the
+## same function and execution runs through the boundary, reaching the transfer
+## in the second range instead of stopping at the end of the first. A symbol
+## that starts a function at the later range, as a split-out part of a function
+## carries, ends the range there instead; the fifth record covers that.
+##
+## The fourth record starts a range in a function whose two debug info ranges
+## are separated by an address gap. Stop at the gap even though both sides
+## belong to one function.
+##
+## The fifth record starts a range in a function that is adjacent to a
+## different function of the same name, as produced by static functions of
+## different compilation units. Debug info groups both under one name, so stop
+## at the entry of the second function rather than run into it.
+##
+## The sixth record has an external source and a destination in code that no
+## debug info describes. Reduce the range to the destination instead of
+## extending it into the next described function, and still count the branch.
+##
+## The seventh and the eighth record start ranges that a call and a return end,
+## which transfer execution just as a branch does.
+##
+## The ninth record starts a range in a described function that is immediately
+## followed by code no debug info describes. Stop at the last instruction of
+## the function rather than continue into code of an unknown function.
+##
+## The tenth record starts a range in a function that is adjacent to a
+## differently named function whose first range carries no symbol, as a part
+## split out of a function does. The change of function is what ends the range
+## there, because a range that no symbol starts is not marked as an entry.
+##
+## The eleventh record starts a range in a function that holds a trap. The trap
+## transfers control nowhere, so it is no branch, call or return, but execution
+## never continues past it, and the range ends there instead of running on to
+## the branch that ends the function.
+##
+## The twelfth record is a branch to itself, so its range starts at the branch
+## and the branch ends it, leaving a range of that one instruction.
+##
+## The thirteenth record is the back edge of a loop whose body holds a branch
+## of its own. The range starts at the top of the body and ends at that inner
+## branch, because what follows it ran only if it was not taken, which the
+## record does not say. A back edge therefore does not attest that the whole
+## body ran.
+# CHECK:      13
+# CHECK-NEXT: 500-500:1
+# CHECK-NEXT: 1000-1004:1
+# CHECK-NEXT: 2000-2004:1
+# CHECK-NEXT: 4000-400c:1
+# CHECK-NEXT: 6000-6004:1
+# CHECK-NEXT: 8000-8004:1
+# CHECK-NEXT: 9000-9004:1
+# CHECK-NEXT: a000-a004:1
+# CHECK-NEXT: b000-b004:1
+# CHECK-NEXT: c000-c004:1
+# CHECK-NEXT: d800-d804:1
+# CHECK-NEXT: d900-d900:1
+# CHECK-NEXT: da00-da04:1
+# CHECK-NEXT: 13
+# CHECK-NEXT: 1->500:1
+# CHECK-NEXT: 100c->1000:1
+# CHECK-NEXT: 3000->2000:1
+# CHECK-NEXT: 400c->4000:1
+# CHECK-NEXT: 7004->6000:1
+# CHECK-NEXT: 800c->8000:1
+# CHECK-NEXT: 900c->9000:1
+# CHECK-NEXT: a00c->a000:1
+# CHECK-NEXT: b00c->b000:1
+# CHECK-NEXT: c00c->c000:1
+# CHECK-NEXT: d80c->d800:1
+# CHECK-NEXT: d900->d900:1
+# CHECK-NEXT: da0c->da00:1
+
+## Function ranges come from debug info, which profiled binaries are expected
+## to carry. Where they are missing, an inferred range must not run past an
+## unknown function boundary, so it is reduced to the sampled instruction while
+## the branch itself is still counted. Both counters weigh two, because the
+## trace holds the record twice.
+# NO-DEBUG:      1
+# NO-DEBUG-NEXT: 1000-1000:2
+# NO-DEBUG-NEXT: 1
+# NO-DEBUG-NEXT: 1008->1000:2
+
+## Instructions are collected in section-header order, which an object file
+## need not keep in address order. Range inference orders them itself, so each
+## sampled branch destination is bounded by the branch that ends the function it
+## belongs to. Both functions are sampled: without that ordering the backward
+## walk over instructions loses the function whose section is listed first, and
+## its range would shrink to the sampled instruction alone.
+# UNORDERED: 2
+# UNORDERED-NEXT: 1000-1004:1
+# UNORDERED-NEXT: 2000-2004:1
+# UNORDERED-NEXT: 2
+# UNORDERED-NEXT: 1004->1000:1
+# UNORDERED-NEXT: 2004->2000:1
+
+#--- binary.yaml
+--- !ELF
+FileHeader:
+  Class:   ELFCLASS64
+  Data:    ELFDATA2LSB
+  Type:    ET_DYN
+  Machine: EM_AARCH64
+  Entry:   0x1000
+Sections:
+  ## Code that no debug info describes. The symbol makes it disassembled, and
+  ## the described functions that follow provide range ends that must not be
+  ## reached from here.
+  - Name:         .text.undescribed
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x500
+    Offset:       0x500
+    AddressAlign: 0x4
+    ## nop
+    ## nop
+    Content:      1F2003D51F2003D5
+  - Name:         .text.first
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x1000
+    Offset:       0x1000
+    AddressAlign: 0x4
+    ## nop
+    ## nop
+    Content:      1F2003D51F2003D5
+  - Name:         .text.adjacent
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x1008
+    Offset:       0x1008
+    AddressAlign: 0x4
+    ## nop
+    ## b 0x1000
+    Content:      1F2003D5FDFFFF17
+  - Name:         .text.gap
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x2000
+    Offset:       0x2000
+    AddressAlign: 0x4
+    ## nop
+    ## nop
+    Content:      1F2003D51F2003D5
+  - Name:         .text.transfer
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x3000
+    Offset:       0x3000
+    AddressAlign: 0x4
+    ## b 0x2000
+    Content:      00FCFF17
+  - Name:         .text.split
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x4000
+    Offset:       0x4000
+    AddressAlign: 0x4
+    ## nop
+    ## nop
+    ## nop
+    ## b 0x4000
+    Content:      1F2003D51F2003D51F2003D5FDFFFF17
+  - Name:         .text.split-gap-head
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x6000
+    Offset:       0x6000
+    AddressAlign: 0x4
+    ## nop
+    ## nop
+    Content:      1F2003D51F2003D5
+  - Name:         .text.split-gap-tail
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x7000
+    Offset:       0x7000
+    AddressAlign: 0x4
+    ## nop
+    ## b 0x6000
+    Content:      1F2003D5FFFBFF17
+  - Name:         .text.duplicate
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x8000
+    Offset:       0x8000
+    AddressAlign: 0x4
+    ## nop
+    ## nop
+    ## nop
+    ## b 0x8000
+    Content:      1F2003D51F2003D51F2003D5FDFFFF17
+  - Name:         .text.call
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x9000
+    Offset:       0x9000
+    AddressAlign: 0x4
+    ## nop
+    ## bl 0x9000
+    ## nop
+    ## b 0x9000
+    Content:      1F2003D5FFFFFF971F2003D5FDFFFF17
+  - Name:         .text.ret
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0xA000
+    Offset:       0xA000
+    AddressAlign: 0x4
+    ## nop
+    ## ret
+    ## nop
+    ## b 0xa000
+    Content:      1F2003D5C0035FD61F2003D5FDFFFF17
+  - Name:         .text.tail
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0xB000
+    Offset:       0xB000
+    AddressAlign: 0x4
+    ## nop
+    ## nop
+    Content:      1F2003D51F2003D5
+  ## Code that no debug info describes, adjacent to the described function
+  ## above, so that only the unknown function can end the inferred range.
+  - Name:         .text.tail-undescribed
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0xB008
+    Offset:       0xB008
+    AddressAlign: 0x4
+    ## nop
+    ## b 0xb000
+    Content:      1F2003D5FDFFFF17
+  - Name:         .text.func-change
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0xC000
+    Offset:       0xC000
+    AddressAlign: 0x4
+    ## nop
+    ## nop
+    ## nop
+    ## b 0xc000
+    Content:      1F2003D51F2003D51F2003D5FDFFFF17
+  - Name:         .text.func-change-tail
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0xD000
+    Offset:       0xD000
+    AddressAlign: 0x4
+    ## nop
+    ## nop
+    Content:      1F2003D51F2003D5
+  ## A trap in the middle of a function. It transfers control nowhere, so it is
+  ## no branch, call or return, but execution never continues past it either.
+  - Name:         .text.trap
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0xD800
+    Offset:       0xD800
+    AddressAlign: 0x4
+    ## nop
+    ## brk #0
+    ## nop
+    ## b 0xd800
+    Content:      1F2003D5000020D41F2003D5FDFFFF17
+  ## A branch to itself, which is both the source and the destination of the
+  ## record below.
+  - Name:         .text.self
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0xD900
+    Offset:       0xD900
+    AddressAlign: 0x4
+    ## b 0xd900
+    Content:      00000014
+  ## A loop whose body holds a branch of its own, so that the range inferred
+  ## from a record of the loop back edge covers a part of the body only.
+  - Name:         .text.inner
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0xDA00
+    Offset:       0xDA00
+    AddressAlign: 0x4
+    ## nop
+    ## b.eq 0xda00
+    ## nop
+    ## b 0xda00
+    Content:      1F2003D5E0FFFF541F2003D5FDFFFF17
+ProgramHeaders:
+  - Type:     PT_LOAD
+    Flags:    [ PF_X, PF_R ]
+    Offset:   0
+    VAddr:    0
+    FileSize: 0xDA10
+    MemSize:  0xDA10
+    Align:    0x10000
+Symbols:
+  - Name:    undescribed
+    Type:    STT_FUNC
+    Section: .text.undescribed
+    Binding: STB_GLOBAL
+    Value:   0x500
+    Size:    0x8
+  - Name:    first
+    Type:    STT_FUNC
+    Section: .text.first
+    Binding: STB_GLOBAL
+    Value:   0x1000
+    Size:    0x8
+  - Name:    second
+    Type:    STT_FUNC
+    Section: .text.adjacent
+    Binding: STB_GLOBAL
+    Value:   0x1008
+    Size:    0x8
+  - Name:    third
+    Type:    STT_FUNC
+    Section: .text.gap
+    Binding: STB_GLOBAL
+    Value:   0x2000
+    Size:    0x8
+  - Name:    fourth
+    Type:    STT_FUNC
+    Section: .text.transfer
+    Binding: STB_GLOBAL
+    Value:   0x3000
+    Size:    0x4
+  - Name:    split
+    Type:    STT_FUNC
+    Section: .text.split
+    Binding: STB_GLOBAL
+    Value:   0x4000
+    Size:    0x10
+  - Name:    split_gap
+    Type:    STT_FUNC
+    Section: .text.split-gap-head
+    Binding: STB_GLOBAL
+    Value:   0x6000
+    Size:    0x8
+  ## A name of its own, so that this range is not marked as a function entry
+  ## and only the address gap can end the inferred range.
+  - Name:    split_gap_tail
+    Type:    STT_FUNC
+    Section: .text.split-gap-tail
+    Binding: STB_GLOBAL
+    Value:   0x7000
+    Size:    0x8
+  - Name:    duplicate
+    Type:    STT_FUNC
+    Section: .text.duplicate
+    Binding: STB_GLOBAL
+    Value:   0x8000
+    Size:    0x8
+  ## Two functions cannot share a symbol name, so the second one carries the
+  ## internal suffix that link-time renaming appends. Debug info keeps the
+  ## source name for both, and the suffix is stripped before the symbol is
+  ## compared with it, so this range is still marked as a function entry.
+  - Name:    duplicate.llvm.1
+    Type:    STT_FUNC
+    Section: .text.duplicate
+    Binding: STB_GLOBAL
+    Value:   0x8008
+    Size:    0x8
+  - Name:    call_stop
+    Type:    STT_FUNC
+    Section: .text.call
+    Binding: STB_GLOBAL
+    Value:   0x9000
+    Size:    0x10
+  - Name:    ret_stop
+    Type:    STT_FUNC
+    Section: .text.ret
+    Binding: STB_GLOBAL
+    Value:   0xA000
+    Size:    0x10
+  - Name:    tail
+    Type:    STT_FUNC
+    Section: .text.tail
+    Binding: STB_GLOBAL
+    Value:   0xB000
+    Size:    0x8
+  - Name:    tail_undescribed
+    Type:    STT_FUNC
+    Section: .text.tail-undescribed
+    Binding: STB_GLOBAL
+    Value:   0xB008
+    Size:    0x8
+  - Name:    preceding
+    Type:    STT_FUNC
+    Section: .text.func-change
+    Binding: STB_GLOBAL
+    Value:   0xC000
+    Size:    0x8
+  ## The function that follows starts at 0xc008, where no symbol sits, so that
+  ## range is not marked as a function entry. Its symbol sits at its later
+  ## range, which keeps the function from having no entry at all.
+  - Name:    following
+    Type:    STT_FUNC
+    Section: .text.func-change-tail
+    Binding: STB_GLOBAL
+    Value:   0xD000
+    Size:    0x8
+  - Name:    trap_stop
+    Type:    STT_FUNC
+    Section: .text.trap
+    Binding: STB_GLOBAL
+    Value:   0xD800
+    Size:    0x10
+  - Name:    self_loop
+    Type:    STT_FUNC
+    Section: .text.self
+    Binding: STB_GLOBAL
+    Value:   0xD900
+    Size:    0x4
+  - Name:    inner_branch
+    Type:    STT_FUNC
+    Section: .text.inner
+    Binding: STB_GLOBAL
+    Value:   0xDA00
+    Size:    0x10
+DWARF:
+  debug_abbrev:
+    - ID: 0
+      Table:
+        - Code:     1
+          Tag:      DW_TAG_compile_unit
+          Children: DW_CHILDREN_yes
+          Attributes:
+            - Attribute: DW_AT_name
+              Form:      DW_FORM_string
+        - Code:     2
+          Tag:      DW_TAG_subprogram
+          Children: DW_CHILDREN_no
+          Attributes:
+            - Attribute: DW_AT_name
+              Form:      DW_FORM_string
+            - Attribute: DW_AT_low_pc
+              Form:      DW_FORM_addr
+            - Attribute: DW_AT_high_pc
+              Form:      DW_FORM_data8
+        ## A function whose code is described by a range list instead of a
+        ## single low_pc/high_pc pair.
+        - Code:     3
+          Tag:      DW_TAG_subprogram
+          Children: DW_CHILDREN_no
+          Attributes:
+            - Attribute: DW_AT_name
+              Form:      DW_FORM_string
+            - Attribute: DW_AT_ranges
+              Form:      DW_FORM_sec_offset
+  debug_ranges:
+    ## Two adjacent ranges of one function.
+    - Offset: 0x0
+      Entries:
+        - LowOffset:  0x4000
+          HighOffset: 0x4008
+        - LowOffset:  0x4008
+          HighOffset: 0x4010
+        ## End of the range list.
+        - LowOffset:  0x0
+          HighOffset: 0x0
+    ## Two ranges of one function separated by an address gap.
+    - Offset: 0x40
+      Entries:
+        - LowOffset:  0x6000
+          HighOffset: 0x6008
+        - LowOffset:  0x7000
+          HighOffset: 0x7008
+        - LowOffset:  0x0
+          HighOffset: 0x0
+    ## Two ranges of one function, the first of them adjacent to the function
+    ## that precedes it and started by no symbol.
+    - Offset: 0x80
+      Entries:
+        - LowOffset:  0xC008
+          HighOffset: 0xC010
+        - LowOffset:  0xD000
+          HighOffset: 0xD008
+        - LowOffset:  0x0
+          HighOffset: 0x0
+  debug_info:
+    - Version:       4
+      AbbrevTableID: 0
+      AddrSize:      8
+      Entries:
+        - AbbrCode: 1
+          Values:
+            - CStr: spe-range-gap.c
+        - AbbrCode: 2
+          Values:
+            - CStr:  first
+            - Value: 0x1000
+            - Value: 0x8
+        - AbbrCode: 2
+          Values:
+            - CStr:  second
+            - Value: 0x1008
+            - Value: 0x8
+        - AbbrCode: 2
+          Values:
+            - CStr:  third
+            - Value: 0x2000
+            - Value: 0x8
+        - AbbrCode: 2
+          Values:
+            - CStr:  fourth
+            - Value: 0x3000
+            - Value: 0x4
+        - AbbrCode: 3
+          Values:
+            - CStr:  split
+            - Value: 0x0
+        - AbbrCode: 3
+          Values:
+            - CStr:  split_gap
+            - Value: 0x40
+        ## Two functions of the same name, each described by its own entry.
+        - AbbrCode: 2
+          Values:
+            - CStr:  duplicate
+            - Value: 0x8000
+            - Value: 0x8
+        - AbbrCode: 2
+          Values:
+            - CStr:  duplicate
+            - Value: 0x8008
+            - Value: 0x8
+        - AbbrCode: 2
+          Values:
+            - CStr:  call_stop
+            - Value: 0x9000
+            - Value: 0x10
+        - AbbrCode: 2
+          Values:
+            - CStr:  ret_stop
+            - Value: 0xA000
+            - Value: 0x10
+        ## The code that follows this function has no entry of its own.
+        - AbbrCode: 2
+          Values:
+            - CStr:  tail
+            - Value: 0xB000
+            - Value: 0x8
+        - AbbrCode: 2
+          Values:
+            - CStr:  preceding
+            - Value: 0xC000
+            - Value: 0x8
+        - AbbrCode: 3
+          Values:
+            - CStr:  following
+            - Value: 0x80
+        - AbbrCode: 2
+          Values:
+            - CStr:  trap_stop
+            - Value: 0xD800
+            - Value: 0x10
+        - AbbrCode: 2
+          Values:
+            - CStr:  self_loop
+            - Value: 0xD900
+            - Value: 0x4
+        - AbbrCode: 2
+          Values:
+            - CStr:  inner_branch
+            - Value: 0xDA00
+            - Value: 0x10
+        - AbbrCode: 0
+
+#--- no-debug.yaml
+## The symbol table describes the function, but the binary carries no debug
+## info. Function ranges must not be recovered from the symbol table here.
+--- !ELF
+FileHeader:
+  Class:   ELFCLASS64
+  Data:    ELFDATA2LSB
+  Type:    ET_DYN
+  Machine: EM_AARCH64
+  Entry:   0x1000
+Sections:
+  - Name:         .text
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x1000
+    Offset:       0x1000
+    AddressAlign: 0x4
+    ## nop
+    ## nop
+    ## b 0x1000
+    Content:      1F2003D51F2003D5FEFFFF17
+ProgramHeaders:
+  - Type:     PT_LOAD
+    Flags:    [ PF_X, PF_R ]
+    Offset:   0
+    VAddr:    0
+    FileSize: 0x100C
+    MemSize:  0x100C
+    Align:    0x10000
+Symbols:
+  - Name:    foo
+    Type:    STT_FUNC
+    Section: .text
+    Binding: STB_GLOBAL
+    Value:   0x1000
+    Size:    0xC
+
+#--- perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0xe000) @ 0 00:00 0 0]: r-xp /tmp/spe-range-gap.exe
+70000000100c 0x70000000100c/0x700000001000/P/-/-/10/UNCOND/-
+
+700000003000 0x700000003000/0x700000002000/P/-/-/10/UNCOND/-
+70000000400c 0x70000000400c/0x700000004000/P/-/-/10/UNCOND/-
+700000007004 0x700000007004/0x700000006000/P/-/-/10/UNCOND/-
+70000000800c 0x70000000800c/0x700000008000/P/-/-/10/UNCOND/-
+70000000f000 0x70000000f000/0x700000000500/P/-/-/10/CALL/-
+70000000900c 0x70000000900c/0x700000009000/P/-/-/10/UNCOND/-
+70000000a00c 0x70000000a00c/0x70000000a000/P/-/-/10/UNCOND/-
+70000000b00c 0x70000000b00c/0x70000000b000/P/-/-/10/UNCOND/-
+70000000c00c 0x70000000c00c/0x70000000c000/P/-/-/10/UNCOND/-
+70000000d80c 0x70000000d80c/0x70000000d800/P/-/-/10/UNCOND/-
+70000000d900 0x70000000d900/0x70000000d900/P/-/-/10/UNCOND/-
+70000000da0c 0x70000000da0c/0x70000000da00/P/-/-/10/UNCOND/-
+
+#--- no-debug.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/no-debug.exe
+700000001008 0x700000001008/0x700000001000/P/-/-/10/UNCOND/-
+700000001008 0x700000001008/0x700000001000/P/-/-/10/UNCOND/-
+
+#--- unordered.yaml
+## Two described functions whose sections are listed in decreasing address
+## order, which leaves the collected instruction addresses unordered. The file
+## offsets still increase, as an object file requires.
+--- !ELF
+FileHeader:
+  Class:   ELFCLASS64
+  Data:    ELFDATA2LSB
+  Type:    ET_DYN
+  Machine: EM_AARCH64
+  Entry:   0x1000
+Sections:
+  - Name:         .text.high
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x2000
+    Offset:       0x1000
+    AddressAlign: 0x4
+    ## nop
+    ## b 0x2000
+    Content:      1F2003D5FFFFFF17
+  - Name:         .text.low
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x1000
+    Offset:       0x1008
+    AddressAlign: 0x4
+    ## nop
+    ## b 0x1000
+    Content:      1F2003D5FFFFFF17
+ProgramHeaders:
+  - Type:     PT_LOAD
+    Flags:    [ PF_X, PF_R ]
+    Offset:   0
+    VAddr:    0
+    FileSize: 0x1010
+    MemSize:  0x1010
+    Align:    0x10000
+Symbols:
+  - Name:    high
+    Type:    STT_FUNC
+    Section: .text.high
+    Binding: STB_GLOBAL
+    Value:   0x2000
+    Size:    0x8
+  - Name:    low
+    Type:    STT_FUNC
+    Section: .text.low
+    Binding: STB_GLOBAL
+    Value:   0x1000
+    Size:    0x8
+DWARF:
+  debug_abbrev:
+    - ID: 0
+      Table:
+        - Code:     1
+          Tag:      DW_TAG_compile_unit
+          Children: DW_CHILDREN_yes
+          Attributes:
+            - Attribute: DW_AT_name
+              Form:      DW_FORM_string
+        - Code:     2
+          Tag:      DW_TAG_subprogram
+          Children: DW_CHILDREN_no
+          Attributes:
+            - Attribute: DW_AT_name
+              Form:      DW_FORM_string
+            - Attribute: DW_AT_low_pc
+              Form:      DW_FORM_addr
+            - Attribute: DW_AT_high_pc
+              Form:      DW_FORM_data8
+  debug_info:
+    - Version:       4
+      AbbrevTableID: 0
+      AddrSize:      8
+      Entries:
+        - AbbrCode: 1
+          Values:
+            - CStr: unordered.c
+        - AbbrCode: 2
+          Values:
+            - CStr:  low
+            - Value: 0x1000
+            - Value: 0x8
+        - AbbrCode: 2
+          Values:
+            - CStr:  high
+            - Value: 0x2000
+            - Value: 0x8
+        - AbbrCode: 0
+
+#--- unordered.perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x3000) @ 0 00:00 0 0]: r-xp /tmp/unordered.exe
+700000001004 0x700000001004/0x700000001000/P/-/-/10/UNCOND/-
+700000002004 0x700000002004/0x700000002000/P/-/-/10/UNCOND/-
diff --git a/llvm/test/tools/llvm-profgen/AArch64/spe-symbolized-profile.test b/llvm/test/tools/llvm-profgen/AArch64/spe-symbolized-profile.test
new file mode 100644
index 0000000000000..a2aef4c3e96cf
--- /dev/null
+++ b/llvm/test/tools/llvm-profgen/AArch64/spe-symbolized-profile.test
@@ -0,0 +1,109 @@
+## The other Arm SPE tests skip symbolization and check the inferred ranges
+## themselves. This one carries a line table, so that the range and branch
+## counters of a single-branch record are checked where they end up: as the
+## line samples and the head count of the enclosing function.
+
+# RUN: split-file %s %t
+# RUN: yaml2obj %t/binary.yaml -o %t/spe-symbolized-profile.exe
+# RUN: llvm-profgen --binary=%t/spe-symbolized-profile.exe \
+# RUN:   --perfscript=%t/perfscript --spe-branch-profile --format=text \
+# RUN:   --output=%t/profile 2>&1 \
+# RUN:   | FileCheck --allow-empty %s --check-prefix=NO-WARN
+# RUN: FileCheck %s --input-file=%t/profile
+
+## The record is a return into a function that carries debug info, so neither
+## its source nor its destination is reported as unexpected, and no record is
+## dropped.
+# NO-WARN-NOT: warning: {{.*}}do not match the binary
+# NO-WARN-NOT: warning: {{.*}}not an instruction of this binary
+# NO-WARN-NOT: warning: No samples in perf script
+
+## The range inferred for the destination covers both instructions of
+## with_debug, which the line table attributes to the two lines that follow its
+## declaration, and the branch to its entry becomes the head count. The total
+## weighs each line sample by the size of the instructions it covers, as it
+## does for a branch-stack profile.
+# CHECK:      with_debug:8:1{{$}}
+# CHECK-NEXT:  1: 1
+# CHECK-NEXT:  2: 1
+
+#--- binary.yaml
+## with_debug is described by both the line table and a subprogram; the
+## instructions of without_debug are described by neither, which leaves the
+## source of the record unsymbolized without affecting the counters.
+--- !ELF
+FileHeader:
+  Class:   ELFCLASS64
+  Data:    ELFDATA2LSB
+  Type:    ET_DYN
+  Machine: EM_AARCH64
+ProgramHeaders:
+  - Type:     PT_LOAD
+    Flags:    [ PF_X, PF_R ]
+    VAddr:    0
+    Offset:   0
+    Align:    0x1000
+    FirstSec: .text
+    LastSec:  .text
+Sections:
+  - Name:         .text
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x27C
+    Offset:       0x27C
+    AddressAlign: 0x4
+    ## with_debug: add w0, w0, #1; ret
+    ## without_debug: add w0, w0, #2; ret
+    Content:      00040011C0035FD600080011C0035FD6
+  - Name:         .debug_abbrev
+    Type:         SHT_PROGBITS
+    AddressAlign: 0x1
+    Content:      011101252513050325721710171B25111B120673170000022E00111B1206401803253A0B3B0B49133F19000003240003253E0B0B0B000000
+  - Name:         .debug_info
+    Type:         SHT_PROGBITS
+    AddressAlign: 0x1
+    Content:      33000000050001080000000001001D0001080000000000000002000800000008000000020008000000016F030001320000000304050400
+  - Name:         .debug_str_offsets
+    Type:         SHT_PROGBITS
+    AddressAlign: 0x1
+    Content:      18000000050000000200000007000000000000001D00000003000000
+  - Name:         .debug_line
+    Type:         SHT_PROGBITS
+    AddressAlign: 0x1
+    Content:      460000000500080025000000010101FB0E0D00010101010000000100000101011F010000000002011F020F0102000000000400050C0A0009027C020000000000001305034B0204000101
+  - Name:         .debug_line_str
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_MERGE, SHF_STRINGS ]
+    AddressAlign: 0x1
+    EntSize:      0x1
+    Content:      2F007761726E2D6E6F742D73796D626F6C697A65642E6300
+Symbols:
+  - Name:    with_debug
+    Type:    STT_FUNC
+    Section: .text
+    Binding: STB_GLOBAL
+    Value:   0x27C
+    Size:    0x8
+  - Name:    without_debug
+    Type:    STT_FUNC
+    Section: .text
+    Binding: STB_GLOBAL
+    Value:   0x284
+    Size:    0x8
+DWARF:
+  debug_str:
+    - '/'
+    - ''
+    - int
+    - warn-not-symbolized.c
+    - with_debug
+  debug_addr:
+    - Length:      0xC
+      Version:     0x5
+      AddressSize: 0x8
+      Entries:
+        - Address: 0x27C
+
+#--- perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x1000) @ 0 00:00 0 0]: r-xp /tmp/spe-symbolized-profile.exe
+700000000288 0x700000000288/0x70000000027c/P/-/-/10/RET/-
diff --git a/llvm/test/tools/llvm-profgen/spe-target-validation.test b/llvm/test/tools/llvm-profgen/spe-target-validation.test
new file mode 100644
index 0000000000000..fbcd792cbbfaf
--- /dev/null
+++ b/llvm/test/tools/llvm-profgen/spe-target-validation.test
@@ -0,0 +1,45 @@
+# REQUIRES: arm-registered-target
+#
+# RUN: split-file %s %t
+# RUN: yaml2obj %t/binary.yaml -o %t/arm.exe
+# RUN: not llvm-profgen --binary=%t/arm.exe --perfscript=%t/perfscript \
+# RUN:   --spe-branch-profile --skip-symbolization --output=/dev/null 2>&1 \
+# RUN:   | FileCheck %s
+#
+# CHECK: error: --spe-branch-profile requires an AArch64 binary.
+
+#--- binary.yaml
+--- !ELF
+FileHeader:
+  Class:   ELFCLASS32
+  Data:    ELFDATA2LSB
+  Type:    ET_DYN
+  Machine: EM_ARM
+Sections:
+  - Name:         .text
+    Type:         SHT_PROGBITS
+    Flags:        [ SHF_ALLOC, SHF_EXECINSTR ]
+    Address:      0x1000
+    Offset:       0x1000
+    AddressAlign: 0x4
+    ## bx lr
+    Content:      1EFF2FE1
+ProgramHeaders:
+  - Type:     PT_LOAD
+    Flags:    [ PF_X, PF_R ]
+    Offset:   0
+    VAddr:    0
+    FileSize: 0x1004
+    MemSize:  0x1004
+    Align:    0x1000
+Symbols:
+  - Name:    foo
+    Type:    STT_FUNC
+    Section: .text
+    Binding: STB_GLOBAL
+    Value:   0x1000
+    Size:    0x4
+
+#--- perfscript
+PERF_RECORD_MMAP2 1/1: [0x700000000000(0x2000) @ 0 00:00 0 0]: r-xp /tmp/arm.exe
+700000001000 0x700000001000/0x700000001000/P/-/-/10/RET/-
diff --git a/llvm/tools/llvm-profgen/Options.h b/llvm/tools/llvm-profgen/Options.h
index 395a55726274d..4886962790bd9 100644
--- a/llvm/tools/llvm-profgen/Options.h
+++ b/llvm/tools/llvm-profgen/Options.h
@@ -24,6 +24,7 @@ extern cl::opt<bool> EnableCSPreInliner;
 extern cl::opt<bool> UseContextCostForPreInliner;
 extern cl::opt<bool> LoadFunctionFromSymbol;
 extern cl::opt<bool> TimeProfGen;
+extern cl::opt<bool> ReadSPEBranchProfile;
 
 } // end namespace llvm
 
diff --git a/llvm/tools/llvm-profgen/PerfReader.cpp b/llvm/tools/llvm-profgen/PerfReader.cpp
index 18297d5f7665b..2b79714ec7cf3 100644
--- a/llvm/tools/llvm-profgen/PerfReader.cpp
+++ b/llvm/tools/llvm-profgen/PerfReader.cpp
@@ -70,6 +70,12 @@ cl::opt<bool> TimeProfGen("time-profgen", cl::desc("Time llvm-profgen phases"),
 static const char *TimerGroupName = "profgen";
 static const char *TimerGroupDesc = "llvm-profgen";
 
+cl::opt<bool> ReadSPEBranchProfile(
+    "spe-branch-profile",
+    cl::desc("Read the input as an Arm SPE branch profile of an AArch64 "
+             "binary; requires FEAT_SPEv1p2 and event filtering."),
+    cl::cat(ProfGenCategory));
+
 namespace sampleprof {
 
 void VirtualUnwinder::unwindCall(UnwindState &State) {
@@ -364,8 +370,15 @@ PerfReaderBase::create(ProfiledBinary *Binary, InputFile &Input,
     return PerfReader;
   }
 
+  // An Arm SPE trace cannot be detected: its lines look like the branch-stack
+  // lines of an LBR trace, so detection would read it as one. The kind is also
+  // needed before a recording is printed, because the perf command differs.
+  if (ReadSPEBranchProfile)
+    Input.Content = PerfContent::ArmSPE;
+
   // For perf data input, we need to convert them into perf script first.
   // If this is a kernel perf file, there is no need for retrieving PIDs.
+  bool ConvertedFromPerfData = Input.Format == InputFormat::PerfData;
   if (Input.Format == InputFormat::PerfData)
     Input = PerfScriptReader::convertPerfDataToTrace(Binary, Binary->isKernel(),
                                                      Input, PIDFilter);
@@ -373,12 +386,16 @@ PerfReaderBase::create(ProfiledBinary *Binary, InputFile &Input,
   assert((Input.Format == InputFormat::PerfScript) &&
          "Should be a perfscript!");
 
-  Input.Content = PerfScriptReader::checkPerfScriptType(Input.InputFilePath);
+  if (Input.Content == PerfContent::UnknownContent)
+    Input.Content = PerfScriptReader::checkPerfScriptType(Input.InputFilePath);
   if (Input.Content == PerfContent::LBRStack) {
     PerfReader.reset(
         new HybridPerfReader(Binary, Input.InputFilePath, PIDFilter));
   } else if (Input.Content == PerfContent::LBR) {
     PerfReader.reset(new LBRPerfReader(Binary, Input.InputFilePath, PIDFilter));
+  } else if (Input.Content == PerfContent::ArmSPE) {
+    PerfReader.reset(new ArmSPEReader(Binary, Input.InputFilePath, PIDFilter,
+                                      ConvertedFromPerfData));
   } else {
     exitWithError("Unsupported perfscript!");
   }
@@ -553,6 +570,10 @@ PerfScriptReader::convertPerfDataToTrace(ProfiledBinary *Binary, bool SkipPID,
   SmallVector<StringRef, 8> ScriptSampleArgs;
   ScriptSampleArgs.push_back(PerfExecutablePath);
   ScriptSampleArgs.push_back("script");
+  // Decode Arm SPE branch events into a branch stack of one entry: a larger
+  // limit adds a synthesized entry with source 0x0, which the reader rejects.
+  if (File.Content == PerfContent::ArmSPE)
+    ScriptSampleArgs.push_back("--itrace=bl1");
   ScriptSampleArgs.push_back("--show-mmap-events");
   ScriptSampleArgs.push_back("-F");
   ScriptSampleArgs.push_back("ip,brstack");
@@ -564,8 +585,7 @@ PerfScriptReader::convertPerfDataToTrace(ProfiledBinary *Binary, bool SkipPID,
   }
   RunPerfScript(ScriptSampleArgs);
 
-  return {std::string(PerfTraceFile), InputFormat::PerfScript,
-          PerfContent::UnknownContent};
+  return {std::string(PerfTraceFile), InputFormat::PerfScript, File.Content};
 }
 
 static StringRef filename(StringRef Path, bool UseBackSlash) {
@@ -731,8 +751,45 @@ static bool parseAddress(StringRef Str, uint64_t &Addr, bool HasPrefix) {
   return Str.getAsInteger(16, Addr);
 }
 
+/// Clear the bit of the branch-stack entry at \p Index in \p TakenMask. Only a
+/// record format that holds a branch that was not taken reaches this, and the
+/// deepest such record holds one branch, so the index stays inside the mask.
+static void clearTaken(uint64_t &TakenMask, size_t Index) {
+  assert(Index < TakenMaskWidth && "Branch-stack entry past the mask");
+  TakenMask &= ~(1ULL << Index);
+}
+
+bool PerfScriptReader::parseBranchEntry(StringRef Token, uint64_t &Source,
+                                        uint64_t &Target, bool &Taken) {
+  // Only the source, the destination and the prediction field are read, so stop
+  // at the third slash and leave the rest of the entry as one field. Every
+  // field is kept, so a field count still counts the slashes before it.
+  SmallVector<StringRef, 4> Fields;
+  Token.split(Fields, "/", /*MaxSplit=*/3);
+  if (Fields.size() < 2 || parseAddress(Fields[0], Source, true) ||
+      parseAddress(Fields[1], Target, true))
+    return false;
+
+  // The third field holds the prediction flags, whose alphabet depends on the
+  // record format, so whether they mark a branch that was taken is left to the
+  // reader. A field the reader cannot read makes the entry unreadable. perf
+  // prints fields of its own after this one, so a prediction field that ends
+  // the entry was cut short and is withheld: its P may be all of a PN.
+  std::optional<bool> Flag =
+      parseTakenFlag(Fields.size() > 3 ? Fields[2] : StringRef());
+  if (!Flag)
+    return false;
+  Taken = *Flag;
+
+  // Canonicalize to use preferred load address as base address.
+  Source = Binary->canonicalizeVirtualAddress(Source);
+  Target = Binary->canonicalizeVirtualAddress(Target);
+  return true;
+}
+
 bool PerfScriptReader::extractLBRStack(TraceStream &TraceIt,
-                                       SmallVectorImpl<LBREntry> &LBRStack) {
+                                       SmallVectorImpl<LBREntry> &LBRStack,
+                                       uint64_t &TakenMask) {
   // The raw format of LBR stack is like:
   // 0x4005c8/0x4005dc/P/-/-/0 0x40062f/0x4005b0/P/-/-/0 ...
   //                           ... 0x4005c8/0x4005dc/P/-/-/0
@@ -765,21 +822,16 @@ bool PerfScriptReader::extractLBRStack(TraceStream &TraceIt,
     if (Token.size() == 0)
       continue;
 
-    SmallVector<StringRef, 8> Addresses;
-    Token.split(Addresses, "/");
     uint64_t Src;
     uint64_t Dst;
+    bool Taken = true;
 
     // Stop at broken LBR records.
-    if (Addresses.size() < 2 || parseAddress(Addresses[0], Src, true) ||
-        parseAddress(Addresses[1], Dst, true)) {
+    if (!parseBranchEntry(Token, Src, Dst, Taken)) {
       WarnInvalidLBR(TraceIt);
       break;
     }
 
-    // Canonicalize to use preferred load address as base address.
-    Src = Binary->canonicalizeVirtualAddress(Src);
-    Dst = Binary->canonicalizeVirtualAddress(Dst);
     bool SrcIsInternal = Binary->addressIsCode(Src);
     bool DstIsInternal = Binary->addressIsCode(Dst);
     if (!SrcIsInternal)
@@ -790,12 +842,245 @@ bool PerfScriptReader::extractLBRStack(TraceStream &TraceIt,
     if (!SrcIsInternal && !DstIsInternal)
       continue;
 
+    if (!Taken)
+      clearTaken(TakenMask, LBRStack.size());
     LBRStack.emplace_back(LBREntry(Src, Dst));
   }
   TraceIt.advance();
   return !LBRStack.empty();
 }
 
+uint64_t ArmSPEReader::parseAggregatedCount(TraceStream &) {
+  // perf prints no repeat count for Arm SPE, so leave the line to
+  // extractLBRStack, which reports it when unreadable, instead of silently
+  // weighing the next record by an address read off it as a count.
+  return 1;
+}
+
+std::optional<bool> ArmSPEReader::parseTakenFlag(StringRef PredictionFlags) {
+  // perf appends N to the prediction flags of a branch it recorded as not
+  // taken, and prints `-`, `P`, `M`, `PN` or `MN` for the field. Anything else
+  // leaves the entry unreadable rather than silently counted as taken.
+  if (PredictionFlags == "-" || PredictionFlags == "P" ||
+      PredictionFlags == "M")
+    return true;
+  if (PredictionFlags == "PN" || PredictionFlags == "MN")
+    return false;
+  return std::nullopt;
+}
+
+bool ArmSPEReader::extractLBRStack(TraceStream &TraceIt,
+                                   SmallVectorImpl<LBREntry> &LBRStack,
+                                   uint64_t &TakenMask) {
+  // perf prints an Arm SPE branch record as a branch stack of one entry,
+  // preceded by the instruction pointer of the sample, which repeats the
+  // source of the branch:
+  //   ffff800080387878 0xffff800080387878/0xffff80008039c524/P/-/-/19/RET/-
+  SmallVector<StringRef, 4> Fields;
+  TraceIt.getCurrentLine().rtrim().split(Fields, " ", -1, false);
+
+  // A blank line carries no record, so skip it without a diagnostic.
+  if (Fields.empty()) {
+    TraceIt.advance();
+    return false;
+  }
+
+  // Report a line that holds no readable record. A trace printed with another
+  // field list holds millions of them, so name a line only on request;
+  // reportTraceLosses reports the share once.
+  auto ReportInvalidRecord = [&]() {
+    NumInvalidRecords++;
+    if (ShowDetailedWarning)
+      WithColor::warning() << "Invalid Arm SPE branch record at line "
+                           << TraceIt.getLineNumber() << ": "
+                           << TraceIt.getCurrentLine() << "\n";
+    TraceIt.advance();
+  };
+
+  // The leading instruction pointer is a bare address: no slash, and none of
+  // the 0x prefix that every address of an entry carries, so a prefixed field
+  // is something else and leaves the record unreadable.
+  size_t Index = 0;
+  uint64_t LeadingAddr = 0;
+  bool HasLeadingAddr = !Fields[0].contains('/');
+  if (HasLeadingAddr) {
+    if (parseAddress(Fields[0], LeadingAddr, /*HasPrefix=*/false)) {
+      ReportInvalidRecord();
+      return false;
+    }
+    Index = 1;
+  }
+
+  // A record holds the entry of the sampled branch right after the instruction
+  // pointer.
+  uint64_t Src = 0;
+  uint64_t Dst = 0;
+  bool Taken = true;
+  if (Fields.size() - Index < 1 ||
+      !parseBranchEntry(Fields[Index], Src, Dst, Taken)) {
+    ReportInvalidRecord();
+    return false;
+  }
+
+  // --itrace=bl1 leaves a record one entry. A second entry that parses is a
+  // deeper stack, whose synthesized source address 0x0 would fabricate a branch
+  // from outside the binary, so stop instead of weighing such a trace. Parse
+  // the field to decide, because a trailing field a wider -F list adds, a
+  // shared-object path above all, holds a slash too and leaves the record
+  // unreadable rather than the stack deeper.
+  if (Fields.size() - Index > 1) {
+    uint64_t ExtraSrc = 0;
+    uint64_t ExtraDst = 0;
+    bool ExtraTaken = true;
+    if (parseBranchEntry(Fields[Index + 1], ExtraSrc, ExtraDst, ExtraTaken))
+      exitWithError(
+          Twine("--spe-branch-profile requires Arm SPE branch records "
+                "created with `perf script --show-mmap-events --itrace=bl1 -F "
+                "ip,brstack`, but the record at line ") +
+          Twine(TraceIt.getLineNumber()) +
+          " holds a branch stack of more than one entry");
+    ReportInvalidRecord();
+    return false;
+  }
+
+  // A source address of zero names no instruction: perf prints it where the
+  // address of the sampled branch was lost, and left in place it reads as a
+  // branch from outside the binary. Compare against the canonical form of zero,
+  // because parseBranchEntry has already shifted every address of the trace
+  // onto the preferred load address of the binary.
+  if (Src == Binary->canonicalizeVirtualAddress(0)) {
+    ReportInvalidRecord();
+    return false;
+  }
+
+  // The leading instruction pointer repeats the source of the branch. Where the
+  // two differ, another field was printed first and neither can be trusted.
+  if (HasLeadingAddr &&
+      Binary->canonicalizeVirtualAddress(LeadingAddr) != Src) {
+    ReportInvalidRecord();
+    return false;
+  }
+
+  NumRecords++;
+
+  // The inherited parseSample reports a missing mmap event only for a record
+  // that yields a sample, and without the event every record is usually
+  // external and dropped below. The repeated call is a no-op once it has
+  // warned.
+  warnIfMissingMMap();
+
+  TraceIt.advance();
+  // A record holds one branch and no preceding one to bound a range with, so a
+  // destination outside the binary yields no counter at all. Count the record
+  // so that reportTraceLosses reports the share.
+  if (!Binary->addressIsCode(Dst)) {
+    NumDroppedRecords++;
+    return false;
+  }
+  // A branch from outside the binary keeps an external source, which the
+  // profile turns into a head sample of the function holding the destination.
+  if (!Binary->addressIsCode(Src))
+    Src = ExternalAddr;
+
+  // A branch that was not taken reports the following instruction as its
+  // destination, so keep the entry for the range that really was executed
+  // there; computeCounterFromLBR leaves out its branch counter from the flag.
+  if (!Taken)
+    clearTaken(TakenMask, LBRStack.size());
+  LBRStack.emplace_back(Src, Dst);
+  return true;
+}
+
+uint64_t ArmSPEReader::inferFirstRangeEnd(const PerfSample *Sample,
+                                          uint64_t Repeat) {
+  assert(!Sample->LBRStack.empty() &&
+         "an aggregated Arm SPE sample holds one branch");
+  // Execution continues at the destination of the recorded branch until the
+  // next transfer, so that transfer ends the range.
+  bool Uncovered = false;
+  uint64_t End =
+      Binary->findRangeEnd(Sample->LBRStack.front().Target, &Uncovered);
+  // Where the binary describes no function around the destination, the range
+  // is reduced to the sampled instruction. Count that here, in the one pass
+  // that looks the range of every sample up, so that reportTraceLosses needs
+  // no lookup of its own.
+  if (Uncovered)
+    NumUncoveredSamples += Repeat;
+  return End;
+}
+
+void ArmSPEReader::validateParsedTrace() {
+  // An input without a single Arm SPE branch record is the wrong input file, so
+  // reject it instead of writing an empty profile. Deciding after the whole
+  // trace was read keeps a damaged first line from discarding what follows.
+  if (NumRecords)
+    return;
+
+  // A recording is turned into a trace by this tool, with a perf command the
+  // user never wrote, so name what to record instead of what to print.
+  if (ConvertedFromPerfData)
+    exitWithError("--spe-branch-profile found no Arm SPE branch record in "
+                  "the perf data; record one with "
+                  "arm_spe/branch_filter=1,event_filter=2/ using perf 6.15 or "
+                  "later");
+
+  // For a trace the user printed, the field list of the perf command is what
+  // has to be corrected, so name it together with what the lines looked like.
+  StringRef Cause = NumInvalidRecords
+                        ? "no line of the input holds a usable one"
+                        : "the trace holds no record at all";
+  exitWithError(
+      Twine("--spe-branch-profile requires Arm SPE branch records created "
+            "with `perf script --show-mmap-events --itrace=bl1 -F "
+            "ip,brstack`, but ") +
+      Cause);
+}
+
+void ArmSPEReader::generateUnsymbolizedProfile() {
+  PerfScriptReader::generateUnsymbolizedProfile();
+  // Computing the counters looks the range of every sample up, which is what
+  // tells how many of them had to be reduced, so the trace is reported on once
+  // that has run rather than from warnInvalidRange.
+  reportTraceLosses();
+}
+
+void ArmSPEReader::reportTraceLosses() {
+  // Every share below is of the same base, the lines of the trace that were
+  // read as records, so that the reports can be read against one another.
+  // Blank lines are left out because they carry no record to lose.
+  uint64_t NumLines = NumInvalidRecords + NumRecords;
+
+  // Report the unreadable lines as a share, the way the other readers report
+  // their losses. Naming each line is left to --show-detailed-warning, because
+  // a wrong field list makes every line of the trace unreadable.
+  emitWarningSummary(NumInvalidRecords, NumLines,
+                     "of the sample lines of the trace are not usable Arm SPE "
+                     "branch records.");
+
+  // A record whose destination is not an instruction of this binary carries no
+  // counter, and a trace of another process or build consists of such records.
+  // Report their share, so a thinned profile is not taken for a complete one.
+  emitWarningSummary(NumDroppedRecords, NumLines,
+                     "of the sample lines of the trace are Arm SPE branch "
+                     "records whose destination is not an instruction of this "
+                     "binary.");
+
+  // The range of a record whose destination no function range covers is reduced
+  // to that one instruction. Report the share reduced this way; the count is of
+  // lines again because inferFirstRangeEnd weighs each sample by its repeats.
+  emitWarningSummary(NumUncoveredSamples, NumLines,
+                     "of the sample lines of the trace are Arm SPE branch "
+                     "records whose destination cannot be associated with a "
+                     "function, so only that instruction is profiled.");
+
+  // A trace whose every record was dropped yields an empty profile. The base
+  // class decides that from ranges formed out of consecutive branch-stack
+  // entries, which a single-branch record never forms, so name the records.
+  if (NumRecords == NumDroppedRecords)
+    WithColor::warning() << "No Arm SPE branch record yields a counter for "
+                            "this binary!\n";
+}
+
 bool PerfScriptReader::extractCallstack(TraceStream &TraceIt,
                                         SmallVectorImpl<uint64_t> &CallStack) {
   // The raw format of call stack is like:
@@ -907,7 +1192,7 @@ void HybridPerfReader::parseSample(TraceStream &TraceIt, uint64_t Count) {
 
   if (!TraceIt.isAtEoF() && isLBRSample(TraceIt.getCurrentLine(), true)) {
     // Parsing LBR stack and populate into PerfSample.LBRStack
-    if (extractLBRStack(TraceIt, Sample->LBRStack)) {
+    if (extractLBRStack(TraceIt, Sample->LBRStack, Sample->TakenMask)) {
       if (IgnoreStackSamples) {
         Sample->CallStack.clear();
       } else {
@@ -1072,20 +1357,26 @@ void UnsymbolizedProfileReader::parsePerfTraces() {
 void PerfScriptReader::computeCounterFromLBR(const PerfSample *Sample,
                                              uint64_t Repeat) {
   SampleCounter &Counter = SampleCounters.begin()->second;
-  uint64_t EndAddress = 0;
-  for (const LBREntry &LBR : Sample->LBRStack) {
+  uint64_t EndAddress = inferFirstRangeEnd(Sample, Repeat);
+  for (size_t I = 0; I < Sample->LBRStack.size(); I++) {
+    const LBREntry &LBR = Sample->LBRStack[I];
     uint64_t SourceAddress = LBR.Source;
     uint64_t TargetAddress = LBR.Target;
 
     // Record the branch if its SourceAddress is external. It can be the case an
     // external source call an internal function, later this branch will be used
     // to generate the function's head sample.
-    if (Binary->addressIsCode(TargetAddress)) {
+    // A branch that was not taken transferred no control, and counting it would
+    // make a fall-through that reaches the entry of the next function a call of
+    // that function and a head sample of it. A hardware branch stack holds
+    // taken branches alone, so every entry of one is flagged as taken.
+    if (Binary->addressIsCode(TargetAddress) && Sample->taken(I))
       Counter.recordBranchCount(SourceAddress, TargetAddress, Repeat);
-    }
 
     // If this not the first LBR, update the range count between TO of current
-    // LBR and FROM of next LBR.
+    // LBR and FROM of next LBR. The range that follows the first LBR is bound
+    // by no other entry, so it is only recorded for an input whose reader can
+    // infer where it ends.
     uint64_t StartAddress = TargetAddress;
     if (Binary->addressIsCode(StartAddress) &&
         Binary->addressIsCode(EndAddress) &&
@@ -1098,7 +1389,7 @@ void PerfScriptReader::computeCounterFromLBR(const PerfSample *Sample,
 void LBRPerfReader::parseSample(TraceStream &TraceIt, uint64_t Count) {
   std::shared_ptr<PerfSample> Sample = std::make_shared<PerfSample>();
   // Parsing LBR stack and populate into PerfSample.LBRStack
-  if (extractLBRStack(TraceIt, Sample->LBRStack)) {
+  if (extractLBRStack(TraceIt, Sample->LBRStack, Sample->TakenMask)) {
     warnIfMissingMMap();
     // Record LBR only samples by aggregation
     AggregatedSamples[Hashable<PerfSample>(Sample)] += Count;
@@ -1427,41 +1718,56 @@ void PerfScriptReader::warnInvalidRange() {
 }
 
 void PerfScriptReader::warnIfBranchTargetMismatch() {
-  // Collect unique branch source and target addresses from LBR samples,
-  // then check what percentage don't match known instructions in the binary.
+  // Check what share of branch samples with both endpoints in the binary
+  // disagrees with its instructions, weighing each aggregated sample by how
+  // often it was recorded.
 
   uint64_t MismatchedBranches = 0;
+  uint64_t NonConditionalNotTaken = 0;
   uint64_t MismatchedIndirectTargets = 0;
   uint64_t MismatchedTargets = 0;
   uint64_t TotalSamples = 0;
 
   for (const auto &Item : AggregatedSamples) {
     const PerfSample *Sample = Item.first.getPtr();
-    for (const LBREntry &LBR : Sample->LBRStack) {
-      uint64_t Source = LBR.Source;
-      uint64_t Target = LBR.Target;
+    uint64_t Repeat = Item.second;
+    for (size_t I = 0; I < Sample->LBRStack.size(); I++) {
+      uint64_t Source = Sample->LBRStack[I].Source;
+      uint64_t Target = Sample->LBRStack[I].Target;
       if (Source == ExternalAddr || Target == ExternalAddr)
         continue;
-      TotalSamples++;
+      TotalSamples += Repeat;
 
       // Validate Branch sources are Call/Branch/Indirect Branch
       if (!Binary->addressIsTransfer(Source))
-        MismatchedBranches++;
-
-      // Validate Indirect Branch targets landed in code. This may over estimate
-      // the vaid targets only because there's no good way to determine jump
-      // table targets
-      if (Binary->addressIsIndirectBranch(Source)) {
+        MismatchedBranches += Repeat;
+
+      // A not-taken branch must be conditional and target its following
+      // instruction, even when the recorded target is otherwise known.
+      if (!Sample->taken(I)) {
+        if (!Binary->addressIsConditionalBranch(Source))
+          NonConditionalNotTaken += Repeat;
+        else if (Target <= Source ||
+                 Target - Source != Binary->getInstSize(Source))
+          MismatchedTargets += Repeat;
+      } else if (Binary->addressIsIndirectBranch(Source)) {
+        // Validate indirect targets as code because jump-table targets cannot
+        // be identified more precisely.
         if (!Binary->addressIsCode(Target))
-          MismatchedIndirectTargets++;
+          MismatchedIndirectTargets += Repeat;
       } else if (!Binary->addressIsBranchTarget(Target) &&
-                 !Binary->findFuncRangeForStartAddr(Target))
-        MismatchedTargets++;
+                 !Binary->findFuncRangeForStartAddr(Target)) {
+        MismatchedTargets += Repeat;
+      }
     }
   }
 
   emitWarningSummary(MismatchedBranches, TotalSamples,
                      "of branch samples do not match the binary.");
+  emitWarningSummary(
+      NonConditionalNotTaken, TotalSamples,
+      "of branch samples are marked not taken but do not name a conditional "
+      "branch.");
   emitWarningSummary(MismatchedTargets, TotalSamples,
                      "of branch targets do not match the binary.");
   emitWarningSummary(MismatchedIndirectTargets, TotalSamples,
@@ -1471,6 +1777,7 @@ void PerfScriptReader::warnIfBranchTargetMismatch() {
 void PerfScriptReader::parsePerfTraces() {
   // Parse perf traces and do aggregation.
   parseAndAggregateTrace();
+  validateParsedTrace();
   if (Binary->isKernel() && !Binary->getIsLoadedByMMap()) {
     exitWithError(
         "Kernel is requested, but no kernel is found in mmap events.");
diff --git a/llvm/tools/llvm-profgen/PerfReader.h b/llvm/tools/llvm-profgen/PerfReader.h
index b9af0f19cb5d3..7aa707924871b 100644
--- a/llvm/tools/llvm-profgen/PerfReader.h
+++ b/llvm/tools/llvm-profgen/PerfReader.h
@@ -73,6 +73,7 @@ enum PerfContent {
   UnknownContent = 0,
   LBR = 1,      // Only LBR sample.
   LBRStack = 2, // Hybrid sample including call stack and LBR stack.
+  ArmSPE = 3,   // Arm SPE branch samples.
 };
 
 struct InputFile {
@@ -95,11 +96,18 @@ struct LBREntry {
 #endif
 };
 
+// Number of branch-stack entries PerfSample::TakenMask can flag, one bit each.
+// Arm BRBE, the deepest branch stack in hardware, holds 64 records.
+constexpr size_t TakenMaskWidth = 64;
+
 #ifndef NDEBUG
-static inline void printLBRStack(const SmallVectorImpl<LBREntry> &LBRStack) {
+static inline void printLBRStack(const SmallVectorImpl<LBREntry> &LBRStack,
+                                 uint64_t TakenMask = ~uint64_t(0)) {
   for (size_t I = 0; I < LBRStack.size(); I++) {
     dbgs() << "[" << I << "] ";
     LBRStack[I].print();
+    if (I < TakenMaskWidth && !((TakenMask >> I) & 1))
+      dbgs() << " (not taken)";
     dbgs() << "\n";
   }
 }
@@ -154,6 +162,19 @@ struct PerfSample {
   // Call stack recorded in FILO(leaf to root) order, it's used for CS-profile
   // generation
   SmallVector<uint64_t, 16> CallStack;
+  // Bit I is clear where LBRStack[I] is a branch that was not taken, which perf
+  // flags with the N suffix of the prediction field; such an entry names the
+  // instruction that follows the branch as its target. Only Arm SPE records
+  // such a branch. Kept beside the stack so that an LBREntry stays two
+  // addresses wide. Every bit starts set, so a reader of taken branches alone
+  // leaves the mask as it is.
+  uint64_t TakenMask = ~uint64_t(0);
+
+  // Whether the branch of LBRStack[Index] was taken. An entry past the width of
+  // the mask carries no flag and counts as taken.
+  bool taken(size_t Index) const {
+    return Index >= TakenMaskWidth || (TakenMask >> Index) & 1;
+  }
 
   virtual ~PerfSample() = default;
   uint64_t getHashCode() const {
@@ -169,6 +190,9 @@ struct PerfSample {
       Hash = HashCombine(Hash, Entry.Source);
       Hash = HashCombine(Hash, Entry.Target);
     }
+    // TakenMask is left out, so a taken and a not-taken record holding the
+    // same two addresses hash alike, which a conditional branch to the
+    // instruction that follows it produces; isEqual tells the two apart.
     return Hash;
   }
 
@@ -176,8 +200,11 @@ struct PerfSample {
     const SmallVector<uint64_t, 16> &OtherCallStack = Other->CallStack;
     const SmallVector<LBREntry, 16> &OtherLBRStack = Other->LBRStack;
 
+    // A branch that was taken and one that was not are separate events even
+    // where they hold the same two addresses, so the masks take part.
     if (CallStack.size() != OtherCallStack.size() ||
-        LBRStack.size() != OtherLBRStack.size())
+        LBRStack.size() != OtherLBRStack.size() ||
+        TakenMask != Other->TakenMask)
       return false;
 
     if (!std::equal(CallStack.begin(), CallStack.end(), OtherCallStack.begin()))
@@ -197,7 +224,7 @@ struct PerfSample {
   void print() const {
     dbgs() << "Line " << Linenum << "\n";
     dbgs() << "LBR stack\n";
-    printLBRStack(LBRStack);
+    printLBRStack(LBRStack, TakenMask);
     dbgs() << "Call stack\n";
     printCallStack(CallStack);
   }
@@ -646,25 +673,53 @@ class PerfScriptReader : public PerfReaderBase {
   void parseEventOrSample(TraceStream &TraceIt);
   // Warn if the relevant mmap event is missing.
   void warnIfMissingMMap();
+  // Reject an input that holds no record this reader can use, once the whole
+  // trace has been read. A branch-stack reader recognizes the shape of its
+  // trace before reading it, so this does nothing for one.
+  virtual void validateParsedTrace() {}
   // Emit accumulate warnings.
   void warnTruncatedStack();
   // Warn if range is invalid.
-  void warnInvalidRange();
+  virtual void warnInvalidRange();
   // Warn if sampled branch/target addresses don't match the binary.
   void warnIfBranchTargetMismatch();
   // Extract call stack from the perf trace lines
   bool extractCallstack(TraceStream &TraceIt,
                         SmallVectorImpl<uint64_t> &CallStack);
-  // Extract LBR stack from one perf trace line
-  bool extractLBRStack(TraceStream &TraceIt,
-                       SmallVectorImpl<LBREntry> &LBRStack);
-  uint64_t parseAggregatedCount(TraceStream &TraceIt);
+  // Parse one branch-stack entry "<source>/<destination>/<flags>...", the way
+  // `perf script -F brstack` prints it, into addresses canonicalized against
+  // the runtime mapping and the one flag the profile depends on. Return false
+  // for any other token, which each reader reports in its own way.
+  bool parseBranchEntry(StringRef Token, uint64_t &Source, uint64_t &Target,
+                        bool &Taken);
+  // Read from the prediction field of a branch-stack entry whether the branch
+  // was taken, or std::nullopt for a field the reader cannot read, which makes
+  // the entry unreadable. A hardware branch stack holds taken branches alone,
+  // so the base implementation answers yes without reading the field.
+  virtual std::optional<bool> parseTakenFlag(StringRef) { return true; }
+  // Extract LBR stack from one perf trace line. Bit I of \p TakenMask flags
+  // LBRStack[I], as PerfSample::TakenMask describes.
+  virtual bool extractLBRStack(TraceStream &TraceIt,
+                               SmallVectorImpl<LBREntry> &LBRStack,
+                               uint64_t &TakenMask);
+  // Consume a line that holds a decimal repeat count alone and return it, or
+  // return 1 and leave the line for the record parser. Override this where a
+  // record cannot be preceded by such a line, so it is not read as a count.
+  virtual uint64_t parseAggregatedCount(TraceStream &TraceIt);
   // Parse one sample from multiple perf lines, override this for different
   // sample type
   void parseSample(TraceStream &TraceIt);
   // An aggregated count is given to indicate how many times the sample is
   // repeated.
   virtual void parseSample(TraceStream &TraceIt, uint64_t Count){};
+  // Return the end of the executed range that follows the most recent branch of
+  // the sample, or zero when the reader cannot infer it. Consecutive entries of
+  // a branch stack bound every other range of the sample, but not this one, so
+  // the LBR and BRBE readers omit it while a reader of single-branch records
+  // infers it from the binary. Repeat is how often the sample was recorded.
+  virtual uint64_t inferFirstRangeEnd(const PerfSample *, uint64_t) {
+    return 0;
+  }
   void computeCounterFromLBR(const PerfSample *Sample, uint64_t Repeat);
   // Post process the profile after trace aggregation, we will do simple range
   // overlap computation for AutoFDO, or unwind for CSSPGO(hybrid sample).
@@ -695,6 +750,65 @@ class LBRPerfReader : public PerfScriptReader {
   void parseSample(TraceStream &TraceIt, uint64_t Count) override;
 };
 
+// The reader of an Arm SPE branch profile, whose record holds a single branch
+// printed as a branch stack of one entry and fed into the same aggregation.
+class ArmSPEReader : public LBRPerfReader {
+public:
+  ArmSPEReader(ProfiledBinary *Binary, StringRef PerfTrace,
+               std::optional<int32_t> PID, bool ConvertedFromPerfData)
+      : LBRPerfReader(Binary, PerfTrace, PID),
+        ConvertedFromPerfData(ConvertedFromPerfData) {}
+
+protected:
+  // Weigh every Arm SPE record by one, since perf prints no repeat count for it
+  // and a record damaged down to its instruction pointer would read as a count.
+  uint64_t parseAggregatedCount(TraceStream &TraceIt) override;
+  // Read the N flag perf appends to the prediction field of a branch that was
+  // not taken, which only Arm SPE records.
+  std::optional<bool> parseTakenFlag(StringRef PredictionFlags) override;
+  // Parse one Arm SPE branch record, and reject a record that holds more than
+  // the sampled branch, which a trace decoded without --itrace=bl1 carries.
+  bool extractLBRStack(TraceStream &TraceIt,
+                       SmallVectorImpl<LBREntry> &LBRStack,
+                       uint64_t &TakenMask) override;
+  // Infer from the binary the range executed after the recorded branch, and
+  // count the sample as reduced when no function range covers its destination.
+  uint64_t inferFirstRangeEnd(const PerfSample *Sample,
+                              uint64_t Repeat) override;
+  // Reject a trace that holds no Arm SPE branch record at all, naming the
+  // record format the lines were read as.
+  void validateParsedTrace() override;
+  // Report nothing: the base report describes ranges bounded by consecutive
+  // branch-stack entries, which a single-branch record cannot form.
+  void warnInvalidRange() override {}
+  // Compute the counters, then report what the trace lost on the way to them.
+  void generateUnsymbolizedProfile() override;
+
+private:
+  // Report, as shares of the sample lines of the trace, the lines that hold no
+  // record, the records that carry no counter and the samples whose range was
+  // reduced, and report a trace that yields no counter at all.
+  void reportTraceLosses();
+
+  // Number of parsable Arm SPE branch records, whether or not they yield a
+  // counter for this binary.
+  uint64_t NumRecords = 0;
+  // Number of lines that hold no readable record, every line of a trace printed
+  // with another field list among them.
+  uint64_t NumInvalidRecords = 0;
+  // Number of records dropped because their destination is not an instruction
+  // of this binary: a branch leaving the binary, a trace of another process or
+  // build, or a record whose branch target address was lost.
+  uint64_t NumDroppedRecords = 0;
+  // Number of samples, weighed by how often each was recorded, whose range was
+  // reduced to the sampled instruction because no function range covers it.
+  uint64_t NumUncoveredSamples = 0;
+  // Whether this tool printed the trace from a --perfdata recording. Only the
+  // report of an input that holds no record depends on it: the perf command to
+  // correct is one the user wrote for a --perfscript input alone.
+  bool ConvertedFromPerfData = false;
+};
+
 /*
   Hybrid perf script includes a group of hybrid samples(LBRs + call stack),
   which is used to generate CS profile. An example of hybrid sample:
diff --git a/llvm/tools/llvm-profgen/ProfiledBinary.cpp b/llvm/tools/llvm-profgen/ProfiledBinary.cpp
index 188fb2c20608c..f1bc278dca988 100644
--- a/llvm/tools/llvm-profgen/ProfiledBinary.cpp
+++ b/llvm/tools/llvm-profgen/ProfiledBinary.cpp
@@ -26,6 +26,7 @@
 #include "llvm/Support/Format.h"
 #include "llvm/Support/TargetSelect.h"
 #include "llvm/TargetParser/Triple.h"
+#include <algorithm>
 #include <optional>
 
 #define DEBUG_TYPE "load-binary"
@@ -659,6 +660,12 @@ bool ProfiledBinary::dissassembleSymbol(std::size_t SI, ArrayRef<uint8_t> Bytes,
         if (MCDesc.isUnconditionalBranch())
           UncondBranchAddrSet.insert(Address);
         BranchAddressSet.insert(Address);
+      } else if (MCDesc.isBarrier() || MCDesc.isTrap()) {
+        // Transfers of control are handled above, so what reaches here stops
+        // execution without naming where it continues, the AArch64 BRK and UDF
+        // above all. Recording it lets range inference end a range there
+        // instead of running it into code the trap never let execute.
+        BarrierAddressSet.insert(Address);
       }
 
       if (MCDesc.isIndirectBranch()) {
@@ -1299,6 +1306,105 @@ void ProfiledBinary::inferMissingFrames(
   MissingContextInferrer->inferMissingFrames(Context, NewContext);
 }
 
+void ProfiledBinary::buildRangeEnds() {
+  // The merge below and findRangeEnd's binary search both need increasing
+  // addresses, and instructions are collected in section-header order, which an
+  // object file need not keep sorted, so copy and sort when it is not.
+  //
+  // TODO: Sort CodeAddressVec in disassemble() and reduce this to an assertion.
+  // getIndexForAddr and InstructionPointer already depend on that order, so a
+  // binary listed out of it yields a wrong profile on those paths as well.
+  std::vector<uint64_t> SortedAddresses;
+  ArrayRef<uint64_t> Addresses = CodeAddressVec;
+  if (!llvm::is_sorted(CodeAddressVec)) {
+    SortedAddresses.assign(CodeAddressVec.begin(), CodeAddressVec.end());
+    llvm::sort(SortedAddresses);
+    Addresses = SortedAddresses;
+  }
+
+  // Walk instructions and function ranges backwards at the same time, so that
+  // classifying every instruction costs one merge step instead of a search.
+  auto FuncIt = StartAddrToFuncRangeMap.rbegin();
+  const FuncRange *NextRange = nullptr;
+  for (size_t I = Addresses.size(); I != 0; --I) {
+    uint64_t Current = Addresses[I - 1];
+
+    // Advance to the last function range that starts at or before Current, then
+    // keep it only if Current is inside that range.
+    while (FuncIt != StartAddrToFuncRangeMap.rend() && Current < FuncIt->first)
+      ++FuncIt;
+    const FuncRange *CurrentRange = nullptr;
+    if (FuncIt != StartAddrToFuncRangeMap.rend() &&
+        Current < FuncIt->second.EndAddress)
+      CurrentRange = &FuncIt->second;
+
+    // An instruction outside every function range is never a valid range end,
+    // because a range must not extend into code whose function is unknown, and
+    // findRangeEnd rejects such an address before the search.
+    if (CurrentRange) {
+      // The following instruction continues the range only when it is adjacent
+      // and belongs to the same function. NextRange is null for the highest
+      // address and for a following instruction no function range covers.
+      bool FollowsAdjacently = I < Addresses.size() &&
+                               Addresses[I] - Current == getInstSize(Current);
+
+      // Comparing the owning function rather than the range keeps a function
+      // described by several adjacent ranges, as function splitting produces,
+      // in one inferred range. Ranges are owned by name, so two same-named
+      // functions laid out back to back are separated only by IsFuncEntry,
+      // which a symbol on the second one sets.
+      bool EndsRange = !NextRange || !FollowsAdjacently ||
+                       NextRange->Func != CurrentRange->Func ||
+                       (NextRange != CurrentRange && NextRange->IsFuncEntry);
+      if (EndsRange || addressIsTransfer(Current) ||
+          BarrierAddressSet.count(Current))
+        RangeEnds.push_back(Current);
+    }
+    NextRange = CurrentRange;
+  }
+  // The traversal is backwards, so restore increasing order for binary search.
+  std::reverse(RangeEnds.begin(), RangeEnds.end());
+  RangeEndsBuilt = true;
+}
+
+uint64_t ProfiledBinary::findRangeEnd(uint64_t Address, bool *Uncovered) {
+  // Assigned on every path, so that a caller never reads a value left over
+  // from an earlier query.
+  if (Uncovered)
+    *Uncovered = false;
+
+  if (!addressIsCode(Address))
+    return 0;
+
+  // Function ranges come from debug info, and from the symbol table only where
+  // the binary carries pseudo probes. Where none covers Address, the executed
+  // range cannot be proven, so fail closed and keep that instruction alone.
+  if (!findFuncRange(Address)) {
+    if (Uncovered)
+      *Uncovered = true;
+    return Address;
+  }
+
+  // Every instruction before a range end has that end as its own, so indexing
+  // the ends alone keeps the index proportional to their number.
+  if (!RangeEndsBuilt)
+    buildRangeEnds();
+
+  // The greatest covered address always ends a range, and every covered address
+  // is at or below it, so the search finds an end.
+  auto End = llvm::lower_bound(RangeEnds, Address);
+  assert(End != RangeEnds.end() &&
+         "an address a function range covers has a range end at or above it");
+  // Should it not, fail closed as an uncovered address does rather than read
+  // past the index or hide the reduction from the caller.
+  if (End == RangeEnds.end()) {
+    if (Uncovered)
+      *Uncovered = true;
+    return Address;
+  }
+  return *End;
+}
+
 InstructionPointer::InstructionPointer(const ProfiledBinary *Binary,
                                        uint64_t Address, bool RoundToNext)
     : Binary(Binary), Address(Address) {
diff --git a/llvm/tools/llvm-profgen/ProfiledBinary.h b/llvm/tools/llvm-profgen/ProfiledBinary.h
index e4af7c0fdf7f4..36949c89c05e3 100644
--- a/llvm/tools/llvm-profgen/ProfiledBinary.h
+++ b/llvm/tools/llvm-profgen/ProfiledBinary.h
@@ -283,6 +283,13 @@ class ProfiledBinary {
   // An array of Addresses of all instructions sorted in increasing order. The
   // sorting is needed to fast advance to the next forward/backward instruction.
   std::vector<uint64_t> CodeAddressVec;
+  // Addresses that end an inferred executable range: transfer instructions and
+  // the last instructions before a code gap or a function boundary. Sorted in
+  // increasing order and filled on the first findRangeEnd query, so an input
+  // that needs no range inference never pays for it.
+  std::vector<uint64_t> RangeEnds;
+  // Whether RangeEnds has been built, which its emptiness does not tell.
+  bool RangeEndsBuilt = false;
   // A set of call instruction addresses. Used by virtual unwinding.
   DenseSet<uint64_t> CallAddressSet;
   // A set of return instruction addresses. Used by virtual unwinding.
@@ -295,6 +302,10 @@ class ProfiledBinary {
   DenseSet<uint64_t> IndirectBranchAddressSet;
   // A set of branch target addresses (destinations of branches/calls).
   DenseSet<uint64_t> BranchTargetAddressSet;
+  // A set of the addresses of instructions that execution never continues past,
+  // other than the calls, returns and branches held above: a trap such as the
+  // AArch64 BRK and UDF above all. buildRangeEnds ends an inferred range there.
+  DenseSet<uint64_t> BarrierAddressSet;
 
   // Estimate and track function prolog and epilog ranges.
   PrologEpilogTracker ProEpilogTracker;
@@ -421,6 +432,9 @@ class ProfiledBinary {
   bool dissassembleSymbol(std::size_t SI, ArrayRef<uint8_t> Bytes,
                           SectionSymbolsTy &Symbols,
                           const object::SectionRef &Section);
+  /// Collect the addresses that end an inferred executable range into
+  /// RangeEnds. Called on the first findRangeEnd query.
+  void buildRangeEnds();
   /// Symbolize a given instruction pointer and return a full call context.
   SampleContextFrameVector symbolize(const InstructionPointer &IP,
                                      bool UseCanonicalFnName = false,
@@ -505,11 +519,27 @@ class ProfiledBinary {
   bool addressIsIndirectBranch(uint64_t Address) const {
     return IndirectBranchAddressSet.count(Address);
   }
+  // Whether Address is a direct branch that may fall through.
+  bool addressIsConditionalBranch(uint64_t Address) const {
+    return BranchAddressSet.count(Address) &&
+           !UncondBranchAddrSet.count(Address) &&
+           !IndirectBranchAddressSet.count(Address);
+  }
   bool addressIsTransfer(uint64_t Address) {
     return BranchAddressSet.count(Address) || RetAddressSet.count(Address) ||
            CallAddressSet.count(Address);
   }
 
+  // Return the end of the executable range that starts at Address: the nearest
+  // instruction at or after it that is a transfer or a trap, without crossing a
+  // code gap or a function boundary, or the last instruction before either
+  // boundary. Return zero when Address is not an instruction, and Address
+  // itself when no function range covers it, which keeps an inferred range out
+  // of code that may belong to another function. Set *Uncovered, when given, to
+  // whether the range had to be reduced to Address; it is assigned on every
+  // path.
+  uint64_t findRangeEnd(uint64_t Address, bool *Uncovered = nullptr);
+
   bool rangeCrossUncondBranch(uint64_t Start, uint64_t End) {
     if (Start >= End)
       return false;
diff --git a/llvm/tools/llvm-profgen/llvm-profgen.cpp b/llvm/tools/llvm-profgen/llvm-profgen.cpp
index 31fda19ee8456..ed4f2db8b8fd4 100644
--- a/llvm/tools/llvm-profgen/llvm-profgen.cpp
+++ b/llvm/tools/llvm-profgen/llvm-profgen.cpp
@@ -33,9 +33,13 @@ static cl::opt<std::string> PerfScriptFilename(
     "perfscript", cl::value_desc("perfscript"),
     cl::desc("Path of a trace created by the Linux `perf script` command. For "
              "LBR or BRBE input, the raw perf data must contain branch "
-             "stacks, for example from recording with -b. "
-             "Cannot be used with --perfdata, --unsymbolized-profile, or "
-             "--llvm-sample-profile."),
+             "stacks, for example from recording with -b. For "
+             "--spe-branch-profile input, the trace must contain only Arm SPE "
+             "branch samples, from recording with "
+             "arm_spe/branch_filter=1,event_filter=2/, and must be generated "
+             "with `--show-mmap-events --itrace=bl1 -F ip,brstack`. Cannot be "
+             "used with --perfdata, "
+             "--unsymbolized-profile, or --llvm-sample-profile."),
     cl::cat(ProfGenCategory));
 static cl::alias PSA("ps", cl::desc("Alias for --perfscript"),
                      cl::aliasopt(PerfScriptFilename));
@@ -44,8 +48,11 @@ static cl::opt<std::string> PerfDataFilename(
     "perfdata", cl::value_desc("perfdata"),
     cl::desc("Path of raw perf data created by the Linux perf tool. For LBR or "
              "BRBE input, it must contain branch stacks, for example from "
-             "recording with -b. Cannot be used with --perfscript, "
-             "--unsymbolized-profile, or --llvm-sample-profile."),
+             "recording with -b. For --spe-branch-profile input, it must "
+             "contain only Arm SPE branch samples, from recording with "
+             "arm_spe/branch_filter=1,event_filter=2/. Cannot be used with "
+             "--perfscript, --unsymbolized-profile, or "
+             "--llvm-sample-profile."),
     cl::cat(ProfGenCategory));
 static cl::alias PDA("pd", cl::desc("Alias for --perfdata"),
                      cl::aliasopt(PerfDataFilename));
@@ -104,11 +111,24 @@ static cl::opt<std::string>
 
 // Validate the command line input.
 static void validateCommandLine() {
+  bool HasPerfData = PerfDataFilename.getNumOccurrences() > 0;
+  bool HasPerfScript = PerfScriptFilename.getNumOccurrences() > 0;
+
+  // Arm SPE branch profiling only describes how a perf data or perfscript
+  // input is parsed. Reject the option for every other mode, where it would
+  // otherwise be silently meaningless.
+  if (ReadSPEBranchProfile) {
+    if (ShowDisassemblyOnly)
+      exitWithError("--spe-branch-profile cannot be used with "
+                    "--show-disassembly-only.");
+    if (!HasPerfData && !HasPerfScript)
+      exitWithError("--spe-branch-profile requires --perfscript or "
+                    "--perfdata input.");
+  }
+
   // Allow the missing perfscript if we only use to show binary disassembly.
   if (!ShowDisassemblyOnly) {
     // Validate input profile is provided only once
-    bool HasPerfData = PerfDataFilename.getNumOccurrences() > 0;
-    bool HasPerfScript = PerfScriptFilename.getNumOccurrences() > 0;
     bool HasUnsymbolizedProfile =
         UnsymbolizedProfFilename.getNumOccurrences() > 0;
     bool HasSampleProfile = SampleProfFilename.getNumOccurrences() > 0;
@@ -192,6 +212,15 @@ int main(int argc, const char *argv[]) {
       std::make_unique<ProfiledBinary>(BinaryPath, DebugBinPath);
   Binary->load(TargetTriple);
 
+  // Arm SPE samples describe AArch64 branch execution, so reject another
+  // target.
+  //
+  // TODO: --target-triple overrides the triple checked here, while the
+  // disassembler is looked up from the object file itself, so an override makes
+  // the two disagree.
+  if (ReadSPEBranchProfile && !Binary->getTriple().isAArch64())
+    exitWithError("--spe-branch-profile requires an AArch64 binary.");
+
   if (ShowDisassemblyOnly)
     return EXIT_SUCCESS;
 



More information about the llvm-commits mailing list