[flang-commits] [clang] [clang-tools-extra] [flang] [libcxx] [lld] [llvm] [mlir] [clang-tidy] Fix performance-noexcept-move-constructor for implicit noexcept(false) (PR #226488)
via flang-commits
flang-commits at lists.llvm.org
Sun Sep 27 06:21:44 PDT 2026
=?utf-8?q?Balázs?= Benics <benicsbalazs at gmail.com>,Dmitry Sidorov
<Dmitry.Sidorov at amd.com>,Konstantinos Parasyris <koparasy at gmail.com>,Mehdi
Amini <joker.eph at gmail.com>,Amit Tiwari <Amit.Tiwari at amd.com>,Florian Hahn
<flo at fhahn.com>,Kazu Hirata <kazu at google.com>,
Yordan =?utf-8?q?Vásquez?= <vyordangiovani at gmail.com>,geoffreygaren
<ggaren at apple.com>,Simon Pilgrim <llvm-dev at redking.me.uk>,Florian Hahn
<flo at fhahn.com>,David Green <david.green at arm.com>,"Michael G. Kazakov"
<mike.kazakov at gmail.com>,Matt Arsenault <Matthew.Arsenault at amd.com>,Alan
Zhao <ayzhao at google.com>,Prajit Rahul <prajitrahul05 at gmail.com>,Florian Hahn
<flo at fhahn.com>,Florian Hahn <flo at fhahn.com>,Florian Hahn <flo at fhahn.com>,Florian
Hahn <flo at fhahn.com>,Dan Salvato <dan at teamsalvato.com>,Alexander Richardson
<alexrichardson at google.com>,Alexander Richardson <alexrichardson at google.com>,Alexander
Richardson <alexrichardson at google.com>,Alexander Richardson
<alexrichardson at google.com>,Alexander Richardson <alexrichardson at google.com>,Aiden
Grossman <aidengrossman at google.com>,Letu Ren <fantasquex at gmail.com>,Alan
Zhao <ayzhao at google.com>,Craig Topper <craig.topper at sifive.com>,Timm Baeder
<tbaeder at redhat.com>,Reid Kleckner <rkleckner at nvidia.com>,Amilendra
Kodithuwakku <amilendra.kodithuwakku at arm.com>,Alexander Richardson
<alexrichardson at google.com>,Lang Hames <lhames at gmail.com>,Matt Arsenault
<Matthew.Arsenault at amd.com>,Lang Hames <lhames at gmail.com>,Letu Ren
<fantasquex at gmail.com>,Mamadou Wane <mamadouswane at gmail.com>,Florian Hahn
<flo at fhahn.com>,Florian Hahn <flo at fhahn.com>,"Chibuoyim (Wilson) Ogbonna"
<chibuoyim.faith.ogbonna at huawei.com>,David Green <david.green at arm.com>,Alexey
Bataev <a.bataev at outlook.com>,"Agaev G." <cpp.thread at gmail.com>
Message-ID:
In-Reply-To: <llvm.org/llvm/llvm-project/pull/226488 at github.com>
https://github.com/cpp-engine updated https://github.com/llvm/llvm-project/pull/226488
>From b513dc568af3516f0f5eb70245073bfcd055b608 Mon Sep 17 00:00:00 2001
From: Agaev G <cpp.thread at gmail.com>
Date: Fri, 25 Sep 2026 13:04:11 +0000
Subject: [PATCH 01/53] [clang-tidy] Fix performance-noexcept-move-constructor
for implicit noexcept(false)
---
.../performance/NoexceptFunctionBaseCheck.cpp | 17 ++++++++-------
.../performance/noexcept-move-constructor.cpp | 21 +++++++++++++++++++
2 files changed, 31 insertions(+), 7 deletions(-)
diff --git a/clang-tools-extra/clang-tidy/performance/NoexceptFunctionBaseCheck.cpp b/clang-tools-extra/clang-tidy/performance/NoexceptFunctionBaseCheck.cpp
index 33d80615fb19f..32d85e18dff95 100644
--- a/clang-tools-extra/clang-tidy/performance/NoexceptFunctionBaseCheck.cpp
+++ b/clang-tools-extra/clang-tidy/performance/NoexceptFunctionBaseCheck.cpp
@@ -26,13 +26,16 @@ void NoexceptFunctionBaseCheck::check(const MatchFinder::MatchResult &Result) {
// Don't complain about nothrow(false), but complain on nothrow(expr)
// where expr evaluates to false.
- const auto *ProtoType = FuncDecl->getType()->castAs<FunctionProtoType>();
- const Expr *NoexceptExpr = ProtoType->getNoexceptExpr();
- if (NoexceptExpr) {
- NoexceptExpr = NoexceptExpr->IgnoreImplicit();
- if (!isa<CXXBoolLiteralExpr>(NoexceptExpr))
- reportNoexceptEvaluatedToFalse(FuncDecl, NoexceptExpr);
- return;
+ const auto ExceptionSpecSourceRange = FuncDecl->getExceptionSpecSourceRange();
+ if (ExceptionSpecSourceRange.isValid()) {
+ const auto *ProtoType = FuncDecl->getType()->castAs<FunctionProtoType>();
+ const Expr *NoexceptExpr = ProtoType->getNoexceptExpr();
+ if (NoexceptExpr) {
+ NoexceptExpr = NoexceptExpr->IgnoreImplicit();
+ if (!isa<CXXBoolLiteralExpr>(NoexceptExpr))
+ reportNoexceptEvaluatedToFalse(FuncDecl, NoexceptExpr);
+ return;
+ }
}
const auto Diag = reportMissingNoexcept(FuncDecl);
diff --git a/clang-tools-extra/test/clang-tidy/checkers/performance/noexcept-move-constructor.cpp b/clang-tools-extra/test/clang-tidy/checkers/performance/noexcept-move-constructor.cpp
index 76e1d9e493421..42a1287163d92 100644
--- a/clang-tools-extra/test/clang-tidy/checkers/performance/noexcept-move-constructor.cpp
+++ b/clang-tools-extra/test/clang-tidy/checkers/performance/noexcept-move-constructor.cpp
@@ -2,6 +2,8 @@
// RUN: %check_clang_tidy -std=c++17-or-later -check-suffixes=,ERR %s performance-noexcept-move-constructor %t \
// RUN: -- --fix-errors -- -fexceptions -DENABLE_ERROR
+#include <utility>
+
namespace std
{
template <typename T>
@@ -181,6 +183,25 @@ struct O : virtual IntWrapper, ThrowOnAnything {
// CHECK-FIXES: O &operator=(O &&) noexcept = default;
};
+struct P {
+ P() = default;
+
+ P(P &&) = default;
+ // CHECK-MESSAGES: :[[@LINE-1]]:3: warning: move constructors should be marked noexcept [performance-noexcept-move-constructor]
+ // CHECK-FIXES: P(P &&) noexcept = default;
+ P &operator=(P &&) = default;
+ // CHECK-MESSAGES: :[[@LINE-1]]:6: warning: move assignment operators should be marked noexcept [performance-noexcept-move-constructor]
+ // CHECK-FIXES: P &operator=(P &&) noexcept = default;
+
+ InheritFromThrowOnAnything IFF;
+};
+
+void p() {
+ P P1{};
+ P P2{std::move(P1)};
+ P P3 = std::move(P2);
+}
+
class OK {};
void f() {
>From a05cc8d35dfa999a70fde51da3b50c9ef588de42 Mon Sep 17 00:00:00 2001
From: David Green <david.green at arm.com>
Date: Sat, 26 Sep 2026 12:14:12 +0100
Subject: [PATCH 02/53] [AArch64] Use SVE rev for full reverse shuffles with
SVE128 (#224589)
When the vscale_range is always 1, we can make use of the SVE rev
instruction to perform a full 128bit vector reverse.
Sidesteps #223597 for SVE128
---
.../Target/AArch64/AArch64ISelLowering.cpp | 12 +++++++-
.../sve-fixed-length-shuffle-reverse.ll | 30 +++++++++++--------
2 files changed, 29 insertions(+), 13 deletions(-)
diff --git a/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp b/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
index 101a3961779fe..d61304154cad7 100644
--- a/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
+++ b/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
@@ -16214,8 +16214,18 @@ SDValue AArch64TargetLowering::LowerVECTOR_SHUFFLE(SDValue Op,
DAG.getNode(AArch64ISD::NVCAST, DL, BSVT, V1)));
}
- if (((NumElts == 8 && EltSize == 16) || (NumElts == 16 && EltSize == 8)) &&
+ if (((NumElts == 8 && EltSize == 16) || (NumElts == 16 && EltSize == 8) ||
+ (NumElts == 4 && EltSize == 32)) &&
ShuffleVectorInst::isReverseMask(ShuffleMask, ShuffleMask.size())) {
+ // For sve128 we can use a REV full vector reverse.
+ if (Subtarget->isSVEorStreamingSVEAvailable() &&
+ Subtarget->getSVEVectorSizeInBits() == 128) {
+ EVT ContainerVT = getContainerForFixedLengthVector(DAG, VT);
+ V1 = convertToScalableVector(DAG, ContainerVT, V1);
+ SDValue Rev = DAG.getNode(ISD::VECTOR_REVERSE, DL, ContainerVT, V1);
+ return convertFromScalableVector(DAG, VT, Rev);
+ }
+
SDValue Rev = DAG.getNode(AArch64ISD::REV64, DL, VT, V1);
return DAG.getNode(AArch64ISD::EXT, DL, VT, Rev, Rev,
DAG.getConstant(8, DL, MVT::i32));
diff --git a/llvm/test/CodeGen/AArch64/sve-fixed-length-shuffle-reverse.ll b/llvm/test/CodeGen/AArch64/sve-fixed-length-shuffle-reverse.ll
index 6dffe8d267353..0bd3afe63dbc0 100644
--- a/llvm/test/CodeGen/AArch64/sve-fixed-length-shuffle-reverse.ll
+++ b/llvm/test/CodeGen/AArch64/sve-fixed-length-shuffle-reverse.ll
@@ -40,8 +40,9 @@ define <2 x double> @testrev_v2f64(<2 x double> %x) {
define <4 x i32> @testrev_v4i32_vscale1(<4 x i32> %x) vscale_range(1,1) {
; CHECK-LABEL: testrev_v4i32_vscale1:
; CHECK: // %bb.0:
-; CHECK-NEXT: rev64 v0.4s, v0.4s
-; CHECK-NEXT: ext v0.16b, v0.16b, v0.16b, #8
+; CHECK-NEXT: // kill: def $q0 killed $q0 def $z0
+; CHECK-NEXT: rev z0.s, z0.s
+; CHECK-NEXT: // kill: def $q0 killed $q0 killed $z0
; CHECK-NEXT: ret
%r = shufflevector <4 x i32> %x, <4 x i32> poison, <4 x i32> <i32 3, i32 2, i32 1, i32 0>
ret <4 x i32> %r
@@ -60,8 +61,9 @@ define <4 x i32> @testrev_v4i32(<4 x i32> %x) {
define <4 x float> @testrev_v4f32_vscale1(<4 x float> %x) vscale_range(1,1) {
; CHECK-LABEL: testrev_v4f32_vscale1:
; CHECK: // %bb.0:
-; CHECK-NEXT: rev64 v0.4s, v0.4s
-; CHECK-NEXT: ext v0.16b, v0.16b, v0.16b, #8
+; CHECK-NEXT: // kill: def $q0 killed $q0 def $z0
+; CHECK-NEXT: rev z0.s, z0.s
+; CHECK-NEXT: // kill: def $q0 killed $q0 killed $z0
; CHECK-NEXT: ret
%r = shufflevector <4 x float> %x, <4 x float> poison, <4 x i32> <i32 3, i32 2, i32 1, i32 0>
ret <4 x float> %r
@@ -80,8 +82,9 @@ define <4 x float> @testrev_v4f32(<4 x float> %x) {
define <8 x i16> @testrev_v8i16_vscale1(<8 x i16> %x) vscale_range(1,1) {
; CHECK-LABEL: testrev_v8i16_vscale1:
; CHECK: // %bb.0:
-; CHECK-NEXT: rev64 v0.8h, v0.8h
-; CHECK-NEXT: ext v0.16b, v0.16b, v0.16b, #8
+; CHECK-NEXT: // kill: def $q0 killed $q0 def $z0
+; CHECK-NEXT: rev z0.h, z0.h
+; CHECK-NEXT: // kill: def $q0 killed $q0 killed $z0
; CHECK-NEXT: ret
%r = shufflevector <8 x i16> %x, <8 x i16> poison, <8 x i32> <i32 7, i32 6, i32 5, i32 4, i32 3, i32 2, i32 1, i32 0>
ret <8 x i16> %r
@@ -100,8 +103,9 @@ define <8 x i16> @testrev_v8i16(<8 x i16> %x) {
define <8 x half> @testrev_v8f16_vscale1(<8 x half> %x) vscale_range(1,1) {
; CHECK-LABEL: testrev_v8f16_vscale1:
; CHECK: // %bb.0:
-; CHECK-NEXT: rev64 v0.8h, v0.8h
-; CHECK-NEXT: ext v0.16b, v0.16b, v0.16b, #8
+; CHECK-NEXT: // kill: def $q0 killed $q0 def $z0
+; CHECK-NEXT: rev z0.h, z0.h
+; CHECK-NEXT: // kill: def $q0 killed $q0 killed $z0
; CHECK-NEXT: ret
%r = shufflevector <8 x half> %x, <8 x half> poison, <8 x i32> <i32 7, i32 6, i32 5, i32 4, i32 3, i32 2, i32 1, i32 0>
ret <8 x half> %r
@@ -120,8 +124,9 @@ define <8 x half> @testrev_v8f16(<8 x half> %x) {
define <8 x bfloat> @testrev_v8bf16_vscale1(<8 x bfloat> %x) vscale_range(1,1) {
; CHECK-LABEL: testrev_v8bf16_vscale1:
; CHECK: // %bb.0:
-; CHECK-NEXT: rev64 v0.8h, v0.8h
-; CHECK-NEXT: ext v0.16b, v0.16b, v0.16b, #8
+; CHECK-NEXT: // kill: def $q0 killed $q0 def $z0
+; CHECK-NEXT: rev z0.h, z0.h
+; CHECK-NEXT: // kill: def $q0 killed $q0 killed $z0
; CHECK-NEXT: ret
%r = shufflevector <8 x bfloat> %x, <8 x bfloat> poison, <8 x i32> <i32 7, i32 6, i32 5, i32 4, i32 3, i32 2, i32 1, i32 0>
ret <8 x bfloat> %r
@@ -140,8 +145,9 @@ define <8 x bfloat> @testrev_v8bf16(<8 x bfloat> %x) {
define <16 x i8> @testrev_v16i8_vscale1(<16 x i8> %x) vscale_range(1,1) {
; CHECK-LABEL: testrev_v16i8_vscale1:
; CHECK: // %bb.0:
-; CHECK-NEXT: rev64 v0.16b, v0.16b
-; CHECK-NEXT: ext v0.16b, v0.16b, v0.16b, #8
+; CHECK-NEXT: // kill: def $q0 killed $q0 def $z0
+; CHECK-NEXT: rev z0.b, z0.b
+; CHECK-NEXT: // kill: def $q0 killed $q0 killed $z0
; CHECK-NEXT: ret
%r = shufflevector <16 x i8> %x, <16 x i8> poison, <16 x i32> <i32 15, i32 14, i32 13, i32 12, i32 11, i32 10, i32 9, i32 8, i32 7, i32 6, i32 5, i32 4, i32 3, i32 2, i32 1, i32 0>
ret <16 x i8> %r
>From 65790a54de852dddbdc0503bbf88c24e72c2e78f Mon Sep 17 00:00:00 2001
From: Lang Hames <lhames at gmail.com>
Date: Sat, 26 Sep 2026 21:23:34 +1000
Subject: [PATCH 03/53] [orc-rt] SimpleRemoteCAOverSocket: require a stream
socket. (#226676)
SimpleRemote framing reads each message in as many parts as the socket
delivers it, which a socket that preserves message boundaries (datagram,
seqpacket) would truncate. Check SO_TYPE up front and reject anything
but SOCK_STREAM; the socket is owned by then, so it is closed on
failure.
Assisted-by: Claude
---
.../bedrock/sps/SimpleRemoteCAOverSocket.h | 3 ++-
.../sys/posix/sps/SimpleRemoteCAOverSocket.cpp | 10 ++++++++++
.../bedrock/sps/SimpleRemoteCAOverSocketTest.cpp | 16 ++++++++++++++++
3 files changed, 28 insertions(+), 1 deletion(-)
diff --git a/orc-rt/include/orc-rt/bedrock/sps/SimpleRemoteCAOverSocket.h b/orc-rt/include/orc-rt/bedrock/sps/SimpleRemoteCAOverSocket.h
index e523933e5cd48..afd2bbfc851d2 100644
--- a/orc-rt/include/orc-rt/bedrock/sps/SimpleRemoteCAOverSocket.h
+++ b/orc-rt/include/orc-rt/bedrock/sps/SimpleRemoteCAOverSocket.h
@@ -23,7 +23,8 @@
namespace orc_rt {
/// Creates a ControllerAccess that carries SimpleRemote messages over Sock,
-/// taking ownership of it. Sock must be a connected stream socket.
+/// taking ownership of it. Fails if Sock is not a stream socket. Sock must be
+/// connected.
///
/// The result is ready to hand to Session::attach, which is what starts the
/// conversation; nothing is sent before then.
diff --git a/orc-rt/lib/bedrock/sys/posix/sps/SimpleRemoteCAOverSocket.cpp b/orc-rt/lib/bedrock/sys/posix/sps/SimpleRemoteCAOverSocket.cpp
index d70ec224b546b..739aa580bf3b5 100644
--- a/orc-rt/lib/bedrock/sys/posix/sps/SimpleRemoteCAOverSocket.cpp
+++ b/orc-rt/lib/bedrock/sys/posix/sps/SimpleRemoteCAOverSocket.cpp
@@ -234,6 +234,16 @@ WrapperFunctionBuffer SocketSimpleRemoteCA::IncomingMessage::take() {
Expected<std::shared_ptr<SocketSimpleRemoteCA>>
SocketSimpleRemoteCA::Create(Session &S, SocketHandle Sock) {
// Sock is owned here, so every early return below closes it.
+
+ // This class assumes a byte stream, so check that this is a SOCK_STREAM.
+ int Type;
+ socklen_t TypeLen = sizeof(Type);
+ if (::getsockopt(Sock.get(), SOL_SOCKET, SO_TYPE, &Type, &TypeLen) != 0)
+ return makeError("getsockopt(SO_TYPE)", errno);
+ if (Type != SOCK_STREAM)
+ return make_error<StringError>(
+ "SimpleRemote over a socket requires a stream socket");
+
if (auto Err = setNonBlocking(Sock.get()))
return std::move(Err);
diff --git a/orc-rt/test/unit/bedrock/sps/SimpleRemoteCAOverSocketTest.cpp b/orc-rt/test/unit/bedrock/sps/SimpleRemoteCAOverSocketTest.cpp
index 4878c98137ed9..9a51eb7a58c2b 100644
--- a/orc-rt/test/unit/bedrock/sps/SimpleRemoteCAOverSocketTest.cpp
+++ b/orc-rt/test/unit/bedrock/sps/SimpleRemoteCAOverSocketTest.cpp
@@ -27,6 +27,7 @@
#include "BedrockTestUtils.h"
#include "CommonTestUtils.h"
+#include "ErrorMatchers.h"
#include "bedrock/SocketTestUtils.h"
#include "orc-rt-internal/support/Endian.h"
@@ -44,6 +45,8 @@
using namespace orc_rt;
using namespace orc_rt::test;
+using ::testing::HasSubstr;
+
namespace orc_rt {
/// Plays the controller against a SimpleRemote CA running over a socket.
@@ -281,6 +284,19 @@ void outOfBandErrorWrapper(orc_rt_SessionRef S,
} // namespace
+TEST_F(SimpleRemoteCAOverSocketTest, RejectsANonStreamSocket) {
+ // The framing reads a message in as many parts as the stream delivers it, so
+ // a socket that preserves message boundaries would truncate one.
+ auto H = makeNativeNonStreamSocket();
+ ASSERT_TRUE(H.has_value()) << "could not create a socket for the test";
+
+ EXPECT_THAT_EXPECTED(
+ createSimpleRemoteCAOverSocket(S, SocketHandle(*H)),
+ FailedWithMessage(HasSubstr("requires a stream socket")));
+ EXPECT_FALSE(isNativeSocketOpen(*H))
+ << "a rejected socket is still owned, and must be closed";
+}
+
TEST_F(SimpleRemoteCAOverSocketTest, SetupIsSentOnConnect) {
ASSERT_FALSE(!!attachOverSocket());
>From e871edaadf4323074e6cd51e33c12b2582b995d7 Mon Sep 17 00:00:00 2001
From: Ben Shi <2283975856 at qq.com>
Date: Sat, 26 Sep 2026 19:31:45 +0800
Subject: [PATCH 04/53] [AVR][NFC] Improve some comment messages (#226423)
---
llvm/lib/Target/AVR/AVRAsmPrinter.cpp | 2 +-
llvm/lib/Target/AVR/AVRDevices.td | 8 ++++----
llvm/lib/Target/AVR/AVRExpandPseudoInsts.cpp | 4 ++--
llvm/lib/Target/AVR/AVRISelLowering.cpp | 13 +++++++------
llvm/lib/Target/AVR/AVRInstrInfo.cpp | 10 +++++-----
llvm/lib/Target/AVR/AVRInstrInfo.td | 19 ++++++++++---------
llvm/lib/Target/AVR/AVRRegisterInfo.cpp | 4 ++--
.../Target/AVR/MCTargetDesc/AVRMCAsmInfo.cpp | 2 +-
.../Target/AVR/MCTargetDesc/AVRMCTargetDesc.h | 3 +--
llvm/lib/Target/AVR/TODO.md | 7 -------
10 files changed, 33 insertions(+), 39 deletions(-)
delete mode 100644 llvm/lib/Target/AVR/TODO.md
diff --git a/llvm/lib/Target/AVR/AVRAsmPrinter.cpp b/llvm/lib/Target/AVR/AVRAsmPrinter.cpp
index e7dd63fd39523..cd653dade917f 100644
--- a/llvm/lib/Target/AVR/AVRAsmPrinter.cpp
+++ b/llvm/lib/Target/AVR/AVRAsmPrinter.cpp
@@ -172,7 +172,7 @@ bool AVRAsmPrinter::PrintAsmMemoryOperand(const MachineInstr *MI,
assert(MO.isReg() && "Unexpected inline asm memory operand");
// TODO: We should be able to look up the alternative name for
- // the register if it's given.
+ // the register if it's given.
// TableGen doesn't expose a way of getting retrieving names
// for registers.
if (MI->getOperand(OpNum).getReg() == AVR::R31R30) {
diff --git a/llvm/lib/Target/AVR/AVRDevices.td b/llvm/lib/Target/AVR/AVRDevices.td
index ad760d7403573..71de1a7c6cf54 100644
--- a/llvm/lib/Target/AVR/AVRDevices.td
+++ b/llvm/lib/Target/AVR/AVRDevices.td
@@ -2,10 +2,10 @@
// AVR Device Definitions
//===---------------------------------------------------------------------===//
-// :TODO: Implement the skip errata, see `gcc/config/avr/avr-arch.h` for details
-// :TODO: We define all devices with SRAM to have all variants of LD/ST/LDD/STD.
-// In reality, avr1 (no SRAM) has one variant each of `LD` and `ST`.
-// avr2 (with SRAM) adds the rest of the variants.
+// TODO: Implement the skip errata, see `gcc/config/avr/avr-arch.h` for details
+// TODO: We define all devices with SRAM to have all variants of LD/ST/LDD/STD.
+// In reality, avr1 (no SRAM) has one variant each of `LD` and `ST`.
+// avr2 (with SRAM) adds the rest of the variants.
// A feature set aggregates features, grouping them. We don't want to create a
// new member in AVRSubtarget (to store a value) for each set because we do not
diff --git a/llvm/lib/Target/AVR/AVRExpandPseudoInsts.cpp b/llvm/lib/Target/AVR/AVRExpandPseudoInsts.cpp
index f9da7b3cfe7e9..f97e7a8ed08b2 100644
--- a/llvm/lib/Target/AVR/AVRExpandPseudoInsts.cpp
+++ b/llvm/lib/Target/AVR/AVRExpandPseudoInsts.cpp
@@ -1184,7 +1184,7 @@ bool AVRExpandPseudo::expand<AVR::STWPtrRr>(Block &MBB, BlockIt MBBI) {
bool SrcIsKill = MI.getOperand(1).isKill();
const AVRSubtarget &STI = MBB.getParent()->getSubtarget<AVRSubtarget>();
- //: TODO: need to reverse this order like inw and stsw?
+ // TODO: Need to reverse this order like inw and stsw?
if (STI.hasTinyEncoding()) {
// Handle this case in the expansion of STDWPtrQRr because it is very
@@ -2680,7 +2680,7 @@ bool AVRExpandPseudo::expandMI(Block &MBB, BlockIt MBBI) {
EXPAND(AVR::LDWRdPtr);
EXPAND(AVR::LDWRdPtrPi);
EXPAND(AVR::LDWRdPtrPd);
- case AVR::LDDWRdYQ: //: FIXME: remove this once PR13375 gets fixed
+ case AVR::LDDWRdYQ: // FIXME: Remove this once PR13375 gets fixed.
EXPAND(AVR::LDDWRdPtrQ);
EXPAND(AVR::LPMBRdZ);
EXPAND(AVR::LPMWRdZ);
diff --git a/llvm/lib/Target/AVR/AVRISelLowering.cpp b/llvm/lib/Target/AVR/AVRISelLowering.cpp
index 79ed56679120b..bb7caa275718a 100644
--- a/llvm/lib/Target/AVR/AVRISelLowering.cpp
+++ b/llvm/lib/Target/AVR/AVRISelLowering.cpp
@@ -195,9 +195,9 @@ AVRTargetLowering::AVRTargetLowering(const AVRTargetMachine &TM,
for (MVT VT : MVT::integer_valuetypes()) {
setOperationAction(ISD::SIGN_EXTEND_INREG, VT, Expand);
// TODO: The generated code is pretty poor. Investigate using the
- // same "shift and subtract with carry" trick that we do for
- // extending 8-bit to 16-bit. This may require infrastructure
- // improvements in how we treat 16-bit "registers" to be feasible.
+ // same "shift and subtract with carry" trick that we do for
+ // extending 8-bit to 16-bit. This may require infrastructure
+ // improvements in how we treat 16-bit "registers" to be feasible.
}
setMinFunctionAlignment(Align(2));
@@ -235,7 +235,7 @@ SDValue AVRTargetLowering::LowerShifts(SDValue Op, SelectionDAG &DAG) const {
if (ShiftAmount == 16) {
// Special case these two operations because they appear to be used by the
// generic codegen parts to lower 32-bit numbers.
- // TODO: perhaps we can lower shift amounts bigger than 16 to a 16-bit
+ // TODO: Perhaps we can lower shift amounts bigger than 16 to a 16-bit
// shift of a part of the 32-bit value?
switch (Op.getOpcode()) {
case ISD::SHL: {
@@ -2267,8 +2267,9 @@ AVRTargetLowering::insertWideShift(MachineInstr &MI,
// - lshr prefers starting from the least significant byte (1st case).
// - for ashr it depends on the number of shifted bytes.
// Some shift operations still don't get the most optimal mov sequences even
- // with this distinction. TODO: figure out why and try to fix it (but we're
- // already equal to or faster than avr-gcc in all cases except ashr 8).
+ // with this distinction.
+ // TODO: Figure out why and try to fix it (but we're
+ // already equal to or faster than avr-gcc in all cases except ashr 8).
if (Opc != ISD::SHL &&
(Opc != ISD::SRA || (ShiftAmt < 16 || ShiftAmt >= 22))) {
// Use the resulting registers starting with the least significant byte.
diff --git a/llvm/lib/Target/AVR/AVRInstrInfo.cpp b/llvm/lib/Target/AVR/AVRInstrInfo.cpp
index f58dbdae46dd9..15f176fafb121 100644
--- a/llvm/lib/Target/AVR/AVRInstrInfo.cpp
+++ b/llvm/lib/Target/AVR/AVRInstrInfo.cpp
@@ -90,7 +90,7 @@ Register AVRInstrInfo::isLoadFromStackSlot(const MachineInstr &MI,
int &FrameIndex) const {
switch (MI.getOpcode()) {
case AVR::LDDRdPtrQ:
- case AVR::LDDWRdYQ: { //: FIXME: remove this once PR13375 gets fixed
+ case AVR::LDDWRdYQ: { // FIXME: Remove this once PR13375 gets fixed.
if (MI.getOperand(1).isFI() && MI.getOperand(2).isImm() &&
MI.getOperand(2).getImm() == 0) {
FrameIndex = MI.getOperand(1).getIndex();
@@ -175,7 +175,7 @@ void AVRInstrInfo::loadRegFromStackSlot(MachineBasicBlock &MBB,
Opcode = AVR::LDDRdPtrQ;
} else if (TRI.isTypeLegalForClass(*RC, MVT::i16)) {
// Opcode = AVR::LDDWRdPtrQ;
- //: FIXME: remove this once PR13375 gets fixed
+ // FIXME: Remove this once PR13375 gets fixed.
Opcode = AVR::LDDWRdYQ;
} else {
llvm_unreachable("Cannot load this register from a stack slot!");
@@ -285,7 +285,7 @@ bool AVRInstrInfo::analyzeBranch(MachineBasicBlock &MBB,
}
// Handle unconditional branches.
- //: TODO: add here jmp
+ // TODO: Add here jmp.
if (I->getOpcode() == AVR::RJMPk) {
UnCondBrIter = I;
@@ -443,8 +443,8 @@ unsigned AVRInstrInfo::removeBranch(MachineBasicBlock &MBB,
if (I->isDebugInstr()) {
continue;
}
- //: TODO: add here the missing jmp instructions once they are implemented
- // like jmp, {e}ijmp, and other cond branches, ...
+ // TODO: Add here the missing jmp instructions once they are implemented
+ // like jmp, {e}ijmp, and other cond branches, ...
if (I->getOpcode() != AVR::RJMPk &&
getCondFromBranchOpc(I->getOpcode()) == AVRCC::COND_INVALID) {
break;
diff --git a/llvm/lib/Target/AVR/AVRInstrInfo.td b/llvm/lib/Target/AVR/AVRInstrInfo.td
index 4b5be413ba669..0f8a7621f56ea 100644
--- a/llvm/lib/Target/AVR/AVRInstrInfo.td
+++ b/llvm/lib/Target/AVR/AVRInstrInfo.td
@@ -692,8 +692,8 @@ let isCall = 1 in {
// SP is marked as a use to prevent stack-pointer assignments that appear
// immediately before calls from potentially appearing dead.
//
- // TODO: the imm field can be either 16 or 22 bits in devices with more
- // than 64k of ROM, fix it once we support the largest devices.
+ // TODO: The imm field can be either 16 or 22 bits in devices with more
+ // than 64k of ROM, fix it once we support the largest devices.
let Uses = [SP] in
def CALLk : F32BRk<0b111, (outs), (ins call_target:$k), "call\t$k",
[(AVRcall imm:$k)]>,
@@ -1407,8 +1407,8 @@ def SWAPRd : FRd<0b1001, 0b0100010, (outs GPR8:$rd), (ins GPR8:$src),
"swap\t$rd", [(set i8:$rd, (AVRSwap i8:$src))]>;
// IO register bit set/clear operations.
-//: TODO: add patterns when popcount(imm)==2 to be expanded with 2 sbi/cbi
-// instead of in+ori+out which requires one more instr.
+// TODO: Add patterns when popcount(imm)==2 to be expanded with 2 sbi/cbi
+// instead of in+ori+out which requires one more instr.
let hasSideEffects = 1, mayStore = 1 in {
def SBIAb : FIOBIT<0b10, (outs), (ins imm_port5:$addr, i8imm:$b),
"sbi\t$addr, $b",
@@ -1526,7 +1526,7 @@ def WDR : F16<0b1001010110101000, (outs), (ins), "wdr", []>;
// Pseudo instructions for later expansion
//===----------------------------------------------------------------------===//
-//: TODO: Optimize this for wider types AND optimize the following code
+// TODO: Optimize this for wider types AND optimize the following code
// compile int foo(char a, char b, char c, char d) {return d+b;}
// looks like a missed sext_inreg opportunity.
def SEXT : ExtensionPseudo<(outs DREGS:$dt), (ins GPR8:$src), "sext\t$dt, $src",
@@ -1656,8 +1656,9 @@ def CopyZero : Pseudo<(outs GPR8:$rd), (ins), "clrz\t$rd", [(set i8:$rd, 0)]>;
// Non-Instruction Patterns
//===----------------------------------------------------------------------===//
-//: TODO: look in x86InstrCompiler.td for odd encoding trick related to
-// add x, 128 -> sub x, -128. Clang is emitting an eor for this (ldi+eor)
+// TODO: Look in x86InstrCompiler.td for odd encoding trick related to
+// `add x, 128` -> `sub x, -128`. Clang is emitting an eor for this
+// (ldi+eor).
// the add instruction always writes the carry flag
def : Pat<(addc i8 : $src, i8 : $src2), (ADDRdRr i8 : $src, i8 : $src2)>;
@@ -1741,8 +1742,8 @@ def : Pat<(i16(AVRWrapper tblockaddress :$dst)), (LDIWRdK tblockaddress:$dst)>;
def : Pat<(i8(trunc(AVRlsrwn DLDREGS:$src, (i16 8)))),
(EXTRACT_SUBREG DREGS:$src, sub_hi)>;
-// :FIXME: DAGCombiner produces an shl node after legalization from these seq:
-// BR_JT -> (mul x, 2) -> (shl x, 1)
+// FIXME: DAGCombiner produces an shl node after legalization from these seq:
+// BR_JT -> (mul x, 2) -> (shl x, 1) .
def : Pat<(shl i16 : $src1, (i8 1)), (LSLWRd i16 : $src1)>;
// Lowering of 'tst' node to 'TST' instruction.
diff --git a/llvm/lib/Target/AVR/AVRRegisterInfo.cpp b/llvm/lib/Target/AVR/AVRRegisterInfo.cpp
index ac86ffb321a32..836faf8edfe19 100644
--- a/llvm/lib/Target/AVR/AVRRegisterInfo.cpp
+++ b/llvm/lib/Target/AVR/AVRRegisterInfo.cpp
@@ -209,8 +209,8 @@ bool AVRRegisterInfo::eliminateFrameIndex(MachineBasicBlock::iterator II,
// If the offset is too big we have to adjust and restore the frame pointer
// to materialize a valid load/store with displacement.
- //: TODO: consider using only one adiw/sbiw chain for more than one frame
- //: index
+ // TODO: Consider using only one adiw/sbiw chain for more than one frame
+ // indexes.
if (Offset > MaxOffset) {
unsigned AddOpc = AVR::ADIWRdK, SubOpc = AVR::SBIWRdK;
int AddOffset = Offset - MaxOffset;
diff --git a/llvm/lib/Target/AVR/MCTargetDesc/AVRMCAsmInfo.cpp b/llvm/lib/Target/AVR/MCTargetDesc/AVRMCAsmInfo.cpp
index b8f2d99442974..2d77e58145f22 100644
--- a/llvm/lib/Target/AVR/MCTargetDesc/AVRMCAsmInfo.cpp
+++ b/llvm/lib/Target/AVR/MCTargetDesc/AVRMCAsmInfo.cpp
@@ -200,7 +200,7 @@ bool AVRMCAsmInfo::evaluateAsRelocatableImpl(const MCSpecifierExpr &Expr,
if (E.getSpecifier() == AVR::S_PM)
Spec = AVR::S_PM;
- // TODO: don't attach specifier to MCSymbolRefExpr.
+ // TODO: Don't attach specifier to MCSymbolRefExpr.
Result =
MCValue::get(Value.getAddSym(), nullptr, Value.getConstant(), Spec);
}
diff --git a/llvm/lib/Target/AVR/MCTargetDesc/AVRMCTargetDesc.h b/llvm/lib/Target/AVR/MCTargetDesc/AVRMCTargetDesc.h
index e83d674f87cc9..e28074cbcd121 100644
--- a/llvm/lib/Target/AVR/MCTargetDesc/AVRMCTargetDesc.h
+++ b/llvm/lib/Target/AVR/MCTargetDesc/AVRMCTargetDesc.h
@@ -32,8 +32,7 @@ class Target;
MCInstrInfo *createAVRMCInstrInfo();
/// Creates a machine code emitter for AVR.
-MCCodeEmitter *createAVRMCCodeEmitter(const MCInstrInfo &MCII,
- MCContext &Ctx);
+MCCodeEmitter *createAVRMCCodeEmitter(const MCInstrInfo &MCII, MCContext &Ctx);
/// Creates an assembly backend for AVR.
MCAsmBackend *createAVRAsmBackend(const Target &T, const MCSubtargetInfo &STI,
diff --git a/llvm/lib/Target/AVR/TODO.md b/llvm/lib/Target/AVR/TODO.md
deleted file mode 100644
index 3a333355646d6..0000000000000
--- a/llvm/lib/Target/AVR/TODO.md
+++ /dev/null
@@ -1,7 +0,0 @@
-# Write an XFAIL test for this `FIXME` in `AVRInstrInfo.td`
-
-```
-// :FIXME: DAGCombiner produces an shl node after legalization from these seq:
-// BR_JT -> (mul x, 2) -> (shl x, 1)
-```
-
>From 8546ddcd811457b1b78df73ca8d8bf55a9c45aba Mon Sep 17 00:00:00 2001
From: Mehdi Amini <joker.eph at gmail.com>
Date: Sat, 26 Sep 2026 14:08:21 +0200
Subject: [PATCH 05/53] [mlir] Enable strict properties in assembly formats by
default (#225742)
Make strict property assembly formats the default and remove redundant
per-dialect settings. Keep legacy test fixtures explicitly opted out and
update the documentation.
Downstream can opt out for now using `let
useStrictPropertiesInAssemblyFormat = 0;` ; This will be removed in a
future release.
Part of #155475
Assisted-by: Codex
---
.../clang/CIR/Dialect/IR/CIRDialect.td | 1 +
.../flang/Optimizer/Dialect/CUF/CUFDialect.td | 1 +
.../flang/Optimizer/Dialect/FIRCG/CGOps.td | 1 +
.../flang/Optimizer/Dialect/FIRDialect.td | 1 +
.../flang/Optimizer/Dialect/MIF/MIFDialect.td | 1 +
.../flang/Optimizer/HLFIR/HLFIROpBase.td | 1 +
mlir/docs/DefiningDialects/Operations.md | 20 +++++++---------
mlir/docs/DefiningDialects/_index.md | 24 +++++++++----------
mlir/docs/ReleaseNotes.md | 12 ++++++++++
mlir/examples/toy/Ch4/include/toy/Ops.td | 1 +
mlir/examples/toy/Ch5/include/toy/Ops.td | 1 +
mlir/examples/toy/Ch6/include/toy/Ops.td | 1 +
mlir/examples/toy/Ch7/include/toy/Ops.td | 1 +
.../mlir/Dialect/AMDGPU/IR/AMDGPUBase.td | 1 -
.../mlir/Dialect/Affine/IR/AffineOps.td | 1 -
.../mlir/Dialect/Arith/IR/ArithBase.td | 1 -
mlir/include/mlir/Dialect/ArmNeon/ArmNeon.td | 1 -
mlir/include/mlir/Dialect/ArmSME/IR/ArmSME.td | 1 -
mlir/include/mlir/Dialect/ArmSVE/IR/ArmSVE.td | 1 -
.../mlir/Dialect/Async/IR/AsyncDialect.td | 1 -
.../Bufferization/IR/BufferizationBase.td | 1 -
.../mlir/Dialect/Complex/IR/ComplexBase.td | 1 -
.../Dialect/ControlFlow/IR/ControlFlowOps.td | 1 -
mlir/include/mlir/Dialect/DLTI/DLTIBase.td | 1 -
.../mlir/Dialect/EmitC/IR/EmitCBase.td | 1 -
mlir/include/mlir/Dialect/Func/IR/FuncOps.td | 1 -
mlir/include/mlir/Dialect/GPU/IR/GPUBase.td | 1 -
mlir/include/mlir/Dialect/IRDL/IR/IRDL.td | 1 -
.../mlir/Dialect/Index/IR/IndexDialect.td | 1 -
.../mlir/Dialect/LLVMIR/LLVMDialect.td | 1 -
.../mlir/Dialect/LLVMIR/NVVMDialect.td | 1 -
.../mlir/Dialect/LLVMIR/ROCDLDialect.td | 1 -
mlir/include/mlir/Dialect/LLVMIR/VCIXOps.td | 1 -
mlir/include/mlir/Dialect/LLVMIR/XeVMOps.td | 1 -
.../mlir/Dialect/Linalg/IR/LinalgBase.td | 1 -
.../Dialect/MLProgram/IR/MLProgramBase.td | 1 -
mlir/include/mlir/Dialect/MPI/IR/MPI.td | 1 -
mlir/include/mlir/Dialect/Math/IR/MathBase.td | 1 -
.../mlir/Dialect/MemRef/IR/MemRefBase.td | 1 -
mlir/include/mlir/Dialect/NVGPU/IR/NVGPU.td | 1 -
.../mlir/Dialect/OpenACC/OpenACCBase.td | 1 -
.../mlir/Dialect/OpenMP/OpenMPDialect.td | 1 -
.../include/mlir/Dialect/PDL/IR/PDLDialect.td | 1 -
.../mlir/Dialect/PDLInterp/IR/PDLInterpOps.td | 1 -
.../include/mlir/Dialect/Ptr/IR/PtrDialect.td | 1 -
.../mlir/Dialect/Quant/IR/QuantBase.td | 1 -
mlir/include/mlir/Dialect/SCF/IR/SCFOps.td | 1 -
.../include/mlir/Dialect/SMT/IR/SMTDialect.td | 1 -
.../mlir/Dialect/SPIRV/IR/SPIRVBase.td | 1 -
.../mlir/Dialect/Shape/IR/ShapeBase.td | 1 -
.../mlir/Dialect/Shard/IR/ShardBase.td | 1 -
.../SparseTensor/IR/SparseTensorBase.td | 1 -
.../mlir/Dialect/Tensor/IR/TensorBase.td | 1 -
.../mlir/Dialect/Tosa/IR/TosaOpBase.td | 1 -
.../Dialect/Transform/IR/TransformDialect.td | 1 -
mlir/include/mlir/Dialect/UB/IR/UBOps.td | 1 -
mlir/include/mlir/Dialect/Vector/IR/Vector.td | 1 -
.../mlir/Dialect/WasmSSA/IR/WasmSSABase.td | 1 -
mlir/include/mlir/Dialect/X86/X86.td | 1 -
.../mlir/Dialect/XeGPU/IR/XeGPUDialect.td | 1 -
mlir/include/mlir/IR/BuiltinDialect.td | 1 -
mlir/include/mlir/IR/DialectBase.td | 11 +++++----
mlir/test/lib/Dialect/Test/TestDialect.td | 2 ++
mlir/test/mlir-tblgen/op-format-invalid.td | 3 ++-
mlir/test/mlir-tblgen/op-format.td | 3 ++-
mlir/test/python/python_test_ops.td | 1 +
66 files changed, 54 insertions(+), 80 deletions(-)
diff --git a/clang/include/clang/CIR/Dialect/IR/CIRDialect.td b/clang/include/clang/CIR/Dialect/IR/CIRDialect.td
index b2a2707a67b66..aee862afc02b4 100644
--- a/clang/include/clang/CIR/Dialect/IR/CIRDialect.td
+++ b/clang/include/clang/CIR/Dialect/IR/CIRDialect.td
@@ -17,6 +17,7 @@ include "mlir/IR/OpBase.td"
def CIR_Dialect : Dialect {
let name = "cir";
+ let useStrictPropertiesInAssemblyFormat = 0;
// A short one-line summary of our dialect.
let summary = "A high-level dialect for analyzing and optimizing Clang "
diff --git a/flang/include/flang/Optimizer/Dialect/CUF/CUFDialect.td b/flang/include/flang/Optimizer/Dialect/CUF/CUFDialect.td
index 0cabd55f1ea5a..984c457f04831 100644
--- a/flang/include/flang/Optimizer/Dialect/CUF/CUFDialect.td
+++ b/flang/include/flang/Optimizer/Dialect/CUF/CUFDialect.td
@@ -20,6 +20,7 @@ include "mlir/IR/OpBase.td"
def CUFDialect : Dialect {
let name = "cuf";
+ let useStrictPropertiesInAssemblyFormat = 0;
let summary = "CUDA Fortran dialect";
diff --git a/flang/include/flang/Optimizer/Dialect/FIRCG/CGOps.td b/flang/include/flang/Optimizer/Dialect/FIRCG/CGOps.td
index 5a7c85e7883fa..5409b81b59aa3 100644
--- a/flang/include/flang/Optimizer/Dialect/FIRCG/CGOps.td
+++ b/flang/include/flang/Optimizer/Dialect/FIRCG/CGOps.td
@@ -22,6 +22,7 @@ include "mlir/IR/BuiltinAttributes.td"
def fircg_Dialect : Dialect {
let name = "fircg";
+ let useStrictPropertiesInAssemblyFormat = 0;
let cppNamespace = "::fir::cg";
}
diff --git a/flang/include/flang/Optimizer/Dialect/FIRDialect.td b/flang/include/flang/Optimizer/Dialect/FIRDialect.td
index 415bf7a6a95fe..3ac9a12d5cf55 100644
--- a/flang/include/flang/Optimizer/Dialect/FIRDialect.td
+++ b/flang/include/flang/Optimizer/Dialect/FIRDialect.td
@@ -23,6 +23,7 @@ include "mlir/Interfaces/SideEffectInterfaces.td"
def FIROpsDialect : Dialect {
let name = "fir";
+ let useStrictPropertiesInAssemblyFormat = 0;
let cppNamespace = "::fir";
let useDefaultTypePrinterParser = 0;
let useDefaultAttributePrinterParser = 0;
diff --git a/flang/include/flang/Optimizer/Dialect/MIF/MIFDialect.td b/flang/include/flang/Optimizer/Dialect/MIF/MIFDialect.td
index f8be6a86a79fe..c4dccc9b77dba 100644
--- a/flang/include/flang/Optimizer/Dialect/MIF/MIFDialect.td
+++ b/flang/include/flang/Optimizer/Dialect/MIF/MIFDialect.td
@@ -20,6 +20,7 @@ include "mlir/IR/OpBase.td"
def MIFDialect : Dialect {
let name = "mif";
+ let useStrictPropertiesInAssemblyFormat = 0;
let summary = "Multi-Image Fortran dialect";
diff --git a/flang/include/flang/Optimizer/HLFIR/HLFIROpBase.td b/flang/include/flang/Optimizer/HLFIR/HLFIROpBase.td
index 1f46b69561d68..a0764077289aa 100644
--- a/flang/include/flang/Optimizer/HLFIR/HLFIROpBase.td
+++ b/flang/include/flang/Optimizer/HLFIR/HLFIROpBase.td
@@ -21,6 +21,7 @@ include "flang/Optimizer/Dialect/FIRTypes.td"
def hlfir_Dialect : Dialect {
let name = "hlfir";
+ let useStrictPropertiesInAssemblyFormat = 0;
let summary = "High Level Fortran IR.";
diff --git a/mlir/docs/DefiningDialects/Operations.md b/mlir/docs/DefiningDialects/Operations.md
index 5615af678d339..f991f5308c5ae 100644
--- a/mlir/docs/DefiningDialects/Operations.md
+++ b/mlir/docs/DefiningDialects/Operations.md
@@ -759,15 +759,12 @@ The available directives are as follows:
* `attr-dict`
- - Represents the attribute dictionary of the operation.
- - Any inherent attributes that are not used elsewhere in the format are
- printed as part of the attribute dictionary unless a `prop-dict` is
- present.
- - Discardable attributes are always part of the `attr-dict`.
- - For dialects that set `useStrictPropertiesInAssemblyFormat`,
- `attr-dict` only carries discardable attributes for property-backed
- operations. Inherent attributes must be bound directly in the format or
- covered by `prop-dict`.
+ - Represents the attribute dictionary of the operation. Under the
+ default strict format rules, it contains only discardable attributes.
+ Inherent attributes must be bound directly in the format or covered by
+ `prop-dict`. The deprecated `useStrictPropertiesInAssemblyFormat = 0`
+ setting temporarily allows inherent attributes to mix with discardable
+ attributes in `attr-dict`.
* `attr-dict-with-keyword`
@@ -1142,9 +1139,8 @@ to:
directives.
1. Unless all non-attribute properties appear in the format, the `prop-dict`
directive must be present.
-1. For dialects that set `useStrictPropertiesInAssemblyFormat`, every inherent
- attribute and property must either appear in the format or be covered by the
- `prop-dict` directive.
+1. Every inherent attribute and property must either appear in the format or
+ be covered by the `prop-dict` directive.
1. The `attr-dict` directive must always be present.
1. Must not contain overlapping information; e.g. multiple instances of
'attr-dict', types, operands, etc.
diff --git a/mlir/docs/DefiningDialects/_index.md b/mlir/docs/DefiningDialects/_index.md
index d4ffd066a61cd..22d067c0a66f6 100644
--- a/mlir/docs/DefiningDialects/_index.md
+++ b/mlir/docs/DefiningDialects/_index.md
@@ -274,22 +274,20 @@ For a more detail description of the expected usages of this hook, view the deta
### Strict Property Assembly Formats
-Dialects can set `useStrictPropertiesInAssemblyFormat` to require declarative
-assembly formats for property-backed operations to account for all inherent
-attributes and properties:
-
-```tablegen
-def MyDialect : Dialect {
- let useStrictPropertiesInAssemblyFormat = 1;
-}
-```
-
-This mode is disabled by default for now. When enabled, an operation format must
+Declarative assembly formats for property-backed operations must account for
+all inherent attributes and properties by default. An operation format must
either bind every inherent attribute and property directly in the format or
include the `prop-dict` directive. Generated parsers also reject inherent
attributes that arrive through `attr-dict`, so `attr-dict` only carries
-discardable attributes for these formats. See the
-[declarative assembly format](Operations.md/#declarative-assembly-format)
+discardable attributes for these formats.
+
+The `useStrictPropertiesInAssemblyFormat` field is deprecated. Setting it to
+`0` temporarily opts a dialect into legacy behavior, allowing inherent
+attributes to mix with discardable attributes in `attr-dict`. Dialects using
+this setting should migrate their formats to bind inherent attributes directly
+or use `prop-dict`.
+
+See the [declarative assembly format](Operations.md/#declarative-assembly-format)
documentation for the corresponding format requirements.
### Default Attribute/Type Parsers and Printers
diff --git a/mlir/docs/ReleaseNotes.md b/mlir/docs/ReleaseNotes.md
index 16b93c8909670..e24b2790d8255 100644
--- a/mlir/docs/ReleaseNotes.md
+++ b/mlir/docs/ReleaseNotes.md
@@ -8,6 +8,18 @@ specifically, it is a snapshot of the MLIR development at the time of the releas
[TOC]
+## LLVM 24
+
+### Strict Property Assembly Formats
+
+Property-backed operations now use strict declarative assembly formats by
+default. Setting the dialect option `useStrictPropertiesInAssemblyFormat = 0`
+is deprecated and temporarily retains the legacy behavior. Dialects using
+this setting must migrate their declarative assembly formats to bind every
+inherent attribute and property directly or include `prop-dict`. Under the
+default rules, `attr-dict` contains only discardable attributes. See
+[the guide](DefiningDialects/_index.md/#strict-property-assembly-formats).
+
## LLVM 21
### GPU/NVVM Changes
diff --git a/mlir/examples/toy/Ch4/include/toy/Ops.td b/mlir/examples/toy/Ch4/include/toy/Ops.td
index e43cca419d847..ba751e0a01fbe 100644
--- a/mlir/examples/toy/Ch4/include/toy/Ops.td
+++ b/mlir/examples/toy/Ch4/include/toy/Ops.td
@@ -24,6 +24,7 @@ include "toy/ShapeInferenceInterface.td"
// can define our operations.
def Toy_Dialect : Dialect {
let name = "toy";
+ let useStrictPropertiesInAssemblyFormat = 0;
let cppNamespace = "::mlir::toy";
}
diff --git a/mlir/examples/toy/Ch5/include/toy/Ops.td b/mlir/examples/toy/Ch5/include/toy/Ops.td
index 3aecebd98bd46..effd58aadc136 100644
--- a/mlir/examples/toy/Ch5/include/toy/Ops.td
+++ b/mlir/examples/toy/Ch5/include/toy/Ops.td
@@ -24,6 +24,7 @@ include "toy/ShapeInferenceInterface.td"
// can define our operations.
def Toy_Dialect : Dialect {
let name = "toy";
+ let useStrictPropertiesInAssemblyFormat = 0;
let cppNamespace = "::mlir::toy";
}
diff --git a/mlir/examples/toy/Ch6/include/toy/Ops.td b/mlir/examples/toy/Ch6/include/toy/Ops.td
index b7818aafb7253..ca06b55e43c97 100644
--- a/mlir/examples/toy/Ch6/include/toy/Ops.td
+++ b/mlir/examples/toy/Ch6/include/toy/Ops.td
@@ -24,6 +24,7 @@ include "toy/ShapeInferenceInterface.td"
// can define our operations.
def Toy_Dialect : Dialect {
let name = "toy";
+ let useStrictPropertiesInAssemblyFormat = 0;
let cppNamespace = "::mlir::toy";
}
diff --git a/mlir/examples/toy/Ch7/include/toy/Ops.td b/mlir/examples/toy/Ch7/include/toy/Ops.td
index 3f6c77dd79726..1dbbf7eda284a 100644
--- a/mlir/examples/toy/Ch7/include/toy/Ops.td
+++ b/mlir/examples/toy/Ch7/include/toy/Ops.td
@@ -24,6 +24,7 @@ include "toy/ShapeInferenceInterface.td"
// can define our operations.
def Toy_Dialect : Dialect {
let name = "toy";
+ let useStrictPropertiesInAssemblyFormat = 0;
let cppNamespace = "::mlir::toy";
// We set this bit to generate a declaration of the `materializeConstant`
diff --git a/mlir/include/mlir/Dialect/AMDGPU/IR/AMDGPUBase.td b/mlir/include/mlir/Dialect/AMDGPU/IR/AMDGPUBase.td
index 8dc4b3a81c334..639dbf6b4a128 100644
--- a/mlir/include/mlir/Dialect/AMDGPU/IR/AMDGPUBase.td
+++ b/mlir/include/mlir/Dialect/AMDGPU/IR/AMDGPUBase.td
@@ -14,7 +14,6 @@ include "mlir/IR/DialectBase.td"
def AMDGPU_Dialect : Dialect {
let name = "amdgpu";
let cppNamespace = "::mlir::amdgpu";
- let useStrictPropertiesInAssemblyFormat = 1;
let description = [{
The `AMDGPU` dialect provides wrappers around AMD-specific functionality
and LLVM intrinsics. These wrappers should be used in conjunction with
diff --git a/mlir/include/mlir/Dialect/Affine/IR/AffineOps.td b/mlir/include/mlir/Dialect/Affine/IR/AffineOps.td
index d37fc62e3ed51..6215d5c292630 100644
--- a/mlir/include/mlir/Dialect/Affine/IR/AffineOps.td
+++ b/mlir/include/mlir/Dialect/Affine/IR/AffineOps.td
@@ -26,7 +26,6 @@ include "mlir/Interfaces/SideEffectInterfaces.td"
def Affine_Dialect : Dialect {
let name = "affine";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::affine";
let hasConstantMaterializer = 1;
let dependentDialects = ["arith::ArithDialect", "ub::UBDialect"];
diff --git a/mlir/include/mlir/Dialect/Arith/IR/ArithBase.td b/mlir/include/mlir/Dialect/Arith/IR/ArithBase.td
index 71e198458d342..0c05c4db79bed 100644
--- a/mlir/include/mlir/Dialect/Arith/IR/ArithBase.td
+++ b/mlir/include/mlir/Dialect/Arith/IR/ArithBase.td
@@ -14,7 +14,6 @@ include "mlir/IR/OpBase.td"
def Arith_Dialect : Dialect {
let name = "arith";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::arith";
let description = [{
The arith dialect is intended to hold basic integer and floating point
diff --git a/mlir/include/mlir/Dialect/ArmNeon/ArmNeon.td b/mlir/include/mlir/Dialect/ArmNeon/ArmNeon.td
index fd0fa7ecf9b0d..ce86ff2cfd922 100644
--- a/mlir/include/mlir/Dialect/ArmNeon/ArmNeon.td
+++ b/mlir/include/mlir/Dialect/ArmNeon/ArmNeon.td
@@ -23,7 +23,6 @@ include "mlir/IR/OpBase.td"
def ArmNeon_Dialect : Dialect {
let name = "arm_neon";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::arm_neon";
// Note: this does not need to depend on LLVMDialect as long as functions in
diff --git a/mlir/include/mlir/Dialect/ArmSME/IR/ArmSME.td b/mlir/include/mlir/Dialect/ArmSME/IR/ArmSME.td
index f937af9c35a71..ffafb2569310e 100644
--- a/mlir/include/mlir/Dialect/ArmSME/IR/ArmSME.td
+++ b/mlir/include/mlir/Dialect/ArmSME/IR/ArmSME.td
@@ -23,7 +23,6 @@ include "mlir/Dialect/LLVMIR/LLVMOpBase.td"
def ArmSME_Dialect : Dialect {
let name = "arm_sme";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::arm_sme";
let summary = "Basic dialect to target Arm SME architectures";
let description = [{
diff --git a/mlir/include/mlir/Dialect/ArmSVE/IR/ArmSVE.td b/mlir/include/mlir/Dialect/ArmSVE/IR/ArmSVE.td
index be4d9b123ac82..2f23404991799 100644
--- a/mlir/include/mlir/Dialect/ArmSVE/IR/ArmSVE.td
+++ b/mlir/include/mlir/Dialect/ArmSVE/IR/ArmSVE.td
@@ -22,7 +22,6 @@ include "mlir/Dialect/LLVMIR/LLVMOpBase.td"
def ArmSVE_Dialect : Dialect {
let name = "arm_sve";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::arm_sve";
let summary = "Basic dialect to target Arm SVE architectures";
let description = [{
diff --git a/mlir/include/mlir/Dialect/Async/IR/AsyncDialect.td b/mlir/include/mlir/Dialect/Async/IR/AsyncDialect.td
index f2c328a61e4cf..eb1d76a180fe2 100644
--- a/mlir/include/mlir/Dialect/Async/IR/AsyncDialect.td
+++ b/mlir/include/mlir/Dialect/Async/IR/AsyncDialect.td
@@ -21,7 +21,6 @@ include "mlir/IR/OpBase.td"
def AsyncDialect : Dialect {
let name = "async";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::async";
let summary = "Types and operations for async dialect";
diff --git a/mlir/include/mlir/Dialect/Bufferization/IR/BufferizationBase.td b/mlir/include/mlir/Dialect/Bufferization/IR/BufferizationBase.td
index 2ef0c93be8bf2..ac19f878656b5 100644
--- a/mlir/include/mlir/Dialect/Bufferization/IR/BufferizationBase.td
+++ b/mlir/include/mlir/Dialect/Bufferization/IR/BufferizationBase.td
@@ -31,7 +31,6 @@ def Bufferization_Dialect : Dialect {
"affine::AffineDialect", "memref::MemRefDialect", "tensor::TensorDialect",
"arith::ArithDialect"
];
- let useStrictPropertiesInAssemblyFormat = 1;
let extraClassDeclaration = [{
/// Verify an attribute from this dialect on the argument at 'argIndex' for
diff --git a/mlir/include/mlir/Dialect/Complex/IR/ComplexBase.td b/mlir/include/mlir/Dialect/Complex/IR/ComplexBase.td
index 4efe1bcc620c2..c8af498f44829 100644
--- a/mlir/include/mlir/Dialect/Complex/IR/ComplexBase.td
+++ b/mlir/include/mlir/Dialect/Complex/IR/ComplexBase.td
@@ -14,7 +14,6 @@ include "mlir/IR/OpBase.td"
def Complex_Dialect : Dialect {
let name = "complex";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::complex";
let description = [{
The complex dialect is intended to hold complex numbers creation and
diff --git a/mlir/include/mlir/Dialect/ControlFlow/IR/ControlFlowOps.td b/mlir/include/mlir/Dialect/ControlFlow/IR/ControlFlowOps.td
index 0e4c4eb78b94b..a441fd82546e3 100644
--- a/mlir/include/mlir/Dialect/ControlFlow/IR/ControlFlowOps.td
+++ b/mlir/include/mlir/Dialect/ControlFlow/IR/ControlFlowOps.td
@@ -23,7 +23,6 @@ def ControlFlow_Dialect : Dialect {
let name = "cf";
let cppNamespace = "::mlir::cf";
let dependentDialects = ["arith::ArithDialect"];
- let useStrictPropertiesInAssemblyFormat = 1;
let description = [{
This dialect contains low-level, i.e. non-region based, control flow
constructs. These constructs generally represent control flow directly
diff --git a/mlir/include/mlir/Dialect/DLTI/DLTIBase.td b/mlir/include/mlir/Dialect/DLTI/DLTIBase.td
index f7c9be4fd7880..3754f3699c7fd 100644
--- a/mlir/include/mlir/Dialect/DLTI/DLTIBase.td
+++ b/mlir/include/mlir/Dialect/DLTI/DLTIBase.td
@@ -76,7 +76,6 @@ def DLTI_Dialect : Dialect {
}];
let useDefaultAttributePrinterParser = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
}
def HasDefaultDLTIDataLayout : NativeOpTrait<"HasDefaultDLTIDataLayout"> {
diff --git a/mlir/include/mlir/Dialect/EmitC/IR/EmitCBase.td b/mlir/include/mlir/Dialect/EmitC/IR/EmitCBase.td
index a7144df36045e..375dbcbce1d03 100644
--- a/mlir/include/mlir/Dialect/EmitC/IR/EmitCBase.td
+++ b/mlir/include/mlir/Dialect/EmitC/IR/EmitCBase.td
@@ -31,7 +31,6 @@ def EmitC_Dialect : Dialect {
let hasConstantMaterializer = 1;
let useDefaultTypePrinterParser = 1;
let useDefaultAttributePrinterParser = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
}
#endif // MLIR_DIALECT_EMITC_IR_EMITCBASE
diff --git a/mlir/include/mlir/Dialect/Func/IR/FuncOps.td b/mlir/include/mlir/Dialect/Func/IR/FuncOps.td
index d7a0d5fc4f277..a99147b380eb3 100644
--- a/mlir/include/mlir/Dialect/Func/IR/FuncOps.td
+++ b/mlir/include/mlir/Dialect/Func/IR/FuncOps.td
@@ -22,7 +22,6 @@ def Func_Dialect : Dialect {
let name = "func";
let cppNamespace = "::mlir::func";
let hasConstantMaterializer = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
}
// Base class for Func dialect ops.
diff --git a/mlir/include/mlir/Dialect/GPU/IR/GPUBase.td b/mlir/include/mlir/Dialect/GPU/IR/GPUBase.td
index 8f3cffcdf43fa..f75d3ae0bd02f 100644
--- a/mlir/include/mlir/Dialect/GPU/IR/GPUBase.td
+++ b/mlir/include/mlir/Dialect/GPU/IR/GPUBase.td
@@ -82,7 +82,6 @@ def GPU_Dialect : Dialect {
let dependentDialects = ["arith::ArithDialect"];
let useDefaultAttributePrinterParser = 1;
let useDefaultTypePrinterParser = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
}
//===----------------------------------------------------------------------===//
diff --git a/mlir/include/mlir/Dialect/IRDL/IR/IRDL.td b/mlir/include/mlir/Dialect/IRDL/IR/IRDL.td
index f554358ed373d..e822969fc575e 100644
--- a/mlir/include/mlir/Dialect/IRDL/IR/IRDL.td
+++ b/mlir/include/mlir/Dialect/IRDL/IR/IRDL.td
@@ -74,7 +74,6 @@ def IRDL_Dialect : Dialect {
let name = "irdl";
let cppNamespace = "::mlir::irdl";
- let useStrictPropertiesInAssemblyFormat = 1;
}
#endif // MLIR_DIALECT_IRDL_IR_IRDL
diff --git a/mlir/include/mlir/Dialect/Index/IR/IndexDialect.td b/mlir/include/mlir/Dialect/Index/IR/IndexDialect.td
index df6087818cc8f..be0fea79ee392 100644
--- a/mlir/include/mlir/Dialect/Index/IR/IndexDialect.td
+++ b/mlir/include/mlir/Dialect/Index/IR/IndexDialect.td
@@ -83,7 +83,6 @@ def IndexDialect : Dialect {
let hasConstantMaterializer = 1;
let useDefaultAttributePrinterParser = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
}
#endif // INDEX_DIALECT
diff --git a/mlir/include/mlir/Dialect/LLVMIR/LLVMDialect.td b/mlir/include/mlir/Dialect/LLVMIR/LLVMDialect.td
index f48cb5385590b..864fc0f647bd1 100644
--- a/mlir/include/mlir/Dialect/LLVMIR/LLVMDialect.td
+++ b/mlir/include/mlir/Dialect/LLVMIR/LLVMDialect.td
@@ -20,7 +20,6 @@ def LLVM_Dialect : Dialect {
let hasRegionArgAttrVerify = 1;
let hasRegionResultAttrVerify = 1;
let hasOperationAttrVerify = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
let discardableAttrs = (ins
/// Attribute encoding size and type of GPU workgroup attributions.
diff --git a/mlir/include/mlir/Dialect/LLVMIR/NVVMDialect.td b/mlir/include/mlir/Dialect/LLVMIR/NVVMDialect.td
index c0f2ee062987e..025e093ebd8b6 100644
--- a/mlir/include/mlir/Dialect/LLVMIR/NVVMDialect.td
+++ b/mlir/include/mlir/Dialect/LLVMIR/NVVMDialect.td
@@ -21,7 +21,6 @@ def NVVM_Dialect : Dialect {
let cppNamespace = "::mlir::NVVM";
let dependentDialects = ["LLVM::LLVMDialect"];
let hasOperationAttrVerify = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
let extraClassDeclaration = [{
/// Get the name of the attribute used to annotate external kernel
diff --git a/mlir/include/mlir/Dialect/LLVMIR/ROCDLDialect.td b/mlir/include/mlir/Dialect/LLVMIR/ROCDLDialect.td
index a05913bb2c15f..5e6cefd977f32 100644
--- a/mlir/include/mlir/Dialect/LLVMIR/ROCDLDialect.td
+++ b/mlir/include/mlir/Dialect/LLVMIR/ROCDLDialect.td
@@ -18,7 +18,6 @@ include "mlir/Dialect/LLVMIR/LLVMOpBase.td"
def ROCDL_Dialect : Dialect {
let name = "rocdl";
let cppNamespace = "::mlir::ROCDL";
- let useStrictPropertiesInAssemblyFormat = 1;
let dependentDialects = ["LLVM::LLVMDialect"];
let summary = "Dialect for wrapping LLVM AMDGPU backend intrinsics and attributes";
let hasOperationAttrVerify = 1;
diff --git a/mlir/include/mlir/Dialect/LLVMIR/VCIXOps.td b/mlir/include/mlir/Dialect/LLVMIR/VCIXOps.td
index 02f80bbd56451..27d9a32dd8e03 100644
--- a/mlir/include/mlir/Dialect/LLVMIR/VCIXOps.td
+++ b/mlir/include/mlir/Dialect/LLVMIR/VCIXOps.td
@@ -31,7 +31,6 @@ def VCIX_Dialect : Dialect {
let name = "vcix";
let cppNamespace = "::mlir::vcix";
let dependentDialects = ["LLVM::LLVMDialect"];
- let useStrictPropertiesInAssemblyFormat = 1;
let description = [{
The SiFive Vector Coprocessor Interface (VCIX) provides a flexible mechanism
to extend application processors with custom coprocessors and
diff --git a/mlir/include/mlir/Dialect/LLVMIR/XeVMOps.td b/mlir/include/mlir/Dialect/LLVMIR/XeVMOps.td
index fdf6263094bb2..995bb758924c2 100644
--- a/mlir/include/mlir/Dialect/LLVMIR/XeVMOps.td
+++ b/mlir/include/mlir/Dialect/LLVMIR/XeVMOps.td
@@ -37,7 +37,6 @@ def XeVM_Dialect : Dialect {
}];
let useDefaultAttributePrinterParser = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
}
class XeVM_Attr<string attrName, string attrMnemonic, list<Trait> traits = []>
diff --git a/mlir/include/mlir/Dialect/Linalg/IR/LinalgBase.td b/mlir/include/mlir/Dialect/Linalg/IR/LinalgBase.td
index 5d53dbc461317..85857d3c11206 100644
--- a/mlir/include/mlir/Dialect/Linalg/IR/LinalgBase.td
+++ b/mlir/include/mlir/Dialect/Linalg/IR/LinalgBase.td
@@ -43,7 +43,6 @@ def Linalg_Dialect : Dialect {
"tensor::TensorDialect",
];
let useDefaultAttributePrinterParser = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
let hasCanonicalizer = 1;
let hasOperationAttrVerify = 1;
let hasConstantMaterializer = 1;
diff --git a/mlir/include/mlir/Dialect/MLProgram/IR/MLProgramBase.td b/mlir/include/mlir/Dialect/MLProgram/IR/MLProgramBase.td
index 5ed346aeade5d..a585059020eaf 100644
--- a/mlir/include/mlir/Dialect/MLProgram/IR/MLProgramBase.td
+++ b/mlir/include/mlir/Dialect/MLProgram/IR/MLProgramBase.td
@@ -13,7 +13,6 @@ include "mlir/IR/OpBase.td"
def MLProgram_Dialect : Dialect {
let name = "ml_program";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::ml_program";
let description = [{
The MLProgram dialect contains structural operations and types for
diff --git a/mlir/include/mlir/Dialect/MPI/IR/MPI.td b/mlir/include/mlir/Dialect/MPI/IR/MPI.td
index 29fd14b3f35f0..ba422273d5354 100644
--- a/mlir/include/mlir/Dialect/MPI/IR/MPI.td
+++ b/mlir/include/mlir/Dialect/MPI/IR/MPI.td
@@ -15,7 +15,6 @@ include "mlir/IR/EnumAttr.td"
def MPI_Dialect : Dialect {
let name = "mpi";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::mpi";
let description = [{
This dialect models the Message Passing Interface (MPI), version
diff --git a/mlir/include/mlir/Dialect/Math/IR/MathBase.td b/mlir/include/mlir/Dialect/Math/IR/MathBase.td
index 162284a1d8363..19fb39d9fd51d 100644
--- a/mlir/include/mlir/Dialect/Math/IR/MathBase.td
+++ b/mlir/include/mlir/Dialect/Math/IR/MathBase.td
@@ -10,7 +10,6 @@
include "mlir/IR/OpBase.td"
def Math_Dialect : Dialect {
let name = "math";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::math";
let description = [{
The math dialect is intended to hold mathematical operations on integer and
diff --git a/mlir/include/mlir/Dialect/MemRef/IR/MemRefBase.td b/mlir/include/mlir/Dialect/MemRef/IR/MemRefBase.td
index 82bd30fedf364..20dd45272898d 100644
--- a/mlir/include/mlir/Dialect/MemRef/IR/MemRefBase.td
+++ b/mlir/include/mlir/Dialect/MemRef/IR/MemRefBase.td
@@ -14,7 +14,6 @@ include "mlir/IR/OpBase.td"
def MemRef_Dialect : Dialect {
let name = "memref";
let cppNamespace = "::mlir::memref";
- let useStrictPropertiesInAssemblyFormat = 1;
let description = [{
The `memref` dialect is intended to hold core memref creation and
manipulation ops, which are not strongly associated with any particular
diff --git a/mlir/include/mlir/Dialect/NVGPU/IR/NVGPU.td b/mlir/include/mlir/Dialect/NVGPU/IR/NVGPU.td
index d43a8c38fd1e0..1c0d7bd1113ea 100644
--- a/mlir/include/mlir/Dialect/NVGPU/IR/NVGPU.td
+++ b/mlir/include/mlir/Dialect/NVGPU/IR/NVGPU.td
@@ -30,7 +30,6 @@ def NVGPU_Dialect : Dialect {
let useDefaultTypePrinterParser = 1;
let useDefaultAttributePrinterParser = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
let extraClassDeclaration = [{
/// Return true if the given MemRefType has an integer address
diff --git a/mlir/include/mlir/Dialect/OpenACC/OpenACCBase.td b/mlir/include/mlir/Dialect/OpenACC/OpenACCBase.td
index 675574547cebc..5810759c54298 100644
--- a/mlir/include/mlir/Dialect/OpenACC/OpenACCBase.td
+++ b/mlir/include/mlir/Dialect/OpenACC/OpenACCBase.td
@@ -19,7 +19,6 @@ include "mlir/IR/AttrTypeBase.td"
def OpenACC_Dialect : Dialect {
let name = "acc";
let useDefaultAttributePrinterParser = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
let useDefaultTypePrinterParser = 1;
let cppNamespace = "::mlir::acc";
let dependentDialects = ["::mlir::memref::MemRefDialect",
diff --git a/mlir/include/mlir/Dialect/OpenMP/OpenMPDialect.td b/mlir/include/mlir/Dialect/OpenMP/OpenMPDialect.td
index dbc851a5e297c..2dfee7120e82a 100644
--- a/mlir/include/mlir/Dialect/OpenMP/OpenMPDialect.td
+++ b/mlir/include/mlir/Dialect/OpenMP/OpenMPDialect.td
@@ -16,7 +16,6 @@ def OpenMP_Dialect : Dialect {
let cppNamespace = "::mlir::omp";
let dependentDialects = ["::mlir::LLVM::LLVMDialect, ::mlir::func::FuncDialect"];
let useDefaultAttributePrinterParser = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
let useDefaultTypePrinterParser = 1;
let hasOperationAttrVerify = 1;
}
diff --git a/mlir/include/mlir/Dialect/PDL/IR/PDLDialect.td b/mlir/include/mlir/Dialect/PDL/IR/PDLDialect.td
index b98940560ef52..d405bec26634c 100644
--- a/mlir/include/mlir/Dialect/PDL/IR/PDLDialect.td
+++ b/mlir/include/mlir/Dialect/PDL/IR/PDLDialect.td
@@ -63,7 +63,6 @@ def PDL_Dialect : Dialect {
}];
let name = "pdl";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::pdl";
let useDefaultTypePrinterParser = 1;
diff --git a/mlir/include/mlir/Dialect/PDLInterp/IR/PDLInterpOps.td b/mlir/include/mlir/Dialect/PDLInterp/IR/PDLInterpOps.td
index 752ba8e3c4e97..d60cd326a7956 100644
--- a/mlir/include/mlir/Dialect/PDLInterp/IR/PDLInterpOps.td
+++ b/mlir/include/mlir/Dialect/PDLInterp/IR/PDLInterpOps.td
@@ -36,7 +36,6 @@ def PDLInterp_Dialect : Dialect {
}];
let name = "pdl_interp";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::pdl_interp";
let dependentDialects = ["pdl::PDLDialect"];
let extraClassDeclaration = [{
diff --git a/mlir/include/mlir/Dialect/Ptr/IR/PtrDialect.td b/mlir/include/mlir/Dialect/Ptr/IR/PtrDialect.td
index bf1f1a3c89f5a..c98df5775195a 100644
--- a/mlir/include/mlir/Dialect/Ptr/IR/PtrDialect.td
+++ b/mlir/include/mlir/Dialect/Ptr/IR/PtrDialect.td
@@ -20,7 +20,6 @@ include "mlir/IR/OpBase.td"
def Ptr_Dialect : Dialect {
let name = "ptr";
- let useStrictPropertiesInAssemblyFormat = 1;
let summary = "Pointer dialect";
let description = [{
The pointer dialect provides types and operations for representing and
diff --git a/mlir/include/mlir/Dialect/Quant/IR/QuantBase.td b/mlir/include/mlir/Dialect/Quant/IR/QuantBase.td
index f11717f835f9b..b129e4b57e353 100644
--- a/mlir/include/mlir/Dialect/Quant/IR/QuantBase.td
+++ b/mlir/include/mlir/Dialect/Quant/IR/QuantBase.td
@@ -17,7 +17,6 @@ include "mlir/IR/OpBase.td"
def Quant_Dialect : Dialect {
let name = "quant";
- let useStrictPropertiesInAssemblyFormat = 1;
let description = [{
The `quant` dialect offers a framework for defining and manipulating
quantized values. Central to this framework is the `!quant.uniform` data
diff --git a/mlir/include/mlir/Dialect/SCF/IR/SCFOps.td b/mlir/include/mlir/Dialect/SCF/IR/SCFOps.td
index b0a34989b6433..31330be70ee2d 100644
--- a/mlir/include/mlir/Dialect/SCF/IR/SCFOps.td
+++ b/mlir/include/mlir/Dialect/SCF/IR/SCFOps.td
@@ -27,7 +27,6 @@ include "mlir/Interfaces/ViewLikeInterface.td"
def SCF_Dialect : Dialect {
let name = "scf";
let cppNamespace = "::mlir::scf";
- let useStrictPropertiesInAssemblyFormat = 1;
let description = [{
The `scf` (structured control flow) dialect contains operations that
diff --git a/mlir/include/mlir/Dialect/SMT/IR/SMTDialect.td b/mlir/include/mlir/Dialect/SMT/IR/SMTDialect.td
index 4b33b07da30c1..00f170659946e 100644
--- a/mlir/include/mlir/Dialect/SMT/IR/SMTDialect.td
+++ b/mlir/include/mlir/Dialect/SMT/IR/SMTDialect.td
@@ -13,7 +13,6 @@ include "mlir/IR/DialectBase.td"
def SMTDialect : Dialect {
let name = "smt";
- let useStrictPropertiesInAssemblyFormat = 1;
let summary = "a dialect that models satisfiability modulo theories";
let cppNamespace = "mlir::smt";
diff --git a/mlir/include/mlir/Dialect/SPIRV/IR/SPIRVBase.td b/mlir/include/mlir/Dialect/SPIRV/IR/SPIRVBase.td
index 47dfd6ece4f5a..a8073d7c848d6 100644
--- a/mlir/include/mlir/Dialect/SPIRV/IR/SPIRVBase.td
+++ b/mlir/include/mlir/Dialect/SPIRV/IR/SPIRVBase.td
@@ -52,7 +52,6 @@ def SPIRV_Dialect : Dialect {
let hasOperationAttrVerify = 1;
let hasRegionArgAttrVerify = 1;
let hasRegionResultAttrVerify = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
let extraClassDeclaration = [{
void registerAttributes();
diff --git a/mlir/include/mlir/Dialect/Shape/IR/ShapeBase.td b/mlir/include/mlir/Dialect/Shape/IR/ShapeBase.td
index d03fb094312a2..9c0257954d3e8 100644
--- a/mlir/include/mlir/Dialect/Shape/IR/ShapeBase.td
+++ b/mlir/include/mlir/Dialect/Shape/IR/ShapeBase.td
@@ -22,7 +22,6 @@ include "mlir/IR/OpBase.td"
def ShapeDialect : Dialect {
let name = "shape";
- let useStrictPropertiesInAssemblyFormat = 1;
let summary = "Types and operations for shape dialect";
let description = [{
diff --git a/mlir/include/mlir/Dialect/Shard/IR/ShardBase.td b/mlir/include/mlir/Dialect/Shard/IR/ShardBase.td
index c398e83f924fe..84c426252f4ab 100644
--- a/mlir/include/mlir/Dialect/Shard/IR/ShardBase.td
+++ b/mlir/include/mlir/Dialect/Shard/IR/ShardBase.td
@@ -21,7 +21,6 @@ include "mlir/IR/EnumAttr.td"
def Shard_Dialect : Dialect {
let name = "shard";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::shard";
let description = [{
diff --git a/mlir/include/mlir/Dialect/SparseTensor/IR/SparseTensorBase.td b/mlir/include/mlir/Dialect/SparseTensor/IR/SparseTensorBase.td
index e29358c6aa558..74e6783e260fa 100644
--- a/mlir/include/mlir/Dialect/SparseTensor/IR/SparseTensorBase.td
+++ b/mlir/include/mlir/Dialect/SparseTensor/IR/SparseTensorBase.td
@@ -91,7 +91,6 @@ def SparseTensor_Dialect : Dialect {
let useDefaultAttributePrinterParser = 1;
let useDefaultTypePrinterParser = 1;
let hasConstantMaterializer = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
}
#endif // SPARSETENSOR_BASE
diff --git a/mlir/include/mlir/Dialect/Tensor/IR/TensorBase.td b/mlir/include/mlir/Dialect/Tensor/IR/TensorBase.td
index 900ad5f40830c..9d0add92737f3 100644
--- a/mlir/include/mlir/Dialect/Tensor/IR/TensorBase.td
+++ b/mlir/include/mlir/Dialect/Tensor/IR/TensorBase.td
@@ -13,7 +13,6 @@ include "mlir/IR/OpBase.td"
def Tensor_Dialect : Dialect {
let name = "tensor";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::tensor";
let description = [{
diff --git a/mlir/include/mlir/Dialect/Tosa/IR/TosaOpBase.td b/mlir/include/mlir/Dialect/Tosa/IR/TosaOpBase.td
index 6a8056769316a..fa0a1dbd60c46 100644
--- a/mlir/include/mlir/Dialect/Tosa/IR/TosaOpBase.td
+++ b/mlir/include/mlir/Dialect/Tosa/IR/TosaOpBase.td
@@ -53,7 +53,6 @@ def Tosa_Dialect : Dialect {
let hasConstantMaterializer = 1;
let useDefaultAttributePrinterParser = 1;
let useDefaultTypePrinterParser = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
}
//===----------------------------------------------------------------------===//
diff --git a/mlir/include/mlir/Dialect/Transform/IR/TransformDialect.td b/mlir/include/mlir/Dialect/Transform/IR/TransformDialect.td
index d53761074db21..ce0ad30ad2c8c 100644
--- a/mlir/include/mlir/Dialect/Transform/IR/TransformDialect.td
+++ b/mlir/include/mlir/Dialect/Transform/IR/TransformDialect.td
@@ -19,7 +19,6 @@ def Transform_Dialect : Dialect {
let cppNamespace = "::mlir::transform";
let hasOperationAttrVerify = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
let extraClassDeclaration = [{
/// Symbol name for the default entry point "named sequence".
constexpr const static ::llvm::StringLiteral
diff --git a/mlir/include/mlir/Dialect/UB/IR/UBOps.td b/mlir/include/mlir/Dialect/UB/IR/UBOps.td
index 666301799256b..1bff39add691e 100644
--- a/mlir/include/mlir/Dialect/UB/IR/UBOps.td
+++ b/mlir/include/mlir/Dialect/UB/IR/UBOps.td
@@ -20,7 +20,6 @@ def UB_Dialect : Dialect {
let hasConstantMaterializer = 1;
let useDefaultAttributePrinterParser = 1;
- let useStrictPropertiesInAssemblyFormat = 1;
}
// Base class for UB dialect attributes.
diff --git a/mlir/include/mlir/Dialect/Vector/IR/Vector.td b/mlir/include/mlir/Dialect/Vector/IR/Vector.td
index f5e76c168f335..5125ae7c13717 100644
--- a/mlir/include/mlir/Dialect/Vector/IR/Vector.td
+++ b/mlir/include/mlir/Dialect/Vector/IR/Vector.td
@@ -18,7 +18,6 @@ include "mlir/IR/OpBase.td"
def Vector_Dialect : Dialect {
let name = "vector";
let cppNamespace = "::mlir::vector";
- let useStrictPropertiesInAssemblyFormat = 1;
let useDefaultAttributePrinterParser = 1;
let hasConstantMaterializer = 1;
diff --git a/mlir/include/mlir/Dialect/WasmSSA/IR/WasmSSABase.td b/mlir/include/mlir/Dialect/WasmSSA/IR/WasmSSABase.td
index 0e93fdd060962..f2777a7b155ed 100644
--- a/mlir/include/mlir/Dialect/WasmSSA/IR/WasmSSABase.td
+++ b/mlir/include/mlir/Dialect/WasmSSA/IR/WasmSSABase.td
@@ -14,7 +14,6 @@ include "mlir/IR/OpBase.td"
def WasmSSA_Dialect : Dialect {
let name = "wasmssa";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::wasmssa";
let description = [{
The `wasmssa` dialect is intended to represent WebAssembly
diff --git a/mlir/include/mlir/Dialect/X86/X86.td b/mlir/include/mlir/Dialect/X86/X86.td
index 681887e894260..193cbdfc1424e 100644
--- a/mlir/include/mlir/Dialect/X86/X86.td
+++ b/mlir/include/mlir/Dialect/X86/X86.td
@@ -25,7 +25,6 @@ include "mlir/IR/BuiltinTypes.td"
def X86_Dialect : Dialect {
let name = "x86";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir::x86";
let useDefaultTypePrinterParser = 1;
diff --git a/mlir/include/mlir/Dialect/XeGPU/IR/XeGPUDialect.td b/mlir/include/mlir/Dialect/XeGPU/IR/XeGPUDialect.td
index 652cf18ade580..b1490c7742a26 100644
--- a/mlir/include/mlir/Dialect/XeGPU/IR/XeGPUDialect.td
+++ b/mlir/include/mlir/Dialect/XeGPU/IR/XeGPUDialect.td
@@ -36,7 +36,6 @@ def XeGPU_Dialect : Dialect {
let useDefaultTypePrinterParser = true;
let useDefaultAttributePrinterParser = true;
- let useStrictPropertiesInAssemblyFormat = 1;
let extraClassDeclaration = [{
/// Checks if the given memref type represents shared local memory (SLM).
diff --git a/mlir/include/mlir/IR/BuiltinDialect.td b/mlir/include/mlir/IR/BuiltinDialect.td
index 35ea65671a7f8..c770dd50f3622 100644
--- a/mlir/include/mlir/IR/BuiltinDialect.td
+++ b/mlir/include/mlir/IR/BuiltinDialect.td
@@ -20,7 +20,6 @@ def Builtin_Dialect : Dialect {
let summary =
"A dialect containing the builtin Attributes, Operations, and Types";
let name = "builtin";
- let useStrictPropertiesInAssemblyFormat = 1;
let cppNamespace = "::mlir";
let useDefaultAttributePrinterParser = 0;
let useDefaultTypePrinterParser = 0;
diff --git a/mlir/include/mlir/IR/DialectBase.td b/mlir/include/mlir/IR/DialectBase.td
index 3b41e841eb3ad..da669c6ada639 100644
--- a/mlir/include/mlir/IR/DialectBase.td
+++ b/mlir/include/mlir/IR/DialectBase.td
@@ -55,11 +55,12 @@ class Dialect {
// dialect declaration.
code extraClassDeclaration = "";
- // If this dialect should require declarative parsers for property-backed
- // operations to bind every inherent attribute and property directly in the
- // custom assembly format, or otherwise cover them with `prop-dict`. This
- // stricter mode is disabled by default for now.
- bit useStrictPropertiesInAssemblyFormat = 0;
+ // Require declarative parsers for property-backed operations to bind every
+ // inherent attribute and property directly in the custom assembly format,
+ // or otherwise cover them with `prop-dict`.
+ // Deprecated: setting this to 0 temporarily opts into legacy behavior that
+ // mixes inherent and discardable attributes in `attr-dict`.
+ bit useStrictPropertiesInAssemblyFormat = 1;
// If this dialect overrides the hook for materializing constants.
bit hasConstantMaterializer = 0;
diff --git a/mlir/test/lib/Dialect/Test/TestDialect.td b/mlir/test/lib/Dialect/Test/TestDialect.td
index 37a263f1d10b8..ed0bedaced126 100644
--- a/mlir/test/lib/Dialect/Test/TestDialect.td
+++ b/mlir/test/lib/Dialect/Test/TestDialect.td
@@ -13,6 +13,8 @@ include "mlir/IR/OpBase.td"
def Test_Dialect : Dialect {
let name = "test";
+ // Keep legacy assembly format coverage for test operations.
+ let useStrictPropertiesInAssemblyFormat = 0;
let cppNamespace = "::test";
let hasCanonicalizer = 1;
let hasConstantMaterializer = 1;
diff --git a/mlir/test/mlir-tblgen/op-format-invalid.td b/mlir/test/mlir-tblgen/op-format-invalid.td
index c13bbedce3785..dc0d1688cc8ec 100644
--- a/mlir/test/mlir-tblgen/op-format-invalid.td
+++ b/mlir/test/mlir-tblgen/op-format-invalid.td
@@ -8,10 +8,11 @@ include "mlir/Interfaces/InferTypeOpInterface.td"
def TestDialect : Dialect {
let name = "test";
+ // Exercise the legacy assembly format behavior.
+ let useStrictPropertiesInAssemblyFormat = 0;
}
def TestStrictPropertiesDialect : Dialect {
let name = "test_strict_properties";
- let useStrictPropertiesInAssemblyFormat = 1;
}
class TestFormat_Op<string fmt, list<Trait> traits = []>
: Op<TestDialect, "format_op", traits> {
diff --git a/mlir/test/mlir-tblgen/op-format.td b/mlir/test/mlir-tblgen/op-format.td
index 54add551c1f18..f98632886449d 100644
--- a/mlir/test/mlir-tblgen/op-format.td
+++ b/mlir/test/mlir-tblgen/op-format.td
@@ -6,10 +6,11 @@ include "mlir/IR/EnumAttr.td"
def TestDialect : Dialect {
let name = "test";
+ // Exercise the legacy assembly format behavior.
+ let useStrictPropertiesInAssemblyFormat = 0;
}
def TestStrictPropertiesDialect : Dialect {
let name = "test_strict_properties";
- let useStrictPropertiesInAssemblyFormat = 1;
}
class TestFormat_Op<string fmt, list<Trait> traits = []>
: Op<TestDialect, "format_op", traits> {
diff --git a/mlir/test/python/python_test_ops.td b/mlir/test/python/python_test_ops.td
index 96e951424ca51..3624506a23a6d 100644
--- a/mlir/test/python/python_test_ops.td
+++ b/mlir/test/python/python_test_ops.td
@@ -15,6 +15,7 @@ include "mlir/Interfaces/InferTypeOpInterface.td"
def Python_Test_Dialect : Dialect {
let name = "python_test";
+ let useStrictPropertiesInAssemblyFormat = 0;
let cppNamespace = "python_test";
let useDefaultTypePrinterParser = 1;
>From 74e74469555000038f41949ad96c272fe4db2e71 Mon Sep 17 00:00:00 2001
From: theRonShark <ron.lieberman at amd.com>
Date: Sat, 26 Sep 2026 08:10:17 -0400
Subject: [PATCH 06/53] Revert "[clang-repl] Implement
IncrementalHIPDeviceParser for HIP device compilation" (#226679)
Reverts llvm/llvm-project#218337 buildbot failures
---
clang/include/clang/CodeGen/CodeGenAction.h | 5 -
clang/include/clang/Interpreter/Interpreter.h | 4 +-
clang/lib/CodeGen/BackendConsumer.h | 8 -
clang/lib/CodeGen/CodeGenAction.cpp | 9 -
clang/lib/Interpreter/DeviceOffload.cpp | 217 +-----------------
clang/lib/Interpreter/DeviceOffload.h | 64 +-----
clang/lib/Interpreter/Interpreter.cpp | 7 +-
clang/unittests/Basic/CMakeLists.txt | 1 -
clang/unittests/Basic/TargetIDTest.cpp | 50 ----
9 files changed, 25 insertions(+), 340 deletions(-)
delete mode 100644 clang/unittests/Basic/TargetIDTest.cpp
diff --git a/clang/include/clang/CodeGen/CodeGenAction.h b/clang/include/clang/CodeGen/CodeGenAction.h
index 319cc8f2b14a1..84fa4549d5033 100644
--- a/clang/include/clang/CodeGen/CodeGenAction.h
+++ b/clang/include/clang/CodeGen/CodeGenAction.h
@@ -63,11 +63,6 @@ class CodeGenAction : public ASTFrontendAction {
CodeGenerator *getCodeGenerator() const;
- /// Reload the -mlink-builtin-bitcode modules into the backend consumer.
- /// LinkInModules() consumes them, so incremental compilation must reload them
- /// before each translation unit (e.g. to re-link HIP device libraries).
- void reloadLinkModules(CompilerInstance &CI);
-
BackendConsumer *BEConsumer = nullptr;
};
diff --git a/clang/include/clang/Interpreter/Interpreter.h b/clang/include/clang/Interpreter/Interpreter.h
index b45c61a199f35..4504b679504e0 100644
--- a/clang/include/clang/Interpreter/Interpreter.h
+++ b/clang/include/clang/Interpreter/Interpreter.h
@@ -44,7 +44,7 @@ class CompilerInstance;
class CXXRecordDecl;
class Decl;
class IncrementalParser;
-class IncrementalDeviceParser;
+class IncrementalCUDADeviceParser;
enum class OffloadType { CUDA, HIP };
@@ -131,7 +131,7 @@ class Interpreter {
std::unique_ptr<IncrementalExecutor> IncrExecutor;
// An optional parser for CUDA offloading
- std::unique_ptr<IncrementalDeviceParser> DeviceParser;
+ std::unique_ptr<IncrementalCUDADeviceParser> DeviceParser;
// An optional action for CUDA offloading
std::unique_ptr<IncrementalAction> DeviceAct;
diff --git a/clang/lib/CodeGen/BackendConsumer.h b/clang/lib/CodeGen/BackendConsumer.h
index d6d713844a195..708658d206baf 100644
--- a/clang/lib/CodeGen/BackendConsumer.h
+++ b/clang/lib/CodeGen/BackendConsumer.h
@@ -92,14 +92,6 @@ class BackendConsumer : public ASTConsumer {
// Links each entry in LinkModules into our module. Returns true on error.
bool LinkInModules(llvm::Module *M);
- /// Replace the set of modules to link in. LinkInModules() consumes the
- /// modules, so incremental compilation (clang-repl) must reload and reseed
- /// them before each translation unit; otherwise later inputs would miss the
- /// linked-in bitcode (e.g. HIP device libraries).
- void setLinkModules(SmallVector<LinkModule, 4> LMs) {
- LinkModules = std::move(LMs);
- }
-
/// Get the best possible source location to represent a diagnostic that
/// may have associated debug info.
const FullSourceLoc getBestLocationFromDebugLoc(
diff --git a/clang/lib/CodeGen/CodeGenAction.cpp b/clang/lib/CodeGen/CodeGenAction.cpp
index c6020dc04a9c2..21e58c4aea8c8 100644
--- a/clang/lib/CodeGen/CodeGenAction.cpp
+++ b/clang/lib/CodeGen/CodeGenAction.cpp
@@ -983,15 +983,6 @@ CodeGenerator *CodeGenAction::getCodeGenerator() const {
return BEConsumer->getCodeGenerator();
}
-void CodeGenAction::reloadLinkModules(CompilerInstance &CI) {
- if (!BEConsumer)
- return;
- SmallVector<LinkModule, 4> LMs;
- if (clang::loadLinkModules(CI, *VMContext, LMs))
- return;
- BEConsumer->setLinkModules(std::move(LMs));
-}
-
bool CodeGenAction::BeginSourceFileAction(CompilerInstance &CI) {
if (CI.getFrontendOpts().GenReducedBMI)
CI.getLangOpts().setCompilingModule(LangOptions::CMK_ModuleInterface);
diff --git a/clang/lib/Interpreter/DeviceOffload.cpp b/clang/lib/Interpreter/DeviceOffload.cpp
index 8f59b569600be..bf7653c518c30 100644
--- a/clang/lib/Interpreter/DeviceOffload.cpp
+++ b/clang/lib/Interpreter/DeviceOffload.cpp
@@ -6,228 +6,32 @@
//
//===----------------------------------------------------------------------===//
//
-// This file implements offloading to HIP and CUDA devices.
+// This file implements offloading to CUDA devices.
//
//===----------------------------------------------------------------------===//
#include "DeviceOffload.h"
-#include "IncrementalAction.h"
-#include "clang/Basic/TargetID.h"
#include "clang/Basic/TargetOptions.h"
-#include "clang/CodeGen/BackendUtil.h"
-#include "clang/CodeGen/CodeGenAction.h"
-#include "clang/Driver/OffloadBundler.h"
+#include "clang/CodeGen/ModuleBuilder.h"
#include "clang/Frontend/CompilerInstance.h"
-#include "clang/Frontend/FrontendAction.h"
#include "clang/Interpreter/PartialTranslationUnit.h"
-#include "llvm/ADT/StringExtras.h"
#include "llvm/IR/LegacyPassManager.h"
#include "llvm/IR/Module.h"
#include "llvm/MC/TargetRegistry.h"
-#include "llvm/Support/FileSystem.h"
-#include "llvm/Support/FileUtilities.h"
-#include "llvm/Support/MemoryBuffer.h"
-#include "llvm/Support/Path.h"
-#include "llvm/Support/Program.h"
#include "llvm/Target/TargetMachine.h"
-#include "llvm/TargetParser/AMDGPUTargetParser.h"
-#include "llvm/TargetParser/Host.h"
namespace clang {
-IncrementalDeviceParser::IncrementalDeviceParser(
- CompilerInstance &DeviceInstance, CompilerInstance &HostInstance,
- IncrementalAction *DeviceAct,
- llvm::IntrusiveRefCntPtr<llvm::vfs::InMemoryFileSystem> FS,
- llvm::Error &Err, std::list<PartialTranslationUnit> &PTUs)
- : IncrementalParser(DeviceInstance, DeviceAct, Err, PTUs),
- DeviceCI(DeviceInstance), VFS(FS),
- CodeGenOpts(HostInstance.getCodeGenOpts()),
- TargetOpts(DeviceInstance.getTargetOpts()) {}
-
-IncrementalDeviceParser::~IncrementalDeviceParser() {}
-
-IncrementalHIPDeviceParser::IncrementalHIPDeviceParser(
- CompilerInstance &DeviceInstance, CompilerInstance &HostInstance,
- IncrementalAction *DeviceAct,
- llvm::IntrusiveRefCntPtr<llvm::vfs::InMemoryFileSystem> FS,
- llvm::Error &Err, std::list<PartialTranslationUnit> &PTUs)
- : IncrementalDeviceParser(DeviceInstance, HostInstance, DeviceAct, FS, Err,
- PTUs) {
- if (Err)
- return;
- StringRef Arch = TargetOpts.CPU;
- if (!Arch.starts_with("gfx")) {
- Err = llvm::joinErrors(std::move(Err), llvm::make_error<llvm::StringError>(
- "Invalid HIP architecture",
- llvm::inconvertibleErrorCode()));
- return;
- }
-}
-
-llvm::Expected<TranslationUnitDecl *>
-IncrementalHIPDeviceParser::Parse(llvm::StringRef Input) {
- if (FrontendAction *WrappedAct = Act->getWrapped())
- if (WrappedAct->hasIRSupport())
- static_cast<CodeGenAction *>(WrappedAct)->reloadLinkModules(DeviceCI);
-
- return IncrementalParser::Parse(Input);
-}
-
-llvm::Expected<llvm::StringRef> IncrementalHIPDeviceParser::GenerateHSACO() {
- auto &PTU = PTUs.back();
-
- CodeGenOptions CodeGenOptsForObj = DeviceCI.getCodeGenOpts();
- CodeGenOptsForObj.DisableLLVMPasses = true;
-
- llvm::SmallVector<char, 0> Object;
- auto ObjOS = std::make_unique<llvm::raw_svector_ostream>(Object);
- clang::emitBackendOutput(DeviceCI, CodeGenOptsForObj, PTU.TheModule.get(),
- Backend_EmitObj, DeviceCI.getVirtualFileSystemPtr(),
- std::move(ObjOS));
-
- if (DeviceCI.getDiagnostics().hasErrorOccurred())
- return llvm::make_error<llvm::StringError>(
- "Backend code generation failed for HIP device code.",
- llvm::inconvertibleErrorCode());
-
- std::string Exe = llvm::sys::fs::getMainExecutable(nullptr, nullptr);
- llvm::StringRef ExeDir = llvm::sys::path::parent_path(Exe);
- llvm::ErrorOr<std::string> LLDPath =
- llvm::sys::findProgramByName("ld.lld", {ExeDir});
- if (!LLDPath)
- LLDPath = llvm::sys::findProgramByName("ld.lld");
- if (!LLDPath)
- return llvm::make_error<llvm::StringError>(
- "Could not find ld.lld next to the executable or on PATH.",
- llvm::inconvertibleErrorCode());
-
- int ObjFD = -1;
- llvm::SmallString<128> ObjFile;
- if (llvm::sys::fs::createTemporaryFile("kernel", "o", ObjFD, ObjFile))
- return llvm::make_error<llvm::StringError>(
- "Failed to create a temporary object file.",
- llvm::inconvertibleErrorCode());
- llvm::FileRemover ObjRemover(ObjFile);
- {
- llvm::raw_fd_ostream OS(ObjFD, /*shouldClose=*/true);
- OS << llvm::StringRef(Object.data(), Object.size());
- }
-
- llvm::SmallString<128> HsacoFile;
- if (llvm::sys::fs::createTemporaryFile("kernel", "hsaco", HsacoFile))
- return llvm::make_error<llvm::StringError>(
- "Failed to create a temporary code object file.",
- llvm::inconvertibleErrorCode());
- llvm::FileRemover HsacoRemover(HsacoFile);
-
- llvm::StringRef Args[] = {"ld.lld", "-shared", "--no-undefined",
- ObjFile, "-o", HsacoFile};
- if (llvm::sys::ExecuteAndWait(*LLDPath, Args) != 0)
- return llvm::make_error<llvm::StringError>("ld.lld invocation failed.",
- llvm::inconvertibleErrorCode());
-
- auto HsacoBuf = llvm::MemoryBuffer::getFile(HsacoFile, /*IsText=*/false);
- if (!HsacoBuf)
- return llvm::make_error<llvm::StringError>(
- "Failed to read the code object.", llvm::inconvertibleErrorCode());
-
- llvm::StringRef Buffer = (*HsacoBuf)->getBuffer();
- HSACOContent.assign(Buffer.begin(), Buffer.end());
- return llvm::StringRef(HSACOContent.data(), HSACOContent.size());
-}
-
-llvm::Error IncrementalHIPDeviceParser::GenerateOffloadBundle() {
- static constexpr unsigned CodeObjectAlign = 4096;
-
- const PartialTranslationUnit &PTU = PTUs.back();
-
- llvm::SmallString<128> HostFile;
- if (llvm::sys::fs::createTemporaryFile("hip-host", "", HostFile))
- return llvm::make_error<llvm::StringError>(
- "Failed to create a temporary host bundle input.",
- llvm::inconvertibleErrorCode());
- llvm::FileRemover HostRemover(HostFile);
-
- llvm::SmallString<128> DeviceFile;
- int DeviceFD = -1;
- if (llvm::sys::fs::createTemporaryFile("hip-device", "hsaco", DeviceFD,
- DeviceFile))
- return llvm::make_error<llvm::StringError>(
- "Failed to create a temporary code object file.",
- llvm::inconvertibleErrorCode());
- llvm::FileRemover DeviceRemover(DeviceFile);
- {
- llvm::raw_fd_ostream OS(DeviceFD, /*shouldClose=*/true);
- OS << llvm::StringRef(HSACOContent.data(), HSACOContent.size());
- }
-
- llvm::SmallString<128> BundleFile;
- if (llvm::sys::fs::createTemporaryFile("hip-bundle", "hipfb", BundleFile))
- return llvm::make_error<llvm::StringError>(
- "Failed to create a temporary offload bundle file.",
- llvm::inconvertibleErrorCode());
- llvm::FileRemover BundleRemover(BundleFile);
-
- std::string TargetID = llvm::AMDGPU::TargetID::createFromSubtargetFeatures(
- DeviceCI.getTarget().getTriple(), TargetOpts.CPU,
- llvm::join(TargetOpts.Features, ","))
- .getCanonicalFeatureString();
-
- llvm::StringRef OffloadKind =
- (TargetOpts.CodeObjectVersion == llvm::CodeObjectVersionKind::COV_2 ||
- TargetOpts.CodeObjectVersion == llvm::CodeObjectVersionKind::COV_3)
- ? "hip"
- : "hipv4";
-
- std::string HostTriple = "host-" + llvm::sys::getProcessTriple() + "-";
- std::string DeviceTriple =
- OffloadKind.str() + "-" +
- normalizeForBundler(PTU.TheModule->getTargetTriple(), TargetID) + "-" +
- TargetID;
-
- OffloadBundlerConfig Config;
- Config.FilesType = "o";
- Config.BundleAlignment = CodeObjectAlign;
- Config.HostInputIndex = 0;
- Config.TargetNames = {HostTriple, DeviceTriple};
- Config.InputFileNames = {std::string(HostFile), std::string(DeviceFile)};
- Config.OutputFileNames = {std::string(BundleFile)};
-
- if (llvm::Error Err = OffloadBundler(Config).BundleFiles())
- return Err;
-
- auto BundleBuf = llvm::MemoryBuffer::getFile(BundleFile, /*IsText=*/false);
- if (!BundleBuf)
- return llvm::make_error<llvm::StringError>(
- "Failed to read the offload bundle.", llvm::inconvertibleErrorCode());
-
- std::string BundleFileName = "/" + PTU.TheModule->getName().str() + ".hipfb";
- VFS->addFile(BundleFileName, 0,
- llvm::MemoryBuffer::getMemBufferCopy((*BundleBuf)->getBuffer()));
-
- CodeGenOpts.OffloadBinaryToEmbedFile = std::move(BundleFileName);
- return llvm::Error::success();
-}
-
-llvm::Error IncrementalHIPDeviceParser::GenerateOffloadBinary() {
- llvm::Expected<llvm::StringRef> HSACO = GenerateHSACO();
- if (!HSACO)
- return HSACO.takeError();
- return GenerateOffloadBundle();
-}
-
-IncrementalHIPDeviceParser::~IncrementalHIPDeviceParser() {}
-
IncrementalCUDADeviceParser::IncrementalCUDADeviceParser(
CompilerInstance &DeviceInstance, CompilerInstance &HostInstance,
IncrementalAction *DeviceAct,
llvm::IntrusiveRefCntPtr<llvm::vfs::InMemoryFileSystem> FS,
llvm::Error &Err, std::list<PartialTranslationUnit> &PTUs)
- : IncrementalDeviceParser(DeviceInstance, HostInstance, DeviceAct, FS, Err,
- PTUs) {
+ : IncrementalParser(DeviceInstance, DeviceAct, Err, PTUs), VFS(FS),
+ CodeGenOpts(HostInstance.getCodeGenOpts()),
+ TargetOpts(DeviceInstance.getTargetOpts()) {
if (Err)
return;
StringRef Arch = TargetOpts.CPU;
@@ -264,7 +68,9 @@ llvm::Expected<llvm::StringRef> IncrementalCUDADeviceParser::GeneratePTX() {
llvm::inconvertibleErrorCode());
}
- PM.run(*PTU.TheModule);
+ if (!PM.run(*PTU.TheModule))
+ return llvm::make_error<llvm::StringError>("Failed to emit PTX code.",
+ llvm::inconvertibleErrorCode());
PTXCode += '\0';
while (PTXCode.size() % 8)
@@ -352,13 +158,6 @@ llvm::Error IncrementalCUDADeviceParser::GenerateFatbinary() {
return llvm::Error::success();
}
-llvm::Error IncrementalCUDADeviceParser::GenerateOffloadBinary() {
- llvm::Expected<llvm::StringRef> PTX = GeneratePTX();
- if (!PTX)
- return PTX.takeError();
- return GenerateFatbinary();
-}
-
IncrementalCUDADeviceParser::~IncrementalCUDADeviceParser() {}
} // namespace clang
diff --git a/clang/lib/Interpreter/DeviceOffload.h b/clang/lib/Interpreter/DeviceOffload.h
index 45a3fc17c065e..a31bd5a0499b8 100644
--- a/clang/lib/Interpreter/DeviceOffload.h
+++ b/clang/lib/Interpreter/DeviceOffload.h
@@ -6,7 +6,7 @@
//
//===----------------------------------------------------------------------===//
//
-// This file implements classes required for offloading to HIP and CUDA devices.
+// This file implements classes required for offloading to CUDA devices.
//
//===----------------------------------------------------------------------===//
@@ -14,11 +14,9 @@
#define LLVM_CLANG_LIB_INTERPRETER_DEVICE_OFFLOAD_H
#include "IncrementalParser.h"
-#include "llvm/Support/Error.h"
+#include "llvm/Support/FileSystem.h"
#include "llvm/Support/VirtualFileSystem.h"
-#include <memory>
-
namespace clang {
struct PartialTranslationUnit;
class CompilerInstance;
@@ -26,52 +24,7 @@ class CodeGenOptions;
class TargetOptions;
class IncrementalAction;
-class IncrementalDeviceParser : public IncrementalParser {
-
-public:
- IncrementalDeviceParser(
- CompilerInstance &DeviceInstance, CompilerInstance &HostInstance,
- IncrementalAction *DeviceAct,
- llvm::IntrusiveRefCntPtr<llvm::vfs::InMemoryFileSystem> VFS,
- llvm::Error &Err, std::list<PartialTranslationUnit> &PTUs);
-
- virtual llvm::Error GenerateOffloadBinary() = 0;
-
- ~IncrementalDeviceParser() override;
-
-protected:
- CompilerInstance &DeviceCI;
- llvm::IntrusiveRefCntPtr<llvm::vfs::InMemoryFileSystem> VFS;
- CodeGenOptions &CodeGenOpts;
- const TargetOptions &TargetOpts;
-};
-
-class IncrementalHIPDeviceParser : public IncrementalDeviceParser {
-
-public:
- IncrementalHIPDeviceParser(
- CompilerInstance &DeviceInstance, CompilerInstance &HostInstance,
- IncrementalAction *DeviceAct,
- llvm::IntrusiveRefCntPtr<llvm::vfs::InMemoryFileSystem> VFS,
- llvm::Error &Err, std::list<PartialTranslationUnit> &PTUs);
-
- llvm::Expected<TranslationUnitDecl *> Parse(llvm::StringRef Input) override;
-
- llvm::Error GenerateOffloadBinary() override;
-
- ~IncrementalHIPDeviceParser();
-
-protected:
- // Generate the HSACO code object for the last PTU.
- llvm::Expected<llvm::StringRef> GenerateHSACO();
-
- // Bundle the HSACO into a HIP offload bundle in memory.
- llvm::Error GenerateOffloadBundle();
-
- llvm::SmallVector<char, 1024> HSACOContent;
-};
-
-class IncrementalCUDADeviceParser : public IncrementalDeviceParser {
+class IncrementalCUDADeviceParser : public IncrementalParser {
public:
IncrementalCUDADeviceParser(
@@ -80,20 +33,21 @@ class IncrementalCUDADeviceParser : public IncrementalDeviceParser {
llvm::IntrusiveRefCntPtr<llvm::vfs::InMemoryFileSystem> VFS,
llvm::Error &Err, std::list<PartialTranslationUnit> &PTUs);
- llvm::Error GenerateOffloadBinary() override;
-
- ~IncrementalCUDADeviceParser();
-
-protected:
// Generate PTX for the last PTU.
llvm::Expected<llvm::StringRef> GeneratePTX();
// Generate fatbinary contents in memory
llvm::Error GenerateFatbinary();
+ ~IncrementalCUDADeviceParser();
+
+protected:
int SMVersion;
llvm::SmallString<1024> PTXCode;
llvm::SmallVector<char, 1024> FatbinContent;
+ llvm::IntrusiveRefCntPtr<llvm::vfs::InMemoryFileSystem> VFS;
+ CodeGenOptions &CodeGenOpts; // Intentionally a reference.
+ const TargetOptions &TargetOpts;
};
} // namespace clang
diff --git a/clang/lib/Interpreter/Interpreter.cpp b/clang/lib/Interpreter/Interpreter.cpp
index 8aafbc867ddaf..655db32477a0c 100644
--- a/clang/lib/Interpreter/Interpreter.cpp
+++ b/clang/lib/Interpreter/Interpreter.cpp
@@ -576,7 +576,12 @@ Interpreter::Parse(llvm::StringRef Code) {
DeviceParser->RegisterPTU(*DeviceTU);
- if (llvm::Error Err = DeviceParser->GenerateOffloadBinary())
+ llvm::Expected<llvm::StringRef> PTX = DeviceParser->GeneratePTX();
+ if (!PTX)
+ return PTX.takeError();
+
+ llvm::Error Err = DeviceParser->GenerateFatbinary();
+ if (Err)
return std::move(Err);
}
diff --git a/clang/unittests/Basic/CMakeLists.txt b/clang/unittests/Basic/CMakeLists.txt
index 3cd843736eb9b..32dce09866892 100644
--- a/clang/unittests/Basic/CMakeLists.txt
+++ b/clang/unittests/Basic/CMakeLists.txt
@@ -13,7 +13,6 @@ add_distinct_clang_unittest(BasicTests
SanitizersTest.cpp
SarifTest.cpp
SourceManagerTest.cpp
- TargetIDTest.cpp
CLANG_LIBS
clangBasic
clangLex
diff --git a/clang/unittests/Basic/TargetIDTest.cpp b/clang/unittests/Basic/TargetIDTest.cpp
deleted file mode 100644
index a4eaa06b074f8..0000000000000
--- a/clang/unittests/Basic/TargetIDTest.cpp
+++ /dev/null
@@ -1,50 +0,0 @@
-//===- unittests/Basic/TargetIDTest.cpp - Test TargetID -----------------===//
-//
-// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
-// See https://llvm.org/LICENSE.txt for license information.
-// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
-//
-//===----------------------------------------------------------------------===//
-
-#include "clang/Basic/TargetID.h"
-#include "llvm/TargetParser/AMDGPUTargetParser.h"
-#include "llvm/TargetParser/Triple.h"
-#include "gtest/gtest.h"
-
-namespace {
-
-static std::string bundleEntryID(const llvm::Triple &T, llvm::StringRef CPU,
- llvm::StringRef Features,
- llvm::StringRef OffloadKind) {
- std::string TargetID =
- llvm::AMDGPU::TargetID::createFromSubtargetFeatures(T, CPU, Features)
- .getCanonicalFeatureString();
- return OffloadKind.str() + "-" + clang::normalizeForBundler(T, TargetID) +
- "-" + TargetID;
-}
-
-TEST(TargetIDTest, HIPBundleEntryPreservesXnack) {
- llvm::Triple T("amdgcn-amd-amdhsa");
- EXPECT_EQ(bundleEntryID(T, "gfx90a", "+xnack,+wavefrontsize64", "hipv4"),
- "hipv4-amdgcn-amd-amdhsa--gfx90a:xnack+");
-}
-
-TEST(TargetIDTest, HIPBundleEntryPreservesDisabledFeature) {
- llvm::Triple T("amdgcn-amd-amdhsa");
- EXPECT_EQ(bundleEntryID(T, "gfx90a", "-xnack", "hipv4"),
- "hipv4-amdgcn-amd-amdhsa--gfx90a:xnack-");
-}
-
-TEST(TargetIDTest, HIPBundleEntryWithoutFeatures) {
- llvm::Triple T("amdgcn-amd-amdhsa");
- EXPECT_EQ(bundleEntryID(T, "gfx90a", "", "hipv4"),
- "hipv4-amdgcn-amd-amdhsa--gfx90a");
-}
-
-TEST(TargetIDTest, HIPBundleEntryLegacyKind) {
- llvm::Triple T("amdgcn-amd-amdhsa");
- EXPECT_EQ(bundleEntryID(T, "gfx908", "+sramecc", "hip"),
- "hip-amdgcn-amd-amdhsa--gfx908:sramecc+");
-}
-
-} // namespace
>From 569a30ffd0f7e66912f8f81e494db10ef5f1bb61 Mon Sep 17 00:00:00 2001
From: Alexey Bataev <a.bataev at outlook.com>
Date: Sat, 26 Sep 2026 08:27:16 -0400
Subject: [PATCH 07/53] [SLP][NFC]Add extra FMA candidates tests, NFC
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/226684
---
.../reduced-value-next-to-fused-fmul.ll | 42 +++++++++++++++
.../AMDGPU/fmul-sunk-into-fadd-block.ll | 52 +++++++++++++++++++
2 files changed, 94 insertions(+)
create mode 100644 llvm/test/Transforms/SLPVectorizer/AArch64/reduced-value-next-to-fused-fmul.ll
create mode 100644 llvm/test/Transforms/SLPVectorizer/AMDGPU/fmul-sunk-into-fadd-block.ll
diff --git a/llvm/test/Transforms/SLPVectorizer/AArch64/reduced-value-next-to-fused-fmul.ll b/llvm/test/Transforms/SLPVectorizer/AArch64/reduced-value-next-to-fused-fmul.ll
new file mode 100644
index 0000000000000..288364e64958b
--- /dev/null
+++ b/llvm/test/Transforms/SLPVectorizer/AArch64/reduced-value-next-to-fused-fmul.ll
@@ -0,0 +1,42 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
+; RUN: opt -passes=slp-vectorizer -mtriple=aarch64-unknown-linux -mcpu=neoverse-v2 -S < %s | FileCheck %s
+
+; The first reduction operation fuses with its fmul operand, not with the
+; reduced fdiv, so the cost of the fdiv is not excluded from the scalar
+; reduction as the cost of the fused fmul.
+define double @fdivs_next_to_fused_fmul(ptr %x, ptr %y, double %a0, double %b0) {
+; CHECK-LABEL: define double @fdivs_next_to_fused_fmul(
+; CHECK-SAME: ptr [[X:%.*]], ptr [[Y:%.*]], double [[A0:%.*]], double [[B0:%.*]]) #[[ATTR0:[0-9]+]] {
+; CHECK-NEXT: [[P0:%.*]] = fmul reassoc contract double [[A0]], [[B0]]
+; CHECK-NEXT: [[TMP1:%.*]] = load <4 x double>, ptr [[X]], align 8
+; CHECK-NEXT: [[TMP2:%.*]] = load <4 x double>, ptr [[Y]], align 8
+; CHECK-NEXT: [[TMP3:%.*]] = fdiv reassoc contract <4 x double> [[TMP1]], [[TMP2]]
+; CHECK-NEXT: [[TMP4:%.*]] = call reassoc contract double @llvm.vector.reduce.fadd.v4f64(double -0.000000e+00, <4 x double> [[TMP3]])
+; CHECK-NEXT: [[OP_RDX:%.*]] = fadd reassoc contract double [[TMP4]], [[P0]]
+; CHECK-NEXT: ret double [[OP_RDX]]
+;
+ %p0 = fmul reassoc contract double %a0, %b0
+ %x1 = load double, ptr %x
+ %y1 = load double, ptr %y
+ %q1 = fdiv reassoc contract double %x1, %y1
+ %x2p = getelementptr double, ptr %x, i64 1
+ %x2 = load double, ptr %x2p
+ %y2p = getelementptr double, ptr %y, i64 1
+ %y2 = load double, ptr %y2p
+ %q2 = fdiv reassoc contract double %x2, %y2
+ %x3p = getelementptr double, ptr %x, i64 2
+ %x3 = load double, ptr %x3p
+ %y3p = getelementptr double, ptr %y, i64 2
+ %y3 = load double, ptr %y3p
+ %q3 = fdiv reassoc contract double %x3, %y3
+ %x4p = getelementptr double, ptr %x, i64 3
+ %x4 = load double, ptr %x4p
+ %y4p = getelementptr double, ptr %y, i64 3
+ %y4 = load double, ptr %y4p
+ %q4 = fdiv reassoc contract double %x4, %y4
+ %r0 = fadd reassoc contract double %p0, %q1
+ %r1 = fadd reassoc contract double %r0, %q2
+ %r2 = fadd reassoc contract double %r1, %q3
+ %r3 = fadd reassoc contract double %r2, %q4
+ ret double %r3
+}
diff --git a/llvm/test/Transforms/SLPVectorizer/AMDGPU/fmul-sunk-into-fadd-block.ll b/llvm/test/Transforms/SLPVectorizer/AMDGPU/fmul-sunk-into-fadd-block.ll
new file mode 100644
index 0000000000000..63597fb54b840
--- /dev/null
+++ b/llvm/test/Transforms/SLPVectorizer/AMDGPU/fmul-sunk-into-fadd-block.ll
@@ -0,0 +1,52 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
+; RUN: opt -S -passes=slp-vectorizer -mtriple=amdgpu9.0a-amd-amdhsa < %s | FileCheck %s
+
+; The fmul operands of the fadds are computed in different blocks, so they
+; cannot form a vector node and are gathered. The target sinks each fmul into
+; the block of its fadd user, where they fuse into an fma in the scalar code.
+; The vectorized fadds keep the gathered fmuls unfused, so the fadds stay
+; scalar.
+define void @fmuls_sunk_into_fadd_block(ptr addrspace(1) %out, ptr addrspace(1) %c, float %a0, float %b0, float %a1, float %b1, i1 %cond) {
+; CHECK-LABEL: define void @fmuls_sunk_into_fadd_block(
+; CHECK-SAME: ptr addrspace(1) [[OUT:%.*]], ptr addrspace(1) [[C:%.*]], float [[A0:%.*]], float [[B0:%.*]], float [[A1:%.*]], float [[B1:%.*]], i1 [[COND:%.*]]) {
+; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: [[M0:%.*]] = fmul contract float [[A0]], [[B0]]
+; CHECK-NEXT: br i1 [[COND]], label %[[MUL:.*]], label %[[EXIT:.*]]
+; CHECK: [[MUL]]:
+; CHECK-NEXT: [[M1:%.*]] = fmul contract float [[A1]], [[B1]]
+; CHECK-NEXT: br label %[[ADD:.*]]
+; CHECK: [[ADD]]:
+; CHECK-NEXT: [[C0:%.*]] = load float, ptr addrspace(1) [[C]], align 8
+; CHECK-NEXT: [[C_1:%.*]] = getelementptr inbounds float, ptr addrspace(1) [[C]], i64 1
+; CHECK-NEXT: [[C1:%.*]] = load float, ptr addrspace(1) [[C_1]], align 4
+; CHECK-NEXT: [[S0:%.*]] = fadd contract float [[M0]], [[C0]]
+; CHECK-NEXT: [[S1:%.*]] = fadd contract float [[M1]], [[C1]]
+; CHECK-NEXT: store float [[S0]], ptr addrspace(1) [[OUT]], align 8
+; CHECK-NEXT: [[OUT_1:%.*]] = getelementptr inbounds float, ptr addrspace(1) [[OUT]], i64 1
+; CHECK-NEXT: store float [[S1]], ptr addrspace(1) [[OUT_1]], align 4
+; CHECK-NEXT: br label %[[EXIT]]
+; CHECK: [[EXIT]]:
+; CHECK-NEXT: ret void
+;
+entry:
+ %m0 = fmul contract float %a0, %b0
+ br i1 %cond, label %mul, label %exit
+
+mul:
+ %m1 = fmul contract float %a1, %b1
+ br label %add
+
+add:
+ %c0 = load float, ptr addrspace(1) %c, align 8
+ %c.1 = getelementptr inbounds float, ptr addrspace(1) %c, i64 1
+ %c1 = load float, ptr addrspace(1) %c.1, align 4
+ %s0 = fadd contract float %m0, %c0
+ %s1 = fadd contract float %m1, %c1
+ store float %s0, ptr addrspace(1) %out, align 8
+ %out.1 = getelementptr inbounds float, ptr addrspace(1) %out, i64 1
+ store float %s1, ptr addrspace(1) %out.1, align 4
+ br label %exit
+
+exit:
+ ret void
+}
>From 710ecd388e25591390179b28a8c02c552f68082d Mon Sep 17 00:00:00 2001
From: Alexey Bataev <a.bataev at outlook.com>
Date: Sat, 26 Sep 2026 08:48:23 -0400
Subject: [PATCH 08/53] [SLP]Do not reuse transformed nodes in gathers emitted
before their user
Gathers of the users with all scalars used outside the block are emitted
before the user, while the transformed nodes are emitted with the user.
Reusing such a transformed node postponed the gather and moved only the
last instruction of the transformed node buildvector, breaking dominance.
Fixes #226674
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/226687
---
.../Transforms/Vectorize/SLPVectorizer.cpp | 6 +-
...mmed-node-matching-gather-emitted-first.ll | 69 +++++++++++++++++++
2 files changed, 73 insertions(+), 2 deletions(-)
create mode 100644 llvm/test/Transforms/SLPVectorizer/X86/trimmed-node-matching-gather-emitted-first.ll
diff --git a/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp b/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
index 9a1b310fd327f..f76a754418422 100644
--- a/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+++ b/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
@@ -21027,11 +21027,13 @@ BoUpSLP::isGatherShuffledSingleRegisterEntry(
GatherNodes.push_back(E);
}
} else if (const TreeEntry *E = getSameValuesTreeEntry(V, TE->Scalars);
- E && TransformedToGatherNodes.contains(E) && E->UserTreeIndex &&
+ E && !TEUserNeedsEmitFirst &&
+ TransformedToGatherNodes.contains(E) && E->UserTreeIndex &&
E->UserTreeIndex.UserTE == TE->UserTreeIndex.UserTE &&
!E->UserTreeIndex.UserTE->isGather()) {
// Regular gathers reuse only perfectly matched transformed nodes of the
- // same user.
+ // same user. Gathers, emitted before their user, cannot reuse transformed
+ // nodes, which are emitted with the user.
GatherNodes.push_back(E);
}
for (const TreeEntry *TEPtr : GatherNodes) {
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/trimmed-node-matching-gather-emitted-first.ll b/llvm/test/Transforms/SLPVectorizer/X86/trimmed-node-matching-gather-emitted-first.ll
new file mode 100644
index 0000000000000..f86a3fc7abc7c
--- /dev/null
+++ b/llvm/test/Transforms/SLPVectorizer/X86/trimmed-node-matching-gather-emitted-first.ll
@@ -0,0 +1,69 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
+; RUN: opt -passes=slp-vectorizer -S -mtriple=x86_64-unknown-linux-gnu < %s | FileCheck %s
+
+define double @test(double %x, i1 %c) {
+; CHECK-LABEL: define double @test(
+; CHECK-SAME: double [[X:%.*]], i1 [[C:%.*]]) {
+; CHECK-NEXT: [[ENTRY:.*]]:
+; CHECK-NEXT: br i1 [[C]], label %[[BB:.*]], label %[[EXIT:.*]]
+; CHECK: [[BB]]:
+; CHECK-NEXT: [[ADD:%.*]] = fadd double [[X]], [[X]]
+; CHECK-NEXT: [[SUB:%.*]] = fsub double 0.000000e+00, [[X]]
+; CHECK-NEXT: [[TMP0:%.*]] = insertelement <2 x double> poison, double [[ADD]], i64 0
+; CHECK-NEXT: [[TMP1:%.*]] = insertelement <2 x double> [[TMP0]], double [[SUB]], i64 1
+; CHECK-NEXT: [[TMP2:%.*]] = call <2 x double> @llvm.fmuladd.v2f64(<2 x double> [[TMP1]], <2 x double> [[TMP1]], <2 x double> zeroinitializer)
+; CHECK-NEXT: br label %[[EXIT]]
+; CHECK: [[EXIT]]:
+; CHECK-NEXT: [[TMP3:%.*]] = phi <2 x double> [ [[TMP2]], %[[BB]] ], [ zeroinitializer, %[[ENTRY]] ]
+; CHECK-NEXT: [[A:%.*]] = fadd double [[X]], 0.000000e+00
+; CHECK-NEXT: [[TMP4:%.*]] = extractelement <2 x double> [[TMP3]], i64 1
+; CHECK-NEXT: [[M:%.*]] = fmul double [[A]], [[TMP4]]
+; CHECK-NEXT: [[TMP5:%.*]] = extractelement <2 x double> [[TMP3]], i64 0
+; CHECK-NEXT: [[R:%.*]] = call double @llvm.fmuladd.f64(double [[TMP5]], double 0.000000e+00, double [[M]])
+; CHECK-NEXT: ret double [[R]]
+;
+entry:
+ br i1 %c, label %bb, label %exit
+
+bb:
+ %add = fadd double %x, %x
+ %f0 = call double @llvm.fmuladd.f64(double %add, double %add, double 0.000000e+00)
+ %sub = fsub double 0.000000e+00, %x
+ %f1 = call double @llvm.fmuladd.f64(double %sub, double %sub, double 0.000000e+00)
+ br label %exit
+
+exit:
+ %p0 = phi double [ %f0, %bb ], [ 0.000000e+00, %entry ]
+ %p1 = phi double [ %f1, %bb ], [ 0.000000e+00, %entry ]
+ %a = fadd double %x, 0.000000e+00
+ %m = fmul double %a, %p1
+ %r = call double @llvm.fmuladd.f64(double %p0, double 0.000000e+00, double %m)
+ ret double %r
+}
+
+define void @test_loop(float %x) {
+; CHECK-LABEL: define void @test_loop(
+; CHECK-SAME: float [[X:%.*]]) {
+; CHECK-NEXT: [[ENTRY:.*]]:
+; CHECK-NEXT: br label %[[LOOP:.*]]
+; CHECK: [[LOOP]]:
+; CHECK-NEXT: [[TMP0:%.*]] = phi <2 x float> [ [[TMP3:%.*]], %[[LOOP]] ], [ zeroinitializer, %[[ENTRY]] ]
+; CHECK-NEXT: [[ADD:%.*]] = fadd float [[X]], [[X]]
+; CHECK-NEXT: [[SUB:%.*]] = fsub float 0.000000e+00, [[X]]
+; CHECK-NEXT: [[TMP1:%.*]] = insertelement <2 x float> poison, float [[ADD]], i64 0
+; CHECK-NEXT: [[TMP2:%.*]] = insertelement <2 x float> [[TMP1]], float [[SUB]], i64 1
+; CHECK-NEXT: [[TMP3]] = call <2 x float> @llvm.fmuladd.v2f32(<2 x float> [[TMP2]], <2 x float> [[TMP2]], <2 x float> [[TMP0]])
+; CHECK-NEXT: br label %[[LOOP]]
+;
+entry:
+ br label %loop
+
+loop:
+ %p0 = phi float [ %f0, %loop ], [ 0.000000e+00, %entry ]
+ %p1 = phi float [ %f1, %loop ], [ 0.000000e+00, %entry ]
+ %add = fadd float %x, %x
+ %f0 = call float @llvm.fmuladd.f32(float %add, float %add, float %p0)
+ %sub = fsub float 0.000000e+00, %x
+ %f1 = call float @llvm.fmuladd.f32(float %sub, float %sub, float %p1)
+ br label %loop
+}
>From f3de7325ad4c48a542a08014af730138e9ef0eb0 Mon Sep 17 00:00:00 2001
From: =?UTF-8?q?Bal=C3=A1zs=20Benics?= <benicsbalazs at gmail.com>
Date: Sat, 26 Sep 2026 09:47:17 -0400
Subject: [PATCH 09/53] [SSAF] Extract the virtual method override relation per
TU (#213316)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
A virtual call may dispatch to any override of its callee, so a whole-program
analysis cannot reason about a method's parameters and return value in
isolation. It needs to know which method overrides which, and which slots
that relates. Collect this per TU, so a later pass can join the related
slots into families.
JSON serialization lands separately, so the summary is not writable via
`--ssaf-extract-summaries` yet.
§1 of rdar://179151603
Assisted-By: claude
---
.../VirtualMethodFamily/VirtualMethodFamily.h | 51 ++++
.../BuiltinAnchorSources.def | 1 +
.../Analyses/CMakeLists.txt | 1 +
.../VirtualMethodEntityExtractor.cpp | 92 ++++++
.../VirtualMethodFamilyExtractorTest.cpp | 264 ++++++++++++++++++
.../VirtualMethodFamilyTestSupport.h | 155 ++++++++++
.../ScalableStaticAnalysis/CMakeLists.txt | 1 +
.../ScalableStaticAnalysis/ParsedAST.h | 145 ++++++++++
8 files changed, 710 insertions(+)
create mode 100644 clang/include/clang/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodFamily.h
create mode 100644 clang/lib/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodEntityExtractor.cpp
create mode 100644 clang/unittests/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodFamilyExtractorTest.cpp
create mode 100644 clang/unittests/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodFamilyTestSupport.h
create mode 100644 clang/unittests/ScalableStaticAnalysis/ParsedAST.h
diff --git a/clang/include/clang/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodFamily.h b/clang/include/clang/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodFamily.h
new file mode 100644
index 0000000000000..a39f2ffa1c9dc
--- /dev/null
+++ b/clang/include/clang/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodFamily.h
@@ -0,0 +1,51 @@
+//===- VirtualMethodFamily.h ------------------------------------*- C++ -*-===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+#ifndef LLVM_CLANG_SCALABLESTATICANALYSIS_ANALYSES_VIRTUALMETHODFAMILY_VIRTUALMETHODFAMILY_H
+#define LLVM_CLANG_SCALABLESTATICANALYSIS_ANALYSES_VIRTUALMETHODFAMILY_VIRTUALMETHODFAMILY_H
+
+#include "clang/ScalableStaticAnalysis/Core/Model/EntityId.h"
+#include "clang/ScalableStaticAnalysis/Core/Model/SummaryName.h"
+#include "clang/ScalableStaticAnalysis/Core/TUSummary/EntitySummary.h"
+#include "llvm/ADT/StringRef.h"
+#include <optional>
+#include <tuple>
+#include <vector>
+
+namespace clang::ssaf {
+
+struct VirtualMethodSummary final : public EntitySummary {
+ static constexpr llvm::StringLiteral Name = "VirtualMethod";
+
+ static SummaryName summaryName() { return SummaryName(Name.str()); }
+
+ SummaryName getSummaryName() const override { return summaryName(); }
+
+ /// EntityIds of each ParmVarDecl, in source order.
+ std::vector<EntityId> ParamEntities;
+
+ /// EntityId of the synthetic return-slot entity for this method.
+ std::optional<EntityId> ReturnEntity;
+
+ /// The result of \c CXXMethodDecl::overridden_methods().
+ std::vector<EntityId> OverriddenMethods;
+
+ bool operator==(const VirtualMethodSummary &Other) const {
+ return std::tie(ParamEntities, ReturnEntity, OverriddenMethods) ==
+ std::tie(Other.ParamEntities, Other.ReturnEntity,
+ Other.OverriddenMethods);
+ }
+
+ bool operator!=(const VirtualMethodSummary &Other) const {
+ return !(*this == Other);
+ }
+};
+
+} // namespace clang::ssaf
+
+#endif // LLVM_CLANG_SCALABLESTATICANALYSIS_ANALYSES_VIRTUALMETHODFAMILY_VIRTUALMETHODFAMILY_H
diff --git a/clang/include/clang/ScalableStaticAnalysis/BuiltinAnchorSources.def b/clang/include/clang/ScalableStaticAnalysis/BuiltinAnchorSources.def
index 4af2184b9d8a6..3ccb902ceb43e 100644
--- a/clang/include/clang/ScalableStaticAnalysis/BuiltinAnchorSources.def
+++ b/clang/include/clang/ScalableStaticAnalysis/BuiltinAnchorSources.def
@@ -31,5 +31,6 @@ ANCHOR(SharedLexicalRepresentationJSONFormatAnchorSource)
ANCHOR(UnsafeBufferUsageAnalysisAnchorSource)
ANCHOR(UnsafeBufferUsageExtractorAnchorSource)
ANCHOR(UnsafeBufferUsageJSONFormatAnchorSource)
+ANCHOR(VirtualMethodEntityExtractorAnchorSource)
#undef ANCHOR
diff --git a/clang/lib/ScalableStaticAnalysis/Analyses/CMakeLists.txt b/clang/lib/ScalableStaticAnalysis/Analyses/CMakeLists.txt
index 98ce8e799e0e0..cc9191922309b 100644
--- a/clang/lib/ScalableStaticAnalysis/Analyses/CMakeLists.txt
+++ b/clang/lib/ScalableStaticAnalysis/Analyses/CMakeLists.txt
@@ -20,6 +20,7 @@ add_clang_library(clangScalableStaticAnalysisAnalyses
UnsafeBufferUsage/UnsafeBufferUsageAnalysis.cpp
UnsafeBufferUsage/UnsafeBufferUsageExtractor.cpp
UnsafeBufferUsage/UnsafeBufferUsageFormat.cpp
+ VirtualMethodFamily/VirtualMethodEntityExtractor.cpp
LINK_LIBS
clangAST
diff --git a/clang/lib/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodEntityExtractor.cpp b/clang/lib/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodEntityExtractor.cpp
new file mode 100644
index 0000000000000..e3a55487d7505
--- /dev/null
+++ b/clang/lib/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodEntityExtractor.cpp
@@ -0,0 +1,92 @@
+//===- VirtualMethodEntityExtractor.cpp ----------------------------------===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+//
+// Extract what virtual methods override what other methods.
+// The parameters might be also important for consumers so collect those as
+// well - alongside with the ID of the return value.
+//
+//===----------------------------------------------------------------------===//
+
+#include "clang/AST/ASTContext.h"
+#include "clang/AST/Decl.h"
+#include "clang/AST/DeclCXX.h"
+#include "clang/AST/DynamicRecursiveASTVisitor.h"
+#include "clang/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodFamily.h"
+#include "clang/ScalableStaticAnalysis/Core/Model/EntityId.h"
+#include "clang/ScalableStaticAnalysis/Core/TUSummary/ExtractorRegistry.h"
+#include "clang/ScalableStaticAnalysis/Core/TUSummary/TUSummaryBuilder.h"
+#include "clang/ScalableStaticAnalysis/Core/TUSummary/TUSummaryExtractor.h"
+#include <memory>
+#include <optional>
+
+using namespace clang;
+using namespace ssaf;
+
+namespace {
+
+class VirtualMethodEntityExtractor final : public TUSummaryExtractor,
+ ConstDynamicRecursiveASTVisitor {
+public:
+ explicit VirtualMethodEntityExtractor(TUSummaryBuilder &Builder)
+ : TUSummaryExtractor(Builder) {
+ ShouldVisitTemplateInstantiations = true;
+ ShouldWalkTypesOfTypeLocs = false;
+ ShouldVisitImplicitCode = false;
+ ShouldVisitLambdaBody = true;
+ }
+
+private:
+ void HandleTranslationUnit(ASTContext &Ctx) override { TraverseAST(Ctx); }
+
+ bool VisitCXXMethodDecl(const CXXMethodDecl *MD) override;
+};
+} // namespace
+
+bool VirtualMethodEntityExtractor::VisitCXXMethodDecl(const CXXMethodDecl *MD) {
+ if (!MD->isVirtual())
+ return true;
+
+ std::optional<EntityId> MethodId = addEntity(MD);
+ if (!MethodId)
+ return true;
+
+ auto Summary = std::make_unique<VirtualMethodSummary>();
+ Summary->ParamEntities.reserve(MD->getNumParams());
+
+ for (const ParmVarDecl *P : MD->parameters()) {
+ auto ParamId = addEntity(P);
+ if (!ParamId) {
+ // If we can't get an EntityId for a parameter, drop the entire summary
+ // rather than leaving a half-populated record.
+ return true;
+ }
+ Summary->ParamEntities.push_back(ParamId.value());
+ }
+
+ if (auto ReturnId = addEntityForReturn(MD))
+ Summary->ReturnEntity = ReturnId.value();
+
+ for (const CXXMethodDecl *Overridden : MD->overridden_methods()) {
+ // We may not be able to convert methods that are coming from system
+ // headers, so skip them gracefully.
+ if (auto OverriddenId = addEntity(Overridden))
+ Summary->OverriddenMethods.push_back(*OverriddenId);
+ }
+
+ SummaryBuilder.addSummary(MethodId.value(), std::move(Summary));
+ return true;
+}
+
+static TUSummaryExtractorRegistry::Add<VirtualMethodEntityExtractor>
+ RegisterExtractor(VirtualMethodSummary::Name,
+ "Extract information about virtual methods");
+
+namespace clang::ssaf {
+// NOLINTNEXTLINE(misc-use-internal-linkage)
+volatile int VirtualMethodEntityExtractorAnchorSource = 0;
+} // namespace clang::ssaf
diff --git a/clang/unittests/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodFamilyExtractorTest.cpp b/clang/unittests/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodFamilyExtractorTest.cpp
new file mode 100644
index 0000000000000..8342ee3e56de7
--- /dev/null
+++ b/clang/unittests/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodFamilyExtractorTest.cpp
@@ -0,0 +1,264 @@
+//===- VirtualMethodFamilyExtractorTest.cpp -------------------------------===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+#include "VirtualMethodFamilyTestSupport.h"
+#include "clang/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodFamily.h"
+#include "clang/ScalableStaticAnalysis/Core/Model/EntityId.h"
+#include "clang/ScalableStaticAnalysis/Core/TUSummary/ExtractorRegistry.h"
+#include "gtest/gtest.h"
+
+#include <set>
+
+using namespace clang;
+using namespace ssaf;
+
+namespace {
+
+using VirtualMethodFamilyExtractorTest = VirtualMethodFamilyTestBase;
+
+TEST_F(VirtualMethodFamilyExtractorTest, Registers) {
+ EXPECT_TRUE(isTUSummaryExtractorRegistered(VirtualMethodSummary::Name));
+}
+
+using VirtualMethodFamilyExtractorBasicFieldPopulationTest =
+ VirtualMethodFamilyExtractorTest;
+
+TEST_F(VirtualMethodFamilyExtractorBasicFieldPopulationTest, BaseVirtual) {
+ ASSERT_TRUE(runVirtualMethodExtractor(R"cpp(
+ class Base {
+ public:
+ virtual void foo(int *p);
+ };
+ )cpp"));
+
+ const auto *S = getMethodSummary(AST.fn("Base::foo"));
+ ASSERT_TRUE(S);
+ EXPECT_TRUE(S->ReturnEntity.has_value());
+ EXPECT_EQ(S->ParamEntities.size(), 1u);
+ // A root virtual method overrides nothing.
+ EXPECT_TRUE(S->OverriddenMethods.empty());
+}
+
+TEST_F(VirtualMethodFamilyExtractorBasicFieldPopulationTest, PureVirtual) {
+ ASSERT_TRUE(runVirtualMethodExtractor(R"cpp(
+ class Interface {
+ public:
+ virtual void foo(int *p) = 0;
+ };
+ )cpp"));
+
+ // A pure-virtual method is still virtual, so a summary is produced for it.
+ ASSERT_TRUE(getMethodSummary(AST.fn("Interface::foo")));
+}
+
+TEST_F(VirtualMethodFamilyExtractorBasicFieldPopulationTest,
+ NonVirtualMethodSkipped) {
+ ASSERT_TRUE(runVirtualMethodExtractor(R"cpp(
+ class C {
+ public:
+ virtual void v();
+ void nv();
+ };
+ )cpp"));
+
+ // Only the virtual method has a summary; non-virtual is skipped.
+ EXPECT_EQ(methodSummaryCount(), 1u);
+ EXPECT_TRUE(getMethodSummary(AST.fn("C::v")));
+}
+
+TEST_F(VirtualMethodFamilyExtractorBasicFieldPopulationTest,
+ OverrideWithoutVirtualKeywordExtracted) {
+ ASSERT_TRUE(runVirtualMethodExtractor(R"cpp(
+ class Base {
+ public:
+ virtual void foo(int *p);
+ };
+ class Derived : public Base {
+ public:
+ void foo(int *p) override;
+ };
+ )cpp"));
+
+ EXPECT_EQ(methodSummaryCount(), 2u);
+ EXPECT_TRUE(getMethodSummary(AST.fn("Base::foo")));
+ EXPECT_TRUE(getMethodSummary(AST.fn("Derived::foo")));
+}
+
+using VirtualMethodFamilyExtractorOverriddenMethodTest =
+ VirtualMethodFamilyExtractorTest;
+
+TEST_F(VirtualMethodFamilyExtractorOverriddenMethodTest,
+ EdgeBaseDerivedOverride) {
+ ASSERT_TRUE(runVirtualMethodExtractor(R"cpp(
+ struct Base {
+ virtual void f(int *p);
+ };
+ struct Derived : Base {
+ void f(int *p) override;
+ };
+ )cpp"));
+ const auto *D = getMethodSummary(AST.fn("Derived::f"));
+ const auto *B = getMethodSummary(AST.fn("Base::f"));
+ ASSERT_TRUE(D);
+ ASSERT_TRUE(B);
+ auto BId = entityIdOf(AST.fn("Base::f"));
+ ASSERT_TRUE(BId.has_value());
+ ASSERT_EQ(D->OverriddenMethods.size(), 1u);
+ EXPECT_EQ(D->OverriddenMethods[0], *BId);
+ EXPECT_TRUE(B->OverriddenMethods.empty());
+}
+
+TEST_F(VirtualMethodFamilyExtractorOverriddenMethodTest,
+ EdgePureVirtualReabstractsOverride) {
+ // Tricky: a pure-virtual method that OVERRIDES a concrete virtual. Its edge
+ // set must be non-empty despite being pure.
+ ASSERT_TRUE(runVirtualMethodExtractor(R"cpp(
+ struct A {
+ virtual void f(int *p);
+ };
+ struct B : A {
+ void f(int *p) = 0;
+ };
+ )cpp"));
+ const auto *Bf = getMethodSummary(AST.fn("B::f"));
+ ASSERT_TRUE(Bf);
+ auto Af = entityIdOf(AST.fn("A::f"));
+ ASSERT_TRUE(Af.has_value());
+ ASSERT_EQ(Bf->OverriddenMethods.size(), 1u);
+ EXPECT_EQ(Bf->OverriddenMethods[0], *Af);
+}
+
+TEST_F(VirtualMethodFamilyExtractorOverriddenMethodTest,
+ EdgeMultipleInheritanceTwoEdges) {
+ // Tricky: one override occupies two independent base slots.
+ ASSERT_TRUE(runVirtualMethodExtractor(R"cpp(
+ struct A {
+ virtual void f(int *p);
+ };
+ struct B {
+ virtual void f(int *p);
+ };
+ struct D : A, B {
+ void f(int *p) override;
+ };
+ )cpp"));
+ const auto *Df = getMethodSummary(AST.fn("D::f"));
+ ASSERT_TRUE(Df);
+ auto Af = entityIdOf(AST.fn("A::f"));
+ auto Bf = entityIdOf(AST.fn("B::f"));
+ ASSERT_TRUE(Af.has_value() && Bf.has_value());
+ ASSERT_EQ(Df->OverriddenMethods.size(), 2u);
+ std::set<EntityId> Edges(Df->OverriddenMethods.begin(),
+ Df->OverriddenMethods.end());
+ EXPECT_EQ(Edges.count(*Af), 1u);
+ EXPECT_EQ(Edges.count(*Bf), 1u);
+}
+
+TEST_F(VirtualMethodFamilyExtractorOverriddenMethodTest,
+ EdgeTransitiveChainRecordsOnlyDirectOverride) {
+ ASSERT_TRUE(runVirtualMethodExtractor(R"cpp(
+ struct C {
+ virtual void f(int *p);
+ };
+ struct B : C {
+ void f(int *p) override;
+ };
+ struct A : B {
+ void f(int *p) override;
+ };
+ )cpp"));
+ const auto *Af = getMethodSummary(AST.fn("A::f"));
+ const auto *Bf = getMethodSummary(AST.fn("B::f"));
+ const auto *Cf = getMethodSummary(AST.fn("C::f"));
+ auto BfId = entityIdOf(AST.fn("B::f"));
+ auto CfId = entityIdOf(AST.fn("C::f"));
+ ASSERT_TRUE(Af);
+ ASSERT_TRUE(Bf);
+ ASSERT_TRUE(Cf);
+ ASSERT_TRUE(BfId.has_value());
+ ASSERT_TRUE(CfId.has_value());
+
+ // A::f overrides only its immediate base B::f, not the transitive C::f.
+ ASSERT_EQ(Af->OverriddenMethods.size(), 1u);
+ EXPECT_EQ(Af->OverriddenMethods[0], *BfId);
+
+ // B::f overrides C::f.
+ ASSERT_EQ(Bf->OverriddenMethods.size(), 1u);
+ EXPECT_EQ(Bf->OverriddenMethods[0], *CfId);
+
+ // C::f is a root and overrides nothing.
+ EXPECT_TRUE(Cf->OverriddenMethods.empty());
+}
+
+TEST_F(VirtualMethodFamilyExtractorOverriddenMethodTest,
+ EdgeOverrideLinksMatchingOverloadOnly) {
+ // Tricky: overloads must not be conflated. B::f(int*) overrides only the
+ // f(int*) base overload, never f(char*).
+ ASSERT_TRUE(runVirtualMethodExtractor(R"cpp(
+ struct A {
+ virtual void f(int *p);
+ virtual void f(char *p);
+ };
+ struct B : A {
+ void f(int *p) override;
+ };
+ )cpp"));
+ const auto *Bf = getMethodSummary(AST.fn("B::f"));
+ ASSERT_TRUE(Bf);
+ auto AfInt = entityIdOf(AST.fn("A::f(int *)"));
+ auto AfChar = entityIdOf(AST.fn("A::f(char *)"));
+ ASSERT_TRUE(AfInt.has_value() && AfChar.has_value());
+ ASSERT_EQ(Bf->OverriddenMethods.size(), 1u);
+ EXPECT_EQ(Bf->OverriddenMethods[0], *AfInt);
+ EXPECT_NE(Bf->OverriddenMethods[0], *AfChar);
+}
+
+TEST_F(VirtualMethodFamilyExtractorOverriddenMethodTest,
+ EdgeCovariantReturnOverride) {
+ ASSERT_TRUE(runVirtualMethodExtractor(R"cpp(
+ struct Base {
+ virtual Base *clone();
+ };
+ struct Deriv : Base {
+ Deriv *clone() override;
+ };
+ )cpp"));
+ const auto *Dc = getMethodSummary(AST.fn("Deriv::clone"));
+ ASSERT_TRUE(Dc);
+ auto Bc = entityIdOf(AST.fn("Base::clone"));
+ ASSERT_TRUE(Bc.has_value());
+ ASSERT_EQ(Dc->OverriddenMethods.size(), 1u);
+ EXPECT_EQ(Dc->OverriddenMethods[0], *Bc);
+ EXPECT_TRUE(Dc->ReturnEntity.has_value());
+}
+
+TEST_F(VirtualMethodFamilyExtractorOverriddenMethodTest,
+ EdgeDependentBaseTemplatePatternNoCrash) {
+ // Tricky: the primary template pattern has a dependent base; overridden_
+ // methods is unresolved there. Must not crash; the instantiation carries the
+ // edge.
+ ASSERT_TRUE(runVirtualMethodExtractor(R"cpp(
+ template <class T>
+ struct Wrapper {
+ virtual void f(int *p);
+ };
+ template <class T> struct DTypeParam : T {
+ virtual void f(int *p);
+ };
+ template <class T> struct DSpec : Wrapper<T> {
+ void f(int *p) override;
+ };
+ struct Concrete {
+ virtual void f(int *p);
+ };
+ template struct DSpec<Concrete>;
+ )cpp"));
+ EXPECT_GT(methodSummaryCount(), 0u);
+}
+
+} // namespace
diff --git a/clang/unittests/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodFamilyTestSupport.h b/clang/unittests/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodFamilyTestSupport.h
new file mode 100644
index 0000000000000..122f827bdda5f
--- /dev/null
+++ b/clang/unittests/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodFamilyTestSupport.h
@@ -0,0 +1,155 @@
+//===- VirtualMethodFamilyTestSupport.h -------------------------*- C++ -*-===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+//
+// Shared fixture for the VirtualMethodFamily tests: parses a snippet, runs the
+// VirtualMethod extractor over it, and resolves declarations to the EntityIds
+// the extractor minted.
+//
+//===----------------------------------------------------------------------===//
+
+#ifndef LLVM_CLANG_UNITTESTS_SCALABLESTATICANALYSIS_ANALYSES_VIRTUALMETHODFAMILY_VIRTUALMETHODFAMILYTESTSUPPORT_H
+#define LLVM_CLANG_UNITTESTS_SCALABLESTATICANALYSIS_ANALYSES_VIRTUALMETHODFAMILY_VIRTUALMETHODFAMILYTESTSUPPORT_H
+
+#include "ParsedAST.h"
+#include "TestFixture.h"
+#include "clang/Frontend/SSAFOptions.h"
+#include "clang/ScalableStaticAnalysis/Analyses/VirtualMethodFamily/VirtualMethodFamily.h"
+#include "clang/ScalableStaticAnalysis/Core/ASTEntityMapping.h"
+#include "clang/ScalableStaticAnalysis/Core/Model/BuildNamespace.h"
+#include "clang/ScalableStaticAnalysis/Core/Model/EntityId.h"
+#include "clang/ScalableStaticAnalysis/Core/Model/EntityName.h"
+#include "clang/ScalableStaticAnalysis/Core/TUSummary/ExtractorRegistry.h"
+#include "clang/ScalableStaticAnalysis/Core/TUSummary/TUSummary.h"
+#include "clang/ScalableStaticAnalysis/Core/TUSummary/TUSummaryBuilder.h"
+#include "llvm/TargetParser/Triple.h"
+
+#include <map>
+#include <optional>
+#include <string>
+
+namespace clang::ssaf {
+
+/// Base fixture for tests that need a TUSummary populated by the
+/// VirtualMethod extractor.
+class VirtualMethodFamilyTestBase : public TestFixture {
+protected:
+ ParsedAST AST;
+
+ /// Parses \p Code and runs the VirtualMethod extractor over it. Returns
+ /// false if the AST could not be built or the extractor is not registered.
+ /// Call once per test, before any of the lookups below.
+ [[nodiscard]] bool runVirtualMethodExtractor(llvm::StringRef Code) {
+ if (!AST.parse(Code))
+ return false;
+ auto Extractor =
+ makeTUSummaryExtractor(VirtualMethodSummary::Name, Builder);
+ if (!Extractor)
+ return false;
+ Extractor->HandleTranslationUnit(AST.getASTContext());
+ return true;
+ }
+
+ /// Resolves \p ND to the EntityId the extractor minted for it, or
+ /// std::nullopt if the extractor produced no entity for it.
+ std::optional<EntityId> entityIdOf(const NamedDecl *ND) const {
+ return ND ? lookup(getEntityName(ND)) : std::nullopt;
+ }
+
+ /// Resolves the return slot of \p FD to its EntityId.
+ std::optional<EntityId> returnEntityIdOf(const FunctionDecl *FD) const {
+ return FD ? lookup(getEntityNameForReturn(FD)) : std::nullopt;
+ }
+
+ /// Resolves an EntityName against the table the extractor populated.
+ std::optional<EntityId> lookup(std::optional<EntityName> Name) const {
+ if (!Name)
+ return std::nullopt;
+ const auto &Entities = getEntities(getIdTable(TUSum));
+ auto It = Entities.find(*Name);
+ if (It == Entities.end())
+ return std::nullopt;
+ return It->second;
+ }
+
+ /// Looks up the extractor's summary for \p FD, or nullptr if it produced
+ /// none for this function.
+ const VirtualMethodSummary *getMethodSummary(const FunctionDecl *FD) const {
+ auto Id = entityIdOf(FD);
+ if (!Id)
+ return nullptr;
+ const auto &Data = getData(TUSum);
+ auto SumIt = Data.find(VirtualMethodSummary::summaryName());
+ if (SumIt == Data.end())
+ return nullptr;
+ auto EIt = SumIt->second.find(*Id);
+ if (EIt == SumIt->second.end())
+ return nullptr;
+ return static_cast<const VirtualMethodSummary *>(EIt->second.get());
+ }
+
+ /// Count of method-summary entries in the TUSummary.
+ std::size_t methodSummaryCount() const {
+ const auto &Data = getData(TUSum);
+ auto It = Data.find(VirtualMethodSummary::summaryName());
+ if (It == Data.end())
+ return 0;
+ return It->second.size();
+ }
+
+ /// Maps every EntityId the extractor minted for the parsed snippet to a
+ /// readable label: "Base::foo(int *)", "Base::foo(int *)#return" or
+ /// "Base::foo(int *)#param0 'p'". Entities the extractor skipped are absent.
+ std::map<EntityId, std::string> entityLabels() const {
+ std::map<EntityId, std::string> Labels;
+ auto Add = [&](std::optional<EntityId> Id, std::string Label) {
+ if (Id)
+ Labels.insert({*Id, std::move(Label)});
+ };
+
+ for (const FunctionDecl *FD : AST.functions()) {
+ std::string Sig = ParsedAST::signatureOf(FD);
+ Add(entityIdOf(FD), Sig);
+ Add(returnEntityIdOf(FD), Sig + "#return");
+ for (unsigned I = 0, E = FD->getNumParams(); I != E; ++I) {
+ const ParmVarDecl *P = FD->getParamDecl(I);
+ std::string Label = Sig + "#param" + std::to_string(I);
+ if (!P->getName().empty())
+ Label += " '" + P->getName().str() + "'";
+ Add(entityIdOf(P), std::move(Label));
+ }
+ }
+ return Labels;
+ }
+
+ /// The entityLabels() mapping as text, to be streamed into a failing
+ /// assertion so that the EntityIds in its message can be decoded.
+ std::string legend() const {
+ std::map<EntityId, std::string> Labels = entityLabels();
+ if (Labels.empty())
+ return "\nid legend: <no entities extracted>";
+
+ std::string Result;
+ llvm::raw_string_ostream OS(Result);
+ OS << "\nid legend:";
+ for (const auto &[Id, Label] : Labels)
+ OS << "\n " << Id << " = " << Label;
+ return Result;
+ }
+
+ TUSummary &tuSummary() { return TUSum; }
+
+private:
+ SSAFOptions Opts;
+ BuildNamespace NS{BuildNamespaceKind::CompilationUnit, "Mock.cpp"};
+ TUSummary TUSum{llvm::Triple("arm64-apple-macosx"), NS};
+ TUSummaryBuilder Builder{TUSum, Opts};
+};
+
+} // namespace clang::ssaf
+
+#endif // LLVM_CLANG_UNITTESTS_SCALABLESTATICANALYSIS_ANALYSES_VIRTUALMETHODFAMILY_VIRTUALMETHODFAMILYTESTSUPPORT_H
diff --git a/clang/unittests/ScalableStaticAnalysis/CMakeLists.txt b/clang/unittests/ScalableStaticAnalysis/CMakeLists.txt
index 8c339bee60f63..aaee5b76a6e73 100644
--- a/clang/unittests/ScalableStaticAnalysis/CMakeLists.txt
+++ b/clang/unittests/ScalableStaticAnalysis/CMakeLists.txt
@@ -7,6 +7,7 @@ add_distinct_clang_unittest(ClangScalableAnalysisTests
Analyses/SharedLexicalRepresentation/EntitySourceLocationExtractorTest.cpp
Analyses/UnsafeBufferUsage/UnsafeBufferUsageTest.cpp
Analyses/UnsafeBufferUsage/UnsafeBufferUsageWPATest.cpp
+ Analyses/VirtualMethodFamily/VirtualMethodFamilyExtractorTest.cpp
ASTEntityMappingTest.cpp
BuildNamespaceTest.cpp
EntityIdTableTest.cpp
diff --git a/clang/unittests/ScalableStaticAnalysis/ParsedAST.h b/clang/unittests/ScalableStaticAnalysis/ParsedAST.h
new file mode 100644
index 0000000000000..71f9decd6ec58
--- /dev/null
+++ b/clang/unittests/ScalableStaticAnalysis/ParsedAST.h
@@ -0,0 +1,145 @@
+//===- ParsedAST.h ----------------------------------------------*- C++ -*-===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+//
+// Owns an AST parsed from a code snippet and resolves declarations in it by
+// qualified name. Held by value in a test fixture, so that fixtures which
+// already have a base class of their own can still get AST lookups.
+//
+// Unlike findDeclByName in FindDecl.h, lookups here take a *qualified* name
+// ("Base::foo") and can pick among overloads.
+//
+//===----------------------------------------------------------------------===//
+
+#ifndef LLVM_CLANG_UNITTESTS_SCALABLESTATICANALYSIS_PARSEDAST_H
+#define LLVM_CLANG_UNITTESTS_SCALABLESTATICANALYSIS_PARSEDAST_H
+
+#include "clang/AST/ASTContext.h"
+#include "clang/AST/DeclCXX.h"
+#include "clang/ASTMatchers/ASTMatchFinder.h"
+#include "clang/ASTMatchers/ASTMatchers.h"
+#include "clang/Frontend/ASTUnit.h"
+#include "clang/Tooling/Tooling.h"
+#include "llvm/ADT/STLExtras.h"
+#include "llvm/ADT/SmallVector.h"
+#include "llvm/ADT/StringRef.h"
+#include "llvm/Support/raw_ostream.h"
+#include "gtest/gtest.h"
+
+#include <memory>
+#include <string>
+
+namespace clang::ssaf {
+
+class ParsedAST {
+ std::unique_ptr<ASTUnit> AST;
+
+public:
+ /// Parses \p Code as C++17. Returns false if the AST could not be built, in
+ /// which case the lookups below all return nullptr.
+ [[nodiscard]] bool parse(llvm::StringRef Code) {
+ AST = tooling::buildASTFromCodeWithArgs(Code, {"-std=c++17"});
+ return AST != nullptr;
+ }
+
+ explicit operator bool() const { return AST != nullptr; }
+
+ ASTContext &getASTContext() const { return AST->getASTContext(); }
+
+ /// Finds a function by qualified name, e.g. "foo" or "Base::foo". Methods
+ /// are functions too, so this finds those as well.
+ ///
+ /// To pick among overloads, append the parameter list as clang prints it:
+ /// "A::f(int *)" and "A::f(char *)" name the two overloads of A::f. A bare
+ /// name that matches more than one function is an error, not a silent pick
+ /// of the first: it reports a gtest failure and returns nullptr, as does a
+ /// name that matches nothing.
+ const FunctionDecl *fn(llvm::StringRef NameOrSignature) const {
+ using namespace ast_matchers;
+ if (!AST)
+ return nullptr;
+
+ auto [Name, Params] = NameOrSignature.split('(');
+ auto Matches =
+ match(functionDecl(hasName(Name)).bind("f"), AST->getASTContext());
+
+ llvm::SmallVector<const FunctionDecl *> Candidates;
+ for (const auto &M : Matches) {
+ const auto *FD = M.getNodeAs<FunctionDecl>("f");
+ if (Params.empty() || paramsOf(FD) == Params.rtrim(')'))
+ Candidates.push_back(FD);
+ }
+
+ if (Candidates.size() == 1)
+ return Candidates.front();
+ if (Candidates.empty()) {
+ ADD_FAILURE() << "no function named '" << NameOrSignature << "'";
+ } else {
+ ADD_FAILURE() << "'" << NameOrSignature << "' is ambiguous; it matches "
+ << Candidates.size()
+ << " overloads. Append the parameter list to select one, "
+ "e.g. '"
+ << Name << "(" << paramsOf(Candidates.front()) << ")'";
+ }
+ return nullptr;
+ }
+
+ /// Finds parameter \p Index of the function named \p NameOrSignature.
+ /// Returns nullptr if there is no such function, or if it has too few
+ /// parameters.
+ const ParmVarDecl *findParam(llvm::StringRef NameOrSignature,
+ unsigned Index) const {
+ const FunctionDecl *FD = fn(NameOrSignature);
+ if (!FD)
+ return nullptr;
+ if (Index >= FD->getNumParams()) {
+ ADD_FAILURE() << "'" << NameOrSignature << "' has no parameter " << Index;
+ return nullptr;
+ }
+ return FD->getParamDecl(Index);
+ }
+
+ /// Every non-implicit function (methods included) in the parsed snippet, in
+ /// the order the matcher walks the AST. Unlike fn(), a miss is not a test
+ /// failure: this is a plain enumeration, meant for building diagnostics.
+ llvm::SmallVector<const FunctionDecl *> functions() const {
+ using namespace ast_matchers;
+ llvm::SmallVector<const FunctionDecl *> Result;
+ if (!AST)
+ return Result;
+ for (const auto &M :
+ match(functionDecl().bind("f"), AST->getASTContext())) {
+ const auto *FD = M.getNodeAs<FunctionDecl>("f");
+ if (FD && !FD->isImplicit())
+ Result.push_back(FD);
+ }
+ return Result;
+ }
+
+ /// \p FD spelled the way fn() takes it: "Base::foo(int *)".
+ static std::string signatureOf(const FunctionDecl *FD) {
+ if (!FD)
+ return "<null>";
+ return FD->getQualifiedNameAsString() + "(" + paramsOf(FD) + ")";
+ }
+
+private:
+ /// The parameter list of \p FD as clang prints it, without the parentheses:
+ /// "int *", or "unsigned long, A &".
+ static std::string paramsOf(const FunctionDecl *FD) {
+ std::string Result;
+ llvm::raw_string_ostream OS(Result);
+ llvm::interleave(
+ FD->parameters(), OS,
+ [&](const ParmVarDecl *P) { OS << P->getType().getAsString(); }, ", ");
+ return Result;
+ }
+};
+
+} // namespace clang::ssaf
+
+#endif // LLVM_CLANG_UNITTESTS_SCALABLESTATICANALYSIS_PARSEDAST_H
>From fd1131bb96ccaa459eeda2c3923c4800f131d545 Mon Sep 17 00:00:00 2001
From: Dmitry Sidorov <Dmitry.Sidorov at amd.com>
Date: Sat, 26 Sep 2026 16:00:18 +0200
Subject: [PATCH 10/53] [NFC][SLP][AMDGPU] Precommit a cross-block fmul fadd
contraction test (#226609)
Two contract fmuls feed contract fadds in different successors. AMDGPU
sinks such an fmul into the block of its user and fuses the pair, so the
scalar form is one fma per path. Record the current behaviour, the pair
is kept scalar when the fmul is operand 0 of the fadd and paired when it
is operand 1.
---
.../AMDGPU/cross-block-fmul-fadd.ll | 124 ++++++++++++++++++
1 file changed, 124 insertions(+)
create mode 100644 llvm/test/Transforms/SLPVectorizer/AMDGPU/cross-block-fmul-fadd.ll
diff --git a/llvm/test/Transforms/SLPVectorizer/AMDGPU/cross-block-fmul-fadd.ll b/llvm/test/Transforms/SLPVectorizer/AMDGPU/cross-block-fmul-fadd.ll
new file mode 100644
index 0000000000000..8af404e1ce3c8
--- /dev/null
+++ b/llvm/test/Transforms/SLPVectorizer/AMDGPU/cross-block-fmul-fadd.ll
@@ -0,0 +1,124 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
+; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgpu9.0a-amd-amdhsa < %s | FileCheck %s
+; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgpu9.42-amd-amdhsa < %s | FileCheck %s
+; RUN: opt -passes=slp-vectorizer -S -mtriple=amdgpu12.50-amd-amdhsa < %s | FileCheck %s
+
+; Two products are computed up front and each one feeds a contractable fadd in
+; a different successor. The backend sinks such a single use fmul into the
+; block of its fadd and fuses the pair, and it merges the adjacent scalar
+; loads on its own, so the scalar code is one fma per path. Pairing the
+; multiplies keeps the loads and trades the fma of each path for a packed
+; multiply before the branch plus an add on the path.
+
+define void @cross_block_fmul_lhs(ptr addrspace(1) %p, ptr addrspace(1) %q, ptr addrspace(1) %r, ptr addrspace(1) %r2, float %x, float %y, i1 %c) {
+; CHECK-LABEL: define void @cross_block_fmul_lhs(
+; CHECK-SAME: ptr addrspace(1) [[P:%.*]], ptr addrspace(1) [[Q:%.*]], ptr addrspace(1) [[R:%.*]], ptr addrspace(1) [[R2:%.*]], float [[X:%.*]], float [[Y:%.*]], i1 [[C:%.*]]) {
+; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: [[TID:%.*]] = call i32 @llvm.amdgcn.workitem.id.x()
+; CHECK-NEXT: [[I0:%.*]] = shl i32 [[TID]], 1
+; CHECK-NEXT: [[P0:%.*]] = getelementptr inbounds float, ptr addrspace(1) [[P]], i32 [[I0]]
+; CHECK-NEXT: [[P1:%.*]] = getelementptr inbounds float, ptr addrspace(1) [[P0]], i32 1
+; CHECK-NEXT: [[Q0:%.*]] = getelementptr inbounds float, ptr addrspace(1) [[Q]], i32 [[I0]]
+; CHECK-NEXT: [[Q1:%.*]] = getelementptr inbounds float, ptr addrspace(1) [[Q0]], i32 1
+; CHECK-NEXT: [[A0:%.*]] = load float, ptr addrspace(1) [[P0]], align 4
+; CHECK-NEXT: [[A1:%.*]] = load float, ptr addrspace(1) [[P1]], align 4
+; CHECK-NEXT: [[B0:%.*]] = load float, ptr addrspace(1) [[Q0]], align 4
+; CHECK-NEXT: [[B1:%.*]] = load float, ptr addrspace(1) [[Q1]], align 4
+; CHECK-NEXT: [[M0:%.*]] = fmul contract float [[A0]], [[B0]]
+; CHECK-NEXT: [[M1:%.*]] = fmul contract float [[A1]], [[B1]]
+; CHECK-NEXT: br i1 [[C]], label %[[T:.*]], label %[[F:.*]]
+; CHECK: [[T]]:
+; CHECK-NEXT: [[S0:%.*]] = fadd contract float [[M0]], [[X]]
+; CHECK-NEXT: [[RT:%.*]] = getelementptr inbounds float, ptr addrspace(1) [[R]], i32 [[TID]]
+; CHECK-NEXT: store float [[S0]], ptr addrspace(1) [[RT]], align 4
+; CHECK-NEXT: ret void
+; CHECK: [[F]]:
+; CHECK-NEXT: [[S1:%.*]] = fadd contract float [[M1]], [[Y]]
+; CHECK-NEXT: [[RF:%.*]] = getelementptr inbounds float, ptr addrspace(1) [[R2]], i32 [[TID]]
+; CHECK-NEXT: store float [[S1]], ptr addrspace(1) [[RF]], align 4
+; CHECK-NEXT: ret void
+;
+entry:
+ %tid = call i32 @llvm.amdgcn.workitem.id.x()
+ %i0 = shl i32 %tid, 1
+ %p0 = getelementptr inbounds float, ptr addrspace(1) %p, i32 %i0
+ %p1 = getelementptr inbounds float, ptr addrspace(1) %p0, i32 1
+ %q0 = getelementptr inbounds float, ptr addrspace(1) %q, i32 %i0
+ %q1 = getelementptr inbounds float, ptr addrspace(1) %q0, i32 1
+ %a0 = load float, ptr addrspace(1) %p0, align 4
+ %a1 = load float, ptr addrspace(1) %p1, align 4
+ %b0 = load float, ptr addrspace(1) %q0, align 4
+ %b1 = load float, ptr addrspace(1) %q1, align 4
+ %m0 = fmul contract float %a0, %b0
+ %m1 = fmul contract float %a1, %b1
+ br i1 %c, label %t, label %f
+
+t:
+ %s0 = fadd contract float %m0, %x
+ %rt = getelementptr inbounds float, ptr addrspace(1) %r, i32 %tid
+ store float %s0, ptr addrspace(1) %rt, align 4
+ ret void
+
+f:
+ %s1 = fadd contract float %m1, %y
+ %rf = getelementptr inbounds float, ptr addrspace(1) %r2, i32 %tid
+ store float %s1, ptr addrspace(1) %rf, align 4
+ ret void
+}
+
+; The same with the fmul in operand 1 of the fadd.
+
+define void @cross_block_fmul_rhs(ptr addrspace(1) %p, ptr addrspace(1) %q, ptr addrspace(1) %r, ptr addrspace(1) %r2, float %x, float %y, i1 %c) {
+; CHECK-LABEL: define void @cross_block_fmul_rhs(
+; CHECK-SAME: ptr addrspace(1) [[P:%.*]], ptr addrspace(1) [[Q:%.*]], ptr addrspace(1) [[R:%.*]], ptr addrspace(1) [[R2:%.*]], float [[X:%.*]], float [[Y:%.*]], i1 [[C:%.*]]) {
+; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: [[TID:%.*]] = call i32 @llvm.amdgcn.workitem.id.x()
+; CHECK-NEXT: [[I0:%.*]] = shl i32 [[TID]], 1
+; CHECK-NEXT: [[P0:%.*]] = getelementptr inbounds float, ptr addrspace(1) [[P]], i32 [[I0]]
+; CHECK-NEXT: [[Q0:%.*]] = getelementptr inbounds float, ptr addrspace(1) [[Q]], i32 [[I0]]
+; CHECK-NEXT: [[TMP0:%.*]] = load <2 x float>, ptr addrspace(1) [[P0]], align 4
+; CHECK-NEXT: [[TMP1:%.*]] = load <2 x float>, ptr addrspace(1) [[Q0]], align 4
+; CHECK-NEXT: [[TMP2:%.*]] = fmul contract <2 x float> [[TMP0]], [[TMP1]]
+; CHECK-NEXT: br i1 [[C]], label %[[T:.*]], label %[[F:.*]]
+; CHECK: [[T]]:
+; CHECK-NEXT: [[TMP3:%.*]] = extractelement <2 x float> [[TMP2]], i64 0
+; CHECK-NEXT: [[S0:%.*]] = fadd contract float [[X]], [[TMP3]]
+; CHECK-NEXT: [[RT:%.*]] = getelementptr inbounds float, ptr addrspace(1) [[R]], i32 [[TID]]
+; CHECK-NEXT: store float [[S0]], ptr addrspace(1) [[RT]], align 4
+; CHECK-NEXT: ret void
+; CHECK: [[F]]:
+; CHECK-NEXT: [[TMP4:%.*]] = extractelement <2 x float> [[TMP2]], i64 1
+; CHECK-NEXT: [[S1:%.*]] = fadd contract float [[Y]], [[TMP4]]
+; CHECK-NEXT: [[RF:%.*]] = getelementptr inbounds float, ptr addrspace(1) [[R2]], i32 [[TID]]
+; CHECK-NEXT: store float [[S1]], ptr addrspace(1) [[RF]], align 4
+; CHECK-NEXT: ret void
+;
+entry:
+ %tid = call i32 @llvm.amdgcn.workitem.id.x()
+ %i0 = shl i32 %tid, 1
+ %p0 = getelementptr inbounds float, ptr addrspace(1) %p, i32 %i0
+ %p1 = getelementptr inbounds float, ptr addrspace(1) %p0, i32 1
+ %q0 = getelementptr inbounds float, ptr addrspace(1) %q, i32 %i0
+ %q1 = getelementptr inbounds float, ptr addrspace(1) %q0, i32 1
+ %a0 = load float, ptr addrspace(1) %p0, align 4
+ %a1 = load float, ptr addrspace(1) %p1, align 4
+ %b0 = load float, ptr addrspace(1) %q0, align 4
+ %b1 = load float, ptr addrspace(1) %q1, align 4
+ %m0 = fmul contract float %a0, %b0
+ %m1 = fmul contract float %a1, %b1
+ br i1 %c, label %t, label %f
+
+t:
+ %s0 = fadd contract float %x, %m0
+ %rt = getelementptr inbounds float, ptr addrspace(1) %r, i32 %tid
+ store float %s0, ptr addrspace(1) %rt, align 4
+ ret void
+
+f:
+ %s1 = fadd contract float %y, %m1
+ %rf = getelementptr inbounds float, ptr addrspace(1) %r2, i32 %tid
+ store float %s1, ptr addrspace(1) %rf, align 4
+ ret void
+}
+
+declare i32 @llvm.amdgcn.workitem.id.x()
>From 8d93cf98e4bcba2fce24720d8f28230b8b866175 Mon Sep 17 00:00:00 2001
From: Konstantinos Parasyris <koparasy at gmail.com>
Date: Sat, 26 Sep 2026 07:17:10 -0700
Subject: [PATCH 11/53] [CIR][CUDA] Read fat binary in CIRGen and store its
bytes on the module (#225971)
Two problems with the current implementation of GPU CIR:
1. `LoweringPrepare` read the fat binary from disk, via
`astCtx->getSourceManager().getFileManager().getVirtualFileSystem()`. A
transform pass should not do I/O, and this is one of the `ASTContext`
dependencies that keeps the post-CIRGen pipeline from being IR-to-IR.
Some of the related discussion on this has been done
[here](https://discourse.llvm.org/t/rfc-clangir-making-cir-pipeline-boundaries-first-class-driver-artifacts/90998)
and PR: #219048
2. `#cir.cu.binary_handle` stored the *file path*. That's a build input,
not a
property of the program: it bakes one machine's directory layout into a
`.cir`.
The PR removes the file path and adds an attribute
CIRGen now reads the file in and records the contents
as `#cir.cu.device_binary`, a `StringAttr` of raw bytes. The
attribute is deleted once LoweringPrepare "consumes" it.
A couple of notes:
1. The GPU dialect has a similar representation. We may want to move
towards that direction slowly, but I don't see any value doing this now.
It will confuse the direction of this PR.
2. The test `clang/test/CIR/Diagnostics/mlir-error-routing.cu` has
completely
changed, because it relied on an error occurring when the file attribute
was missing.
Since we deleted that attribute completely, the test was failing.
---------
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
---
.../clang/CIR/Dialect/IR/CIRCUDAAttrs.td | 18 -----
.../clang/CIR/Dialect/IR/CIRDialect.td | 4 +-
clang/lib/CIR/CodeGen/CIRGenCUDANV.cpp | 68 ++++++++++++++++++-
clang/lib/CIR/CodeGen/CIRGenModule.cpp | 11 ---
clang/lib/CIR/Dialect/IR/CIRDialect.cpp | 15 ++++
.../Dialect/Transforms/LoweringPrepare.cpp | 48 +++++--------
clang/test/CIR/CodeGenCUDA/device-stub.cu | 6 ++
.../test/CIR/CodeGenCUDA/missing-gpubinary.cu | 22 ++++++
.../CIR/Diagnostics/mlir-error-routing.cpp | 21 ++++++
.../CIR/Diagnostics/mlir-error-routing.cu | 20 ------
.../CIR/IR/invalid-cuda-device-binary.cir | 34 ++++++++++
11 files changed, 184 insertions(+), 83 deletions(-)
create mode 100644 clang/test/CIR/CodeGenCUDA/missing-gpubinary.cu
create mode 100644 clang/test/CIR/Diagnostics/mlir-error-routing.cpp
delete mode 100644 clang/test/CIR/Diagnostics/mlir-error-routing.cu
create mode 100644 clang/test/CIR/IR/invalid-cuda-device-binary.cir
diff --git a/clang/include/clang/CIR/Dialect/IR/CIRCUDAAttrs.td b/clang/include/clang/CIR/Dialect/IR/CIRCUDAAttrs.td
index d993e1b2b11eb..ebe8eaf60cc43 100644
--- a/clang/include/clang/CIR/Dialect/IR/CIRCUDAAttrs.td
+++ b/clang/include/clang/CIR/Dialect/IR/CIRCUDAAttrs.td
@@ -50,24 +50,6 @@ def CIR_CUDAExternallyInitializedAttr : CIR_Attr<"CUDAExternallyInitialized",
}];
let canHaveIllegalCXXABIType = 0;
}
-def CIR_CUDABinaryHandleAttr : CIR_Attr<
- "CUDABinaryHandle", "cu.binary_handle"
-> {
- let summary = "Fat binary handle for device code.";
- let description =
- [{
- This attribute is attached to the ModuleOp and records the binary file
- name passed to host.
-
- CUDA first compiles device-side code into a fat binary file. The file
- name is then passed into host-side code, which is used to create a handle
- and then generate various registration functions.
- }];
-
- let parameters = (ins "mlir::StringAttr":$name);
- let assemblyFormat = "`<` $name `>`";
-}
-
// No wrapper attribute: the kind is only ever printed by
// CIR_CUDAVarRegistrationInfoAttr's own assembly format.
def CIR_CUDADeviceVarKind : CIR_I32Enum<"CUDADeviceVarKind",
diff --git a/clang/include/clang/CIR/Dialect/IR/CIRDialect.td b/clang/include/clang/CIR/Dialect/IR/CIRDialect.td
index aee862afc02b4..96ea50b8badd1 100644
--- a/clang/include/clang/CIR/Dialect/IR/CIRDialect.td
+++ b/clang/include/clang/CIR/Dialect/IR/CIRDialect.td
@@ -94,7 +94,9 @@ def CIR_Dialect : Dialect {
static llvm::StringRef getTargetCPUAttrName() { return "cir.target-cpu"; }
static llvm::StringRef getTuneCPUAttrName() { return "cir.tune-cpu"; }
static llvm::StringRef getTargetFeaturesAttrName() { return "cir.target-features"; }
- static llvm::StringRef getCUDABinaryHandleAttrName() { return "cir.cu.binary_handle"; }
+ // Raw bytes of the device-side fat binary, read by CIRGen so LoweringPrepare
+ // can build the runtime-registration globals without doing file I/O.
+ static llvm::StringRef getCUDADeviceBinaryAttrName() { return "cir.cu.device_binary"; }
// Mangled symbol name of the C++20 named-module initializer function,
// precomputed by CIRGen so later passes don't need a live ASTContext.
static llvm::StringRef getCXXModuleInitFnNameAttrName() { return "cir.cxx_module_init_fn_name"; }
diff --git a/clang/lib/CIR/CodeGen/CIRGenCUDANV.cpp b/clang/lib/CIR/CodeGen/CIRGenCUDANV.cpp
index ab4baf336d379..79153ca788756 100644
--- a/clang/lib/CIR/CodeGen/CIRGenCUDANV.cpp
+++ b/clang/lib/CIR/CodeGen/CIRGenCUDANV.cpp
@@ -22,9 +22,14 @@
#include "clang/AST/GlobalDecl.h"
#include "clang/Basic/AddressSpaces.h"
#include "clang/Basic/Cuda.h"
+#include "clang/Basic/DiagnosticFrontend.h"
+#include "clang/Basic/FileManager.h"
+#include "clang/Basic/SourceManager.h"
#include "clang/CIR/Dialect/IR/CIRDialect.h"
#include "clang/CIR/Dialect/IR/CIRTypes.h"
#include "llvm/Support/Casting.h"
+#include "llvm/Support/MemoryBuffer.h"
+#include "llvm/Support/VirtualFileSystem.h"
using namespace clang;
using namespace clang::CIRGen;
@@ -56,6 +61,7 @@ class CIRGenNVCUDARuntime : public CIRGenCUDARuntime {
private:
void emitDeviceStubBodyNew(CIRGenFunction &cgf, cir::FuncOp fn,
FunctionArgList &args);
+ void recordDeviceBinary();
mlir::Value prepareKernelArgs(CIRGenFunction &cgf, mlir::Location loc,
FunctionArgList &args);
mlir::Operation *getKernelHandle(cir::FuncOp fn, GlobalDecl gd) override;
@@ -517,9 +523,69 @@ void CIRGenNVCUDARuntime::handleGlobalReplace(cir::GlobalOp oldGV,
}
}
+/// Whether this translation unit has anything for the CUDA runtime to register.
+/// These are the same two attributes LoweringPrepare collects to decide whether
+/// to build a module ctor, so both sides answer the question from one source.
+static bool hasEntitiesToRegister(mlir::ModuleOp module) {
+ // A walk, not a scan of the module body, so this stays in agreement with the
+ // recursive walk LoweringPrepare collects them with.
+ return module
+ ->walk([](mlir::Operation *op) {
+ if (op->hasAttr(cir::CUDAKernelNameAttr::getMnemonic()) ||
+ op->hasAttr(cir::CUDAVarRegistrationInfoAttr::getMnemonic()))
+ return mlir::WalkResult::interrupt();
+ return mlir::WalkResult::advance();
+ })
+ .wasInterrupted();
+}
+
+/// Read the device-side fat binary and record its contents on the module as
+/// `cir.cu.device_binary`, for LoweringPrepare to build the fatbin global from.
+///
+/// This mirrors the read in CGNVCUDARuntime::makeModuleCtorFunction, guards
+/// included: nothing is read in a compilation that would build no module
+/// constructor.
+void CIRGenNVCUDARuntime::recordDeviceBinary() {
+ StringRef binaryName = cgm.getCodeGenOpts().OffloadBinaryToEmbedFile;
+ if (binaryName.empty())
+ return;
+
+ const LangOptions &langOpts = cgm.getLangOpts();
+ if ((langOpts.HIP || !langOpts.GPURelocatableDeviceCode) &&
+ !hasEntitiesToRegister(cgm.getModule()))
+ return;
+
+ llvm::vfs::FileSystem &fs = cgm.getASTContext()
+ .getSourceManager()
+ .getFileManager()
+ .getVirtualFileSystem();
+ llvm::ErrorOr<std::unique_ptr<llvm::MemoryBuffer>> binaryOrErr =
+ fs.getBufferForFile(binaryName, /*FileSize=*/-1,
+ /*RequiresNullTerminator=*/false);
+ if (std::error_code ec = binaryOrErr.getError()) {
+ cgm.getDiags().Report(diag::err_cannot_open_file)
+ << binaryName << ec.message();
+ return;
+ }
+
+ // Typed as the fatbin global's array type so LoweringPrepare can use this
+ // attribute as the initializer as-is: attributes are uniqued on {value, type}
+ // and never freed, so building a second, typed copy there would keep the fat
+ // binary in memory twice.
+ StringRef bytes = binaryOrErr.get()->getBuffer();
+ mlir::MLIRContext &ctx = cgm.getMLIRContext();
+ auto charTy = cir::IntType::get(&ctx, cgm.getTarget().getCharWidth(),
+ /*isSigned=*/false);
+ auto fatbinTy = cir::ArrayType::get(charTy, bytes.size());
+ cgm.getModule()->setAttr(cir::CIRDialect::getCUDADeviceBinaryAttrName(),
+ mlir::StringAttr::get(bytes, fatbinTy));
+}
+
void CIRGenNVCUDARuntime::finalizeModule() {
- if (!cgm.getLangOpts().CUDAIsDevice)
+ if (!cgm.getLangOpts().CUDAIsDevice) {
+ recordDeviceBinary();
return;
+ }
// Mark ODR-used device variables as compiler used to prevent them from being
// eliminated by optimization. This is necessary for device variables
diff --git a/clang/lib/CIR/CodeGen/CIRGenModule.cpp b/clang/lib/CIR/CodeGen/CIRGenModule.cpp
index adffa7dfe2969..007788d7e27c7 100644
--- a/clang/lib/CIR/CodeGen/CIRGenModule.cpp
+++ b/clang/lib/CIR/CodeGen/CIRGenModule.cpp
@@ -216,17 +216,6 @@ CIRGenModule::CIRGenModule(mlir::MLIRContext &mlirContext,
/*line=*/0,
/*column=*/0));
}
-
- // Set CUDA GPU binary handle.
- if (langOpts.CUDA) {
- llvm::StringRef cudaBinaryName = codeGenOpts.OffloadBinaryToEmbedFile;
- if (!cudaBinaryName.empty()) {
- theModule->setAttr(cir::CIRDialect::getCUDABinaryHandleAttrName(),
- cir::CUDABinaryHandleAttr::get(
- &mlirContext, mlir::StringAttr::get(
- &mlirContext, cudaBinaryName)));
- }
- }
}
CIRGenModule::~CIRGenModule() = default;
diff --git a/clang/lib/CIR/Dialect/IR/CIRDialect.cpp b/clang/lib/CIR/Dialect/IR/CIRDialect.cpp
index ce406707f5942..d5a587ff6d81e 100644
--- a/clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+++ b/clang/lib/CIR/Dialect/IR/CIRDialect.cpp
@@ -270,6 +270,21 @@ cir::CIRDialect::verifyOperationAttribute(mlir::Operation *op,
<< "' attribute to be attached to '"
<< mlir::ModuleOp::getOperationName() << "'";
+ // LoweringPrepare uses this attribute directly as the fatbin global's
+ // initializer, so it must be a valid #cir.const_array payload for its type.
+ if (attrName == getCUDADeviceBinaryAttrName()) {
+ auto bytes = mlir::dyn_cast<mlir::StringAttr>(attr.getValue());
+ auto arrayTy =
+ bytes ? mlir::dyn_cast<cir::ArrayType>(bytes.getType()) : nullptr;
+ if (!arrayTy || arrayTy.getSize() != bytes.size())
+ return op->emitOpError()
+ << "expects '" << getCUDADeviceBinaryAttrName()
+ << "' to be a string typed as an array of its length";
+ return cir::ConstArrayAttr::verify([&] { return op->emitOpError(); },
+ arrayTy, bytes,
+ /*trailingZerosNum=*/0);
+ }
+
return success();
}
diff --git a/clang/lib/CIR/Dialect/Transforms/LoweringPrepare.cpp b/clang/lib/CIR/Dialect/Transforms/LoweringPrepare.cpp
index dff7646c640b0..9a3c9c9745eaa 100644
--- a/clang/lib/CIR/Dialect/Transforms/LoweringPrepare.cpp
+++ b/clang/lib/CIR/Dialect/Transforms/LoweringPrepare.cpp
@@ -32,10 +32,8 @@
#include "llvm/ADT/TypeSwitch.h"
#include "llvm/IR/Instructions.h"
#include "llvm/Support/ErrorHandling.h"
-#include "llvm/Support/MemoryBuffer.h"
#include "llvm/Support/Path.h"
#include "llvm/Support/VersionTuple.h"
-#include "llvm/Support/VirtualFileSystem.h"
#include <map>
#include <memory>
@@ -2538,31 +2536,14 @@ void LoweringPreparePass::buildCUDAModuleCtor() {
// There's no device-side binary, so no need to proceed for CUDA.
// HIP has to create an external symbol in this case, which is NYI.
- mlir::Attribute cudaBinaryHandleAttr =
- mlirModule->getAttr(CIRDialect::getCUDABinaryHandleAttrName());
- if (!cudaBinaryHandleAttr) {
+ auto deviceBinaryAttr = mlirModule->getAttrOfType<mlir::StringAttr>(
+ CIRDialect::getCUDADeviceBinaryAttrName());
+ if (!deviceBinaryAttr) {
if (isHIP)
assert(!cir::MissingFeatures::hipModuleCtor());
return;
}
- llvm::StringRef cudaGPUBinaryName =
- mlir::cast<CUDABinaryHandleAttr>(cudaBinaryHandleAttr)
- .getName()
- .getValue();
-
- llvm::vfs::FileSystem &vfs =
- astCtx->getSourceManager().getFileManager().getVirtualFileSystem();
- llvm::ErrorOr<std::unique_ptr<llvm::MemoryBuffer>> gpuBinaryOrErr =
- vfs.getBufferForFile(cudaGPUBinaryName);
- if (std::error_code ec = gpuBinaryOrErr.getError()) {
- mlirModule->emitError("cannot open GPU binary file: " + cudaGPUBinaryName +
- ": " + ec.message());
- return;
- }
- std::unique_ptr<llvm::MemoryBuffer> gpuBinary =
- std::move(gpuBinaryOrErr.get());
-
// Set up common types and builder.
llvm::StringRef cudaPrefix = getCUDAPrefix(getLangOpts());
mlir::Location loc = mlirModule->getLoc();
@@ -2573,9 +2554,6 @@ void LoweringPreparePass::buildCUDAModuleCtor() {
PointerType voidPtrTy = builder.getVoidPtrTy();
PointerType voidPtrPtrTy = builder.getPointerTo(voidPtrTy);
IntType intTy = builder.getSIntNTy(32);
- IntType charTy =
- cir::IntType::get(&getContext(), getTargetInfo().getCharWidth(),
- /*isSigned=*/false);
// --- Create fatbin globals ---
@@ -2587,8 +2565,8 @@ void LoweringPreparePass::buildCUDAModuleCtor() {
getLangOpts().HIP ? ".hipFatBinSegment" : ".nvFatBinSegment";
// Create the fatbin string constant with GPU binary contents.
- auto fatbinType =
- ArrayType::get(&getContext(), charTy, gpuBinary->getBuffer().size());
+ // The dialect verifier guarantees the attribute is typed as the array.
+ auto fatbinType = mlir::cast<ArrayType>(deviceBinaryAttr.getType());
std::string fatbinStrName = addUnderscoredPrefix(cudaPrefix, "_fatbin_str");
GlobalOp fatbinStr = GlobalOp::create(builder, loc, fatbinStrName, fatbinType,
/*isConstant=*/true, {},
@@ -2600,8 +2578,8 @@ void LoweringPreparePass::buildCUDAModuleCtor() {
fatbinStr.setAlignment(8);
}
- fatbinStr.setInitialValueAttr(cir::ConstArrayAttr::get(
- fatbinType, StringAttr::get(gpuBinary->getBuffer(), fatbinType)));
+ fatbinStr.setInitialValueAttr(
+ cir::ConstArrayAttr::get(fatbinType, deviceBinaryAttr));
fatbinStr.setSection(fatbinConstName);
fatbinStr.setPrivate();
@@ -2781,7 +2759,7 @@ void LoweringPreparePass::buildCUDAModuleCtor() {
}
std::optional<FuncOp> LoweringPreparePass::buildCUDAModuleDtor() {
- if (!mlirModule->getAttr(CIRDialect::getCUDABinaryHandleAttrName()))
+ if (!mlirModule->getAttr(CIRDialect::getCUDADeviceBinaryAttrName()))
return {};
llvm::StringRef prefix = getCUDAPrefix(getLangOpts());
@@ -2838,7 +2816,7 @@ std::optional<FuncOp> LoweringPreparePass::buildCUDAModuleDtor() {
/// the dtor list would cause a double-free. It is meant to be registered via
/// atexit() at the end of the module ctor.
std::optional<FuncOp> LoweringPreparePass::buildHIPModuleDtor() {
- if (!mlirModule->getAttr(CIRDialect::getCUDABinaryHandleAttrName()))
+ if (!mlirModule->getAttr(CIRDialect::getCUDADeviceBinaryAttrName()))
return {};
llvm::StringRef prefix = getCUDAPrefix(getLangOpts());
@@ -3118,8 +3096,14 @@ void LoweringPreparePass::runOnOperation() {
buildCXXGlobalInitFunc();
buildCXXGlobalTlsFunc();
- if (getLangOpts().CUDA && !getLangOpts().CUDAIsDevice)
+ if (getLangOpts().CUDA && !getLangOpts().CUDAIsDevice) {
buildCUDAModuleCtor();
+ // The fatbin global now references the same attribute; drop the module's
+ // reference so an emitted .cir doesn't print the bytes twice. This has to
+ // happen out here because the ctor and both dtor builders test the
+ // attribute to decide whether a device-side binary exists at all.
+ mlirModule->removeAttr(CIRDialect::getCUDADeviceBinaryAttrName());
+ }
buildGlobalCtorDtorList();
}
diff --git a/clang/test/CIR/CodeGenCUDA/device-stub.cu b/clang/test/CIR/CodeGenCUDA/device-stub.cu
index 768a013dc8319..a480cee5d8ebd 100644
--- a/clang/test/CIR/CodeGenCUDA/device-stub.cu
+++ b/clang/test/CIR/CodeGenCUDA/device-stub.cu
@@ -130,6 +130,10 @@ __device__ _BitInt(36) c;
// CIR: cir.global "private" constant cir_private @__cuda_fatbin_str = #cir.const_array<"GPU binary would be here." : !cir.array<!u8i x 25>> : !cir.array<!u8i x 25> {alignment = 8 : i64, section = ".nv_fatbin"}
+// The bytes arrive as the #cir.cu.device_binary module attribute, which
+// LoweringPrepare erases once they are in the global above.
+// CIR-NOT: cir.cu.device_binary
+
// Check the fatbin wrapper struct: { magic, version, ptr to fatbin, null }, with section.
// CIR: cir.global constant cir_private @__cuda_fatbin_wrapper = #cir.const_record<{
// CIR-SAME: #cir.int<1180844977> : !s32i,
@@ -213,6 +217,7 @@ __device__ _BitInt(36) c;
// LLVM: call i32 @atexit(ptr @__cuda_module_dtor)
// No GPU binary — no registration infrastructure at all.
+// NOGPUBIN-NOT: cir.cu.device_binary
// NOGPUBIN-NOT: fatbin
// NOGPUBIN-NOT: gpubin
// NOGPUBIN-NOT: __cuda_register_globals
@@ -354,6 +359,7 @@ __device__ _BitInt(36) c;
// HIP-LLVM: ret void
// No GPU binary: no fatbin, no handle, no registration scaffolding.
+// HIP-NOGPUBIN-NOT: cir.cu.device_binary
// HIP-NOGPUBIN-NOT: __hip_fatbin
// HIP-NOGPUBIN-NOT: __hip_gpubin_handle
// HIP-NOGPUBIN-NOT: __hip_register_globals
diff --git a/clang/test/CIR/CodeGenCUDA/missing-gpubinary.cu b/clang/test/CIR/CodeGenCUDA/missing-gpubinary.cu
new file mode 100644
index 0000000000000..964ec4a40bffd
--- /dev/null
+++ b/clang/test/CIR/CodeGenCUDA/missing-gpubinary.cu
@@ -0,0 +1,22 @@
+// A missing fat binary must produce the same clang diagnostic from ClangIR as
+// from classic codegen, not an MLIR pass error.
+
+// RUN: not %clang_cc1 -triple x86_64-linux-gnu -emit-cir %s -x cuda \
+// RUN: -target-sdk-version=12.3 -fcuda-include-gpubinary %t.nonexistent \
+// RUN: -o %t.cir 2>&1 | FileCheck %s
+
+// RUN: not %clang_cc1 -triple x86_64-linux-gnu -fclangir -emit-llvm %s -x cuda \
+// RUN: -target-sdk-version=12.3 -fcuda-include-gpubinary %t.nonexistent \
+// RUN: -o %t.ll 2>&1 | FileCheck %s
+
+// RUN: not %clang_cc1 -triple x86_64-linux-gnu -emit-llvm %s -x cuda \
+// RUN: -target-sdk-version=12.3 -fcuda-include-gpubinary %t.nonexistent \
+// RUN: -o %t-ogcg.ll 2>&1 | FileCheck %s
+
+#include "Inputs/cuda.h"
+
+// A kernel is needed: with nothing to register neither CIRGen nor classic
+// codegen reads the fat binary at all.
+__global__ void kernel() {}
+
+// CHECK: fatal error: cannot open file '{{.*}}.nonexistent':
diff --git a/clang/test/CIR/Diagnostics/mlir-error-routing.cpp b/clang/test/CIR/Diagnostics/mlir-error-routing.cpp
new file mode 100644
index 0000000000000..5ce9da67f1f25
--- /dev/null
+++ b/clang/test/CIR/Diagnostics/mlir-error-routing.cpp
@@ -0,0 +1,21 @@
+// RUN: not %clang_cc1 -triple x86_64-apple-macosx10.15 -fclangir -emit-llvm \
+// RUN: %s -o %t.ll 2>&1 | FileCheck %s
+
+// LoweringPrepare emits an MLIR-side error via mlir::Operation::emitError for
+// a thread_local variable on a target whose thread wrapper is replaceable.
+// CIRDiagnosticHandler must surface it in clang's `file:line:col: error: ...`
+// format rather than MLIR's `loc("file":N:M): error: ...` default.
+
+// CHECK: mlir-error-routing.cpp:[[#@LINE+7]]:1: error: Unhandled thread wrapper attributes for CC and Nounwind
+// CHECK-NOT: loc({{.*}}): error: Unhandled thread wrapper attributes
+
+struct S {
+ S();
+ ~S();
+};
+thread_local S s;
+S *use() { return &s; }
+
+// The generic CIR-to-CIR transform fatal error must not be reported on top of
+// the specific one: CIRGenAction gates it on hasErrorOccurred().
+// CHECK-NOT: error: CIR-to-CIR transformation failed
diff --git a/clang/test/CIR/Diagnostics/mlir-error-routing.cu b/clang/test/CIR/Diagnostics/mlir-error-routing.cu
deleted file mode 100644
index 25759ace8b5bd..0000000000000
--- a/clang/test/CIR/Diagnostics/mlir-error-routing.cu
+++ /dev/null
@@ -1,20 +0,0 @@
-// RUN: not %clang_cc1 -triple x86_64-linux-gnu -fclangir -emit-llvm -x cuda \
-// RUN: -target-sdk-version=12.3 -fcuda-include-gpubinary %t.missing.bin \
-// RUN: %s -o %t.ll 2>&1 | FileCheck %s
-
-// LoweringPrepare emits an MLIR-side error via mlir::Operation::emitError when
-// the requested CUDA gpubinary cannot be opened. With CIRDiagnosticHandler
-// installed, that diagnostic surfaces through clang's DiagnosticsEngine in
-// clang's standard format (`error: ...`) rather than MLIR's
-// `loc("file":N:M): error: ...` default-handler format.
-
-// CHECK: error: cannot open GPU binary file: {{.*}}.missing.bin
-// CHECK-NOT: loc({{.*}}): error: cannot open GPU binary file
-
-// The generic CIR-to-CIR transform fatal error must NOT be reported on top of
-// the specific MLIR-relayed error. CIRGenAction gates the fallback diag on
-// clang::DiagnosticsEngine::hasErrorOccurred() so users see one root cause,
-// not two.
-// CHECK-NOT: error: CIR-to-CIR transformation failed
-
-__attribute__((global)) void kernel() {}
diff --git a/clang/test/CIR/IR/invalid-cuda-device-binary.cir b/clang/test/CIR/IR/invalid-cuda-device-binary.cir
new file mode 100644
index 0000000000000..ef4d2c6b22374
--- /dev/null
+++ b/clang/test/CIR/IR/invalid-cuda-device-binary.cir
@@ -0,0 +1,34 @@
+// RUN: cir-opt %s -verify-diagnostics -split-input-file
+
+!u8i = !cir.int<u, 8>
+
+module attributes {cir.cu.device_binary = "abc" : !cir.array<!u8i x 3>} {
+}
+
+// -----
+
+// expected-error at +1 {{expects 'cir.cu.device_binary' to be a string typed as an array of its length}}
+module attributes {cir.cu.device_binary = "abc"} {
+}
+
+// -----
+
+!u8i = !cir.int<u, 8>
+
+// expected-error at +1 {{expects 'cir.cu.device_binary' to be a string typed as an array of its length}}
+module attributes {cir.cu.device_binary = "abc" : !cir.array<!u8i x 4>} {
+}
+
+// -----
+
+!u16i = !cir.int<u, 16>
+
+// expected-error at +1 {{expects !cir.int<u, 8> element type}}
+module attributes {cir.cu.device_binary = "abc" : !cir.array<!u16i x 3>} {
+}
+
+// -----
+
+// expected-error at +1 {{expects 'cir.cu.device_binary' to be a string typed as an array of its length}}
+module attributes {cir.cu.device_binary = 42 : i32} {
+}
>From 472a18e5af7e74cd5d6ce17025ccef3b06a57188 Mon Sep 17 00:00:00 2001
From: Mehdi Amini <joker.eph at gmail.com>
Date: Sat, 26 Sep 2026 16:19:56 +0200
Subject: [PATCH 12/53] [mlir] Remove remaining strict property assembly
opt-outs (#226702)
Migrate Toy, test, and Python dialect formats to use prop-dict for
inherent properties. Update assembly fixtures and expected output for
strict property printing.
Assisted-by: Codex
---
mlir/examples/toy/Ch4/include/toy/Ops.td | 3 +-
mlir/examples/toy/Ch5/include/toy/Ops.td | 3 +-
mlir/examples/toy/Ch6/include/toy/Ops.td | 3 +-
mlir/examples/toy/Ch7/include/toy/Ops.td | 3 +-
.../test-last-modified-callgraph.mlir | 4 +-
.../Analysis/DataFlow/test-last-modified.mlir | 16 +-
.../Analysis/DataFlow/test-next-access.mlir | 12 +-
.../test-strided-metadata-range-analysis.mlir | 4 +-
.../Dialect/Affine/affine-loop-normalize.mlir | 14 +-
.../Dialect/Affine/int-range-interface.mlir | 110 ++++++------
.../Dialect/Affine/simplify-min-max-ops.mlir | 36 ++--
.../Dialect/Affine/simplify-with-bounds.mlir | 12 +-
.../transform-op-simplify-min-max-ops.mlir | 36 ++--
.../Arith/int-range-analysis-convergence.mlir | 4 +-
.../Dialect/Arith/int-range-interface.mlir | 68 +++----
.../Arith/int-range-loop-iter-args.mlir | 6 +-
.../Dialect/Arith/int-range-narrowing.mlir | 140 +++++++--------
.../Arith/int-range-opts-bug-119045.mlir | 2 +-
mlir/test/Dialect/Arith/int-range-opts.mlir | 28 +--
.../GPU/int-range-interface-cluster.mlir | 24 +--
.../test/Dialect/GPU/int-range-interface.mlir | 166 +++++++++---------
.../Dialect/MLProgram/pipeline-globals.mlir | 2 +-
.../Dialect/MemRef/int-range-inference.mlir | 10 +-
.../Dialect/OpenACC/acc-implicit-routine.mlir | 2 +-
.../Dialect/Tensor/int-range-inference.mlir | 10 +-
.../Dialect/Vector/int-range-interface.mlir | 52 +++---
mlir/test/IR/attribute.mlir | 2 +-
.../infer-int-range-test-ops-invalid.mlir | 4 +-
.../infer-int-range-test-ops.mlir | 70 ++++----
.../loop-invariant-code-motion.mlir | 36 ++--
.../make-composed-folded-affine-apply.mlir | 26 +--
mlir/test/lib/Dialect/Test/TestDialect.td | 2 -
mlir/test/lib/Dialect/Test/TestOps.td | 30 ++--
mlir/test/lib/Dialect/Test/TestOpsSyntax.td | 6 +-
.../constant-str-attr-invalid.mlir | 2 +-
mlir/test/mlir-tblgen/op-format-invalid.td | 2 -
mlir/test/mlir-tblgen/op-format.mlir | 8 +-
mlir/test/mlir-tblgen/op-format.td | 33 +---
mlir/test/mlir-tblgen/pattern.mlir | 4 +-
mlir/test/python/dialects/python_test.py | 83 ++++-----
mlir/test/python/python_test_ops.td | 4 +-
41 files changed, 525 insertions(+), 557 deletions(-)
diff --git a/mlir/examples/toy/Ch4/include/toy/Ops.td b/mlir/examples/toy/Ch4/include/toy/Ops.td
index ba751e0a01fbe..0568396c5b2c3 100644
--- a/mlir/examples/toy/Ch4/include/toy/Ops.td
+++ b/mlir/examples/toy/Ch4/include/toy/Ops.td
@@ -24,7 +24,6 @@ include "toy/ShapeInferenceInterface.td"
// can define our operations.
def Toy_Dialect : Dialect {
let name = "toy";
- let useStrictPropertiesInAssemblyFormat = 0;
let cppNamespace = "::mlir::toy";
}
@@ -229,7 +228,7 @@ def GenericCallOp : Toy_Op<"generic_call",
// Specialize assembly printing and parsing using a declarative format.
let assemblyFormat = [{
- $callee `(` $inputs `)` attr-dict `:` functional-type($inputs, results)
+ $callee `(` $inputs `)` prop-dict attr-dict `:` functional-type($inputs, results)
}];
// Add custom build methods for the generic call operation.
diff --git a/mlir/examples/toy/Ch5/include/toy/Ops.td b/mlir/examples/toy/Ch5/include/toy/Ops.td
index effd58aadc136..af7dc471d1708 100644
--- a/mlir/examples/toy/Ch5/include/toy/Ops.td
+++ b/mlir/examples/toy/Ch5/include/toy/Ops.td
@@ -24,7 +24,6 @@ include "toy/ShapeInferenceInterface.td"
// can define our operations.
def Toy_Dialect : Dialect {
let name = "toy";
- let useStrictPropertiesInAssemblyFormat = 0;
let cppNamespace = "::mlir::toy";
}
@@ -228,7 +227,7 @@ def GenericCallOp : Toy_Op<"generic_call",
// Specialize assembly printing and parsing using a declarative format.
let assemblyFormat = [{
- $callee `(` $inputs `)` attr-dict `:` functional-type($inputs, results)
+ $callee `(` $inputs `)` prop-dict attr-dict `:` functional-type($inputs, results)
}];
// Add custom build methods for the generic call operation.
diff --git a/mlir/examples/toy/Ch6/include/toy/Ops.td b/mlir/examples/toy/Ch6/include/toy/Ops.td
index ca06b55e43c97..19111d600135a 100644
--- a/mlir/examples/toy/Ch6/include/toy/Ops.td
+++ b/mlir/examples/toy/Ch6/include/toy/Ops.td
@@ -24,7 +24,6 @@ include "toy/ShapeInferenceInterface.td"
// can define our operations.
def Toy_Dialect : Dialect {
let name = "toy";
- let useStrictPropertiesInAssemblyFormat = 0;
let cppNamespace = "::mlir::toy";
}
@@ -228,7 +227,7 @@ def GenericCallOp : Toy_Op<"generic_call",
// Specialize assembly printing and parsing using a declarative format.
let assemblyFormat = [{
- $callee `(` $inputs `)` attr-dict `:` functional-type($inputs, results)
+ $callee `(` $inputs `)` prop-dict attr-dict `:` functional-type($inputs, results)
}];
// Add custom build methods for the generic call operation.
diff --git a/mlir/examples/toy/Ch7/include/toy/Ops.td b/mlir/examples/toy/Ch7/include/toy/Ops.td
index 1dbbf7eda284a..cc25615590dc1 100644
--- a/mlir/examples/toy/Ch7/include/toy/Ops.td
+++ b/mlir/examples/toy/Ch7/include/toy/Ops.td
@@ -24,7 +24,6 @@ include "toy/ShapeInferenceInterface.td"
// can define our operations.
def Toy_Dialect : Dialect {
let name = "toy";
- let useStrictPropertiesInAssemblyFormat = 0;
let cppNamespace = "::mlir::toy";
// We set this bit to generate a declaration of the `materializeConstant`
@@ -252,7 +251,7 @@ def GenericCallOp : Toy_Op<"generic_call",
// Specialize assembly printing and parsing using a declarative format.
let assemblyFormat = [{
- $callee `(` $inputs `)` attr-dict `:` functional-type($inputs, results)
+ $callee `(` $inputs `)` prop-dict attr-dict `:` functional-type($inputs, results)
}];
// Add custom build methods for the generic call operation.
diff --git a/mlir/test/Analysis/DataFlow/test-last-modified-callgraph.mlir b/mlir/test/Analysis/DataFlow/test-last-modified-callgraph.mlir
index a5eba43ac68ab..1237733954088 100644
--- a/mlir/test/Analysis/DataFlow/test-last-modified-callgraph.mlir
+++ b/mlir/test/Analysis/DataFlow/test-last-modified-callgraph.mlir
@@ -169,7 +169,7 @@ func.func @call_and_store_before(%arg0: memref<f32>) -> memref<f32> {
%1 = arith.constant 1.0 : f32
memref.store %0, %arg0[] {tag_name = "pre"} : memref<f32>
memref.load %arg0[] {tag = "before_call"} : memref<f32>
- test.call_and_store @callee(%arg0), %arg0 {tag_name = "call", store_before_call = true} : (memref<f32>, memref<f32>) -> ()
+ test.call_and_store @callee(%arg0), %arg0 <store_before_call = true> {tag_name = "call"} : (memref<f32>, memref<f32>) -> ()
memref.load %arg0[] {tag = "after_call"} : memref<f32>
memref.store %1, %arg0[] {tag_name = "post"} : memref<f32>
return {tag = "return"} %arg0 : memref<f32>
@@ -221,7 +221,7 @@ func.func @call_and_store_after(%arg0: memref<f32>) -> memref<f32> {
%1 = arith.constant 1.0 : f32
memref.store %0, %arg0[] {tag_name = "pre"} : memref<f32>
memref.load %arg0[] {tag = "before_call"} : memref<f32>
- test.call_and_store @callee(%arg0), %arg0 {tag_name = "call", store_before_call = false} : (memref<f32>, memref<f32>) -> ()
+ test.call_and_store @callee(%arg0), %arg0 <store_before_call = false> {tag_name = "call"} : (memref<f32>, memref<f32>) -> ()
memref.load %arg0[] {tag = "after_call"} : memref<f32>
memref.store %1, %arg0[] {tag_name = "post"} : memref<f32>
return {tag = "return"} %arg0 : memref<f32>
diff --git a/mlir/test/Analysis/DataFlow/test-last-modified.mlir b/mlir/test/Analysis/DataFlow/test-last-modified.mlir
index da22fdbf56e84..157a287801ad0 100644
--- a/mlir/test/Analysis/DataFlow/test-last-modified.mlir
+++ b/mlir/test/Analysis/DataFlow/test-last-modified.mlir
@@ -131,7 +131,7 @@ func.func @store_with_a_region_before(%arg0: memref<f32>) -> memref<f32> {
%1 = arith.constant 1.0 : f32
memref.store %0, %arg0[] {tag_name = "pre"} : memref<f32>
memref.load %arg0[] {tag = "store_with_a_region_before::before"} : memref<f32>
- test.store_with_a_region %arg0 attributes { tag_name = "region", store_before_region = true } {
+ test.store_with_a_region %arg0 <store_before_region = true> attributes {tag_name = "region"} {
memref.load %arg0[] {tag = "inside_region"} : memref<f32>
test.store_with_a_region_terminator
} : memref<f32>
@@ -157,7 +157,7 @@ func.func @store_with_a_region_after(%arg0: memref<f32>) -> memref<f32> {
%1 = arith.constant 1.0 : f32
memref.store %0, %arg0[] {tag_name = "pre"} : memref<f32>
memref.load %arg0[] {tag = "store_with_a_region_after::before"} : memref<f32>
- test.store_with_a_region %arg0 attributes { tag_name = "region", store_before_region = false } {
+ test.store_with_a_region %arg0 <store_before_region = false> attributes {tag_name = "region"} {
memref.load %arg0[] {tag = "inside_region"} : memref<f32>
test.store_with_a_region_terminator
} : memref<f32>
@@ -186,7 +186,7 @@ func.func @store_with_a_region_before_containing_a_store(%arg0: memref<f32>) ->
%1 = arith.constant 1.0 : f32
memref.store %0, %arg0[] {tag_name = "pre"} : memref<f32>
memref.load %arg0[] {tag = "store_with_a_region_before_containing_a_store::before"} : memref<f32>
- test.store_with_a_region %arg0 attributes { tag_name = "region", store_before_region = true } {
+ test.store_with_a_region %arg0 <store_before_region = true> attributes {tag_name = "region"} {
memref.load %arg0[] {tag = "enter_region"} : memref<f32>
%2 = arith.constant 2.0 : f32
memref.store %2, %arg0[] {tag_name = "inner"} : memref<f32>
@@ -218,7 +218,7 @@ func.func @store_with_a_region_after_containing_a_store(%arg0: memref<f32>) -> m
%1 = arith.constant 1.0 : f32
memref.store %0, %arg0[] {tag_name = "pre"} : memref<f32>
memref.load %arg0[] {tag = "store_with_a_region_after_containing_a_store::before"} : memref<f32>
- test.store_with_a_region %arg0 attributes { tag_name = "region", store_before_region = false } {
+ test.store_with_a_region %arg0 <store_before_region = false> attributes {tag_name = "region"} {
memref.load %arg0[] {tag = "enter_region"} : memref<f32>
%2 = arith.constant 2.0 : f32
memref.store %2, %arg0[] {tag_name = "inner"} : memref<f32>
@@ -247,7 +247,7 @@ func.func @store_with_a_loop_region_before(%arg0: memref<f32>) -> memref<f32> {
%1 = arith.constant 1.0 : f32
memref.store %0, %arg0[] {tag_name = "pre"} : memref<f32>
memref.load %arg0[] {tag = "store_with_a_loop_region_before::before"} : memref<f32>
- test.store_with_a_loop_region %arg0 attributes { tag_name = "region", store_before_region = true } {
+ test.store_with_a_loop_region %arg0 <store_before_region = true> attributes {tag_name = "region"} {
memref.load %arg0[] {tag = "inside_region"} : memref<f32>
test.store_with_a_region_terminator
} : memref<f32>
@@ -273,7 +273,7 @@ func.func @store_with_a_loop_region_after(%arg0: memref<f32>) -> memref<f32> {
%1 = arith.constant 1.0 : f32
memref.store %0, %arg0[] {tag_name = "pre"} : memref<f32>
memref.load %arg0[] {tag = "store_with_a_loop_region_after::before"} : memref<f32>
- test.store_with_a_loop_region %arg0 attributes { tag_name = "region", store_before_region = false } {
+ test.store_with_a_loop_region %arg0 <store_before_region = false> attributes {tag_name = "region"} {
memref.load %arg0[] {tag = "inside_region"} : memref<f32>
test.store_with_a_region_terminator
} : memref<f32>
@@ -304,7 +304,7 @@ func.func @store_with_a_loop_region_before_containing_a_store(%arg0: memref<f32>
%1 = arith.constant 1.0 : f32
memref.store %0, %arg0[] {tag_name = "pre"} : memref<f32>
memref.load %arg0[] {tag = "store_with_a_loop_region_before_containing_a_store::before"} : memref<f32>
- test.store_with_a_loop_region %arg0 attributes { tag_name = "region", store_before_region = true } {
+ test.store_with_a_loop_region %arg0 <store_before_region = true> attributes {tag_name = "region"} {
memref.load %arg0[] {tag = "enter_region"} : memref<f32>
%2 = arith.constant 2.0 : f32
memref.store %2, %arg0[] {tag_name = "inner"} : memref<f32>
@@ -337,7 +337,7 @@ func.func @store_with_a_loop_region_after_containing_a_store(%arg0: memref<f32>)
%1 = arith.constant 1.0 : f32
memref.store %0, %arg0[] {tag_name = "pre"} : memref<f32>
memref.load %arg0[] {tag = "store_with_a_loop_region_after_containing_a_store::before"} : memref<f32>
- test.store_with_a_loop_region %arg0 attributes { tag_name = "region", store_before_region = false } {
+ test.store_with_a_loop_region %arg0 <store_before_region = false> attributes {tag_name = "region"} {
memref.load %arg0[] {tag = "enter_region"} : memref<f32>
%2 = arith.constant 2.0 : f32
memref.store %2, %arg0[] {tag_name = "inner"} : memref<f32>
diff --git a/mlir/test/Analysis/DataFlow/test-next-access.mlir b/mlir/test/Analysis/DataFlow/test-next-access.mlir
index 700a23aa8bc40..3a945898b2315 100644
--- a/mlir/test/Analysis/DataFlow/test-next-access.mlir
+++ b/mlir/test/Analysis/DataFlow/test-next-access.mlir
@@ -418,7 +418,7 @@ func.func @call_and_store_before(%arg0: memref<f32>) {
// Note that the access after the entire call is "post".
// CHECK: name = "call"
// CHECK-SAME: next_access = {{\[}}["post"], ["post"]]
- test.call_and_store @callee(%arg0), %arg0 {name = "call", store_before_call = true} : (memref<f32>, memref<f32>) -> ()
+ test.call_and_store @callee(%arg0), %arg0 <store_before_call = true> {name = "call"} : (memref<f32>, memref<f32>) -> ()
// CHECK: name = "post"
// CHECK-SAME: next_access = ["unknown"]
memref.load %arg0[] {name = "post"} : memref<f32>
@@ -451,7 +451,7 @@ func.func @call_and_store_after(%arg0: memref<f32>) {
memref.load %arg0[] {name = "caller"} : memref<f32>
// CHECK: name = "call"
// CHECK-SAME: next_access = {{\[}}["post"], ["post"]]
- test.call_and_store @callee(%arg0), %arg0 {name = "call", store_before_call = false} : (memref<f32>, memref<f32>) -> ()
+ test.call_and_store @callee(%arg0), %arg0 <store_before_call = false> {name = "call"} : (memref<f32>, memref<f32>) -> ()
// CHECK: name = "post"
// CHECK-SAME: next_access = ["unknown"]
memref.load %arg0[] {name = "post"} : memref<f32>
@@ -472,7 +472,7 @@ func.func @store_with_a_region_before(%arg0: memref<f32>) {
// CHECK: name = "region"
// CHECK-SAME: next_access = {{\[}}["post"]]
// CHECK-SAME: next_at_entry_point = {{\[}}{{\[}}["post"]]]
- test.store_with_a_region %arg0 attributes { name = "region", store_before_region = true } {
+ test.store_with_a_region %arg0 <store_before_region = true> attributes {name = "region"} {
test.store_with_a_region_terminator
} : memref<f32>
memref.load %arg0[] {name = "post"} : memref<f32>
@@ -491,7 +491,7 @@ func.func @store_with_a_region_after(%arg0: memref<f32>) {
// CHECK: name = "region"
// CHECK-SAME: next_access = {{\[}}["post"]]
// CHECK-SAME: next_at_entry_point = {{\[}}{{\[}}["region"]]]
- test.store_with_a_region %arg0 attributes { name = "region", store_before_region = false } {
+ test.store_with_a_region %arg0 <store_before_region = false> attributes {name = "region"} {
test.store_with_a_region_terminator
} : memref<f32>
memref.load %arg0[] {name = "post"} : memref<f32>
@@ -515,7 +515,7 @@ func.func @store_with_a_region_before_containing_a_load(%arg0: memref<f32>) {
// CHECK: name = "region"
// CHECK-SAME: next_access = {{\[}}["post"]]
// CHECK-SAME: next_at_entry_point = {{\[}}{{\[}}["inner"]]]
- test.store_with_a_region %arg0 attributes { name = "region", store_before_region = true } {
+ test.store_with_a_region %arg0 <store_before_region = true> attributes {name = "region"} {
// CHECK: name = "inner"
// CHECK-SAME: next_access = {{\[}}["post"]]
memref.load %arg0[] {name = "inner"} : memref<f32>
@@ -544,7 +544,7 @@ func.func @store_with_a_region_after_containing_a_load(%arg0: memref<f32>) {
// CHECK: name = "region"
// CHECK-SAME: next_access = {{\[}}["post"]]
// CHECK-SAME: next_at_entry_point = {{\[}}{{\[}}["inner"]]]
- test.store_with_a_region %arg0 attributes { name = "region", store_before_region = false } {
+ test.store_with_a_region %arg0 <store_before_region = false> attributes {name = "region"} {
// CHECK: name = "inner"
// CHECK-SAME: next_access = {{\[}}["region"]]
memref.load %arg0[] {name = "inner"} : memref<f32>
diff --git a/mlir/test/Analysis/DataFlow/test-strided-metadata-range-analysis.mlir b/mlir/test/Analysis/DataFlow/test-strided-metadata-range-analysis.mlir
index 808c1c2bfd2a8..b0a984e4dae45 100644
--- a/mlir/test/Analysis/DataFlow/test-strided-metadata-range-analysis.mlir
+++ b/mlir/test/Analysis/DataFlow/test-strided-metadata-range-analysis.mlir
@@ -4,8 +4,8 @@ func.func @memref_subview(%arg0: memref<8x16x4xf32, strided<[64, 4, 1]>>, %arg1:
%c0 = arith.constant 0 : index
%c1 = arith.constant 1 : index
%c2 = arith.constant 2 : index
- %0 = test.with_bounds {smax = 13 : index, smin = 11 : index, umax = 13 : index, umin = 11 : index} : index
- %1 = test.with_bounds {smax = 7 : index, smin = 5 : index, umax = 7 : index, umin = 5 : index} : index
+ %0 = test.with_bounds <smax = 13 : index, smin = 11 : index, umax = 13 : index, umin = 11 : index> : index
+ %1 = test.with_bounds <smax = 7 : index, smin = 5 : index, umax = 7 : index, umin = 5 : index> : index
// Test subview with unknown sizes, and constant offsets and strides.
// CHECK: Op: %[[SV0:.*]] = memref.subview
diff --git a/mlir/test/Dialect/Affine/affine-loop-normalize.mlir b/mlir/test/Dialect/Affine/affine-loop-normalize.mlir
index 879ccd798eba3..ab2c3d056e134 100644
--- a/mlir/test/Dialect/Affine/affine-loop-normalize.mlir
+++ b/mlir/test/Dialect/Affine/affine-loop-normalize.mlir
@@ -333,7 +333,7 @@ func.func @multi_level_tiled_matmul() {
// USE-EXPENSIVE-MATH-LABEL: func @peeling_main_loop
func.func @peeling_main_loop() {
%c0 = arith.constant 0 : index
- %bound = test.value_with_bounds { min = 0 : index, max = 1 : index}
+ %bound = test.value_with_bounds < min = 0, max = 1>
affine.for %iv = %bound to 9 step 2 iter_args(%arg = %c0) -> index {
%sum = arith.addi %arg, %bound : index
affine.yield %sum : index
@@ -342,7 +342,7 @@ func.func @peeling_main_loop() {
}
// USE-EXPENSIVE-MATH: %[[C0:.*]] = arith.constant 0 : index
-// USE-EXPENSIVE-MATH: %[[BOUND:.*]] = test.value_with_bounds {max = 1 : index, min = 0 : index}
+// USE-EXPENSIVE-MATH: %[[BOUND:.*]] = test.value_with_bounds <min = 0, max = 1>
// USE-EXPENSIVE-MATH: %[[MAIN_RES:.*]] = affine.for %[[IV_MAIN:.*]] = 0 to 4 iter_args(%[[ARG_MAIN:.*]] = %[[C0]]) -> (index)
// USE-EXPENSIVE-MATH: %{{.*}} = affine.apply #[[$MAP_APPLY]](%[[IV_MAIN]])[%[[BOUND]]]
// USE-EXPENSIVE-MATH: %[[TAIL_RES:.*]] = affine.for %[[IV_TAIL:.*]] = 4 to #[[$MAP_UB]]()[%[[BOUND]]] iter_args(%[[ARG_TAIL:.*]] = %[[MAIN_RES]]) -> (index)
@@ -354,7 +354,7 @@ func.func @peeling_main_loop() {
// USE-EXPENSIVE-MATH-LABEL: func @fully_constantized_no_peeling
func.func @fully_constantized_no_peeling() {
%c0 = arith.constant 0 : index
- %bound = test.value_with_bounds { min = 0 : index, max = 1 : index}
+ %bound = test.value_with_bounds < min = 0, max = 1>
affine.for %iv = %bound to 6 step 2 iter_args(%arg = %c0) -> index {
%sum = arith.addi %arg, %bound : index
affine.yield %sum : index
@@ -363,7 +363,7 @@ func.func @fully_constantized_no_peeling() {
}
// USE-EXPENSIVE-MATH: %[[C0:.*]] = arith.constant 0 : index
-// USE-EXPENSIVE-MATH: %[[BOUND:.*]] = test.value_with_bounds {max = 1 : index, min = 0 : index}
+// USE-EXPENSIVE-MATH: %[[BOUND:.*]] = test.value_with_bounds <min = 0, max = 1>
// USE-EXPENSIVE-MATH: %{{.*}} = affine.for %[[IV:.*]] = 0 to 3 iter_args(%{{.*}} = %[[C0]]) -> (index)
// USE-EXPENSIVE-MATH: %{{.*}} = affine.apply #[[$MAP_APPLY]](%[[IV]])[%[[BOUND]]]
@@ -373,7 +373,7 @@ func.func @fully_constantized_no_peeling() {
// USE-EXPENSIVE-MATH-LABEL: func @peel_nested_loops
func.func @peel_nested_loops() {
%c0 = arith.constant 0 : index
- %bound = test.value_with_bounds { min = 0 : index, max = 1 : index}
+ %bound = test.value_with_bounds < min = 0, max = 1>
affine.for %i = %bound to 7 step 2 {
affine.for %j = %bound to 7 step 2 {
"test.foo"() : () -> ()
@@ -395,7 +395,7 @@ func.func @peel_nested_loops() {
// USE-EXPENSIVE-MATH-AND-PROMOTE-LABEL: func @single_iter_promoted
func.func @single_iter_promoted() {
%c0 = arith.constant 0 :index
- %bound = test.value_with_bounds { min = 0 : index, max = 1 : index}
+ %bound = test.value_with_bounds < min = 0, max = 1>
affine.for %iv = %bound to 2 step 2 {
"test.foo"() : () -> ()
}
@@ -409,7 +409,7 @@ func.func @single_iter_promoted() {
// USE-EXPENSIVE-MATH-AND-PROMOTE-LABEL: func @single_iter_promoted_with_remainder
func.func @single_iter_promoted_with_remainder() {
%c0 = arith.constant 0 : index
- %bound = test.value_with_bounds { min = 0 : index, max = 1 : index}
+ %bound = test.value_with_bounds < min = 0, max = 1>
affine.for %i = %bound to 3 step 2 {
"test.foo"() : () -> ()
}
diff --git a/mlir/test/Dialect/Affine/int-range-interface.mlir b/mlir/test/Dialect/Affine/int-range-interface.mlir
index ac64ad09ee244..8dc12401ca142 100644
--- a/mlir/test/Dialect/Affine/int-range-interface.mlir
+++ b/mlir/test/Dialect/Affine/int-range-interface.mlir
@@ -1,7 +1,7 @@
// RUN: mlir-opt --int-range-optimizations %s | FileCheck %s
// CHECK-LABEL: func @affine_apply_constant
-// CHECK: test.reflect_bounds {smax = 42 : index, smin = 42 : index, umax = 42 : index, umin = 42 : index}
+// CHECK: test.reflect_bounds <umin = 42 : index, umax = 42 : index, smin = 42 : index, smax = 42 : index>
func.func @affine_apply_constant() -> index {
%0 = affine.apply affine_map<() -> (42)>()
%1 = test.reflect_bounds %0 : index
@@ -9,66 +9,66 @@ func.func @affine_apply_constant() -> index {
}
// CHECK-LABEL: func @affine_apply_add
-// CHECK: test.reflect_bounds {smax = 15 : index, smin = 6 : index, umax = 15 : index, umin = 6 : index}
+// CHECK: test.reflect_bounds <umin = 6 : index, umax = 15 : index, smin = 6 : index, smax = 15 : index>
func.func @affine_apply_add() -> index {
- %d0 = test.with_bounds { umin = 2 : index, umax = 5 : index,
- smin = 2 : index, smax = 5 : index } : index
- %d1 = test.with_bounds { umin = 4 : index, umax = 10 : index,
- smin = 4 : index, smax = 10 : index } : index
+ %d0 = test.with_bounds < umin = 2 : index, umax = 5 : index,
+ smin = 2 : index, smax = 5 : index > : index
+ %d1 = test.with_bounds < umin = 4 : index, umax = 10 : index,
+ smin = 4 : index, smax = 10 : index > : index
%0 = affine.apply affine_map<(d0, d1) -> (d0 + d1)>(%d0, %d1)
%1 = test.reflect_bounds %0 : index
func.return %1 : index
}
// CHECK-LABEL: func @affine_apply_mul
-// CHECK: test.reflect_bounds {smax = 30 : index, smin = 12 : index, umax = 30 : index, umin = 12 : index}
+// CHECK: test.reflect_bounds <umin = 12 : index, umax = 30 : index, smin = 12 : index, smax = 30 : index>
func.func @affine_apply_mul() -> index {
- %d0 = test.with_bounds { umin = 2 : index, umax = 5 : index,
- smin = 2 : index, smax = 5 : index } : index
- %s0 = test.with_bounds { umin = 6 : index, umax = 6 : index,
- smin = 6 : index, smax = 6 : index } : index
+ %d0 = test.with_bounds < umin = 2 : index, umax = 5 : index,
+ smin = 2 : index, smax = 5 : index > : index
+ %s0 = test.with_bounds < umin = 6 : index, umax = 6 : index,
+ smin = 6 : index, smax = 6 : index > : index
%0 = affine.apply affine_map<(d0)[s0] -> (d0 * s0)>(%d0)[%s0]
%1 = test.reflect_bounds %0 : index
func.return %1 : index
}
// CHECK-LABEL: func @affine_apply_floordiv
-// CHECK: test.reflect_bounds {smax = 2 : index, smin = 1 : index, umax = 2 : index, umin = 1 : index}
+// CHECK: test.reflect_bounds <umin = 1 : index, umax = 2 : index, smin = 1 : index, smax = 2 : index>
func.func @affine_apply_floordiv() -> index {
- %d0 = test.with_bounds { umin = 5 : index, umax = 10 : index,
- smin = 5 : index, smax = 10 : index } : index
+ %d0 = test.with_bounds < umin = 5 : index, umax = 10 : index,
+ smin = 5 : index, smax = 10 : index > : index
%0 = affine.apply affine_map<(d0) -> (d0 floordiv 4)>(%d0)
%1 = test.reflect_bounds %0 : index
func.return %1 : index
}
// CHECK-LABEL: func @affine_apply_ceildiv
-// CHECK: test.reflect_bounds {smax = 3 : index, smin = 2 : index, umax = 3 : index, umin = 2 : index}
+// CHECK: test.reflect_bounds <umin = 2 : index, umax = 3 : index, smin = 2 : index, smax = 3 : index>
func.func @affine_apply_ceildiv() -> index {
- %d0 = test.with_bounds { umin = 5 : index, umax = 10 : index,
- smin = 5 : index, smax = 10 : index } : index
+ %d0 = test.with_bounds < umin = 5 : index, umax = 10 : index,
+ smin = 5 : index, smax = 10 : index > : index
%0 = affine.apply affine_map<(d0) -> (d0 ceildiv 4)>(%d0)
%1 = test.reflect_bounds %0 : index
func.return %1 : index
}
// CHECK-LABEL: func @affine_apply_mod
-// CHECK: test.reflect_bounds {smax = 3 : index, smin = 0 : index, umax = 3 : index, umin = 0 : index}
+// CHECK: test.reflect_bounds <umin = 0 : index, umax = 3 : index, smin = 0 : index, smax = 3 : index>
func.func @affine_apply_mod() -> index {
- %d0 = test.with_bounds { umin = 5 : index, umax = 27 : index,
- smin = 5 : index, smax = 27 : index } : index
+ %d0 = test.with_bounds < umin = 5 : index, umax = 27 : index,
+ smin = 5 : index, smax = 27 : index > : index
%0 = affine.apply affine_map<(d0) -> (d0 mod 4)>(%d0)
%1 = test.reflect_bounds %0 : index
func.return %1 : index
}
// CHECK-LABEL: func @affine_apply_complex
-// CHECK: test.reflect_bounds {smax = 13 : index, smin = 5 : index, umax = 13 : index, umin = 5 : index}
+// CHECK: test.reflect_bounds <umin = 5 : index, umax = 13 : index, smin = 5 : index, smax = 13 : index>
func.func @affine_apply_complex() -> index {
- %d0 = test.with_bounds { umin = 10 : index, umax = 20 : index,
- smin = 10 : index, smax = 20 : index } : index
- %d1 = test.with_bounds { umin = 3 : index, umax = 7 : index,
- smin = 3 : index, smax = 7 : index } : index
+ %d0 = test.with_bounds < umin = 10 : index, umax = 20 : index,
+ smin = 10 : index, smax = 20 : index > : index
+ %d1 = test.with_bounds < umin = 3 : index, umax = 7 : index,
+ smin = 3 : index, smax = 7 : index > : index
// (d0 floordiv 2) + (d1 mod 4) = [5, 10] + [0, 3] = [5, 13]
%0 = affine.apply affine_map<(d0, d1) -> (d0 floordiv 2 + d1 mod 4)>(%d0, %d1)
%1 = test.reflect_bounds %0 : index
@@ -76,12 +76,12 @@ func.func @affine_apply_complex() -> index {
}
// CHECK-LABEL: func @affine_apply_with_symbols
-// CHECK: test.reflect_bounds {smax = 24 : index, smin = 9 : index, umax = 24 : index, umin = 9 : index}
+// CHECK: test.reflect_bounds <umin = 9 : index, umax = 24 : index, smin = 9 : index, smax = 24 : index>
func.func @affine_apply_with_symbols() -> index {
- %d0 = test.with_bounds { umin = 2 : index, umax = 5 : index,
- smin = 2 : index, smax = 5 : index } : index
- %s0 = test.with_bounds { umin = 3 : index, umax = 4 : index,
- smin = 3 : index, smax = 4 : index } : index
+ %d0 = test.with_bounds < umin = 2 : index, umax = 5 : index,
+ smin = 2 : index, smax = 5 : index > : index
+ %s0 = test.with_bounds < umin = 3 : index, umax = 4 : index,
+ smin = 3 : index, smax = 4 : index > : index
// d0 * s0 + s0 = s0 * (d0 + 1) = [3, 4] * [3, 6] = [9, 24]
%0 = affine.apply affine_map<(d0)[s0] -> (d0 * s0 + s0)>(%d0)[%s0]
%1 = test.reflect_bounds %0 : index
@@ -89,12 +89,12 @@ func.func @affine_apply_with_symbols() -> index {
}
// CHECK-LABEL: func @affine_apply_sub
-// CHECK: test.reflect_bounds {smax = 1 : index, smin = -8 : index
+// CHECK: test.reflect_bounds <{{.*}}smin = -8 : index, smax = 1 : index>
func.func @affine_apply_sub() -> index {
- %d0 = test.with_bounds { umin = 2 : index, umax = 5 : index,
- smin = 2 : index, smax = 5 : index } : index
- %d1 = test.with_bounds { umin = 4 : index, umax = 10 : index,
- smin = 4 : index, smax = 10 : index } : index
+ %d0 = test.with_bounds < umin = 2 : index, umax = 5 : index,
+ smin = 2 : index, smax = 5 : index > : index
+ %d1 = test.with_bounds < umin = 4 : index, umax = 10 : index,
+ smin = 4 : index, smax = 10 : index > : index
// d0 - d1 = [2, 5] - [4, 10] = [2-10, 5-4] = [-8, 1]
%0 = affine.apply affine_map<(d0, d1) -> (d0 - d1)>(%d0, %d1)
%1 = test.reflect_bounds %0 : index
@@ -102,10 +102,10 @@ func.func @affine_apply_sub() -> index {
}
// CHECK-LABEL: func @affine_apply_mul_constant
-// CHECK: test.reflect_bounds {smax = 20 : index, smin = 8 : index, umax = 20 : index, umin = 8 : index}
+// CHECK: test.reflect_bounds <umin = 8 : index, umax = 20 : index, smin = 8 : index, smax = 20 : index>
func.func @affine_apply_mul_constant() -> index {
- %d0 = test.with_bounds { umin = 2 : index, umax = 5 : index,
- smin = 2 : index, smax = 5 : index } : index
+ %d0 = test.with_bounds < umin = 2 : index, umax = 5 : index,
+ smin = 2 : index, smax = 5 : index > : index
// d0 * 4 = [2, 5] * 4 = [8, 20]
%0 = affine.apply affine_map<(d0) -> (d0 * 4)>(%d0)
%1 = test.reflect_bounds %0 : index
@@ -113,10 +113,10 @@ func.func @affine_apply_mul_constant() -> index {
}
// CHECK-LABEL: func @affine_apply_mod_small_range
-// CHECK: test.reflect_bounds {smax = 2 : index, smin = 1 : index, umax = 2 : index, umin = 1 : index}
+// CHECK: test.reflect_bounds <umin = 1 : index, umax = 2 : index, smin = 1 : index, smax = 2 : index>
func.func @affine_apply_mod_small_range() -> index {
- %d0 = test.with_bounds { umin = 5 : index, umax = 6 : index,
- smin = 5 : index, smax = 6 : index } : index
+ %d0 = test.with_bounds < umin = 5 : index, umax = 6 : index,
+ smin = 5 : index, smax = 6 : index > : index
// Small range optimization: 5 mod 4 = 1, 6 mod 4 = 2, so [1, 2]
%0 = affine.apply affine_map<(d0) -> (d0 mod 4)>(%d0)
%1 = test.reflect_bounds %0 : index
@@ -124,10 +124,10 @@ func.func @affine_apply_mod_small_range() -> index {
}
// CHECK-LABEL: func @affine_apply_mod_already_in_range
-// CHECK: test.reflect_bounds {smax = 7 : index, smin = 5 : index, umax = 7 : index, umin = 5 : index}
+// CHECK: test.reflect_bounds <umin = 5 : index, umax = 7 : index, smin = 5 : index, smax = 7 : index>
func.func @affine_apply_mod_already_in_range() -> index {
- %d0 = test.with_bounds { umin = 5 : index, umax = 7 : index,
- smin = 5 : index, smax = 7 : index } : index
+ %d0 = test.with_bounds < umin = 5 : index, umax = 7 : index,
+ smin = 5 : index, smax = 7 : index > : index
// Dividend [5, 7] already in [0, 10), result equals dividend: [5, 7]
%0 = affine.apply affine_map<(d0) -> (d0 mod 10)>(%d0)
%1 = test.reflect_bounds %0 : index
@@ -135,12 +135,12 @@ func.func @affine_apply_mod_already_in_range() -> index {
}
// CHECK-LABEL: func @affine_apply_mod_variable_divisor
-// CHECK: test.reflect_bounds {smax = 4 : index, smin = 0 : index, umax = 4 : index, umin = 0 : index}
+// CHECK: test.reflect_bounds <umin = 0 : index, umax = 4 : index, smin = 0 : index, smax = 4 : index>
func.func @affine_apply_mod_variable_divisor() -> index {
- %d0 = test.with_bounds { umin = 10 : index, umax = 20 : index,
- smin = 10 : index, smax = 20 : index } : index
- %s0 = test.with_bounds { umin = 3 : index, umax = 5 : index,
- smin = 3 : index, smax = 5 : index } : index
+ %d0 = test.with_bounds < umin = 10 : index, umax = 20 : index,
+ smin = 10 : index, smax = 20 : index > : index
+ %s0 = test.with_bounds < umin = 3 : index, umax = 5 : index,
+ smin = 3 : index, smax = 5 : index > : index
// s0 can be 3, 4, or 5, so result is [0, max(s0)-1] = [0, 4]
%0 = affine.apply affine_map<(d0)[s0] -> (d0 mod s0)>(%d0)[%s0]
%1 = test.reflect_bounds %0 : index
@@ -148,10 +148,10 @@ func.func @affine_apply_mod_variable_divisor() -> index {
}
// CHECK-LABEL: func @affine_apply_mod_cross_boundary
-// CHECK: test.reflect_bounds {smax = 7 : index, smin = 0 : index, umax = 7 : index, umin = 0 : index}
+// CHECK: test.reflect_bounds <umin = 0 : index, umax = 7 : index, smin = 0 : index, smax = 7 : index>
func.func @affine_apply_mod_cross_boundary() -> index {
- %d0 = test.with_bounds { umin = 14 : index, umax = 17 : index,
- smin = 14 : index, smax = 17 : index } : index
+ %d0 = test.with_bounds < umin = 14 : index, umax = 17 : index,
+ smin = 14 : index, smax = 17 : index > : index
// Range [14, 17] spans a mod-8 boundary (at 16): 14%8=6, 17%8=1.
// Span 3 < 8 but the range wraps, so we fall back to [0, 7].
%0 = affine.apply affine_map<(d0) -> (d0 mod 8)>(%d0)
@@ -160,10 +160,10 @@ func.func @affine_apply_mod_cross_boundary() -> index {
}
// CHECK-LABEL: func @affine_apply_mod_negative_dividend
-// CHECK: test.reflect_bounds {smax = 3 : index, smin = 0 : index, umax = 3 : index, umin = 0 : index}
+// CHECK: test.reflect_bounds <umin = 0 : index, umax = 3 : index, smin = 0 : index, smax = 3 : index>
func.func @affine_apply_mod_negative_dividend() -> index {
- %d0 = test.with_bounds { umin = 0 : index, umax = 2 : index,
- smin = -2 : index, smax = 2 : index } : index
+ %d0 = test.with_bounds < umin = 0 : index, umax = 2 : index,
+ smin = -2 : index, smax = 2 : index > : index
// Negative dividend: signed range [-2, 2] mod 4
// Actual results: -2->2, -1->3, 0->0, 1->1, 2->2 (Euclidean mod)
// Range is NOT contiguous, so we return conservative [0, 3]
diff --git a/mlir/test/Dialect/Affine/simplify-min-max-ops.mlir b/mlir/test/Dialect/Affine/simplify-min-max-ops.mlir
index 300486eb1366b..96e3c32f6d49f 100644
--- a/mlir/test/Dialect/Affine/simplify-min-max-ops.mlir
+++ b/mlir/test/Dialect/Affine/simplify-min-max-ops.mlir
@@ -6,10 +6,10 @@
// CHECK: @min_max_full_simplify
func.func @min_max_full_simplify() -> (index, index) {
- %0 = test.value_with_bounds {max = 128 : index, min = 0 : index}
- %1 = test.value_with_bounds {max = 512 : index, min = 256 : index}
- // CHECK: %[[V0:.*]] = test.value_with_bounds {max = 128 : index, min = 0 : index}
- // CHECK: %[[V1:.*]] = test.value_with_bounds {max = 512 : index, min = 256 : index}
+ %0 = test.value_with_bounds <min = 0, max = 128>
+ %1 = test.value_with_bounds <min = 256, max = 512>
+ // CHECK: %[[V0:.*]] = test.value_with_bounds <min = 0, max = 128>
+ // CHECK: %[[V1:.*]] = test.value_with_bounds <min = 256, max = 512>
// CHECK-NOT: affine.min
// CHECK-NOT: affine.max
// CHECK: return %[[V0]], %[[V1]]
@@ -20,12 +20,12 @@ func.func @min_max_full_simplify() -> (index, index) {
// CHECK: @min_only_simplify
func.func @min_only_simplify() -> (index, index) {
- // CHECK: %[[V0:.*]] = test.value_with_bounds {max = 512 : index, min = 0 : index}
- // CHECK: %[[V1:.*]] = test.value_with_bounds {max = 512 : index, min = 256 : index}
+ // CHECK: %[[V0:.*]] = test.value_with_bounds <min = 0, max = 512>
+ // CHECK: %[[V1:.*]] = test.value_with_bounds <min = 256, max = 512>
// CHECK: affine.min #[[MAP_0]]()[%[[V0]]]
// CHECK: affine.max #[[MAP_1]]()[%[[V0]], %[[V1]]]
- %0 = test.value_with_bounds {max = 512 : index, min = 0 : index}
- %1 = test.value_with_bounds {max = 512 : index, min = 256 : index}
+ %0 = test.value_with_bounds <min = 0, max = 512>
+ %1 = test.value_with_bounds <min = 256, max = 512>
%r0 = affine.min affine_map<()[s0, s1] -> (s0, 32, s1)>()[%0, %1]
%r1 = affine.max affine_map<()[s0, s1] -> (s0, 32, s1)>()[%0, %1]
return %r0, %r1 : index, index
@@ -33,12 +33,12 @@ func.func @min_only_simplify() -> (index, index) {
// CHECK: @max_only_simplify
func.func @max_only_simplify() -> (index, index) {
- // CHECK: %[[V0:.*]] = test.value_with_bounds {max = 128 : index, min = 0 : index}
- // CHECK: %[[V1:.*]] = test.value_with_bounds {max = 512 : index, min = 0 : index}
+ // CHECK: %[[V0:.*]] = test.value_with_bounds <min = 0, max = 128>
+ // CHECK: %[[V1:.*]] = test.value_with_bounds <min = 0, max = 512>
// CHECK: affine.min #[[MAP_1]]()[%[[V0]], %[[V1]]]
// CHECK: affine.max #[[MAP_2]]()[%[[V1]]]
- %0 = test.value_with_bounds {max = 128 : index, min = 0 : index}
- %1 = test.value_with_bounds {max = 512 : index, min = 0 : index}
+ %0 = test.value_with_bounds <min = 0, max = 128>
+ %1 = test.value_with_bounds <min = 0, max = 512>
%r0 = affine.min affine_map<()[s0, s1] -> (s0, 256, s1)>()[%0, %1]
%r1 = affine.max affine_map<()[s0, s1] -> (s0, 256, s1)>()[%0, %1]
return %r0, %r1 : index, index
@@ -46,12 +46,12 @@ func.func @max_only_simplify() -> (index, index) {
// CHECK: @overlapping_constraints
func.func @overlapping_constraints() -> (index, index) {
- %0 = test.value_with_bounds {max = 192 : index, min = 0 : index}
- %1 = test.value_with_bounds {max = 384 : index, min = 128 : index}
- %2 = test.value_with_bounds {max = 512 : index, min = 256 : index}
- // CHECK: %[[V0:.*]] = test.value_with_bounds {max = 192 : index, min = 0 : index}
- // CHECK: %[[V1:.*]] = test.value_with_bounds {max = 384 : index, min = 128 : index}
- // CHECK: %[[V2:.*]] = test.value_with_bounds {max = 512 : index, min = 256 : index}
+ %0 = test.value_with_bounds <min = 0, max = 192>
+ %1 = test.value_with_bounds <min = 128, max = 384>
+ %2 = test.value_with_bounds <min = 256, max = 512>
+ // CHECK: %[[V0:.*]] = test.value_with_bounds <min = 0, max = 192>
+ // CHECK: %[[V1:.*]] = test.value_with_bounds <min = 128, max = 384>
+ // CHECK: %[[V2:.*]] = test.value_with_bounds <min = 256, max = 512>
// CHECK: affine.min #[[MAP_1]]()[%[[V0]], %[[V1]]]
// CHECK: affine.max #[[MAP_1]]()[%[[V1]], %[[V2]]]
%r0 = affine.min affine_map<()[s0, s1, s2] -> (s0, s1, s2)>()[%0, %1, %2]
diff --git a/mlir/test/Dialect/Affine/simplify-with-bounds.mlir b/mlir/test/Dialect/Affine/simplify-with-bounds.mlir
index 6cf474076f6fa..347b65ca0c987 100644
--- a/mlir/test/Dialect/Affine/simplify-with-bounds.mlir
+++ b/mlir/test/Dialect/Affine/simplify-with-bounds.mlir
@@ -174,9 +174,9 @@ func.func @input_not_linearize(%x: index) -> (index, index) {
func.func @simplify_loop_bound() -> index{
%c0 = arith.constant 0 :index
%c1 = arith.constant 1 : index
- %bound = test.value_with_bounds { min = 0 : index, max = 1 : index}
- %bound1 = test.value_with_bounds { min = 2 : index, max = 3 : index}
- %bound2 = test.value_with_bounds { min = 2 : index, max = 3 : index}
+ %bound = test.value_with_bounds < min = 0, max = 1>
+ %bound1 = test.value_with_bounds < min = 2, max = 3>
+ %bound2 = test.value_with_bounds < min = 2, max = 3>
%res = affine.for %iv = max affine_map<(d0, d1, d2) -> (d0, d1,d2)>(%bound, %bound1,%bound2) to min affine_map<(d0, d1, d2) -> (d0, d1,d2)>(%bound, %bound1,%bound2) step 2 iter_args(%arg = %c0) -> index {
%sum = arith.addi %arg, %c1 : index
affine.yield %sum : index
@@ -186,9 +186,9 @@ func.func @simplify_loop_bound() -> index{
// CHECK-DAG: %[[C0:.*]] = arith.constant 0 : index
// CHECK-DAG: %[[C1:.*]] = arith.constant 1 : index
-// CHECK-DAG: %[[BOUND0:.*]] = test.value_with_bounds {max = 1 : index, min = 0 : index}
-// CHECK-DAG: %[[BOUND1:.*]] = test.value_with_bounds {max = 3 : index, min = 2 : index}
-// CHECK-DAG: %[[BOUND2:.*]] = test.value_with_bounds {max = 3 : index, min = 2 : index}
+// CHECK-DAG: %[[BOUND0:.*]] = test.value_with_bounds <min = 0, max = 1>
+// CHECK-DAG: %[[BOUND1:.*]] = test.value_with_bounds <min = 2, max = 3>
+// CHECK-DAG: %[[BOUND2:.*]] = test.value_with_bounds <min = 2, max = 3>
// CHECK: affine.for %[[IV:.*]] = max #[[$MAP]]()[%[[BOUND1]], %[[BOUND2]]] to %[[BOUND0]] step 2 iter_args(%[[ARG:.*]] = %[[C0]]) -> (index) {
// CHECK: %[[SUM:.*]] = arith.addi %[[ARG]], %[[C1]] : index
// CHECK: affine.yield %[[SUM]] : index
diff --git a/mlir/test/Dialect/Affine/transform-op-simplify-min-max-ops.mlir b/mlir/test/Dialect/Affine/transform-op-simplify-min-max-ops.mlir
index aa9bdf2b34eb9..64f8ddc78361b 100644
--- a/mlir/test/Dialect/Affine/transform-op-simplify-min-max-ops.mlir
+++ b/mlir/test/Dialect/Affine/transform-op-simplify-min-max-ops.mlir
@@ -6,10 +6,10 @@
// CHECK: @min_max_full_simplify
func.func @min_max_full_simplify() -> (index, index) {
- %0 = test.value_with_bounds {max = 128 : index, min = 0 : index}
- %1 = test.value_with_bounds {max = 512 : index, min = 256 : index}
- // CHECK: %[[V0:.*]] = test.value_with_bounds {max = 128 : index, min = 0 : index}
- // CHECK: %[[V1:.*]] = test.value_with_bounds {max = 512 : index, min = 256 : index}
+ %0 = test.value_with_bounds <min = 0, max = 128>
+ %1 = test.value_with_bounds <min = 256, max = 512>
+ // CHECK: %[[V0:.*]] = test.value_with_bounds <min = 0, max = 128>
+ // CHECK: %[[V1:.*]] = test.value_with_bounds <min = 256, max = 512>
// CHECK-NOT: affine.min
// CHECK-NOT: affine.max
// CHECK: return %[[V0]], %[[V1]]
@@ -20,12 +20,12 @@ func.func @min_max_full_simplify() -> (index, index) {
// CHECK: @min_only_simplify
func.func @min_only_simplify() -> (index, index) {
- // CHECK: %[[V0:.*]] = test.value_with_bounds {max = 512 : index, min = 0 : index}
- // CHECK: %[[V1:.*]] = test.value_with_bounds {max = 512 : index, min = 256 : index}
+ // CHECK: %[[V0:.*]] = test.value_with_bounds <min = 0, max = 512>
+ // CHECK: %[[V1:.*]] = test.value_with_bounds <min = 256, max = 512>
// CHECK: affine.min #[[MAP_0]]()[%[[V0]]]
// CHECK: affine.max #[[MAP_1]]()[%[[V0]], %[[V1]]]
- %0 = test.value_with_bounds {max = 512 : index, min = 0 : index}
- %1 = test.value_with_bounds {max = 512 : index, min = 256 : index}
+ %0 = test.value_with_bounds <min = 0, max = 512>
+ %1 = test.value_with_bounds <min = 256, max = 512>
%r0 = affine.min affine_map<()[s0, s1] -> (s0, 32, s1)>()[%0, %1]
%r1 = affine.max affine_map<()[s0, s1] -> (s0, 32, s1)>()[%0, %1]
return %r0, %r1 : index, index
@@ -33,12 +33,12 @@ func.func @min_only_simplify() -> (index, index) {
// CHECK: @max_only_simplify
func.func @max_only_simplify() -> (index, index) {
- // CHECK: %[[V0:.*]] = test.value_with_bounds {max = 128 : index, min = 0 : index}
- // CHECK: %[[V1:.*]] = test.value_with_bounds {max = 512 : index, min = 0 : index}
+ // CHECK: %[[V0:.*]] = test.value_with_bounds <min = 0, max = 128>
+ // CHECK: %[[V1:.*]] = test.value_with_bounds <min = 0, max = 512>
// CHECK: affine.min #[[MAP_1]]()[%[[V0]], %[[V1]]]
// CHECK: affine.max #[[MAP_2]]()[%[[V1]]]
- %0 = test.value_with_bounds {max = 128 : index, min = 0 : index}
- %1 = test.value_with_bounds {max = 512 : index, min = 0 : index}
+ %0 = test.value_with_bounds <min = 0, max = 128>
+ %1 = test.value_with_bounds <min = 0, max = 512>
%r0 = affine.min affine_map<()[s0, s1] -> (s0, 256, s1)>()[%0, %1]
%r1 = affine.max affine_map<()[s0, s1] -> (s0, 256, s1)>()[%0, %1]
return %r0, %r1 : index, index
@@ -46,12 +46,12 @@ func.func @max_only_simplify() -> (index, index) {
// CHECK: @overlapping_constraints
func.func @overlapping_constraints() -> (index, index) {
- %0 = test.value_with_bounds {max = 192 : index, min = 0 : index}
- %1 = test.value_with_bounds {max = 384 : index, min = 128 : index}
- %2 = test.value_with_bounds {max = 512 : index, min = 256 : index}
- // CHECK: %[[V0:.*]] = test.value_with_bounds {max = 192 : index, min = 0 : index}
- // CHECK: %[[V1:.*]] = test.value_with_bounds {max = 384 : index, min = 128 : index}
- // CHECK: %[[V2:.*]] = test.value_with_bounds {max = 512 : index, min = 256 : index}
+ %0 = test.value_with_bounds <min = 0, max = 192>
+ %1 = test.value_with_bounds <min = 128, max = 384>
+ %2 = test.value_with_bounds <min = 256, max = 512>
+ // CHECK: %[[V0:.*]] = test.value_with_bounds <min = 0, max = 192>
+ // CHECK: %[[V1:.*]] = test.value_with_bounds <min = 128, max = 384>
+ // CHECK: %[[V2:.*]] = test.value_with_bounds <min = 256, max = 512>
// CHECK: affine.min #[[MAP_1]]()[%[[V0]], %[[V1]]]
// CHECK: affine.max #[[MAP_1]]()[%[[V1]], %[[V2]]]
%r0 = affine.min affine_map<()[s0, s1, s2] -> (s0, s1, s2)>()[%0, %1, %2]
diff --git a/mlir/test/Dialect/Arith/int-range-analysis-convergence.mlir b/mlir/test/Dialect/Arith/int-range-analysis-convergence.mlir
index a932d4b699a89..93fc0ac32f48f 100644
--- a/mlir/test/Dialect/Arith/int-range-analysis-convergence.mlir
+++ b/mlir/test/Dialect/Arith/int-range-analysis-convergence.mlir
@@ -70,7 +70,7 @@ func.func @grouped_gemm_while_hang(%n: i32, %flag: i1) -> i32 {
// The yielded `arith.remui` result stays at [0, 126]: the widening
// budget only fires on virtual `Lattice::join` at framework merge
// sites, not on transfer-function joins for inferrable ops.
- // CHECK: test.reflect_bounds {smax = 126 : si32, smin = 0 : si32, umax = 126 : ui32, umin = 0 : ui32}
+ // CHECK: test.reflect_bounds <umin = 0 : ui32, umax = 126 : ui32, smin = 0 : si32, smax = 126 : si32>
%r_l1 = test.reflect_bounds %L1 : i32
scf.yield %L1, %nic : i32, i1
}
@@ -83,7 +83,7 @@ func.func @grouped_gemm_while_hang(%n: i32, %flag: i1) -> i32 {
// presence of these bounds here is the convergence assertion: without
// the patch the analysis would not terminate to print this attribute.
// CHECK: %[[BOUNDED:.*]] = test.reflect_bounds
- // CHECK-SAME: {smax = 2147483647 : si32, smin = -2147483648 : si32, umax = 4294967295 : ui32, umin = 0 : ui32}
+ // CHECK-SAME: <umin = 0 : ui32, umax = 4294967295 : ui32, smin = -2147483648 : si32, smax = 2147483647 : si32>
// CHECK-SAME: %[[OUTER]]#0 : i32
%r = test.reflect_bounds %res#0 : i32
// CHECK: return %[[BOUNDED]] : i32
diff --git a/mlir/test/Dialect/Arith/int-range-interface.mlir b/mlir/test/Dialect/Arith/int-range-interface.mlir
index dd8240299ef7e..e2f8ac72d56e9 100644
--- a/mlir/test/Dialect/Arith/int-range-interface.mlir
+++ b/mlir/test/Dialect/Arith/int-range-interface.mlir
@@ -229,7 +229,7 @@ func.func @ceil_divui(%arg0 : index) -> i1 {
// CHECK-LABEL: func @ceil_divui_by_zero_issue_131273
// CHECK-NEXT: return
func.func @ceil_divui_by_zero_issue_131273() {
- %0 = test.with_bounds {smax = 0 : i32, smin = -1 : i32, umax = 0 : i32, umin = -1 : i32} : i32
+ %0 = test.with_bounds <smax = 0 : i32, smin = -1 : i32, umax = 0 : i32, umin = -1 : i32> : i32
%c7_i32 = arith.constant 7 : i32
%1 = arith.ceildivui %c7_i32, %0 : i32
return
@@ -276,9 +276,9 @@ func.func @ceil_divsi_full_range(%6: index) -> index {
// CHECK: %[[ret:.*]] = arith.constant true
// CHECK: return %[[ret]]
func.func @ceil_divsi_intmin_bug_115293() -> i1 {
- %intMin_i64 = test.with_bounds { smin = -9223372036854775808 : si64, smax = -9223372036854775808 : si64, umin = 9223372036854775808 : ui64, umax = 9223372036854775808 : ui64 } : i64
- %denom_i64 = test.with_bounds { smin = 1189465982 : si64, smax = 1189465982 : si64, umin = 1189465982 : ui64, umax = 1189465982 : ui64 } : i64
- %res_i64 = test.with_bounds { smin = 7754212542 : si64, smax = 7754212542 : si64, umin = 7754212542 : ui64, umax = 7754212542 : ui64 } : i64
+ %intMin_i64 = test.with_bounds < smin = -9223372036854775808 : si64, smax = -9223372036854775808 : si64, umin = 9223372036854775808 : ui64, umax = 9223372036854775808 : ui64 > : i64
+ %denom_i64 = test.with_bounds < smin = 1189465982 : si64, smax = 1189465982 : si64, umin = 1189465982 : ui64, umax = 1189465982 : ui64 > : i64
+ %res_i64 = test.with_bounds < smin = 7754212542 : si64, smax = 7754212542 : si64, umin = 7754212542 : ui64, umax = 7754212542 : ui64 > : i64
%0 = arith.ceildivsi %intMin_i64, %denom_i64 : i64
%1 = arith.cmpi eq, %0, %res_i64 : i64
@@ -484,7 +484,7 @@ func.func @ori(%arg0 : i128, %arg1 : i128) -> i1 {
func.func @xori_issue_82168() -> i1 {
%c0_i64 = arith.constant 0 : i64
%c2060639849_i64 = arith.constant 2060639849 : i64
- %2 = test.with_bounds { umin = 2060639849 : i64, umax = 2060639850 : i64, smin = 2060639849 : i64, smax = 2060639850 : i64 } : i64
+ %2 = test.with_bounds < umin = 2060639849 : i64, umax = 2060639850 : i64, smin = 2060639849 : i64, smax = 2060639850 : i64 > : i64
%3 = arith.xori %2, %c2060639849_i64 : i64
%4 = arith.cmpi eq, %3, %c0_i64 : i64
func.return %4 : i1
@@ -496,8 +496,8 @@ func.func @xori_issue_82168() -> i1 {
// CHECK: return %[[true]], %[[false]]
func.func @xori_i1() -> (i1, i1) {
%true = arith.constant true
- %1 = test.with_bounds { umin = 0 : i1, umax = 0 : i1, smin = 0 : i1, smax = 0 : i1 } : i1
- %2 = test.with_bounds { umin = 1 : i1, umax = 1 : i1, smin = 1 : i1, smax = 1 : i1 } : i1
+ %1 = test.with_bounds < umin = 0 : i1, umax = 0 : i1, smin = 0 : i1, smax = 0 : i1 > : i1
+ %2 = test.with_bounds < umin = 1 : i1, umax = 1 : i1, smin = 1 : i1, smax = 1 : i1 > : i1
%3 = arith.xori %1, %true : i1
%4 = arith.xori %2, %true : i1
func.return %3, %4 : i1, i1
@@ -864,20 +864,20 @@ func.func private @callee(%arg0: memref<?xindex, 4>) {
}
// CHECK-LABEL: func @test_i8_bounds
-// CHECK: test.reflect_bounds {smax = 127 : si8, smin = -128 : si8, umax = 255 : ui8, umin = 0 : ui8}
+// CHECK: test.reflect_bounds <umin = 0 : ui8, umax = 255 : ui8, smin = -128 : si8, smax = 127 : si8>
func.func @test_i8_bounds() -> i8 {
%cst1 = arith.constant 1 : i8
- %0 = test.with_bounds { umin = 0 : i8, umax = 255 : i8, smin = -128 : i8, smax = 127 : i8 } : i8
+ %0 = test.with_bounds < umin = 0 : i8, umax = 255 : i8, smin = -128 : i8, smax = 127 : i8 > : i8
%1 = arith.addi %0, %cst1 : i8
%2 = test.reflect_bounds %1 : i8
return %2: i8
}
// CHECK-LABEL: func @test_add_1
-// CHECK: test.reflect_bounds {smax = 127 : si8, smin = -128 : si8, umax = 255 : ui8, umin = 0 : ui8}
+// CHECK: test.reflect_bounds <umin = 0 : ui8, umax = 255 : ui8, smin = -128 : si8, smax = 127 : si8>
func.func @test_add_1() -> i8 {
%cst1 = arith.constant 1 : i8
- %0 = test.with_bounds { umin = 0 : i8, umax = 255 : i8, smin = -128 : i8, smax = 127 : i8 } : i8
+ %0 = test.with_bounds < umin = 0 : i8, umax = 255 : i8, smin = -128 : i8, smax = 127 : i8 > : i8
%1 = arith.addi %0, %cst1 : i8
%2 = test.reflect_bounds %1 : i8
return %2: i8
@@ -886,10 +886,10 @@ func.func @test_add_1() -> i8 {
// Tests below check inference with overflow flags.
// CHECK-LABEL: func @test_add_i8_wrap1
-// CHECK: test.reflect_bounds {smax = 127 : si8, smin = -128 : si8, umax = 128 : ui8, umin = 1 : ui8}
+// CHECK: test.reflect_bounds <umin = 1 : ui8, umax = 128 : ui8, smin = -128 : si8, smax = 127 : si8>
func.func @test_add_i8_wrap1() -> i8 {
%cst1 = arith.constant 1 : i8
- %0 = test.with_bounds { umin = 0 : i8, umax = 127 : i8, smin = 0 : i8, smax = 127 : i8 } : i8
+ %0 = test.with_bounds < umin = 0 : i8, umax = 127 : i8, smin = 0 : i8, smax = 127 : i8 > : i8
// smax overflow
%1 = arith.addi %0, %cst1 : i8
%2 = test.reflect_bounds %1 : i8
@@ -897,10 +897,10 @@ func.func @test_add_i8_wrap1() -> i8 {
}
// CHECK-LABEL: func @test_add_i8_wrap2
-// CHECK: test.reflect_bounds {smax = 127 : si8, smin = -128 : si8, umax = 128 : ui8, umin = 1 : ui8}
+// CHECK: test.reflect_bounds <umin = 1 : ui8, umax = 128 : ui8, smin = -128 : si8, smax = 127 : si8>
func.func @test_add_i8_wrap2() -> i8 {
%cst1 = arith.constant 1 : i8
- %0 = test.with_bounds { umin = 0 : i8, umax = 127 : i8, smin = 0 : i8, smax = 127 : i8 } : i8
+ %0 = test.with_bounds < umin = 0 : i8, umax = 127 : i8, smin = 0 : i8, smax = 127 : i8 > : i8
// smax overflow
%1 = arith.addi %0, %cst1 overflow<nuw> : i8
%2 = test.reflect_bounds %1 : i8
@@ -908,10 +908,10 @@ func.func @test_add_i8_wrap2() -> i8 {
}
// CHECK-LABEL: func @test_add_i8_nowrap
-// CHECK: test.reflect_bounds {smax = 127 : si8, smin = 1 : si8, umax = 127 : ui8, umin = 1 : ui8}
+// CHECK: test.reflect_bounds <umin = 1 : ui8, umax = 127 : ui8, smin = 1 : si8, smax = 127 : si8>
func.func @test_add_i8_nowrap() -> i8 {
%cst1 = arith.constant 1 : i8
- %0 = test.with_bounds { umin = 0 : i8, umax = 127 : i8, smin = 0 : i8, smax = 127 : i8 } : i8
+ %0 = test.with_bounds < umin = 0 : i8, umax = 127 : i8, smin = 0 : i8, smax = 127 : i8 > : i8
// nsw flag stops smax from overflowing
%1 = arith.addi %0, %cst1 overflow<nsw> : i8
%2 = test.reflect_bounds %1 : i8
@@ -919,10 +919,10 @@ func.func @test_add_i8_nowrap() -> i8 {
}
// CHECK-LABEL: func @test_sub_i8_wrap1
-// CHECK: test.reflect_bounds {smax = 5 : si8, smin = -10 : si8, umax = 255 : ui8, umin = 0 : ui8} %1 : i8
+// CHECK: test.reflect_bounds <umin = 0 : ui8, umax = 255 : ui8, smin = -10 : si8, smax = 5 : si8> %1 : i8
func.func @test_sub_i8_wrap1() -> i8 {
%cst10 = arith.constant 10 : i8
- %0 = test.with_bounds { umin = 0 : i8, umax = 15 : i8, smin = 0 : i8, smax = 15 : i8 } : i8
+ %0 = test.with_bounds < umin = 0 : i8, umax = 15 : i8, smin = 0 : i8, smax = 15 : i8 > : i8
// umin underflows
%1 = arith.subi %0, %cst10 : i8
%2 = test.reflect_bounds %1 : i8
@@ -930,10 +930,10 @@ func.func @test_sub_i8_wrap1() -> i8 {
}
// CHECK-LABEL: func @test_sub_i8_wrap2
-// CHECK: test.reflect_bounds {smax = 5 : si8, smin = -10 : si8, umax = 255 : ui8, umin = 0 : ui8} %1 : i8
+// CHECK: test.reflect_bounds <umin = 0 : ui8, umax = 255 : ui8, smin = -10 : si8, smax = 5 : si8> %1 : i8
func.func @test_sub_i8_wrap2() -> i8 {
%cst10 = arith.constant 10 : i8
- %0 = test.with_bounds { umin = 0 : i8, umax = 15 : i8, smin = 0 : i8, smax = 15 : i8 } : i8
+ %0 = test.with_bounds < umin = 0 : i8, umax = 15 : i8, smin = 0 : i8, smax = 15 : i8 > : i8
// umin underflows
%1 = arith.subi %0, %cst10 overflow<nsw> : i8
%2 = test.reflect_bounds %1 : i8
@@ -941,10 +941,10 @@ func.func @test_sub_i8_wrap2() -> i8 {
}
// CHECK-LABEL: func @test_sub_i8_nowrap
-// CHECK: test.reflect_bounds {smax = 5 : si8, smin = 0 : si8, umax = 5 : ui8, umin = 0 : ui8}
+// CHECK: test.reflect_bounds <umin = 0 : ui8, umax = 5 : ui8, smin = 0 : si8, smax = 5 : si8>
func.func @test_sub_i8_nowrap() -> i8 {
%cst10 = arith.constant 10 : i8
- %0 = test.with_bounds { umin = 0 : i8, umax = 15 : i8, smin = 0 : i8, smax = 15 : i8 } : i8
+ %0 = test.with_bounds < umin = 0 : i8, umax = 15 : i8, smin = 0 : i8, smax = 15 : i8 > : i8
// nuw flag stops umin from underflowing
%1 = arith.subi %0, %cst10 overflow<nuw> : i8
%2 = test.reflect_bounds %1 : i8
@@ -952,10 +952,10 @@ func.func @test_sub_i8_nowrap() -> i8 {
}
// CHECK-LABEL: func @test_mul_i8_wrap
-// CHECK: test.reflect_bounds {smax = 127 : si8, smin = -128 : si8, umax = 200 : ui8, umin = 100 : ui8}
+// CHECK: test.reflect_bounds <umin = 100 : ui8, umax = 200 : ui8, smin = -128 : si8, smax = 127 : si8>
func.func @test_mul_i8_wrap() -> i8 {
%cst10 = arith.constant 10 : i8
- %0 = test.with_bounds { umin = 10 : i8, umax = 20 : i8, smin = 10 : i8, smax = 20 : i8 } : i8
+ %0 = test.with_bounds < umin = 10 : i8, umax = 20 : i8, smin = 10 : i8, smax = 20 : i8 > : i8
// smax overflows
%1 = arith.muli %0, %cst10 : i8
%2 = test.reflect_bounds %1 : i8
@@ -963,10 +963,10 @@ func.func @test_mul_i8_wrap() -> i8 {
}
// CHECK-LABEL: func @test_mul_i8_nowrap
-// CHECK: test.reflect_bounds {smax = 127 : si8, smin = 100 : si8, umax = 127 : ui8, umin = 100 : ui8}
+// CHECK: test.reflect_bounds <umin = 100 : ui8, umax = 127 : ui8, smin = 100 : si8, smax = 127 : si8>
func.func @test_mul_i8_nowrap() -> i8 {
%cst10 = arith.constant 10 : i8
- %0 = test.with_bounds { umin = 10 : i8, umax = 20 : i8, smin = 10 : i8, smax = 20 : i8 } : i8
+ %0 = test.with_bounds < umin = 10 : i8, umax = 20 : i8, smin = 10 : i8, smax = 20 : i8 > : i8
// nsw stops overflow
%1 = arith.muli %0, %cst10 overflow<nsw> : i8
%2 = test.reflect_bounds %1 : i8
@@ -974,10 +974,10 @@ func.func @test_mul_i8_nowrap() -> i8 {
}
// CHECK-LABEL: func @test_shl_i8_wrap1
-// CHECK: test.reflect_bounds {smax = 127 : si8, smin = -128 : si8, umax = 160 : ui8, umin = 80 : ui8}
+// CHECK: test.reflect_bounds <umin = 80 : ui8, umax = 160 : ui8, smin = -128 : si8, smax = 127 : si8>
func.func @test_shl_i8_wrap1() -> i8 {
%cst3 = arith.constant 3 : i8
- %0 = test.with_bounds { umin = 10 : i8, umax = 20 : i8, smin = 10 : i8, smax = 20 : i8 } : i8
+ %0 = test.with_bounds < umin = 10 : i8, umax = 20 : i8, smin = 10 : i8, smax = 20 : i8 > : i8
// smax overflows
%1 = arith.shli %0, %cst3 : i8
%2 = test.reflect_bounds %1 : i8
@@ -985,10 +985,10 @@ func.func @test_shl_i8_wrap1() -> i8 {
}
// CHECK-LABEL: func @test_shl_i8_wrap2
-// CHECK: test.reflect_bounds {smax = 127 : si8, smin = -128 : si8, umax = 160 : ui8, umin = 80 : ui8}
+// CHECK: test.reflect_bounds <umin = 80 : ui8, umax = 160 : ui8, smin = -128 : si8, smax = 127 : si8>
func.func @test_shl_i8_wrap2() -> i8 {
%cst3 = arith.constant 3 : i8
- %0 = test.with_bounds { umin = 10 : i8, umax = 20 : i8, smin = 10 : i8, smax = 20 : i8 } : i8
+ %0 = test.with_bounds < umin = 10 : i8, umax = 20 : i8, smin = 10 : i8, smax = 20 : i8 > : i8
// smax overflows
%1 = arith.shli %0, %cst3 overflow<nuw> : i8
%2 = test.reflect_bounds %1 : i8
@@ -996,10 +996,10 @@ func.func @test_shl_i8_wrap2() -> i8 {
}
// CHECK-LABEL: func @test_shl_i8_nowrap
-// CHECK: test.reflect_bounds {smax = 127 : si8, smin = 80 : si8, umax = 127 : ui8, umin = 80 : ui8}
+// CHECK: test.reflect_bounds <umin = 80 : ui8, umax = 127 : ui8, smin = 80 : si8, smax = 127 : si8>
func.func @test_shl_i8_nowrap() -> i8 {
%cst3 = arith.constant 3 : i8
- %0 = test.with_bounds { umin = 10 : i8, umax = 20 : ui8, smin = 10 : i8, smax = 20 : i8 } : i8
+ %0 = test.with_bounds < umin = 10 : i8, umax = 20 : ui8, smin = 10 : i8, smax = 20 : i8 > : i8
// nsw stops smax overflow
%1 = arith.shli %0, %cst3 overflow<nsw> : i8
%2 = test.reflect_bounds %1 : i8
@@ -1013,7 +1013,7 @@ func.func @test_shl_i8_nowrap() -> i8 {
/// though it has an integer valued result.
// CHECK-LABEL: func @test_cmpf_propagates
-// CHECK: test.reflect_bounds {smax = 2 : index, smin = 1 : index, umax = 2 : index, umin = 1 : index}
+// CHECK: test.reflect_bounds <umin = 1 : index, umax = 2 : index, smin = 1 : index, smax = 2 : index>
func.func @test_cmpf_propagates(%a: f32, %b: f32) -> index {
%c1 = arith.constant 1 : index
%c2 = arith.constant 2 : index
diff --git a/mlir/test/Dialect/Arith/int-range-loop-iter-args.mlir b/mlir/test/Dialect/Arith/int-range-loop-iter-args.mlir
index 24801875e257c..2d0575f756531 100644
--- a/mlir/test/Dialect/Arith/int-range-loop-iter-args.mlir
+++ b/mlir/test/Dialect/Arith/int-range-loop-iter-args.mlir
@@ -8,7 +8,7 @@
// without being widened to `[INT_MIN, INT_MAX]`.
// CHECK-LABEL: func @bounded_acc_for
-// CHECK: test.reflect_bounds {smax = 10 : si32, smin = 0 : si32, umax = 10 : ui32, umin = 0 : ui32}
+// CHECK: test.reflect_bounds <umin = 0 : ui32, umax = 10 : ui32, smin = 0 : si32, smax = 10 : si32>
func.func @bounded_acc_for(%n: i32) -> i32 {
%c0 = arith.constant 0 : i32
%c1 = arith.constant 1 : i32
@@ -28,7 +28,7 @@ func.func @bounded_acc_for(%n: i32) -> i32 {
// CHECK-LABEL: func @bounded_acc_while
// CHECK: %[[TRUE:.*]] = arith.constant true
// CHECK: scf.condition(%[[TRUE]])
-// CHECK: test.reflect_bounds {smax = 10 : si32, smin = 0 : si32, umax = 10 : ui32, umin = 0 : ui32}
+// CHECK: test.reflect_bounds <umin = 0 : ui32, umax = 10 : ui32, smin = 0 : si32, smax = 10 : si32>
func.func @bounded_acc_while() -> i32 {
%c0 = arith.constant 0 : i32
%c1 = arith.constant 1 : i32
@@ -48,7 +48,7 @@ func.func @bounded_acc_while() -> i32 {
}
// CHECK-LABEL: func @bounded_mask_for
-// CHECK: test.reflect_bounds {smax = 15 : si32, smin = 0 : si32, umax = 15 : ui32, umin = 0 : ui32}
+// CHECK: test.reflect_bounds <umin = 0 : ui32, umax = 15 : ui32, smin = 0 : si32, smax = 15 : si32>
func.func @bounded_mask_for(%n: i32) -> i32 {
%c0 = arith.constant 0 : i32
%c1 = arith.constant 1 : i32
diff --git a/mlir/test/Dialect/Arith/int-range-narrowing.mlir b/mlir/test/Dialect/Arith/int-range-narrowing.mlir
index 06ca5d19de9c9..780e4a918d159 100644
--- a/mlir/test/Dialect/Arith/int-range-narrowing.mlir
+++ b/mlir/test/Dialect/Arith/int-range-narrowing.mlir
@@ -6,75 +6,75 @@
// Truncate possibly-negative values in a signed way
// CHECK-LABEL: func @test_addi_neg
-// CHECK: %[[POS:.*]] = test.with_bounds {smax = 1 : index, smin = 0 : index, umax = 1 : index, umin = 0 : index} : index
-// CHECK: %[[NEG:.*]] = test.with_bounds {smax = 0 : index, smin = -1 : index, umax = -1 : index, umin = 0 : index} : index
+// CHECK: %[[POS:.*]] = test.with_bounds <umin = 0 : index, umax = 1 : index, smin = 0 : index, smax = 1 : index> : index
+// CHECK: %[[NEG:.*]] = test.with_bounds <umin = 0 : index, umax = -1 : index, smin = -1 : index, smax = 0 : index> : index
// CHECK: %[[POS_I8:.*]] = arith.index_castui %[[POS]] : index to i8
// CHECK: %[[NEG_I8:.*]] = arith.index_cast %[[NEG]] : index to i8
// CHECK: %[[RES_I8:.*]] = arith.addi %[[POS_I8]], %[[NEG_I8]] : i8
// CHECK: %[[RES:.*]] = arith.index_cast %[[RES_I8]] : i8 to index
// CHECK: return %[[RES]] : index
func.func @test_addi_neg() -> index {
- %0 = test.with_bounds { umin = 0 : index, umax = 1 : index, smin = 0 : index, smax = 1 : index } : index
- %1 = test.with_bounds { umin = 0 : index, umax = -1 : index, smin = -1 : index, smax = 0 : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 1 : index, smin = 0 : index, smax = 1 : index > : index
+ %1 = test.with_bounds < umin = 0 : index, umax = -1 : index, smin = -1 : index, smax = 0 : index > : index
%2 = arith.addi %0, %1 : index
return %2 : index
}
// CHECK-LABEL: func @test_addi
-// CHECK: %[[A:.*]] = test.with_bounds {smax = 5 : index, smin = 4 : index, umax = 5 : index, umin = 4 : index} : index
-// CHECK: %[[B:.*]] = test.with_bounds {smax = 7 : index, smin = 6 : index, umax = 7 : index, umin = 6 : index} : index
+// CHECK: %[[A:.*]] = test.with_bounds <umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index> : index
+// CHECK: %[[B:.*]] = test.with_bounds <umin = 6 : index, umax = 7 : index, smin = 6 : index, smax = 7 : index> : index
// CHECK: %[[A_CASTED:.*]] = arith.index_castui %[[A]] : index to i8
// CHECK: %[[B_CASTED:.*]] = arith.index_castui %[[B]] : index to i8
// CHECK: %[[RES:.*]] = arith.addi %[[A_CASTED]], %[[B_CASTED]] : i8
// CHECK: %[[RES_CASTED:.*]] = arith.index_castui %[[RES]] : i8 to index
// CHECK: return %[[RES_CASTED]] : index
func.func @test_addi() -> index {
- %0 = test.with_bounds { umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index } : index
- %1 = test.with_bounds { umin = 6 : index, umax = 7 : index, smin = 6 : index, smax = 7 : index } : index
+ %0 = test.with_bounds < umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index > : index
+ %1 = test.with_bounds < umin = 6 : index, umax = 7 : index, smin = 6 : index, smax = 7 : index > : index
%2 = arith.addi %0, %1 : index
return %2 : index
}
// CHECK-LABEL: func @test_addi_vec
-// CHECK: %[[A:.*]] = test.with_bounds {smax = 5 : index, smin = 4 : index, umax = 5 : index, umin = 4 : index} : vector<4xindex>
-// CHECK: %[[B:.*]] = test.with_bounds {smax = 7 : index, smin = 6 : index, umax = 7 : index, umin = 6 : index} : vector<4xindex>
+// CHECK: %[[A:.*]] = test.with_bounds <umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index> : vector<4xindex>
+// CHECK: %[[B:.*]] = test.with_bounds <umin = 6 : index, umax = 7 : index, smin = 6 : index, smax = 7 : index> : vector<4xindex>
// CHECK: %[[A_CASTED:.*]] = arith.index_castui %[[A]] : vector<4xindex> to vector<4xi8>
// CHECK: %[[B_CASTED:.*]] = arith.index_castui %[[B]] : vector<4xindex> to vector<4xi8>
// CHECK: %[[RES:.*]] = arith.addi %[[A_CASTED]], %[[B_CASTED]] : vector<4xi8>
// CHECK: %[[RES_CASTED:.*]] = arith.index_castui %[[RES]] : vector<4xi8> to vector<4xindex>
// CHECK: return %[[RES_CASTED]] : vector<4xindex>
func.func @test_addi_vec() -> vector<4xindex> {
- %0 = test.with_bounds { umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index } : vector<4xindex>
- %1 = test.with_bounds { umin = 6 : index, umax = 7 : index, smin = 6 : index, smax = 7 : index } : vector<4xindex>
+ %0 = test.with_bounds < umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index > : vector<4xindex>
+ %1 = test.with_bounds < umin = 6 : index, umax = 7 : index, smin = 6 : index, smax = 7 : index > : vector<4xindex>
%2 = arith.addi %0, %1 : vector<4xindex>
return %2 : vector<4xindex>
}
// CHECK-LABEL: func @test_addi_i64
-// CHECK: %[[A:.*]] = test.with_bounds {smax = 5 : i64, smin = 4 : i64, umax = 5 : i64, umin = 4 : i64} : i64
-// CHECK: %[[B:.*]] = test.with_bounds {smax = 7 : i64, smin = 6 : i64, umax = 7 : i64, umin = 6 : i64} : i64
+// CHECK: %[[A:.*]] = test.with_bounds <umin = 4 : i64, umax = 5 : i64, smin = 4 : i64, smax = 5 : i64> : i64
+// CHECK: %[[B:.*]] = test.with_bounds <umin = 6 : i64, umax = 7 : i64, smin = 6 : i64, smax = 7 : i64> : i64
// CHECK: %[[A_CASTED:.*]] = arith.trunci %[[A]] : i64 to i8
// CHECK: %[[B_CASTED:.*]] = arith.trunci %[[B]] : i64 to i8
// CHECK: %[[RES:.*]] = arith.addi %[[A_CASTED]], %[[B_CASTED]] : i8
// CHECK: %[[RES_CASTED:.*]] = arith.extui %[[RES]] : i8 to i64
// CHECK: return %[[RES_CASTED]] : i64
func.func @test_addi_i64() -> i64 {
- %0 = test.with_bounds { umin = 4 : i64, umax = 5 : i64, smin = 4 : i64, smax = 5 : i64 } : i64
- %1 = test.with_bounds { umin = 6 : i64, umax = 7 : i64, smin = 6 : i64, smax = 7 : i64 } : i64
+ %0 = test.with_bounds < umin = 4 : i64, umax = 5 : i64, smin = 4 : i64, smax = 5 : i64 > : i64
+ %1 = test.with_bounds < umin = 6 : i64, umax = 7 : i64, smin = 6 : i64, smax = 7 : i64 > : i64
%2 = arith.addi %0, %1 : i64
return %2 : i64
}
// CHECK-LABEL: func @test_cmpi
-// CHECK: %[[A:.*]] = test.with_bounds {smax = 10 : index, smin = 0 : index, umax = 10 : index, umin = 0 : index} : index
-// CHECK: %[[B:.*]] = test.with_bounds {smax = 10 : index, smin = 0 : index, umax = 10 : index, umin = 0 : index} : index
+// CHECK: %[[A:.*]] = test.with_bounds <umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index> : index
+// CHECK: %[[B:.*]] = test.with_bounds <umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index> : index
// CHECK: %[[A_CASTED:.*]] = arith.index_castui %[[A]] : index to i8
// CHECK: %[[B_CASTED:.*]] = arith.index_castui %[[B]] : index to i8
// CHECK: %[[RES:.*]] = arith.cmpi slt, %[[A_CASTED]], %[[B_CASTED]] : i8
// CHECK: return %[[RES]] : i1
func.func @test_cmpi() -> i1 {
- %0 = test.with_bounds { umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index } : index
- %1 = test.with_bounds { umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index > : index
+ %1 = test.with_bounds < umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index > : index
%2 = arith.cmpi slt, %0, %1 : index
return %2 : i1
}
@@ -84,11 +84,11 @@ func.func @test_cmpi() -> i1 {
// CHECK-NOT: arith.cmpi sgt, {{.*}} : i32
// CHECK-NOT: arith.cmpi sle, {{.*}} : i32
// CHECK-NOT: arith.cmpi sge, {{.*}} : i32
-// CHECK: %[[A:.*]] = test.with_bounds {smax = 4292870144 : index, smin = 0 : index, umax = 4292870144 : index, umin = 0 : index} : index
-// CHECK: %[[B:.*]] = test.with_bounds {smax = 4292870144 : index, smin = 0 : index, umax = 4292870144 : index, umin = 0 : index} : index
+// CHECK: %[[A:.*]] = test.with_bounds <umin = 0 : index, umax = 4292870144 : index, smin = 0 : index, smax = 4292870144 : index> : index
+// CHECK: %[[B:.*]] = test.with_bounds <umin = 0 : index, umax = 4292870144 : index, smin = 0 : index, smax = 4292870144 : index> : index
// CHECK: %[[SLT:.*]] = arith.cmpi slt, %[[B]], %[[A]] : index
-// CHECK: %[[C:.*]] = test.with_bounds {smax = 2147483648 : index, smin = 0 : index, umax = 2147483648 : index, umin = 0 : index} : index
-// CHECK: %[[ZERO:.*]] = test.with_bounds {smax = 0 : index, smin = 0 : index, umax = 0 : index, umin = 0 : index} : index
+// CHECK: %[[C:.*]] = test.with_bounds <umin = 0 : index, umax = 2147483648 : index, smin = 0 : index, smax = 2147483648 : index> : index
+// CHECK: %[[ZERO:.*]] = test.with_bounds <umin = 0 : index, umax = 0 : index, smin = 0 : index, smax = 0 : index> : index
// CHECK: %[[SGT:.*]] = arith.cmpi sgt, %[[C]], %[[ZERO]] : index
// CHECK: %[[SLE:.*]] = arith.cmpi sle, %[[A]], %[[ZERO]] : index
// CHECK: %[[SGE:.*]] = arith.cmpi sge, %[[C]], %[[ZERO]] : index
@@ -97,11 +97,11 @@ func.func @test_cmpi() -> i1 {
// CHECK: %[[AND2:.*]] = arith.andi %[[AND0]], %[[AND1]] : i1
// CHECK: return %[[AND2]] : i1
func.func @test_cmpi_si_pred_out_of_signed_bounds() -> i1 {
- %0 = test.with_bounds { umin = 0 : index, umax = 4292870144 : index, smin = 0 : index, smax = 4292870144 : index } : index
- %1 = test.with_bounds { umin = 0 : index, umax = 4292870144 : index, smin = 0 : index, smax = 4292870144 : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 4292870144 : index, smin = 0 : index, smax = 4292870144 : index > : index
+ %1 = test.with_bounds < umin = 0 : index, umax = 4292870144 : index, smin = 0 : index, smax = 4292870144 : index > : index
%2 = arith.cmpi slt, %1, %0 : index
- %3 = test.with_bounds { umin = 0 : index, umax = 2147483648 : index, smin = 0 : index, smax = 2147483648 : index } : index
- %4 = test.with_bounds { umin = 0 : index, umax = 0 : index, smin = 0 : index, smax = 0 : index } : index
+ %3 = test.with_bounds < umin = 0 : index, umax = 2147483648 : index, smin = 0 : index, smax = 2147483648 : index > : index
+ %4 = test.with_bounds < umin = 0 : index, umax = 0 : index, smin = 0 : index, smax = 0 : index > : index
%5 = arith.cmpi sgt, %3, %4 : index
%6 = arith.cmpi sle, %0, %4 : index
%7 = arith.cmpi sge, %3, %4 : index
@@ -112,37 +112,37 @@ func.func @test_cmpi_si_pred_out_of_signed_bounds() -> i1 {
}
// CHECK-LABEL: func @test_cmpi_ui_pred_out_of_signed_bounds
-// CHECK: %[[A:.*]] = test.with_bounds {smax = 4292870144 : index, smin = 0 : index, umax = 4292870144 : index, umin = 0 : index} : index
-// CHECK: %[[B:.*]] = test.with_bounds {smax = 4292870144 : index, smin = 0 : index, umax = 4292870144 : index, umin = 0 : index} : index
+// CHECK: %[[A:.*]] = test.with_bounds <umin = 0 : index, umax = 4292870144 : index, smin = 0 : index, smax = 4292870144 : index> : index
+// CHECK: %[[B:.*]] = test.with_bounds <umin = 0 : index, umax = 4292870144 : index, smin = 0 : index, smax = 4292870144 : index> : index
// CHECK: %[[A_I32:.*]] = arith.index_castui %[[A]] : index to i32
// CHECK: %[[B_I32:.*]] = arith.index_castui %[[B]] : index to i32
// CHECK: %[[RES:.*]] = arith.cmpi ult, %[[A_I32]], %[[B_I32]] : i32
// CHECK: return %[[RES]] : i1
func.func @test_cmpi_ui_pred_out_of_signed_bounds() -> i1 {
- %0 = test.with_bounds { umin = 0 : index, umax = 4292870144 : index, smin = 0 : index, smax = 4292870144 : index } : index
- %1 = test.with_bounds { umin = 0 : index, umax = 4292870144 : index, smin = 0 : index, smax = 4292870144 : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 4292870144 : index, smin = 0 : index, smax = 4292870144 : index > : index
+ %1 = test.with_bounds < umin = 0 : index, umax = 4292870144 : index, smin = 0 : index, smax = 4292870144 : index > : index
%2 = arith.cmpi ult, %0, %1 : index
return %2 : i1
}
// CHECK-LABEL: func @test_cmpi_vec
-// CHECK: %[[A:.*]] = test.with_bounds {smax = 10 : index, smin = 0 : index, umax = 10 : index, umin = 0 : index} : vector<4xindex>
-// CHECK: %[[B:.*]] = test.with_bounds {smax = 10 : index, smin = 0 : index, umax = 10 : index, umin = 0 : index} : vector<4xindex>
+// CHECK: %[[A:.*]] = test.with_bounds <umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index> : vector<4xindex>
+// CHECK: %[[B:.*]] = test.with_bounds <umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index> : vector<4xindex>
// CHECK: %[[A_CASTED:.*]] = arith.index_castui %[[A]] : vector<4xindex> to vector<4xi8>
// CHECK: %[[B_CASTED:.*]] = arith.index_castui %[[B]] : vector<4xindex> to vector<4xi8>
// CHECK: %[[RES:.*]] = arith.cmpi slt, %[[A_CASTED]], %[[B_CASTED]] : vector<4xi8>
// CHECK: return %[[RES]] : vector<4xi1>
func.func @test_cmpi_vec() -> vector<4xi1> {
- %0 = test.with_bounds { umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index } : vector<4xindex>
- %1 = test.with_bounds { umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index } : vector<4xindex>
+ %0 = test.with_bounds < umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index > : vector<4xindex>
+ %1 = test.with_bounds < umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index > : vector<4xindex>
%2 = arith.cmpi slt, %0, %1 : vector<4xindex>
return %2 : vector<4xi1>
}
// CHECK-LABEL: func @test_add_cmpi
-// CHECK: %[[A:.*]] = test.with_bounds {smax = 10 : index, smin = 0 : index, umax = 10 : index, umin = 0 : index} : index
-// CHECK: %[[B:.*]] = test.with_bounds {smax = 10 : index, smin = 0 : index, umax = 10 : index, umin = 0 : index} : index
-// CHECK: %[[C:.*]] = test.with_bounds {smax = 10 : index, smin = 0 : index, umax = 10 : index, umin = 0 : index} : index
+// CHECK: %[[A:.*]] = test.with_bounds <umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index> : index
+// CHECK: %[[B:.*]] = test.with_bounds <umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index> : index
+// CHECK: %[[C:.*]] = test.with_bounds <umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index> : index
// CHECK: %[[A_CASTED:.*]] = arith.index_castui %[[A]] : index to i8
// CHECK: %[[B_CASTED:.*]] = arith.index_castui %[[B]] : index to i8
// CHECK: %[[RES1:.*]] = arith.addi %[[A_CASTED]], %[[B_CASTED]] : i8
@@ -150,18 +150,18 @@ func.func @test_cmpi_vec() -> vector<4xi1> {
// CHECK: %[[RES2:.*]] = arith.cmpi slt, %[[C_CASTED]], %[[RES1]] : i8
// CHECK: return %[[RES2]] : i1
func.func @test_add_cmpi() -> i1 {
- %0 = test.with_bounds { umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index } : index
- %1 = test.with_bounds { umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index } : index
- %3 = test.with_bounds { umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index > : index
+ %1 = test.with_bounds < umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index > : index
+ %3 = test.with_bounds < umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index > : index
%4 = arith.addi %0, %1 : index
%5 = arith.cmpi slt, %3, %4 : index
return %5 : i1
}
// CHECK-LABEL: func @test_add_cmpi_i64
-// CHECK: %[[A:.*]] = test.with_bounds {smax = 10 : i64, smin = 0 : i64, umax = 10 : i64, umin = 0 : i64} : i64
-// CHECK: %[[B:.*]] = test.with_bounds {smax = 10 : i64, smin = 0 : i64, umax = 10 : i64, umin = 0 : i64} : i64
-// CHECK: %[[C:.*]] = test.with_bounds {smax = 10 : i64, smin = 0 : i64, umax = 10 : i64, umin = 0 : i64} : i64
+// CHECK: %[[A:.*]] = test.with_bounds <umin = 0 : i64, umax = 10 : i64, smin = 0 : i64, smax = 10 : i64> : i64
+// CHECK: %[[B:.*]] = test.with_bounds <umin = 0 : i64, umax = 10 : i64, smin = 0 : i64, smax = 10 : i64> : i64
+// CHECK: %[[C:.*]] = test.with_bounds <umin = 0 : i64, umax = 10 : i64, smin = 0 : i64, smax = 10 : i64> : i64
// CHECK: %[[A_CASTED:.*]] = arith.trunci %[[A]] : i64 to i8
// CHECK: %[[B_CASTED:.*]] = arith.trunci %[[B]] : i64 to i8
// CHECK: %[[RES1:.*]] = arith.addi %[[A_CASTED]], %[[B_CASTED]] : i8
@@ -169,9 +169,9 @@ func.func @test_add_cmpi() -> i1 {
// CHECK: %[[RES2:.*]] = arith.cmpi slt, %[[C_CASTED]], %[[RES1]] : i8
// CHECK: return %[[RES2]] : i1
func.func @test_add_cmpi_i64() -> i1 {
- %0 = test.with_bounds { umin = 0 : i64, umax = 10 : i64, smin = 0 : i64, smax = 10 : i64 } : i64
- %1 = test.with_bounds { umin = 0 : i64, umax = 10 : i64, smin = 0 : i64, smax = 10 : i64 } : i64
- %3 = test.with_bounds { umin = 0 : i64, umax = 10 : i64, smin = 0 : i64, smax = 10 : i64 } : i64
+ %0 = test.with_bounds < umin = 0 : i64, umax = 10 : i64, smin = 0 : i64, smax = 10 : i64 > : i64
+ %1 = test.with_bounds < umin = 0 : i64, umax = 10 : i64, smin = 0 : i64, smax = 10 : i64 > : i64
+ %3 = test.with_bounds < umin = 0 : i64, umax = 10 : i64, smin = 0 : i64, smax = 10 : i64 > : i64
%4 = arith.addi %0, %1 : i64
%5 = arith.cmpi slt, %3, %4 : i64
return %5 : i1
@@ -287,9 +287,9 @@ func.func @subi_mixed_ext_i8(%lhs: i8, %rhs: i8) -> i32 {
// CHECK: %[[SHR_I64:.*]] = arith.shrsi {{.*}} : i64
// CHECK: return %{{.*}} : i64, i64, i64, i64, i64, i64, i64
func.func @signed_ops_out_of_narrowed_signed_range() -> (i64, i64, i64, i64, i64, i64, i64) {
- %0 = test.with_bounds { umin = 0 : i64, umax = 4292870144 : i64, smin = 0 : i64, smax = 4292870144 : i64 } : i64
- %1 = test.with_bounds { umin = 1 : i64, umax = 8 : i64, smin = 1 : i64, smax = 8 : i64 } : i64
- %2 = test.with_bounds { umin = 0 : i64, umax = 0 : i64, smin = 0 : i64, smax = 0 : i64 } : i64
+ %0 = test.with_bounds < umin = 0 : i64, umax = 4292870144 : i64, smin = 0 : i64, smax = 4292870144 : i64 > : i64
+ %1 = test.with_bounds < umin = 1 : i64, umax = 8 : i64, smin = 1 : i64, smax = 8 : i64 > : i64
+ %2 = test.with_bounds < umin = 0 : i64, umax = 0 : i64, smin = 0 : i64, smax = 0 : i64 > : i64
%3 = arith.divsi %0, %1 : i64
%4 = arith.ceildivsi %0, %1 : i64
%5 = arith.floordivsi %0, %1 : i64
@@ -309,9 +309,9 @@ func.func @signed_ops_out_of_narrowed_signed_range() -> (i64, i64, i64, i64, i64
// CHECK: arith.shrui {{.*}} : i32
// CHECK: return %{{.*}} : i64, i64, i64, i64, i64, i64
func.func @unsigned_ops_out_of_narrowed_signed_range() -> (i64, i64, i64, i64, i64, i64) {
- %0 = test.with_bounds { umin = 0 : i64, umax = 4292870144 : i64, smin = 0 : i64, smax = 4292870144 : i64 } : i64
- %1 = test.with_bounds { umin = 1 : i64, umax = 8 : i64, smin = 1 : i64, smax = 8 : i64 } : i64
- %2 = test.with_bounds { umin = 0 : i64, umax = 0 : i64, smin = 0 : i64, smax = 0 : i64 } : i64
+ %0 = test.with_bounds < umin = 0 : i64, umax = 4292870144 : i64, smin = 0 : i64, smax = 4292870144 : i64 > : i64
+ %1 = test.with_bounds < umin = 1 : i64, umax = 8 : i64, smin = 1 : i64, smax = 8 : i64 > : i64
+ %2 = test.with_bounds < umin = 0 : i64, umax = 0 : i64, smin = 0 : i64, smax = 0 : i64 > : i64
%3 = arith.divui %0, %1 : i64
%4 = arith.ceildivui %0, %1 : i64
%5 = arith.remui %0, %1 : i64
@@ -331,8 +331,8 @@ func.func @unsigned_ops_out_of_narrowed_signed_range() -> (i64, i64, i64, i64, i
// CHECK: arith.shrsi {{.*}} : i64
// CHECK-NOT: arith.shrsi {{.*}} : i32
func.func @shrsi_amount_out_of_narrowed_range() -> i64 {
- %0 = test.with_bounds { umin = 0 : i64, umax = 4 : i64, smin = 0 : i64, smax = 4 : i64 } : i64
- %1 = test.with_bounds { umin = 0 : i64, umax = 63 : i64, smin = 0 : i64, smax = 63 : i64 } : i64
+ %0 = test.with_bounds < umin = 0 : i64, umax = 4 : i64, smin = 0 : i64, smax = 4 : i64 > : i64
+ %1 = test.with_bounds < umin = 0 : i64, umax = 63 : i64, smin = 0 : i64, smax = 63 : i64 > : i64
%2 = arith.shrsi %0, %1 : i64
return %2 : i64
}
@@ -342,8 +342,8 @@ func.func @shrsi_amount_out_of_narrowed_range() -> i64 {
// CHECK-LABEL: func.func @shrui_amount_in_narrowed_range
// CHECK: arith.shrui {{.*}} : i32
func.func @shrui_amount_in_narrowed_range() -> i64 {
- %0 = test.with_bounds { umin = 0 : i64, umax = 4 : i64, smin = 0 : i64, smax = 4 : i64 } : i64
- %1 = test.with_bounds { umin = 0 : i64, umax = 31 : i64, smin = 0 : i64, smax = 31 : i64 } : i64
+ %0 = test.with_bounds < umin = 0 : i64, umax = 4 : i64, smin = 0 : i64, smax = 4 : i64 > : i64
+ %1 = test.with_bounds < umin = 0 : i64, umax = 31 : i64, smin = 0 : i64, smax = 31 : i64 > : i64
%2 = arith.shrui %0, %1 : i64
return %2 : i64
}
@@ -465,8 +465,8 @@ func.func @clamp_to_loop_bound_and_id() {
%c16 = arith.constant 16 : index
%c64 = arith.constant 64 : index
- %tid = test.with_bounds {smin = 0 : index, smax = 63 : index, umin = 0 : index, umax = 63 : index} : index
- %bound = test.with_bounds {smin = 16 : index, smax = 112 : index, umin = 16 : index, umax = 112 : index} : index
+ %tid = test.with_bounds <smin = 0 : index, smax = 63 : index, umin = 0 : index, umax = 63 : index> : index
+ %bound = test.with_bounds <smin = 16 : index, smax = 112 : index, umin = 16 : index, umax = 112 : index> : index
scf.for %arg0 = %c16 to %bound step %c64 {
%0 = arith.subi %bound, %arg0 : index
%1 = arith.minsi %0, %c64 : index
@@ -490,8 +490,8 @@ func.func @loop_with_iter_arg() {
%cst = arith.constant dense<0.000000e+00> : vector<4xf32>
-// CHECK: %[[POS:.*]] = test.with_bounds {smax = 1 : index, smin = 0 : index, umax = 1 : index, umin = 0 : index} : index
-// CHECK: %[[NEG:.*]] = test.with_bounds {smax = 0 : index, smin = -1 : index, umax = -1 : index, umin = 0 : index} : index
+// CHECK: %[[POS:.*]] = test.with_bounds <umin = 0 : index, umax = 1 : index, smin = 0 : index, smax = 1 : index> : index
+// CHECK: %[[NEG:.*]] = test.with_bounds <umin = 0 : index, umax = -1 : index, smin = -1 : index, smax = 0 : index> : index
// Check iter args are still present
// CHECK: scf.for {{.*}} iter_args({{.*}})
// CHECK: %[[POS_I8:.*]] = arith.index_castui %[[POS]] : index to i8
@@ -500,8 +500,8 @@ func.func @loop_with_iter_arg() {
// CHECK: %[[RES:.*]] = arith.index_cast %[[RES_I8]] : i8 to index
// CHECK: call @use(%[[RES]])
- %0 = test.with_bounds { umin = 0 : index, umax = 1 : index, smin = 0 : index, smax = 1 : index } : index
- %1 = test.with_bounds { umin = 0 : index, umax = -1 : index, smin = -1 : index, smax = 0 : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 1 : index, smin = 0 : index, smax = 1 : index > : index
+ %1 = test.with_bounds < umin = 0 : index, umax = -1 : index, smin = -1 : index, smax = 0 : index > : index
%res = scf.for %arg0 = %c0 to %c16 step %c1 iter_args(%arg1 = %cst) -> (vector<4xf32>) {
%2 = arith.addi %0, %1 : index
func.call @use(%2) : (index) -> ()
@@ -521,12 +521,12 @@ func.func @narrow_loop_bounds() {
%c1_i64 = arith.constant 1 : i64
// CHECK-DAG: %[[C1_I8:.*]] = arith.constant 1 : i8
- // CHECK-DAG: %[[LB:.*]] = test.with_bounds {smax = 0 : i64, smin = 0 : i64, umax = 0 : i64, umin = 0 : i64} : i64
- // CHECK-DAG: %[[UB:.*]] = test.with_bounds {smax = 10 : i64, smin = 10 : i64, umax = 10 : i64, umin = 10 : i64} : i64
- // CHECK-DAG: %[[STEP:.*]] = test.with_bounds {smax = 1 : i64, smin = 1 : i64, umax = 1 : i64, umin = 1 : i64} : i64
- %lb = test.with_bounds {smin = 0 : i64, smax = 0 : i64, umin = 0 : i64, umax = 0 : i64} : i64
- %ub = test.with_bounds {smin = 10 : i64, smax = 10 : i64, umin = 10 : i64, umax = 10 : i64} : i64
- %step = test.with_bounds {smin = 1 : i64, smax = 1 : i64, umin = 1 : i64, umax = 1 : i64} : i64
+ // CHECK-DAG: %[[LB:.*]] = test.with_bounds <umin = 0 : i64, umax = 0 : i64, smin = 0 : i64, smax = 0 : i64> : i64
+ // CHECK-DAG: %[[UB:.*]] = test.with_bounds <umin = 10 : i64, umax = 10 : i64, smin = 10 : i64, smax = 10 : i64> : i64
+ // CHECK-DAG: %[[STEP:.*]] = test.with_bounds <umin = 1 : i64, umax = 1 : i64, smin = 1 : i64, smax = 1 : i64> : i64
+ %lb = test.with_bounds <smin = 0 : i64, smax = 0 : i64, umin = 0 : i64, umax = 0 : i64> : i64
+ %ub = test.with_bounds <smin = 10 : i64, smax = 10 : i64, umin = 10 : i64, umax = 10 : i64> : i64
+ %step = test.with_bounds <smin = 1 : i64, smax = 1 : i64, umin = 1 : i64, umax = 1 : i64> : i64
// CHECK-DAG: %[[LB_I8:.*]] = arith.trunci %[[LB]] : i64 to i8
// CHECK-DAG: %[[UB_I8:.*]] = arith.trunci %[[UB]] : i64 to i8
diff --git a/mlir/test/Dialect/Arith/int-range-opts-bug-119045.mlir b/mlir/test/Dialect/Arith/int-range-opts-bug-119045.mlir
index a2cf72a00e019..9b9f3a78659c8 100644
--- a/mlir/test/Dialect/Arith/int-range-opts-bug-119045.mlir
+++ b/mlir/test/Dialect/Arith/int-range-opts-bug-119045.mlir
@@ -15,7 +15,7 @@ func.func @blocks_prematurely_declared_dead_bug(%mem: memref<?xf16>) {
%c0 = arith.constant 0 : index
%c64 = arith.constant 64 : index
%thread_id_x = gpu.thread_id x upper_bound 64
- %6 = test.with_bounds { smin = 16 : index, smax = 112 : index, umin = 16 : index, umax = 112 : index } : index
+ %6 = test.with_bounds < smin = 16 : index, smax = 112 : index, umin = 16 : index, umax = 112 : index > : index
%8 = arith.divui %6, %c16 : index
%9 = arith.muli %8, %c16 : index
cf.br ^bb1(%c0 : index)
diff --git a/mlir/test/Dialect/Arith/int-range-opts.mlir b/mlir/test/Dialect/Arith/int-range-opts.mlir
index 0cfbdf3e33c93..ff6292ed35635 100644
--- a/mlir/test/Dialect/Arith/int-range-opts.mlir
+++ b/mlir/test/Dialect/Arith/int-range-opts.mlir
@@ -5,7 +5,7 @@
// CHECK: return %[[C]]
func.func @test() -> i1 {
%cst1 = arith.constant -1 : index
- %0 = test.with_bounds { umin = 0 : index, umax = 0x7fffffffffffffff : index, smin = 0 : index, smax = 0x7fffffffffffffff : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 0x7fffffffffffffff : index, smin = 0 : index, smax = 0x7fffffffffffffff : index > : index
%1 = arith.cmpi eq, %0, %cst1 : index
return %1: i1
}
@@ -17,7 +17,7 @@ func.func @test() -> i1 {
// CHECK: return %[[C]]
func.func @test() -> i1 {
%cst1 = arith.constant -1 : index
- %0 = test.with_bounds { umin = 0 : index, umax = 0x7fffffffffffffff : index, smin = 0 : index, smax = 0x7fffffffffffffff : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 0x7fffffffffffffff : index, smin = 0 : index, smax = 0x7fffffffffffffff : index > : index
%1 = arith.cmpi ne, %0, %cst1 : index
return %1: i1
}
@@ -30,7 +30,7 @@ func.func @test() -> i1 {
// CHECK: return %[[C]]
func.func @test() -> i1 {
%cst = arith.constant 0 : index
- %0 = test.with_bounds { umin = 0 : index, umax = 0x7fffffffffffffff : index, smin = 0 : index, smax = 0x7fffffffffffffff : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 0x7fffffffffffffff : index, smin = 0 : index, smax = 0x7fffffffffffffff : index > : index
%1 = arith.cmpi sge, %0, %cst : index
return %1: i1
}
@@ -42,7 +42,7 @@ func.func @test() -> i1 {
// CHECK: return %[[C]]
func.func @test() -> i1 {
%cst = arith.constant 0 : index
- %0 = test.with_bounds { umin = 0 : index, umax = 0x7fffffffffffffff : index, smin = 0 : index, smax = 0x7fffffffffffffff : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 0x7fffffffffffffff : index, smin = 0 : index, smax = 0x7fffffffffffffff : index > : index
%1 = arith.cmpi slt, %0, %cst : index
return %1: i1
}
@@ -55,7 +55,7 @@ func.func @test() -> i1 {
// CHECK: return %[[C]]
func.func @test() -> i1 {
%cst1 = arith.constant -1 : index
- %0 = test.with_bounds { umin = 0 : index, umax = 0x7fffffffffffffff : index, smin = 0 : index, smax = 0x7fffffffffffffff : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 0x7fffffffffffffff : index, smin = 0 : index, smax = 0x7fffffffffffffff : index > : index
%1 = arith.cmpi sgt, %0, %cst1 : index
return %1: i1
}
@@ -67,7 +67,7 @@ func.func @test() -> i1 {
// CHECK: return %[[C]]
func.func @test() -> i1 {
%cst1 = arith.constant -1 : index
- %0 = test.with_bounds { umin = 0 : index, umax = 0x7fffffffffffffff : index, smin = 0 : index, smax = 0x7fffffffffffffff : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 0x7fffffffffffffff : index, smin = 0 : index, smax = 0x7fffffffffffffff : index > : index
%1 = arith.cmpi sle, %0, %cst1 : index
return %1: i1
}
@@ -75,10 +75,10 @@ func.func @test() -> i1 {
// -----
// CHECK-LABEL: func @test
-// CHECK: test.reflect_bounds {smax = 24 : si8, smin = 0 : si8, umax = 24 : ui8, umin = 0 : ui8}
+// CHECK: test.reflect_bounds <umin = 0 : ui8, umax = 24 : ui8, smin = 0 : si8, smax = 24 : si8>
func.func @test() -> i8 {
%cst1 = arith.constant 1 : i8
- %i8val = test.with_bounds { umin = 0 : i8, umax = 12 : i8, smin = 0 : i8, smax = 12 : i8 } : i8
+ %i8val = test.with_bounds < umin = 0 : i8, umax = 12 : i8, smin = 0 : i8, smax = 12 : i8 > : i8
%shifted = arith.shli %i8val, %cst1 : i8
%1 = test.reflect_bounds %shifted : i8
return %1: i8
@@ -87,10 +87,10 @@ func.func @test() -> i8 {
// -----
// CHECK-LABEL: func @test
-// CHECK: test.reflect_bounds {smax = 127 : si8, smin = -128 : si8, umax = 254 : ui8, umin = 0 : ui8}
+// CHECK: test.reflect_bounds <umin = 0 : ui8, umax = 254 : ui8, smin = -128 : si8, smax = 127 : si8>
func.func @test() -> i8 {
%cst1 = arith.constant 1 : i8
- %i8val = test.with_bounds { umin = 0 : i8, umax = 127 : i8, smin = 0 : i8, smax = 127 : i8 } : i8
+ %i8val = test.with_bounds < umin = 0 : i8, umax = 127 : i8, smin = 0 : i8, smax = 127 : i8 > : i8
%shifted = arith.shli %i8val, %cst1 : i8
%1 = test.reflect_bounds %shifted : i8
return %1: i8
@@ -103,7 +103,7 @@ func.func @test() -> i8 {
// CHECK: return [[val]]
func.func @trivial_rem() -> i8 {
%c64 = arith.constant 64 : i8
- %val = test.with_bounds { umin = 0 : ui8, umax = 63 : ui8, smin = 0 : si8, smax = 63 : si8 } : i8
+ %val = test.with_bounds < umin = 0 : ui8, umax = 63 : ui8, smin = 0 : si8, smax = 63 : si8 > : i8
%mod = arith.remsi %val, %c64 : i8
return %mod : i8
}
@@ -115,8 +115,8 @@ func.func @trivial_rem() -> i8 {
// CHECK: return [[mod]]
func.func @non_const_rhs() -> i8 {
%c64 = arith.constant 64 : i8
- %val = test.with_bounds { umin = 0 : ui8, umax = 2 : ui8, smin = 0 : si8, smax = 2 : si8 } : i8
- %rhs = test.with_bounds { umin = 63 : ui8, umax = 64 : ui8, smin = 63 : si8, smax = 64 : si8 } : i8
+ %val = test.with_bounds < umin = 0 : ui8, umax = 2 : ui8, smin = 0 : si8, smax = 2 : si8 > : i8
+ %rhs = test.with_bounds < umin = 63 : ui8, umax = 64 : ui8, smin = 63 : si8, smax = 64 : si8 > : i8
%mod = arith.remui %val, %rhs : i8
return %mod : i8
}
@@ -128,7 +128,7 @@ func.func @non_const_rhs() -> i8 {
// CHECK: return [[mod]]
func.func @wraps() -> i8 {
%c64 = arith.constant 64 : i8
- %val = test.with_bounds { umin = 63 : ui8, umax = 65 : ui8, smin = 63 : si8, smax = 65 : si8 } : i8
+ %val = test.with_bounds < umin = 63 : ui8, umax = 65 : ui8, smin = 63 : si8, smax = 65 : si8 > : i8
%mod = arith.remsi %val, %c64 : i8
return %mod : i8
}
diff --git a/mlir/test/Dialect/GPU/int-range-interface-cluster.mlir b/mlir/test/Dialect/GPU/int-range-interface-cluster.mlir
index ff8ddd147ee91..08be9c8c5a083 100644
--- a/mlir/test/Dialect/GPU/int-range-interface-cluster.mlir
+++ b/mlir/test/Dialect/GPU/int-range-interface-cluster.mlir
@@ -3,23 +3,23 @@
gpu.module @test_module {
gpu.func @test_cluster_ranges() kernel attributes {known_cluster_size = array<i32: 8, 4, 1>} {
%c0 = gpu.cluster_block_id x
- // CHECK: test.reflect_bounds {smax = 7 : index, smin = 0 : index, umax = 7 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 7 : index, smin = 0 : index, smax = 7 : index>
%c0_0 = test.reflect_bounds %c0 : index
%c1 = gpu.cluster_block_id y
- // CHECK: test.reflect_bounds {smax = 3 : index, smin = 0 : index, umax = 3 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 3 : index, smin = 0 : index, smax = 3 : index>
%c1_0 = test.reflect_bounds %c1 : index
%c2 = gpu.cluster_block_id z
- // CHECK: test.reflect_bounds {smax = 0 : index, smin = 0 : index, umax = 0 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 0 : index, smin = 0 : index, smax = 0 : index>
%c2_0 = test.reflect_bounds %c2 : index
%d0 = gpu.cluster_dim_blocks x
- // CHECK: test.reflect_bounds {smax = 8 : index, smin = 8 : index, umax = 8 : index, umin = 8 : index}
+ // CHECK: test.reflect_bounds <umin = 8 : index, umax = 8 : index, smin = 8 : index, smax = 8 : index>
%d0_0 = test.reflect_bounds %d0 : index
%d1 = gpu.cluster_dim_blocks y
- // CHECK: test.reflect_bounds {smax = 4 : index, smin = 4 : index, umax = 4 : index, umin = 4 : index}
+ // CHECK: test.reflect_bounds <umin = 4 : index, umax = 4 : index, smin = 4 : index, smax = 4 : index>
%d1_0 = test.reflect_bounds %d1 : index
%d2 = gpu.cluster_dim_blocks z
- // CHECK: test.reflect_bounds {smax = 1 : index, smin = 1 : index, umax = 1 : index, umin = 1 : index}
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 1 : index, smin = 1 : index, smax = 1 : index>
%d2_0 = test.reflect_bounds %d2 : index
gpu.return
@@ -27,23 +27,23 @@ gpu.module @test_module {
gpu.func @test_cluster_unknown() kernel {
%c0 = gpu.cluster_block_id x
- // CHECK: test.reflect_bounds {smax = 15 : index, smin = 0 : index, umax = 15 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 15 : index, smin = 0 : index, smax = 15 : index>
%c0_0 = test.reflect_bounds %c0 : index
%c1 = gpu.cluster_block_id y
- // CHECK: test.reflect_bounds {smax = 15 : index, smin = 0 : index, umax = 15 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 15 : index, smin = 0 : index, smax = 15 : index>
%c1_0 = test.reflect_bounds %c1 : index
%c2 = gpu.cluster_block_id z
- // CHECK: test.reflect_bounds {smax = 15 : index, smin = 0 : index, umax = 15 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 15 : index, smin = 0 : index, smax = 15 : index>
%c2_0 = test.reflect_bounds %c2 : index
%d0 = gpu.cluster_dim_blocks x
- // CHECK: test.reflect_bounds {smax = 16 : index, smin = 1 : index, umax = 16 : index, umin = 1 : index}
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 16 : index, smin = 1 : index, smax = 16 : index>
%d0_0 = test.reflect_bounds %d0 : index
%d1 = gpu.cluster_dim_blocks y
- // CHECK: test.reflect_bounds {smax = 16 : index, smin = 1 : index, umax = 16 : index, umin = 1 : index}
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 16 : index, smin = 1 : index, smax = 16 : index>
%d1_0 = test.reflect_bounds %d1 : index
%d2 = gpu.cluster_dim_blocks z
- // CHECK: test.reflect_bounds {smax = 16 : index, smin = 1 : index, umax = 16 : index, umin = 1 : index}
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 16 : index, smin = 1 : index, smax = 16 : index>
%d2_0 = test.reflect_bounds %d2 : index
gpu.return
diff --git a/mlir/test/Dialect/GPU/int-range-interface.mlir b/mlir/test/Dialect/GPU/int-range-interface.mlir
index 628b668a65969..caa8c2d7d05a4 100644
--- a/mlir/test/Dialect/GPU/int-range-interface.mlir
+++ b/mlir/test/Dialect/GPU/int-range-interface.mlir
@@ -2,47 +2,47 @@
// CHECK-LABEL: func @launch_func
func.func @launch_func(%arg0 : index) {
- %0 = test.with_bounds {
+ %0 = test.with_bounds <
umin = 3 : index, umax = 5 : index,
smin = 3 : index, smax = 5 : index
- } : index
- %1 = test.with_bounds {
+ > : index
+ %1 = test.with_bounds <
umin = 7 : index, umax = 11 : index,
smin = 7 : index, smax = 11 : index
- } : index
+ > : index
gpu.launch blocks(%block_id_x, %block_id_y, %block_id_z) in (%grid_dim_x = %0, %grid_dim_y = %1, %grid_dim_z = %arg0)
threads(%thread_id_x, %thread_id_y, %thread_id_z) in (%block_dim_x = %arg0, %block_dim_y = %0, %block_dim_z = %1) {
- // CHECK: test.reflect_bounds {smax = 5 : index, smin = 3 : index, umax = 5 : index, umin = 3 : index}
- // CHECK: test.reflect_bounds {smax = 11 : index, smin = 7 : index, umax = 11 : index, umin = 7 : index}
- // CHECK: test.reflect_bounds {smax = 4294967295 : index, smin = 1 : index, umax = 4294967295 : index, umin = 1 : index}
+ // CHECK: test.reflect_bounds <umin = 3 : index, umax = 5 : index, smin = 3 : index, smax = 5 : index>
+ // CHECK: test.reflect_bounds <umin = 7 : index, umax = 11 : index, smin = 7 : index, smax = 11 : index>
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 4294967295 : index, smin = 1 : index, smax = 4294967295 : index>
%grid_dim_x0 = test.reflect_bounds %grid_dim_x : index
%grid_dim_y0 = test.reflect_bounds %grid_dim_y : index
%grid_dim_z0 = test.reflect_bounds %grid_dim_z : index
- // CHECK: test.reflect_bounds {smax = 4 : index, smin = 0 : index, umax = 4 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 10 : index, smin = 0 : index, umax = 10 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 4294967294 : index, smin = 0 : index, umax = 4294967294 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 4 : index, smin = 0 : index, smax = 4 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 4294967294 : index, smin = 0 : index, smax = 4294967294 : index>
%block_id_x0 = test.reflect_bounds %block_id_x : index
%block_id_y0 = test.reflect_bounds %block_id_y : index
%block_id_z0 = test.reflect_bounds %block_id_z : index
- // CHECK: test.reflect_bounds {smax = 4294967295 : index, smin = 1 : index, umax = 4294967295 : index, umin = 1 : index}
- // CHECK: test.reflect_bounds {smax = 5 : index, smin = 3 : index, umax = 5 : index, umin = 3 : index}
- // CHECK: test.reflect_bounds {smax = 11 : index, smin = 7 : index, umax = 11 : index, umin = 7 : index}
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 4294967295 : index, smin = 1 : index, smax = 4294967295 : index>
+ // CHECK: test.reflect_bounds <umin = 3 : index, umax = 5 : index, smin = 3 : index, smax = 5 : index>
+ // CHECK: test.reflect_bounds <umin = 7 : index, umax = 11 : index, smin = 7 : index, smax = 11 : index>
%block_dim_x0 = test.reflect_bounds %block_dim_x : index
%block_dim_y0 = test.reflect_bounds %block_dim_y : index
%block_dim_z0 = test.reflect_bounds %block_dim_z : index
- // CHECK: test.reflect_bounds {smax = 4294967294 : index, smin = 0 : index, umax = 4294967294 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 4 : index, smin = 0 : index, umax = 4 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 10 : index, smin = 0 : index, umax = 10 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 4294967294 : index, smin = 0 : index, smax = 4294967294 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 4 : index, smin = 0 : index, smax = 4 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index>
%thread_id_x0 = test.reflect_bounds %thread_id_x : index
%thread_id_y0 = test.reflect_bounds %thread_id_y : index
%thread_id_z0 = test.reflect_bounds %thread_id_z : index
// The launch bounds are not constant, and so this can't infer anything
- // CHECK: test.reflect_bounds {smax = 4294967294 : index, smin = 0 : index, umax = 4294967294 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 4294967294 : index, smin = 0 : index, smax = 4294967294 : index>
%thread_id_op = gpu.thread_id y
%thread_id_op0 = test.reflect_bounds %thread_id_op : index
gpu.terminator
@@ -62,9 +62,9 @@ module attributes {gpu.container_module} {
%grid_dim_y = gpu.grid_dim y
%grid_dim_z = gpu.grid_dim z
- // CHECK: test.reflect_bounds {smax = 4294967295 : index, smin = 1 : index, umax = 4294967295 : index, umin = 1 : index}
- // CHECK: test.reflect_bounds {smax = 4294967295 : index, smin = 1 : index, umax = 4294967295 : index, umin = 1 : index}
- // CHECK: test.reflect_bounds {smax = 4294967295 : index, smin = 1 : index, umax = 4294967295 : index, umin = 1 : index}
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 4294967295 : index, smin = 1 : index, smax = 4294967295 : index>
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 4294967295 : index, smin = 1 : index, smax = 4294967295 : index>
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 4294967295 : index, smin = 1 : index, smax = 4294967295 : index>
%grid_dim_x0 = test.reflect_bounds %grid_dim_x : index
%grid_dim_y0 = test.reflect_bounds %grid_dim_y : index
%grid_dim_z0 = test.reflect_bounds %grid_dim_z : index
@@ -73,9 +73,9 @@ module attributes {gpu.container_module} {
%block_id_y = gpu.block_id y
%block_id_z = gpu.block_id z
- // CHECK: test.reflect_bounds {smax = 4294967294 : index, smin = 0 : index, umax = 4294967294 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 4294967294 : index, smin = 0 : index, umax = 4294967294 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 4294967294 : index, smin = 0 : index, umax = 4294967294 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 4294967294 : index, smin = 0 : index, smax = 4294967294 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 4294967294 : index, smin = 0 : index, smax = 4294967294 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 4294967294 : index, smin = 0 : index, smax = 4294967294 : index>
%block_id_x0 = test.reflect_bounds %block_id_x : index
%block_id_y0 = test.reflect_bounds %block_id_y : index
%block_id_z0 = test.reflect_bounds %block_id_z : index
@@ -84,9 +84,9 @@ module attributes {gpu.container_module} {
%block_dim_y = gpu.block_dim y
%block_dim_z = gpu.block_dim z
- // CHECK: test.reflect_bounds {smax = 4294967295 : index, smin = 1 : index, umax = 4294967295 : index, umin = 1 : index}
- // CHECK: test.reflect_bounds {smax = 4294967295 : index, smin = 1 : index, umax = 4294967295 : index, umin = 1 : index}
- // CHECK: test.reflect_bounds {smax = 4294967295 : index, smin = 1 : index, umax = 4294967295 : index, umin = 1 : index}
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 4294967295 : index, smin = 1 : index, smax = 4294967295 : index>
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 4294967295 : index, smin = 1 : index, smax = 4294967295 : index>
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 4294967295 : index, smin = 1 : index, smax = 4294967295 : index>
%block_dim_x0 = test.reflect_bounds %block_dim_x : index
%block_dim_y0 = test.reflect_bounds %block_dim_y : index
%block_dim_z0 = test.reflect_bounds %block_dim_z : index
@@ -95,9 +95,9 @@ module attributes {gpu.container_module} {
%thread_id_y = gpu.thread_id y
%thread_id_z = gpu.thread_id z
- // CHECK: test.reflect_bounds {smax = 4294967294 : index, smin = 0 : index, umax = 4294967294 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 4294967294 : index, smin = 0 : index, umax = 4294967294 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 4294967294 : index, smin = 0 : index, umax = 4294967294 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 4294967294 : index, smin = 0 : index, smax = 4294967294 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 4294967294 : index, smin = 0 : index, smax = 4294967294 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 4294967294 : index, smin = 0 : index, smax = 4294967294 : index>
%thread_id_x0 = test.reflect_bounds %thread_id_x : index
%thread_id_y0 = test.reflect_bounds %thread_id_y : index
%thread_id_z0 = test.reflect_bounds %thread_id_z : index
@@ -106,9 +106,9 @@ module attributes {gpu.container_module} {
%global_id_y = gpu.global_id y
%global_id_z = gpu.global_id z
- // CHECK: test.reflect_bounds {smax = 9223372036854775807 : index, smin = -9223372036854775808 : index, umax = -8589934592 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 9223372036854775807 : index, smin = -9223372036854775808 : index, umax = -8589934592 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 9223372036854775807 : index, smin = -9223372036854775808 : index, umax = -8589934592 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = -8589934592 : index, smin = -9223372036854775808 : index, smax = 9223372036854775807 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = -8589934592 : index, smin = -9223372036854775808 : index, smax = 9223372036854775807 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = -8589934592 : index, smin = -9223372036854775808 : index, smax = 9223372036854775807 : index>
%global_id_x0 = test.reflect_bounds %global_id_x : index
%global_id_y0 = test.reflect_bounds %global_id_y : index
%global_id_z0 = test.reflect_bounds %global_id_z : index
@@ -118,10 +118,10 @@ module attributes {gpu.container_module} {
%num_subgroups = gpu.num_subgroups : index
%subgroup_id = gpu.subgroup_id : index
- // CHECK: test.reflect_bounds {smax = 128 : index, smin = 1 : index, umax = 128 : index, umin = 1 : index}
- // CHECK: test.reflect_bounds {smax = 127 : index, smin = 0 : index, umax = 127 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 4294967295 : index, smin = 1 : index, umax = 4294967295 : index, umin = 1 : index}
- // CHECK: test.reflect_bounds {smax = 4294967294 : index, smin = 0 : index, umax = 4294967294 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 128 : index, smin = 1 : index, smax = 128 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 127 : index, smin = 0 : index, smax = 127 : index>
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 4294967295 : index, smin = 1 : index, smax = 4294967295 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 4294967294 : index, smin = 0 : index, smax = 4294967294 : index>
%subgroup_size0 = test.reflect_bounds %subgroup_size : index
%lane_id0 = test.reflect_bounds %lane_id : index
%num_subgroups0 = test.reflect_bounds %num_subgroups : index
@@ -145,9 +145,9 @@ module attributes {gpu.container_module} {
%grid_dim_y = gpu.grid_dim y
%grid_dim_z = gpu.grid_dim z
- // CHECK: test.reflect_bounds {smax = 20 : index, smin = 20 : index, umax = 20 : index, umin = 20 : index}
- // CHECK: test.reflect_bounds {smax = 24 : index, smin = 24 : index, umax = 24 : index, umin = 24 : index}
- // CHECK: test.reflect_bounds {smax = 28 : index, smin = 28 : index, umax = 28 : index, umin = 28 : index}
+ // CHECK: test.reflect_bounds <umin = 20 : index, umax = 20 : index, smin = 20 : index, smax = 20 : index>
+ // CHECK: test.reflect_bounds <umin = 24 : index, umax = 24 : index, smin = 24 : index, smax = 24 : index>
+ // CHECK: test.reflect_bounds <umin = 28 : index, umax = 28 : index, smin = 28 : index, smax = 28 : index>
%grid_dim_x0 = test.reflect_bounds %grid_dim_x : index
%grid_dim_y0 = test.reflect_bounds %grid_dim_y : index
%grid_dim_z0 = test.reflect_bounds %grid_dim_z : index
@@ -156,9 +156,9 @@ module attributes {gpu.container_module} {
%block_id_y = gpu.block_id y
%block_id_z = gpu.block_id z
- // CHECK: test.reflect_bounds {smax = 19 : index, smin = 0 : index, umax = 19 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 23 : index, smin = 0 : index, umax = 23 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 27 : index, smin = 0 : index, umax = 27 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 19 : index, smin = 0 : index, smax = 19 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 23 : index, smin = 0 : index, smax = 23 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 27 : index, smin = 0 : index, smax = 27 : index>
%block_id_x0 = test.reflect_bounds %block_id_x : index
%block_id_y0 = test.reflect_bounds %block_id_y : index
%block_id_z0 = test.reflect_bounds %block_id_z : index
@@ -167,9 +167,9 @@ module attributes {gpu.container_module} {
%block_dim_y = gpu.block_dim y
%block_dim_z = gpu.block_dim z
- // CHECK: test.reflect_bounds {smax = 8 : index, smin = 8 : index, umax = 8 : index, umin = 8 : index}
- // CHECK: test.reflect_bounds {smax = 12 : index, smin = 12 : index, umax = 12 : index, umin = 12 : index}
- // CHECK: test.reflect_bounds {smax = 16 : index, smin = 16 : index, umax = 16 : index, umin = 16 : index}
+ // CHECK: test.reflect_bounds <umin = 8 : index, umax = 8 : index, smin = 8 : index, smax = 8 : index>
+ // CHECK: test.reflect_bounds <umin = 12 : index, umax = 12 : index, smin = 12 : index, smax = 12 : index>
+ // CHECK: test.reflect_bounds <umin = 16 : index, umax = 16 : index, smin = 16 : index, smax = 16 : index>
%block_dim_x0 = test.reflect_bounds %block_dim_x : index
%block_dim_y0 = test.reflect_bounds %block_dim_y : index
%block_dim_z0 = test.reflect_bounds %block_dim_z : index
@@ -178,9 +178,9 @@ module attributes {gpu.container_module} {
%thread_id_y = gpu.thread_id y
%thread_id_z = gpu.thread_id z
- // CHECK: test.reflect_bounds {smax = 7 : index, smin = 0 : index, umax = 7 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 11 : index, smin = 0 : index, umax = 11 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 15 : index, smin = 0 : index, umax = 15 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 7 : index, smin = 0 : index, smax = 7 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 11 : index, smin = 0 : index, smax = 11 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 15 : index, smin = 0 : index, smax = 15 : index>
%thread_id_x0 = test.reflect_bounds %thread_id_x : index
%thread_id_y0 = test.reflect_bounds %thread_id_y : index
%thread_id_z0 = test.reflect_bounds %thread_id_z : index
@@ -189,9 +189,9 @@ module attributes {gpu.container_module} {
%global_id_y = gpu.global_id y
%global_id_z = gpu.global_id z
- // CHECK: test.reflect_bounds {smax = 159 : index, smin = 0 : index, umax = 159 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 287 : index, smin = 0 : index, umax = 287 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 447 : index, smin = 0 : index, umax = 447 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 159 : index, smin = 0 : index, smax = 159 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 287 : index, smin = 0 : index, smax = 287 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 447 : index, smin = 0 : index, smax = 447 : index>
%global_id_x0 = test.reflect_bounds %global_id_x : index
%global_id_y0 = test.reflect_bounds %global_id_y : index
%global_id_z0 = test.reflect_bounds %global_id_z : index
@@ -201,10 +201,10 @@ module attributes {gpu.container_module} {
%num_subgroups = gpu.num_subgroups : index
%subgroup_id = gpu.subgroup_id : index
- // CHECK: test.reflect_bounds {smax = 128 : index, smin = 1 : index, umax = 128 : index, umin = 1 : index}
- // CHECK: test.reflect_bounds {smax = 127 : index, smin = 0 : index, umax = 127 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 4294967295 : index, smin = 1 : index, umax = 4294967295 : index, umin = 1 : index}
- // CHECK: test.reflect_bounds {smax = 4294967294 : index, smin = 0 : index, umax = 4294967294 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 128 : index, smin = 1 : index, smax = 128 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 127 : index, smin = 0 : index, smax = 127 : index>
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 4294967295 : index, smin = 1 : index, smax = 4294967295 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 4294967294 : index, smin = 0 : index, smax = 4294967294 : index>
%subgroup_size0 = test.reflect_bounds %subgroup_size : index
%lane_id0 = test.reflect_bounds %lane_id : index
%num_subgroups0 = test.reflect_bounds %num_subgroups : index
@@ -227,9 +227,9 @@ module {
%block_id_y = gpu.block_id y
%block_id_z = gpu.block_id z
- // CHECK: test.reflect_bounds {smax = 19 : index, smin = 0 : index, umax = 19 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 23 : index, smin = 0 : index, umax = 23 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 27 : index, smin = 0 : index, umax = 27 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 19 : index, smin = 0 : index, smax = 19 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 23 : index, smin = 0 : index, smax = 23 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 27 : index, smin = 0 : index, smax = 27 : index>
%block_id_x0 = test.reflect_bounds %block_id_x : index
%block_id_y0 = test.reflect_bounds %block_id_y : index
%block_id_z0 = test.reflect_bounds %block_id_z : index
@@ -238,9 +238,9 @@ module {
%thread_id_y = gpu.thread_id y
%thread_id_z = gpu.thread_id z
- // CHECK: test.reflect_bounds {smax = 7 : index, smin = 0 : index, umax = 7 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 11 : index, smin = 0 : index, umax = 11 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 15 : index, smin = 0 : index, umax = 15 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 7 : index, smin = 0 : index, smax = 7 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 11 : index, smin = 0 : index, smax = 11 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 15 : index, smin = 0 : index, smax = 15 : index>
%thread_id_x0 = test.reflect_bounds %thread_id_x : index
%thread_id_y0 = test.reflect_bounds %thread_id_y : index
%thread_id_z0 = test.reflect_bounds %thread_id_z : index
@@ -260,9 +260,9 @@ module attributes {gpu.container_module} {
%grid_dim_y = gpu.grid_dim y upper_bound 24
%grid_dim_z = gpu.grid_dim z upper_bound 28
- // CHECK: test.reflect_bounds {smax = 20 : index, smin = 1 : index, umax = 20 : index, umin = 1 : index}
- // CHECK: test.reflect_bounds {smax = 24 : index, smin = 1 : index, umax = 24 : index, umin = 1 : index}
- // CHECK: test.reflect_bounds {smax = 28 : index, smin = 1 : index, umax = 28 : index, umin = 1 : index}
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 20 : index, smin = 1 : index, smax = 20 : index>
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 24 : index, smin = 1 : index, smax = 24 : index>
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 28 : index, smin = 1 : index, smax = 28 : index>
%grid_dim_x0 = test.reflect_bounds %grid_dim_x : index
%grid_dim_y0 = test.reflect_bounds %grid_dim_y : index
%grid_dim_z0 = test.reflect_bounds %grid_dim_z : index
@@ -271,9 +271,9 @@ module attributes {gpu.container_module} {
%block_id_y = gpu.block_id y upper_bound 24
%block_id_z = gpu.block_id z upper_bound 28
- // CHECK: test.reflect_bounds {smax = 19 : index, smin = 0 : index, umax = 19 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 23 : index, smin = 0 : index, umax = 23 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 27 : index, smin = 0 : index, umax = 27 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 19 : index, smin = 0 : index, smax = 19 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 23 : index, smin = 0 : index, smax = 23 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 27 : index, smin = 0 : index, smax = 27 : index>
%block_id_x0 = test.reflect_bounds %block_id_x : index
%block_id_y0 = test.reflect_bounds %block_id_y : index
%block_id_z0 = test.reflect_bounds %block_id_z : index
@@ -282,9 +282,9 @@ module attributes {gpu.container_module} {
%block_dim_y = gpu.block_dim y upper_bound 12
%block_dim_z = gpu.block_dim z upper_bound 16
- // CHECK: test.reflect_bounds {smax = 8 : index, smin = 1 : index, umax = 8 : index, umin = 1 : index}
- // CHECK: test.reflect_bounds {smax = 12 : index, smin = 1 : index, umax = 12 : index, umin = 1 : index}
- // CHECK: test.reflect_bounds {smax = 16 : index, smin = 1 : index, umax = 16 : index, umin = 1 : index}
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 8 : index, smin = 1 : index, smax = 8 : index>
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 12 : index, smin = 1 : index, smax = 12 : index>
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 16 : index, smin = 1 : index, smax = 16 : index>
%block_dim_x0 = test.reflect_bounds %block_dim_x : index
%block_dim_y0 = test.reflect_bounds %block_dim_y : index
%block_dim_z0 = test.reflect_bounds %block_dim_z : index
@@ -293,9 +293,9 @@ module attributes {gpu.container_module} {
%thread_id_y = gpu.thread_id y upper_bound 12
%thread_id_z = gpu.thread_id z upper_bound 16
- // CHECK: test.reflect_bounds {smax = 7 : index, smin = 0 : index, umax = 7 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 11 : index, smin = 0 : index, umax = 11 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 15 : index, smin = 0 : index, umax = 15 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 7 : index, smin = 0 : index, smax = 7 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 11 : index, smin = 0 : index, smax = 11 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 15 : index, smin = 0 : index, smax = 15 : index>
%thread_id_x0 = test.reflect_bounds %thread_id_x : index
%thread_id_y0 = test.reflect_bounds %thread_id_y : index
%thread_id_z0 = test.reflect_bounds %thread_id_z : index
@@ -304,9 +304,9 @@ module attributes {gpu.container_module} {
%global_id_y = gpu.global_id y upper_bound 288
%global_id_z = gpu.global_id z upper_bound 448
- // CHECK: test.reflect_bounds {smax = 159 : index, smin = 0 : index, umax = 159 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 287 : index, smin = 0 : index, umax = 287 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 447 : index, smin = 0 : index, umax = 447 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 159 : index, smin = 0 : index, smax = 159 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 287 : index, smin = 0 : index, smax = 287 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 447 : index, smin = 0 : index, smax = 447 : index>
%global_id_x0 = test.reflect_bounds %global_id_x : index
%global_id_y0 = test.reflect_bounds %global_id_y : index
%global_id_z0 = test.reflect_bounds %global_id_z : index
@@ -316,10 +316,10 @@ module attributes {gpu.container_module} {
%num_subgroups = gpu.num_subgroups upper_bound 8 : index
%lane_id = gpu.lane_id upper_bound 64
- // CHECK: test.reflect_bounds {smax = 32 : index, smin = 1 : index, umax = 32 : index, umin = 1 : index}
- // CHECK: test.reflect_bounds {smax = 31 : index, smin = 0 : index, umax = 31 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 8 : index, smin = 1 : index, umax = 8 : index, umin = 1 : index}
- // CHECK: test.reflect_bounds {smax = 63 : index, smin = 0 : index, umax = 63 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 32 : index, smin = 1 : index, smax = 32 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 31 : index, smin = 0 : index, smax = 31 : index>
+ // CHECK: test.reflect_bounds <umin = 1 : index, umax = 8 : index, smin = 1 : index, smax = 8 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 63 : index, smin = 0 : index, smax = 63 : index>
%subgroup_size0 = test.reflect_bounds %subgroup_size : index
%subgroup_id0 = test.reflect_bounds %subgroup_id : index
%num_subgroups0 = test.reflect_bounds %num_subgroups : index
@@ -334,12 +334,12 @@ module attributes {gpu.container_module} {
// CHECK-LABEL: func @broadcast
func.func @broadcast(%idx: i32) {
- %0 = test.with_bounds { umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index > : index
%1 = gpu.subgroup_broadcast %0, first_active_lane : index
%2 = gpu.subgroup_broadcast %0, specific_lane %idx : index
- // CHECK: test.reflect_bounds {smax = 10 : index, smin = 0 : index, umax = 10 : index, umin = 0 : index}
- // CHECK: test.reflect_bounds {smax = 10 : index, smin = 0 : index, umax = 10 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index>
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 10 : index, smin = 0 : index, smax = 10 : index>
%4 = test.reflect_bounds %1 : index
%5 = test.reflect_bounds %2 : index
diff --git a/mlir/test/Dialect/MLProgram/pipeline-globals.mlir b/mlir/test/Dialect/MLProgram/pipeline-globals.mlir
index 3f45d0ff2eb38..37d59cda80805 100644
--- a/mlir/test/Dialect/MLProgram/pipeline-globals.mlir
+++ b/mlir/test/Dialect/MLProgram/pipeline-globals.mlir
@@ -233,7 +233,7 @@ func.func @call_with_unresolvable_callee(%arg0: memref<f32>) {
// CHECK: ml_program.global_load @global_variable
%0 = ml_program.global_load @global_variable : tensor<4xi32>
// @callee is not defined anywhere in this module.
- test.call_and_store @callee(%arg0), %arg0 {store_before_call = false} : (memref<f32>, memref<f32>) -> ()
+ test.call_and_store @callee(%arg0), %arg0 <store_before_call = false> : (memref<f32>, memref<f32>) -> ()
// CHECK: ml_program.global_load @global_variable
%1 = ml_program.global_load @global_variable : tensor<4xi32>
func.return
diff --git a/mlir/test/Dialect/MemRef/int-range-inference.mlir b/mlir/test/Dialect/MemRef/int-range-inference.mlir
index 34568d1d1d520..1228f64b9a056 100644
--- a/mlir/test/Dialect/MemRef/int-range-inference.mlir
+++ b/mlir/test/Dialect/MemRef/int-range-inference.mlir
@@ -13,7 +13,7 @@ func.func @dim_const(%m: memref<3x5xi32>) -> index {
// CHECK-LABEL: @dim_any_static
// CHECK: %[[op:.+]] = memref.dim
-// CHECK: %[[ret:.+]] = test.reflect_bounds {smax = 5 : index, smin = 3 : index, umax = 5 : index, umin = 3 : index} %[[op]]
+// CHECK: %[[ret:.+]] = test.reflect_bounds <umin = 3 : index, umax = 5 : index, smin = 3 : index, smax = 5 : index> %[[op]]
// CHECK: return %[[ret]]
func.func @dim_any_static(%m: memref<3x5xi32>, %x: index) -> index {
%0 = memref.dim %m, %x : memref<3x5xi32>
@@ -25,7 +25,7 @@ func.func @dim_any_static(%m: memref<3x5xi32>, %x: index) -> index {
// CHECK-LABEL: @dim_dynamic
// CHECK: %[[op:.+]] = memref.dim
-// CHECK: %[[ret:.+]] = test.reflect_bounds {smax = 9223372036854775807 : index, smin = 0 : index, umax = 9223372036854775807 : index, umin = 0 : index} %[[op]]
+// CHECK: %[[ret:.+]] = test.reflect_bounds <umin = 0 : index, umax = 9223372036854775807 : index, smin = 0 : index, smax = 9223372036854775807 : index> %[[op]]
// CHECK: return %[[ret]]
func.func @dim_dynamic(%m: memref<?x5xi32>) -> index {
%c0 = arith.constant 0 : index
@@ -38,7 +38,7 @@ func.func @dim_dynamic(%m: memref<?x5xi32>) -> index {
// CHECK-LABEL: @dim_any_dynamic
// CHECK: %[[op:.+]] = memref.dim
-// CHECK: %[[ret:.+]] = test.reflect_bounds {smax = 9223372036854775807 : index, smin = 0 : index, umax = 9223372036854775807 : index, umin = 0 : index} %[[op]]
+// CHECK: %[[ret:.+]] = test.reflect_bounds <umin = 0 : index, umax = 9223372036854775807 : index, smin = 0 : index, smax = 9223372036854775807 : index> %[[op]]
// CHECK: return %[[ret]]
func.func @dim_any_dynamic(%m: memref<?x5xi32>, %x: index) -> index {
%0 = memref.dim %m, %x : memref<?x5xi32>
@@ -50,7 +50,7 @@ func.func @dim_any_dynamic(%m: memref<?x5xi32>, %x: index) -> index {
// CHECK-LABEL: @dim_some_omitting_dynamic
// CHECK: %[[op:.+]] = memref.dim
-// CHECK: %[[ret:.+]] = test.reflect_bounds {smax = 5 : index, smin = 3 : index, umax = 5 : index, umin = 3 : index} %[[op]]
+// CHECK: %[[ret:.+]] = test.reflect_bounds <umin = 3 : index, umax = 5 : index, smin = 3 : index, smax = 5 : index> %[[op]]
// CHECK: return %[[ret]]
func.func @dim_some_omitting_dynamic(%m: memref<?x3x5xi32>, %x: index) -> index {
%c1 = arith.constant 1 : index
@@ -64,7 +64,7 @@ func.func @dim_some_omitting_dynamic(%m: memref<?x3x5xi32>, %x: index) -> index
// CHECK-LABEL: @dim_unranked
// CHECK: %[[op:.+]] = memref.dim
-// CHECK: %[[ret:.+]] = test.reflect_bounds {smax = 9223372036854775807 : index, smin = 0 : index, umax = 9223372036854775807 : index, umin = 0 : index} %[[op]]
+// CHECK: %[[ret:.+]] = test.reflect_bounds <umin = 0 : index, umax = 9223372036854775807 : index, smin = 0 : index, smax = 9223372036854775807 : index> %[[op]]
// CHECK: return %[[ret]]
func.func @dim_unranked(%m: memref<*xi32>) -> index {
%c0 = arith.constant 0 : index
diff --git a/mlir/test/Dialect/OpenACC/acc-implicit-routine.mlir b/mlir/test/Dialect/OpenACC/acc-implicit-routine.mlir
index 88853dfb307ee..cef8fc663291e 100644
--- a/mlir/test/Dialect/OpenACC/acc-implicit-routine.mlir
+++ b/mlir/test/Dialect/OpenACC/acc-implicit-routine.mlir
@@ -244,7 +244,7 @@ module {
func.func @test_undefined_callee() {
%buf = memref.alloca() : memref<f32>
acc.serial {
- test.call_and_store @undefined_callee(%buf), %buf {store_before_call = false} : (memref<f32>, memref<f32>) -> ()
+ test.call_and_store @undefined_callee(%buf), %buf <store_before_call = false> : (memref<f32>, memref<f32>) -> ()
acc.yield
}
return
diff --git a/mlir/test/Dialect/Tensor/int-range-inference.mlir b/mlir/test/Dialect/Tensor/int-range-inference.mlir
index e90ebf5fccb8e..946416558e31d 100644
--- a/mlir/test/Dialect/Tensor/int-range-inference.mlir
+++ b/mlir/test/Dialect/Tensor/int-range-inference.mlir
@@ -13,7 +13,7 @@ func.func @dim_const(%t: tensor<3x5xi32>) -> index {
// CHECK-LABEL: @dim_any_static
// CHECK: %[[op:.+]] = tensor.dim
-// CHECK: %[[ret:.+]] = test.reflect_bounds {smax = 5 : index, smin = 3 : index, umax = 5 : index, umin = 3 : index} %[[op]]
+// CHECK: %[[ret:.+]] = test.reflect_bounds <umin = 3 : index, umax = 5 : index, smin = 3 : index, smax = 5 : index> %[[op]]
// CHECK: return %[[ret]]
func.func @dim_any_static(%t: tensor<3x5xi32>, %x: index) -> index {
%0 = tensor.dim %t, %x : tensor<3x5xi32>
@@ -25,7 +25,7 @@ func.func @dim_any_static(%t: tensor<3x5xi32>, %x: index) -> index {
// CHECK-LABEL: @dim_dynamic
// CHECK: %[[op:.+]] = tensor.dim
-// CHECK: %[[ret:.+]] = test.reflect_bounds {smax = 9223372036854775807 : index, smin = 0 : index, umax = 9223372036854775807 : index, umin = 0 : index} %[[op]]
+// CHECK: %[[ret:.+]] = test.reflect_bounds <umin = 0 : index, umax = 9223372036854775807 : index, smin = 0 : index, smax = 9223372036854775807 : index> %[[op]]
// CHECK: return %[[ret]]
func.func @dim_dynamic(%t: tensor<?x5xi32>) -> index {
%c0 = arith.constant 0 : index
@@ -38,7 +38,7 @@ func.func @dim_dynamic(%t: tensor<?x5xi32>) -> index {
// CHECK-LABEL: @dim_any_dynamic
// CHECK: %[[op:.+]] = tensor.dim
-// CHECK: %[[ret:.+]] = test.reflect_bounds {smax = 9223372036854775807 : index, smin = 0 : index, umax = 9223372036854775807 : index, umin = 0 : index} %[[op]]
+// CHECK: %[[ret:.+]] = test.reflect_bounds <umin = 0 : index, umax = 9223372036854775807 : index, smin = 0 : index, smax = 9223372036854775807 : index> %[[op]]
// CHECK: return %[[ret]]
func.func @dim_any_dynamic(%t: tensor<?x5xi32>, %x: index) -> index {
%0 = tensor.dim %t, %x : tensor<?x5xi32>
@@ -50,7 +50,7 @@ func.func @dim_any_dynamic(%t: tensor<?x5xi32>, %x: index) -> index {
// CHECK-LABEL: @dim_some_omitting_dynamic
// CHECK: %[[op:.+]] = tensor.dim
-// CHECK: %[[ret:.+]] = test.reflect_bounds {smax = 5 : index, smin = 3 : index, umax = 5 : index, umin = 3 : index} %[[op]]
+// CHECK: %[[ret:.+]] = test.reflect_bounds <umin = 3 : index, umax = 5 : index, smin = 3 : index, smax = 5 : index> %[[op]]
// CHECK: return %[[ret]]
func.func @dim_some_omitting_dynamic(%t: tensor<?x3x5xi32>, %x: index) -> index {
%c1 = arith.constant 1 : index
@@ -64,7 +64,7 @@ func.func @dim_some_omitting_dynamic(%t: tensor<?x3x5xi32>, %x: index) -> index
// CHECK-LABEL: @dim_unranked
// CHECK: %[[op:.+]] = tensor.dim
-// CHECK: %[[ret:.+]] = test.reflect_bounds {smax = 9223372036854775807 : index, smin = 0 : index, umax = 9223372036854775807 : index, umin = 0 : index} %[[op]]
+// CHECK: %[[ret:.+]] = test.reflect_bounds <umin = 0 : index, umax = 9223372036854775807 : index, smin = 0 : index, smax = 9223372036854775807 : index> %[[op]]
// CHECK: return %[[ret]]
func.func @dim_unranked(%t: tensor<*xi32>) -> index {
%c0 = arith.constant 0 : index
diff --git a/mlir/test/Dialect/Vector/int-range-interface.mlir b/mlir/test/Dialect/Vector/int-range-interface.mlir
index 6e2aa064a67f8..003d0d36f26b0 100644
--- a/mlir/test/Dialect/Vector/int-range-interface.mlir
+++ b/mlir/test/Dialect/Vector/int-range-interface.mlir
@@ -2,7 +2,7 @@
// CHECK-LABEL: func @constant_vec
-// CHECK: test.reflect_bounds {smax = 7 : index, smin = 0 : index, umax = 7 : index, umin = 0 : index}
+// CHECK: test.reflect_bounds <umin = 0 : index, umax = 7 : index, smin = 0 : index, smax = 7 : index>
func.func @constant_vec() -> vector<8xindex> {
%0 = arith.constant dense<[0, 1, 2, 3, 4, 5, 6, 7]> : vector<8xindex>
%1 = test.reflect_bounds %0 : vector<8xindex>
@@ -10,7 +10,7 @@ func.func @constant_vec() -> vector<8xindex> {
}
// CHECK-LABEL: func @constant_splat
-// CHECK: test.reflect_bounds {smax = 3 : si32, smin = 3 : si32, umax = 3 : ui32, umin = 3 : ui32}
+// CHECK: test.reflect_bounds <umin = 3 : ui32, umax = 3 : ui32, smin = 3 : si32, smax = 3 : si32>
func.func @constant_splat() -> vector<8xi32> {
%0 = arith.constant dense<3> : vector<8xi32>
%1 = test.reflect_bounds %0 : vector<8xi32>
@@ -25,65 +25,65 @@ func.func @float_constant_splat() -> vector<8xf32> {
}
// CHECK-LABEL: func @vector_splat
-// CHECK: test.reflect_bounds {smax = 5 : index, smin = 4 : index, umax = 5 : index, umin = 4 : index}
+// CHECK: test.reflect_bounds <umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index>
func.func @vector_splat() -> vector<4xindex> {
- %0 = test.with_bounds { umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index } : index
+ %0 = test.with_bounds < umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index > : index
%1 = vector.broadcast %0 : index to vector<4xindex>
%2 = test.reflect_bounds %1 : vector<4xindex>
func.return %2 : vector<4xindex>
}
// CHECK-LABEL: func @vector_broadcast
-// CHECK: test.reflect_bounds {smax = 5 : index, smin = 4 : index, umax = 5 : index, umin = 4 : index}
+// CHECK: test.reflect_bounds <umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index>
func.func @vector_broadcast() -> vector<4x16xindex> {
- %0 = test.with_bounds { umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index } : vector<16xindex>
+ %0 = test.with_bounds < umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index > : vector<16xindex>
%1 = vector.broadcast %0 : vector<16xindex> to vector<4x16xindex>
%2 = test.reflect_bounds %1 : vector<4x16xindex>
func.return %2 : vector<4x16xindex>
}
// CHECK-LABEL: func @vector_shape_cast
-// CHECK: test.reflect_bounds {smax = 5 : index, smin = 4 : index, umax = 5 : index, umin = 4 : index}
+// CHECK: test.reflect_bounds <umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index>
func.func @vector_shape_cast() -> vector<4x4xindex> {
- %0 = test.with_bounds { umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index } : vector<16xindex>
+ %0 = test.with_bounds < umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index > : vector<16xindex>
%1 = vector.shape_cast %0 : vector<16xindex> to vector<4x4xindex>
%2 = test.reflect_bounds %1 : vector<4x4xindex>
func.return %2 : vector<4x4xindex>
}
// CHECK-LABEL: func @vector_transpose
-// CHECK: test.reflect_bounds {smax = 8 : index, smin = 7 : index, umax = 8 : index, umin = 7 : index}
+// CHECK: test.reflect_bounds <umin = 7 : index, umax = 8 : index, smin = 7 : index, smax = 8 : index>
func.func @vector_transpose() -> vector<2x4xindex> {
- %0 = test.with_bounds { smax = 8 : index, smin = 7 : index, umax = 8 : index, umin = 7 : index } : vector<4x2xindex>
+ %0 = test.with_bounds < smax = 8 : index, smin = 7 : index, umax = 8 : index, umin = 7 : index > : vector<4x2xindex>
%1 = vector.transpose %0, [1, 0] : vector<4x2xindex> to vector<2x4xindex>
%2 = test.reflect_bounds %1 : vector<2x4xindex>
func.return %2 : vector<2x4xindex>
}
// CHECK-LABEL: func @vector_extract
-// CHECK: test.reflect_bounds {smax = 6 : index, smin = 5 : index, umax = 6 : index, umin = 5 : index}
+// CHECK: test.reflect_bounds <umin = 5 : index, umax = 6 : index, smin = 5 : index, smax = 6 : index>
func.func @vector_extract() -> index {
- %0 = test.with_bounds { umin = 5 : index, umax = 6 : index, smin = 5 : index, smax = 6 : index } : vector<4xindex>
+ %0 = test.with_bounds < umin = 5 : index, umax = 6 : index, smin = 5 : index, smax = 6 : index > : vector<4xindex>
%1 = vector.extract %0[0] : index from vector<4xindex>
%2 = test.reflect_bounds %1 : index
func.return %2 : index
}
// CHECK-LABEL: func @vector_add
-// CHECK: test.reflect_bounds {smax = 12 : index, smin = 10 : index, umax = 12 : index, umin = 10 : index}
+// CHECK: test.reflect_bounds <umin = 10 : index, umax = 12 : index, smin = 10 : index, smax = 12 : index>
func.func @vector_add() -> vector<4xindex> {
- %0 = test.with_bounds { umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index } : vector<4xindex>
- %1 = test.with_bounds { umin = 6 : index, umax = 7 : index, smin = 6 : index, smax = 7 : index } : vector<4xindex>
+ %0 = test.with_bounds < umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index > : vector<4xindex>
+ %1 = test.with_bounds < umin = 6 : index, umax = 7 : index, smin = 6 : index, smax = 7 : index > : vector<4xindex>
%2 = arith.addi %0, %1 : vector<4xindex>
%3 = test.reflect_bounds %2 : vector<4xindex>
func.return %3 : vector<4xindex>
}
// CHECK-LABEL: func @vector_insert
-// CHECK: test.reflect_bounds {smax = 8 : index, smin = 5 : index, umax = 8 : index, umin = 5 : index}
+// CHECK: test.reflect_bounds <umin = 5 : index, umax = 8 : index, smin = 5 : index, smax = 8 : index>
func.func @vector_insert() -> vector<4xindex> {
- %0 = test.with_bounds { umin = 5 : index, umax = 7 : index, smin = 5 : index, smax = 7 : index } : vector<4xindex>
- %1 = test.with_bounds { umin = 6 : index, umax = 8 : index, smin = 6 : index, smax = 8 : index } : index
+ %0 = test.with_bounds < umin = 5 : index, umax = 7 : index, smin = 5 : index, smax = 7 : index > : vector<4xindex>
+ %1 = test.with_bounds < umin = 6 : index, umax = 8 : index, smin = 6 : index, smax = 8 : index > : index
%2 = vector.insert %1, %0[0] : index into vector<4xindex>
%3 = test.reflect_bounds %2 : vector<4xindex>
func.return %3 : vector<4xindex>
@@ -91,7 +91,7 @@ func.func @vector_insert() -> vector<4xindex> {
// CHECK-LABEL: func @test_loaded_vector_extract
// No bounds
-// CHECK: test.reflect_bounds {smax = 2147483647 : si32, smin = -2147483648 : si32, umax = 4294967295 : ui32, umin = 0 : ui32} %{{.*}} : i32
+// CHECK: test.reflect_bounds <umin = 0 : ui32, umax = 4294967295 : ui32, smin = -2147483648 : si32, smax = 2147483647 : si32> %{{.*}} : i32
func.func @test_loaded_vector_extract(%memref : memref<16xi32>) -> i32 {
%c0 = arith.constant 0 : index
%v = vector.load %memref[%c0] : memref<16xi32>, vector<4xi32>
@@ -101,16 +101,16 @@ func.func @test_loaded_vector_extract(%memref : memref<16xi32>) -> i32 {
}
// CHECK-LABEL: func @test_vector_extsi
-// CHECK: test.reflect_bounds {smax = 5 : si32, smin = 1 : si32, umax = 5 : ui32, umin = 1 : ui32}
+// CHECK: test.reflect_bounds <umin = 1 : ui32, umax = 5 : ui32, smin = 1 : si32, smax = 5 : si32>
func.func @test_vector_extsi() -> vector<2xi32> {
- %0 = test.with_bounds {smax = 5 : si8, smin = 1 : si8, umax = 5 : ui8, umin = 1 : ui8 } : vector<2xi8>
+ %0 = test.with_bounds <smax = 5 : si8, smin = 1 : si8, umax = 5 : ui8, umin = 1 : ui8 > : vector<2xi8>
%1 = arith.extsi %0 : vector<2xi8> to vector<2xi32>
%2 = test.reflect_bounds %1 : vector<2xi32>
func.return %2 : vector<2xi32>
}
// CHECK-LABEL: func @vector_step
-// CHECK: test.reflect_bounds {smax = 7 : index, smin = 0 : index, umax = 7 : index, umin = 0 : index}
+// CHECK: test.reflect_bounds <umin = 0 : index, umax = 7 : index, smin = 0 : index, smax = 7 : index>
func.func @vector_step() -> vector<8xindex> {
%0 = vector.step : vector<8xindex>
%1 = test.reflect_bounds %0 : vector<8xindex>
@@ -118,7 +118,7 @@ func.func @vector_step() -> vector<8xindex> {
}
// CHECK-LABEL: func @vector_step_i32
-// CHECK: test.reflect_bounds {smax = 7 : si32, smin = 0 : si32, umax = 7 : ui32, umin = 0 : ui32}
+// CHECK: test.reflect_bounds <umin = 0 : ui32, umax = 7 : ui32, smin = 0 : si32, smax = 7 : si32>
func.func @vector_step_i32() -> vector<8xi32> {
%0 = vector.step : vector<8xi32>
%1 = test.reflect_bounds %0 : vector<8xi32>
@@ -127,7 +127,7 @@ func.func @vector_step_i32() -> vector<8xi32> {
// Boundary: 255 lanes do not wrap, so the upper bound is the last lane (254).
// CHECK-LABEL: func @vector_step_i8_no_wrap_boundary
-// CHECK: test.reflect_bounds {smax = 127 : si8, smin = -128 : si8, umax = 254 : ui8, umin = 0 : ui8}
+// CHECK: test.reflect_bounds <umin = 0 : ui8, umax = 254 : ui8, smin = -128 : si8, smax = 127 : si8>
func.func @vector_step_i8_no_wrap_boundary() -> vector<255xi8> {
%0 = vector.step : vector<255xi8>
%1 = test.reflect_bounds %0 : vector<255xi8>
@@ -136,7 +136,7 @@ func.func @vector_step_i8_no_wrap_boundary() -> vector<255xi8> {
// Boundary: 256 lanes exactly fill the i8 range without wrapping (last lane 255).
// CHECK-LABEL: func @vector_step_i8_exact_fit
-// CHECK: test.reflect_bounds {smax = 127 : si8, smin = -128 : si8, umax = 255 : ui8, umin = 0 : ui8}
+// CHECK: test.reflect_bounds <umin = 0 : ui8, umax = 255 : ui8, smin = -128 : si8, smax = 127 : si8>
func.func @vector_step_i8_exact_fit() -> vector<256xi8> {
%0 = vector.step : vector<256xi8>
%1 = test.reflect_bounds %0 : vector<256xi8>
@@ -145,7 +145,7 @@ func.func @vector_step_i8_exact_fit() -> vector<256xi8> {
// The sequence wraps (300 > 256), so the result spans the entire i8 range.
// CHECK-LABEL: func @vector_step_i8_wrap
-// CHECK: test.reflect_bounds {smax = 127 : si8, smin = -128 : si8, umax = 255 : ui8, umin = 0 : ui8}
+// CHECK: test.reflect_bounds <umin = 0 : ui8, umax = 255 : ui8, smin = -128 : si8, smax = 127 : si8>
func.func @vector_step_i8_wrap() -> vector<300xi8> {
%0 = vector.step : vector<300xi8>
%1 = test.reflect_bounds %0 : vector<300xi8>
diff --git a/mlir/test/IR/attribute.mlir b/mlir/test/IR/attribute.mlir
index 7448ba9ddf6a3..ea378439c0f3b 100644
--- a/mlir/test/IR/attribute.mlir
+++ b/mlir/test/IR/attribute.mlir
@@ -946,7 +946,7 @@ func.func @default_value_printing(%arg0 : i32) {
// CHECK: test.default_value_print %arg0
"test.default_value_print"(%arg0) {"value_with_default" = 0 : i32} : (i32) -> ()
// The attribute SHOULD be printed because it is not equal to the default
- // CHECK: test.default_value_print {value_with_default = 1 : i32} %arg0
+ // CHECK: test.default_value_print <value_with_default = 1> %arg0
"test.default_value_print"(%arg0) {"value_with_default" = 1 : i32} : (i32) -> ()
return
}
diff --git a/mlir/test/Interfaces/InferIntRangeInterface/infer-int-range-test-ops-invalid.mlir b/mlir/test/Interfaces/InferIntRangeInterface/infer-int-range-test-ops-invalid.mlir
index 7392b9d2ec712..73085684d7e0b 100644
--- a/mlir/test/Interfaces/InferIntRangeInterface/infer-int-range-test-ops-invalid.mlir
+++ b/mlir/test/Interfaces/InferIntRangeInterface/infer-int-range-test-ops-invalid.mlir
@@ -5,8 +5,8 @@
// See: https://github.com/llvm/llvm-project/issues/120882
func.func @with_bounds_mismatched_width() -> i8 {
// expected-error at +1 {{'test.with_bounds' op bound attribute width (64) does not match result type width (8)}}
- %0 = test.with_bounds { umin = 10 : i64, umax = 15 : i64,
- smin = 10 : i64, smax = 15 : i64 } : i8
+ %0 = test.with_bounds < umin = 10 : i64, umax = 15 : i64,
+ smin = 10 : i64, smax = 15 : i64 > : i8
%1 = test.reflect_bounds %0 : i8
return %1 : i8
}
diff --git a/mlir/test/Interfaces/InferIntRangeInterface/infer-int-range-test-ops.mlir b/mlir/test/Interfaces/InferIntRangeInterface/infer-int-range-test-ops.mlir
index c6344447d9f74..8908f923ce6ba 100644
--- a/mlir/test/Interfaces/InferIntRangeInterface/infer-int-range-test-ops.mlir
+++ b/mlir/test/Interfaces/InferIntRangeInterface/infer-int-range-test-ops.mlir
@@ -4,8 +4,8 @@
// CHECK: %[[cst:.*]] = "test.constant"() <{value = 3 : index}
// CHECK: return %[[cst]]
func.func @constant() -> index {
- %0 = test.with_bounds { umin = 3 : index, umax = 3 : index,
- smin = 3 : index, smax = 3 : index} : index
+ %0 = test.with_bounds < umin = 3 : index, umax = 3 : index,
+ smin = 3 : index, smax = 3 : index> : index
func.return %0 : index
}
@@ -13,16 +13,16 @@ func.func @constant() -> index {
// CHECK: %[[cst:.*]] = "test.constant"() <{value = 4 : index}
// CHECK: return %[[cst]]
func.func @increment() -> index {
- %0 = test.with_bounds { umin = 3 : index, umax = 3 : index, smin = 0 : index, smax = 0x7fffffffffffffff : index } : index
+ %0 = test.with_bounds < umin = 3 : index, umax = 3 : index, smin = 0 : index, smax = 0x7fffffffffffffff : index > : index
%1 = test.increment %0 : index
func.return %1 : index
}
// CHECK-LABEL: func @maybe_increment
-// CHECK: test.reflect_bounds {smax = 4 : index, smin = 3 : index, umax = 4 : index, umin = 3 : index}
+// CHECK: test.reflect_bounds <umin = 3 : index, umax = 4 : index, smin = 3 : index, smax = 4 : index>
func.func @maybe_increment(%arg0 : i1) -> index {
- %0 = test.with_bounds { umin = 3 : index, umax = 3 : index,
- smin = 3 : index, smax = 3 : index} : index
+ %0 = test.with_bounds < umin = 3 : index, umax = 3 : index,
+ smin = 3 : index, smax = 3 : index> : index
%1 = scf.if %arg0 -> index {
scf.yield %0 : index
} else {
@@ -34,10 +34,10 @@ func.func @maybe_increment(%arg0 : i1) -> index {
}
// CHECK-LABEL: func @maybe_increment_br
-// CHECK: test.reflect_bounds {smax = 4 : index, smin = 3 : index, umax = 4 : index, umin = 3 : index}
+// CHECK: test.reflect_bounds <umin = 3 : index, umax = 4 : index, smin = 3 : index, smax = 4 : index>
func.func @maybe_increment_br(%arg0 : i1) -> index {
- %0 = test.with_bounds { umin = 3 : index, umax = 3 : index,
- smin = 3 : index, smax = 3 : index} : index
+ %0 = test.with_bounds < umin = 3 : index, umax = 3 : index,
+ smin = 3 : index, smax = 3 : index> : index
cf.cond_br %arg0, ^bb0, ^bb1
^bb0:
%1 = test.increment %0 : index
@@ -50,14 +50,14 @@ func.func @maybe_increment_br(%arg0 : i1) -> index {
}
// CHECK-LABEL: func @for_bounds
-// CHECK: test.reflect_bounds {smax = 1 : index, smin = 0 : index, umax = 1 : index, umin = 0 : index}
+// CHECK: test.reflect_bounds <umin = 0 : index, umax = 1 : index, smin = 0 : index, smax = 1 : index>
func.func @for_bounds() -> index {
- %c0 = test.with_bounds { umin = 0 : index, umax = 0 : index,
- smin = 0 : index, smax = 0 : index} : index
- %c1 = test.with_bounds { umin = 1 : index, umax = 1 : index,
- smin = 1 : index, smax = 1 : index} : index
- %c2 = test.with_bounds { umin = 2 : index, umax = 2 : index,
- smin = 2 : index, smax = 2 : index} : index
+ %c0 = test.with_bounds < umin = 0 : index, umax = 0 : index,
+ smin = 0 : index, smax = 0 : index> : index
+ %c1 = test.with_bounds < umin = 1 : index, umax = 1 : index,
+ smin = 1 : index, smax = 1 : index> : index
+ %c2 = test.with_bounds < umin = 2 : index, umax = 2 : index,
+ smin = 2 : index, smax = 2 : index> : index
%0 = scf.for %arg0 = %c0 to %c2 step %c1 iter_args(%arg2 = %c0) -> index {
scf.yield %arg0 : index
@@ -67,14 +67,14 @@ func.func @for_bounds() -> index {
}
// CHECK-LABEL: func @no_analysis_of_loop_variants
-// CHECK: test.reflect_bounds {smax = 9223372036854775807 : index, smin = -9223372036854775808 : index, umax = -1 : index, umin = 0 : index}
+// CHECK: test.reflect_bounds <umin = 0 : index, umax = -1 : index, smin = -9223372036854775808 : index, smax = 9223372036854775807 : index>
func.func @no_analysis_of_loop_variants() -> index {
- %c0 = test.with_bounds { umin = 0 : index, umax = 0 : index,
- smin = 0 : index, smax = 0 : index} : index
- %c1 = test.with_bounds { umin = 1 : index, umax = 1 : index,
- smin = 1 : index, smax = 1 : index} : index
- %c2 = test.with_bounds { umin = 2 : index, umax = 2 : index,
- smin = 2 : index, smax = 2 : index} : index
+ %c0 = test.with_bounds < umin = 0 : index, umax = 0 : index,
+ smin = 0 : index, smax = 0 : index> : index
+ %c1 = test.with_bounds < umin = 1 : index, umax = 1 : index,
+ smin = 1 : index, smax = 1 : index> : index
+ %c2 = test.with_bounds < umin = 2 : index, umax = 2 : index,
+ smin = 2 : index, smax = 2 : index> : index
%0 = scf.for %arg0 = %c0 to %c2 step %c1 iter_args(%arg2 = %c0) -> index {
%1 = test.increment %arg2 : index
@@ -85,7 +85,7 @@ func.func @no_analysis_of_loop_variants() -> index {
}
// CHECK-LABEL: func @region_args
-// CHECK: test.reflect_bounds {smax = 4 : index, smin = 3 : index, umax = 4 : index, umin = 3 : index}
+// CHECK: test.reflect_bounds <umin = 3 : index, umax = 4 : index, smin = 3 : index, smax = 4 : index>
func.func @region_args() {
test.with_bounds_region { umin = 3 : index, umax = 4 : index,
smin = 3 : index, smax = 4 : index } %arg0 : index {
@@ -95,7 +95,7 @@ func.func @region_args() {
}
// CHECK-LABEL: func @func_args_unbound
-// CHECK: test.reflect_bounds {smax = 9223372036854775807 : index, smin = -9223372036854775808 : index, umax = -1 : index, umin = 0 : index}
+// CHECK: test.reflect_bounds <umin = 0 : index, umax = -1 : index, smin = -9223372036854775808 : index, smax = 9223372036854775807 : index>
func.func @func_args_unbound(%arg0 : index) -> index {
%0 = test.reflect_bounds %arg0 : index
func.return %0 : index
@@ -104,8 +104,8 @@ func.func @func_args_unbound(%arg0 : index) -> index {
// CHECK-LABEL: func @propagate_across_while_loop_false()
func.func @propagate_across_while_loop_false() -> index {
// CHECK: %[[C1:.*]] = "test.constant"() <{value = 1
- %0 = test.with_bounds { umin = 0 : index, umax = 0 : index,
- smin = 0 : index, smax = 0 : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 0 : index,
+ smin = 0 : index, smax = 0 : index > : index
%1 = scf.while : () -> index {
%false = arith.constant false
scf.condition(%false) %0 : index
@@ -121,8 +121,8 @@ func.func @propagate_across_while_loop_false() -> index {
// CHECK-LABEL: func @propagate_across_while_loop
func.func @propagate_across_while_loop(%arg0 : i1) -> index {
// CHECK: %[[C1:.*]] = "test.constant"() <{value = 1
- %0 = test.with_bounds { umin = 0 : index, umax = 0 : index,
- smin = 0 : index, smax = 0 : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 0 : index,
+ smin = 0 : index, smax = 0 : index > : index
%1 = scf.while : () -> index {
scf.condition(%arg0) %0 : index
} do {
@@ -137,8 +137,8 @@ func.func @propagate_across_while_loop(%arg0 : i1) -> index {
// CHECK-LABEL: func @dont_propagate_across_infinite_loop()
func.func @dont_propagate_across_infinite_loop() -> index {
// CHECK: %[[C0:.*]] = "test.constant"() <{value = 0
- %0 = test.with_bounds { umin = 0 : index, umax = 0 : index,
- smin = 0 : index, smax = 0 : index } : index
+ %0 = test.with_bounds < umin = 0 : index, umax = 0 : index,
+ smin = 0 : index, smax = 0 : index > : index
// CHECK: %[[loopRes:.*]] = scf.while
%1 = scf.while : () -> index {
%true = arith.constant true
@@ -187,14 +187,14 @@ func.func @propagate_from_block_to_iterarg(%arg0: index, %arg1: i1) {
// CHECK-LABEL: func @multiple_loop_ivs
func.func @multiple_loop_ivs(%arg0: memref<?x64xi32>) {
- %ub1 = test.with_bounds { umin = 1 : index, umax = 32 : index,
- smin = 1 : index, smax = 32 : index } : index
+ %ub1 = test.with_bounds < umin = 1 : index, umax = 32 : index,
+ smin = 1 : index, smax = 32 : index > : index
%c0_i32 = arith.constant 0 : i32
// CHECK: scf.forall
scf.forall (%arg1, %arg2) in (%ub1, 64) {
- // CHECK: test.reflect_bounds {smax = 31 : index, smin = 0 : index, umax = 31 : index, umin = 0 : index}
+ // CHECK: test.reflect_bounds <umin = 0 : index, umax = 31 : index, smin = 0 : index, smax = 31 : index>
%1 = test.reflect_bounds %arg1 : index
- // CHECK-NEXT: test.reflect_bounds {smax = 63 : index, smin = 0 : index, umax = 63 : index, umin = 0 : index}
+ // CHECK-NEXT: test.reflect_bounds <umin = 0 : index, umax = 63 : index, smin = 0 : index, smax = 63 : index>
%2 = test.reflect_bounds %arg2 : index
memref.store %c0_i32, %arg0[%1, %2] : memref<?x64xi32>
}
diff --git a/mlir/test/Transforms/loop-invariant-code-motion.mlir b/mlir/test/Transforms/loop-invariant-code-motion.mlir
index 31a4f64dd7de0..ac4fbf72a12d3 100644
--- a/mlir/test/Transforms/loop-invariant-code-motion.mlir
+++ b/mlir/test/Transforms/loop-invariant-code-motion.mlir
@@ -1135,7 +1135,7 @@ func.func @speculate_ceildivsi_const(
func.func @no_speculate_divui_range(
// CHECK-LABEL: @no_speculate_divui_range(
%num: i8, %lb: index, %ub: index, %step: index) {
- %denom = test.with_bounds {smax = 127 : i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8} : i8
+ %denom = test.with_bounds <smax = 127 : i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8> : i8
scf.for %i = %lb to %ub step %step {
// CHECK: scf.for
// CHECK: arith.divui
@@ -1148,7 +1148,7 @@ func.func @no_speculate_divui_range(
func.func @no_speculate_udiv_range(
// CHECK-LABEL: @no_speculate_udiv_range(
%num: i8, %lb: index, %ub: index, %step: index) {
- %denom = test.with_bounds {smax = 127 : i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8} : i8
+ %denom = test.with_bounds <smax = 127 : i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8> : i8
scf.for %i = %lb to %ub step %step {
// CHECK: scf.for
// CHECK: llvm.udiv
@@ -1161,8 +1161,8 @@ func.func @no_speculate_udiv_range(
func.func @no_speculate_divsi_range(
// CHECK-LABEL: @no_speculate_divsi_range(
%num: i8, %lb: index, %ub: index, %step: index) {
- %denom0 = test.with_bounds {smax = -1: i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8} : i8
- %denom1 = test.with_bounds {smax = 127 : i8, smin = 0 : i8, umax = 255 : i8, umin = 0 : i8} : i8
+ %denom0 = test.with_bounds <smax = -1: i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8> : i8
+ %denom1 = test.with_bounds <smax = 127 : i8, smin = 0 : i8, umax = 255 : i8, umin = 0 : i8> : i8
scf.for %i = %lb to %ub step %step {
// CHECK: scf.for
// CHECK-COUNT-2: arith.divsi
@@ -1176,8 +1176,8 @@ func.func @no_speculate_divsi_range(
func.func @no_speculate_sdiv_range(
// CHECK-LABEL: @no_speculate_sdiv_range(
%num: i8, %lb: index, %ub: index, %step: index) {
- %denom0 = test.with_bounds {smax = -1: i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8} : i8
- %denom1 = test.with_bounds {smax = 127 : i8, smin = 0 : i8, umax = 255 : i8, umin = 0 : i8} : i8
+ %denom0 = test.with_bounds <smax = -1: i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8> : i8
+ %denom1 = test.with_bounds <smax = 127 : i8, smin = 0 : i8, umax = 255 : i8, umin = 0 : i8> : i8
scf.for %i = %lb to %ub step %step {
// CHECK: scf.for
// CHECK-COUNT-2: llvm.sdiv
@@ -1191,7 +1191,7 @@ func.func @no_speculate_sdiv_range(
func.func @no_speculate_ceildivui_range(
// CHECK-LABEL: @no_speculate_ceildivui_range(
%num: i8, %lb: index, %ub: index, %step: index) {
- %denom = test.with_bounds {smax = 127 : i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8} : i8
+ %denom = test.with_bounds <smax = 127 : i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8> : i8
scf.for %i = %lb to %ub step %step {
// CHECK: scf.for
// CHECK: arith.ceildivui
@@ -1204,8 +1204,8 @@ func.func @no_speculate_ceildivui_range(
func.func @no_speculate_ceildivsi_range(
// CHECK-LABEL: @no_speculate_ceildivsi_range(
%num: i8, %lb: index, %ub: index, %step: index) {
- %denom0 = test.with_bounds {smax = -1 : i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8} : i8
- %denom1 = test.with_bounds {smax = 127 : i8, smin = 0 : i8, umax = 255 : i8, umin = 0 : i8} : i8
+ %denom0 = test.with_bounds <smax = -1 : i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8> : i8
+ %denom1 = test.with_bounds <smax = 127 : i8, smin = 0 : i8, umax = 255 : i8, umin = 0 : i8> : i8
scf.for %i = %lb to %ub step %step {
// CHECK: scf.for
// CHECK-COUNT-2: arith.ceildivsi
@@ -1219,7 +1219,7 @@ func.func @no_speculate_ceildivsi_range(
func.func @speculate_divui_range(
// CHECK-LABEL: @speculate_divui_range(
%num: i8, %lb: index, %ub: index, %step: index) {
- %denom = test.with_bounds {smax = 127 : i8, smin = -128 : i8, umax = 255 : i8, umin = 1 : i8} : i8
+ %denom = test.with_bounds <smax = 127 : i8, smin = -128 : i8, umax = 255 : i8, umin = 1 : i8> : i8
scf.for %i = %lb to %ub step %step {
// CHECK: arith.divui
// CHECK: scf.for
@@ -1232,7 +1232,7 @@ func.func @speculate_divui_range(
func.func @speculate_udiv_range(
// CHECK-LABEL: @speculate_udiv_range(
%num: i8, %lb: index, %ub: index, %step: index) {
- %denom = test.with_bounds {smax = 127 : i8, smin = -128 : i8, umax = 255 : i8, umin = 1 : i8} : i8
+ %denom = test.with_bounds <smax = 127 : i8, smin = -128 : i8, umax = 255 : i8, umin = 1 : i8> : i8
scf.for %i = %lb to %ub step %step {
// CHECK: llvm.udiv
// CHECK: scf.for
@@ -1245,8 +1245,8 @@ func.func @speculate_udiv_range(
func.func @speculate_divsi_range(
// CHECK-LABEL: @speculate_divsi_range(
%num: i8, %lb: index, %ub: index, %step: index) {
- %denom0 = test.with_bounds {smax = 127 : i8, smin = 1 : i8, umax = 255 : i8, umin = 0 : i8} : i8
- %denom1 = test.with_bounds {smax = -2 : i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8} : i8
+ %denom0 = test.with_bounds <smax = 127 : i8, smin = 1 : i8, umax = 255 : i8, umin = 0 : i8> : i8
+ %denom1 = test.with_bounds <smax = -2 : i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8> : i8
scf.for %i = %lb to %ub step %step {
// CHECK-COUNT-2: arith.divsi
// CHECK: scf.for
@@ -1261,8 +1261,8 @@ func.func @speculate_divsi_range(
func.func @speculate_sdiv_range(
// CHECK-LABEL: @speculate_sdiv_range(
%num: i8, %lb: index, %ub: index, %step: index) {
- %denom0 = test.with_bounds {smax = 127 : i8, smin = 1 : i8, umax = 255 : i8, umin = 0 : i8} : i8
- %denom1 = test.with_bounds {smax = -2 : i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8} : i8
+ %denom0 = test.with_bounds <smax = 127 : i8, smin = 1 : i8, umax = 255 : i8, umin = 0 : i8> : i8
+ %denom1 = test.with_bounds <smax = -2 : i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8> : i8
scf.for %i = %lb to %ub step %step {
// CHECK-COUNT-2: llvm.sdiv
// CHECK: scf.for
@@ -1277,7 +1277,7 @@ func.func @speculate_sdiv_range(
func.func @speculate_ceildivui_range(
// CHECK-LABEL: @speculate_ceildivui_range(
%num: i8, %lb: index, %ub: index, %step: index) {
- %denom = test.with_bounds {smax = 127 : i8, smin = -128 : i8, umax = 255 : i8, umin = 1 : i8} : i8
+ %denom = test.with_bounds <smax = 127 : i8, smin = -128 : i8, umax = 255 : i8, umin = 1 : i8> : i8
scf.for %i = %lb to %ub step %step {
// CHECK: arith.ceildivui
// CHECK: scf.for
@@ -1290,8 +1290,8 @@ func.func @speculate_ceildivui_range(
func.func @speculate_ceildivsi_range(
// CHECK-LABEL: @speculate_ceildivsi_range(
%num: i8, %lb: index, %ub: index, %step: index) {
- %denom0 = test.with_bounds {smax = 127 : i8, smin = 1 : i8, umax = 255 : i8, umin = 0 : i8} : i8
- %denom1 = test.with_bounds {smax = -2 : i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8} : i8
+ %denom0 = test.with_bounds <smax = 127 : i8, smin = 1 : i8, umax = 255 : i8, umin = 0 : i8> : i8
+ %denom1 = test.with_bounds <smax = -2 : i8, smin = -128 : i8, umax = 255 : i8, umin = 0 : i8> : i8
scf.for %i = %lb to %ub step %step {
// CHECK-COUNT-2: arith.ceildivsi
// CHECK: scf.for
diff --git a/mlir/test/Transforms/make-composed-folded-affine-apply.mlir b/mlir/test/Transforms/make-composed-folded-affine-apply.mlir
index 138426827fa15..d81ddeb4f950a 100644
--- a/mlir/test/Transforms/make-composed-folded-affine-apply.mlir
+++ b/mlir/test/Transforms/make-composed-folded-affine-apply.mlir
@@ -16,8 +16,8 @@
// Check the apply gets simplified.
// CHECK: @apply_simplification
func.func @apply_simplification_1() -> index {
- %0 = test.value_with_bounds {max = 64 : index, min = 32 : index}
- %1 = test.value_with_bounds {max = 64 : index, min = 32 : index}
+ %0 = test.value_with_bounds <min = 32, max = 64>
+ %1 = test.value_with_bounds <min = 32, max = 64>
%2 = affine.min #map()[%0, %1]
// CHECK-NOT: affine.apply
// CHECK: arith.constant 2 : index
@@ -27,9 +27,9 @@ func.func @apply_simplification_1() -> index {
// Check the simplification can match non-trivial affine expressions like s1 + s2.
func.func @apply_simplification_2() -> index {
- %0 = test.value_with_bounds {max = 64 : index, min = 32 : index}
- %1 = test.value_with_bounds {max = 64 : index, min = 32 : index}
- %2 = test.value_with_bounds {max = 64 : index, min = 32 : index}
+ %0 = test.value_with_bounds <min = 32, max = 64>
+ %1 = test.value_with_bounds <min = 32, max = 64>
+ %2 = test.value_with_bounds <min = 32, max = 64>
%3 = affine.min #map2()[%0, %1, %2]
// CHECK-NOT: affine.apply
// CHECK: arith.constant 4 : index
@@ -41,13 +41,13 @@ func.func @apply_simplification_2() -> index {
// The apply cannot be simplified because `s1 = %0` doesn't appear in the input min.
// CHECK: @no_simplification_0
func.func @no_simplification_0() -> index {
- // CHECK: %[[V0:.*]] = test.value_with_bounds {max = 64 : index, min = 32 : index}
- // CHECK: %[[V1:.*]] = test.value_with_bounds {max = 64 : index, min = 16 : index}
+ // CHECK: %[[V0:.*]] = test.value_with_bounds <min = 32, max = 64>
+ // CHECK: %[[V1:.*]] = test.value_with_bounds <min = 16, max = 64>
// CHECK: %[[V2:.*]] = affine.min #{{.*}}()[%[[V0]], %[[V1]]]
// CHECK: %[[V3:.*]] = affine.apply #[[MAP5]]()[%[[V2]], %[[V0]]]
// CHECK: return %[[V3]] : index
- %0 = test.value_with_bounds {max = 64 : index, min = 32 : index}
- %1 = test.value_with_bounds {max = 64 : index, min = 16 : index}
+ %0 = test.value_with_bounds <min = 32, max = 64>
+ %1 = test.value_with_bounds <min = 16, max = 64>
%2 = affine.min #map4()[%0, %1]
%3 = affine.apply #map5()[%2, %0]
return %3 : index
@@ -56,13 +56,13 @@ func.func @no_simplification_0() -> index {
// The apply cannot be simplified because the min cannot be proven to be greater than 0.
// CHECK: @no_simplification_1
func.func @no_simplification_1() -> index {
- // CHECK: %[[V0:.*]] = test.value_with_bounds {max = 64 : index, min = 32 : index}
- // CHECK: %[[V1:.*]] = test.value_with_bounds {max = 64 : index, min = 16 : index}
+ // CHECK: %[[V0:.*]] = test.value_with_bounds <min = 32, max = 64>
+ // CHECK: %[[V1:.*]] = test.value_with_bounds <min = 16, max = 64>
// CHECK: %[[V2:.*]] = affine.min #{{.*}}()[%[[V0]], %[[V1]]]
// CHECK: %[[V3:.*]] = affine.apply #[[MAP1]]()[%[[V2]], %[[V1]]]
// CHECK: return %[[V3]] : index
- %0 = test.value_with_bounds {max = 64 : index, min = 32 : index}
- %1 = test.value_with_bounds {max = 64 : index, min = 16 : index}
+ %0 = test.value_with_bounds <min = 32, max = 64>
+ %1 = test.value_with_bounds <min = 16, max = 64>
%2 = affine.min #map6()[%0, %1]
%3 = affine.apply #map1()[%2, %1]
return %3 : index
diff --git a/mlir/test/lib/Dialect/Test/TestDialect.td b/mlir/test/lib/Dialect/Test/TestDialect.td
index ed0bedaced126..37a263f1d10b8 100644
--- a/mlir/test/lib/Dialect/Test/TestDialect.td
+++ b/mlir/test/lib/Dialect/Test/TestDialect.td
@@ -13,8 +13,6 @@ include "mlir/IR/OpBase.td"
def Test_Dialect : Dialect {
let name = "test";
- // Keep legacy assembly format coverage for test operations.
- let useStrictPropertiesInAssemblyFormat = 0;
let cppNamespace = "::test";
let hasCanonicalizer = 1;
let hasConstantMaterializer = 1;
diff --git a/mlir/test/lib/Dialect/Test/TestOps.td b/mlir/test/lib/Dialect/Test/TestOps.td
index c3e06dc968a27..a6b05b55804b7 100644
--- a/mlir/test/lib/Dialect/Test/TestOps.td
+++ b/mlir/test/lib/Dialect/Test/TestOps.td
@@ -2386,7 +2386,7 @@ def TestCrashingReturnOp : TEST_Op<"crashing_return", [
["getMutableSuccessorOperands"]>,
Terminator]> {
let arguments = (ins Variadic<AnyType>:$args, UnitAttr:$valid);
- let assemblyFormat = "($args^ `:` type($args))? attr-dict";
+ let assemblyFormat = "($args^ `:` type($args))? prop-dict attr-dict";
let hasVerifier = 1;
}
def TestCastOp : TEST_Op<"cast">,
@@ -2733,12 +2733,12 @@ def TestValueWithBoundsOp : TEST_Op<"value_with_bounds", [
Example:
```mlir
- %0 = test.value_with_bounds { min = 4 : index, max = 5 : index}
+ %0 = test.value_with_bounds <min = 4, max = 5>
```
}];
let arguments = (ins IndexAttr:$min, IndexAttr:$max);
let results = (outs Index:$result);
- let assemblyFormat = "attr-dict";
+ let assemblyFormat = "prop-dict attr-dict";
}
@@ -3422,12 +3422,12 @@ def TestNVVMRequiresSMArchOrFamilyCondOp :
def TestDefaultStrAttrNoValueOp : TEST_Op<"no_str_value"> {
let arguments = (ins DefaultValuedAttr<StrAttr, "">:$value);
- let assemblyFormat = "attr-dict";
+ let assemblyFormat = "prop-dict attr-dict";
}
def TestDefaultStrAttrHasValueOp : TEST_Op<"has_str_value"> {
let arguments = (ins DefaultValuedStrAttr<StrAttr, "">:$value);
- let assemblyFormat = "attr-dict";
+ let assemblyFormat = "prop-dict attr-dict";
}
def : Pat<(TestDefaultStrAttrNoValueOp $value),
@@ -3463,7 +3463,7 @@ def : Pat<(TestVariadicRewriteSrcOp $arg, $brg, $crg),
def TestDefaultAttrPrintOp : TEST_Op<"default_value_print"> {
let arguments = (ins DefaultValuedAttr<I32Attr, "0">:$value_with_default,
I32:$operand);
- let assemblyFormat = "attr-dict $operand";
+ let assemblyFormat = "prop-dict attr-dict $operand";
}
//===----------------------------------------------------------------------===//
@@ -3554,7 +3554,7 @@ def TestWithBoundsOp : TEST_Op<"with_bounds",
Example:
```mlir
- %0 = test.with_bounds { umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index } : index
+ %0 = test.with_bounds <umin = 4 : index, umax = 5 : index, smin = 4 : index, smax = 5 : index> : index
```
}];
@@ -3564,7 +3564,7 @@ def TestWithBoundsOp : TEST_Op<"with_bounds",
APIntAttr:$smax);
let results = (outs InferIntRangeType:$fakeVal);
- let assemblyFormat = "attr-dict `:` type($fakeVal)";
+ let assemblyFormat = "prop-dict attr-dict `:` type($fakeVal)";
}
def TestWithoutBoundsOp : TEST_Op<"without_bounds",
@@ -3615,7 +3615,7 @@ def TestReflectBoundsOp : TEST_Op<"reflect_bounds",
Example:
```mlir
- CHECK: test.reflect_bounds {smax = 7 : index, smin = 0 : index, umax = 7 : index, umin = 0 : index}
+ CHECK: test.reflect_bounds <umin = 0 : index, umax = 7 : index, smin = 0 : index, smax = 7 : index>
%1 = test.reflect_bounds %0 : index
```
}];
@@ -3627,7 +3627,7 @@ def TestReflectBoundsOp : TEST_Op<"reflect_bounds",
OptionalAttr<APIntAttr>:$smax);
let results = (outs InferIntRangeType:$result);
- let assemblyFormat = "attr-dict $value `:` type($result)";
+ let assemblyFormat = "prop-dict attr-dict $value `:` type($result)";
}
//===----------------------------------------------------------------------===//
@@ -4430,7 +4430,7 @@ def TestCallAndStoreOp : TEST_Op<"call_and_store",
Variadic<AnyType>:$results
);
let assemblyFormat =
- "$callee `(` $callee_operands `)` `,` $address attr-dict "
+ "$callee `(` $callee_operands `)` `,` $address prop-dict attr-dict "
"`:` functional-type(operands, results)";
}
@@ -4448,7 +4448,7 @@ def TestCallOnDeviceOp : TEST_Op<"call_on_device",
);
let assemblyFormat =
"$callee `(` $forwarded_operands `)` `,` $non_forwarded_device_operand "
- "attr-dict `:` functional-type(operands, results)";
+ "prop-dict attr-dict `:` functional-type(operands, results)";
}
// A call-like operation whose first result is produced by the operation itself
@@ -4468,7 +4468,7 @@ def TestCallAndProduceOp : TEST_Op<"call_and_produce",
Variadic<AnyType>:$forwarded
);
let assemblyFormat =
- "$callee `(` $forwarded_operands `)` attr-dict "
+ "$callee `(` $forwarded_operands `)` prop-dict attr-dict "
"`:` functional-type(operands, results)";
}
@@ -4481,7 +4481,7 @@ def TestStoreWithARegion : TEST_Op<"store_with_a_region",
);
let regions = (region AnyRegion:$body);
let assemblyFormat =
- "$address attr-dict-with-keyword regions `:` type($address)";
+ "$address prop-dict attr-dict-with-keyword regions `:` type($address)";
}
def TestStoreWithALoopRegion : TEST_Op<"store_with_a_loop_region",
@@ -4493,7 +4493,7 @@ def TestStoreWithALoopRegion : TEST_Op<"store_with_a_loop_region",
);
let regions = (region AnyRegion:$body);
let assemblyFormat =
- "$address attr-dict-with-keyword regions `:` type($address)";
+ "$address prop-dict attr-dict-with-keyword regions `:` type($address)";
}
def TestStoreWithARegionTerminator : TEST_Op<"store_with_a_region_terminator",
diff --git a/mlir/test/lib/Dialect/Test/TestOpsSyntax.td b/mlir/test/lib/Dialect/Test/TestOpsSyntax.td
index 3ab3c8b9db6e7..8b0ded9ad27d9 100644
--- a/mlir/test/lib/Dialect/Test/TestOpsSyntax.td
+++ b/mlir/test/lib/Dialect/Test/TestOpsSyntax.td
@@ -184,7 +184,7 @@ def FormatLiteralOp : TEST_Op<"format_literal_op"> {
}];
}
-// Test that we elide attributes that are within the syntax.
+// Test that inherent and discardable attributes use separate dictionaries.
def FormatAttrOp : TEST_Op<"format_attr_op"> {
let arguments = (ins I64Attr:$attr);
let assemblyFormat = "$attr attr-dict";
@@ -221,7 +221,7 @@ def FormatOptSymbolRefAttrOp : TEST_Op<"format_opt_symbol_ref_attr_op"> {
// Test that we elide attributes that are within the syntax.
def FormatAttrDictWithKeywordOp : TEST_Op<"format_attr_dict_w_keyword"> {
let arguments = (ins I64Attr:$attr, OptionalAttr<I64Attr>:$opt_attr);
- let assemblyFormat = "attr-dict-with-keyword";
+ let assemblyFormat = "prop-dict attr-dict-with-keyword";
}
// Test that we don't need to provide types in the format if they are buildable.
@@ -637,7 +637,7 @@ def FormatCustomDirectiveAttrDict
: TEST_Op<"format_custom_directive_attrdict"> {
let arguments = (ins I64Attr:$attr, OptionalAttr<I64Attr>:$optAttr);
let assemblyFormat = [{
- custom<CustomDirectiveAttrDict>( attr-dict )
+ prop-dict custom<CustomDirectiveAttrDict>( attr-dict )
}];
}
diff --git a/mlir/test/mlir-tblgen/constant-str-attr-invalid.mlir b/mlir/test/mlir-tblgen/constant-str-attr-invalid.mlir
index 1d0e62d5dec09..d0106b4ea3151 100644
--- a/mlir/test/mlir-tblgen/constant-str-attr-invalid.mlir
+++ b/mlir/test/mlir-tblgen/constant-str-attr-invalid.mlir
@@ -1,4 +1,4 @@
// RUN: mlir-opt -verify-diagnostics %s
// Test DefaultValuedAttr<StrAttr, ""> is recognized as "no default value"
-test.no_str_value {} // expected-error {{'test.no_str_value' op requires attribute 'value'}}
+test.no_str_value // expected-error {{properties dictionary is missing required attribute: value}}
diff --git a/mlir/test/mlir-tblgen/op-format-invalid.td b/mlir/test/mlir-tblgen/op-format-invalid.td
index dc0d1688cc8ec..1b78d8177091f 100644
--- a/mlir/test/mlir-tblgen/op-format-invalid.td
+++ b/mlir/test/mlir-tblgen/op-format-invalid.td
@@ -8,8 +8,6 @@ include "mlir/Interfaces/InferTypeOpInterface.td"
def TestDialect : Dialect {
let name = "test";
- // Exercise the legacy assembly format behavior.
- let useStrictPropertiesInAssemblyFormat = 0;
}
def TestStrictPropertiesDialect : Dialect {
let name = "test_strict_properties";
diff --git a/mlir/test/mlir-tblgen/op-format.mlir b/mlir/test/mlir-tblgen/op-format.mlir
index 1c1e2c68ce0d3..d01273b5384c2 100644
--- a/mlir/test/mlir-tblgen/op-format.mlir
+++ b/mlir/test/mlir-tblgen/op-format.mlir
@@ -40,11 +40,11 @@ test.format_opt_symbol_name_attr_op
test.format_opt_symbol_ref_attr_op @foo {test.unit}
test.format_opt_symbol_ref_attr_op {test.unit}
-// CHECK: test.format_attr_dict_w_keyword attributes {attr = 10 : i64}
-test.format_attr_dict_w_keyword attributes {attr = 10 : i64}
+// CHECK: test.format_attr_dict_w_keyword <attr = 10>
+test.format_attr_dict_w_keyword <attr = 10>
-// CHECK: test.format_attr_dict_w_keyword attributes {attr = 10 : i64, opt_attr = 10 : i64}
-test.format_attr_dict_w_keyword attributes {attr = 10 : i64, opt_attr = 10 : i64}
+// CHECK: test.format_attr_dict_w_keyword <attr = 10, opt_attr = 10> attributes {tag = "test"}
+test.format_attr_dict_w_keyword <attr = 10, opt_attr = 10> attributes {tag = "test"}
// CHECK: test.format_buildable_type_op %[[I64]]
%ignored = test.format_buildable_type_op %i64
diff --git a/mlir/test/mlir-tblgen/op-format.td b/mlir/test/mlir-tblgen/op-format.td
index f98632886449d..679d45b7fa0de 100644
--- a/mlir/test/mlir-tblgen/op-format.td
+++ b/mlir/test/mlir-tblgen/op-format.td
@@ -6,8 +6,6 @@ include "mlir/IR/EnumAttr.td"
def TestDialect : Dialect {
let name = "test";
- // Exercise the legacy assembly format behavior.
- let useStrictPropertiesInAssemblyFormat = 0;
}
def TestStrictPropertiesDialect : Dialect {
let name = "test_strict_properties";
@@ -33,32 +31,11 @@ class TestStrictPropertiesFormat_Op<string fmt, list<Trait> traits = []>
// CHECK: // coverity[ARRAY_VS_SINGLETON]
// CHECK-NEXT: if (parser.resolveOperands(
// CHECK-LABEL: AttrDictContextOperand::print
-// CHECK: ::llvm::ArrayRef<::llvm::StringRef> elidedAttrs;
-// CHECK: _odsPrinter.printOptionalAttrDict(_odsAttrs.getDictionary((*this)->getContext()).getValue(), elidedAttrs);
+// CHECK: _odsPrinter.printOptionalAttrDict((*this)->getDiscardableAttrDictionary().getValue());
def AttrDictContextOperand : TestFormat_Op<[{
$context `:` type($context) attr-dict
}]>, Arguments<(ins AnyType:$context)>;
-// CHECK-LABEL: AttrDictDefaultInherentAttr::parse
-// CHECK: verifyInherentAttrs
-def AttrDictDefaultInherentAttr : TestFormat_Op<[{
- $attr attr-dict
-}]>, Arguments<(ins I64Attr:$attr)>;
-
-// CHECK-LABEL: AttrDictElideDefaultAndFixedAttr::print
-// CHECK: ::llvm::SmallVector<::llvm::StringRef, 2> elidedAttrs = {"used"};
-// CHECK: elidedAttrs.push_back("attr");
-def AttrDictElideDefaultAndFixedAttr : TestFormat_Op<"$used attr-dict">,
- Arguments<(ins I64Attr:$used,
- DefaultValuedStrAttr<StrAttr, "default">:$attr)>;
-
-// CHECK-LABEL: AttrDictElideDefaultAttr::print
-// CHECK: ::llvm::SmallVector<::llvm::StringRef, 2> elidedAttrs = {};
-// CHECK: if(attr && (attr ==
-// CHECK: elidedAttrs.push_back("attr");
-def AttrDictElideDefaultAttr : TestFormat_Op<"attr-dict">,
- Arguments<(ins DefaultValuedStrAttr<StrAttr, "default">:$attr)>;
-
// CHECK-LABEL: AttrDictStrictDefaultAttr::print
// CHECK-NOT: elidedAttrs
// CHECK: _odsPrinter.printOptionalAttrDict((*this)->getDiscardableAttrDictionary().getValue());
@@ -109,7 +86,7 @@ def AttrDictStrictPropDictInherentAttr : TestStrictPropertiesFormat_Op<[{
// CHECK: parseFoo({{.*}}, parser.getBuilder().getI1Type())
// CHECK-LABEL: CustomStringLiteralA::print
// CHECK: printFoo({{.*}}, ::mlir::Builder(getContext()).getI1Type())
-// CHECK: ::llvm::ArrayRef<::llvm::StringRef> elidedAttrs;
+// CHECK: _odsPrinter.printOptionalAttrDict((*this)->getDiscardableAttrDictionary().getValue());
def CustomStringLiteralA : TestFormat_Op<[{
custom<Foo>("$_builder.getI1Type()") attr-dict
}]>;
@@ -382,8 +359,7 @@ def OptionalGroupB : TestFormat_Op<[{
// CHECK-NEXT: odsPrinter << ' ';
// CHECK-NEXT: odsPrinter.printAttributeWithoutType(getAAttr());
// CHECK-NEXT: }
-// CHECK: static constexpr ::llvm::StringRef elidedAttrs[] = {"a"};
-// CHECK: _odsPrinter.printOptionalAttrDict
+// CHECK: _odsPrinter.printOptionalAttrDict((*this)->getDiscardableAttrDictionary().getValue());
def OptionalGroupC : TestFormat_Op<[{
($a^)? attr-dict
}]>, Arguments<(ins DefaultValuedStrAttr<StrAttr, "default">:$a)>;
@@ -392,8 +368,7 @@ def OptionalGroupC : TestFormat_Op<[{
// CHECK: elidedProps.push_back("a");
// CHECK-NOT: elidedProps.push_back("a");
// CHECK: printProperties
-// CHECK: static constexpr ::llvm::StringRef elidedAttrs[] = {"a"};
-// CHECK: _odsPrinter.printOptionalAttrDict
+// CHECK: _odsPrinter.printOptionalAttrDict((*this)->getDiscardableAttrDictionary().getValue());
def OptionalGroupCPropDict : TestFormat_Op<[{
($a^)? prop-dict attr-dict
}]>, Arguments<(ins DefaultValuedStrAttr<StrAttr, "default">:$a)>;
diff --git a/mlir/test/mlir-tblgen/pattern.mlir b/mlir/test/mlir-tblgen/pattern.mlir
index ffb78c28412ce..9d1a88dbffe07 100644
--- a/mlir/test/mlir-tblgen/pattern.mlir
+++ b/mlir/test/mlir-tblgen/pattern.mlir
@@ -759,8 +759,8 @@ func.func @returnTypeAndLocation(%arg0 : i32) -> i1 {
// CHECK-LABEL: @testConstantStrAttr
func.func @testConstantStrAttr() -> () {
- // CHECK: test.has_str_value {value = "foo"}
- test.no_str_value {value = "bar"}
+ // CHECK: test.has_str_value <value = "foo">
+ test.no_str_value <value = "bar">
return
}
diff --git a/mlir/test/python/dialects/python_test.py b/mlir/test/python/dialects/python_test.py
index 655862f957771..917d1eda747d0 100644
--- a/mlir/test/python/dialects/python_test.py
+++ b/mlir/test/python/dialects/python_test.py
@@ -39,17 +39,17 @@ def testAttributes():
two = IntegerAttr.get(i32, 2)
unit = UnitAttr.get()
- # CHECK: python_test.attributed_op {
- # CHECK-DAG: mandatory_i32 = 1 : i32
- # CHECK-DAG: optional_i32 = 2 : i32
+ # CHECK: python_test.attributed_op <
+ # CHECK-DAG: mandatory_i32 = 1
+ # CHECK-DAG: optional_i32 = 2
# CHECK-DAG: unit
- # CHECK: }
+ # CHECK: >
op = test.AttributedOp(one, optional_i32=two, unit=unit)
print(f"{op}")
- # CHECK: python_test.attributed_op {
- # CHECK: mandatory_i32 = 2 : i32
- # CHECK: }
+ # CHECK: python_test.attributed_op <
+ # CHECK: mandatory_i32 = 2
+ # CHECK: >
op2 = test.AttributedOp(two)
print(f"{op2}")
@@ -59,24 +59,25 @@ def testAttributes():
assert "additional" not in op.attributes
- # CHECK: python_test.attributed_op {
+ # CHECK: python_test.attributed_op <
# CHECK-DAG: additional = 1 : i32
- # CHECK-DAG: mandatory_i32 = 2 : i32
+ # CHECK-DAG: mandatory_i32 = 2
# CHECK: }
op2.attributes["additional"] = one
print(f"{op2}")
- # CHECK: python_test.attributed_op {
+ # CHECK: python_test.attributed_op <
# CHECK-DAG: additional = 2 : i32
- # CHECK-DAG: mandatory_i32 = 2 : i32
+ # CHECK-DAG: mandatory_i32 = 2
# CHECK: }
op2.attributes["additional"] = two
print(f"{op2}")
- # CHECK: python_test.attributed_op {
+ # CHECK: python_test.attributed_op <
+ # CHECK-NOT: additional = 2 : i32
+ # CHECK: mandatory_i32 = 2
+ # CHECK: >
# CHECK-NOT: additional = 2 : i32
- # CHECK: mandatory_i32 = 2 : i32
- # CHECK: }
del op2.attributes["additional"]
print(f"{op2}")
@@ -160,23 +161,23 @@ def attrBuilder():
x_arr=[BoolAttr.get(True), StringAttr.get("x")],
x_boolarr=[False, True], # CHECK-DAG: x_boolarr = [false, true]
x_bool=True, # CHECK-DAG: x_bool = true
- x_dboolarr=[True, False], # CHECK-DAG: x_dboolarr = array<i1: true, false>
- x_df16arr=[21, 22], # CHECK-DAG: x_df16arr = array<i16: 21, 22>
- # CHECK-DAG: x_df32arr = array<f32: 2.300000e+01, 2.400000e+01>
+ x_dboolarr=[True, False], # CHECK-DAG: x_dboolarr = [true, false]
+ x_df16arr=[21, 22], # CHECK-DAG: x_df16arr = [21, 22]
+ # CHECK-DAG: x_df32arr = [2.300000e+01, 2.400000e+01]
x_df32arr=[23, 24],
- # CHECK-DAG: x_df64arr = array<f64: 2.500000e+01, 2.600000e+01>
+ # CHECK-DAG: x_df64arr = [2.500000e+01, 2.600000e+01]
x_df64arr=[25, 26],
- x_di32arr=[0, 1], # CHECK-DAG: x_di32arr = array<i32: 0, 1>
- # CHECK-DAG: x_di64arr = array<i64: 1, 2>
+ x_di32arr=[0, 1], # CHECK-DAG: x_di32arr = [0, 1]
+ # CHECK-DAG: x_di64arr = [1, 2]
x_di64arr=[1, 2],
- x_di8arr=[2, 3], # CHECK-DAG: x_di8arr = array<i8: 2, 3>
+ x_di8arr=[2, 3], # CHECK-DAG: x_di8arr = [2, 3]
# CHECK-DAG: x_dictarr = [{a = false}]
x_dictarr=[{"a": BoolAttr.get(False)}],
x_dict={"b": BoolAttr.get(True)}, # CHECK-DAG: x_dict = {b = true}
- x_f32=-2.25, # CHECK-DAG: x_f32 = -2.250000e+00 : f32
+ x_f32=-2.25, # CHECK-DAG: x_f32 = -2.250000e+00
# CHECK-DAG: x_f32arr = [2.000000e+00 : f32, 3.000000e+00 : f32]
x_f32arr=[2.0, 3.0],
- x_f64=4.25, # CHECK-DAG: x_f64 = 4.250000e+00 : f64
+ x_f64=4.25, # CHECK-DAG: x_f64 = 4.250000e+00
x_f64arr=[4.0, 8.0], # CHECK-DAG: x_f64arr = [4.000000e+00, 8.000000e+00]
# CHECK-DAG: x_f64elems = dense<[8.000000e+00, 1.600000e+01]> : tensor<2xf64>
x_f64elems=[8.0, 16.0],
@@ -184,38 +185,38 @@ def attrBuilder():
x_flatsymrefarr=["symbol1", "symbol2"],
x_flatsymref="symbol3", # CHECK-DAG: x_flatsymref = @symbol3
x_i1=0, # CHECK-DAG: x_i1 = false
- x_i16=42, # CHECK-DAG: x_i16 = 42 : i16
- x_i32=6, # CHECK-DAG: x_i32 = 6 : i32
+ x_i16=42, # CHECK-DAG: x_i16 = 42
+ x_i32=6, # CHECK-DAG: x_i32 = 6
x_i32arr=[4, 5], # CHECK-DAG: x_i32arr = [4 : i32, 5 : i32]
x_i32elems=[5, 6], # CHECK-DAG: x_i32elems = dense<[5, 6]> : tensor<2xi32>
- x_i64=9, # CHECK-DAG: x_i64 = 9 : i64
+ x_i64=9, # CHECK-DAG: x_i64 = 9
x_i64arr=[7, 8], # CHECK-DAG: x_i64arr = [7, 8]
x_i64elems=[8, 9], # CHECK-DAG: x_i64elems = dense<[8, 9]> : tensor<2xi64>
x_i64svecarr=[10, 11], # CHECK-DAG: x_i64svecarr = [10, 11]
- x_i8=11, # CHECK-DAG: x_i8 = 11 : i8
- x_idx=10, # CHECK-DAG: x_idx = 10 : index
+ x_i8=11, # CHECK-DAG: x_i8 = 11
+ x_idx=10, # CHECK-DAG: x_idx = 10
# CHECK-DAG: x_idxelems = dense<[11, 12]> : tensor<2xindex>
x_idxelems=[11, 12],
# CHECK-DAG: x_idxlistarr = [{{\[}}13], [14, 15]]
x_idxlistarr=[[13], [14, 15]],
- x_si1=-1, # CHECK-DAG: x_si1 = -1 : si1
- x_si16=-2, # CHECK-DAG: x_si16 = -2 : si16
- x_si32=-3, # CHECK-DAG: x_si32 = -3 : si32
- x_si64=-123, # CHECK-DAG: x_si64 = -123 : si64
- x_si8=-4, # CHECK-DAG: x_si8 = -4 : si8
+ x_si1=-1, # CHECK-DAG: x_si1 = -1
+ x_si16=-2, # CHECK-DAG: x_si16 = -2
+ x_si32=-3, # CHECK-DAG: x_si32 = -3
+ x_si64=-123, # CHECK-DAG: x_si64 = -123
+ x_si8=-4, # CHECK-DAG: x_si8 = -4
x_strarr=["hello", "world"], # CHECK-DAG: x_strarr = ["hello", "world"]
x_str="hello world!", # CHECK-DAG: x_str = "hello world!"
# CHECK-DAG: x_symrefarr = [@flatsym, @deep::@sym]
x_symrefarr=["flatsym", ["deep", "sym"]],
x_symref=["deep", "sym2"], # CHECK-DAG: x_symref = @deep::@sym2
- x_sym="symbol", # CHECK-DAG: x_sym = "symbol"
+ x_sym="symbol", # CHECK-DAG: x_sym = @symbol
x_typearr=[F32Type.get()], # CHECK-DAG: x_typearr = [f32]
x_type=F64Type.get(), # CHECK-DAG: x_type = f64
- x_ui1=1, # CHECK-DAG: x_ui1 = 1 : ui1
- x_ui16=2, # CHECK-DAG: x_ui16 = 2 : ui16
- x_ui32=3, # CHECK-DAG: x_ui32 = 3 : ui32
- x_ui64=4, # CHECK-DAG: x_ui64 = 4 : ui64
- x_ui8=5, # CHECK-DAG: x_ui8 = 5 : ui8
+ x_ui1=1, # CHECK-DAG: x_ui1 = 1
+ x_ui16=2, # CHECK-DAG: x_ui16 = 2
+ x_ui32=3, # CHECK-DAG: x_ui32 = 3
+ x_ui64=4, # CHECK-DAG: x_ui64 = 4
+ x_ui8=5, # CHECK-DAG: x_ui8 = 5
x_unit=True, # CHECK-DAG: x_unit
)
op.verify()
@@ -559,9 +560,9 @@ def testCustomAttribute():
# CHECK: #python_test.test_attr
print(a)
- # CHECK: python_test.custom_attributed_op {
+ # CHECK: python_test.custom_attributed_op <
# CHECK: #python_test.test_attr
- # CHECK: }
+ # CHECK: >
op2 = test.CustomAttributedOp(a)
print(f"{op2}")
diff --git a/mlir/test/python/python_test_ops.td b/mlir/test/python/python_test_ops.td
index 3624506a23a6d..aabd6c22efb1f 100644
--- a/mlir/test/python/python_test_ops.td
+++ b/mlir/test/python/python_test_ops.td
@@ -15,7 +15,6 @@ include "mlir/Interfaces/InferTypeOpInterface.td"
def Python_Test_Dialect : Dialect {
let name = "python_test";
- let useStrictPropertiesInAssemblyFormat = 0;
let cppNamespace = "python_test";
let useDefaultTypePrinterParser = 1;
@@ -34,7 +33,8 @@ class TestAttr<string name, string attrMnemonic>
class TestOp<string mnemonic, list<Trait> traits = []>
: Op<Python_Test_Dialect, mnemonic, traits> {
- let assemblyFormat = "operands attr-dict functional-type(operands, results)";
+ let assemblyFormat =
+ "operands prop-dict attr-dict functional-type(operands, results)";
}
//===----------------------------------------------------------------------===//
>From 05973cccd52568420cf09f9241542f2eb8e6f7d5 Mon Sep 17 00:00:00 2001
From: Amit Tiwari <Amit.Tiwari at amd.com>
Date: Sat, 26 Sep 2026 19:52:24 +0530
Subject: [PATCH 13/53] [Clang][OpenMP] Optimize `collapse` IV bit-width
precision expression (#225612)
`collapse` used to build both a 32-bit and a 64-bit trip-count, then
keep one.
This patch does the same in `checkOpenMPLoop`:
If we know 32-bit is enough, build only 32-bit.
Else build 64-bit first.
Build 32-bit only when the product is a compile-time constant and fits.
Origin: `flatten` already builds only the width it keeps.
---
clang/lib/Sema/SemaOpenMP.cpp | 102 +++++++++---------
.../test/OpenMP/collapse_iv_width_codegen.cpp | 52 +++++++++
2 files changed, 105 insertions(+), 49 deletions(-)
create mode 100644 clang/test/OpenMP/collapse_iv_width_codegen.cpp
diff --git a/clang/lib/Sema/SemaOpenMP.cpp b/clang/lib/Sema/SemaOpenMP.cpp
index 4b49a75d1f84b..41cfc1dcd3649 100644
--- a/clang/lib/Sema/SemaOpenMP.cpp
+++ b/clang/lib/Sema/SemaOpenMP.cpp
@@ -10731,29 +10731,10 @@ checkOpenMPLoop(OpenMPDirectiveKind DKind, Expr *CollapseLoopCountExpr,
// Precondition tests if there is at least one iteration (all conditions are
// true).
auto PreCond = ExprResult(IterSpaces[0].PreCond);
- Expr *N0 = IterSpaces[0].NumIterations;
- ExprResult LastIteration32 = widenIterationCount(
- /*Bits=*/32,
- SemaRef
- .PerformImplicitConversion(N0->IgnoreImpCasts(), N0->getType(),
- AssignmentAction::Converting,
- /*AllowExplicit=*/true)
- .get(),
- SemaRef);
- ExprResult LastIteration64 = widenIterationCount(
- /*Bits=*/64,
- SemaRef
- .PerformImplicitConversion(N0->IgnoreImpCasts(), N0->getType(),
- AssignmentAction::Converting,
- /*AllowExplicit=*/true)
- .get(),
- SemaRef);
-
- if (!LastIteration32.isUsable() || !LastIteration64.isUsable())
- return NestedLoopCount;
-
ASTContext &C = SemaRef.Context;
- bool AllCountsNeedLessThan32Bits = C.getTypeSize(N0->getType()) < 32;
+ unsigned FirstCountBits =
+ C.getTypeSize(IterSpaces[0].NumIterations->getType());
+ bool AllCountsNeedLessThan32Bits = FirstCountBits < 32;
Scope *CurScope = DSA.getCurScope();
for (unsigned Cnt = 1; Cnt < NestedLoopCount; ++Cnt) {
@@ -10763,37 +10744,63 @@ checkOpenMPLoop(OpenMPDirectiveKind DKind, Expr *CollapseLoopCountExpr,
PreCond.get(), IterSpaces[Cnt].PreCond);
}
Expr *N = IterSpaces[Cnt].NumIterations;
- SourceLocation Loc = N->getExprLoc();
AllCountsNeedLessThan32Bits &= C.getTypeSize(N->getType()) < 32;
- if (LastIteration32.isUsable())
- LastIteration32 = SemaRef.BuildBinOp(
- CurScope, Loc, BO_Mul, LastIteration32.get(),
- SemaRef
- .PerformImplicitConversion(N->IgnoreImpCasts(), N->getType(),
- AssignmentAction::Converting,
- /*AllowExplicit=*/true)
- .get());
- if (LastIteration64.isUsable())
- LastIteration64 = SemaRef.BuildBinOp(
- CurScope, Loc, BO_Mul, LastIteration64.get(),
+ }
+
+ auto BuildLastIteration = [&](unsigned Bits) -> ExprResult {
+ ExprResult Result;
+ for (unsigned Cnt : llvm::seq<unsigned>(NestedLoopCount)) {
+ Expr *N = IterSpaces[Cnt].NumIterations;
+ ExprResult Count = widenIterationCount(
+ Bits,
SemaRef
.PerformImplicitConversion(N->IgnoreImpCasts(), N->getType(),
AssignmentAction::Converting,
/*AllowExplicit=*/true)
- .get());
- }
+ .get(),
+ SemaRef);
+ if (!Count.isUsable())
+ return ExprError();
+ if (Cnt == 0)
+ Result = Count;
+ else
+ Result = SemaRef.BuildBinOp(CurScope, N->getExprLoc(), BO_Mul,
+ Result.get(), Count.get());
+ if (!Result.isUsable())
+ return ExprError();
+ }
+ return Result;
+ };
- // Choose either the 32-bit or 64-bit version.
- ExprResult LastIteration = LastIteration64;
+ // Build the 32-bit tree immediately only when it is always selected.
+ // Otherwise, build the 64-bit tree first and build the 32-bit tree only when
+ // the constant product may fit.
+ ExprResult LastIteration;
if (SemaRef.getLangOpts().OpenMPOptimisticCollapse ||
- (LastIteration32.isUsable() &&
- C.getTypeSize(LastIteration32.get()->getType()) == 32 &&
- (AllCountsNeedLessThan32Bits || NestedLoopCount == 1 ||
- fitsInto(
- /*Bits=*/32,
- LastIteration32.get()->getType()->hasSignedIntegerRepresentation(),
- LastIteration64.get(), SemaRef))))
- LastIteration = LastIteration32;
+ AllCountsNeedLessThan32Bits ||
+ (NestedLoopCount == 1 && FirstCountBits == 32)) {
+ LastIteration = BuildLastIteration(/*Bits=*/32);
+ } else {
+ ExprResult LastIteration64 = BuildLastIteration(/*Bits=*/64);
+ if (!LastIteration64.isUsable())
+ return NestedLoopCount;
+ LastIteration = LastIteration64;
+ if (LastIteration64.get()->isIntegerConstantExpr(C)) {
+ ExprResult LastIteration32 = BuildLastIteration(/*Bits=*/32);
+ if (LastIteration32.isUsable() &&
+ C.getTypeSize(LastIteration32.get()->getType()) == 32 &&
+ fitsInto(
+ /*Bits=*/32,
+ LastIteration32.get()
+ ->getType()
+ ->hasSignedIntegerRepresentation(),
+ LastIteration64.get(), SemaRef))
+ LastIteration = LastIteration32;
+ }
+ }
+ if (!LastIteration.isUsable())
+ return NestedLoopCount;
+
QualType VType = LastIteration.get()->getType();
QualType RealVType = VType;
QualType StrideVType = VType;
@@ -10804,9 +10811,6 @@ checkOpenMPLoop(OpenMPDirectiveKind DKind, Expr *CollapseLoopCountExpr,
SemaRef.Context.getIntTypeForBitwidth(/*DestWidth=*/64, /*Signed=*/1);
}
- if (!LastIteration.isUsable())
- return 0;
-
// Save the number of iterations.
ExprResult NumIterations = LastIteration;
{
diff --git a/clang/test/OpenMP/collapse_iv_width_codegen.cpp b/clang/test/OpenMP/collapse_iv_width_codegen.cpp
new file mode 100644
index 0000000000000..2a1d0dcc9f19a
--- /dev/null
+++ b/clang/test/OpenMP/collapse_iv_width_codegen.cpp
@@ -0,0 +1,52 @@
+// RUN: %clang_cc1 -verify -fopenmp -std=c++20 -x c++ -triple x86_64-unknown-unknown \
+// RUN: -Wno-bit-int-extension -emit-llvm %s -o - | FileCheck %s
+
+// expected-no-diagnostics
+
+void one_i32(unsigned n) {
+#pragma omp parallel for collapse(1)
+ for (unsigned i = 0; i < n; ++i)
+ ;
+}
+
+// CHECK-LABEL: define internal void @_Z7one_i32j.omp_outlined(
+// CHECK: call void @__kmpc_for_static_init_4u(
+
+void one_i40(_BitInt(40) n) {
+#pragma omp parallel for collapse(1)
+ for (_BitInt(40) i = 0; i < n; ++i)
+ ;
+}
+
+// CHECK-LABEL: define internal void @_Z7one_i40DB40_.omp_outlined(
+// CHECK: call void @__kmpc_for_static_init_8(
+
+void dynamic_two(unsigned n, unsigned m) {
+#pragma omp parallel for collapse(2)
+ for (unsigned i = 0; i < n; ++i)
+ for (unsigned j = 0; j < m; ++j)
+ ;
+}
+
+// CHECK-LABEL: define internal void @_Z11dynamic_twojj.omp_outlined(
+// CHECK: call void @__kmpc_for_static_init_8(
+
+void fit_constant() {
+#pragma omp parallel for collapse(2)
+ for (int i = 0; i < 100; ++i)
+ for (int j = 0; j < 100; ++j)
+ ;
+}
+
+// CHECK-LABEL: define internal void @_Z12fit_constantv.omp_outlined(
+// CHECK: call void @__kmpc_for_static_init_4(
+
+void wide_constant() {
+#pragma omp parallel for collapse(2)
+ for (int i = 0; i < 100000; ++i)
+ for (int j = 0; j < 100000; ++j)
+ ;
+}
+
+// CHECK-LABEL: define internal void @_Z13wide_constantv.omp_outlined(
+// CHECK: call void @__kmpc_for_static_init_8(
>From 9eec0a8d7c8ec0d37711aec482b98b2ce52f0a55 Mon Sep 17 00:00:00 2001
From: Florian Hahn <flo at fhahn.com>
Date: Sat, 26 Sep 2026 17:14:59 +0100
Subject: [PATCH 14/53] [VPlan] Mark default value or ExtractLastActive as only
first lane used. (#226153)
The default value (operand 0) of ExtractLastActive is the scalar value
of @llvm.experimental.vector.extract.last.active. Only the first lane is
used, mark accordingly.
PR: https://github.com/llvm/llvm-project/pull/226153
---
.../lib/Transforms/Vectorize/VPlanRecipes.cpp | 2 ++
...conditional-scalar-assignment-fold-tail.ll | 3 +-
.../AArch64/conditional-scalar-assignment.ll | 28 ++++++++-----------
...conditional-scalar-assignment-fold-tail.ll | 3 +-
.../RISCV/conditional-scalar-assignment.ll | 7 ++---
.../X86/conditional-scalar-assignment.ll | 3 +-
.../LoopVectorize/find-last-ptr-induction.ll | 7 ++---
.../iv-select-cmp-non-const-iv-start.ll | 10 +++----
.../LoopVectorize/iv-select-cmp-trunc.ll | 20 ++++++-------
.../Transforms/LoopVectorize/iv-select-cmp.ll | 10 +++----
10 files changed, 39 insertions(+), 54 deletions(-)
diff --git a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
index b2df7c33227e9..4b68e3e71afea 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
@@ -1746,6 +1746,8 @@ bool VPInstruction::usesFirstLaneOnly(const VPValue *Op) const {
return Op == getOperand(1);
case Instruction::InsertElement:
return Op == getOperand(1) || Op == getOperand(2);
+ case VPInstruction::ExtractLastActive:
+ return Op == getOperand(0);
case Instruction::PHI:
return true;
case Instruction::FCmp:
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/conditional-scalar-assignment-fold-tail.ll b/llvm/test/Transforms/LoopVectorize/AArch64/conditional-scalar-assignment-fold-tail.ll
index 8c98cd2e9663d..4515d331feb3c 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/conditional-scalar-assignment-fold-tail.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/conditional-scalar-assignment-fold-tail.ll
@@ -102,8 +102,7 @@ define i32 @non_speculatable_find_last_reduction(ptr noalias %a, ptr noalias %b,
; CHECK-NEXT: [[TMP16:%.*]] = xor i1 [[TMP15]], true
; CHECK-NEXT: br i1 [[TMP16]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP3:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
-; CHECK-NEXT: [[TMP17:%.*]] = extractelement <vscale x 4 x i32> [[BROADCAST_SPLAT2]], i64 0
-; CHECK-NEXT: [[TMP18:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.nxv4i32(<vscale x 4 x i32> [[TMP14]], <vscale x 4 x i1> [[TMP13]], i32 [[TMP17]])
+; CHECK-NEXT: [[TMP18:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.nxv4i32(<vscale x 4 x i32> [[TMP14]], <vscale x 4 x i1> [[TMP13]], i32 [[DEFAULT_VAL]])
; CHECK-NEXT: br label %[[EXIT:.*]]
; CHECK: [[EXIT]]:
; CHECK-NEXT: ret i32 [[TMP18]]
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/conditional-scalar-assignment.ll b/llvm/test/Transforms/LoopVectorize/AArch64/conditional-scalar-assignment.ll
index ee5b164eae8c3..d197d727eae66 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/conditional-scalar-assignment.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/conditional-scalar-assignment.ll
@@ -141,13 +141,12 @@ define ptr @simple_csa_ptr_select(i64 %N, ptr %data, i64 %a, ptr %init) {
; SVE-NEXT: [[TMP12:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; SVE-NEXT: br i1 [[TMP12]], label %[[MIDDLE_BLOCK:.*]], label %[[LOOP]], !llvm.loop [[LOOP4:![0-9]+]]
; SVE: [[MIDDLE_BLOCK]]:
-; SVE-NEXT: [[TMP13:%.*]] = extractelement <vscale x 2 x ptr> [[BROADCAST_SPLAT2]], i64 0
-; SVE-NEXT: [[TMP14:%.*]] = call ptr @llvm.experimental.vector.extract.last.active.nxv2p0(<vscale x 2 x ptr> [[TMP11]], <vscale x 2 x i1> [[TMP10]], ptr [[TMP13]])
+; SVE-NEXT: [[TMP13:%.*]] = call ptr @llvm.experimental.vector.extract.last.active.nxv2p0(<vscale x 2 x ptr> [[TMP11]], <vscale x 2 x i1> [[TMP10]], ptr [[INIT]])
; SVE-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
; SVE-NEXT: br i1 [[CMP_N]], label %[[EXIT:.*]], label %[[SCALAR_PH]]
; SVE: [[SCALAR_PH]]:
; SVE-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
-; SVE-NEXT: [[BC_MERGE_RDX:%.*]] = phi ptr [ [[TMP14]], %[[MIDDLE_BLOCK]] ], [ [[INIT]], %[[ENTRY]] ]
+; SVE-NEXT: [[BC_MERGE_RDX:%.*]] = phi ptr [ [[TMP13]], %[[MIDDLE_BLOCK]] ], [ [[INIT]], %[[ENTRY]] ]
; SVE-NEXT: br label %[[LOOP1:.*]]
; SVE: [[LOOP1]]:
; SVE-NEXT: [[IV1:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[LOOP1]] ]
@@ -161,7 +160,7 @@ define ptr @simple_csa_ptr_select(i64 %N, ptr %data, i64 %a, ptr %init) {
; SVE-NEXT: [[EXIT_CMP:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
; SVE-NEXT: br i1 [[EXIT_CMP]], label %[[EXIT]], label %[[LOOP1]], !llvm.loop [[LOOP5:![0-9]+]]
; SVE: [[EXIT]]:
-; SVE-NEXT: [[SELECT_DATA_LCSSA:%.*]] = phi ptr [ [[SELECT_DATA]], %[[LOOP1]] ], [ [[TMP14]], %[[MIDDLE_BLOCK]] ]
+; SVE-NEXT: [[SELECT_DATA_LCSSA:%.*]] = phi ptr [ [[SELECT_DATA]], %[[LOOP1]] ], [ [[TMP13]], %[[MIDDLE_BLOCK]] ]
; SVE-NEXT: ret ptr [[SELECT_DATA_LCSSA]]
;
entry:
@@ -1038,13 +1037,12 @@ define i32 @simple_csa_int_load(ptr noalias %a, ptr noalias %b, i32 %default_val
; SVE-NEXT: [[TMP12:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; SVE-NEXT: br i1 [[TMP12]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP16:![0-9]+]]
; SVE: [[MIDDLE_BLOCK]]:
-; SVE-NEXT: [[TMP13:%.*]] = extractelement <vscale x 4 x i32> [[BROADCAST_SPLAT2]], i64 0
-; SVE-NEXT: [[TMP14:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.nxv4i32(<vscale x 4 x i32> [[TMP11]], <vscale x 4 x i1> [[TMP10]], i32 [[TMP13]])
+; SVE-NEXT: [[TMP13:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.nxv4i32(<vscale x 4 x i32> [[TMP11]], <vscale x 4 x i1> [[TMP10]], i32 [[DEFAULT_VAL]])
; SVE-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
; SVE-NEXT: br i1 [[CMP_N]], label %[[EXIT:.*]], label %[[SCALAR_PH]]
; SVE: [[SCALAR_PH]]:
; SVE-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
-; SVE-NEXT: [[BC_MERGE_RDX:%.*]] = phi i32 [ [[TMP14]], %[[MIDDLE_BLOCK]] ], [ [[DEFAULT_VAL]], %[[ENTRY]] ]
+; SVE-NEXT: [[BC_MERGE_RDX:%.*]] = phi i32 [ [[TMP13]], %[[MIDDLE_BLOCK]] ], [ [[DEFAULT_VAL]], %[[ENTRY]] ]
; SVE-NEXT: br label %[[LOOP:.*]]
; SVE: [[LOOP]]:
; SVE-NEXT: [[IV:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[LATCH:.*]] ]
@@ -1063,7 +1061,7 @@ define i32 @simple_csa_int_load(ptr noalias %a, ptr noalias %b, i32 %default_val
; SVE-NEXT: [[EXIT_CMP:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
; SVE-NEXT: br i1 [[EXIT_CMP]], label %[[EXIT]], label %[[LOOP]], !llvm.loop [[LOOP17:![0-9]+]]
; SVE: [[EXIT]]:
-; SVE-NEXT: [[SELECT_DATA_LCSSA:%.*]] = phi i32 [ [[SELECT_DATA]], %[[LATCH]] ], [ [[TMP14]], %[[MIDDLE_BLOCK]] ]
+; SVE-NEXT: [[SELECT_DATA_LCSSA:%.*]] = phi i32 [ [[SELECT_DATA]], %[[LATCH]] ], [ [[TMP13]], %[[MIDDLE_BLOCK]] ]
; SVE-NEXT: ret i32 [[SELECT_DATA_LCSSA]]
;
entry:
@@ -1155,13 +1153,12 @@ define i32 @simple_csa_int_divide(ptr noalias %a, ptr noalias %b, i32 %default_v
; SVE-NEXT: [[TMP13:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; SVE-NEXT: br i1 [[TMP13]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP18:![0-9]+]]
; SVE: [[MIDDLE_BLOCK]]:
-; SVE-NEXT: [[TMP14:%.*]] = extractelement <vscale x 4 x i32> [[BROADCAST_SPLAT2]], i64 0
-; SVE-NEXT: [[TMP15:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.nxv4i32(<vscale x 4 x i32> [[TMP12]], <vscale x 4 x i1> [[TMP11]], i32 [[TMP14]])
+; SVE-NEXT: [[TMP14:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.nxv4i32(<vscale x 4 x i32> [[TMP12]], <vscale x 4 x i1> [[TMP11]], i32 [[DEFAULT_VAL]])
; SVE-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
; SVE-NEXT: br i1 [[CMP_N]], label %[[EXIT:.*]], label %[[SCALAR_PH]]
; SVE: [[SCALAR_PH]]:
; SVE-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
-; SVE-NEXT: [[BC_MERGE_RDX:%.*]] = phi i32 [ [[TMP15]], %[[MIDDLE_BLOCK]] ], [ [[DEFAULT_VAL]], %[[ENTRY]] ]
+; SVE-NEXT: [[BC_MERGE_RDX:%.*]] = phi i32 [ [[TMP14]], %[[MIDDLE_BLOCK]] ], [ [[DEFAULT_VAL]], %[[ENTRY]] ]
; SVE-NEXT: br label %[[LOOP:.*]]
; SVE: [[LOOP]]:
; SVE-NEXT: [[IV:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[LATCH:.*]] ]
@@ -1179,7 +1176,7 @@ define i32 @simple_csa_int_divide(ptr noalias %a, ptr noalias %b, i32 %default_v
; SVE-NEXT: [[EXIT_CMP:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
; SVE-NEXT: br i1 [[EXIT_CMP]], label %[[EXIT]], label %[[LOOP]], !llvm.loop [[LOOP19:![0-9]+]]
; SVE: [[EXIT]]:
-; SVE-NEXT: [[SELECT_DATA_LCSSA:%.*]] = phi i32 [ [[SELECT_DATA]], %[[LATCH]] ], [ [[TMP15]], %[[MIDDLE_BLOCK]] ]
+; SVE-NEXT: [[SELECT_DATA_LCSSA:%.*]] = phi i32 [ [[SELECT_DATA]], %[[LATCH]] ], [ [[TMP14]], %[[MIDDLE_BLOCK]] ]
; SVE-NEXT: ret i32 [[SELECT_DATA_LCSSA]]
;
entry:
@@ -1283,13 +1280,12 @@ define i32 @csa_load_nested_ifs(ptr noalias %a, ptr noalias %b, i32 %default_val
; SVE-NEXT: [[TMP15:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; SVE-NEXT: br i1 [[TMP15]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP20:![0-9]+]]
; SVE: [[MIDDLE_BLOCK]]:
-; SVE-NEXT: [[TMP16:%.*]] = extractelement <vscale x 4 x i32> [[BROADCAST_SPLAT2]], i64 0
-; SVE-NEXT: [[TMP17:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.nxv4i32(<vscale x 4 x i32> [[TMP14]], <vscale x 4 x i1> [[TMP13]], i32 [[TMP16]])
+; SVE-NEXT: [[TMP16:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.nxv4i32(<vscale x 4 x i32> [[TMP14]], <vscale x 4 x i1> [[TMP13]], i32 [[DEFAULT_VAL]])
; SVE-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
; SVE-NEXT: br i1 [[CMP_N]], label %[[EXIT:.*]], label %[[SCALAR_PH]]
; SVE: [[SCALAR_PH]]:
; SVE-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
-; SVE-NEXT: [[BC_MERGE_RDX:%.*]] = phi i32 [ [[TMP17]], %[[MIDDLE_BLOCK]] ], [ [[DEFAULT_VAL]], %[[ENTRY]] ]
+; SVE-NEXT: [[BC_MERGE_RDX:%.*]] = phi i32 [ [[TMP16]], %[[MIDDLE_BLOCK]] ], [ [[DEFAULT_VAL]], %[[ENTRY]] ]
; SVE-NEXT: br label %[[LOOP:.*]]
; SVE: [[LOOP]]:
; SVE-NEXT: [[IV:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[LATCH:.*]] ]
@@ -1311,7 +1307,7 @@ define i32 @csa_load_nested_ifs(ptr noalias %a, ptr noalias %b, i32 %default_val
; SVE-NEXT: [[EXIT_CMP:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
; SVE-NEXT: br i1 [[EXIT_CMP]], label %[[EXIT]], label %[[LOOP]], !llvm.loop [[LOOP21:![0-9]+]]
; SVE: [[EXIT]]:
-; SVE-NEXT: [[SELECT_DATA_LCSSA:%.*]] = phi i32 [ [[SELECT_DATA]], %[[LATCH]] ], [ [[TMP17]], %[[MIDDLE_BLOCK]] ]
+; SVE-NEXT: [[SELECT_DATA_LCSSA:%.*]] = phi i32 [ [[SELECT_DATA]], %[[LATCH]] ], [ [[TMP16]], %[[MIDDLE_BLOCK]] ]
; SVE-NEXT: ret i32 [[SELECT_DATA_LCSSA]]
;
entry:
diff --git a/llvm/test/Transforms/LoopVectorize/RISCV/conditional-scalar-assignment-fold-tail.ll b/llvm/test/Transforms/LoopVectorize/RISCV/conditional-scalar-assignment-fold-tail.ll
index 06758467221c5..0e4e110ca1997 100644
--- a/llvm/test/Transforms/LoopVectorize/RISCV/conditional-scalar-assignment-fold-tail.ll
+++ b/llvm/test/Transforms/LoopVectorize/RISCV/conditional-scalar-assignment-fold-tail.ll
@@ -96,8 +96,7 @@ define i32 @non_speculatable_find_last_reduction(ptr noalias %a, ptr noalias %b,
; CHECK-NEXT: [[TMP14:%.*]] = icmp eq i64 [[AVL_NEXT]], 0
; CHECK-NEXT: br i1 [[TMP14]], label %[[IF_THEN:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP3:![0-9]+]]
; CHECK: [[IF_THEN]]:
-; CHECK-NEXT: [[TMP15:%.*]] = extractelement <vscale x 4 x i32> [[BROADCAST_SPLAT2]], i64 0
-; CHECK-NEXT: [[SELECT_DATA_LCSSA:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.nxv4i32(<vscale x 4 x i32> [[TMP12]], <vscale x 4 x i1> [[TMP11]], i32 [[TMP15]])
+; CHECK-NEXT: [[SELECT_DATA_LCSSA:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.nxv4i32(<vscale x 4 x i32> [[TMP12]], <vscale x 4 x i1> [[TMP11]], i32 [[DEFAULT_VAL]])
; CHECK-NEXT: br label %[[LATCH:.*]]
; CHECK: [[LATCH]]:
; CHECK-NEXT: ret i32 [[SELECT_DATA_LCSSA]]
diff --git a/llvm/test/Transforms/LoopVectorize/RISCV/conditional-scalar-assignment.ll b/llvm/test/Transforms/LoopVectorize/RISCV/conditional-scalar-assignment.ll
index f4b3bf607fa9d..d685a238e7d9b 100644
--- a/llvm/test/Transforms/LoopVectorize/RISCV/conditional-scalar-assignment.ll
+++ b/llvm/test/Transforms/LoopVectorize/RISCV/conditional-scalar-assignment.ll
@@ -114,13 +114,12 @@ define i32 @non_speculatable_find_last_reduction(ptr noalias %a, ptr noalias %b,
; CHECK-NEXT: [[TMP13:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-NEXT: br i1 [[TMP13]], label %[[IF_THEN:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
; CHECK: [[IF_THEN]]:
-; CHECK-NEXT: [[TMP15:%.*]] = extractelement <vscale x 4 x i32> [[BROADCAST_SPLAT2]], i64 0
-; CHECK-NEXT: [[SELECT_DATA_LCSSA:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.nxv4i32(<vscale x 4 x i32> [[TMP12]], <vscale x 4 x i1> [[TMP11]], i32 [[TMP15]])
+; CHECK-NEXT: [[TMP14:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.nxv4i32(<vscale x 4 x i32> [[TMP12]], <vscale x 4 x i1> [[TMP11]], i32 [[DEFAULT_VAL]])
; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
; CHECK-NEXT: br i1 [[CMP_N]], label %[[EXIT1:.*]], label %[[SCALAR_PH1]]
; CHECK: [[SCALAR_PH1]]:
; CHECK-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[IF_THEN]] ], [ 0, %[[SCALAR_PH]] ]
-; CHECK-NEXT: [[BC_MERGE_RDX:%.*]] = phi i32 [ [[SELECT_DATA_LCSSA]], %[[IF_THEN]] ], [ [[DEFAULT_VAL]], %[[SCALAR_PH]] ]
+; CHECK-NEXT: [[BC_MERGE_RDX:%.*]] = phi i32 [ [[TMP14]], %[[IF_THEN]] ], [ [[DEFAULT_VAL]], %[[SCALAR_PH]] ]
; CHECK-NEXT: br label %[[LATCH:.*]]
; CHECK: [[LATCH]]:
; CHECK-NEXT: [[IV1:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[SCALAR_PH1]] ], [ [[IV_NEXT:%.*]], %[[LATCH1:.*]] ]
@@ -139,7 +138,7 @@ define i32 @non_speculatable_find_last_reduction(ptr noalias %a, ptr noalias %b,
; CHECK-NEXT: [[EXIT_CMP:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
; CHECK-NEXT: br i1 [[EXIT_CMP]], label %[[EXIT1]], label %[[LATCH]], !llvm.loop [[LOOP5:![0-9]+]]
; CHECK: [[EXIT1]]:
-; CHECK-NEXT: [[SELECT_DATA_LCSSA1:%.*]] = phi i32 [ [[SELECT_DATA]], %[[LATCH1]] ], [ [[SELECT_DATA_LCSSA]], %[[IF_THEN]] ]
+; CHECK-NEXT: [[SELECT_DATA_LCSSA1:%.*]] = phi i32 [ [[SELECT_DATA]], %[[LATCH1]] ], [ [[TMP14]], %[[IF_THEN]] ]
; CHECK-NEXT: ret i32 [[SELECT_DATA_LCSSA1]]
;
entry:
diff --git a/llvm/test/Transforms/LoopVectorize/X86/conditional-scalar-assignment.ll b/llvm/test/Transforms/LoopVectorize/X86/conditional-scalar-assignment.ll
index f6bd69acb775b..f5c51e4c1b071 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/conditional-scalar-assignment.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/conditional-scalar-assignment.ll
@@ -167,8 +167,7 @@ define ptr @simple_csa_ptr_select(i64 %N, ptr %data, i64 %a, ptr %init) {
; AVX512-NEXT: [[TMP8:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; AVX512-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK:.*]], label %[[LOOP]], !llvm.loop [[LOOP4:![0-9]+]]
; AVX512: [[MIDDLE_BLOCK]]:
-; AVX512-NEXT: [[TMP9:%.*]] = extractelement <8 x ptr> [[BROADCAST_SPLAT2]], i64 0
-; AVX512-NEXT: [[TMP10:%.*]] = call ptr @llvm.experimental.vector.extract.last.active.v8p0(<8 x ptr> [[TMP7]], <8 x i1> [[TMP6]], ptr [[TMP9]])
+; AVX512-NEXT: [[TMP10:%.*]] = call ptr @llvm.experimental.vector.extract.last.active.v8p0(<8 x ptr> [[TMP7]], <8 x i1> [[TMP6]], ptr [[INIT]])
; AVX512-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
; AVX512-NEXT: br i1 [[CMP_N]], label %[[EXIT:.*]], label %[[SCALAR_PH]]
; AVX512: [[SCALAR_PH]]:
diff --git a/llvm/test/Transforms/LoopVectorize/find-last-ptr-induction.ll b/llvm/test/Transforms/LoopVectorize/find-last-ptr-induction.ll
index 2cc0581e6b629..815452573858b 100644
--- a/llvm/test/Transforms/LoopVectorize/find-last-ptr-induction.ll
+++ b/llvm/test/Transforms/LoopVectorize/find-last-ptr-induction.ll
@@ -110,14 +110,13 @@ define ptr @find_last_ptr_induction_separate_array(ptr %a, ptr %b, ptr %init, i6
; CHECK-NEXT: [[TMP10:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-NEXT: br i1 [[TMP10]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
-; CHECK-NEXT: [[TMP11:%.*]] = extractelement <4 x ptr> [[BROADCAST_SPLAT]], i64 0
-; CHECK-NEXT: [[TMP12:%.*]] = call ptr @llvm.experimental.vector.extract.last.active.v4p0(<4 x ptr> [[TMP9]], <4 x i1> [[TMP8]], ptr [[TMP11]])
+; CHECK-NEXT: [[TMP11:%.*]] = call ptr @llvm.experimental.vector.extract.last.active.v4p0(<4 x ptr> [[TMP9]], <4 x i1> [[TMP8]], ptr [[INIT]])
; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
; CHECK-NEXT: br i1 [[CMP_N]], label %[[EXIT:.*]], label %[[SCALAR_PH]]
; CHECK: [[SCALAR_PH]]:
; CHECK-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
; CHECK-NEXT: [[BC_RESUME_VAL1:%.*]] = phi ptr [ [[TMP2]], %[[MIDDLE_BLOCK]] ], [ [[B]], %[[ENTRY]] ]
-; CHECK-NEXT: [[BC_MERGE_RDX:%.*]] = phi ptr [ [[TMP12]], %[[MIDDLE_BLOCK]] ], [ [[INIT]], %[[ENTRY]] ]
+; CHECK-NEXT: [[BC_MERGE_RDX:%.*]] = phi ptr [ [[TMP11]], %[[MIDDLE_BLOCK]] ], [ [[INIT]], %[[ENTRY]] ]
; CHECK-NEXT: br label %[[LOOP:.*]]
; CHECK: [[LOOP]]:
; CHECK-NEXT: [[IV:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[LOOP]] ]
@@ -132,7 +131,7 @@ define ptr @find_last_ptr_induction_separate_array(ptr %a, ptr %b, ptr %init, i6
; CHECK-NEXT: [[EC:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
; CHECK-NEXT: br i1 [[EC]], label %[[EXIT]], label %[[LOOP]], !llvm.loop [[LOOP5:![0-9]+]]
; CHECK: [[EXIT]]:
-; CHECK-NEXT: [[R:%.*]] = phi ptr [ [[LAST_NEXT]], %[[LOOP]] ], [ [[TMP12]], %[[MIDDLE_BLOCK]] ]
+; CHECK-NEXT: [[R:%.*]] = phi ptr [ [[LAST_NEXT]], %[[LOOP]] ], [ [[TMP11]], %[[MIDDLE_BLOCK]] ]
; CHECK-NEXT: ret ptr [[R]]
;
entry:
diff --git a/llvm/test/Transforms/LoopVectorize/iv-select-cmp-non-const-iv-start.ll b/llvm/test/Transforms/LoopVectorize/iv-select-cmp-non-const-iv-start.ll
index 967c70f094346..ad84f30810955 100644
--- a/llvm/test/Transforms/LoopVectorize/iv-select-cmp-non-const-iv-start.ll
+++ b/llvm/test/Transforms/LoopVectorize/iv-select-cmp-non-const-iv-start.ll
@@ -297,13 +297,12 @@ define i32 @select_trunc_non_const_iv_start_signed_guard(ptr %a, i32 %rdx_start,
; CHECK-VF4IC1-NEXT: [[TMP9:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-VF4IC1-NEXT: br i1 [[TMP9]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
; CHECK-VF4IC1: [[MIDDLE_BLOCK]]:
-; CHECK-VF4IC1-NEXT: [[TMP10:%.*]] = extractelement <4 x i32> [[BROADCAST_SPLAT]], i64 0
-; CHECK-VF4IC1-NEXT: [[TMP11:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.v4i32(<4 x i32> [[TMP8]], <4 x i1> [[TMP7]], i32 [[TMP10]])
+; CHECK-VF4IC1-NEXT: [[TMP14:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.v4i32(<4 x i32> [[TMP8]], <4 x i1> [[TMP7]], i32 [[RDX_START]])
; CHECK-VF4IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[TMP1]], [[N_VEC]]
; CHECK-VF4IC1-NEXT: br i1 [[CMP_N]], label %[[EXIT_LOOPEXIT:.*]], label %[[SCALAR_PH]]
; CHECK-VF4IC1: [[SCALAR_PH]]:
; CHECK-VF4IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[TMP2]], %[[MIDDLE_BLOCK]] ], [ [[TMP0]], %[[FOR_BODY_PREHEADER]] ]
-; CHECK-VF4IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i32 [ [[TMP11]], %[[MIDDLE_BLOCK]] ], [ [[RDX_START]], %[[FOR_BODY_PREHEADER]] ]
+; CHECK-VF4IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i32 [ [[TMP14]], %[[MIDDLE_BLOCK]] ], [ [[RDX_START]], %[[FOR_BODY_PREHEADER]] ]
; CHECK-VF4IC1-NEXT: br label %[[FOR_BODY:.*]]
; CHECK-VF4IC1: [[FOR_BODY]]:
; CHECK-VF4IC1-NEXT: [[IV:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[FOR_BODY]] ]
@@ -317,7 +316,7 @@ define i32 @select_trunc_non_const_iv_start_signed_guard(ptr %a, i32 %rdx_start,
; CHECK-VF4IC1-NEXT: [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], [[WIDE_TRIP_COUNT]]
; CHECK-VF4IC1-NEXT: br i1 [[EXITCOND_NOT]], label %[[EXIT_LOOPEXIT]], label %[[FOR_BODY]], !llvm.loop [[LOOP5:![0-9]+]]
; CHECK-VF4IC1: [[EXIT_LOOPEXIT]]:
-; CHECK-VF4IC1-NEXT: [[COND_LCSSA:%.*]] = phi i32 [ [[COND]], %[[FOR_BODY]] ], [ [[TMP11]], %[[MIDDLE_BLOCK]] ]
+; CHECK-VF4IC1-NEXT: [[COND_LCSSA:%.*]] = phi i32 [ [[COND]], %[[FOR_BODY]] ], [ [[TMP14]], %[[MIDDLE_BLOCK]] ]
; CHECK-VF4IC1-NEXT: br label %[[EXIT]]
; CHECK-VF4IC1: [[EXIT]]:
; CHECK-VF4IC1-NEXT: [[IDX_0_LCSSA:%.*]] = phi i32 [ [[RDX_START]], %[[ENTRY]] ], [ [[COND_LCSSA]], %[[EXIT_LOOPEXIT]] ]
@@ -392,8 +391,7 @@ define i32 @select_trunc_non_const_iv_start_signed_guard(ptr %a, i32 %rdx_start,
; CHECK-VF4IC4-NEXT: [[TMP9:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-VF4IC4-NEXT: br i1 [[TMP9]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
; CHECK-VF4IC4: [[MIDDLE_BLOCK]]:
-; CHECK-VF4IC4-NEXT: [[TMP10:%.*]] = extractelement <4 x i32> [[BROADCAST_SPLAT]], i64 0
-; CHECK-VF4IC4-NEXT: [[TMP11:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.v4i32(<4 x i32> [[TMP8]], <4 x i1> [[TMP7]], i32 [[TMP10]])
+; CHECK-VF4IC4-NEXT: [[TMP11:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.v4i32(<4 x i32> [[TMP8]], <4 x i1> [[TMP7]], i32 [[RDX_START]])
; CHECK-VF4IC4-NEXT: [[TMP34:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.v4i32(<4 x i32> [[TMP28]], <4 x i1> [[TMP24]], i32 [[TMP11]])
; CHECK-VF4IC4-NEXT: [[TMP35:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.v4i32(<4 x i32> [[TMP29]], <4 x i1> [[TMP25]], i32 [[TMP34]])
; CHECK-VF4IC4-NEXT: [[TMP36:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.v4i32(<4 x i32> [[TMP30]], <4 x i1> [[TMP26]], i32 [[TMP35]])
diff --git a/llvm/test/Transforms/LoopVectorize/iv-select-cmp-trunc.ll b/llvm/test/Transforms/LoopVectorize/iv-select-cmp-trunc.ll
index ede9f3eb7ce04..c0dfbbb7ffca9 100644
--- a/llvm/test/Transforms/LoopVectorize/iv-select-cmp-trunc.ll
+++ b/llvm/test/Transforms/LoopVectorize/iv-select-cmp-trunc.ll
@@ -1413,13 +1413,12 @@ define i32 @not_vectorized_select_iv_icmp_no_guard(ptr %a, ptr %b, i32 %start, i
; CHECK-VF4IC1-NEXT: [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-VF4IC1-NEXT: br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP12:![0-9]+]]
; CHECK-VF4IC1: [[MIDDLE_BLOCK]]:
-; CHECK-VF4IC1-NEXT: [[TMP8:%.*]] = extractelement <4 x i32> [[BROADCAST_SPLAT]], i64 0
-; CHECK-VF4IC1-NEXT: [[TMP9:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.v4i32(<4 x i32> [[TMP6]], <4 x i1> [[TMP5]], i32 [[TMP8]])
+; CHECK-VF4IC1-NEXT: [[TMP12:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.v4i32(<4 x i32> [[TMP6]], <4 x i1> [[TMP5]], i32 [[START]])
; CHECK-VF4IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[WIDE_TRIP_COUNT]], [[N_VEC]]
; CHECK-VF4IC1-NEXT: br i1 [[CMP_N]], label %[[EXIT:.*]], label %[[SCALAR_PH]]
; CHECK-VF4IC1: [[SCALAR_PH]]:
; CHECK-VF4IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
-; CHECK-VF4IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i32 [ [[TMP9]], %[[MIDDLE_BLOCK]] ], [ [[START]], %[[ENTRY]] ]
+; CHECK-VF4IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i32 [ [[TMP12]], %[[MIDDLE_BLOCK]] ], [ [[START]], %[[ENTRY]] ]
; CHECK-VF4IC1-NEXT: br label %[[FOR_BODY:.*]]
; CHECK-VF4IC1: [[FOR_BODY]]:
; CHECK-VF4IC1-NEXT: [[IV1:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[SCALAR_PH]] ], [ [[INC:%.*]], %[[FOR_BODY]] ]
@@ -1435,7 +1434,7 @@ define i32 @not_vectorized_select_iv_icmp_no_guard(ptr %a, ptr %b, i32 %start, i
; CHECK-VF4IC1-NEXT: [[EXITCOND_NOT:%.*]] = icmp eq i64 [[INC]], [[WIDE_TRIP_COUNT]]
; CHECK-VF4IC1-NEXT: br i1 [[EXITCOND_NOT]], label %[[EXIT]], label %[[FOR_BODY]], !llvm.loop [[LOOP13:![0-9]+]]
; CHECK-VF4IC1: [[EXIT]]:
-; CHECK-VF4IC1-NEXT: [[COND_LCSSA:%.*]] = phi i32 [ [[COND]], %[[FOR_BODY]] ], [ [[TMP9]], %[[MIDDLE_BLOCK]] ]
+; CHECK-VF4IC1-NEXT: [[COND_LCSSA:%.*]] = phi i32 [ [[COND]], %[[FOR_BODY]] ], [ [[TMP12]], %[[MIDDLE_BLOCK]] ]
; CHECK-VF4IC1-NEXT: ret i32 [[COND_LCSSA]]
;
; CHECK-VF4IC4-LABEL: define i32 @not_vectorized_select_iv_icmp_no_guard(
@@ -1505,8 +1504,7 @@ define i32 @not_vectorized_select_iv_icmp_no_guard(ptr %a, ptr %b, i32 %start, i
; CHECK-VF4IC4-NEXT: [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-VF4IC4-NEXT: br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP12:![0-9]+]]
; CHECK-VF4IC4: [[MIDDLE_BLOCK]]:
-; CHECK-VF4IC4-NEXT: [[TMP8:%.*]] = extractelement <4 x i32> [[BROADCAST_SPLAT]], i64 0
-; CHECK-VF4IC4-NEXT: [[TMP9:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.v4i32(<4 x i32> [[TMP6]], <4 x i1> [[TMP5]], i32 [[TMP8]])
+; CHECK-VF4IC4-NEXT: [[TMP9:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.v4i32(<4 x i32> [[TMP6]], <4 x i1> [[TMP5]], i32 [[START]])
; CHECK-VF4IC4-NEXT: [[TMP35:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.v4i32(<4 x i32> [[TMP29]], <4 x i1> [[TMP25]], i32 [[TMP9]])
; CHECK-VF4IC4-NEXT: [[TMP36:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.v4i32(<4 x i32> [[TMP30]], <4 x i1> [[TMP26]], i32 [[TMP35]])
; CHECK-VF4IC4-NEXT: [[TMP37:%.*]] = call i32 @llvm.experimental.vector.extract.last.active.v4i32(<4 x i32> [[TMP31]], <4 x i1> [[TMP27]], i32 [[TMP36]])
@@ -1900,13 +1898,12 @@ define i16 @not_vectorized_select_iv_icmp_overflow_unwidened_tripcount(ptr %a, p
; CHECK-VF4IC1-NEXT: [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-VF4IC1-NEXT: br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP16:![0-9]+]]
; CHECK-VF4IC1: [[MIDDLE_BLOCK]]:
-; CHECK-VF4IC1-NEXT: [[TMP8:%.*]] = extractelement <4 x i16> [[BROADCAST_SPLAT]], i64 0
-; CHECK-VF4IC1-NEXT: [[TMP9:%.*]] = call i16 @llvm.experimental.vector.extract.last.active.v4i16(<4 x i16> [[TMP6]], <4 x i1> [[TMP5]], i16 [[TMP8]])
+; CHECK-VF4IC1-NEXT: [[TMP12:%.*]] = call i16 @llvm.experimental.vector.extract.last.active.v4i16(<4 x i16> [[TMP6]], <4 x i1> [[TMP5]], i16 [[START]])
; CHECK-VF4IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[WIDE_TRIP_COUNT]], [[N_VEC]]
; CHECK-VF4IC1-NEXT: br i1 [[CMP_N]], label %[[EXIT_LOOPEXIT:.*]], label %[[SCALAR_PH]]
; CHECK-VF4IC1: [[SCALAR_PH]]:
; CHECK-VF4IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[FOR_BODY_PREHEADER]] ]
-; CHECK-VF4IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i16 [ [[TMP9]], %[[MIDDLE_BLOCK]] ], [ [[START]], %[[FOR_BODY_PREHEADER]] ]
+; CHECK-VF4IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i16 [ [[TMP12]], %[[MIDDLE_BLOCK]] ], [ [[START]], %[[FOR_BODY_PREHEADER]] ]
; CHECK-VF4IC1-NEXT: br label %[[FOR_BODY:.*]]
; CHECK-VF4IC1: [[FOR_BODY]]:
; CHECK-VF4IC1-NEXT: [[IV1:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[SCALAR_PH]] ], [ [[INC:%.*]], %[[FOR_BODY]] ]
@@ -1922,7 +1919,7 @@ define i16 @not_vectorized_select_iv_icmp_overflow_unwidened_tripcount(ptr %a, p
; CHECK-VF4IC1-NEXT: [[EXITCOND_NOT:%.*]] = icmp eq i64 [[INC]], [[WIDE_TRIP_COUNT]]
; CHECK-VF4IC1-NEXT: br i1 [[EXITCOND_NOT]], label %[[EXIT_LOOPEXIT]], label %[[FOR_BODY]], !llvm.loop [[LOOP17:![0-9]+]]
; CHECK-VF4IC1: [[EXIT_LOOPEXIT]]:
-; CHECK-VF4IC1-NEXT: [[COND_LCSSA:%.*]] = phi i16 [ [[COND]], %[[FOR_BODY]] ], [ [[TMP9]], %[[MIDDLE_BLOCK]] ]
+; CHECK-VF4IC1-NEXT: [[COND_LCSSA:%.*]] = phi i16 [ [[COND]], %[[FOR_BODY]] ], [ [[TMP12]], %[[MIDDLE_BLOCK]] ]
; CHECK-VF4IC1-NEXT: br label %[[EXIT]]
; CHECK-VF4IC1: [[EXIT]]:
; CHECK-VF4IC1-NEXT: [[RDX_0_LCSSA:%.*]] = phi i16 [ [[START]], %[[ENTRY]] ], [ [[COND_LCSSA]], %[[EXIT_LOOPEXIT]] ]
@@ -1998,8 +1995,7 @@ define i16 @not_vectorized_select_iv_icmp_overflow_unwidened_tripcount(ptr %a, p
; CHECK-VF4IC4-NEXT: [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-VF4IC4-NEXT: br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP16:![0-9]+]]
; CHECK-VF4IC4: [[MIDDLE_BLOCK]]:
-; CHECK-VF4IC4-NEXT: [[TMP8:%.*]] = extractelement <4 x i16> [[BROADCAST_SPLAT]], i64 0
-; CHECK-VF4IC4-NEXT: [[TMP9:%.*]] = call i16 @llvm.experimental.vector.extract.last.active.v4i16(<4 x i16> [[TMP6]], <4 x i1> [[TMP5]], i16 [[TMP8]])
+; CHECK-VF4IC4-NEXT: [[TMP9:%.*]] = call i16 @llvm.experimental.vector.extract.last.active.v4i16(<4 x i16> [[TMP6]], <4 x i1> [[TMP5]], i16 [[START]])
; CHECK-VF4IC4-NEXT: [[TMP35:%.*]] = call i16 @llvm.experimental.vector.extract.last.active.v4i16(<4 x i16> [[TMP29]], <4 x i1> [[TMP25]], i16 [[TMP9]])
; CHECK-VF4IC4-NEXT: [[TMP36:%.*]] = call i16 @llvm.experimental.vector.extract.last.active.v4i16(<4 x i16> [[TMP30]], <4 x i1> [[TMP26]], i16 [[TMP35]])
; CHECK-VF4IC4-NEXT: [[TMP37:%.*]] = call i16 @llvm.experimental.vector.extract.last.active.v4i16(<4 x i16> [[TMP31]], <4 x i1> [[TMP27]], i16 [[TMP36]])
diff --git a/llvm/test/Transforms/LoopVectorize/iv-select-cmp.ll b/llvm/test/Transforms/LoopVectorize/iv-select-cmp.ll
index efbbe480e0041..c9c4615c36760 100644
--- a/llvm/test/Transforms/LoopVectorize/iv-select-cmp.ll
+++ b/llvm/test/Transforms/LoopVectorize/iv-select-cmp.ll
@@ -1892,14 +1892,13 @@ define float @not_vectorized_select_float_induction_icmp(ptr %a, ptr %b, float %
; CHECK-VF4IC1-NEXT: [[TMP10:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-VF4IC1-NEXT: br i1 [[TMP10]], label %[[MIDDLE_BLOCK:.*]], label %[[FOR_BODY]], !llvm.loop [[LOOP20:![0-9]+]]
; CHECK-VF4IC1: [[MIDDLE_BLOCK]]:
-; CHECK-VF4IC1-NEXT: [[TMP11:%.*]] = extractelement <4 x float> [[BROADCAST_SPLAT]], i64 0
-; CHECK-VF4IC1-NEXT: [[TMP12:%.*]] = call float @llvm.experimental.vector.extract.last.active.v4f32(<4 x float> [[TMP9]], <4 x i1> [[TMP8]], float [[TMP11]])
+; CHECK-VF4IC1-NEXT: [[TMP14:%.*]] = call float @llvm.experimental.vector.extract.last.active.v4f32(<4 x float> [[TMP9]], <4 x i1> [[TMP8]], float [[RDX_START]])
; CHECK-VF4IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
; CHECK-VF4IC1-NEXT: br i1 [[CMP_N]], label %[[EXIT:.*]], label %[[SCALAR_PH]]
; CHECK-VF4IC1: [[SCALAR_PH]]:
; CHECK-VF4IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
; CHECK-VF4IC1-NEXT: [[BC_RESUME_VAL2:%.*]] = phi float [ [[TMP13]], %[[MIDDLE_BLOCK]] ], [ 0.000000e+00, %[[ENTRY]] ]
-; CHECK-VF4IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi float [ [[TMP12]], %[[MIDDLE_BLOCK]] ], [ [[RDX_START]], %[[ENTRY]] ]
+; CHECK-VF4IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi float [ [[TMP14]], %[[MIDDLE_BLOCK]] ], [ [[RDX_START]], %[[ENTRY]] ]
; CHECK-VF4IC1-NEXT: br label %[[FOR_BODY1:.*]]
; CHECK-VF4IC1: [[FOR_BODY1]]:
; CHECK-VF4IC1-NEXT: [[IV1:%.*]] = phi i64 [ [[INC:%.*]], %[[FOR_BODY1]] ], [ [[BC_RESUME_VAL]], %[[SCALAR_PH]] ]
@@ -1916,7 +1915,7 @@ define float @not_vectorized_select_float_induction_icmp(ptr %a, ptr %b, float %
; CHECK-VF4IC1-NEXT: [[EXITCOND_NOT:%.*]] = icmp eq i64 [[INC]], [[N]]
; CHECK-VF4IC1-NEXT: br i1 [[EXITCOND_NOT]], label %[[EXIT]], label %[[FOR_BODY1]], !llvm.loop [[LOOP21:![0-9]+]]
; CHECK-VF4IC1: [[EXIT]]:
-; CHECK-VF4IC1-NEXT: [[COND_LCSSA:%.*]] = phi float [ [[COND]], %[[FOR_BODY1]] ], [ [[TMP12]], %[[MIDDLE_BLOCK]] ]
+; CHECK-VF4IC1-NEXT: [[COND_LCSSA:%.*]] = phi float [ [[COND]], %[[FOR_BODY1]] ], [ [[TMP14]], %[[MIDDLE_BLOCK]] ]
; CHECK-VF4IC1-NEXT: ret float [[COND_LCSSA]]
;
; CHECK-VF4IC4-LABEL: define float @not_vectorized_select_float_induction_icmp(
@@ -1988,8 +1987,7 @@ define float @not_vectorized_select_float_induction_icmp(ptr %a, ptr %b, float %
; CHECK-VF4IC4-NEXT: [[TMP10:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-VF4IC4-NEXT: br i1 [[TMP10]], label %[[MIDDLE_BLOCK:.*]], label %[[FOR_BODY]], !llvm.loop [[LOOP20:![0-9]+]]
; CHECK-VF4IC4: [[MIDDLE_BLOCK]]:
-; CHECK-VF4IC4-NEXT: [[TMP11:%.*]] = extractelement <4 x float> [[BROADCAST_SPLAT]], i64 0
-; CHECK-VF4IC4-NEXT: [[TMP12:%.*]] = call float @llvm.experimental.vector.extract.last.active.v4f32(<4 x float> [[TMP9]], <4 x i1> [[TMP8]], float [[TMP11]])
+; CHECK-VF4IC4-NEXT: [[TMP12:%.*]] = call float @llvm.experimental.vector.extract.last.active.v4f32(<4 x float> [[TMP9]], <4 x i1> [[TMP8]], float [[RDX_START]])
; CHECK-VF4IC4-NEXT: [[TMP37:%.*]] = call float @llvm.experimental.vector.extract.last.active.v4f32(<4 x float> [[TMP31]], <4 x i1> [[TMP27]], float [[TMP12]])
; CHECK-VF4IC4-NEXT: [[TMP38:%.*]] = call float @llvm.experimental.vector.extract.last.active.v4f32(<4 x float> [[TMP32]], <4 x i1> [[TMP28]], float [[TMP37]])
; CHECK-VF4IC4-NEXT: [[TMP39:%.*]] = call float @llvm.experimental.vector.extract.last.active.v4f32(<4 x float> [[TMP33]], <4 x i1> [[TMP29]], float [[TMP38]])
>From f7a254fca87c6ab789d9cb305dde3f886d858a75 Mon Sep 17 00:00:00 2001
From: Kazu Hirata <kazu at google.com>
Date: Sat, 26 Sep 2026 09:27:15 -0700
Subject: [PATCH 15/53] [Transforms] Remove unused functions (NFC) (#226655)
createAnyOfReduction:
The last caller was removed on January 18, 2026 in commit
ae1bd068db293c494c4c6314da3b9d138706460d.
canHaveUnrollRemainder:
The last caller, in an assert, was removed on May 1, 2026 in commit
316f0d3bfeaf7eee7b6d4ae60d357a8216ec5264.
Assisted-by: Antigravity
---
.../include/llvm/Transforms/Utils/LoopUtils.h | 5 ----
llvm/lib/Transforms/Utils/LoopUnroll.cpp | 20 -------------
llvm/lib/Transforms/Utils/LoopUtils.cpp | 30 -------------------
3 files changed, 55 deletions(-)
diff --git a/llvm/include/llvm/Transforms/Utils/LoopUtils.h b/llvm/include/llvm/Transforms/Utils/LoopUtils.h
index 74c549be35ddf..a68b928985316 100644
--- a/llvm/include/llvm/Transforms/Utils/LoopUtils.h
+++ b/llvm/include/llvm/Transforms/Utils/LoopUtils.h
@@ -550,11 +550,6 @@ LLVM_ABI Value *createSimpleReduction(IRBuilderBase &B, Value *Src,
RecurKind RdxKind, Value *Mask,
Value *EVL);
-/// Create a reduction of the given vector \p Src for a reduction of kind
-/// RecurKind::AnyOf. The start value of the reduction is \p InitVal.
-LLVM_ABI Value *createAnyOfReduction(IRBuilderBase &B, Value *Src,
- Value *InitVal, PHINode *OrigPhi);
-
/// Create an ordered reduction intrinsic using the given recurrence
/// kind \p RdxKind.
LLVM_ABI Value *createOrderedReduction(IRBuilderBase &B, RecurKind RdxKind,
diff --git a/llvm/lib/Transforms/Utils/LoopUnroll.cpp b/llvm/lib/Transforms/Utils/LoopUnroll.cpp
index e80dd18a34fb7..d7517fd1660ea 100644
--- a/llvm/lib/Transforms/Utils/LoopUnroll.cpp
+++ b/llvm/lib/Transforms/Utils/LoopUnroll.cpp
@@ -426,26 +426,6 @@ void llvm::simplifyLoopAfterUnroll(Loop *L, bool SimplifyIVs, LoopInfo *LI,
}
}
-// Loops containing convergent instructions that are uncontrolled or controlled
-// from outside the loop must have a count that divides their TripMultiple.
-LLVM_ATTRIBUTE_USED
-static bool canHaveUnrollRemainder(const Loop *L) {
- if (getLoopConvergenceHeart(L))
- return false;
-
- // Check for uncontrolled convergent operations.
- for (auto &BB : L->blocks()) {
- for (auto &I : *BB) {
- if (isa<ConvergenceControlInst>(I))
- return true;
- if (auto *CB = dyn_cast<CallBase>(&I))
- if (CB->isConvergent())
- return CB->getConvergenceControlToken();
- }
- }
- return true;
-}
-
// If LoopUnroll has proven OriginalLoopProb is incorrect for some iterations
// of the original loop, adjust latch probabilities in the unrolled loop to
// maintain the original total frequency of the original loop body.
diff --git a/llvm/lib/Transforms/Utils/LoopUtils.cpp b/llvm/lib/Transforms/Utils/LoopUtils.cpp
index 784c833152611..0599f045c4f84 100644
--- a/llvm/lib/Transforms/Utils/LoopUtils.cpp
+++ b/llvm/lib/Transforms/Utils/LoopUtils.cpp
@@ -1497,36 +1497,6 @@ Value *llvm::getShuffleReduction(IRBuilderBase &Builder, Value *Src,
return Builder.CreateExtractElement(TmpVec, Builder.getInt32(0));
}
-Value *llvm::createAnyOfReduction(IRBuilderBase &Builder, Value *Src,
- Value *InitVal, PHINode *OrigPhi) {
- Value *NewVal = nullptr;
-
- // First use the original phi to determine the new value we're trying to
- // select from in the loop.
- SelectInst *SI = nullptr;
- for (auto *U : OrigPhi->users()) {
- if ((SI = dyn_cast<SelectInst>(U)))
- break;
- }
- assert(SI && "One user of the original phi should be a select");
-
- if (SI->getTrueValue() == OrigPhi)
- NewVal = SI->getFalseValue();
- else {
- assert(SI->getFalseValue() == OrigPhi &&
- "At least one input to the select should be the original Phi");
- NewVal = SI->getTrueValue();
- }
-
- // If any predicate is true it means that we want to select the new value.
- Value *AnyOf =
- Src->getType()->isVectorTy() ? Builder.CreateOrReduce(Src) : Src;
- // The compares in the loop may yield poison, which propagates through the
- // bitwise ORs. Freeze it here before the condition is used.
- AnyOf = Builder.CreateFreeze(AnyOf);
- return Builder.CreateSelect(AnyOf, NewVal, InitVal, "rdx.select");
-}
-
Value *llvm::getReductionIdentity(Intrinsic::ID RdxID, Type *Ty,
FastMathFlags Flags) {
bool Negative = false;
>From 0d053e7137ac34f86ba398f1bd25fcb075f9676c Mon Sep 17 00:00:00 2001
From: =?UTF-8?q?Yordan=20V=C3=A1squez?= <vyordangiovani at gmail.com>
Date: Sat, 26 Sep 2026 10:27:50 -0600
Subject: [PATCH 16/53] [libc++] Add static_assert diagnostics for LWG3133
named requirements (#212360)
Add a static_assert to both std::complex<T> and std::valarray<T>
requiring that T be a cv-unqualified object type that satisfies the
Cpp17DefaultConstructible, Cpp17CopyConstructible, Cpp17CopyAssignable,
and Cpp17Destructible named requirements, per the revised wording in
[numeric.requirements]. This mirrors the existing pattern already used
by std::optional<T>.
Non-_v (class-style) trait forms are used throughout so that the
assertion is well-formed even when <complex>/<valarray> are included in
C++03/11/14 mode.
Test coverage:
- A .verify.cpp for complex<T> and one for valarray<T>, each covering
six failure modes: cv-qualified types, and one type violating each of
the four named requirements individually.
Follows-up e062a29cf865bb7cadea6cb605c9f3515e5b883f.
---
libcxx/include/complex | 11 ++++
libcxx/include/valarray | 11 ++++
..._requires_cv_unqualified_object.verify.cpp | 63 +++++++++++++++++++
..._requires_cv_unqualified_object.verify.cpp | 62 ++++++++++++++++++
4 files changed, 147 insertions(+)
create mode 100644 libcxx/test/std/numerics/complex.number/complex/complex_requires_cv_unqualified_object.verify.cpp
create mode 100644 libcxx/test/std/numerics/numarray/template.valarray/valarray_requires_cv_unqualified_object.verify.cpp
diff --git a/libcxx/include/complex b/libcxx/include/complex
index b407bb70156e3..e763688a23bc2 100644
--- a/libcxx/include/complex
+++ b/libcxx/include/complex
@@ -267,9 +267,14 @@ template<class T> complex<T> tanh (const complex<T>&);
# include <__tuple/tuple_size.h>
# include <__type_traits/conditional.h>
# include <__type_traits/is_arithmetic.h>
+# include <__type_traits/is_assignable.h>
+# include <__type_traits/is_constructible.h>
+# include <__type_traits/is_destructible.h>
# include <__type_traits/is_floating_point.h>
# include <__type_traits/is_integral.h>
+# include <__type_traits/is_object.h>
# include <__type_traits/is_same.h>
+# include <__type_traits/is_unqualified.h>
# include <__type_traits/promote.h>
# include <__utility/move.h>
# include <cmath>
@@ -312,6 +317,12 @@ class complex {
public:
typedef _Tp value_type;
+ static_assert(is_object<_Tp>::value && __is_unqualified_v<_Tp> && is_default_constructible<_Tp>::value &&
+ is_copy_constructible<_Tp>::value && is_copy_assignable<_Tp>::value && is_destructible<_Tp>::value,
+ "std::complex<T> requires T to be a cv-unqualified object type that satisfies the "
+ "Cpp17DefaultConstructible, Cpp17CopyConstructible, Cpp17CopyAssignable, and "
+ "Cpp17Destructible requirements");
+
private:
value_type __re_;
value_type __im_;
diff --git a/libcxx/include/valarray b/libcxx/include/valarray
index d0997ebf2caae..df3ce3252e3d2 100644
--- a/libcxx/include/valarray
+++ b/libcxx/include/valarray
@@ -364,6 +364,11 @@ template<class T> valarray<T> tanh (const valarray<T>& x);
# include <__memory/allocator.h>
# include <__memory/uninitialized_algorithms.h>
# include <__type_traits/decay.h>
+# include <__type_traits/is_assignable.h>
+# include <__type_traits/is_constructible.h>
+# include <__type_traits/is_destructible.h>
+# include <__type_traits/is_object.h>
+# include <__type_traits/is_unqualified.h>
# include <__type_traits/remove_reference.h>
# include <__utility/exception_guard.h>
# include <__utility/move.h>
@@ -783,6 +788,12 @@ template <class _Tp>
class valarray {
public:
typedef _Tp value_type;
+
+ static_assert(is_object<_Tp>::value && __is_unqualified_v<_Tp> && is_default_constructible<_Tp>::value &&
+ is_copy_constructible<_Tp>::value && is_copy_assignable<_Tp>::value && is_destructible<_Tp>::value,
+ "std::valarray<T> requires T to be a cv-unqualified object type that satisfies the "
+ "Cpp17DefaultConstructible, Cpp17CopyConstructible, Cpp17CopyAssignable, and "
+ "Cpp17Destructible requirements");
typedef _Tp __result_type;
using iterator = _Tp*;
diff --git a/libcxx/test/std/numerics/complex.number/complex/complex_requires_cv_unqualified_object.verify.cpp b/libcxx/test/std/numerics/complex.number/complex/complex_requires_cv_unqualified_object.verify.cpp
new file mode 100644
index 0000000000000..146cebd63107e
--- /dev/null
+++ b/libcxx/test/std/numerics/complex.number/complex/complex_requires_cv_unqualified_object.verify.cpp
@@ -0,0 +1,63 @@
+//===----------------------------------------------------------------------===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+// XFAIL: FROZEN-CXX03-HEADERS-FIXME
+
+// <complex>
+
+// template<class T>
+// class complex
+//
+// T shall be a cv-unqualified object type (LWG3133)
+
+#include <complex>
+
+struct NotDefaultConstructible {
+ NotDefaultConstructible() = delete;
+};
+
+struct NotCopyConstructible {
+ NotCopyConstructible() = default;
+ NotCopyConstructible(const NotCopyConstructible&) = delete;
+};
+
+struct NotCopyAssignable {
+ NotCopyAssignable() = default;
+ NotCopyAssignable(const NotCopyAssignable&) = default;
+ NotCopyAssignable& operator=(const NotCopyAssignable&) = delete;
+};
+
+struct NotDestructible {
+ NotDestructible() = default;
+ NotDestructible(const NotDestructible&) = default;
+ NotDestructible& operator=(const NotDestructible&) = default;
+
+private:
+ ~NotDestructible() = default;
+};
+
+void test() {
+ // expected-error-re@*:* {{static assertion failed{{.*}}std::complex<T> requires T to be a cv-unqualified object type}}
+ (void)sizeof(std::complex<const double>);
+ // expected-error-re@*:* {{static assertion failed{{.*}}std::complex<T> requires T to be a cv-unqualified object type}}
+ (void)sizeof(std::complex<volatile int>);
+ // expected-error-re@*:* {{static assertion failed{{.*}}std::complex<T> requires T to be a cv-unqualified object type}}
+ (void)sizeof(std::complex<const volatile float>);
+
+ // expected-error-re@*:* {{static assertion failed{{.*}}std::complex<T> requires T to be a cv-unqualified object type}}
+ (void)sizeof(std::complex<NotDefaultConstructible>);
+
+ // expected-error-re@*:* {{static assertion failed{{.*}}std::complex<T> requires T to be a cv-unqualified object type}}
+ (void)sizeof(std::complex<NotCopyConstructible>);
+
+ // expected-error-re@*:* {{static assertion failed{{.*}}std::complex<T> requires T to be a cv-unqualified object type}}
+ (void)sizeof(std::complex<NotCopyAssignable>);
+
+ // expected-error-re@*:* {{static assertion failed{{.*}}std::complex<T> requires T to be a cv-unqualified object type}}
+ (void)sizeof(std::complex<NotDestructible>);
+}
diff --git a/libcxx/test/std/numerics/numarray/template.valarray/valarray_requires_cv_unqualified_object.verify.cpp b/libcxx/test/std/numerics/numarray/template.valarray/valarray_requires_cv_unqualified_object.verify.cpp
new file mode 100644
index 0000000000000..ac7f0a1f7e865
--- /dev/null
+++ b/libcxx/test/std/numerics/numarray/template.valarray/valarray_requires_cv_unqualified_object.verify.cpp
@@ -0,0 +1,62 @@
+//===----------------------------------------------------------------------===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+// XFAIL: FROZEN-CXX03-HEADERS-FIXME
+
+// <valarray>
+
+// template<class T>
+// class valarray
+//
+// T shall be a cv-unqualified object type that satisfies the Cpp17DefaultConstructible,
+// Cpp17CopyConstructible, Cpp17CopyAssignable, and Cpp17Destructible requirements (LWG3133)
+
+#include <valarray>
+
+struct NotDefaultConstructible {
+ NotDefaultConstructible() = delete;
+};
+
+struct NotCopyConstructible {
+ NotCopyConstructible() = default;
+ NotCopyConstructible(const NotCopyConstructible&) = delete;
+};
+
+struct NotCopyAssignable {
+ NotCopyAssignable() = default;
+ NotCopyAssignable(const NotCopyAssignable&) = default;
+ NotCopyAssignable& operator=(const NotCopyAssignable&) = delete;
+};
+
+struct NotDestructible {
+ NotDestructible() = default;
+ NotDestructible(const NotDestructible&) = default;
+ NotDestructible& operator=(const NotDestructible&) = default;
+
+private:
+ ~NotDestructible() = default;
+};
+
+void test() {
+ // expected-error-re@*:* {{static assertion failed{{.*}}std::valarray<T> requires T to be a cv-unqualified object type}}
+ (void)sizeof(std::valarray<const int>);
+ // expected-error-re@*:* {{static assertion failed{{.*}}std::valarray<T> requires T to be a cv-unqualified object type}}
+ (void)sizeof(std::valarray<volatile int>);
+
+ // expected-error-re@*:* {{static assertion failed{{.*}}std::valarray<T> requires T to be a cv-unqualified object type}}
+ (void)sizeof(std::valarray<NotDefaultConstructible>);
+
+ // expected-error-re@*:* {{static assertion failed{{.*}}std::valarray<T> requires T to be a cv-unqualified object type}}
+ (void)sizeof(std::valarray<NotCopyConstructible>);
+
+ // expected-error-re@*:* {{static assertion failed{{.*}}std::valarray<T> requires T to be a cv-unqualified object type}}
+ (void)sizeof(std::valarray<NotCopyAssignable>);
+
+ // expected-error-re@*:* {{static assertion failed{{.*}}std::valarray<T> requires T to be a cv-unqualified object type}}
+ (void)sizeof(std::valarray<NotDestructible>);
+}
>From 5fe49f2e00e57e806ac4ff244726c4b36a9b468a Mon Sep 17 00:00:00 2001
From: geoffreygaren <ggaren at apple.com>
Date: Sat, 26 Sep 2026 09:52:51 -0700
Subject: [PATCH 17/53] [WebKit Checkers] Add built-in recognition for standard
view types (#226350)
libc++ doesn't fully annotate `[[clang::lifetimebound]]` for all view
types. This results in false negatives in borrow checking.
Ultimately we need to fix this in libc++, but for now we can work around
the most common / most important false negatives. For example, borrow
checking can now check
for (auto& x : vector | std::views::reverse) { ... }
Assisted-by: Claude
---
clang/docs/analyzer/checkers.md | 20 ++
.../Checkers/WebKit/ASTUtils.cpp | 104 +++++---
.../Checkers/WebKit/PtrTypesSemantics.cpp | 47 +++-
.../Checkers/WebKit/PtrTypesSemantics.h | 7 +
.../Analysis/Checkers/WebKit/mock-canborrow.h | 30 +++
.../Checkers/WebKit/unborrowed-call-args.cpp | 1 +
.../WebKit/unborrowed-local-vars-cxx23.cpp | 243 +++++++++++++++++-
.../Checkers/WebKit/unborrowed-local-vars.cpp | 45 ++++
8 files changed, 447 insertions(+), 50 deletions(-)
diff --git a/clang/docs/analyzer/checkers.md b/clang/docs/analyzer/checkers.md
index f6b6aa3212c12..b3a64ccd564d4 100644
--- a/clang/docs/analyzer/checkers.md
+++ b/clang/docs/analyzer/checkers.md
@@ -4284,6 +4284,26 @@ The cost is that an identity function is reported even though its result really
> }
> ```
+Includes built-in recognition for std view types. For example:
+
+> ```cpp
+> void foo8(Vector<char>& buffer) {
+> for (char& c : buffer | std::views::reverse) // warn
+> someFunction();
+> }
+>
+> void foo9(Vector<char>& buffer) {
+> // ok, C++23 extends the borrow() temporary across the loop
+> for (char& c : borrow(buffer).get() | std::views::reverse)
+> someFunction();
+> }
+>
+> void foo10(Vector<char>& buffer) {
+> char* p = std::data(buffer); // warn
+> someFunction();
+> }
+> ```
+
#### alpha.webkit.UnborrowedCallArgsChecker
The same rule as alpha.webkit.UnborrowedLocalVarsChecker, applied to function arguments.
diff --git a/clang/lib/StaticAnalyzer/Checkers/WebKit/ASTUtils.cpp b/clang/lib/StaticAnalyzer/Checkers/WebKit/ASTUtils.cpp
index 995667225c961..a1dae480f85b2 100644
--- a/clang/lib/StaticAnalyzer/Checkers/WebKit/ASTUtils.cpp
+++ b/clang/lib/StaticAnalyzer/Checkers/WebKit/ASTUtils.cpp
@@ -38,11 +38,30 @@ static bool tryToFindPtrOriginImpl(
namespace {
+bool isStdViewType(QualType T) {
+ return !T.isNull() &&
+ isStdView(T.getNonReferenceType()->getAsCXXRecordDecl());
+}
+
+void appendPresumedBorrowSources(
+ const FunctionDecl *Callee, ArrayRef<const Expr *> Args,
+ SmallVectorImpl<const Expr *> &LifetimeBoundArgs) {
+ for (unsigned I = 0; I < Args.size(); ++I) {
+ QualType ParamType;
+ if (Callee && I < Callee->getNumParams())
+ ParamType = Callee->getParamDecl(I)->getType();
+ QualType ArgType = Args[I]->getType();
+ if ((!ParamType.isNull() && ParamType->isReferenceType()) ||
+ (!ArgType.isNull() && isView(ArgType)))
+ LifetimeBoundArgs.push_back(Args[I]);
+ }
+}
+
/// Collects the entries of \p Args that \p Callee declares
/// [[clang::lifetimebound]].
void findLifetimeBoundArgs(const FunctionDecl *Callee,
ArrayRef<const Expr *> Args,
- SmallVectorImpl<const Expr *> &BoundArgs) {
+ SmallVectorImpl<const Expr *> &LifetimeBoundArgs) {
if (!Callee)
return;
const FunctionDecl *Canon =
@@ -50,23 +69,29 @@ void findLifetimeBoundArgs(const FunctionDecl *Callee,
unsigned Count = std::min<unsigned>(Canon->getNumParams(), Args.size());
for (unsigned I = 0; I < Count; ++I) {
if (Canon->getParamDecl(I)->hasAttr<LifetimeBoundAttr>())
- BoundArgs.push_back(Args[I]);
+ LifetimeBoundArgs.push_back(Args[I]);
}
}
/// Collects the arguments that \p Construct declares [[clang::lifetimebound]].
+/// Absent annotations, a std view constructor is treated as if libc++ had
+/// annotated it.
void findLifetimeBoundArgs(const CXXConstructExpr *Construct,
- SmallVectorImpl<const Expr *> &BoundArgs) {
- findLifetimeBoundArgs(
- Construct->getConstructor(),
- ArrayRef<const Expr *>(Construct->getArgs(), Construct->getNumArgs()),
- BoundArgs);
+ SmallVectorImpl<const Expr *> &LifetimeBoundArgs) {
+ const auto *Ctor = Construct->getConstructor();
+ ArrayRef<const Expr *> Args(Construct->getArgs(), Construct->getNumArgs());
+ findLifetimeBoundArgs(Ctor, Args, LifetimeBoundArgs);
+ if (!LifetimeBoundArgs.empty() || !Ctor || !isStdView(Ctor->getParent()))
+ return;
+ appendPresumedBorrowSources(Ctor, Args, LifetimeBoundArgs);
}
/// Collects the arguments that \p Call declares [[clang::lifetimebound]],
-/// including the implicit 'this' argument.
+/// including the implicit 'this' argument. Absent annotations, a call that
+/// returns or operates on a std view, or to std::data or std::get, is treated
+/// as if libc++ had annotated it.
void findLifetimeBoundArgs(const CallExpr *Call,
- SmallVectorImpl<const Expr *> &BoundArgs) {
+ SmallVectorImpl<const Expr *> &LifetimeBoundArgs) {
const FunctionDecl *Callee = Call->getDirectCallee();
const Expr *ObjectArg = nullptr;
@@ -77,16 +102,27 @@ void findLifetimeBoundArgs(const CallExpr *Call,
ArgOffset = 1;
} else if (auto *MemberCall = dyn_cast<CXXMemberCallExpr>(Call))
ObjectArg = MemberCall->getImplicitObjectArgument();
+ ArrayRef<const Expr *> Args(Call->getArgs() + ArgOffset,
+ Call->getNumArgs() - ArgOffset);
if (auto *MD = dyn_cast_or_null<CXXMethodDecl>(Callee)) {
if (ObjectArg && lifetimes::implicitObjectParamIsLifetimeBound(MD))
- BoundArgs.push_back(ObjectArg);
+ LifetimeBoundArgs.push_back(ObjectArg);
}
+ findLifetimeBoundArgs(Callee, Args, LifetimeBoundArgs);
+ if (!LifetimeBoundArgs.empty() || !Callee)
+ return;
- findLifetimeBoundArgs(Callee,
- ArrayRef<const Expr *>(Call->getArgs() + ArgOffset,
- Call->getNumArgs() - ArgOffset),
- BoundArgs);
+ bool IsStdAccessor =
+ Callee->isInStdNamespace() &&
+ (safeGetName(Callee) == "data" || safeGetName(Callee) == "get");
+ if (!isStdViewType(Callee->getReturnType()) &&
+ !(ObjectArg && isStdViewType(ObjectArg->getType())) && !IsStdAccessor)
+ return;
+
+ if (ObjectArg)
+ LifetimeBoundArgs.push_back(ObjectArg);
+ appendPresumedBorrowSources(Callee, Args, LifetimeBoundArgs);
}
/// Traces each of \p Args independently and requires every one to be safe.
@@ -165,20 +201,14 @@ static bool tryToFindPtrOriginImpl(
PtrIsLifetimeBoundToOrigin);
if (FollowLifetimeBound) {
- SmallVector<const Expr *, 2> BoundArgs;
- findLifetimeBoundArgs(tempExpr, BoundArgs);
- if (!BoundArgs.empty())
- PtrIsLifetimeBoundToOrigin = true;
- if (BoundArgs.size() == 1) {
- E = BoundArgs.front();
- continue;
- }
- if (BoundArgs.size() > 1)
+ SmallVector<const Expr *, 2> LifetimeBoundArgs;
+ findLifetimeBoundArgs(tempExpr, LifetimeBoundArgs);
+ if (!LifetimeBoundArgs.empty())
return tryToFindPtrOriginOfEach(
- BoundArgs, StopAtFirstRefCountedObj, isSafePtr, isSafePtrType,
- isSafeGlobalDecl, callback,
- OriginDependsOnFullExpressionTemporary,
- PtrIsLifetimeBoundToOrigin);
+ LifetimeBoundArgs, StopAtFirstRefCountedObj, isSafePtr,
+ isSafePtrType, isSafeGlobalDecl, callback,
+ /*OriginDependsOnFullExpressionTemporary=*/false,
+ /*PtrIsLifetimeBoundToOrigin=*/true);
}
break;
}
@@ -370,20 +400,14 @@ static bool tryToFindPtrOriginImpl(
}
if (FollowLifetimeBound) {
- SmallVector<const Expr *, 2> BoundArgs;
- findLifetimeBoundArgs(call, BoundArgs);
- if (!BoundArgs.empty())
- PtrIsLifetimeBoundToOrigin = true;
- if (BoundArgs.size() == 1) {
- E = BoundArgs.front();
- continue;
- }
- if (BoundArgs.size() > 1)
+ SmallVector<const Expr *, 2> LifetimeBoundArgs;
+ findLifetimeBoundArgs(call, LifetimeBoundArgs);
+ if (!LifetimeBoundArgs.empty())
return tryToFindPtrOriginOfEach(
- BoundArgs, StopAtFirstRefCountedObj, isSafePtr, isSafePtrType,
- isSafeGlobalDecl, callback,
- OriginDependsOnFullExpressionTemporary,
- PtrIsLifetimeBoundToOrigin);
+ LifetimeBoundArgs, StopAtFirstRefCountedObj, isSafePtr,
+ isSafePtrType, isSafeGlobalDecl, callback,
+ /*OriginDependsOnFullExpressionTemporary=*/false,
+ /*PtrIsLifetimeBoundToOrigin=*/true);
}
}
if (auto *ObjCMsgExpr = dyn_cast<ObjCMessageExpr>(E)) {
diff --git a/clang/lib/StaticAnalyzer/Checkers/WebKit/PtrTypesSemantics.cpp b/clang/lib/StaticAnalyzer/Checkers/WebKit/PtrTypesSemantics.cpp
index b38bbba2c173f..a89c5b432bd75 100644
--- a/clang/lib/StaticAnalyzer/Checkers/WebKit/PtrTypesSemantics.cpp
+++ b/clang/lib/StaticAnalyzer/Checkers/WebKit/PtrTypesSemantics.cpp
@@ -17,6 +17,7 @@
#include "clang/AST/StmtVisitor.h"
#include "clang/Analysis/Analyses/LifetimeSafety/LifetimeAnnotations.h"
#include "clang/Analysis/DomainSpecific/CocoaConventions.h"
+#include "llvm/ADT/StringSet.h"
#include <optional>
using namespace clang;
@@ -183,12 +184,56 @@ static bool hasLifetimeBoundCtor(const clang::CXXRecordDecl *R) {
return false;
}
+static bool isStdRangesViewInterface(const clang::CXXRecordDecl *R) {
+ if (!R || !R->getIdentifier() || R->getName() != "view_interface")
+ return false;
+ const auto *NS = dyn_cast<NamespaceDecl>(R->getDeclContext());
+ return NS && NS->getIdentifier() && NS->getName() == "ranges" &&
+ NS->getParent()->isStdNamespace();
+}
+
+static bool derivesFromViewInterface(const clang::CXXRecordDecl *R) {
+ if (!R)
+ return false;
+ R = R->getDefinition();
+ if (!R)
+ return false;
+ if (isStdRangesViewInterface(R))
+ return true;
+ for (const CXXBaseSpecifier &Base : R->bases()) {
+ if (derivesFromViewInterface(Base.getType()->getAsCXXRecordDecl()))
+ return true;
+ }
+ return false;
+}
+
+bool isStdView(const clang::CXXRecordDecl *R) {
+ if (!R)
+ return false;
+ if (R->hasAttr<PointerAttr>())
+ return true;
+ static const llvm::StringSet<> StdIterators{
+ "reverse_iterator", "move_iterator", "common_iterator",
+ "counted_iterator", "basic_const_iterator"};
+ if (R->isInStdNamespace() && R->getIdentifier() &&
+ StdIterators.contains(R->getName()))
+ return true;
+ if (derivesFromViewInterface(R))
+ return true;
+ if (const auto *Parent = dyn_cast<CXXRecordDecl>(R->getDeclContext()))
+ return isStdView(Parent);
+ return false;
+}
+
bool isView(const clang::QualType T) {
if (T->isReferenceType())
return true;
if (lifetimes::isPointerLikeType(T))
return true;
- return hasLifetimeBoundCtor(T->getAsCXXRecordDecl());
+ auto *Record = T->getAsCXXRecordDecl();
+ if (isStdView(Record))
+ return true;
+ return hasLifetimeBoundCtor(Record);
}
bool isRefType(const std::string &Name) {
diff --git a/clang/lib/StaticAnalyzer/Checkers/WebKit/PtrTypesSemantics.h b/clang/lib/StaticAnalyzer/Checkers/WebKit/PtrTypesSemantics.h
index 9e9bb995f7ca2..7fb78233a6288 100644
--- a/clang/lib/StaticAnalyzer/Checkers/WebKit/PtrTypesSemantics.h
+++ b/clang/lib/StaticAnalyzer/Checkers/WebKit/PtrTypesSemantics.h
@@ -73,6 +73,13 @@ clang::QualType borrowedType(clang::QualType T);
/// \returns true if a value of type \p T is a pointer/reference/view.
bool isView(const clang::QualType T);
+/// \returns true if \p Class declares reference semantics structurally: it is
+/// annotated [[gsl::Pointer]] (explicitly, or by Sema's inference for
+/// standard types), derives from std::ranges::view_interface, is a standard
+/// iterator adaptor, or is nested inside such a class, as the iterators of
+/// standard views are.
+bool isStdView(const clang::CXXRecordDecl *Class);
+
/// \returns true if \p Class is ref-counted, false if not.
bool isRefCounted(const clang::CXXRecordDecl *Class);
diff --git a/clang/test/Analysis/Checkers/WebKit/mock-canborrow.h b/clang/test/Analysis/Checkers/WebKit/mock-canborrow.h
index 7b68bba4999a3..fc756714dc4c1 100644
--- a/clang/test/Analysis/Checkers/WebKit/mock-canborrow.h
+++ b/clang/test/Analysis/Checkers/WebKit/mock-canborrow.h
@@ -255,4 +255,34 @@ class Function {
void callEscaping(const Function &);
void callNoEscape([[clang::noescape]] const Function &);
+namespace std {
+inline namespace __1 {
+using size_t = decltype(sizeof(0));
+
+namespace ranges {
+template <typename Derived> class view_interface {};
+} // namespace ranges
+
+template <typename Iterator> class reverse_iterator {
+public:
+ reverse_iterator(Iterator);
+ auto &operator*() const { return *m_it; }
+ reverse_iterator &operator++();
+ bool operator!=(const reverse_iterator &) const;
+
+private:
+ Iterator m_it;
+};
+
+template <typename A, typename B> struct pair {
+ A first;
+ B second;
+};
+
+template <size_t I, typename A, typename B> A &get(pair<A, B> &);
+
+template <typename T> T *data(Vector<T> &);
+} // namespace __1
+} // namespace std
+
#endif
diff --git a/clang/test/Analysis/Checkers/WebKit/unborrowed-call-args.cpp b/clang/test/Analysis/Checkers/WebKit/unborrowed-call-args.cpp
index 72cd785725b41..5de9efce09423 100644
--- a/clang/test/Analysis/Checkers/WebKit/unborrowed-call-args.cpp
+++ b/clang/test/Analysis/Checkers/WebKit/unborrowed-call-args.cpp
@@ -192,6 +192,7 @@ namespace known_gaps {
void unannotated_intermediate() {
Vector<char> vec;
takeSpan(makeSpanUnannotated(vec));
+ // expected-warning at -1{{Function argument 'makeSpanUnannotated(vec)' (to 'takeSpan') is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedCallArgsChecker]}}
}
inline void trivialSink(char &c) {}
diff --git a/clang/test/Analysis/Checkers/WebKit/unborrowed-local-vars-cxx23.cpp b/clang/test/Analysis/Checkers/WebKit/unborrowed-local-vars-cxx23.cpp
index 0f1ae4e58d6ac..3fc233b0c0abb 100644
--- a/clang/test/Analysis/Checkers/WebKit/unborrowed-local-vars-cxx23.cpp
+++ b/clang/test/Analysis/Checkers/WebKit/unborrowed-local-vars-cxx23.cpp
@@ -11,24 +11,249 @@ void borrow_function_get_loop(Vector<char> &vec) {
}
}
-struct ReversedChars {
- char *b;
- char *e;
- char *begin() const;
- char *end() const;
-};
struct ReverseAdaptor {};
-ReversedChars operator|(Vector<char> &vec, ReverseAdaptor);
+struct ReversedChars : std::ranges::view_interface<ReversedChars> {
+ explicit ReversedChars(Vector<char> &);
+ struct Iterator {
+ char &operator*() const;
+ Iterator &operator++();
+ bool operator!=(const Iterator &) const;
+ };
+ Iterator begin() const;
+ Iterator end() const;
+ ReversedChars zipWith(Vector<int> &) const;
+};
+inline constexpr ReverseAdaptor reversed{};
+ReversedChars operator|(Vector<char> &, const ReverseAdaptor &);
+ReversedChars operator|(ReversedChars &&, const ReverseAdaptor &);
-void borrow_get_pipe_loop(Vector<char> &vec) {
- for (char &c : borrow(vec).get() | ReverseAdaptor()) {
+void unguarded_global_adaptor_pipe_loop(Vector<char> &vec) {
+ for (char &c : vec | reversed) {
+ // expected-warning at -1{{Local variable 'c' is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ someFunction();
+ (void)c;
+ }
+}
+
+void borrowed_global_adaptor_pipe_loop(Vector<char> &vec) {
+ for (char &c : borrow(vec).get() | reversed) {
someFunction();
(void)c;
}
}
+void chained_pipe_unguarded(Vector<char> &vec) {
+ ReversedChars rv = vec | reversed | reversed;
+ // expected-warning at -1{{Local variable 'rv' is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ someFunction();
+ (void)rv;
+}
+
+void chained_pipe_borrowed(Vector<char> &vec) {
+ Borrow<Vector<char>> b(vec);
+ ReversedChars rv = b.get() | reversed | reversed;
+ someFunction();
+ (void)rv;
+}
+
void unguarded_pipe_loop(Vector<char> &vec) {
for (char &c : vec | ReverseAdaptor()) {
+ // expected-warning at -1{{Local variable 'c' is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ someFunction();
+ (void)c;
+ }
+}
+
+void borrow_get_pipe_loop(Vector<char> &vec) {
+ for (char &c : borrow(vec).get() | ReverseAdaptor()) {
+ someFunction();
+ (void)c;
+ }
+}
+
+void named_view(Vector<char> &vec) {
+ ReversedChars rv = vec | ReverseAdaptor();
+ // expected-warning at -1{{Local variable 'rv' is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ someFunction();
+ (void)rv;
+}
+
+void named_view_borrowed(Vector<char> &vec) {
+ Borrow<Vector<char>> b(vec);
+ ReversedChars rv = b.get() | ReverseAdaptor();
+ someFunction();
+ (void)rv;
+}
+
+void constructed_view(Vector<char> &vec) {
+ ReversedChars rv(vec);
+ // expected-warning at -1{{Local variable 'rv' is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ someFunction();
+ (void)rv;
+}
+
+struct PlainReversed : std::ranges::view_interface<PlainReversed> {
+ explicit PlainReversed(Vector<char> &);
+ std::reverse_iterator<char *> begin() const;
+ std::reverse_iterator<char *> end() const;
+};
+
+void std_reverse_iterator_loop(Vector<char> &vec) {
+ for (char &c : PlainReversed(vec)) {
+ // expected-warning at -1{{Local variable 'c' is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ someFunction();
+ (void)c;
+ }
+}
+
+void std_reverse_iterator_borrowed(Vector<char> &vec) {
+ Borrow<Vector<char>> b(vec);
+ for (char &c : PlainReversed(b.get())) {
+ someFunction();
+ (void)c;
+ }
+}
+
+void data_from_vector(Vector<char> &vec) {
+ char *p = std::data(vec);
+ // expected-warning at -1{{Local variable 'p' is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ someFunction();
+ (void)p;
+}
+
+void data_from_borrow(Vector<char> &vec) {
+ Borrow<Vector<char>> b(vec);
+ char *p = std::data(b.get());
+ someFunction();
+ (void)p;
+}
+
+void get_from_element(Vector<std::pair<int, int>> &vec) {
+ auto &first = std::get<0>(vec[0]);
+ // expected-warning at -1{{Local variable 'first' is a loan on CanBorrow type 'Vector<std::pair<int, int>>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ someFunction();
+ (void)first;
+}
+
+void get_from_borrowed_element(Vector<std::pair<int, int>> &vec) {
+ Borrow<Vector<std::pair<int, int>>> b(vec);
+ auto &first = std::get<0>(b.get()[0]);
+ someFunction();
+ (void)first;
+}
+
+struct AnnotatedView : std::ranges::view_interface<AnnotatedView> {
+ AnnotatedView(Vector<char> &tracked LIFETIME_BOUND, Vector<char> &untracked);
+};
+
+void trusted_annotations(Vector<char> &tracked, Vector<char> &untracked) {
+ Borrow<Vector<char>> b(tracked);
+ AnnotatedView v(b.get(), untracked);
+ someFunction();
+ (void)v;
+}
+
+void member_arg_unguarded(Vector<char> &vec, Vector<int> &ints) {
+ Borrow<Vector<char>> b(vec);
+ ReversedChars rv = (b.get() | reversed).zipWith(ints);
+ // expected-warning at -1{{Local variable 'rv' is a loan on CanBorrow type 'Vector<int>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ someFunction();
+ (void)rv;
+}
+
+void member_object_unguarded(Vector<char> &vec, Vector<int> &ints) {
+ Borrow<Vector<int>> b(ints);
+ ReversedChars rv = (vec | reversed).zipWith(b.get());
+ // expected-warning at -1{{Local variable 'rv' is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ someFunction();
+ (void)rv;
+}
+
+void member_both_borrowed(Vector<char> &vec, Vector<int> &ints) {
+ Borrow<Vector<char>> bc(vec);
+ Borrow<Vector<int>> bi(ints);
+ ReversedChars rv = (bc.get() | reversed).zipWith(bi.get());
+ someFunction();
+ (void)rv;
+}
+
+struct ZipView : std::ranges::view_interface<ZipView> {
+ ZipView(Vector<char> &, Vector<int> &);
+ struct Iterator {
+ char &operator*() const;
+ Iterator &operator++();
+ bool operator!=(const Iterator &) const;
+ };
+ Iterator begin() const;
+ Iterator end() const;
+};
+struct ZipAdaptor {
+ ZipView operator()(Vector<char> &, Vector<int> &) const;
+};
+inline constexpr ZipAdaptor zip{};
+
+void zip_unguarded(Vector<char> &vec, Vector<int> &ints) {
+ for (char &c : zip(vec, ints)) {
+ // expected-warning at -1{{Local variable 'c' is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ someFunction();
+ (void)c;
+ }
+}
+
+void zip_first_borrowed(Vector<char> &vec, Vector<int> &ints) {
+ Borrow<Vector<char>> b(vec);
+ for (char &c : zip(b.get(), ints)) {
+ // expected-warning at -1{{Local variable 'c' is a loan on CanBorrow type 'Vector<int>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ someFunction();
+ (void)c;
+ }
+}
+
+void zip_second_borrowed(Vector<char> &vec, Vector<int> &ints) {
+ Borrow<Vector<int>> b(ints);
+ for (char &c : zip(vec, b.get())) {
+ // expected-warning at -1{{Local variable 'c' is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ someFunction();
+ (void)c;
+ }
+}
+
+void zip_both_borrowed(Vector<char> &vec, Vector<int> &ints) {
+ Borrow<Vector<char>> bc(vec);
+ Borrow<Vector<int>> bi(ints);
+ for (char &c : zip(bc.get(), bi.get())) {
+ someFunction();
+ (void)c;
+ }
+}
+
+void zip_constructed_second_borrowed(Vector<char> &vec, Vector<int> &ints) {
+ Borrow<Vector<int>> b(ints);
+ ZipView z(vec, b.get());
+ // expected-warning at -1{{Local variable 'z' is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ someFunction();
+ (void)z;
+}
+
+void zip_constructed_both_borrowed(Vector<char> &vec, Vector<int> &ints) {
+ Borrow<Vector<char>> bc(vec);
+ Borrow<Vector<int>> bi(ints);
+ ZipView z(bc.get(), bi.get());
+ someFunction();
+ (void)z;
+}
+
+struct NonStdAdaptor {};
+struct NonStdReversed {
+ char *b;
+ char *e;
+ char *begin() const;
+ char *end() const;
+};
+NonStdReversed operator|(Vector<char> &, NonStdAdaptor);
+
+void non_std_pipe_loop(Vector<char> &vec) {
+ for (char &c : vec | NonStdAdaptor()) {
someFunction();
(void)c;
}
diff --git a/clang/test/Analysis/Checkers/WebKit/unborrowed-local-vars.cpp b/clang/test/Analysis/Checkers/WebKit/unborrowed-local-vars.cpp
index 8de699f967146..ac2c7dc082e97 100644
--- a/clang/test/Analysis/Checkers/WebKit/unborrowed-local-vars.cpp
+++ b/clang/test/Analysis/Checkers/WebKit/unborrowed-local-vars.cpp
@@ -368,15 +368,60 @@ void guarded_store_through_out_pointer(Vector<char> &vec, char **out) {
}
} // namespace escape_paths
+struct OwningBuffer : CanBorrow {
+ char *data() LIFETIME_BOUND;
+};
+OwningBuffer makeOwningBuffer(const Vector<char> &vec LIFETIME_BOUND);
+extern const Vector<char> globalVec;
+
+void owning_temporary_from_global() {
+ char *p = makeOwningBuffer(globalVec).data();
+ // expected-warning at -1{{temporary whose address is used as value of local variable 'p' will be destroyed at the end of the full-expression}}
+ someFunction();
+ (void)p;
+}
+
+void owning_temporary_from_borrow(Vector<char> &vec) {
+ Borrow<Vector<char>> b(vec);
+ char *p = makeOwningBuffer(b.get()).data();
+ // expected-warning at -1{{temporary whose address is used as value of local variable 'p' will be destroyed at the end of the full-expression}}
+ someFunction();
+ (void)p;
+}
+
+void owning_temporary_from_unguarded(Vector<char> &vec) {
+ char *p = makeOwningBuffer(vec).data();
+ // expected-warning at -1{{Local variable 'p' is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ // expected-warning at -2{{temporary whose address is used as value of local variable 'p' will be destroyed at the end of the full-expression}}
+ someFunction();
+ (void)p;
+}
+
+void extended_owning_from_unguarded(Vector<char> &vec) {
+ const OwningBuffer &buf = makeOwningBuffer(vec);
+ // expected-warning at -1{{Local variable 'buf' is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
+ someFunction();
+ (void)buf;
+}
+
+void extended_owning_from_borrow(Vector<char> &vec) {
+ Borrow<Vector<char>> b(vec);
+ const OwningBuffer &buf = makeOwningBuffer(b.get());
+ someFunction();
+ (void)buf;
+}
+
namespace known_gaps {
void unannotated_view_constructor(Vector<char> &vec) {
CharSpan s(vec.data());
+ // expected-warning at -1{{Local variable 's' is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
someFunction();
(void)s;
}
void unannotated_function_parameter(Vector<char> &vec) {
CharSpan s = makeSpanUnannotated(vec);
+ // expected-warning at -1{{Local variable 's' is a loan on CanBorrow type 'Vector<char>' that is not guarded by a Borrow [alpha.webkit.UnborrowedLocalVarsChecker]}}
someFunction();
(void)s;
}
>From 49f73a70eb3b170c440ccb726299a3bddbdbe8a1 Mon Sep 17 00:00:00 2001
From: Simon Pilgrim <llvm-dev at redking.me.uk>
Date: Sat, 26 Sep 2026 17:58:02 +0100
Subject: [PATCH 18/53] [CostModel][X86] arith-fp.ll - test AVX512DQ instead of
AVX512BW (#226519)
AVX512DQ has instructions relevant to fp arithmetic (vXi64 fp2int in particular)
---
llvm/test/Analysis/CostModel/X86/arith-fp.ll | 35 +++++++++++++-------
1 file changed, 23 insertions(+), 12 deletions(-)
diff --git a/llvm/test/Analysis/CostModel/X86/arith-fp.ll b/llvm/test/Analysis/CostModel/X86/arith-fp.ll
index e3eb3e60b844c..cd2679d1e0336 100644
--- a/llvm/test/Analysis/CostModel/X86/arith-fp.ll
+++ b/llvm/test/Analysis/CostModel/X86/arith-fp.ll
@@ -4,8 +4,8 @@
; RUN: opt < %s -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mtriple=x86_64-- -mattr=+sse4.2 | FileCheck %s --check-prefixes=SSE42
; RUN: opt < %s -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mtriple=x86_64-- -mattr=+avx | FileCheck %s --check-prefixes=AVX,AVX1
; RUN: opt < %s -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mtriple=x86_64-- -mattr=+avx2 | FileCheck %s --check-prefixes=AVX,AVX2
-; RUN: opt < %s -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mtriple=x86_64-- -mattr=+avx512f | FileCheck %s --check-prefixes=AVX512
-; RUN: opt < %s -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mtriple=x86_64-- -mattr=+avx512f,+avx512bw | FileCheck %s --check-prefixes=AVX512
+; RUN: opt < %s -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mtriple=x86_64-- -mattr=+avx512f | FileCheck %s --check-prefixes=AVX512,AVX512F
+; RUN: opt < %s -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mtriple=x86_64-- -mattr=+avx512f,+avx512dq | FileCheck %s --check-prefixes=AVX512,AVX512DQ
;
; RUN: opt < %s -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mtriple=x86_64-- -mcpu=slm | FileCheck %s --check-prefixes=SLM
; RUN: opt < %s -passes="print<cost-model>" 2>&1 -disable-output -cost-kind=all -mtriple=x86_64-- -mcpu=goldmont | FileCheck %s --check-prefixes=GLM
@@ -1327,16 +1327,27 @@ define i32 @llrint(i32 %arg) {
; AVX-NEXT: Cost Model: Found costs of RThru:28 CodeSize:1 Lat:1 SizeLat:1 for: %V8F64 = call <8 x i64> @llvm.llrint.v8i64.v8f64(<8 x double> undef)
; AVX-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret i32 undef
;
-; AVX512-LABEL: 'llrint'
-; AVX512-NEXT: Cost Model: Found costs of 1 for: %F32 = call i64 @llvm.llrint.i64.f32(float undef)
-; AVX512-NEXT: Cost Model: Found costs of RThru:14 CodeSize:1 Lat:1 SizeLat:1 for: %V4F32 = call <4 x i64> @llvm.llrint.v4i64.v4f32(<4 x float> undef)
-; AVX512-NEXT: Cost Model: Found costs of RThru:30 CodeSize:1 Lat:1 SizeLat:1 for: %V8F32 = call <8 x i64> @llvm.llrint.v8i64.v8f32(<8 x float> undef)
-; AVX512-NEXT: Cost Model: Found costs of RThru:61 CodeSize:1 Lat:1 SizeLat:1 for: %V16F32 = call <16 x i64> @llvm.llrint.v16i64.v16f32(<16 x float> undef)
-; AVX512-NEXT: Cost Model: Found costs of 1 for: %F64 = call i64 @llvm.llrint.i64.f64(double undef)
-; AVX512-NEXT: Cost Model: Found costs of RThru:6 CodeSize:1 Lat:1 SizeLat:1 for: %V2F64 = call <2 x i64> @llvm.llrint.v2i64.v2f64(<2 x double> undef)
-; AVX512-NEXT: Cost Model: Found costs of RThru:14 CodeSize:1 Lat:1 SizeLat:1 for: %V4F64 = call <4 x i64> @llvm.llrint.v4i64.v4f64(<4 x double> undef)
-; AVX512-NEXT: Cost Model: Found costs of RThru:30 CodeSize:1 Lat:1 SizeLat:1 for: %V8F64 = call <8 x i64> @llvm.llrint.v8i64.v8f64(<8 x double> undef)
-; AVX512-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret i32 undef
+; AVX512F-LABEL: 'llrint'
+; AVX512F-NEXT: Cost Model: Found costs of 1 for: %F32 = call i64 @llvm.llrint.i64.f32(float undef)
+; AVX512F-NEXT: Cost Model: Found costs of RThru:14 CodeSize:1 Lat:1 SizeLat:1 for: %V4F32 = call <4 x i64> @llvm.llrint.v4i64.v4f32(<4 x float> undef)
+; AVX512F-NEXT: Cost Model: Found costs of RThru:30 CodeSize:1 Lat:1 SizeLat:1 for: %V8F32 = call <8 x i64> @llvm.llrint.v8i64.v8f32(<8 x float> undef)
+; AVX512F-NEXT: Cost Model: Found costs of RThru:61 CodeSize:1 Lat:1 SizeLat:1 for: %V16F32 = call <16 x i64> @llvm.llrint.v16i64.v16f32(<16 x float> undef)
+; AVX512F-NEXT: Cost Model: Found costs of 1 for: %F64 = call i64 @llvm.llrint.i64.f64(double undef)
+; AVX512F-NEXT: Cost Model: Found costs of RThru:6 CodeSize:1 Lat:1 SizeLat:1 for: %V2F64 = call <2 x i64> @llvm.llrint.v2i64.v2f64(<2 x double> undef)
+; AVX512F-NEXT: Cost Model: Found costs of RThru:14 CodeSize:1 Lat:1 SizeLat:1 for: %V4F64 = call <4 x i64> @llvm.llrint.v4i64.v4f64(<4 x double> undef)
+; AVX512F-NEXT: Cost Model: Found costs of RThru:30 CodeSize:1 Lat:1 SizeLat:1 for: %V8F64 = call <8 x i64> @llvm.llrint.v8i64.v8f64(<8 x double> undef)
+; AVX512F-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret i32 undef
+;
+; AVX512DQ-LABEL: 'llrint'
+; AVX512DQ-NEXT: Cost Model: Found costs of 1 for: %F32 = call i64 @llvm.llrint.i64.f32(float undef)
+; AVX512DQ-NEXT: Cost Model: Found costs of 1 for: %V4F32 = call <4 x i64> @llvm.llrint.v4i64.v4f32(<4 x float> undef)
+; AVX512DQ-NEXT: Cost Model: Found costs of 1 for: %V8F32 = call <8 x i64> @llvm.llrint.v8i64.v8f32(<8 x float> undef)
+; AVX512DQ-NEXT: Cost Model: Found costs of RThru:3 CodeSize:1 Lat:1 SizeLat:1 for: %V16F32 = call <16 x i64> @llvm.llrint.v16i64.v16f32(<16 x float> undef)
+; AVX512DQ-NEXT: Cost Model: Found costs of 1 for: %F64 = call i64 @llvm.llrint.i64.f64(double undef)
+; AVX512DQ-NEXT: Cost Model: Found costs of 1 for: %V2F64 = call <2 x i64> @llvm.llrint.v2i64.v2f64(<2 x double> undef)
+; AVX512DQ-NEXT: Cost Model: Found costs of 1 for: %V4F64 = call <4 x i64> @llvm.llrint.v4i64.v4f64(<4 x double> undef)
+; AVX512DQ-NEXT: Cost Model: Found costs of 1 for: %V8F64 = call <8 x i64> @llvm.llrint.v8i64.v8f64(<8 x double> undef)
+; AVX512DQ-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret i32 undef
;
; SLM-LABEL: 'llrint'
; SLM-NEXT: Cost Model: Found costs of 1 for: %F32 = call i64 @llvm.llrint.i64.f32(float undef)
>From 82c9d0f64176e40255ca75a062a60e6a2e52933d Mon Sep 17 00:00:00 2001
From: Florian Hahn <flo at fhahn.com>
Date: Sat, 26 Sep 2026 18:07:08 +0100
Subject: [PATCH 19/53] [LV] Add tests for wide induction wrap flags (NFC).
(#226715)
Add tests for:
* inductions where the lane offsets may signed-overflow, even if the
scalar induction values do not,
* inductions whose increment does not directly update the phi, where the
increment's wrap flags do not apply to the induction,
* narrow inductions where the VF may exceed the signed maximum of the
induction type.
---
.../AArch64/sve-induction-wrapflags.ll | 162 +++++++++++++++
.../LoopVectorize/induction-wrapflags.ll | 190 ++++++++++++++++++
2 files changed, 352 insertions(+)
create mode 100644 llvm/test/Transforms/LoopVectorize/AArch64/sve-induction-wrapflags.ll
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/sve-induction-wrapflags.ll b/llvm/test/Transforms/LoopVectorize/AArch64/sve-induction-wrapflags.ll
new file mode 100644
index 0000000000000..f39516c3a327d
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/sve-induction-wrapflags.ll
@@ -0,0 +1,162 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --check-globals none --filter-out-after "^scalar.ph:" --version 6
+; RUN: opt -p loop-vectorize -force-vector-width="vscale x 8" -force-vector-interleave=1 -S %s | FileCheck %s
+
+target triple = "aarch64-linux-gnu"
+
+; The VF may exceed the signed maximum of i8, so the lane indices may wrap.
+define void @i8_induction_vf_may_exceed_signed_max(ptr noalias %dst, i64 %n) #0 {
+; CHECK-LABEL: define void @i8_induction_vf_may_exceed_signed_max(
+; CHECK-SAME: ptr noalias [[DST:%.*]], i64 [[N:%.*]]) #[[ATTR0:[0-9]+]] {
+; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: [[TMP0:%.*]] = call i64 @llvm.vscale.i64()
+; CHECK-NEXT: [[TMP1:%.*]] = shl nuw i64 [[TMP0]], 3
+; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], [[TMP1]]
+; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: [[N_MOD_VF:%.*]] = urem i64 [[N]], [[TMP1]]
+; CHECK-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[N_MOD_VF]]
+; CHECK-NEXT: [[TMP2:%.*]] = trunc i64 [[N_VEC]] to i8
+; CHECK-NEXT: [[TMP3:%.*]] = sub i8 0, [[TMP2]]
+; CHECK-NEXT: [[TMP4:%.*]] = call <vscale x 8 x i8> @llvm.stepvector.nxv8i8()
+; CHECK-NEXT: [[TMP5:%.*]] = sub nsw <vscale x 8 x i8> zeroinitializer, [[TMP4]]
+; CHECK-NEXT: [[TMP6:%.*]] = trunc i64 [[TMP1]] to i8
+; CHECK-NEXT: [[TMP7:%.*]] = sub nsw i8 0, [[TMP6]]
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <vscale x 8 x i8> poison, i8 [[TMP7]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <vscale x 8 x i8> [[BROADCAST_SPLATINSERT]], <vscale x 8 x i8> poison, <vscale x 8 x i32> zeroinitializer
+; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK: [[VECTOR_BODY]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_IND:%.*]] = phi <vscale x 8 x i8> [ [[TMP5]], %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[TMP8:%.*]] = getelementptr i8, ptr [[DST]], i64 [[INDEX]]
+; CHECK-NEXT: store <vscale x 8 x i8> [[VEC_IND]], ptr [[TMP8]], align 1
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], [[TMP1]]
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add nsw <vscale x 8 x i8> [[VEC_IND]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT: [[TMP9:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[TMP9]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK]]:
+; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK: [[SCALAR_PH]]:
+;
+entry:
+ br label %loop
+
+loop:
+ %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+ %iv = phi i8 [ 0, %entry ], [ %iv.next, %loop ]
+ %gep = getelementptr i8, ptr %dst, i64 %i
+ store i8 %iv, ptr %gep
+ %iv.next = add nsw i8 %iv, -1
+ %i.next = add i64 %i, 1
+ %ec = icmp eq i64 %i.next, %n
+ br i1 %ec, label %exit, label %loop
+
+exit:
+ ret void
+}
+
+define void @i8_induction_unknown_max_vscale(ptr noalias %dst, i64 %n) #1 {
+; CHECK-LABEL: define void @i8_induction_unknown_max_vscale(
+; CHECK-SAME: ptr noalias [[DST:%.*]], i64 [[N:%.*]]) #[[ATTR1:[0-9]+]] {
+; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: [[TMP0:%.*]] = call i64 @llvm.vscale.i64()
+; CHECK-NEXT: [[TMP1:%.*]] = shl nuw i64 [[TMP0]], 3
+; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], [[TMP1]]
+; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: [[N_MOD_VF:%.*]] = urem i64 [[N]], [[TMP1]]
+; CHECK-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[N_MOD_VF]]
+; CHECK-NEXT: [[TMP2:%.*]] = trunc i64 [[N_VEC]] to i8
+; CHECK-NEXT: [[TMP3:%.*]] = sub i8 0, [[TMP2]]
+; CHECK-NEXT: [[TMP4:%.*]] = call <vscale x 8 x i8> @llvm.stepvector.nxv8i8()
+; CHECK-NEXT: [[TMP5:%.*]] = sub nsw <vscale x 8 x i8> zeroinitializer, [[TMP4]]
+; CHECK-NEXT: [[TMP6:%.*]] = trunc i64 [[TMP1]] to i8
+; CHECK-NEXT: [[TMP7:%.*]] = sub nsw i8 0, [[TMP6]]
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <vscale x 8 x i8> poison, i8 [[TMP7]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <vscale x 8 x i8> [[BROADCAST_SPLATINSERT]], <vscale x 8 x i8> poison, <vscale x 8 x i32> zeroinitializer
+; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK: [[VECTOR_BODY]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_IND:%.*]] = phi <vscale x 8 x i8> [ [[TMP5]], %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[TMP8:%.*]] = getelementptr i8, ptr [[DST]], i64 [[INDEX]]
+; CHECK-NEXT: store <vscale x 8 x i8> [[VEC_IND]], ptr [[TMP8]], align 1
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], [[TMP1]]
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add nsw <vscale x 8 x i8> [[VEC_IND]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT: [[TMP9:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[TMP9]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK]]:
+; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK: [[SCALAR_PH]]:
+;
+entry:
+ br label %loop
+
+loop:
+ %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+ %iv = phi i8 [ 0, %entry ], [ %iv.next, %loop ]
+ %gep = getelementptr i8, ptr %dst, i64 %i
+ store i8 %iv, ptr %gep
+ %iv.next = add nsw i8 %iv, -1
+ %i.next = add i64 %i, 1
+ %ec = icmp eq i64 %i.next, %n
+ br i1 %ec, label %exit, label %loop
+
+exit:
+ ret void
+}
+
+define void @i8_induction_vf_fits(ptr noalias %dst, i64 %n) #2 {
+; CHECK-LABEL: define void @i8_induction_vf_fits(
+; CHECK-SAME: ptr noalias [[DST:%.*]], i64 [[N:%.*]]) #[[ATTR2:[0-9]+]] {
+; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: [[TMP0:%.*]] = call i64 @llvm.vscale.i64()
+; CHECK-NEXT: [[TMP1:%.*]] = shl nuw i64 [[TMP0]], 3
+; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], [[TMP1]]
+; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: [[N_MOD_VF:%.*]] = urem i64 [[N]], [[TMP1]]
+; CHECK-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[N_MOD_VF]]
+; CHECK-NEXT: [[TMP2:%.*]] = trunc i64 [[N_VEC]] to i8
+; CHECK-NEXT: [[TMP3:%.*]] = sub i8 0, [[TMP2]]
+; CHECK-NEXT: [[TMP4:%.*]] = call <vscale x 8 x i8> @llvm.stepvector.nxv8i8()
+; CHECK-NEXT: [[TMP5:%.*]] = sub nsw <vscale x 8 x i8> zeroinitializer, [[TMP4]]
+; CHECK-NEXT: [[TMP6:%.*]] = trunc i64 [[TMP1]] to i8
+; CHECK-NEXT: [[TMP7:%.*]] = sub nsw i8 0, [[TMP6]]
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <vscale x 8 x i8> poison, i8 [[TMP7]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <vscale x 8 x i8> [[BROADCAST_SPLATINSERT]], <vscale x 8 x i8> poison, <vscale x 8 x i32> zeroinitializer
+; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK: [[VECTOR_BODY]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_IND:%.*]] = phi <vscale x 8 x i8> [ [[TMP5]], %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[TMP8:%.*]] = getelementptr i8, ptr [[DST]], i64 [[INDEX]]
+; CHECK-NEXT: store <vscale x 8 x i8> [[VEC_IND]], ptr [[TMP8]], align 1
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], [[TMP1]]
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add nsw <vscale x 8 x i8> [[VEC_IND]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT: [[TMP9:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[TMP9]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP6:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK]]:
+; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK: [[SCALAR_PH]]:
+;
+entry:
+ br label %loop
+
+loop:
+ %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+ %iv = phi i8 [ 0, %entry ], [ %iv.next, %loop ]
+ %gep = getelementptr i8, ptr %dst, i64 %i
+ store i8 %iv, ptr %gep
+ %iv.next = add nsw i8 %iv, -1
+ %i.next = add i64 %i, 1
+ %ec = icmp eq i64 %i.next, %n
+ br i1 %ec, label %exit, label %loop
+
+exit:
+ ret void
+}
+
+attributes #0 = { "target-features"="+sve" vscale_range(1,16) }
+attributes #1 = { "target-features"="+sve" }
+attributes #2 = { "target-features"="+sve" vscale_range(1,8) }
diff --git a/llvm/test/Transforms/LoopVectorize/induction-wrapflags.ll b/llvm/test/Transforms/LoopVectorize/induction-wrapflags.ll
index fc0f331c039d5..20ba7f8366e30 100644
--- a/llvm/test/Transforms/LoopVectorize/induction-wrapflags.ll
+++ b/llvm/test/Transforms/LoopVectorize/induction-wrapflags.ll
@@ -322,5 +322,195 @@ exit:
ret void
}
+; NSW cannot be retained, as the lane offsets (e.g. 3 * 60) may overflow, even if
+; the scalar induction values do not.
+define i8 @add_induction_nsw_lane_offset_overflow(i8 %start, i8 %n) {
+; CHECK-LABEL: define i8 @add_induction_nsw_lane_offset_overflow(
+; CHECK-SAME: i8 [[START:%.*]], i8 [[N:%.*]]) {
+; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: [[TMP0:%.*]] = add i8 [[N]], -1
+; CHECK-NEXT: [[TMP1:%.*]] = zext i8 [[TMP0]] to i32
+; CHECK-NEXT: [[TMP2:%.*]] = add nuw nsw i32 [[TMP1]], 1
+; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i32 [[TMP2]], 4
+; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: [[TMP3:%.*]] = and i32 [[TMP2]], 3
+; CHECK-NEXT: [[N_VEC:%.*]] = sub i32 [[TMP2]], [[TMP3]]
+; CHECK-NEXT: [[TMP4:%.*]] = trunc i32 [[N_VEC]] to i8
+; CHECK-NEXT: [[TMP5:%.*]] = mul i8 [[TMP4]], 60
+; CHECK-NEXT: [[TMP6:%.*]] = add i8 [[START]], [[TMP5]]
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i8> poison, i8 [[START]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i8> [[BROADCAST_SPLATINSERT]], <4 x i8> poison, <4 x i32> zeroinitializer
+; CHECK-NEXT: [[INDUCTION:%.*]] = add nuw nsw <4 x i8> [[BROADCAST_SPLAT]], <i8 0, i8 60, i8 120, i8 -76>
+; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK: [[VECTOR_BODY]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i32 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_PHI:%.*]] = phi <4 x i8> [ zeroinitializer, %[[VECTOR_PH]] ], [ [[TMP7:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_IND:%.*]] = phi <4 x i8> [ [[INDUCTION]], %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[TMP7]] = add <4 x i8> [[VEC_PHI]], [[VEC_IND]]
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 4
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add nuw nsw <4 x i8> [[VEC_IND]], splat (i8 -16)
+; CHECK-NEXT: [[TMP8:%.*]] = icmp eq i32 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP14:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK]]:
+; CHECK-NEXT: [[TMP9:%.*]] = call i8 @llvm.vector.reduce.add.v4i8(<4 x i8> [[TMP7]])
+; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i32 [[TMP2]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK: [[SCALAR_PH]]:
+;
+entry:
+ br label %loop
+
+loop:
+ %iv = phi i8 [ 0, %entry ], [ %iv.next, %loop ]
+ %red = phi i8 [ 0, %entry ], [ %red.next, %loop ]
+ %f = phi i8 [ %start, %entry ], [ %f.next, %loop ]
+ %red.next = add i8 %red, %f
+ %f.next = add nuw nsw i8 %f, 60
+ %iv.next = add i8 %iv, 1
+ %ec = icmp eq i8 %iv.next, %n
+ br i1 %ec, label %exit, label %loop
+
+exit:
+ ret i8 %red.next
+}
+
+; NSW can be retained, as start and step are both non-positive.
+define i8 @add_induction_nsw_start_and_step_non_positive(i8 %n) {
+; CHECK-LABEL: define i8 @add_induction_nsw_start_and_step_non_positive(
+; CHECK-SAME: i8 [[N:%.*]]) {
+; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: [[TMP0:%.*]] = add i8 [[N]], -1
+; CHECK-NEXT: [[TMP1:%.*]] = zext i8 [[TMP0]] to i32
+; CHECK-NEXT: [[TMP2:%.*]] = add nuw nsw i32 [[TMP1]], 1
+; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i32 [[TMP2]], 4
+; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: [[TMP3:%.*]] = and i32 [[TMP2]], 3
+; CHECK-NEXT: [[N_VEC:%.*]] = sub i32 [[TMP2]], [[TMP3]]
+; CHECK-NEXT: [[TMP4:%.*]] = trunc i32 [[N_VEC]] to i8
+; CHECK-NEXT: [[TMP5:%.*]] = mul i8 [[TMP4]], -3
+; CHECK-NEXT: [[TMP6:%.*]] = add i8 -1, [[TMP5]]
+; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK: [[VECTOR_BODY]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i32 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_PHI:%.*]] = phi <4 x i8> [ zeroinitializer, %[[VECTOR_PH]] ], [ [[TMP7:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_IND:%.*]] = phi <4 x i8> [ <i8 -1, i8 -4, i8 -7, i8 -10>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[TMP7]] = add <4 x i8> [[VEC_PHI]], [[VEC_IND]]
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 4
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add nsw <4 x i8> [[VEC_IND]], splat (i8 -12)
+; CHECK-NEXT: [[TMP8:%.*]] = icmp eq i32 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP16:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK]]:
+; CHECK-NEXT: [[TMP9:%.*]] = call i8 @llvm.vector.reduce.add.v4i8(<4 x i8> [[TMP7]])
+; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i32 [[TMP2]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK: [[SCALAR_PH]]:
+;
+entry:
+ br label %loop
+
+loop:
+ %iv = phi i8 [ 0, %entry ], [ %iv.next, %loop ]
+ %red = phi i8 [ 0, %entry ], [ %red.next, %loop ]
+ %f = phi i8 [ -1, %entry ], [ %f.next, %loop ]
+ %red.next = add i8 %red, %f
+ %f.next = add nsw i8 %f, -3
+ %iv.next = add i8 %iv, 1
+ %ec = icmp eq i8 %iv.next, %n
+ br i1 %ec, label %exit, label %loop
+
+exit:
+ ret i8 %red.next
+}
+
+; The increment does not update %iv directly, so its nsw does not bound the
+; induction.
+define void @sub_induction_nsw_increment_not_on_phi(ptr noalias %dst, i32 range(i32 0, 8) %b) {
+; CHECK-LABEL: define void @sub_induction_nsw_increment_not_on_phi(
+; CHECK-SAME: ptr noalias [[DST:%.*]], i32 range(i32 0, 8) [[B:%.*]]) {
+; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: [[TMP0:%.*]] = sub nsw i32 -5, [[B]]
+; CHECK-NEXT: br label %[[VECTOR_PH:.*]]
+; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[TMP0]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-NEXT: [[TMP1:%.*]] = mul <4 x i32> <i32 0, i32 1, i32 2, i32 3>, [[BROADCAST_SPLAT]]
+; CHECK-NEXT: [[INDUCTION:%.*]] = add <4 x i32> splat (i32 -2147483638), [[TMP1]]
+; CHECK-NEXT: [[TMP2:%.*]] = shl i32 [[TMP0]], 2
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT1:%.*]] = insertelement <4 x i32> poison, i32 [[TMP2]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT2:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT1]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK: [[VECTOR_BODY]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_IND:%.*]] = phi <4 x i32> [ [[INDUCTION]], %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[TMP3:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-NEXT: store <4 x i32> [[VEC_IND]], ptr [[TMP3]], align 4
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <4 x i32> [[VEC_IND]], [[BROADCAST_SPLAT2]]
+; CHECK-NEXT: [[TMP4:%.*]] = icmp eq i64 [[INDEX_NEXT]], 8
+; CHECK-NEXT: br i1 [[TMP4]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP18:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK]]:
+; CHECK-NEXT: br label %[[EXIT:.*]]
+; CHECK: [[EXIT]]:
+; CHECK-NEXT: ret void
+;
+entry:
+ br label %loop
+
+loop:
+ %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+ %iv = phi i32 [ -2147483638, %entry ], [ %iv.next, %loop ]
+ %x = add i32 %iv, -5
+ %iv.next = sub nsw i32 %x, %b
+ %gep = getelementptr i32, ptr %dst, i64 %i
+ store i32 %iv, ptr %gep
+ %i.next = add i64 %i, 1
+ %ec = icmp eq i64 %i.next, 8
+ br i1 %ec, label %exit, label %loop
+
+exit:
+ ret void
+}
+
+define void @add_induction_nuw_nsw_increment_not_on_phi(ptr noalias %dst) {
+; CHECK-LABEL: define void @add_induction_nuw_nsw_increment_not_on_phi(
+; CHECK-SAME: ptr noalias [[DST:%.*]]) {
+; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: br label %[[VECTOR_PH:.*]]
+; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK: [[VECTOR_BODY]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_IND:%.*]] = phi <4 x i32> [ <i32 2147483637, i32 2147483643, i32 -2147483647, i32 -2147483641>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[TMP0:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-NEXT: store <4 x i32> [[VEC_IND]], ptr [[TMP0]], align 4
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add nuw nsw <4 x i32> [[VEC_IND]], splat (i32 24)
+; CHECK-NEXT: [[TMP1:%.*]] = icmp eq i64 [[INDEX_NEXT]], 16
+; CHECK-NEXT: br i1 [[TMP1]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP19:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK]]:
+; CHECK-NEXT: br label %[[EXIT:.*]]
+; CHECK: [[EXIT]]:
+; CHECK-NEXT: ret void
+;
+entry:
+ br label %loop
+
+loop:
+ %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+ %iv = phi i32 [ 2147483637, %entry ], [ %iv.next, %loop ]
+ %x = add i32 %iv, 5
+ %iv.next = add nuw nsw i32 %x, 1
+ %gep = getelementptr i32, ptr %dst, i64 %i
+ store i32 %iv, ptr %gep
+ %i.next = add i64 %i, 1
+ %ec = icmp eq i64 %i.next, 16
+ br i1 %ec, label %exit, label %loop
+
+exit:
+ ret void
+}
+
!0 = distinct !{!0, !1}
!1 = !{!"llvm.loop.interleave.count", i32 2}
>From 4eadb0fecaf5dbb55acf14bbdfee8e8bd1c02ddc Mon Sep 17 00:00:00 2001
From: David Green <david.green at arm.com>
Date: Sat, 26 Sep 2026 18:17:58 +0100
Subject: [PATCH 20/53] [AArch64] Reorganise perfect shuffle generation. NFC
(#224526)
This adds a generatePerfectShuffle implementation for parsing through
the perfect shuffle tables, generating a list of ShuffleEntry's that
represent the sequence of shuffles that need to be performed. This is
intended to be a NFC as-is, allowing it to be reused in global isel and
extended in the future to handle shuffles that are not part of the
shuffle table. The number of instructions generated can also be used for
costing shuffles, as we do for immediate generation.
---
.../Target/AArch64/AArch64ISelLowering.cpp | 321 ++++++++----------
.../Target/AArch64/AArch64PerfectShuffle.h | 119 +++++++
2 files changed, 259 insertions(+), 181 deletions(-)
diff --git a/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp b/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
index d61304154cad7..c98d607c4a22c 100644
--- a/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
+++ b/llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
@@ -15669,174 +15669,6 @@ static SDValue tryFormConcatFromShuffle(SDValue Op, SelectionDAG &DAG) {
return DAG.getNode(ISD::CONCAT_VECTORS, DL, VT, V0, V1);
}
-/// GeneratePerfectShuffle - Given an entry in the perfect-shuffle table, emit
-/// the specified operations to build the shuffle. ID is the perfect-shuffle
-//ID, V1 and V2 are the original shuffle inputs. PFEntry is the Perfect shuffle
-//table entry and LHS/RHS are the immediate inputs for this stage of the
-//shuffle.
-static SDValue GeneratePerfectShuffle(unsigned ID, SDValue V1, SDValue V2,
- unsigned PFEntry, SDValue LHS,
- SDValue RHS, SelectionDAG &DAG,
- const SDLoc &DL) {
- unsigned OpNum = (PFEntry >> 26) & 0x0F;
- unsigned LHSID = (PFEntry >> 13) & ((1 << 13) - 1);
- unsigned RHSID = (PFEntry >> 0) & ((1 << 13) - 1);
-
- enum {
- OP_COPY = 0, // Copy, used for things like <u,u,u,3> to say it is <0,1,2,3>
- OP_VREV,
- OP_VDUP0,
- OP_VDUP1,
- OP_VDUP2,
- OP_VDUP3,
- OP_VEXT1,
- OP_VEXT2,
- OP_VEXT3,
- OP_VUZPL, // VUZP, left result
- OP_VUZPR, // VUZP, right result
- OP_VZIPL, // VZIP, left result
- OP_VZIPR, // VZIP, right result
- OP_VTRNL, // VTRN, left result
- OP_VTRNR, // VTRN, right result
- OP_MOVLANE // Move lane. RHSID is the lane to move into
- };
-
- if (OpNum == OP_COPY) {
- if (LHSID == (1 * 9 + 2) * 9 + 3)
- return LHS;
- assert(LHSID == ((4 * 9 + 5) * 9 + 6) * 9 + 7 && "Illegal OP_COPY!");
- return RHS;
- }
-
- if (OpNum == OP_MOVLANE) {
- // Decompose a PerfectShuffle ID to get the Mask for lane Elt
- auto getPFIDLane = [](unsigned ID, int Elt) -> int {
- assert(Elt < 4 && "Expected Perfect Lanes to be less than 4");
- Elt = 3 - Elt;
- while (Elt > 0) {
- ID /= 9;
- Elt--;
- }
- return (ID % 9 == 8) ? -1 : ID % 9;
- };
-
- // For OP_MOVLANE shuffles, the RHSID represents the lane to move into. We
- // get the lane to move from the PFID, which is always from the
- // original vectors (V1 or V2).
- SDValue OpLHS = GeneratePerfectShuffle(
- LHSID, V1, V2, PerfectShuffleTable[LHSID], LHS, RHS, DAG, DL);
- EVT VT = OpLHS.getValueType();
- assert(RHSID < 8 && "Expected a lane index for RHSID!");
- unsigned ExtLane = 0;
- SDValue Input;
-
- // OP_MOVLANE are either D movs (if bit 0x4 is set) or S movs. D movs
- // convert into a higher type.
- if (RHSID & 0x4) {
- int MaskElt = getPFIDLane(ID, (RHSID & 0x01) << 1) >> 1;
- if (MaskElt == -1)
- MaskElt = (getPFIDLane(ID, ((RHSID & 0x01) << 1) + 1) - 1) >> 1;
- assert(MaskElt >= 0 && "Didn't expect an undef movlane index!");
- ExtLane = MaskElt < 2 ? MaskElt : (MaskElt - 2);
- Input = MaskElt < 2 ? V1 : V2;
- if (VT.getScalarSizeInBits() == 16) {
- Input = DAG.getBitcast(MVT::v2f32, Input);
- OpLHS = DAG.getBitcast(MVT::v2f32, OpLHS);
- } else {
- assert(VT.getScalarSizeInBits() == 32 &&
- "Expected 16 or 32 bit shuffle elements");
- Input = DAG.getBitcast(MVT::v2f64, Input);
- OpLHS = DAG.getBitcast(MVT::v2f64, OpLHS);
- }
- } else {
- int MaskElt = getPFIDLane(ID, RHSID);
- assert(MaskElt >= 0 && "Didn't expect an undef movlane index!");
- ExtLane = MaskElt < 4 ? MaskElt : (MaskElt - 4);
- Input = MaskElt < 4 ? V1 : V2;
- // Be careful about creating illegal types. Use f16 instead of i16.
- if (VT == MVT::v4i16) {
- Input = DAG.getBitcast(MVT::v4f16, Input);
- OpLHS = DAG.getBitcast(MVT::v4f16, OpLHS);
- }
- }
- SDValue Ext = DAG.getExtractVectorElt(
- DL, Input.getValueType().getVectorElementType(), Input, ExtLane);
- SDValue Ins = DAG.getInsertVectorElt(DL, OpLHS, Ext, RHSID & 0x3);
- return DAG.getBitcast(VT, Ins);
- }
-
- SDValue OpLHS, OpRHS;
- OpLHS = GeneratePerfectShuffle(LHSID, V1, V2, PerfectShuffleTable[LHSID], LHS,
- RHS, DAG, DL);
- OpRHS = GeneratePerfectShuffle(RHSID, V1, V2, PerfectShuffleTable[RHSID], LHS,
- RHS, DAG, DL);
- EVT VT = OpLHS.getValueType();
-
- switch (OpNum) {
- default:
- llvm_unreachable("Unknown shuffle opcode!");
- case OP_VREV: {
- // VREV divides the vector in half and swaps within the half.
- if (VT.getVectorElementType() == MVT::i32 ||
- VT.getVectorElementType() == MVT::f32)
- return DAG.getNode(AArch64ISD::REV64, DL, VT, OpLHS);
- // vrev <4 x i16> -> REV32
- if (VT.getVectorElementType() == MVT::i16 ||
- VT.getVectorElementType() == MVT::f16 ||
- VT.getVectorElementType() == MVT::bf16)
- return DAG.getNode(AArch64ISD::REV32, DL, VT, OpLHS);
- // vrev <4 x i8> -> BSWAP which is REV16
- assert(VT == MVT::v8i8 || VT == MVT::v16i8);
- EVT BSVT = VT == MVT::v8i8 ? MVT::v4i16 : MVT::v8i16;
- return DAG.getNode(
- AArch64ISD::NVCAST, DL, VT,
- DAG.getNode(ISD::BSWAP, DL, BSVT,
- DAG.getNode(AArch64ISD::NVCAST, DL, BSVT, OpLHS)));
- }
- case OP_VDUP0:
- case OP_VDUP1:
- case OP_VDUP2:
- case OP_VDUP3: {
- EVT EltTy = VT.getVectorElementType();
- unsigned Opcode;
- if (EltTy == MVT::i8)
- Opcode = AArch64ISD::DUPLANE8;
- else if (EltTy == MVT::i16 || EltTy == MVT::f16 || EltTy == MVT::bf16)
- Opcode = AArch64ISD::DUPLANE16;
- else if (EltTy == MVT::i32 || EltTy == MVT::f32)
- Opcode = AArch64ISD::DUPLANE32;
- else if (EltTy == MVT::i64 || EltTy == MVT::f64)
- Opcode = AArch64ISD::DUPLANE64;
- else
- llvm_unreachable("Invalid vector element type?");
-
- if (VT.getSizeInBits() == 64)
- OpLHS = WidenVector(OpLHS, DAG);
- SDValue Lane = DAG.getConstant(OpNum - OP_VDUP0, DL, MVT::i64);
- return DAG.getNode(Opcode, DL, VT, OpLHS, Lane);
- }
- case OP_VEXT1:
- case OP_VEXT2:
- case OP_VEXT3: {
- unsigned Imm = (OpNum - OP_VEXT1 + 1) * getExtFactor(OpLHS);
- return DAG.getNode(AArch64ISD::EXT, DL, VT, OpLHS, OpRHS,
- DAG.getConstant(Imm, DL, MVT::i32));
- }
- case OP_VUZPL:
- return DAG.getNode(AArch64ISD::UZP1, DL, VT, OpLHS, OpRHS);
- case OP_VUZPR:
- return DAG.getNode(AArch64ISD::UZP2, DL, VT, OpLHS, OpRHS);
- case OP_VZIPL:
- return DAG.getNode(AArch64ISD::ZIP1, DL, VT, OpLHS, OpRHS);
- case OP_VZIPR:
- return DAG.getNode(AArch64ISD::ZIP2, DL, VT, OpLHS, OpRHS);
- case OP_VTRNL:
- return DAG.getNode(AArch64ISD::TRN1, DL, VT, OpLHS, OpRHS);
- case OP_VTRNR:
- return DAG.getNode(AArch64ISD::TRN2, DL, VT, OpLHS, OpRHS);
- }
-}
-
static SDValue GenerateTBL(SDValue Op, ArrayRef<int> ShuffleMask,
SelectionDAG &DAG) {
// Check to see if we can use the TBL instruction.
@@ -16348,20 +16180,147 @@ SDValue AArch64TargetLowering::LowerVECTOR_SHUFFLE(SDValue Op,
// If the shuffle is not directly supported and it has 4 elements, use
// the PerfectShuffle-generated table to synthesize it from other shuffles.
if (NumElts == 4) {
- unsigned PFIndexes[4];
- for (unsigned i = 0; i != 4; ++i) {
- if (ShuffleMask[i] < 0)
- PFIndexes[i] = 8;
- else
- PFIndexes[i] = ShuffleMask[i];
+ SmallVector<ShuffleEntry> Entries;
+ if (generatePerfectShuffle(ShuffleMask, NumElts, Entries)) {
+ SmallVector<SDValue> Vals;
+ auto getValue = [](unsigned Idx, SDValue LHS, SDValue RHS,
+ SmallVector<SDValue> &Vals) {
+ if (Idx == ShuffleEntry::LHS)
+ return LHS;
+ if (Idx == ShuffleEntry::RHS)
+ return RHS;
+ assert(Idx < Vals.size());
+ return Vals[Idx];
+ };
+ for (const ShuffleEntry &Entry : Entries) {
+ SDValue OpLHS = getValue(Entry.LHSID, V1, V2, Vals);
+
+ switch (Entry.Op) {
+ case ShuffleEntry::OP_COPY:
+ case ShuffleEntry::OP_MOVLANE:
+ llvm_unreachable("Did not expect a OP_COPY or OP_MOVLANE");
+ case ShuffleEntry::OP_MOVLANE64: {
+ unsigned ExtLane = (Entry.RHSID >> 8) & 0xff;
+ unsigned ToLane = Entry.RHSID & 0xff;
+ bool Input2 = Entry.RHSID >> 16;
+
+ SDValue Input = Input2 ? V2 : V1;
+ if (VT.getScalarSizeInBits() == 16) {
+ Input = DAG.getBitcast(MVT::v2f32, Input);
+ OpLHS = DAG.getBitcast(MVT::v2f32, OpLHS);
+ } else {
+ assert(VT.getScalarSizeInBits() == 32 &&
+ "Expected 16 or 32 bit shuffle elements");
+ Input = DAG.getBitcast(MVT::v2f64, Input);
+ OpLHS = DAG.getBitcast(MVT::v2f64, OpLHS);
+ }
+ SDValue Ext = DAG.getExtractVectorElt(
+ DL, Input.getValueType().getVectorElementType(), Input, ExtLane);
+ SDValue Ins = DAG.getInsertVectorElt(DL, OpLHS, Ext, ToLane);
+ Vals.push_back(DAG.getBitcast(VT, Ins));
+ break;
+ }
+ case ShuffleEntry::OP_MOVLANE32: {
+ unsigned ExtLane = (Entry.RHSID >> 8) & 0xff;
+ unsigned ToLane = Entry.RHSID & 0xff;
+ bool Input2 = Entry.RHSID >> 16;
+
+ SDValue Input = Input2 ? V2 : V1;
+ // Be careful about creating illegal types. Use f16 instead of i16.
+ if (VT == MVT::v4i16) {
+ Input = DAG.getBitcast(MVT::v4f16, Input);
+ OpLHS = DAG.getBitcast(MVT::v4f16, OpLHS);
+ }
+ SDValue Ext = DAG.getExtractVectorElt(
+ DL, Input.getValueType().getVectorElementType(), Input, ExtLane);
+ SDValue Ins = DAG.getInsertVectorElt(DL, OpLHS, Ext, ToLane);
+ Vals.push_back(DAG.getBitcast(VT, Ins));
+ break;
+ }
+ case ShuffleEntry::OP_VREV: {
+ // VREV divides the vector in half and swaps within the half.
+ if (VT.getVectorElementType() == MVT::i32 ||
+ VT.getVectorElementType() == MVT::f32)
+ Vals.push_back(DAG.getNode(AArch64ISD::REV64, DL, VT, OpLHS));
+ // vrev <4 x i16> -> REV32
+ else if (VT.getVectorElementType() == MVT::i16 ||
+ VT.getVectorElementType() == MVT::f16 ||
+ VT.getVectorElementType() == MVT::bf16)
+ Vals.push_back(DAG.getNode(AArch64ISD::REV32, DL, VT, OpLHS));
+ else {
+ // vrev <4 x i8> -> BSWAP which is REV16
+ assert(VT == MVT::v8i8 || VT == MVT::v16i8);
+ EVT BSVT = VT == MVT::v8i8 ? MVT::v4i16 : MVT::v8i16;
+ Vals.push_back(DAG.getNode(
+ AArch64ISD::NVCAST, DL, VT,
+ DAG.getNode(ISD::BSWAP, DL, BSVT,
+ DAG.getNode(AArch64ISD::NVCAST, DL, BSVT, OpLHS))));
+ }
+ break;
+ }
+ case ShuffleEntry::OP_VDUP0:
+ case ShuffleEntry::OP_VDUP1:
+ case ShuffleEntry::OP_VDUP2:
+ case ShuffleEntry::OP_VDUP3: {
+ EVT EltTy = VT.getVectorElementType();
+ unsigned Opcode;
+ if (EltTy == MVT::i8)
+ Opcode = AArch64ISD::DUPLANE8;
+ else if (EltTy == MVT::i16 || EltTy == MVT::f16 || EltTy == MVT::bf16)
+ Opcode = AArch64ISD::DUPLANE16;
+ else if (EltTy == MVT::i32 || EltTy == MVT::f32)
+ Opcode = AArch64ISD::DUPLANE32;
+ else if (EltTy == MVT::i64 || EltTy == MVT::f64)
+ Opcode = AArch64ISD::DUPLANE64;
+ else
+ llvm_unreachable("Invalid vector element type?");
+
+ if (VT.getSizeInBits() == 64)
+ OpLHS = WidenVector(OpLHS, DAG);
+ SDValue Lane =
+ DAG.getConstant(Entry.Op - ShuffleEntry::OP_VDUP0, DL, MVT::i64);
+ Vals.push_back(DAG.getNode(Opcode, DL, VT, OpLHS, Lane));
+ break;
+ }
+ case ShuffleEntry::OP_VEXT1:
+ case ShuffleEntry::OP_VEXT2:
+ case ShuffleEntry::OP_VEXT3: {
+ unsigned Imm =
+ (Entry.Op - ShuffleEntry::OP_VEXT1 + 1) * getExtFactor(OpLHS);
+ Vals.push_back(DAG.getNode(AArch64ISD::EXT, DL, VT, OpLHS,
+ getValue(Entry.RHSID, V1, V2, Vals),
+ DAG.getConstant(Imm, DL, MVT::i32)));
+ break;
+ }
+ case ShuffleEntry::OP_VUZPL:
+ Vals.push_back(DAG.getNode(AArch64ISD::UZP1, DL, VT, OpLHS,
+ getValue(Entry.RHSID, V1, V2, Vals)));
+ break;
+ case ShuffleEntry::OP_VUZPR:
+ Vals.push_back(DAG.getNode(AArch64ISD::UZP2, DL, VT, OpLHS,
+ getValue(Entry.RHSID, V1, V2, Vals)));
+ break;
+ case ShuffleEntry::OP_VZIPL:
+ Vals.push_back(DAG.getNode(AArch64ISD::ZIP1, DL, VT, OpLHS,
+ getValue(Entry.RHSID, V1, V2, Vals)));
+ break;
+ case ShuffleEntry::OP_VZIPR:
+ Vals.push_back(DAG.getNode(AArch64ISD::ZIP2, DL, VT, OpLHS,
+ getValue(Entry.RHSID, V1, V2, Vals)));
+ break;
+ case ShuffleEntry::OP_VTRNL:
+ Vals.push_back(DAG.getNode(AArch64ISD::TRN1, DL, VT, OpLHS,
+ getValue(Entry.RHSID, V1, V2, Vals)));
+ break;
+ case ShuffleEntry::OP_VTRNR:
+ Vals.push_back(DAG.getNode(AArch64ISD::TRN2, DL, VT, OpLHS,
+ getValue(Entry.RHSID, V1, V2, Vals)));
+ break;
+ }
+ }
+ assert(Vals.size() == Entries.size());
+ return Vals.back();
}
-
- // Compute the index in the perfect shuffle table.
- unsigned PFTableIndex = PFIndexes[0] * 9 * 9 * 9 + PFIndexes[1] * 9 * 9 +
- PFIndexes[2] * 9 + PFIndexes[3];
- unsigned PFEntry = PerfectShuffleTable[PFTableIndex];
- return GeneratePerfectShuffle(PFTableIndex, V1, V2, PFEntry, V1, V2, DAG,
- DL);
}
// Check for a "select shuffle", generating a BSL to pick between lanes in
diff --git a/llvm/lib/Target/AArch64/AArch64PerfectShuffle.h b/llvm/lib/Target/AArch64/AArch64PerfectShuffle.h
index 00d9edbbc02ea..3f9c9558d04b3 100644
--- a/llvm/lib/Target/AArch64/AArch64PerfectShuffle.h
+++ b/llvm/lib/Target/AArch64/AArch64PerfectShuffle.h
@@ -305,6 +305,125 @@ inline bool isDUPFirstSegmentMask(ArrayRef<int> Mask, unsigned Segments,
});
}
+/// ShuffleEntry - Represent a shuffle entry in the decomposion of a vector
+/// shuffle. i.e. a vector shuffle LHS, RHS, Mask can be built using a list of
+/// ShuffleEntry, by performing Op to either LHS, RHS or one of the previous
+/// ShuffleEntrys in the lists.
+struct ShuffleEntry {
+ // The supported operations. The first 16 match those generated by
+ // PerfectShuffle.cpp.
+ enum Operation {
+ OP_COPY = 0, // Copy, used for things like <u,u,u,3> to say it is <0,1,2,3>
+ OP_VREV,
+ OP_VDUP0,
+ OP_VDUP1,
+ OP_VDUP2,
+ OP_VDUP3,
+ OP_VEXT1,
+ OP_VEXT2,
+ OP_VEXT3,
+ OP_VUZPL, // VUZP, left result
+ OP_VUZPR, // VUZP, right result
+ OP_VZIPL, // VZIP, left result
+ OP_VZIPR, // VZIP, right result
+ OP_VTRNL, // VTRN, left result
+ OP_VTRNR, // VTRN, right result
+ OP_MOVLANE, // Move lane. RHSID is the lane to move into
+
+ OP_MOVLANE32,
+ OP_MOVLANE64,
+ };
+
+ /// Special IDs for the LHS and RHS values.
+ enum IDs {
+ LHS = 0xfe,
+ RHS = 0xff,
+ };
+
+ Operation Op;
+ unsigned LHSID;
+ unsigned RHSID;
+};
+
+inline unsigned
+generatePerfectShuffleFromTable(unsigned PFTableIndex,
+ SmallVector<ShuffleEntry> &Entries) {
+ unsigned PFEntry = PerfectShuffleTable[PFTableIndex];
+ ShuffleEntry::Operation OpNum =
+ (ShuffleEntry::Operation)((PFEntry >> 26) & 0x0F);
+ unsigned LHSID = (PFEntry >> 13) & ((1 << 13) - 1);
+ unsigned RHSID = (PFEntry >> 0) & ((1 << 13) - 1);
+
+ if (OpNum == ShuffleEntry::OP_COPY) {
+ if (LHSID == (1 * 9 + 2) * 9 + 3)
+ return ShuffleEntry::LHS;
+ assert(LHSID == ((4 * 9 + 5) * 9 + 6) * 9 + 7 && "Illegal OP_COPY!");
+ return ShuffleEntry::RHS;
+ }
+
+ unsigned OpLHS = generatePerfectShuffleFromTable(LHSID, Entries);
+
+ if (OpNum == ShuffleEntry::OP_MOVLANE) {
+ // Decompose a PerfectShuffle ID to get the Mask for lane Elt
+ auto getPFIDLane = [](unsigned ID, int Elt) -> int {
+ assert(Elt < 4 && "Expected Perfect Lanes to be less than 4");
+ Elt = 3 - Elt;
+ while (Elt > 0) {
+ ID /= 9;
+ Elt--;
+ }
+ return (ID % 9 == 8) ? -1 : ID % 9;
+ };
+
+ // OP_MOVLANE are either D movs (if bit 0x4 is set) or S movs. D movs
+ // convert into a higher type.
+ if (RHSID & 0x4) {
+ int MaskElt = getPFIDLane(PFTableIndex, (RHSID & 0x01) << 1) >> 1;
+ if (MaskElt == -1)
+ MaskElt =
+ (getPFIDLane(PFTableIndex, ((RHSID & 0x01) << 1) + 1) - 1) >> 1;
+ assert(MaskElt >= 0 && "Didn't expect an undef movlane index!");
+ unsigned ExtLane = MaskElt < 2 ? MaskElt : (MaskElt - 2);
+ Entries.push_back({ShuffleEntry::OP_MOVLANE64, OpLHS,
+ (MaskElt >= 2) << 16 | ExtLane << 8 | (RHSID & 0x3)});
+ } else {
+ int MaskElt = getPFIDLane(PFTableIndex, RHSID);
+ assert(MaskElt >= 0 && "Didn't expect an undef movlane index!");
+ unsigned ExtLane = MaskElt < 4 ? MaskElt : (MaskElt - 4);
+ Entries.push_back({ShuffleEntry::OP_MOVLANE32, OpLHS,
+ (MaskElt >= 4) << 16 | ExtLane << 8 | (RHSID & 0x3)});
+ }
+ return Entries.size() - 1;
+ }
+
+ unsigned OpRHS = generatePerfectShuffleFromTable(RHSID, Entries);
+
+ Entries.push_back({OpNum, OpLHS, OpRHS});
+ return Entries.size() - 1;
+}
+
+/// generatePerfectShuffle - Given a Mask, attempt to generate the optimal
+/// sequence of instructions using zip/uzp/trn/dup/etc. Currently uses the
+/// perfect shuffle tables.
+inline bool generatePerfectShuffle(ArrayRef<int> Mask, unsigned NumElts,
+ SmallVector<ShuffleEntry> &Entries) {
+ assert(NumElts == 4 && "Only 4 element masks supported at the moment");
+
+ // Compute the index in the perfect shuffle table.
+ unsigned PFIndexes[4];
+ for (unsigned i = 0; i != 4; ++i) {
+ if (Mask[i] < 0)
+ PFIndexes[i] = 8;
+ else
+ PFIndexes[i] = Mask[i];
+ }
+
+ unsigned PFTableIndex = PFIndexes[0] * 9 * 9 * 9 + PFIndexes[1] * 9 * 9 +
+ PFIndexes[2] * 9 + PFIndexes[3];
+ unsigned Idx = generatePerfectShuffleFromTable(PFTableIndex, Entries);
+ return Idx == Entries.size() - 1;
+}
+
} // namespace llvm
#endif
>From 5b8821ed663a8b2afbbc9808593d61ddd27f6076 Mon Sep 17 00:00:00 2001
From: "Michael G. Kazakov" <mike.kazakov at gmail.com>
Date: Sat, 26 Sep 2026 18:52:34 +0100
Subject: [PATCH 21/53] [libc++][pstl] Add more benchmarks of the parallel
algorithms (#225908)
This PR adds benchmarks of these 3 parallel algorithms:
- `std::find(policy, ...)`
- `std::sort(policy, ...)`
- `std::transform_reduce(policy, ...)`
---------
Co-authored-by: Louis Dionne <ldionne.2 at gmail.com>
---
.../nonmodifying/pstl.find.bench.cpp | 62 ++++++++++++++
.../algorithms/sorting/pstl.sort.bench.cpp | 79 ++++++++++++++++++
.../numeric/pstl.transform_reduce.bench.cpp | 82 +++++++++++++++++++
3 files changed, 223 insertions(+)
create mode 100644 libcxx/test/benchmarks/algorithms/nonmodifying/pstl.find.bench.cpp
create mode 100644 libcxx/test/benchmarks/algorithms/sorting/pstl.sort.bench.cpp
create mode 100644 libcxx/test/benchmarks/numeric/pstl.transform_reduce.bench.cpp
diff --git a/libcxx/test/benchmarks/algorithms/nonmodifying/pstl.find.bench.cpp b/libcxx/test/benchmarks/algorithms/nonmodifying/pstl.find.bench.cpp
new file mode 100644
index 0000000000000..98642ad7aa87e
--- /dev/null
+++ b/libcxx/test/benchmarks/algorithms/nonmodifying/pstl.find.bench.cpp
@@ -0,0 +1,62 @@
+//===----------------------------------------------------------------------===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+// REQUIRES: std-at-least-c++20
+
+// UNSUPPORTED: libcpp-has-no-incomplete-pstl
+
+#include <algorithm>
+#include <cstddef>
+#include <cstdint>
+#include <string>
+#include <vector>
+#include <execution>
+#include <utility>
+
+#include <benchmark/benchmark.h>
+#include "../../GenerateInput.h"
+
+int main(int argc, char** argv) {
+ auto bm = [](std::string name, auto&& policy, bool has_needle) {
+ benchmark::RegisterBenchmark(
+ name,
+ [&policy, has_needle](auto& st) mutable {
+ std::size_t size = st.range(0);
+ double x = Generate<double>::random();
+ double y = random_different_from({x});
+ std::vector<double> c(size, x);
+
+ if (has_needle) {
+ // put the element we're searching for at 25% of the sequence
+ *std::next(c.begin(), size / 4) = y;
+ }
+
+ for ([[maybe_unused]] auto _ : st) {
+ benchmark::DoNotOptimize(c);
+ benchmark::DoNotOptimize(y);
+ auto result = std::find(policy, c.begin(), c.end(), y);
+ benchmark::DoNotOptimize(result);
+ }
+ })
+ ->Arg(1 << 6) // 64
+ ->Arg(1 << 16) // 65'536
+ ->Arg(1 << 26) // 67'108'864
+ ->UseRealTime();
+ };
+#if defined(TEST_PSTL_ENABLE_SEQ_BASELINES)
+ bm("std::find(std::execution::seq, vector<double>) (bail 25%)", std::execution::seq, true);
+ bm("std::find(std::execution::seq, vector<double>) (process all)", std::execution::seq, false);
+#endif
+ bm("std::find(std::execution::par, vector<double>) (bail 25%)", std::execution::par, true);
+ bm("std::find(std::execution::par, vector<double>) (process all)", std::execution::par, false);
+
+ benchmark::Initialize(&argc, argv);
+ benchmark::RunSpecifiedBenchmarks();
+ benchmark::Shutdown();
+ return 0;
+}
diff --git a/libcxx/test/benchmarks/algorithms/sorting/pstl.sort.bench.cpp b/libcxx/test/benchmarks/algorithms/sorting/pstl.sort.bench.cpp
new file mode 100644
index 0000000000000..e080deb5fa722
--- /dev/null
+++ b/libcxx/test/benchmarks/algorithms/sorting/pstl.sort.bench.cpp
@@ -0,0 +1,79 @@
+//===----------------------------------------------------------------------===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+// REQUIRES: std-at-least-c++20
+
+// UNSUPPORTED: libcpp-has-no-incomplete-pstl
+
+#include <algorithm>
+#include <array>
+#include <cstddef>
+#include <cstdint>
+#include <string>
+#include <vector>
+#include <execution>
+#include <utility>
+
+#include <benchmark/benchmark.h>
+
+#include "common.h"
+#include "test_macros.h"
+
+int main(int argc, char** argv) {
+ auto bm = [](std::string name, auto&& policy, auto generate_data) {
+ benchmark::RegisterBenchmark(
+ name,
+ [&policy, generate_data](auto& st) mutable {
+ constexpr std::size_t BatchSize = 32;
+ std::size_t size = st.range(0);
+ std::vector<int> data = generate_data(size);
+ std::array<std::vector<int>, BatchSize> c;
+ std::fill_n(c.begin(), BatchSize, data);
+
+ while (st.KeepRunningBatch(BatchSize)) {
+ for (std::size_t i = 0; i != BatchSize; ++i) {
+ benchmark::DoNotOptimize(c[i]);
+ std::sort(policy, c[i].begin(), c[i].end());
+ benchmark::DoNotOptimize(c[i]);
+ }
+
+ // Reset c to its original unsorted state
+ st.PauseTiming();
+ for (std::size_t i = 0; i != BatchSize; ++i) {
+ std::copy(data.begin(), data.end(), c[i].begin());
+ }
+ st.ResumeTiming();
+ }
+ })
+ ->Arg(1 << 6) // 64
+ ->Arg(1 << 16) // 65'536
+ ->Arg(1 << 26) // 67'108'864
+ ->UseRealTime();
+ };
+
+ auto register_bm = [&](auto generate, std::string variant) {
+ auto name = [variant](std::string op) { return op + " (" + variant + ")"; };
+#if defined(TEST_PSTL_ENABLE_SEQ_BASELINES)
+ bm.operator()(name("std::sort(std::execution::seq, vector<int>)"), std::execution::seq, generate);
+#endif
+ bm.operator()(name("std::sort(std::execution::par, vector<int>)"), std::execution::par, generate);
+ };
+
+ register_bm(support::quicksort_adversarial_data<int>, "qsort adversarial");
+ register_bm(support::ascending_sorted_data<int>, "ascending");
+ register_bm(support::descending_sorted_data<int>, "descending");
+ register_bm(support::pipe_organ_data<int>, "pipe-organ");
+ register_bm(support::heap_data<int>, "heap");
+ register_bm(support::shuffled_data<int>, "shuffled");
+ register_bm(support::single_element_data<int>, "repeated");
+
+ benchmark::Initialize(&argc, argv);
+ benchmark::RunSpecifiedBenchmarks();
+ benchmark::Shutdown();
+ return 0;
+}
diff --git a/libcxx/test/benchmarks/numeric/pstl.transform_reduce.bench.cpp b/libcxx/test/benchmarks/numeric/pstl.transform_reduce.bench.cpp
new file mode 100644
index 0000000000000..9ae2ad3cc0860
--- /dev/null
+++ b/libcxx/test/benchmarks/numeric/pstl.transform_reduce.bench.cpp
@@ -0,0 +1,82 @@
+//===----------------------------------------------------------------------===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+// REQUIRES: std-at-least-c++17
+
+// UNSUPPORTED: libcpp-has-no-incomplete-pstl
+
+#include <algorithm>
+#include <cstddef>
+#include <cstdint>
+#include <cmath>
+#include <numeric>
+#include <string>
+#include <type_traits>
+#include <vector>
+#include <execution>
+#include <utility>
+
+#include <benchmark/benchmark.h>
+
+int main(int argc, char** argv) {
+ // A transformation/reduction that does just a bit more than a no-op
+ struct minimal {
+ double operator()(double x) const { return -x; }
+ double operator()(double x, double y) const { return x + y; }
+ };
+
+ // A transformation/reduction that does miniscule work per element
+ struct cheap {
+ double operator()(double x) const { return -x + 0.25; }
+ double operator()(double x, double y) const {
+ // (compilers aren't allowed to optimize away the FP divisions and multiplications without fast-math)
+ return x / 2. * 2. / 2. * 2. + y / 2. * 2. / 2. * 2.;
+ }
+ };
+
+ // A transformation/reduction that does significant work per element
+ struct expensive {
+ double operator()(double x) const { return (-x + 0.25) * (x + 10.); }
+ double operator()(double x, double y) const {
+ // (compilers aren't allowed to optimize away the FP divisions and multiplications without fast-math)
+ return x / 2. * 2. / 3. * 3. / 4. * 4. / 5. * 5. + y / 2. * 2. / 2. * 2. / 4. * 4. / 5. * 5.;
+ }
+ };
+
+ auto bm = [](std::string name, auto&& policy, auto func) {
+ benchmark::RegisterBenchmark(
+ name,
+ [&policy, func](auto& st) {
+ std::size_t size = st.range(0);
+ std::vector<double> c(size);
+ std::iota(c.begin(), c.end(), 1.);
+ for ([[maybe_unused]] auto _ : st) {
+ benchmark::ClobberMemory();
+ auto result = std::transform_reduce(policy, c.begin(), c.end(), 0.0, func, func);
+ benchmark::DoNotOptimize(result);
+ }
+ })
+ ->Arg(1 << 6) // 64
+ ->Arg(1 << 16) // 65'536
+ ->Arg(1 << 26) // 67'108'864
+ ->UseRealTime();
+ };
+#if defined(TEST_PSTL_ENABLE_SEQ_BASELINES)
+ bm("std::transform_reduce(std::execution::seq, vector<double>) (minimal)", std::execution::seq, minimal{});
+ bm("std::transform_reduce(std::execution::seq, vector<double>) (cheap)", std::execution::seq, cheap{});
+ bm("std::transform_reduce(std::execution::seq, vector<double>) (expensive)", std::execution::seq, expensive{});
+#endif
+ bm("std::transform_reduce(std::execution::par, vector<double>) (minimal)", std::execution::par, minimal{});
+ bm("std::transform_reduce(std::execution::par, vector<double>) (cheap)", std::execution::par, cheap{});
+ bm("std::transform_reduce(std::execution::par, vector<double>) (expensive)", std::execution::par, expensive{});
+
+ benchmark::Initialize(&argc, argv);
+ benchmark::RunSpecifiedBenchmarks();
+ benchmark::Shutdown();
+ return 0;
+}
>From f336092bc46a489dd9183db97025e6f322c11ff8 Mon Sep 17 00:00:00 2001
From: Matt Arsenault <Matthew.Arsenault at amd.com>
Date: Sat, 26 Sep 2026 19:59:39 +0200
Subject: [PATCH 22/53] ARM: Remove cached TargetABI from ARMBaseTargetMachine
(#226482)
This cannot account for the "target-abi" module flag, so
let the uses query that. ARMElfTargetObjectFile was the one user of
this value, so this failed to respect the module flag.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
---
llvm/lib/Target/ARM/ARMTargetMachine.cpp | 5 ++---
llvm/lib/Target/ARM/ARMTargetMachine.h | 18 ------------------
llvm/lib/Target/ARM/ARMTargetObjectFile.cpp | 20 +++++++++++++-------
llvm/lib/Target/ARM/ARMTargetObjectFile.h | 2 ++
llvm/test/CodeGen/ARM/module-target-abi.ll | 13 ++++++++++++-
5 files changed, 29 insertions(+), 29 deletions(-)
diff --git a/llvm/lib/Target/ARM/ARMTargetMachine.cpp b/llvm/lib/Target/ARM/ARMTargetMachine.cpp
index 667f6c00d5996..e1153ed22ab91 100644
--- a/llvm/lib/Target/ARM/ARMTargetMachine.cpp
+++ b/llvm/lib/Target/ARM/ARMTargetMachine.cpp
@@ -155,7 +155,6 @@ ARMBaseTargetMachine::ARMBaseTargetMachine(const Target &T, const Triple &TT,
: CodeGenTargetMachineImpl(T, TT, CPU, FS, Options,
getEffectiveRelocModel(TT, RM),
getEffectiveCodeModel(CM, CodeModel::Small), OL),
- TargetABI(ARM::computeTargetABI(TT, Options.MCOptions.ABIName)),
TLOF(createTLOF(getTargetTriple())), isLittle(TT.isLittleEndian()) {
if (TT.isOSBinFormatMachO()) {
@@ -224,7 +223,7 @@ FloatABI::ABIType ARMBaseTargetMachine::getFloatABI(const Module &M) const {
// With no explicit ABI, an explicit -target-abi=aapcs16 forces hard float
// even on triples whose default float ABI is soft (the triple default only
// detects AAPCS16 when it is the triple's own default ABI).
- if (TargetABI == ARM::ARM_ABI_AAPCS16)
+ if (getEffectiveABI(M) == ARM::ARM_ABI_AAPCS16)
return FloatABI::Hard;
// Otherwise fall back to the ABI implied by the target triple.
return M.getTargetTriple().getDefaultFloatABI();
@@ -234,7 +233,7 @@ ARM::ARMABI ARMBaseTargetMachine::getEffectiveABI(const Module &M) const {
// Consistency of "target-abi" and -target-abi is validated elsewhere.
if (const auto *MD = cast_or_null<MDString>(M.getModuleFlag("target-abi")))
return ARM::computeTargetABI(TargetTriple, MD->getString());
- return TargetABI;
+ return ARM::computeTargetABI(TargetTriple, Options.MCOptions.getABIName());
}
const ARMSubtarget *
diff --git a/llvm/lib/Target/ARM/ARMTargetMachine.h b/llvm/lib/Target/ARM/ARMTargetMachine.h
index 1771fcaa26c66..51960634bf993 100644
--- a/llvm/lib/Target/ARM/ARMTargetMachine.h
+++ b/llvm/lib/Target/ARM/ARMTargetMachine.h
@@ -27,9 +27,6 @@
namespace llvm {
class ARMBaseTargetMachine : public CodeGenTargetMachineImpl {
-public:
- ARM::ARMABI TargetABI;
-
protected:
std::unique_ptr<TargetLoweringObjectFile> TLOF;
bool isLittle;
@@ -73,21 +70,6 @@ class ARMBaseTargetMachine : public CodeGenTargetMachineImpl {
return TLOF.get();
}
- bool isAPCS_ABI() const {
- assert(TargetABI != ARM::ARM_ABI_UNKNOWN);
- return TargetABI == ARM::ARM_ABI_APCS;
- }
-
- bool isAAPCS_ABI() const {
- assert(TargetABI != ARM::ARM_ABI_UNKNOWN);
- return TargetABI == ARM::ARM_ABI_AAPCS || TargetABI == ARM::ARM_ABI_AAPCS16;
- }
-
- bool isAAPCS16_ABI() const {
- assert(TargetABI != ARM::ARM_ABI_UNKNOWN);
- return TargetABI == ARM::ARM_ABI_AAPCS16;
- }
-
bool targetSchedulesPostRAScheduling() const override { return true; };
MachineFunctionInfo *
diff --git a/llvm/lib/Target/ARM/ARMTargetObjectFile.cpp b/llvm/lib/Target/ARM/ARMTargetObjectFile.cpp
index d42a6484076d9..44040be9c8db9 100644
--- a/llvm/lib/Target/ARM/ARMTargetObjectFile.cpp
+++ b/llvm/lib/Target/ARM/ARMTargetObjectFile.cpp
@@ -35,17 +35,12 @@ ARMElfTargetObjectFile::ARMElfTargetObjectFile() {
void ARMElfTargetObjectFile::Initialize(MCContext &Ctx,
const TargetMachine &TM) {
- const ARMBaseTargetMachine &ARM_TM = static_cast<const ARMBaseTargetMachine &>(TM);
- bool isAAPCS_ABI = ARM_TM.TargetABI == ARM::ARMABI::ARM_ABI_AAPCS;
+ const ARMBaseTargetMachine &ARM_TM =
+ static_cast<const ARMBaseTargetMachine &>(TM);
bool genExecuteOnly =
ARM_TM.getMCSubtargetInfo().hasFeature(ARM::FeatureExecuteOnly);
TargetLoweringObjectFileELF::Initialize(Ctx, TM);
- InitializeELF(isAAPCS_ABI);
-
- if (isAAPCS_ABI) {
- LSDASection = nullptr;
- }
// Make code section unreadable when in execute-only mode
if (genExecuteOnly) {
@@ -60,6 +55,17 @@ void ARMElfTargetObjectFile::Initialize(MCContext &Ctx,
}
}
+void ARMElfTargetObjectFile::getModuleMetadata(Module &M) {
+ TargetLoweringObjectFileELF::getModuleMetadata(M);
+
+ const auto &ARM_TM = static_cast<const ARMBaseTargetMachine &>(*TM);
+ bool isAAPCS_ABI = ARM_TM.getEffectiveABI(M) == ARM::ARMABI::ARM_ABI_AAPCS;
+ InitializeELF(isAAPCS_ABI);
+
+ if (isAAPCS_ABI)
+ LSDASection = nullptr;
+}
+
MCRegister ARMElfTargetObjectFile::getStaticBase() const { return ARM::R9; }
const MCExpr *ARMElfTargetObjectFile::getIndirectSymViaGOTPCRel(
diff --git a/llvm/lib/Target/ARM/ARMTargetObjectFile.h b/llvm/lib/Target/ARM/ARMTargetObjectFile.h
index e2cce9b79e057..02cd138cf5c04 100644
--- a/llvm/lib/Target/ARM/ARMTargetObjectFile.h
+++ b/llvm/lib/Target/ARM/ARMTargetObjectFile.h
@@ -20,6 +20,8 @@ class ARMElfTargetObjectFile : public TargetLoweringObjectFileELF {
ARMElfTargetObjectFile();
void Initialize(MCContext &Ctx, const TargetMachine &TM) override;
+ void getModuleMetadata(Module &M) override;
+
MCRegister getStaticBase() const override;
const MCExpr *getIndirectSymViaGOTPCRel(const GlobalValue *GV,
diff --git a/llvm/test/CodeGen/ARM/module-target-abi.ll b/llvm/test/CodeGen/ARM/module-target-abi.ll
index 0cb41cae6c67c..72c59ad85c487 100644
--- a/llvm/test/CodeGen/ARM/module-target-abi.ll
+++ b/llvm/test/CodeGen/ARM/module-target-abi.ll
@@ -1,6 +1,7 @@
; The "target-abi" module flag selects the ABI used for codegen. APCS uses
; 4-byte stack alignment while AAPCS uses 8-byte alignment, which is observable
-; in the emitted prologue. The flag drives this with no -target-abi option.
+; in the emitted prologue. AAPCS also selects .init_array over .ctors. The flag
+; drives both with no -target-abi option.
; RUN: split-file %s %t
; RUN: llc -mtriple=armv7-none-eabi -filetype=asm < %t/apcs.ll | FileCheck %s --check-prefix=APCS
; RUN: llc -mtriple=armv7-none-eabi -filetype=asm < %t/aapcs.ll | FileCheck %s --check-prefix=AAPCS
@@ -8,23 +9,33 @@
;--- apcs.ll
; APCS: push {lr}
; APCS: sub sp, sp, #4
+; APCS: .section .ctors,"aw",%progbits
declare void @use(ptr)
define void @f() {
%a = alloca i32
call void @use(ptr %a)
ret void
}
+define void @ctor() {
+ ret void
+}
+ at llvm.global_ctors = appending global [1 x { i32, ptr, ptr }] [{ i32, ptr, ptr } { i32 65535, ptr @ctor, ptr null }]
!llvm.module.flags = !{!0}
!0 = !{i32 1, !"target-abi", !"apcs"}
;--- aapcs.ll
; AAPCS: push {r11, lr}
; AAPCS: sub sp, sp, #8
+; AAPCS: .section .init_array,"aw",%init_array
declare void @use(ptr)
define void @f() {
%a = alloca i32
call void @use(ptr %a)
ret void
}
+define void @ctor() {
+ ret void
+}
+ at llvm.global_ctors = appending global [1 x { i32, ptr, ptr }] [{ i32, ptr, ptr } { i32 65535, ptr @ctor, ptr null }]
!llvm.module.flags = !{!0}
!0 = !{i32 1, !"target-abi", !"aapcs"}
>From a3ba94b90992285ef52972202272952e237245e6 Mon Sep 17 00:00:00 2001
From: Alan Zhao <ayzhao at google.com>
Date: Sat, 26 Sep 2026 11:10:31 -0700
Subject: [PATCH 23/53] [cmake][Windows] Do not strip library prefixes on
Windows (#226635)
---
llvm/cmake/modules/GetLibraryName.cmake | 4 +++-
1 file changed, 3 insertions(+), 1 deletion(-)
diff --git a/llvm/cmake/modules/GetLibraryName.cmake b/llvm/cmake/modules/GetLibraryName.cmake
index 13c0080671a3c..5aa42c343a3d3 100644
--- a/llvm/cmake/modules/GetLibraryName.cmake
+++ b/llvm/cmake/modules/GetLibraryName.cmake
@@ -5,7 +5,9 @@ function(get_library_name path name)
set(suffixes ${CMAKE_FIND_LIBRARY_SUFFIXES})
list(FILTER prefixes EXCLUDE REGEX "^\\s*$")
list(FILTER suffixes EXCLUDE REGEX "^\\s*$")
- if(prefixes)
+ # Do not strip the "lib" prefix for Windows because MSVC-style linkers don't
+ # implicitly add the "lib" prefix.
+ if(prefixes AND NOT Win32)
string(REPLACE ";" "|" prefixes "${prefixes}")
string(REGEX REPLACE "^(${prefixes})" "" path ${path})
endif()
>From ea4860b150562af43d36097ff5f0bec77a2d0af1 Mon Sep 17 00:00:00 2001
From: Prajit Rahul <prajitrahul05 at gmail.com>
Date: Sat, 26 Sep 2026 12:40:05 -0700
Subject: [PATCH 24/53] [LLVM][Docs] Define IR in the lexicon (#221388)
Define IR in the LLVM lexicon using terminology from the LLVM Language
Reference Manual, including its SSA-based structure and three equivalent
representations.
Fixes #139867
AI Usage: ChatGPT
---
llvm/docs/Lexicon.md | 7 +++++++
1 file changed, 7 insertions(+)
diff --git a/llvm/docs/Lexicon.md b/llvm/docs/Lexicon.md
index 9faad0bbcbb83..7bd5ffc8b4b2e 100644
--- a/llvm/docs/Lexicon.md
+++ b/llvm/docs/Lexicon.md
@@ -162,6 +162,13 @@ This document is a work in progress!
: Inter-Procedural Optimization. Refers to any variety of code optimization
that occurs between procedures, functions or compilation units (modules).
+**IR**
+: Intermediate Representation. LLVM IR is the common code representation
+ used throughout LLVM's compilation process. It is typed and based on
+ Static Single Assignment (SSA), with three equivalent forms: in-memory
+ IR, on-disk bitcode, and human-readable assembly. See the
+ {doc}`LLVM Language Reference Manual <LangRef>`.
+
**ISel**
: Instruction Selection
>From 48be11705edfb92be58de11e8058535246e40ec5 Mon Sep 17 00:00:00 2001
From: Florian Hahn <flo at fhahn.com>
Date: Sat, 26 Sep 2026 21:02:25 +0100
Subject: [PATCH 25/53] [VPlan] Take wide induction wrap flags from the
increment of the phi. (#226726)
The binary operator of an integer InductionDescriptor is the incoming
value from the latch, which is not required to have the phi as operand.
E.g. for
%x = add i32 %iv, 5
%iv.next = add nuw nsw i32 %x, 1
the flags only apply to %x + 1, while the induction step is 6, and %iv +
6 may wrap even though %iv.next does not. Determine the wrap flags from
the increment of the header phi instead, and only if it adds to or
subtracts from the phi directly.
---
.../Vectorize/VPlanConstruction.cpp | 2 +-
llvm/lib/Transforms/Vectorize/VPlanUtils.cpp | 23 ++++++++++++++
llvm/lib/Transforms/Vectorize/VPlanUtils.h | 28 +++--------------
.../epilog-vectorization-widen-inductions.ll | 6 ++--
.../Transforms/LoopVectorize/X86/pr36524.ll | 2 +-
.../LoopVectorize/induction-wrapflags.ll | 2 +-
.../Transforms/LoopVectorize/induction.ll | 30 +++++++++----------
.../pr30654-phiscev-sext-trunc.ll | 18 +++++------
llvm/test/Transforms/LoopVectorize/pr35773.ll | 2 +-
.../LoopVectorize/predicated-inductions.ll | 4 +--
.../LoopVectorize/single-value-blend-phis.ll | 2 +-
11 files changed, 61 insertions(+), 58 deletions(-)
diff --git a/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp b/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
index 96653c9a1c3ad..6c5f688835c15 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
@@ -719,7 +719,7 @@ createWidenInductionRecipe(PHINode *Phi, VPPhi *PhiR, VPIRValue *Start,
// It is always safe to copy over the NoWrap and FastMath flags. In
// particular, when folding tail by masking, the masked-off lanes are never
// used, so it is safe.
- VPIRFlags Flags = vputils::getFlagsFromIndDesc(IndDesc);
+ VPIRFlags Flags = vputils::getFlagsForInduction(IndDesc, PhiR);
auto *WideIV = new VPWidenIntOrFpInductionRecipe(
Phi, Start, Step, &Plan.getVF(), IndDesc, Flags, DL);
diff --git a/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp b/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp
index 9a07697f66762..a56689848efdf 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp
@@ -717,6 +717,29 @@ VPBasicBlock *VPBlockUtils::getPlainCFGMiddleBlock(const VPlan &Plan) {
return cast<VPBasicBlock>(Plan.getScalarPreheader()->getPredecessors()[0]);
}
+VPIRFlags vputils::getFlagsForInduction(const InductionDescriptor &ID,
+ const VPPhi *PhiR) {
+ if (ID.getKind() == InductionDescriptor::IK_FpInduction)
+ return ID.getInductionBinOp()->getFastMathFlags();
+
+ // The flags only bound the induction values if the increment directly
+ // updates PhiR.
+ VPValue *Inc = PhiR->getOperand(1);
+ if (match(Inc, m_c_Add(m_Specific(PhiR), m_VPValue())))
+ return cast<VPInstruction>(Inc)->getNoWrapFlagsOrNone();
+
+ if (match(Inc, m_Sub(m_Specific(PhiR), m_VPValue()))) {
+ // The step of a sub induction is negated, so NUW cannot be preserved. NSW
+ // can, if the step is not the signed minimum.
+ ConstantInt *Step = ID.getConstIntStepValue();
+ bool NSW = cast<VPInstruction>(Inc)->getNoWrapFlagsOrNone().HasNSW &&
+ Step && !Step->isMinValue(/*IsSigned=*/true);
+ return VPIRFlags::WrapFlagsTy(/*NUW*/ false, NSW);
+ }
+
+ return VPIRFlags::WrapFlagsTy(false, false);
+}
+
std::optional<MemoryLocation>
vputils::getMemoryLocation(const VPRecipeBase &R) {
auto *M = dyn_cast<VPIRMetadata>(&R);
diff --git a/llvm/lib/Transforms/Vectorize/VPlanUtils.h b/llvm/lib/Transforms/Vectorize/VPlanUtils.h
index 7cdc846fced99..967e612ab7114 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanUtils.h
+++ b/llvm/lib/Transforms/Vectorize/VPlanUtils.h
@@ -135,30 +135,10 @@ getOpcodeOrIntrinsicID(const VPValue *V);
/// the location is conservatively set to nullptr.
std::optional<MemoryLocation> getMemoryLocation(const VPRecipeBase &R);
-/// Extracts and returns NoWrap and FastMath flags from the induction binop in
-/// \p ID, for use on a wide induction, which adds the step.
-inline VPIRFlags getFlagsFromIndDesc(const InductionDescriptor &ID) {
- if (ID.getKind() == InductionDescriptor::IK_FpInduction)
- return ID.getInductionBinOp()->getFastMathFlags();
-
- if (auto *AddO = dyn_cast_if_present<AddOperator>(ID.getInductionBinOp())) {
- return VPIRFlags::WrapFlagsTy(AddO->hasNoUnsignedWrap(),
- AddO->hasNoSignedWrap());
- }
-
- // The step of a sub induction is negated, so NUW cannot be preserved. NSW
- // can, if the step is not the signed minimum.
- if (auto *SubO = dyn_cast_if_present<SubOperator>(ID.getInductionBinOp())) {
- ConstantInt *Step = ID.getConstIntStepValue();
- return VPIRFlags::WrapFlagsTy(false,
- SubO->hasNoSignedWrap() && Step &&
- !Step->isMinValue(/*IsSigned=*/true));
- }
-
- assert(ID.getKind() == InductionDescriptor::IK_IntInduction &&
- "Expected int induction");
- return VPIRFlags::WrapFlagsTy(false, false);
-}
+/// Extracts and returns NoWrap flags from \p PhiR and fast-math flags from \p
+/// ID.
+VPIRFlags getFlagsForInduction(const InductionDescriptor &ID,
+ const VPPhi *PhiR);
/// Search \p Start's users for a recipe satisfying \p Pred, looking through
/// recipes with definitions.
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/epilog-vectorization-widen-inductions.ll b/llvm/test/Transforms/LoopVectorize/AArch64/epilog-vectorization-widen-inductions.ll
index b28e642155b64..5a2a403174f63 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/epilog-vectorization-widen-inductions.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/epilog-vectorization-widen-inductions.ll
@@ -294,7 +294,7 @@ define void @test_widen_induction_step_2(ptr %A, i64 %N, i32 %step) {
; CHECK-NEXT: store <4 x i64> [[TMP4]], ptr [[TMP1]], align 4
; CHECK-NEXT: store <4 x i64> [[TMP2]], ptr [[TMP3]], align 4
; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 8
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add nuw nsw <4 x i64> [[STEP_ADD]], splat (i64 4)
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <4 x i64> [[STEP_ADD]], splat (i64 4)
; CHECK-NEXT: [[TMP6:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[IND_END4]]
; CHECK-NEXT: br i1 [[TMP6]], label [[MIDDLE_BLOCK:%.*]], label [[VECTOR_BODY]], {{!llvm.loop ![0-9]+}}
; CHECK: middle.block:
@@ -309,7 +309,7 @@ define void @test_widen_induction_step_2(ptr %A, i64 %N, i32 %step) {
; CHECK-NEXT: [[IND_END:%.*]] = sub i64 [[N]], [[N_MOD_VF2]]
; CHECK-NEXT: [[DOTSPLATINSERT:%.*]] = insertelement <2 x i64> poison, i64 [[VEC_EPILOG_RESUME_VAL]], i64 0
; CHECK-NEXT: [[DOTSPLAT:%.*]] = shufflevector <2 x i64> [[DOTSPLATINSERT]], <2 x i64> poison, <2 x i32> zeroinitializer
-; CHECK-NEXT: [[INDUCTION:%.*]] = add nuw nsw <2 x i64> [[DOTSPLAT]], <i64 0, i64 1>
+; CHECK-NEXT: [[INDUCTION:%.*]] = add <2 x i64> [[DOTSPLAT]], <i64 0, i64 1>
; CHECK-NEXT: br label [[VEC_EPILOG_VECTOR_BODY:%.*]]
; CHECK: vec.epilog.vector.body:
; CHECK-NEXT: [[INDEX7:%.*]] = phi i64 [ [[VEC_EPILOG_RESUME_VAL]], [[VEC_EPILOG_PH]] ], [ [[INDEX_NEXT10:%.*]], [[VEC_EPILOG_VECTOR_BODY]] ]
@@ -318,7 +318,7 @@ define void @test_widen_induction_step_2(ptr %A, i64 %N, i32 %step) {
; CHECK-NEXT: [[TMP9:%.*]] = add <2 x i64> [[VEC_IND8]], splat (i64 10)
; CHECK-NEXT: store <2 x i64> [[TMP9]], ptr [[TMP8]], align 4
; CHECK-NEXT: [[INDEX_NEXT10]] = add nuw i64 [[INDEX7]], 2
-; CHECK-NEXT: [[VEC_IND_NEXT9]] = add nuw nsw <2 x i64> [[VEC_IND8]], splat (i64 2)
+; CHECK-NEXT: [[VEC_IND_NEXT9]] = add <2 x i64> [[VEC_IND8]], splat (i64 2)
; CHECK-NEXT: [[TMP11:%.*]] = icmp eq i64 [[INDEX_NEXT10]], [[IND_END]]
; CHECK-NEXT: br i1 [[TMP11]], label [[VEC_EPILOG_MIDDLE_BLOCK:%.*]], label [[VEC_EPILOG_VECTOR_BODY]], {{!llvm.loop ![0-9]+}}
; CHECK: vec.epilog.middle.block:
diff --git a/llvm/test/Transforms/LoopVectorize/X86/pr36524.ll b/llvm/test/Transforms/LoopVectorize/X86/pr36524.ll
index 81e3fda7f44b2..4b9611e6ded96 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/pr36524.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/pr36524.ll
@@ -22,7 +22,7 @@ define void @foo(ptr %ptr, ptr %ptr.2) {
; CHECK-NEXT: [[TMP6:%.*]] = getelementptr inbounds i64, ptr [[PTR]], i64 [[INDEX]]
; CHECK-NEXT: store <4 x i64> [[VEC_IND]], ptr [[TMP6]], align 8, !alias.scope [[META0:![0-9]+]]
; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add nuw nsw <4 x i64> [[VEC_IND]], splat (i64 4)
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <4 x i64> [[VEC_IND]], splat (i64 4)
; CHECK-NEXT: [[TMP8:%.*]] = icmp eq i64 [[INDEX_NEXT]], 80
; CHECK-NEXT: br i1 [[TMP8]], label [[MIDDLE_BLOCK:%.*]], label [[VECTOR_BODY]], !llvm.loop [[LOOP3:![0-9]+]]
; CHECK: middle.block:
diff --git a/llvm/test/Transforms/LoopVectorize/induction-wrapflags.ll b/llvm/test/Transforms/LoopVectorize/induction-wrapflags.ll
index 20ba7f8366e30..8be1725227b2b 100644
--- a/llvm/test/Transforms/LoopVectorize/induction-wrapflags.ll
+++ b/llvm/test/Transforms/LoopVectorize/induction-wrapflags.ll
@@ -486,7 +486,7 @@ define void @add_induction_nuw_nsw_increment_not_on_phi(ptr noalias %dst) {
; CHECK-NEXT: [[TMP0:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
; CHECK-NEXT: store <4 x i32> [[VEC_IND]], ptr [[TMP0]], align 4
; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add nuw nsw <4 x i32> [[VEC_IND]], splat (i32 24)
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <4 x i32> [[VEC_IND]], splat (i32 24)
; CHECK-NEXT: [[TMP1:%.*]] = icmp eq i64 [[INDEX_NEXT]], 16
; CHECK-NEXT: br i1 [[TMP1]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP19:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
diff --git a/llvm/test/Transforms/LoopVectorize/induction.ll b/llvm/test/Transforms/LoopVectorize/induction.ll
index 1683f784daf75..f7c93cda82056 100644
--- a/llvm/test/Transforms/LoopVectorize/induction.ll
+++ b/llvm/test/Transforms/LoopVectorize/induction.ll
@@ -5952,8 +5952,8 @@ define void @test_optimized_cast_induction_feeding_first_order_recurrence(i64 %n
; CHECK-NEXT: [[IND_END:%.*]] = mul i32 [[DOTCAST]], [[STEP]]
; CHECK-NEXT: [[DOTSPLATINSERT:%.*]] = insertelement <2 x i32> poison, i32 [[STEP]], i64 0
; CHECK-NEXT: [[DOTSPLAT:%.*]] = shufflevector <2 x i32> [[DOTSPLATINSERT]], <2 x i32> poison, <2 x i32> zeroinitializer
-; CHECK-NEXT: [[TMP19:%.*]] = mul nsw <2 x i32> <i32 0, i32 1>, [[DOTSPLAT]]
-; CHECK-NEXT: [[TMP18:%.*]] = shl nsw i32 [[STEP]], 1
+; CHECK-NEXT: [[TMP19:%.*]] = mul <2 x i32> <i32 0, i32 1>, [[DOTSPLAT]]
+; CHECK-NEXT: [[TMP18:%.*]] = shl i32 [[STEP]], 1
; CHECK-NEXT: [[DOTSPLATINSERT2:%.*]] = insertelement <2 x i32> poison, i32 [[TMP18]], i64 0
; CHECK-NEXT: [[DOTSPLAT3:%.*]] = shufflevector <2 x i32> [[DOTSPLATINSERT2]], <2 x i32> poison, <2 x i32> zeroinitializer
; CHECK-NEXT: br label [[VECTOR_BODY:%.*]]
@@ -5965,7 +5965,7 @@ define void @test_optimized_cast_induction_feeding_first_order_recurrence(i64 %n
; CHECK-NEXT: [[TMP21:%.*]] = getelementptr inbounds i32, ptr [[PTR:%.*]], i64 [[INDEX]]
; CHECK-NEXT: store <2 x i32> [[TMP20]], ptr [[TMP21]], align 4
; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 2
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add nsw <2 x i32> [[VEC_IND]], [[DOTSPLAT3]]
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <2 x i32> [[VEC_IND]], [[DOTSPLAT3]]
; CHECK-NEXT: [[TMP23:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-NEXT: br i1 [[TMP23]], label [[MIDDLE_BLOCK:%.*]], label [[VECTOR_BODY]], !llvm.loop [[LOOP48:![0-9]+]]
; CHECK: middle.block:
@@ -6026,8 +6026,8 @@ define void @test_optimized_cast_induction_feeding_first_order_recurrence(i64 %n
; IND-NEXT: [[IND_END:%.*]] = mul i32 [[DOTCAST]], [[STEP]]
; IND-NEXT: [[DOTSPLATINSERT:%.*]] = insertelement <2 x i32> poison, i32 [[STEP]], i64 0
; IND-NEXT: [[DOTSPLAT:%.*]] = shufflevector <2 x i32> [[DOTSPLATINSERT]], <2 x i32> poison, <2 x i32> zeroinitializer
-; IND-NEXT: [[TMP15:%.*]] = mul nsw <2 x i32> <i32 0, i32 1>, [[DOTSPLAT]]
-; IND-NEXT: [[TMP16:%.*]] = shl nsw i32 [[STEP]], 1
+; IND-NEXT: [[TMP15:%.*]] = mul <2 x i32> <i32 0, i32 1>, [[DOTSPLAT]]
+; IND-NEXT: [[TMP16:%.*]] = shl i32 [[STEP]], 1
; IND-NEXT: [[DOTSPLATINSERT2:%.*]] = insertelement <2 x i32> poison, i32 [[TMP16]], i64 0
; IND-NEXT: [[DOTSPLAT3:%.*]] = shufflevector <2 x i32> [[DOTSPLATINSERT2]], <2 x i32> poison, <2 x i32> zeroinitializer
; IND-NEXT: br label [[VECTOR_BODY:%.*]]
@@ -6039,7 +6039,7 @@ define void @test_optimized_cast_induction_feeding_first_order_recurrence(i64 %n
; IND-NEXT: [[TMP18:%.*]] = getelementptr inbounds i32, ptr [[PTR:%.*]], i64 [[INDEX]]
; IND-NEXT: store <2 x i32> [[TMP17]], ptr [[TMP18]], align 4
; IND-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 2
-; IND-NEXT: [[VEC_IND_NEXT]] = add nsw <2 x i32> [[VEC_IND]], [[DOTSPLAT3]]
+; IND-NEXT: [[VEC_IND_NEXT]] = add <2 x i32> [[VEC_IND]], [[DOTSPLAT3]]
; IND-NEXT: [[TMP19:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; IND-NEXT: br i1 [[TMP19]], label [[MIDDLE_BLOCK:%.*]], label [[VECTOR_BODY]], !llvm.loop [[LOOP48:![0-9]+]]
; IND: middle.block:
@@ -6101,13 +6101,13 @@ define void @test_optimized_cast_induction_feeding_first_order_recurrence(i64 %n
; UNROLL-NEXT: [[DOTCAST:%.*]] = trunc i64 [[N_VEC]] to i32
; UNROLL-NEXT: [[IND_END:%.*]] = mul i32 [[DOTCAST]], [[STEP]]
; UNROLL-NEXT: [[TMP16:%.*]] = shl <2 x i32> [[DOTSPLAT]], splat (i32 1)
-; UNROLL-NEXT: [[TMP17:%.*]] = mul nsw <2 x i32> <i32 0, i32 1>, [[DOTSPLAT]]
+; UNROLL-NEXT: [[TMP17:%.*]] = mul <2 x i32> <i32 0, i32 1>, [[DOTSPLAT]]
; UNROLL-NEXT: br label [[VECTOR_BODY:%.*]]
; UNROLL: vector.body:
; UNROLL-NEXT: [[INDEX:%.*]] = phi i64 [ 0, [[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], [[VECTOR_BODY]] ]
; UNROLL-NEXT: [[VECTOR_RECUR:%.*]] = phi <2 x i32> [ <i32 poison, i32 0>, [[VECTOR_PH]] ], [ [[STEP_ADD:%.*]], [[VECTOR_BODY]] ]
; UNROLL-NEXT: [[VEC_IND:%.*]] = phi <2 x i32> [ [[TMP17]], [[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], [[VECTOR_BODY]] ]
-; UNROLL-NEXT: [[STEP_ADD]] = add nsw <2 x i32> [[VEC_IND]], [[TMP16]]
+; UNROLL-NEXT: [[STEP_ADD]] = add <2 x i32> [[VEC_IND]], [[TMP16]]
; UNROLL-NEXT: [[TMP18:%.*]] = shufflevector <2 x i32> [[VECTOR_RECUR]], <2 x i32> [[VEC_IND]], <2 x i32> <i32 1, i32 2>
; UNROLL-NEXT: [[TMP19:%.*]] = shufflevector <2 x i32> [[VEC_IND]], <2 x i32> [[STEP_ADD]], <2 x i32> <i32 1, i32 2>
; UNROLL-NEXT: [[TMP20:%.*]] = getelementptr inbounds i32, ptr [[PTR:%.*]], i64 [[INDEX]]
@@ -6115,7 +6115,7 @@ define void @test_optimized_cast_induction_feeding_first_order_recurrence(i64 %n
; UNROLL-NEXT: store <2 x i32> [[TMP18]], ptr [[TMP20]], align 4
; UNROLL-NEXT: store <2 x i32> [[TMP19]], ptr [[TMP21]], align 4
; UNROLL-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; UNROLL-NEXT: [[VEC_IND_NEXT]] = add nsw <2 x i32> [[STEP_ADD]], [[TMP16]]
+; UNROLL-NEXT: [[VEC_IND_NEXT]] = add <2 x i32> [[STEP_ADD]], [[TMP16]]
; UNROLL-NEXT: [[TMP22:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; UNROLL-NEXT: br i1 [[TMP22]], label [[MIDDLE_BLOCK:%.*]], label [[VECTOR_BODY]], !llvm.loop [[LOOP48:![0-9]+]]
; UNROLL: middle.block:
@@ -6177,13 +6177,13 @@ define void @test_optimized_cast_induction_feeding_first_order_recurrence(i64 %n
; UNROLL-NO-IC-NEXT: [[DOTCAST:%.*]] = trunc i64 [[N_VEC]] to i32
; UNROLL-NO-IC-NEXT: [[IND_END:%.*]] = mul i32 [[DOTCAST]], [[STEP]]
; UNROLL-NO-IC-NEXT: [[TMP17:%.*]] = shl <2 x i32> [[BROADCAST_SPLAT]], splat (i32 1)
-; UNROLL-NO-IC-NEXT: [[TMP19:%.*]] = mul nsw <2 x i32> <i32 0, i32 1>, [[BROADCAST_SPLAT]]
+; UNROLL-NO-IC-NEXT: [[TMP19:%.*]] = mul <2 x i32> <i32 0, i32 1>, [[BROADCAST_SPLAT]]
; UNROLL-NO-IC-NEXT: br label [[VECTOR_BODY:%.*]]
; UNROLL-NO-IC: vector.body:
; UNROLL-NO-IC-NEXT: [[INDEX:%.*]] = phi i64 [ 0, [[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], [[VECTOR_BODY]] ]
; UNROLL-NO-IC-NEXT: [[VECTOR_RECUR:%.*]] = phi <2 x i32> [ <i32 poison, i32 0>, [[VECTOR_PH]] ], [ [[STEP_ADD:%.*]], [[VECTOR_BODY]] ]
; UNROLL-NO-IC-NEXT: [[VEC_IND:%.*]] = phi <2 x i32> [ [[TMP19]], [[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], [[VECTOR_BODY]] ]
-; UNROLL-NO-IC-NEXT: [[STEP_ADD]] = add nsw <2 x i32> [[VEC_IND]], [[TMP17]]
+; UNROLL-NO-IC-NEXT: [[STEP_ADD]] = add <2 x i32> [[VEC_IND]], [[TMP17]]
; UNROLL-NO-IC-NEXT: [[TMP20:%.*]] = shufflevector <2 x i32> [[VECTOR_RECUR]], <2 x i32> [[VEC_IND]], <2 x i32> <i32 1, i32 2>
; UNROLL-NO-IC-NEXT: [[TMP21:%.*]] = shufflevector <2 x i32> [[VEC_IND]], <2 x i32> [[STEP_ADD]], <2 x i32> <i32 1, i32 2>
; UNROLL-NO-IC-NEXT: [[TMP22:%.*]] = getelementptr inbounds i32, ptr [[PTR:%.*]], i64 [[INDEX]]
@@ -6191,7 +6191,7 @@ define void @test_optimized_cast_induction_feeding_first_order_recurrence(i64 %n
; UNROLL-NO-IC-NEXT: store <2 x i32> [[TMP20]], ptr [[TMP22]], align 4
; UNROLL-NO-IC-NEXT: store <2 x i32> [[TMP21]], ptr [[TMP24]], align 4
; UNROLL-NO-IC-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; UNROLL-NO-IC-NEXT: [[VEC_IND_NEXT]] = add nsw <2 x i32> [[STEP_ADD]], [[TMP17]]
+; UNROLL-NO-IC-NEXT: [[VEC_IND_NEXT]] = add <2 x i32> [[STEP_ADD]], [[TMP17]]
; UNROLL-NO-IC-NEXT: [[TMP25:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; UNROLL-NO-IC-NEXT: br i1 [[TMP25]], label [[MIDDLE_BLOCK:%.*]], label [[VECTOR_BODY]], !llvm.loop [[LOOP48:![0-9]+]]
; UNROLL-NO-IC: middle.block:
@@ -6253,13 +6253,13 @@ define void @test_optimized_cast_induction_feeding_first_order_recurrence(i64 %n
; INTERLEAVE-NEXT: [[DOTCAST:%.*]] = trunc i64 [[N_VEC]] to i32
; INTERLEAVE-NEXT: [[IND_END:%.*]] = mul i32 [[DOTCAST]], [[STEP]]
; INTERLEAVE-NEXT: [[TMP16:%.*]] = shl <4 x i32> [[DOTSPLAT]], splat (i32 2)
-; INTERLEAVE-NEXT: [[TMP17:%.*]] = mul nsw <4 x i32> <i32 0, i32 1, i32 2, i32 3>, [[DOTSPLAT]]
+; INTERLEAVE-NEXT: [[TMP17:%.*]] = mul <4 x i32> <i32 0, i32 1, i32 2, i32 3>, [[DOTSPLAT]]
; INTERLEAVE-NEXT: br label [[VECTOR_BODY:%.*]]
; INTERLEAVE: vector.body:
; INTERLEAVE-NEXT: [[INDEX:%.*]] = phi i64 [ 0, [[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], [[VECTOR_BODY]] ]
; INTERLEAVE-NEXT: [[VECTOR_RECUR:%.*]] = phi <4 x i32> [ <i32 poison, i32 poison, i32 poison, i32 0>, [[VECTOR_PH]] ], [ [[STEP_ADD:%.*]], [[VECTOR_BODY]] ]
; INTERLEAVE-NEXT: [[VEC_IND:%.*]] = phi <4 x i32> [ [[TMP17]], [[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], [[VECTOR_BODY]] ]
-; INTERLEAVE-NEXT: [[STEP_ADD]] = add nsw <4 x i32> [[VEC_IND]], [[TMP16]]
+; INTERLEAVE-NEXT: [[STEP_ADD]] = add <4 x i32> [[VEC_IND]], [[TMP16]]
; INTERLEAVE-NEXT: [[TMP18:%.*]] = shufflevector <4 x i32> [[VECTOR_RECUR]], <4 x i32> [[VEC_IND]], <4 x i32> <i32 3, i32 4, i32 5, i32 6>
; INTERLEAVE-NEXT: [[TMP19:%.*]] = shufflevector <4 x i32> [[VEC_IND]], <4 x i32> [[STEP_ADD]], <4 x i32> <i32 3, i32 4, i32 5, i32 6>
; INTERLEAVE-NEXT: [[TMP20:%.*]] = getelementptr inbounds i32, ptr [[PTR:%.*]], i64 [[INDEX]]
@@ -6267,7 +6267,7 @@ define void @test_optimized_cast_induction_feeding_first_order_recurrence(i64 %n
; INTERLEAVE-NEXT: store <4 x i32> [[TMP18]], ptr [[TMP20]], align 4
; INTERLEAVE-NEXT: store <4 x i32> [[TMP19]], ptr [[TMP21]], align 4
; INTERLEAVE-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 8
-; INTERLEAVE-NEXT: [[VEC_IND_NEXT]] = add nsw <4 x i32> [[STEP_ADD]], [[TMP16]]
+; INTERLEAVE-NEXT: [[VEC_IND_NEXT]] = add <4 x i32> [[STEP_ADD]], [[TMP16]]
; INTERLEAVE-NEXT: [[TMP22:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; INTERLEAVE-NEXT: br i1 [[TMP22]], label [[MIDDLE_BLOCK:%.*]], label [[VECTOR_BODY]], !llvm.loop [[LOOP50:![0-9]+]]
; INTERLEAVE: middle.block:
diff --git a/llvm/test/Transforms/LoopVectorize/pr30654-phiscev-sext-trunc.ll b/llvm/test/Transforms/LoopVectorize/pr30654-phiscev-sext-trunc.ll
index 5b42fec52269d..1d19e85c1c1d3 100644
--- a/llvm/test/Transforms/LoopVectorize/pr30654-phiscev-sext-trunc.ll
+++ b/llvm/test/Transforms/LoopVectorize/pr30654-phiscev-sext-trunc.ll
@@ -70,8 +70,8 @@ define void @doit1(i32 %n, i32 %step) {
; CHECK-NEXT: [[IND_END:%.*]] = mul i32 [[DOTCAST]], [[STEP]]
; CHECK-NEXT: [[DOTSPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[STEP]], i64 0
; CHECK-NEXT: [[DOTSPLAT:%.*]] = shufflevector <4 x i32> [[DOTSPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
-; CHECK-NEXT: [[TMP19:%.*]] = mul nsw <4 x i32> <i32 0, i32 1, i32 2, i32 3>, [[DOTSPLAT]]
-; CHECK-NEXT: [[TMP18:%.*]] = shl nsw i32 [[STEP]], 2
+; CHECK-NEXT: [[TMP19:%.*]] = mul <4 x i32> <i32 0, i32 1, i32 2, i32 3>, [[DOTSPLAT]]
+; CHECK-NEXT: [[TMP18:%.*]] = shl i32 [[STEP]], 2
; CHECK-NEXT: [[DOTSPLATINSERT2:%.*]] = insertelement <4 x i32> poison, i32 [[TMP18]], i64 0
; CHECK-NEXT: [[DOTSPLAT3:%.*]] = shufflevector <4 x i32> [[DOTSPLATINSERT2]], <4 x i32> poison, <4 x i32> zeroinitializer
; CHECK-NEXT: br label [[VECTOR_BODY:%.*]]
@@ -81,7 +81,7 @@ define void @doit1(i32 %n, i32 %step) {
; CHECK-NEXT: [[TMP20:%.*]] = getelementptr inbounds [250 x i32], ptr @a, i64 0, i64 [[INDEX]]
; CHECK-NEXT: store <4 x i32> [[VEC_IND]], ptr [[TMP20]], align 4
; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add nsw <4 x i32> [[VEC_IND]], [[DOTSPLAT3]]
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <4 x i32> [[VEC_IND]], [[DOTSPLAT3]]
; CHECK-NEXT: [[TMP22:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-NEXT: br i1 [[TMP22]], label [[MIDDLE_BLOCK:%.*]], label [[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
; CHECK: middle.block:
@@ -189,8 +189,8 @@ define void @doit2(i32 %n, i32 %step) {
; CHECK-NEXT: [[IND_END:%.*]] = mul i32 [[DOTCAST]], [[STEP]]
; CHECK-NEXT: [[DOTSPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[STEP]], i64 0
; CHECK-NEXT: [[DOTSPLAT:%.*]] = shufflevector <4 x i32> [[DOTSPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
-; CHECK-NEXT: [[TMP18:%.*]] = mul nsw <4 x i32> <i32 0, i32 1, i32 2, i32 3>, [[DOTSPLAT]]
-; CHECK-NEXT: [[TMP17:%.*]] = shl nsw i32 [[STEP]], 2
+; CHECK-NEXT: [[TMP18:%.*]] = mul <4 x i32> <i32 0, i32 1, i32 2, i32 3>, [[DOTSPLAT]]
+; CHECK-NEXT: [[TMP17:%.*]] = shl i32 [[STEP]], 2
; CHECK-NEXT: [[DOTSPLATINSERT2:%.*]] = insertelement <4 x i32> poison, i32 [[TMP17]], i64 0
; CHECK-NEXT: [[DOTSPLAT3:%.*]] = shufflevector <4 x i32> [[DOTSPLATINSERT2]], <4 x i32> poison, <4 x i32> zeroinitializer
; CHECK-NEXT: br label [[VECTOR_BODY:%.*]]
@@ -200,7 +200,7 @@ define void @doit2(i32 %n, i32 %step) {
; CHECK-NEXT: [[TMP19:%.*]] = getelementptr inbounds [250 x i32], ptr @a, i64 0, i64 [[INDEX]]
; CHECK-NEXT: store <4 x i32> [[VEC_IND]], ptr [[TMP19]], align 4
; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add nsw <4 x i32> [[VEC_IND]], [[DOTSPLAT3]]
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <4 x i32> [[VEC_IND]], [[DOTSPLAT3]]
; CHECK-NEXT: [[TMP21:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-NEXT: br i1 [[TMP21]], label [[MIDDLE_BLOCK:%.*]], label [[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
; CHECK: middle.block:
@@ -379,8 +379,8 @@ define void @doit4(i32 %n, i8 signext %cstep) {
; CHECK-NEXT: [[IND_END:%.*]] = mul i32 [[DOTCAST]], [[CONV]]
; CHECK-NEXT: [[DOTSPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[CONV]], i64 0
; CHECK-NEXT: [[DOTSPLAT:%.*]] = shufflevector <4 x i32> [[DOTSPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
-; CHECK-NEXT: [[TMP16:%.*]] = mul nsw <4 x i32> <i32 0, i32 1, i32 2, i32 3>, [[DOTSPLAT]]
-; CHECK-NEXT: [[TMP15:%.*]] = shl nsw i32 [[CONV]], 2
+; CHECK-NEXT: [[TMP16:%.*]] = mul <4 x i32> <i32 0, i32 1, i32 2, i32 3>, [[DOTSPLAT]]
+; CHECK-NEXT: [[TMP15:%.*]] = shl i32 [[CONV]], 2
; CHECK-NEXT: [[DOTSPLATINSERT2:%.*]] = insertelement <4 x i32> poison, i32 [[TMP15]], i64 0
; CHECK-NEXT: [[DOTSPLAT3:%.*]] = shufflevector <4 x i32> [[DOTSPLATINSERT2]], <4 x i32> poison, <4 x i32> zeroinitializer
; CHECK-NEXT: br label [[VECTOR_BODY:%.*]]
@@ -390,7 +390,7 @@ define void @doit4(i32 %n, i8 signext %cstep) {
; CHECK-NEXT: [[TMP17:%.*]] = getelementptr inbounds [250 x i32], ptr @a, i64 0, i64 [[INDEX]]
; CHECK-NEXT: store <4 x i32> [[VEC_IND]], ptr [[TMP17]], align 4
; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add nsw <4 x i32> [[VEC_IND]], [[DOTSPLAT3]]
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <4 x i32> [[VEC_IND]], [[DOTSPLAT3]]
; CHECK-NEXT: [[TMP19:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-NEXT: br i1 [[TMP19]], label [[MIDDLE_BLOCK:%.*]], label [[VECTOR_BODY]], !llvm.loop [[LOOP6:![0-9]+]]
; CHECK: middle.block:
diff --git a/llvm/test/Transforms/LoopVectorize/pr35773.ll b/llvm/test/Transforms/LoopVectorize/pr35773.ll
index 6b8830d916e5b..9787e86834f69 100644
--- a/llvm/test/Transforms/LoopVectorize/pr35773.ll
+++ b/llvm/test/Transforms/LoopVectorize/pr35773.ll
@@ -17,7 +17,7 @@ define void @doit1(ptr %ptr) {
; CHECK-NEXT: store <4 x i32> [[I32_IV]], ptr [[GEP1]], align 4
; CHECK-NEXT: [[MAIN_IV_NEXT]] = add nuw i32 [[MAIN_IV]], 4
-; CHECK-NEXT: [[I32_IV_NEXT]] = add nuw nsw <4 x i32> [[I32_IV]], splat (i32 36)
+; CHECK-NEXT: [[I32_IV_NEXT]] = add <4 x i32> [[I32_IV]], splat (i32 36)
; CHECK-NEXT: [[IV_FROM_TRUNC_NEXT]] = add <4 x i8> [[IV_FROM_TRUNC]], splat (i8 36)
; CHECK-NEXT: [[TMP9:%.*]] = icmp eq i32 [[MAIN_IV_NEXT]], 16
; CHECK-NEXT: br i1 [[TMP9]], label [[MIDDLE_BLOCK:%.*]], label [[VECTOR_BODY]], !llvm.loop !0
diff --git a/llvm/test/Transforms/LoopVectorize/predicated-inductions.ll b/llvm/test/Transforms/LoopVectorize/predicated-inductions.ll
index 80d32ca536ced..c4f1338f6e418 100644
--- a/llvm/test/Transforms/LoopVectorize/predicated-inductions.ll
+++ b/llvm/test/Transforms/LoopVectorize/predicated-inductions.ll
@@ -1002,8 +1002,8 @@ define void @two_used_predicated_ivs(ptr %dst1, ptr %dst2, i64 %n) {
; CHECK-NEXT: [[TMP14:%.*]] = getelementptr inbounds i32, ptr [[DST2]], i64 [[INDEX]]
; CHECK-NEXT: store <4 x i32> [[TMP12]], ptr [[TMP14]], align 4
; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add nuw nsw <4 x i32> [[VEC_IND]], splat (i32 36)
-; CHECK-NEXT: [[VEC_IND_NEXT9]] = add nuw nsw <4 x i32> [[VEC_IND8]], splat (i32 20)
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <4 x i32> [[VEC_IND]], splat (i32 36)
+; CHECK-NEXT: [[VEC_IND_NEXT9]] = add <4 x i32> [[VEC_IND8]], splat (i32 20)
; CHECK-NEXT: [[TMP15:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-NEXT: br i1 [[TMP15]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP16:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
diff --git a/llvm/test/Transforms/LoopVectorize/single-value-blend-phis.ll b/llvm/test/Transforms/LoopVectorize/single-value-blend-phis.ll
index 6bd259c189c0c..48cd2b3e62db6 100644
--- a/llvm/test/Transforms/LoopVectorize/single-value-blend-phis.ll
+++ b/llvm/test/Transforms/LoopVectorize/single-value-blend-phis.ll
@@ -282,7 +282,7 @@ define void @duplicated_incoming_blocks_blend(i32 %x, ptr %ptr) {
; CHECK-NEXT: [[TMP1:%.*]] = getelementptr i32, ptr [[PTR:%.*]], i32 [[INDEX]]
; CHECK-NEXT: store <2 x i32> [[VEC_IND]], ptr [[TMP1]], align 4
; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 2
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add nsw <2 x i32> [[VEC_IND]], splat (i32 2)
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <2 x i32> [[VEC_IND]], splat (i32 2)
; CHECK-NEXT: [[TMP3:%.*]] = icmp eq i32 [[INDEX_NEXT]], 1000
; CHECK-NEXT: br i1 [[TMP3]], label [[MIDDLE_BLOCK:%.*]], label [[VECTOR_BODY]], !llvm.loop [[LOOP6:![0-9]+]]
; CHECK: middle.block:
>From f0efa076587bd8ffa0563f75cc377f6871b969f9 Mon Sep 17 00:00:00 2001
From: Florian Hahn <flo at fhahn.com>
Date: Sat, 26 Sep 2026 21:13:37 +0100
Subject: [PATCH 26/53] [LV] Skip low-trip count logic there is no scalar tail.
(#225633)
https://github.com/llvm/llvm-project/pull/195823 added logic consider
vectorization of low trip count loops if there was no or a single
iteration remaining.
This causes loops to be vectorized with a VF where no scalar tail
remains, even if it is required for legality (loop with multiple
countable exits require scalar epilogue to pick the exit).
For now, limit to cases where there's a scalar iteration remaining.
PR: https://github.com/llvm/llvm-project/pull/225633
---
.../Transforms/Vectorize/LoopVectorize.cpp | 12 ++---
.../AArch64/sve-low-trip-count.ll | 44 ++++++++++++++++++-
.../RISCV/countable-early-exit-no-epilogue.ll | 24 +++++-----
3 files changed, 62 insertions(+), 18 deletions(-)
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
index 73588103e26d7..e934fec366331 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
@@ -3067,11 +3067,11 @@ LoopVectorizationCostModel::computeMaxVF(ElementCount UserVF, unsigned UserIC) {
return MaxFactors;
}
- // Allow cases where the ExactTC == (VF * IC) or ExactTC == (VF * IC) + 1.
+ // Allow cases where the ExactTC == (VF * IC) + 1.
//
- // This produces at most 1 vector iteration, and at most 1 scalar iteration
- // with no remainder. Later passes will eliminate the loop and leave
- // straight-line code as the both iteration counts are statically known.
+ // This produces 1 vector iteration, and 1 scalar iteration with no
+ // remainder. Later passes will eliminate the loop and leave straight-line
+ // code as the both iteration counts are statically known.
//
// If a function is marked as minsize/optsize or OptForSize is set, do not
// allow this form of transformation as this will increase CodeSize.
@@ -3080,7 +3080,7 @@ LoopVectorizationCostModel::computeMaxVF(ElementCount UserVF, unsigned UserIC) {
// enough to accurately determine if vectorization is beneficial.
unsigned EffectiveIC = UserIC > 0 ? UserIC : 1;
unsigned MaxVFForTC = llvm::bit_floor(TC.getFixedValue());
- if (TC.getFixedValue() - MaxVFForTC <= 1 && MaxVFForTC / EffectiveIC > 1 &&
+ if (TC.getFixedValue() - MaxVFForTC == 1 && MaxVFForTC / EffectiveIC > 1 &&
MaxVFForTC <= (MaxFactors.FixedVF.getFixedValue() * EffectiveIC) &&
!Config.OptForSize) {
unsigned NumOfInstructions = llvm::sum_of(
@@ -3090,7 +3090,7 @@ LoopVectorizationCostModel::computeMaxVF(ElementCount UserVF, unsigned UserIC) {
if (NumOfInstructions > LowTripCountLoopBodySizeLimit) {
unsigned VF = MaxVFForTC / EffectiveIC;
LLVM_DEBUG(dbgs() << "LV: Picking MaxVF=" << VF
- << " with at most 1 scalar iteration remaining.\n");
+ << " with 1 scalar iteration remaining.\n");
MaxFactors.FixedVF = ElementCount::getFixed(VF);
MaxFactors.ScalableVF = ElementCount::getScalable(0);
return MaxFactors;
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/sve-low-trip-count.ll b/llvm/test/Transforms/LoopVectorize/AArch64/sve-low-trip-count.ll
index be740c8f0fa16..24d94e5354bda 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/sve-low-trip-count.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/sve-low-trip-count.ll
@@ -617,7 +617,7 @@ exit:
define void @tc4_vf4(ptr noalias %a, ptr noalias %b) #0 {
; CHECK-LABEL: define void @tc4_vf4(
-; CHECK-SAME: ptr noalias [[A:%.*]], ptr noalias [[B:%.*]]) #[[ATTR0:[0-9]+]] {
+; CHECK-SAME: ptr noalias [[A:%.*]], ptr noalias [[B:%.*]]) #[[ATTR0]] {
; CHECK-NEXT: [[ENTRY:.*:]]
; CHECK-NEXT: br label %[[VECTOR_PH:.*]]
; CHECK: [[VECTOR_PH]]:
@@ -690,6 +690,48 @@ exit:
ret void
}
+define i32 @countable_early_exit_tc_4(ptr noalias %b) #0 {
+; CHECK-LABEL: define i32 @countable_early_exit_tc_4(
+; CHECK-SAME: ptr noalias [[B:%.*]]) #[[ATTR0]] {
+; CHECK-NEXT: [[ENTRY:.*]]:
+; CHECK-NEXT: br label %[[LOOP:.*]]
+; CHECK: [[LOOP]]:
+; CHECK-NEXT: [[IV:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IV_NEXT:%.*]], %[[LATCH:.*]] ]
+; CHECK-NEXT: [[C:%.*]] = icmp eq i64 [[IV]], 3
+; CHECK-NEXT: br i1 [[C]], label %[[EXIT1:.*]], label %[[LATCH]]
+; CHECK: [[LATCH]]:
+; CHECK-NEXT: [[GEP:%.*]] = getelementptr inbounds i32, ptr [[B]], i64 [[IV]]
+; CHECK-NEXT: store i32 1, ptr [[GEP]], align 4
+; CHECK-NEXT: [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; CHECK-NEXT: [[EC:%.*]] = icmp eq i64 [[IV_NEXT]], 100
+; CHECK-NEXT: br i1 [[EC]], label %[[EXIT2:.*]], label %[[LOOP]]
+; CHECK: [[EXIT1]]:
+; CHECK-NEXT: ret i32 1
+; CHECK: [[EXIT2]]:
+; CHECK-NEXT: ret i32 2
+;
+entry:
+ br label %loop
+
+loop:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %latch ]
+ %c = icmp eq i64 %iv, 3
+ br i1 %c, label %exit1, label %latch
+
+latch:
+ %gep = getelementptr inbounds i32, ptr %b, i64 %iv
+ store i32 1, ptr %gep, align 4
+ %iv.next = add nuw nsw i64 %iv, 1
+ %ec = icmp eq i64 %iv.next, 100
+ br i1 %ec, label %exit2, label %loop
+
+exit1:
+ ret i32 1
+
+exit2:
+ ret i32 2
+}
+
attributes #0 = { vscale_range(1,16) "target-features"="+sve" }
attributes #1 = { vscale_range(1,16) "target-features"="+sve" minsize }
attributes #2 = { vscale_range(1,16) "target-features"="+sve" optsize }
diff --git a/llvm/test/Transforms/LoopVectorize/RISCV/countable-early-exit-no-epilogue.ll b/llvm/test/Transforms/LoopVectorize/RISCV/countable-early-exit-no-epilogue.ll
index 01cc2c2f7cff0..c1e2913e5be38 100644
--- a/llvm/test/Transforms/LoopVectorize/RISCV/countable-early-exit-no-epilogue.ll
+++ b/llvm/test/Transforms/LoopVectorize/RISCV/countable-early-exit-no-epilogue.ll
@@ -3,20 +3,22 @@
; RUN: opt -passes=loop-vectorize -mtriple=riscv64 -mattr=+v -low-trip-count-loop-body-size-limit=0 -tail-folding-policy=must-fold-tail -S %s | FileCheck %s --check-prefix=NO-EPILOGUE
; RUN: opt -passes=loop-vectorize -mtriple=riscv64 -mattr=+v -vectorizer-min-trip-count=0 -tail-folding-policy=dont-fold-tail -force-vector-width=2 -S %s | FileCheck %s --check-prefix=EPILOGUE
-; FIXME: Currently this gets miscompiled, as the forced scalar epilogue is missing.
define i32 @countable_early_exit(ptr noalias %b) {
; NO-EPILOGUE-LABEL: define i32 @countable_early_exit(
; NO-EPILOGUE-SAME: ptr noalias [[B:%.*]]) #[[ATTR0:[0-9]+]] {
-; NO-EPILOGUE-NEXT: [[ENTRY:.*:]]
-; NO-EPILOGUE-NEXT: br label %[[VECTOR_PH:.*]]
-; NO-EPILOGUE: [[VECTOR_PH]]:
-; NO-EPILOGUE-NEXT: br label %[[VECTOR_BODY:.*]]
-; NO-EPILOGUE: [[VECTOR_BODY]]:
-; NO-EPILOGUE-NEXT: store <4 x i32> splat (i32 1), ptr [[B]], align 4
-; NO-EPILOGUE-NEXT: br label %[[MIDDLE_BLOCK:.*]]
-; NO-EPILOGUE: [[MIDDLE_BLOCK]]:
-; NO-EPILOGUE-NEXT: br label %[[EXIT2:.*]]
-; NO-EPILOGUE: [[EXIT1:.*:]]
+; NO-EPILOGUE-NEXT: [[ENTRY:.*]]:
+; NO-EPILOGUE-NEXT: br label %[[LOOP:.*]]
+; NO-EPILOGUE: [[LOOP]]:
+; NO-EPILOGUE-NEXT: [[IV:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IV_NEXT:%.*]], %[[LATCH:.*]] ]
+; NO-EPILOGUE-NEXT: [[C:%.*]] = icmp eq i64 [[IV]], 3
+; NO-EPILOGUE-NEXT: br i1 [[C]], label %[[EXIT1:.*]], label %[[LATCH]]
+; NO-EPILOGUE: [[LATCH]]:
+; NO-EPILOGUE-NEXT: [[GEP:%.*]] = getelementptr inbounds i32, ptr [[B]], i64 [[IV]]
+; NO-EPILOGUE-NEXT: store i32 1, ptr [[GEP]], align 4
+; NO-EPILOGUE-NEXT: [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; NO-EPILOGUE-NEXT: [[EC:%.*]] = icmp eq i64 [[IV_NEXT]], 100
+; NO-EPILOGUE-NEXT: br i1 [[EC]], label %[[EXIT2:.*]], label %[[LOOP]]
+; NO-EPILOGUE: [[EXIT1]]:
; NO-EPILOGUE-NEXT: ret i32 1
; NO-EPILOGUE: [[EXIT2]]:
; NO-EPILOGUE-NEXT: ret i32 2
>From e7b0dc7e15fb064706bfe657b6d210ed7f637ab6 Mon Sep 17 00:00:00 2001
From: Florian Hahn <flo at fhahn.com>
Date: Sat, 26 Sep 2026 23:02:42 +0100
Subject: [PATCH 27/53] [LAA] Remove unused functions (NFC) (#226733)
Remove RuntimePointerChecking::empty and LoopAccessInfo::getNumLoads /
getNumStores, which have no callers, together with the NumLoads and
NumStores counters that were only read by the latter.
---
llvm/include/llvm/Analysis/LoopAccessAnalysis.h | 9 ---------
llvm/lib/Analysis/LoopAccessAnalysis.cpp | 2 --
2 files changed, 11 deletions(-)
diff --git a/llvm/include/llvm/Analysis/LoopAccessAnalysis.h b/llvm/include/llvm/Analysis/LoopAccessAnalysis.h
index 792b8ba9cabc5..0e5f5a1e81b77 100644
--- a/llvm/include/llvm/Analysis/LoopAccessAnalysis.h
+++ b/llvm/include/llvm/Analysis/LoopAccessAnalysis.h
@@ -598,9 +598,6 @@ class RuntimePointerChecking {
unsigned ASId, PredicatedScalarEvolution &PSE,
bool NeedsFreeze);
- /// No run-time memory checking is necessary.
- bool empty() const { return Pointers.empty(); }
-
/// Generate the checks and store it. This also performs the grouping
/// of pointers to reduce the number of memchecks necessary.
LLVM_ABI void generateChecks(MemoryDepChecker::DepCandidates &DepCands);
@@ -770,9 +767,6 @@ class LoopAccessInfo {
/// Returns true if value \p V is loop invariant.
LLVM_ABI bool isInvariant(Value *V) const;
- unsigned getNumStores() const { return NumStores; }
- unsigned getNumLoads() const { return NumLoads;}
-
/// The diagnostics report generated for the analysis. E.g. why we
/// couldn't analyze the loop.
const OptimizationRemarkAnalysis *getReport() const { return Report.get(); }
@@ -874,9 +868,6 @@ class LoopAccessInfo {
/// memory accesses could be analyzed.
bool AllowPartial;
- unsigned NumLoads = 0;
- unsigned NumStores = 0;
-
/// Cache the result of analyzeLoop.
bool CanVecMem = false;
bool HasConvergentOp = false;
diff --git a/llvm/lib/Analysis/LoopAccessAnalysis.cpp b/llvm/lib/Analysis/LoopAccessAnalysis.cpp
index 409d9ceb5b812..804ffe5bb2edd 100644
--- a/llvm/lib/Analysis/LoopAccessAnalysis.cpp
+++ b/llvm/lib/Analysis/LoopAccessAnalysis.cpp
@@ -2789,7 +2789,6 @@ bool LoopAccessInfo::analyzeLoop(AAResults *AA, const LoopInfo *LI,
HasComplexMemInst = true;
continue;
}
- NumLoads++;
Loads.push_back(Ld);
DepChecker->addAccess(Ld);
if (EnableMemAccessVersioningOfLoop)
@@ -2813,7 +2812,6 @@ bool LoopAccessInfo::analyzeLoop(AAResults *AA, const LoopInfo *LI,
HasComplexMemInst = true;
continue;
}
- NumStores++;
Stores.push_back(St);
DepChecker->addAccess(St);
if (EnableMemAccessVersioningOfLoop)
>From 9160c39339dd951272b4b13176003c031dc392b9 Mon Sep 17 00:00:00 2001
From: Florian Hahn <flo at fhahn.com>
Date: Sat, 26 Sep 2026 23:31:45 +0100
Subject: [PATCH 28/53] [SCEV] Remove unused classof(const SCEVUse *) overloads
(NFC) (#226734)
isa/cast/dyn_cast on SCEVUse go through simplify_type or CastInfo.
classof(SCEVUse) is unused.
---
llvm/include/llvm/Analysis/ScalarEvolutionExpressions.h | 4 ----
1 file changed, 4 deletions(-)
diff --git a/llvm/include/llvm/Analysis/ScalarEvolutionExpressions.h b/llvm/include/llvm/Analysis/ScalarEvolutionExpressions.h
index 12f201f3ae50e..feb62d73079bb 100644
--- a/llvm/include/llvm/Analysis/ScalarEvolutionExpressions.h
+++ b/llvm/include/llvm/Analysis/ScalarEvolutionExpressions.h
@@ -228,7 +228,6 @@ class SCEVNAryExpr : public SCEV {
S->getSCEVType() == scSequentialUMinExpr ||
S->getSCEVType() == scAddRecExpr;
}
- static bool classof(const SCEVUse *U) { return classof(U->getPointer()); }
};
/// This node is the base class for n'ary commutative operators.
@@ -273,7 +272,6 @@ class SCEVAddExpr : public SCEVCommutativeExpr {
public:
/// Methods for support type inquiry through isa, cast, and dyn_cast:
static bool classof(const SCEV *S) { return S->getSCEVType() == scAddExpr; }
- static bool classof(const SCEVUse *U) { return classof(U->getPointer()); }
};
/// This node represents multiplication of some number of SCEVs.
@@ -286,7 +284,6 @@ class SCEVMulExpr : public SCEVCommutativeExpr {
public:
/// Methods for support type inquiry through isa, cast, and dyn_cast:
static bool classof(const SCEV *S) { return S->getSCEVType() == scMulExpr; }
- static bool classof(const SCEVUse *U) { return classof(U->getPointer()); }
};
/// This class represents a binary unsigned division operation.
@@ -539,7 +536,6 @@ class SCEVSequentialMinMaxExpr : public SCEVNAryExpr {
static bool classof(const SCEV *S) {
return isSequentialMinMaxType(S->getSCEVType());
}
- static bool classof(const SCEVUse *U) { return classof(U->getPointer()); }
};
/// This class represents a sequential/in-order unsigned minimum selection.
>From 23b06a3cf6947d4764c79de79258de97f7f020af Mon Sep 17 00:00:00 2001
From: Dan Salvato <dan at teamsalvato.com>
Date: Sat, 26 Sep 2026 18:40:51 -0400
Subject: [PATCH 29/53] [M68k] Finish implementation of `MOVX` (move and
extend) pseudo-instructions and fix related errors (#218938)
This patch adds remaining addressing modes to the "move and extend"
pseudo-instructions, and fixes a few errors and inconsistencies in the
logic that caused inefficient code generation.
- Remaining addressing modes were implemented to match non-pseudo `MOVE`
instructions.
- Names of the pseudos now correctly reflect the register classes they
operate on, e.g. `MOVZXd32r16` for XR16 -> DR32.
- Fixed an issue where `MOVZX` could be allocated to an address
register, which has no way to zero-extend the result. (There were even
some of these in the test `register-spills.ll`, emitted as illegal
instructions, but the test doesn't have instruction verification
enabled, so it wasn't caught.)
- Fixed some `extload` patterns where the register was unnecessarily
extended to 32 bits before being truncated to its final size, causing
redundant instructions to be omitted.
- Fixed `anyext` patterns lowering to `MOVZX` instead of `MOVX`, causing
unnecessary zero-extensions.
- An inconsistency was fixed where the expansion logic would perform a
source-sized move when loading from memory, but a destination-sized move
when moving between registers. They now always perform a source-sized
move, which improves clarity in the assembly output and has more
opportunity for future optimizations.
---
llvm/lib/Target/M68k/M68kExpandPseudo.cpp | 237 ++++++++++++------
llvm/lib/Target/M68k/M68kInstrArithmetic.td | 30 +--
llvm/lib/Target/M68k/M68kInstrData.td | 169 +++++--------
llvm/lib/Target/M68k/M68kInstrInfo.cpp | 120 +++++----
llvm/lib/Target/M68k/M68kInstrInfo.h | 8 +-
.../CodeGen/M68k/Arith/divide-by-constant.ll | 6 +-
.../CodeGen/M68k/Arith/smul-with-overflow.ll | 10 +-
.../CodeGen/M68k/Arith/umul-with-overflow.ll | 10 +-
llvm/test/CodeGen/M68k/Bits/btst.ll | 3 -
llvm/test/CodeGen/M68k/Control/cmp.ll | 7 -
.../CodeGen/M68k/Control/non-cmov-switch.ll | 4 -
llvm/test/CodeGen/M68k/Control/setcc.ll | 3 -
llvm/test/CodeGen/M68k/Data/load-extend.ll | 4 -
llvm/test/CodeGen/M68k/Data/sext-i1.ll | 2 -
llvm/test/CodeGen/M68k/register-spills.ll | 81 ------
15 files changed, 322 insertions(+), 372 deletions(-)
diff --git a/llvm/lib/Target/M68k/M68kExpandPseudo.cpp b/llvm/lib/Target/M68k/M68kExpandPseudo.cpp
index 39e6eeb912a6e..39e2e5c5aa919 100644
--- a/llvm/lib/Target/M68k/M68kExpandPseudo.cpp
+++ b/llvm/lib/Target/M68k/M68kExpandPseudo.cpp
@@ -86,106 +86,179 @@ bool M68kExpandPseudo::ExpandMI(MachineBasicBlock &MBB,
case M68k::MOVI32ri:
return TII->ExpandMOVI(MIB, MVT::i32);
- case M68k::MOVXd16d8:
- return TII->ExpandMOVX_RR(MIB, MVT::i16, MVT::i8);
- case M68k::MOVXd32d8:
- return TII->ExpandMOVX_RR(MIB, MVT::i32, MVT::i8);
- case M68k::MOVXd32d16:
- return TII->ExpandMOVX_RR(MIB, MVT::i32, MVT::i16);
-
case M68k::MOVSXd16d8:
return TII->ExpandMOVSZX_RR(MIB, true, MVT::i16, MVT::i8);
case M68k::MOVSXd32d8:
return TII->ExpandMOVSZX_RR(MIB, true, MVT::i32, MVT::i8);
- case M68k::MOVSXd32d16:
+ case M68k::MOVSXr32r16:
return TII->ExpandMOVSZX_RR(MIB, true, MVT::i32, MVT::i16);
case M68k::MOVZXd16d8:
return TII->ExpandMOVSZX_RR(MIB, false, MVT::i16, MVT::i8);
case M68k::MOVZXd32d8:
return TII->ExpandMOVSZX_RR(MIB, false, MVT::i32, MVT::i8);
- case M68k::MOVZXd32d16:
+ case M68k::MOVZXd32r16:
return TII->ExpandMOVSZX_RR(MIB, false, MVT::i32, MVT::i16);
- case M68k::MOVSXd16j8:
- return TII->ExpandMOVSZX_RM(MIB, true, TII->get(M68k::MOV8dj), MVT::i16,
- MVT::i8);
- case M68k::MOVSXd32j8:
- return TII->ExpandMOVSZX_RM(MIB, true, TII->get(M68k::MOV8dj), MVT::i32,
- MVT::i8);
- case M68k::MOVSXd32j16:
- return TII->ExpandMOVSZX_RM(MIB, true, TII->get(M68k::MOV16rj), MVT::i32,
- MVT::i16);
-
- case M68k::MOVZXd16j8:
- return TII->ExpandMOVSZX_RM(MIB, false, TII->get(M68k::MOV8dj), MVT::i16,
- MVT::i8);
- case M68k::MOVZXd32j8:
- return TII->ExpandMOVSZX_RM(MIB, false, TII->get(M68k::MOV8dj), MVT::i32,
- MVT::i8);
- case M68k::MOVZXd32j16:
- return TII->ExpandMOVSZX_RM(MIB, false, TII->get(M68k::MOV16rj), MVT::i32,
- MVT::i16);
-
- case M68k::MOVSXd16p8:
- return TII->ExpandMOVSZX_RM(MIB, true, TII->get(M68k::MOV8dp), MVT::i16,
- MVT::i8);
- case M68k::MOVSXd32p8:
- return TII->ExpandMOVSZX_RM(MIB, true, TII->get(M68k::MOV8dp), MVT::i32,
- MVT::i8);
- case M68k::MOVSXd32p16:
- return TII->ExpandMOVSZX_RM(MIB, true, TII->get(M68k::MOV16rp), MVT::i32,
- MVT::i16);
-
- case M68k::MOVZXd16p8:
- return TII->ExpandMOVSZX_RM(MIB, false, TII->get(M68k::MOV8dp), MVT::i16,
- MVT::i8);
- case M68k::MOVZXd32p8:
- return TII->ExpandMOVSZX_RM(MIB, false, TII->get(M68k::MOV8dp), MVT::i32,
- MVT::i8);
- case M68k::MOVZXd32p16:
- return TII->ExpandMOVSZX_RM(MIB, false, TII->get(M68k::MOV16rp), MVT::i32,
- MVT::i16);
-
- case M68k::MOVSXd16f8:
- return TII->ExpandMOVSZX_RM(MIB, true, TII->get(M68k::MOV8df), MVT::i16,
- MVT::i8);
- case M68k::MOVSXd32f8:
- return TII->ExpandMOVSZX_RM(MIB, true, TII->get(M68k::MOV8df), MVT::i32,
- MVT::i8);
- case M68k::MOVSXd32f16:
- return TII->ExpandMOVSZX_RM(MIB, true, TII->get(M68k::MOV16rf), MVT::i32,
- MVT::i16);
-
- case M68k::MOVZXd16f8:
- return TII->ExpandMOVSZX_RM(MIB, false, TII->get(M68k::MOV8df), MVT::i16,
- MVT::i8);
- case M68k::MOVZXd32f8:
- return TII->ExpandMOVSZX_RM(MIB, false, TII->get(M68k::MOV8df), MVT::i32,
- MVT::i8);
- case M68k::MOVZXd32f16:
- return TII->ExpandMOVSZX_RM(MIB, false, TII->get(M68k::MOV16rf), MVT::i32,
- MVT::i16);
+ case M68k::MOVXd16d8:
+ return TII->ExpandMOVX_RR(MIB, MVT::i16, MVT::i8);
+ case M68k::MOVXd32d8:
+ return TII->ExpandMOVX_RR(MIB, MVT::i32, MVT::i8);
+ case M68k::MOVXr32r16:
+ return TII->ExpandMOVX_RR(MIB, MVT::i32, MVT::i16);
+ case M68k::MOVSXd16o8:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8do, MVT::i16, MVT::i8);
+ case M68k::MOVSXd16e8:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8de, MVT::i16, MVT::i8);
+ case M68k::MOVSXd16k8:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8dk, MVT::i16, MVT::i8);
case M68k::MOVSXd16q8:
- return TII->ExpandMOVSZX_RM(MIB, true, TII->get(M68k::MOV8dq), MVT::i16,
- MVT::i8);
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8dq, MVT::i16, MVT::i8);
+ case M68k::MOVSXd16f8:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8df, MVT::i16, MVT::i8);
+ case M68k::MOVSXd16p8:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8dp, MVT::i16, MVT::i8);
+ case M68k::MOVSXd16b8:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8db, MVT::i16, MVT::i8);
+ case M68k::MOVSXd16j8:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8dj, MVT::i16, MVT::i8);
+
+ case M68k::MOVSXd32o8:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8do, MVT::i32, MVT::i8);
+ case M68k::MOVSXd32e8:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8de, MVT::i32, MVT::i8);
+ case M68k::MOVSXd32k8:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8dk, MVT::i32, MVT::i8);
case M68k::MOVSXd32q8:
- return TII->ExpandMOVSZX_RM(MIB, true, TII->get(M68k::MOV8dq), MVT::i32,
- MVT::i8);
- case M68k::MOVSXd32q16:
- return TII->ExpandMOVSZX_RM(MIB, true, TII->get(M68k::MOV16dq), MVT::i32,
- MVT::i16);
-
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8dq, MVT::i32, MVT::i8);
+ case M68k::MOVSXd32f8:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8df, MVT::i32, MVT::i8);
+ case M68k::MOVSXd32p8:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8dp, MVT::i32, MVT::i8);
+ case M68k::MOVSXd32b8:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8db, MVT::i32, MVT::i8);
+ case M68k::MOVSXd32j8:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV8dj, MVT::i32, MVT::i8);
+
+ case M68k::MOVSXr32o16:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV16ro, MVT::i32, MVT::i16);
+ case M68k::MOVSXr32e16:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV16re, MVT::i32, MVT::i16);
+ case M68k::MOVSXr32k16:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV16rk, MVT::i32, MVT::i16);
+ case M68k::MOVSXr32q16:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV16rq, MVT::i32, MVT::i16);
+ case M68k::MOVSXr32f16:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV16rf, MVT::i32, MVT::i16);
+ case M68k::MOVSXr32p16:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV16rp, MVT::i32, MVT::i16);
+ case M68k::MOVSXr32b16:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV16rb, MVT::i32, MVT::i16);
+ case M68k::MOVSXr32j16:
+ return TII->ExpandMOVSZX_RM(MIB, true, M68k::MOV16rj, MVT::i32, MVT::i16);
+
+ case M68k::MOVZXd16o8:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8do, MVT::i16, MVT::i8);
+ case M68k::MOVZXd16e8:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8de, MVT::i16, MVT::i8);
+ case M68k::MOVZXd16k8:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8dk, MVT::i16, MVT::i8);
case M68k::MOVZXd16q8:
- return TII->ExpandMOVSZX_RM(MIB, false, TII->get(M68k::MOV8dq), MVT::i16,
- MVT::i8);
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8dq, MVT::i16, MVT::i8);
+ case M68k::MOVZXd16f8:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8df, MVT::i16, MVT::i8);
+ case M68k::MOVZXd16p8:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8dp, MVT::i16, MVT::i8);
+ case M68k::MOVZXd16b8:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8db, MVT::i16, MVT::i8);
+ case M68k::MOVZXd16j8:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8dj, MVT::i16, MVT::i8);
+
+ case M68k::MOVZXd32o8:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8do, MVT::i32, MVT::i8);
+ case M68k::MOVZXd32e8:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8de, MVT::i32, MVT::i8);
+ case M68k::MOVZXd32k8:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8dk, MVT::i32, MVT::i8);
case M68k::MOVZXd32q8:
- return TII->ExpandMOVSZX_RM(MIB, false, TII->get(M68k::MOV8dq), MVT::i32,
- MVT::i8);
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8dq, MVT::i32, MVT::i8);
+ case M68k::MOVZXd32f8:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8df, MVT::i32, MVT::i8);
+ case M68k::MOVZXd32p8:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8dp, MVT::i32, MVT::i8);
+ case M68k::MOVZXd32b8:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8db, MVT::i32, MVT::i8);
+ case M68k::MOVZXd32j8:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV8dj, MVT::i32, MVT::i8);
+
+ case M68k::MOVZXd32o16:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV16ro, MVT::i32, MVT::i16);
+ case M68k::MOVZXd32e16:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV16re, MVT::i32, MVT::i16);
+ case M68k::MOVZXd32k16:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV16rk, MVT::i32, MVT::i16);
case M68k::MOVZXd32q16:
- return TII->ExpandMOVSZX_RM(MIB, false, TII->get(M68k::MOV16dq), MVT::i32,
- MVT::i16);
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV16rq, MVT::i32, MVT::i16);
+ case M68k::MOVZXd32f16:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV16rf, MVT::i32, MVT::i16);
+ case M68k::MOVZXd32p16:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV16rp, MVT::i32, MVT::i16);
+ case M68k::MOVZXd32b16:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV16rb, MVT::i32, MVT::i16);
+ case M68k::MOVZXd32j16:
+ return TII->ExpandMOVSZX_RM(MIB, false, M68k::MOV16rj, MVT::i32, MVT::i16);
+
+ case M68k::MOVXd16o8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8do, MVT::i16, MVT::i8);
+ case M68k::MOVXd16e8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8de, MVT::i16, MVT::i8);
+ case M68k::MOVXd16k8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8dk, MVT::i16, MVT::i8);
+ case M68k::MOVXd16q8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8dq, MVT::i16, MVT::i8);
+ case M68k::MOVXd16f8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8df, MVT::i16, MVT::i8);
+ case M68k::MOVXd16p8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8dp, MVT::i16, MVT::i8);
+ case M68k::MOVXd16b8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8db, MVT::i16, MVT::i8);
+ case M68k::MOVXd16j8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8dj, MVT::i16, MVT::i8);
+
+ case M68k::MOVXd32o8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8do, MVT::i32, MVT::i8);
+ case M68k::MOVXd32e8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8de, MVT::i32, MVT::i8);
+ case M68k::MOVXd32k8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8dk, MVT::i32, MVT::i8);
+ case M68k::MOVXd32q8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8dq, MVT::i32, MVT::i8);
+ case M68k::MOVXd32f8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8df, MVT::i32, MVT::i8);
+ case M68k::MOVXd32p8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8dp, MVT::i32, MVT::i8);
+ case M68k::MOVXd32b8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8db, MVT::i32, MVT::i8);
+ case M68k::MOVXd32j8:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV8dj, MVT::i32, MVT::i8);
+
+ case M68k::MOVXr32o16:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV16ro, MVT::i32, MVT::i16);
+ case M68k::MOVXr32e16:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV16re, MVT::i32, MVT::i16);
+ case M68k::MOVXr32k16:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV16rk, MVT::i32, MVT::i16);
+ case M68k::MOVXr32q16:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV16rq, MVT::i32, MVT::i16);
+ case M68k::MOVXr32f16:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV16rf, MVT::i32, MVT::i16);
+ case M68k::MOVXr32p16:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV16rp, MVT::i32, MVT::i16);
+ case M68k::MOVXr32b16:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV16rb, MVT::i32, MVT::i16);
+ case M68k::MOVXr32j16:
+ return TII->ExpandMOVX_RM(MIB, M68k::MOV16rj, MVT::i32, MVT::i16);
case M68k::MOVM8jm_P:
case M68k::MOVM16jm_P:
diff --git a/llvm/lib/Target/M68k/M68kInstrArithmetic.td b/llvm/lib/Target/M68k/M68kInstrArithmetic.td
index e448cd5b031b3..b9e92dde3d0ad 100644
--- a/llvm/lib/Target/M68k/M68kInstrArithmetic.td
+++ b/llvm/lib/Target/M68k/M68kInstrArithmetic.td
@@ -561,7 +561,7 @@ def EXT32 : MxExt<MxType32d, MxType16d>;
def : Pat<(sext_inreg i16:$src, i8), (EXT16 $src)>;
def : Pat<(sext_inreg i32:$src, i16), (EXT32 $src)>;
def : Pat<(sext_inreg i32:$src, i8),
- (EXT32 (MOVXd32d16 (EXT16 (EXTRACT_SUBREG $src, MxSubRegIndex16Lo))))>;
+ (EXT32 (MOVXr32r16 (EXT16 (EXTRACT_SUBREG $src, MxSubRegIndex16Lo))))>;
//===----------------------------------------------------------------------===//
@@ -669,22 +669,22 @@ def : Pat<(urem i8:$dst, i8:$opd),
// RR i16
def : Pat<(sdiv i16:$dst, i16:$opd),
(EXTRACT_SUBREG
- (SDIVd32d16 (MOVSXd32d16 $dst), $opd),
+ (SDIVd32d16 (MOVSXr32r16 $dst), $opd),
MxSubRegIndex16Lo)>;
def : Pat<(udiv i16:$dst, i16:$opd),
(EXTRACT_SUBREG
- (UDIVd32d16 (MOVZXd32d16 $dst), $opd),
+ (UDIVd32d16 (MOVZXd32r16 $dst), $opd),
MxSubRegIndex16Lo)>;
def : Pat<(srem i16:$dst, i16:$opd),
(EXTRACT_SUBREG
- (SWAP (SDIVd32d16 (MOVSXd32d16 $dst), $opd)),
+ (SWAP (SDIVd32d16 (MOVSXr32r16 $dst), $opd)),
MxSubRegIndex16Lo)>;
def : Pat<(urem i16:$dst, i16:$opd),
(EXTRACT_SUBREG
- (SWAP (UDIVd32d16 (MOVZXd32d16 $dst), $opd)),
+ (SWAP (UDIVd32d16 (MOVZXd32r16 $dst), $opd)),
MxSubRegIndex16Lo)>;
// RI i8
@@ -711,22 +711,22 @@ def : Pat<(urem i8:$dst, Mxi8immSExt8:$opd),
// RI i16
def : Pat<(sdiv i16:$dst, Mxi16immSExt16:$opd),
(EXTRACT_SUBREG
- (SDIVd32i16 (MOVSXd32d16 $dst), imm:$opd),
+ (SDIVd32i16 (MOVSXr32r16 $dst), imm:$opd),
MxSubRegIndex16Lo)>;
def : Pat<(udiv i16:$dst, Mxi16immSExt16:$opd),
(EXTRACT_SUBREG
- (UDIVd32i16 (MOVZXd32d16 $dst), imm:$opd),
+ (UDIVd32i16 (MOVZXd32r16 $dst), imm:$opd),
MxSubRegIndex16Lo)>;
def : Pat<(srem i16:$dst, Mxi16immSExt16:$opd),
(EXTRACT_SUBREG
- (SWAP (SDIVd32i16 (MOVSXd32d16 $dst), imm:$opd)),
+ (SWAP (SDIVd32i16 (MOVSXr32r16 $dst), imm:$opd)),
MxSubRegIndex16Lo)>;
def : Pat<(urem i16:$dst, Mxi16immSExt16:$opd),
(EXTRACT_SUBREG
- (SWAP (UDIVd32i16 (MOVZXd32d16 $dst), imm:$opd)),
+ (SWAP (UDIVd32i16 (MOVZXd32r16 $dst), imm:$opd)),
MxSubRegIndex16Lo)>;
@@ -738,17 +738,17 @@ def UMULd32d32 : MxDiMuOp_DD_Long<"mulu.l", MxUMul, 0x130, /*SIGNED*/false>;
// RR
def : Pat<(mul i16:$dst, i16:$opd),
(EXTRACT_SUBREG
- (SMULd32d16 (MOVXd32d16 $dst), $opd),
+ (SMULd32d16 (MOVXr32r16 $dst), $opd),
MxSubRegIndex16Lo)>;
def : Pat<(mulhs i16:$dst, i16:$opd),
(EXTRACT_SUBREG
- (SWAP (SMULd32d16 (MOVXd32d16 $dst), $opd)),
+ (SWAP (SMULd32d16 (MOVXr32r16 $dst), $opd)),
MxSubRegIndex16Lo)>;
def : Pat<(mulhu i16:$dst, i16:$opd),
(EXTRACT_SUBREG
- (SWAP (UMULd32d16 (MOVXd32d16 $dst), $opd)),
+ (SWAP (UMULd32d16 (MOVXr32r16 $dst), $opd)),
MxSubRegIndex16Lo)>;
def : Pat<(mul i32:$dst, i32:$opd), (SMULd32d32 $dst, $opd)>;
@@ -757,17 +757,17 @@ def : Pat<(mul i32:$dst, i32:$opd), (SMULd32d32 $dst, $opd)>;
// RI
def : Pat<(mul i16:$dst, Mxi16immSExt16:$opd),
(EXTRACT_SUBREG
- (SMULd32i16 (MOVXd32d16 $dst), imm:$opd),
+ (SMULd32i16 (MOVXr32r16 $dst), imm:$opd),
MxSubRegIndex16Lo)>;
def : Pat<(mulhs i16:$dst, Mxi16immSExt16:$opd),
(EXTRACT_SUBREG
- (SWAP (SMULd32i16 (MOVXd32d16 $dst), imm:$opd)),
+ (SWAP (SMULd32i16 (MOVXr32r16 $dst), imm:$opd)),
MxSubRegIndex16Lo)>;
def : Pat<(mulhu i16:$dst, Mxi16immSExt16:$opd),
(EXTRACT_SUBREG
- (SWAP (UMULd32i16 (MOVXd32d16 $dst), imm:$opd)),
+ (SWAP (UMULd32i16 (MOVXr32r16 $dst), imm:$opd)),
MxSubRegIndex16Lo)>;
diff --git a/llvm/lib/Target/M68k/M68kInstrData.td b/llvm/lib/Target/M68k/M68kInstrData.td
index 5a294e1fb3cda..7bf30e05d5d95 100644
--- a/llvm/lib/Target/M68k/M68kInstrData.td
+++ b/llvm/lib/Target/M68k/M68kInstrData.td
@@ -625,130 +625,85 @@ def MOVI32ri : MxPseudoMove_DI<MxType32r>;
/// what registers are allocated for the operands and if they overlap we just
/// extend the value if the registers are completely different we need to move
/// first.
-foreach EXT = ["S", "Z"] in {
- let hasSideEffects = 0 in {
-
- def MOV#EXT#Xd16d8 : MxPseudoMove_RR<MxType16d, MxType8d>;
- def MOV#EXT#Xd32d8 : MxPseudoMove_RR<MxType32d, MxType8d>;
- def MOV#EXT#Xd32d16 : MxPseudoMove_RR<MxType32r, MxType16r>;
-
+/// The MOVX group is similar to the others but does NOT do any value extension,
+/// they just load a smaller register into the lower part of another register
+/// if operands' real registers are different or does nothing if they are the
+/// same.
+let hasSideEffects = 0 in {
+ foreach AM = MxMoveSrcAMs in {
+ defvar OpB8 = !cast<MxOpBundle>("MxOp8AddrMode_"#AM);
+ defvar OpB16 = !cast<MxOpBundle>("MxOp16AddrMode_"#AM);
let mayLoad = 1 in {
-
- def MOV#EXT#Xd16j8 : MxPseudoMove_RM<MxType16d, MxType8.JOp>;
- def MOV#EXT#Xd32j8 : MxPseudoMove_RM<MxType32d, MxType8.JOp>;
- def MOV#EXT#Xd32j16 : MxPseudoMove_RM<MxType32d, MxType16.JOp>;
-
- def MOV#EXT#Xd16p8 : MxPseudoMove_RM<MxType16d, MxType8.POp>;
- def MOV#EXT#Xd32p8 : MxPseudoMove_RM<MxType32d, MxType8.POp>;
- def MOV#EXT#Xd32p16 : MxPseudoMove_RM<MxType32d, MxType16.POp>;
-
- def MOV#EXT#Xd16f8 : MxPseudoMove_RM<MxType16d, MxType8.FOp>;
- def MOV#EXT#Xd32f8 : MxPseudoMove_RM<MxType32d, MxType8.FOp>;
- def MOV#EXT#Xd32f16 : MxPseudoMove_RM<MxType32d, MxType16.FOp>;
-
- def MOV#EXT#Xd16q8 : MxPseudoMove_RM<MxType16d, MxType8.QOp>;
- def MOV#EXT#Xd32q8 : MxPseudoMove_RM<MxType32d, MxType8.QOp>;
- def MOV#EXT#Xd32q16 : MxPseudoMove_RM<MxType32d, MxType16.QOp>;
-
+ def MOVSXd16#AM#8 : MxPseudoMove_RM<MxType16d, OpB8.Op>;
+ def MOVSXd32#AM#8 : MxPseudoMove_RM<MxType32d, OpB8.Op>;
+ def MOVSXr32#AM#16 : MxPseudoMove_RM<MxType32r, OpB16.Op>;
+ def MOVZXd16#AM#8 : MxPseudoMove_RM<MxType16d, OpB8.Op>;
+ def MOVZXd32#AM#8 : MxPseudoMove_RM<MxType32d, OpB8.Op>;
+ def MOVZXd32#AM#16 : MxPseudoMove_RM<MxType32d, OpB16.Op>;
+ def MOVXd16#AM#8 : MxPseudoMove_RM<MxType16d, OpB8.Op>;
+ def MOVXd32#AM#8 : MxPseudoMove_RM<MxType32d, OpB8.Op>;
+ def MOVXr32#AM#16 : MxPseudoMove_RM<MxType32r, OpB16.Op>;
}
}
-}
-/// This group of instructions is similar to the group above but DOES NOT do
-/// any value extension, they just load a smaller register into the lower part
-/// of another register if operands' real registers are different or does
-/// nothing if they are the same.
-def MOVXd16d8 : MxPseudoMove_RR<MxType16d, MxType8d>;
-def MOVXd32d8 : MxPseudoMove_RR<MxType32d, MxType8d>;
-def MOVXd32d16 : MxPseudoMove_RR<MxType32r, MxType16r>;
+ def MOVSXd16d8 : MxPseudoMove_RR<MxType16d, MxType8d>;
+ def MOVSXd32d8 : MxPseudoMove_RR<MxType32d, MxType8d>;
+ def MOVSXr32r16 : MxPseudoMove_RR<MxType32r, MxType16r>;
+ def MOVZXd16d8 : MxPseudoMove_RR<MxType16d, MxType8d>;
+ def MOVZXd32d8 : MxPseudoMove_RR<MxType32d, MxType8d>;
+ def MOVZXd32r16 : MxPseudoMove_RR<MxType32d, MxType16r>;
+ // TODO: MOVX d8 -> r16/32 is possible, but it causes regalloc to make bad
+ // decisions in handling 8-bit register spills. Test in register-spills.ll
+ // for future experimentation.
+ def MOVXd16d8 : MxPseudoMove_RR<MxType16d, MxType8d>;
+ def MOVXd32d8 : MxPseudoMove_RR<MxType32d, MxType8d>;
+ def MOVXr32r16 : MxPseudoMove_RR<MxType32r, MxType16r>;
+}
//===----------------------------------------------------------------------===//
// Extend/Truncate Patterns
//===----------------------------------------------------------------------===//
-// i16 <- sext i8
-def: Pat<(i16 (sext i8:$src)),
- (EXTRACT_SUBREG (MOVSXd32d8 MxDRD8:$src), MxSubRegIndex16Lo)>;
-def: Pat<(MxSExtLoadi16i8 MxCP_ARI:$src),
- (EXTRACT_SUBREG (MOVSXd32j8 MxARI8:$src), MxSubRegIndex16Lo)>;
-def: Pat<(MxSExtLoadi16i8 MxCP_ARID:$src),
- (EXTRACT_SUBREG (MOVSXd32p8 MxARID8:$src), MxSubRegIndex16Lo)>;
-def: Pat<(MxSExtLoadi16i8 MxCP_ARII:$src),
- (EXTRACT_SUBREG (MOVSXd32f8 MxARII8:$src), MxSubRegIndex16Lo)>;
-def: Pat<(MxSExtLoadi16i8 MxCP_PCD:$src), (MOVSXd16q8 MxPCD8:$src)>;
-
-// i32 <- sext i8
+foreach AM = MxMoveSrcAMs in {
+ defvar OpB8 = !cast<MxOpBundle>("MxOp8AddrMode_"#AM);
+ defvar OpB16 = !cast<MxOpBundle>("MxOp16AddrMode_"#AM);
+ def : Pat<(MxSExtLoadi16i8 OpB8.Pat:$src),
+ (!cast<MxInst>("MOVSXd16"#AM#"8") OpB8.Op:$src)>;
+ def : Pat<(MxSExtLoadi32i8 OpB8.Pat:$src),
+ (!cast<MxInst>("MOVSXd32"#AM#"8") OpB8.Op:$src)>;
+ def : Pat<(MxSExtLoadi32i16 OpB16.Pat:$src),
+ (!cast<MxInst>("MOVSXr32"#AM#"16") OpB16.Op:$src)>;
+ def : Pat<(MxZExtLoadi16i8 OpB8.Pat:$src),
+ (!cast<MxInst>("MOVZXd16"#AM#"8") OpB8.Op:$src)>;
+ def : Pat<(MxZExtLoadi32i8 OpB8.Pat:$src),
+ (!cast<MxInst>("MOVZXd32"#AM#"8") OpB8.Op:$src)>;
+ def : Pat<(MxZExtLoadi32i16 OpB16.Pat:$src),
+ (!cast<MxInst>("MOVZXd32"#AM#"16") OpB16.Op:$src)>;
+ def : Pat<(MxExtLoadi16i8 OpB8.Pat:$src),
+ (!cast<MxInst>("MOVXd16"#AM#"8") OpB8.Op:$src)>;
+ def : Pat<(MxExtLoadi32i8 OpB8.Pat:$src),
+ (!cast<MxInst>("MOVXd32"#AM#"8") OpB8.Op:$src)>;
+ def : Pat<(MxExtLoadi32i16 OpB16.Pat:$src),
+ (!cast<MxInst>("MOVXr32"#AM#"16") OpB16.Op:$src)>;
+}
+
+def: Pat<(i16 (sext i8:$src)), (MOVSXd16d8 MxDRD8:$src)>;
def: Pat<(i32 (sext i8:$src)), (MOVSXd32d8 MxDRD8:$src)>;
-def: Pat<(MxSExtLoadi32i8 MxCP_ARI :$src), (MOVSXd32j8 MxARI8 :$src)>;
-def: Pat<(MxSExtLoadi32i8 MxCP_ARID:$src), (MOVSXd32p8 MxARID8:$src)>;
-def: Pat<(MxSExtLoadi32i8 MxCP_ARII:$src), (MOVSXd32f8 MxARII8:$src)>;
-def: Pat<(MxSExtLoadi32i8 MxCP_PCD:$src), (MOVSXd32q8 MxPCD8:$src)>;
-
-// i32 <- sext i16
-def: Pat<(i32 (sext i16:$src)), (MOVSXd32d16 MxDRD16:$src)>;
-def: Pat<(MxSExtLoadi32i16 MxCP_ARI :$src), (MOVSXd32j16 MxARI16 :$src)>;
-def: Pat<(MxSExtLoadi32i16 MxCP_ARID:$src), (MOVSXd32p16 MxARID16:$src)>;
-def: Pat<(MxSExtLoadi32i16 MxCP_ARII:$src), (MOVSXd32f16 MxARII16:$src)>;
-def: Pat<(MxSExtLoadi32i16 MxCP_PCD:$src), (MOVSXd32q16 MxPCD16:$src)>;
-
-// i16 <- zext i8
-def: Pat<(i16 (zext i8:$src)),
- (EXTRACT_SUBREG (MOVZXd32d8 MxDRD8:$src), MxSubRegIndex16Lo)>;
-def: Pat<(MxZExtLoadi16i8 MxCP_ARI:$src),
- (EXTRACT_SUBREG (MOVZXd32j8 MxARI8:$src), MxSubRegIndex16Lo)>;
-def: Pat<(MxZExtLoadi16i8 MxCP_ARID:$src),
- (EXTRACT_SUBREG (MOVZXd32p8 MxARID8:$src), MxSubRegIndex16Lo)>;
-def: Pat<(MxZExtLoadi16i8 MxCP_ARII:$src),
- (EXTRACT_SUBREG (MOVZXd32f8 MxARII8:$src), MxSubRegIndex16Lo)>;
-def: Pat<(MxZExtLoadi16i8 MxCP_PCD :$src), (MOVZXd16q8 MxPCD8 :$src)>;
-
-// i32 <- zext i8
+def: Pat<(i32 (sext i16:$src)), (MOVSXr32r16 MxXRD16:$src)>;
+def: Pat<(i16 (zext i8:$src)), (MOVZXd16d8 MxDRD8:$src)>;
def: Pat<(i32 (zext i8:$src)), (MOVZXd32d8 MxDRD8:$src)>;
-def: Pat<(MxZExtLoadi32i8 MxCP_ARI :$src), (MOVZXd32j8 MxARI8 :$src)>;
-def: Pat<(MxZExtLoadi32i8 MxCP_ARID:$src), (MOVZXd32p8 MxARID8:$src)>;
-def: Pat<(MxZExtLoadi32i8 MxCP_ARII:$src), (MOVZXd32f8 MxARII8:$src)>;
-def: Pat<(MxZExtLoadi32i8 MxCP_PCD :$src), (MOVZXd32q8 MxPCD8 :$src)>;
-
-// i32 <- zext i16
-def: Pat<(i32 (zext i16:$src)), (MOVZXd32d16 MxDRD16:$src)>;
-def: Pat<(MxZExtLoadi32i16 MxCP_ARI :$src), (MOVZXd32j16 MxARI16 :$src)>;
-def: Pat<(MxZExtLoadi32i16 MxCP_ARID:$src), (MOVZXd32p16 MxARID16:$src)>;
-def: Pat<(MxZExtLoadi32i16 MxCP_ARII:$src), (MOVZXd32f16 MxARII16:$src)>;
-def: Pat<(MxZExtLoadi32i16 MxCP_PCD :$src), (MOVZXd32q16 MxPCD16 :$src)>;
-
-// i16 <- anyext i8
-def: Pat<(i16 (anyext i8:$src)),
- (EXTRACT_SUBREG (MOVZXd32d8 MxDRD8:$src), MxSubRegIndex16Lo)>;
-def: Pat<(MxExtLoadi16i8 MxCP_ARI:$src),
- (EXTRACT_SUBREG (MOVZXd32j8 MxARI8:$src), MxSubRegIndex16Lo)>;
-def: Pat<(MxExtLoadi16i8 MxCP_ARID:$src),
- (EXTRACT_SUBREG (MOVZXd32p8 MxARID8:$src), MxSubRegIndex16Lo)>;
-def: Pat<(MxExtLoadi16i8 MxCP_ARII:$src),
- (EXTRACT_SUBREG (MOVZXd32f8 MxARII8:$src), MxSubRegIndex16Lo)>;
-def: Pat<(MxExtLoadi16i8 MxCP_PCD:$src),
- (EXTRACT_SUBREG (MOVZXd32q8 MxPCD8:$src), MxSubRegIndex16Lo)>;
-
-// i32 <- anyext i8
-def: Pat<(i32 (anyext i8:$src)), (MOVZXd32d8 MxDRD8:$src)>;
-def: Pat<(MxExtLoadi32i8 MxCP_ARI :$src), (MOVZXd32j8 MxARI8 :$src)>;
-def: Pat<(MxExtLoadi32i8 MxCP_ARID:$src), (MOVZXd32p8 MxARID8:$src)>;
-def: Pat<(MxExtLoadi32i8 MxCP_ARII:$src), (MOVZXd32f8 MxARII8:$src)>;
-def: Pat<(MxExtLoadi32i8 MxCP_PCD:$src), (MOVZXd32q8 MxPCD8:$src)>;
-
-// i32 <- anyext i16
-def: Pat<(i32 (anyext i16:$src)), (MOVZXd32d16 MxDRD16:$src)>;
-def: Pat<(MxExtLoadi32i16 MxCP_ARI :$src), (MOVZXd32j16 MxARI16 :$src)>;
-def: Pat<(MxExtLoadi32i16 MxCP_ARID:$src), (MOVZXd32p16 MxARID16:$src)>;
-def: Pat<(MxExtLoadi32i16 MxCP_ARII:$src), (MOVZXd32f16 MxARII16:$src)>;
-def: Pat<(MxExtLoadi32i16 MxCP_PCD:$src), (MOVZXd32q16 MxPCD16:$src)>;
+def: Pat<(i32 (zext i16:$src)), (MOVZXd32r16 MxXRD16:$src)>;
+def: Pat<(i16 (anyext i8:$src)), (MOVXd16d8 MxDRD8:$src)>;
+def: Pat<(i32 (anyext i8:$src)), (MOVXd32d8 MxDRD8:$src)>;
+def: Pat<(i32 (anyext i16:$src)), (MOVXr32r16 MxXRD16:$src)>;
// trunc patterns
def : Pat<(i16 (trunc i32:$src)),
(EXTRACT_SUBREG MxXRD32:$src, MxSubRegIndex16Lo)>;
def : Pat<(i8 (trunc i32:$src)),
- (EXTRACT_SUBREG MxXRD32:$src, MxSubRegIndex8Lo)>;
+ (EXTRACT_SUBREG MxDRD32:$src, MxSubRegIndex8Lo)>;
def : Pat<(i8 (trunc i16:$src)),
- (EXTRACT_SUBREG MxXRD16:$src, MxSubRegIndex8Lo)>;
+ (EXTRACT_SUBREG MxDRD16:$src, MxSubRegIndex8Lo)>;
//===----------------------------------------------------------------------===//
// FMOVE
diff --git a/llvm/lib/Target/M68k/M68kInstrInfo.cpp b/llvm/lib/Target/M68k/M68kInstrInfo.cpp
index 61693465af708..d4ebeab9b4b0a 100644
--- a/llvm/lib/Target/M68k/M68kInstrInfo.cpp
+++ b/llvm/lib/Target/M68k/M68kInstrInfo.cpp
@@ -435,7 +435,7 @@ bool M68kInstrInfo::ExpandMOVI(MachineInstrBuilder &MIB, MVT MVTSize) const {
bool M68kInstrInfo::ExpandMOVX_RR(MachineInstrBuilder &MIB, MVT MVTDst,
MVT MVTSrc) const {
- unsigned Move = MVTDst == MVT::i16 ? M68k::MOV16rr : M68k::MOV32rr;
+ unsigned Move = MVTSrc == MVT::i8 ? M68k::MOV8dd : M68k::MOV16rr;
Register Dst = MIB->getOperand(0).getReg();
Register Src = MIB->getOperand(1).getReg();
@@ -450,19 +450,19 @@ bool M68kInstrInfo::ExpandMOVX_RR(MachineInstrBuilder &MIB, MVT MVTDst,
assert(RCDst != RCSrc && "You cannot use the same Reg Classes with MOVX_RR");
(void)RCSrc;
- // We need to find the super source register that matches the size of Dst
- unsigned SSrc = RI.getMatchingMegaReg(Src, RCDst);
- assert(SSrc && "No viable MEGA register available");
+ unsigned SubDst =
+ RI.getSubReg(Dst, MVTSrc == MVT::i8 ? M68k::MxSubRegIndex8Lo
+ : M68k::MxSubRegIndex16Lo);
+ assert(SubDst && "No viable SUB register available");
- // If it happens to that super source register is the destination register
- // we do nothing
- if (Dst == SSrc) {
+ // If source is a subregister of destination, we do nothing
+ if (SubDst == Src) {
LLVM_DEBUG(dbgs() << "Remove " << *MIB.getInstr() << '\n');
MIB->eraseFromParent();
} else { // otherwise we need to MOV
LLVM_DEBUG(dbgs() << "Expand " << *MIB.getInstr() << " to MOV\n");
MIB->setDesc(get(Move));
- MIB->getOperand(1).setReg(SSrc);
+ MIB->getOperand(0).setReg(SubDst);
}
return true;
@@ -488,47 +488,53 @@ bool M68kInstrInfo::ExpandMOVSZX_RR(MachineInstrBuilder &MIB, bool IsSigned,
assert(RCDst != RCSrc && "You cannot use the same Reg Classes with MOVSX_RR");
(void)RCSrc;
- // We need to find the super source register that matches the size of Dst
- unsigned SSrc = RI.getMatchingMegaReg(Src, RCDst);
- assert(SSrc && "No viable MEGA register available");
+ // Move source into subreg of the destination.
+ unsigned SubDst =
+ RI.getSubReg(Dst, MVTSrc == MVT::i8 ? M68k::MxSubRegIndex8Lo
+ : M68k::MxSubRegIndex16Lo);
+ assert(SubDst && "No viable SUB register available");
MachineBasicBlock &MBB = *MIB->getParent();
DebugLoc DL = MIB->getDebugLoc();
+ unsigned Move;
+ if (MVTSrc == MVT::i8)
+ Move = M68k::MOV8dd;
+ else
+ Move = M68k::MOV16rr;
+
// It's more efficient to clear the destination and *then* move, rather than
// move and zext.
- if (Dst != SSrc && !IsSigned) {
+ if (SubDst != Src && !IsSigned) {
LLVM_DEBUG(dbgs() << "Clear and Move" << '\n');
buildClearRegister(Dst, MBB, MIB.getInstr(), DL);
- if (MVTSrc == MVT::i8) {
- unsigned SubDst = RI.getSubReg(Dst, M68k::MxSubRegIndex8Lo);
- BuildMI(MBB, MIB.getInstr(), DL, get(M68k::MOV8dd), SubDst).addReg(Src);
- } else { // i16
- unsigned SubDst = RI.getSubReg(Dst, M68k::MxSubRegIndex16Lo);
- BuildMI(MBB, MIB.getInstr(), DL, get(M68k::MOV16dd), SubDst).addReg(Src);
- }
- } else {
+ MIB->setDesc(get(Move));
+ MIB->getOperand(0).setReg(SubDst);
+ return true;
+ }
- unsigned Move;
- if (MVTDst == MVT::i16)
- Move = M68k::MOV16dd;
- else // i32
- Move = M68k::MOV32dd;
+ // Special case where move to AR16 automatically sign-extends to 32 bits.
+ // Even an in-place move (move.w a0,a0) will do so.
+ if (M68k::AR32RegClass.contains(Dst) && IsSigned) {
+ LLVM_DEBUG(dbgs() << "Move (implicit Sign Extend)" << '\n');
+ MIB->setDesc(get(M68k::MOV16ar));
+ MIB->getOperand(0).setReg(SubDst);
+ return true;
+ }
- if (Dst != SSrc) {
- LLVM_DEBUG(dbgs() << "Move and " << '\n');
- BuildMI(MBB, MIB.getInstr(), DL, get(Move), Dst).addReg(SSrc);
- }
+ if (SubDst != Src) {
+ LLVM_DEBUG(dbgs() << "Move and " << '\n');
+ BuildMI(MBB, MIB.getInstr(), DL, get(Move), SubDst).addReg(Src);
+ }
- if (IsSigned) {
- LLVM_DEBUG(dbgs() << "Sign Extend" << '\n');
- AddSExt(MBB, MIB.getInstr(), DL, Dst, MVTSrc, MVTDst);
- } else {
- LLVM_DEBUG(dbgs() << "Zero Extend" << '\n');
- AddZExt(MBB, MIB.getInstr(), DL, Dst, MVTSrc, MVTDst);
- }
+ if (IsSigned) {
+ LLVM_DEBUG(dbgs() << "Sign Extend" << '\n');
+ AddSExt(MBB, MIB.getInstr(), DL, Dst, MVTSrc, MVTDst);
+ } else {
+ LLVM_DEBUG(dbgs() << "Zero Extend" << '\n');
+ AddZExt(MBB, MIB.getInstr(), DL, Dst, MVTSrc, MVTDst);
}
MIB->eraseFromParent();
@@ -536,18 +542,35 @@ bool M68kInstrInfo::ExpandMOVSZX_RR(MachineInstrBuilder &MIB, bool IsSigned,
return true;
}
+bool M68kInstrInfo::ExpandMOVX_RM(MachineInstrBuilder &MIB, unsigned Opc,
+ MVT MVTDst, MVT MVTSrc) const {
+ LLVM_DEBUG(dbgs() << "Expand " << *MIB.getInstr() << " to LOAD" << '\n');
+
+ const MCInstrDesc &Desc = get(Opc);
+ Register Dst = MIB->getOperand(0).getReg();
+
+ // Load source into subreg of the destination.
+ unsigned SubDst =
+ RI.getSubReg(Dst, MVTSrc == MVT::i8 ? M68k::MxSubRegIndex8Lo
+ : M68k::MxSubRegIndex16Lo);
+ assert(SubDst && "No viable SUB register available");
+
+ // Make this a plain move
+ MIB->setDesc(Desc);
+ MIB->getOperand(0).setReg(SubDst);
+
+ return true;
+}
+
bool M68kInstrInfo::ExpandMOVSZX_RM(MachineInstrBuilder &MIB, bool IsSigned,
- const MCInstrDesc &Desc, MVT MVTDst,
+ unsigned Opc, MVT MVTDst,
MVT MVTSrc) const {
LLVM_DEBUG(dbgs() << "Expand " << *MIB.getInstr() << " to ");
+ const MCInstrDesc &Desc = get(Opc);
Register Dst = MIB->getOperand(0).getReg();
- // We need the subreg of Dst to make instruction verifier happy because the
- // real machine instruction consumes and produces values of the same size and
- // the registers the will be used here fall into different classes and this
- // makes IV cry. We could use a bigger operation, but this will put some
- // pressure on cache and memory, so no.
+ // Load source into subreg of the destination.
unsigned SubDst =
RI.getSubReg(Dst, MVTSrc == MVT::i8 ? M68k::MxSubRegIndex8Lo
: M68k::MxSubRegIndex16Lo);
@@ -557,6 +580,13 @@ bool M68kInstrInfo::ExpandMOVSZX_RM(MachineInstrBuilder &MIB, bool IsSigned,
MIB->setDesc(Desc);
MIB->getOperand(0).setReg(SubDst);
+ // Special case where move to AR16 automatically sign-extends to 32 bits. In
+ // that case, we're already done.
+ if (M68k::AR32RegClass.contains(Dst) && IsSigned) {
+ LLVM_DEBUG(dbgs() << "LOAD (implicit Sign Extend)" << '\n');
+ return true;
+ }
+
MachineBasicBlock::iterator I = MIB.getInstr();
MachineBasicBlock &MBB = *MIB->getParent();
DebugLoc DL = MIB->getDebugLoc();
@@ -755,16 +785,16 @@ void M68kInstrInfo::copyPhysReg(MachineBasicBlock &MBB,
// upper bits will be undefined.
// 8 -> 16
else if (M68k::DR8RegClass.contains(SrcReg) &&
- M68k::XR16RegClass.contains(DstReg)) {
+ M68k::DR16RegClass.contains(DstReg)) {
Opc = M68k::MOVXd16d8;
// 8 -> 32
} else if (M68k::DR8RegClass.contains(SrcReg) &&
- M68k::XR32RegClass.contains(DstReg)) {
+ M68k::DR32RegClass.contains(DstReg)) {
Opc = M68k::MOVXd32d8;
// 16 -> 32
} else if (M68k::XR16RegClass.contains(SrcReg) &&
M68k::XR32RegClass.contains(DstReg)) {
- Opc = M68k::MOVXd32d16;
+ Opc = M68k::MOVXr32r16;
}
// Copy from CCR
diff --git a/llvm/lib/Target/M68k/M68kInstrInfo.h b/llvm/lib/Target/M68k/M68kInstrInfo.h
index a8148dd583380..1af261e64b5ef 100644
--- a/llvm/lib/Target/M68k/M68kInstrInfo.h
+++ b/llvm/lib/Target/M68k/M68kInstrInfo.h
@@ -312,9 +312,13 @@ class M68kInstrInfo : public M68kGenInstrInfo {
bool ExpandMOVSZX_RR(MachineInstrBuilder &MIB, bool IsSigned, MVT MVTDst,
MVT MVTSrc) const;
+ /// Move from memory and expand register class without extension
+ bool ExpandMOVX_RM(MachineInstrBuilder &MIB, unsigned Opc, MVT MVTDst,
+ MVT MVTSrc) const;
+
/// Move from memory and extend
- bool ExpandMOVSZX_RM(MachineInstrBuilder &MIB, bool IsSigned,
- const MCInstrDesc &Desc, MVT MVTDst, MVT MVTSrc) const;
+ bool ExpandMOVSZX_RM(MachineInstrBuilder &MIB, bool IsSigned, unsigned Opc,
+ MVT MVTDst, MVT MVTSrc) const;
/// Push/Pop to/from stack
bool ExpandPUSH_POP(MachineInstrBuilder &MIB, const MCInstrDesc &Desc,
diff --git a/llvm/test/CodeGen/M68k/Arith/divide-by-constant.ll b/llvm/test/CodeGen/M68k/Arith/divide-by-constant.ll
index 65525efe01a8d..e2f9ac38ebd97 100644
--- a/llvm/test/CodeGen/M68k/Arith/divide-by-constant.ll
+++ b/llvm/test/CodeGen/M68k/Arith/divide-by-constant.ll
@@ -38,7 +38,7 @@ define zeroext i8 @test3(i8 zeroext %x, i8 zeroext %c) {
; CHECK-LABEL: test3:
; CHECK: .cfi_startproc
; CHECK-NEXT: ; %bb.0: ; %entry
-; CHECK-NEXT: moveq #0, %d0
+; CHECK-NEXT: clr.w %d0
; CHECK-NEXT: move.b (11,%sp), %d0
; CHECK-NEXT: muls #171, %d0
; CHECK-NEXT: moveq #9, %d1
@@ -128,7 +128,7 @@ define i8 @test8(i8 %x) nounwind {
; CHECK: ; %bb.0:
; CHECK-NEXT: move.b (7,%sp), %d0
; CHECK-NEXT: lsr.b #1, %d0
-; CHECK-NEXT: and.l #255, %d0
+; CHECK-NEXT: and.w #255, %d0
; CHECK-NEXT: muls #211, %d0
; CHECK-NEXT: moveq #13, %d1
; CHECK-NEXT: lsr.w %d1, %d0
@@ -143,7 +143,7 @@ define i8 @test9(i8 %x) nounwind {
; CHECK: ; %bb.0:
; CHECK-NEXT: move.b (7,%sp), %d0
; CHECK-NEXT: lsr.b #2, %d0
-; CHECK-NEXT: and.l #255, %d0
+; CHECK-NEXT: and.w #255, %d0
; CHECK-NEXT: muls #71, %d0
; CHECK-NEXT: moveq #11, %d1
; CHECK-NEXT: lsr.w %d1, %d0
diff --git a/llvm/test/CodeGen/M68k/Arith/smul-with-overflow.ll b/llvm/test/CodeGen/M68k/Arith/smul-with-overflow.ll
index d10e1b12bd083..ed66043267c5c 100644
--- a/llvm/test/CodeGen/M68k/Arith/smul-with-overflow.ll
+++ b/llvm/test/CodeGen/M68k/Arith/smul-with-overflow.ll
@@ -4,13 +4,9 @@
define zeroext i8 @smul_i8(i8 signext %a, i8 signext %b) nounwind ssp {
; CHECK-LABEL: smul_i8:
; CHECK: ; %bb.0: ; %entry
-; CHECK-NEXT: moveq #0, %d0
-; CHECK-NEXT: move.b (11,%sp), %d0
-; CHECK-NEXT: moveq #0, %d1
-; CHECK-NEXT: move.b (7,%sp), %d1
-; CHECK-NEXT: muls %d0, %d1
-; CHECK-NEXT: moveq #0, %d0
-; CHECK-NEXT: move.w %d1, %d0
+; CHECK-NEXT: move.b (7,%sp), %d0
+; CHECK-NEXT: move.b (11,%sp), %d1
+; CHECK-NEXT: muls %d1, %d0
; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: rts
entry:
diff --git a/llvm/test/CodeGen/M68k/Arith/umul-with-overflow.ll b/llvm/test/CodeGen/M68k/Arith/umul-with-overflow.ll
index 42131dfa6b413..30f00fb95e8e6 100644
--- a/llvm/test/CodeGen/M68k/Arith/umul-with-overflow.ll
+++ b/llvm/test/CodeGen/M68k/Arith/umul-with-overflow.ll
@@ -4,13 +4,9 @@
define zeroext i8 @umul_i8(i8 signext %a, i8 signext %b) nounwind ssp {
; CHECK-LABEL: umul_i8:
; CHECK: ; %bb.0: ; %entry
-; CHECK-NEXT: moveq #0, %d0
-; CHECK-NEXT: move.b (11,%sp), %d0
-; CHECK-NEXT: moveq #0, %d1
-; CHECK-NEXT: move.b (7,%sp), %d1
-; CHECK-NEXT: muls %d0, %d1
-; CHECK-NEXT: moveq #0, %d0
-; CHECK-NEXT: move.w %d1, %d0
+; CHECK-NEXT: move.b (7,%sp), %d0
+; CHECK-NEXT: move.b (11,%sp), %d1
+; CHECK-NEXT: muls %d1, %d0
; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: rts
entry:
diff --git a/llvm/test/CodeGen/M68k/Bits/btst.ll b/llvm/test/CodeGen/M68k/Bits/btst.ll
index 63eb07e961e42..cb3dd2e9956d4 100644
--- a/llvm/test/CodeGen/M68k/Bits/btst.ll
+++ b/llvm/test/CodeGen/M68k/Bits/btst.ll
@@ -7,9 +7,6 @@ define fastcc i16 @switch_to_btst(i16 %a) nounwind {
; CHECK-NEXT: cmpi.w #11, %d0
; CHECK-NEXT: bhi .LBB0_3
; CHECK-NEXT: ; %bb.1: ; %entry
-; CHECK-NEXT: swap %d0
-; CHECK-NEXT: clr.w %d0
-; CHECK-NEXT: swap %d0
; CHECK-NEXT: move.l #3612, %d1
; CHECK-NEXT: btst %d0, %d1
; CHECK-NEXT: beq .LBB0_3
diff --git a/llvm/test/CodeGen/M68k/Control/cmp.ll b/llvm/test/CodeGen/M68k/Control/cmp.ll
index 7988e44672dea..e4940ac919244 100644
--- a/llvm/test/CodeGen/M68k/Control/cmp.ll
+++ b/llvm/test/CodeGen/M68k/Control/cmp.ll
@@ -82,7 +82,6 @@ define i64 @test3(i64 %x) nounwind {
; CHECK-NEXT: move.l (8,%sp), %d0
; CHECK-NEXT: or.l (4,%sp), %d0
; CHECK-NEXT: seq %d0
-; CHECK-NEXT: moveq #0, %d1
; CHECK-NEXT: move.b %d0, %d1
; CHECK-NEXT: and.l #1, %d1
; CHECK-NEXT: moveq #0, %d0
@@ -103,7 +102,6 @@ define i64 @test4(i64 %x) nounwind {
; CHECK-NEXT: sub.l #1, %d2
; CHECK-NEXT: subx.l %d0, %d1
; CHECK-NEXT: slt %d1
-; CHECK-NEXT: and.l #255, %d1
; CHECK-NEXT: and.l #1, %d1
; CHECK-NEXT: movem.l (0,%sp), %d2 ; 8-byte Folded Reload
; CHECK-NEXT: adda.l #4, %sp
@@ -145,7 +143,6 @@ define i32 @test7(i64 %res) nounwind {
; CHECK: ; %bb.0: ; %entry
; CHECK-NEXT: cmpi.l #0, (4,%sp)
; CHECK-NEXT: seq %d0
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: and.l #1, %d0
; CHECK-NEXT: rts
entry:
@@ -159,7 +156,6 @@ define i32 @test8(i64 %res) nounwind {
; CHECK: ; %bb.0: ; %entry
; CHECK-NEXT: cmpi.l #3, (4,%sp)
; CHECK-NEXT: scs %d0
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: and.l #1, %d0
; CHECK-NEXT: rts
entry:
@@ -175,7 +171,6 @@ define i32 @test11(i64 %l) nounwind {
; CHECK-NEXT: and.l #-32768, %d0
; CHECK-NEXT: cmpi.l #32768, %d0
; CHECK-NEXT: seq %d0
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: and.l #1, %d0
; CHECK-NEXT: rts
entry:
@@ -241,7 +236,6 @@ define zeroext i1 @test15(i32 %bf.load, i32 %n) {
; CHECK-NEXT: cmpi.l #0, %d1
; CHECK-NEXT: seq %d1
; CHECK-NEXT: or.b %d0, %d1
-; CHECK-NEXT: moveq #0, %d0
; CHECK-NEXT: move.b %d1, %d0
; CHECK-NEXT: and.l #1, %d0
; CHECK-NEXT: rts
@@ -296,7 +290,6 @@ define void @test20(i32 %bf.load, i8 %x1, ptr %b_addr) {
; CHECK-NEXT: move.l #16777215, %d0
; CHECK-NEXT: and.l (8,%sp), %d0
; CHECK-NEXT: sne %d1
-; CHECK-NEXT: and.l #255, %d1
; CHECK-NEXT: and.l #1, %d1
; CHECK-NEXT: moveq #0, %d2
; CHECK-NEXT: move.b (15,%sp), %d2
diff --git a/llvm/test/CodeGen/M68k/Control/non-cmov-switch.ll b/llvm/test/CodeGen/M68k/Control/non-cmov-switch.ll
index 2ce75d380e51e..6efb6f7ca53f8 100644
--- a/llvm/test/CodeGen/M68k/Control/non-cmov-switch.ll
+++ b/llvm/test/CodeGen/M68k/Control/non-cmov-switch.ll
@@ -16,7 +16,6 @@ define internal void @select_i32(i32 %self, ptr nonnull %value) {
; M68000-NEXT: move.w %d2, %ccr
; M68000-NEXT: bne .LBB0_2
; M68000-NEXT: ; %bb.1: ; %start
-; M68000-NEXT: and.l #255, %d1
; M68000-NEXT: and.l #1, %d1
; M68000-NEXT: cmpi.l #0, %d1
; M68000-NEXT: bne .LBB0_3
@@ -41,7 +40,6 @@ define internal void @select_i32(i32 %self, ptr nonnull %value) {
; M68020-NEXT: move.w %d2, %ccr
; M68020-NEXT: bne .LBB0_2
; M68020-NEXT: ; %bb.1: ; %start
-; M68020-NEXT: and.l #255, %d1
; M68020-NEXT: and.l #1, %d1
; M68020-NEXT: cmpi.l #0, %d1
; M68020-NEXT: bne .LBB0_3
@@ -86,7 +84,6 @@ define internal void @select_i16(i16 %self, ptr nonnull %value) {
; M68000-NEXT: move.w %d2, %ccr
; M68000-NEXT: bne .LBB1_2
; M68000-NEXT: ; %bb.1: ; %start
-; M68000-NEXT: and.l #255, %d1
; M68000-NEXT: and.w #1, %d1
; M68000-NEXT: cmpi.w #0, %d1
; M68000-NEXT: bne .LBB1_3
@@ -111,7 +108,6 @@ define internal void @select_i16(i16 %self, ptr nonnull %value) {
; M68020-NEXT: move.w %d2, %ccr
; M68020-NEXT: bne .LBB1_2
; M68020-NEXT: ; %bb.1: ; %start
-; M68020-NEXT: and.l #255, %d1
; M68020-NEXT: and.w #1, %d1
; M68020-NEXT: cmpi.w #0, %d1
; M68020-NEXT: bne .LBB1_3
diff --git a/llvm/test/CodeGen/M68k/Control/setcc.ll b/llvm/test/CodeGen/M68k/Control/setcc.ll
index 4be853b0067b4..73e0480bf94ce 100644
--- a/llvm/test/CodeGen/M68k/Control/setcc.ll
+++ b/llvm/test/CodeGen/M68k/Control/setcc.ll
@@ -8,7 +8,6 @@ define zeroext i16 @t1(i16 zeroext %x) nounwind readnone ssp {
; CHECK: ; %bb.0: ; %entry
; CHECK-NEXT: cmpi.w #26, (6,%sp)
; CHECK-NEXT: shi %d0
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: and.l #1, %d0
; CHECK-NEXT: lsl.l #5, %d0
; CHECK-NEXT: rts
@@ -23,7 +22,6 @@ define zeroext i16 @t2(i16 zeroext %x) nounwind readnone ssp {
; CHECK: ; %bb.0: ; %entry
; CHECK-NEXT: cmpi.w #26, (6,%sp)
; CHECK-NEXT: scs %d0
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: and.l #1, %d0
; CHECK-NEXT: lsl.l #5, %d0
; CHECK-NEXT: rts
@@ -42,7 +40,6 @@ define fastcc i64 @t3(i64 %x) nounwind readnone ssp {
; CHECK-NEXT: sub.l #18, %d1
; CHECK-NEXT: subx.l %d2, %d0
; CHECK-NEXT: scs %d0
-; CHECK-NEXT: moveq #0, %d1
; CHECK-NEXT: move.b %d0, %d1
; CHECK-NEXT: and.l #1, %d1
; CHECK-NEXT: lsl.l #6, %d1
diff --git a/llvm/test/CodeGen/M68k/Data/load-extend.ll b/llvm/test/CodeGen/M68k/Data/load-extend.ll
index 0a4d53643de29..fea7482ef9356 100644
--- a/llvm/test/CodeGen/M68k/Data/load-extend.ll
+++ b/llvm/test/CodeGen/M68k/Data/load-extend.ll
@@ -45,10 +45,8 @@ define i32 @"test_zext_pcd_i16_to_i32"() {
define i16 @test_anyext_pcd_i8_to_i16() nounwind {
; CHECK-LABEL: test_anyext_pcd_i8_to_i16:
; CHECK: ; %bb.0:
-; CHECK-NEXT: moveq #0, %d0
; CHECK-NEXT: move.b (__unnamed_1+4,%pc), %d0
; CHECK-NEXT: lsl.w #8, %d0
-; CHECK-NEXT: ; kill: def $wd0 killed $wd0 killed $d0
; CHECK-NEXT: rts
%copyload = load i8, ptr getelementptr inbounds nuw (i8, ptr @0, i32 4)
%insert_ext = zext i8 %copyload to i16
@@ -60,7 +58,6 @@ define i32 @test_anyext_pcd_i8_to_i32() nounwind {
; CHECK-LABEL: test_anyext_pcd_i8_to_i32:
; CHECK: ; %bb.0:
; CHECK-NEXT: moveq #24, %d1
-; CHECK-NEXT: moveq #0, %d0
; CHECK-NEXT: move.b (__unnamed_1+4,%pc), %d0
; CHECK-NEXT: lsl.l %d1, %d0
; CHECK-NEXT: rts
@@ -73,7 +70,6 @@ define i32 @test_anyext_pcd_i8_to_i32() nounwind {
define i32 @test_anyext_pcd_i16_to_i32() nounwind {
; CHECK-LABEL: test_anyext_pcd_i16_to_i32:
; CHECK: ; %bb.0:
-; CHECK-NEXT: moveq #0, %d0
; CHECK-NEXT: move.w (__unnamed_1+4,%pc), %d0
; CHECK-NEXT: swap %d0
; CHECK-NEXT: clr.w %d0
diff --git a/llvm/test/CodeGen/M68k/Data/sext-i1.ll b/llvm/test/CodeGen/M68k/Data/sext-i1.ll
index a7f97fa5829ce..e35ab50846938 100644
--- a/llvm/test/CodeGen/M68k/Data/sext-i1.ll
+++ b/llvm/test/CodeGen/M68k/Data/sext-i1.ll
@@ -21,7 +21,6 @@ define void @sext_inreg_i1_to_i16(i1 %val, ptr %ptr) {
; CHECK-LABEL: sext_inreg_i1_to_i16:
; CHECK: .cfi_startproc
; CHECK-NEXT: ; %bb.0: ; %entry
-; CHECK-NEXT: moveq #0, %d0
; CHECK-NEXT: move.b (7,%sp), %d0
; CHECK-NEXT: and.w #1, %d0
; CHECK-NEXT: neg.w %d0
@@ -38,7 +37,6 @@ define void @sext_inreg_i1_to_i32(i1 %val, ptr %ptr) {
; CHECK-LABEL: sext_inreg_i1_to_i32:
; CHECK: .cfi_startproc
; CHECK-NEXT: ; %bb.0: ; %entry
-; CHECK-NEXT: moveq #0, %d0
; CHECK-NEXT: move.b (7,%sp), %d0
; CHECK-NEXT: and.l #1, %d0
; CHECK-NEXT: neg.l %d0
diff --git a/llvm/test/CodeGen/M68k/register-spills.ll b/llvm/test/CodeGen/M68k/register-spills.ll
index 4055871632c93..c55da8fc9c97c 100644
--- a/llvm/test/CodeGen/M68k/register-spills.ll
+++ b/llvm/test/CodeGen/M68k/register-spills.ll
@@ -83,46 +83,30 @@ define void @test_force_spill_8() {
; CHECK-NEXT: movem.w %d0, (68,%sp)
; CHECK-NEXT: jsr get8
; CHECK-NEXT: movem.w (66,%sp), %d1
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: move.l %sp, %a0
; CHECK-NEXT: move.l %d0, (60,%a0)
; CHECK-NEXT: movem.w (68,%sp), %d0
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: move.l %d0, (56,%a0)
; CHECK-NEXT: movem.w (70,%sp), %d0
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: move.l %d0, (52,%a0)
; CHECK-NEXT: movem.w (72,%sp), %d0
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: move.l %d0, (48,%a0)
; CHECK-NEXT: movem.w (74,%sp), %d0
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: move.l %d0, (44,%a0)
; CHECK-NEXT: movem.w (76,%sp), %d0
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: move.l %d0, (40,%a0)
; CHECK-NEXT: movem.w (78,%sp), %d0
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: move.l %d0, (36,%a0)
; CHECK-NEXT: movem.w (80,%sp), %d0
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: move.l %d0, (32,%a0)
; CHECK-NEXT: movem.w (82,%sp), %d0
-; CHECK-NEXT: and.l #255, %d7
; CHECK-NEXT: move.l %d7, (28,%a0)
-; CHECK-NEXT: and.l #255, %d6
; CHECK-NEXT: move.l %d6, (24,%a0)
-; CHECK-NEXT: and.l #255, %d5
; CHECK-NEXT: move.l %d5, (20,%a0)
-; CHECK-NEXT: and.l #255, %d4
; CHECK-NEXT: move.l %d4, (16,%a0)
-; CHECK-NEXT: and.l #255, %d3
; CHECK-NEXT: move.l %d3, (12,%a0)
-; CHECK-NEXT: and.l #255, %d2
; CHECK-NEXT: move.l %d2, (8,%a0)
-; CHECK-NEXT: and.l #255, %d1
; CHECK-NEXT: move.l %d1, (4,%a0)
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: move.l %d0, (%a0)
; CHECK-NEXT: jsr test_force_spill_8_consumer
; CHECK-NEXT: movem.l (84,%sp), %d2-%d7 ; 28-byte Folded Reload
@@ -191,60 +175,24 @@ define void @test_force_spill_16() {
; CHECK-NEXT: jsr get16
; CHECK-NEXT: movem.w (64,%sp), %a1
; CHECK-NEXT: movem.w (66,%sp), %d1
-; CHECK-NEXT: swap %d0
-; CHECK-NEXT: clr.w %d0
-; CHECK-NEXT: swap %d0
; CHECK-NEXT: move.l %sp, %a0
; CHECK-NEXT: move.l %d0, (60,%a0)
; CHECK-NEXT: movem.w (68,%sp), %d0
-; CHECK-NEXT: swap %d0
-; CHECK-NEXT: clr.w %d0
-; CHECK-NEXT: swap %d0
; CHECK-NEXT: move.l %d0, (56,%a0)
; CHECK-NEXT: movem.w (70,%sp), %d0
-; CHECK-NEXT: and.l #65535, %a6
; CHECK-NEXT: move.l %a6, (52,%a0)
-; CHECK-NEXT: and.l #65535, %a5
; CHECK-NEXT: move.l %a5, (48,%a0)
-; CHECK-NEXT: and.l #65535, %a4
; CHECK-NEXT: move.l %a4, (44,%a0)
-; CHECK-NEXT: and.l #65535, %a3
; CHECK-NEXT: move.l %a3, (40,%a0)
-; CHECK-NEXT: and.l #65535, %a2
; CHECK-NEXT: move.l %a2, (36,%a0)
-; CHECK-NEXT: swap %d7
-; CHECK-NEXT: clr.w %d7
-; CHECK-NEXT: swap %d7
; CHECK-NEXT: move.l %d7, (32,%a0)
-; CHECK-NEXT: swap %d6
-; CHECK-NEXT: clr.w %d6
-; CHECK-NEXT: swap %d6
; CHECK-NEXT: move.l %d6, (28,%a0)
-; CHECK-NEXT: swap %d5
-; CHECK-NEXT: clr.w %d5
-; CHECK-NEXT: swap %d5
; CHECK-NEXT: move.l %d5, (24,%a0)
-; CHECK-NEXT: swap %d4
-; CHECK-NEXT: clr.w %d4
-; CHECK-NEXT: swap %d4
; CHECK-NEXT: move.l %d4, (20,%a0)
-; CHECK-NEXT: swap %d3
-; CHECK-NEXT: clr.w %d3
-; CHECK-NEXT: swap %d3
; CHECK-NEXT: move.l %d3, (16,%a0)
-; CHECK-NEXT: swap %d2
-; CHECK-NEXT: clr.w %d2
-; CHECK-NEXT: swap %d2
; CHECK-NEXT: move.l %d2, (12,%a0)
-; CHECK-NEXT: and.l #65535, %a1
; CHECK-NEXT: move.l %a1, (8,%a0)
-; CHECK-NEXT: swap %d1
-; CHECK-NEXT: clr.w %d1
-; CHECK-NEXT: swap %d1
; CHECK-NEXT: move.l %d1, (4,%a0)
-; CHECK-NEXT: swap %d0
-; CHECK-NEXT: clr.w %d0
-; CHECK-NEXT: swap %d0
; CHECK-NEXT: move.l %d0, (%a0)
; CHECK-NEXT: jsr test_force_spill_16_consumer
; CHECK-NEXT: movem.l (72,%sp), %d2-%d7/%a2-%a6 ; 48-byte Folded Reload
@@ -407,61 +355,32 @@ define void @test_force_spill_mixed() {
; CHECK-NEXT: jsr get16
; CHECK-NEXT: movem.l (80,%sp), %a1
; CHECK-NEXT: movem.w (86,%sp), %d1
-; CHECK-NEXT: swap %d0
-; CHECK-NEXT: clr.w %d0
-; CHECK-NEXT: swap %d0
; CHECK-NEXT: move.l %sp, %a0
; CHECK-NEXT: move.l %d0, (76,%a0)
; CHECK-NEXT: movem.l (88,%sp), %d0
; CHECK-NEXT: move.l %d0, (72,%a0)
; CHECK-NEXT: movem.w (94,%sp), %d0
-; CHECK-NEXT: swap %d0
-; CHECK-NEXT: clr.w %d0
-; CHECK-NEXT: swap %d0
; CHECK-NEXT: move.l %d0, (68,%a0)
; CHECK-NEXT: movem.w (96,%sp), %d0
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: move.l %d0, (64,%a0)
; CHECK-NEXT: movem.w (98,%sp), %d0
-; CHECK-NEXT: swap %d0
-; CHECK-NEXT: clr.w %d0
-; CHECK-NEXT: swap %d0
; CHECK-NEXT: move.l %d0, (60,%a0)
; CHECK-NEXT: movem.w (100,%sp), %d0
; CHECK-NEXT: move.l %a6, (56,%a0)
-; CHECK-NEXT: and.l #65535, %a5
; CHECK-NEXT: move.l %a5, (52,%a0)
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: move.l %d0, (48,%a0)
; CHECK-NEXT: movem.w (102,%sp), %d0
-; CHECK-NEXT: and.l #65535, %a4
; CHECK-NEXT: move.l %a4, (44,%a0)
; CHECK-NEXT: move.l %a3, (40,%a0)
-; CHECK-NEXT: and.l #65535, %a2
; CHECK-NEXT: move.l %a2, (36,%a0)
-; CHECK-NEXT: and.l #255, %d7
; CHECK-NEXT: move.l %d7, (32,%a0)
-; CHECK-NEXT: swap %d6
-; CHECK-NEXT: clr.w %d6
-; CHECK-NEXT: swap %d6
; CHECK-NEXT: move.l %d6, (28,%a0)
; CHECK-NEXT: move.l %d5, (24,%a0)
-; CHECK-NEXT: swap %d4
-; CHECK-NEXT: clr.w %d4
-; CHECK-NEXT: swap %d4
; CHECK-NEXT: move.l %d4, (20,%a0)
-; CHECK-NEXT: and.l #255, %d3
; CHECK-NEXT: move.l %d3, (16,%a0)
-; CHECK-NEXT: swap %d2
-; CHECK-NEXT: clr.w %d2
-; CHECK-NEXT: swap %d2
; CHECK-NEXT: move.l %d2, (12,%a0)
; CHECK-NEXT: move.l %a1, (8,%a0)
-; CHECK-NEXT: swap %d1
-; CHECK-NEXT: clr.w %d1
-; CHECK-NEXT: swap %d1
; CHECK-NEXT: move.l %d1, (4,%a0)
-; CHECK-NEXT: and.l #255, %d0
; CHECK-NEXT: move.l %d0, (%a0)
; CHECK-NEXT: jsr test_force_spill_mixed_consumer
; CHECK-NEXT: movem.l (104,%sp), %d2-%d7/%a2-%a6 ; 48-byte Folded Reload
>From cc11215f70b28d7895ecb50cee36dded4b257a2e Mon Sep 17 00:00:00 2001
From: Alexander Richardson <alexrichardson at google.com>
Date: Sat, 26 Sep 2026 19:00:43 -0700
Subject: [PATCH 30/53] [RISC-V][LTO] Add baseline tests for LTO inline
assembly and mapping symbols (#225129)
No functional change intended here, just adding test coverage for RISC-V LTO
inline assembly ABI handling (following up on
https://github.com/llvm/llvm-project/pull/223606) and for the `$x<arch>` ELF
mapping symbols emitted for module and function target features.
The `TODO`s for `.lto_discard` dropping module inline asm target features and
for the missing/duplicate `$x<arch>` mapping symbols will be addressed in the
following commits.
This commit was created with the help of AI tools
Pull-Request: https://github.com/llvm/llvm-project/pull/225129
---
clang/test/CodeGen/RISCV/lto-module-asm-abi.c | 36 ++++++
cross-project-tests/.clang-format | 2 +
cross-project-tests/riscv/lit.local.cfg | 6 +
.../riscv/lto-inline-asm-abi.c | 89 +++++++++++++
lld/test/ELF/lto/riscv-target-abi.ll | 121 ++++++++++++++----
.../test/CodeGen/RISCV/module-asm-features.ll | 23 +++-
.../RISCV/riscv-func-target-feature.ll | 60 ++++++---
llvm/test/LTO/RISCV/module-asm.ll | 15 ++-
llvm/test/MC/RISCV/mapping-across-sections.s | 27 +++-
9 files changed, 328 insertions(+), 51 deletions(-)
create mode 100644 clang/test/CodeGen/RISCV/lto-module-asm-abi.c
create mode 100644 cross-project-tests/.clang-format
create mode 100644 cross-project-tests/riscv/lit.local.cfg
create mode 100644 cross-project-tests/riscv/lto-inline-asm-abi.c
diff --git a/clang/test/CodeGen/RISCV/lto-module-asm-abi.c b/clang/test/CodeGen/RISCV/lto-module-asm-abi.c
new file mode 100644
index 0000000000000..cceafaac3ba0e
--- /dev/null
+++ b/clang/test/CodeGen/RISCV/lto-module-asm-abi.c
@@ -0,0 +1,36 @@
+// REQUIRES: riscv-registered-target
+
+/// Regression test for https://github.com/llvm/llvm-project/pull/213410:
+/// Check that -march=rv64gcv -flto records +d in module asm and function
+/// target-features even though the driver only passes mcpu=generic-rv64 to lld.
+
+// RUN: %clang --target=riscv64-linux-android -march=rv64gcv -flto %s -S -emit-llvm -o - \
+// RUN: | FileCheck %s --check-prefix=IR
+
+// IR: module asm(target_features: "{{.*}}+d{{.*}}", target_cpu: "generic-rv64")
+// IR-NEXT: "nop"
+// IR: define dso_local void @_start() #[[#ATTR:]] {{.*}} {
+// IR-NEXT: entry:
+// IR-NEXT: call void asm sideeffect "nop", ""()
+// IR-NEXT: ret void
+// IR-NEXT: }
+// IR-EMPTY:
+// IR-NEXT: attributes #[[#ATTR]] = { {{.*}}"target-cpu"="generic-rv64" "target-features"="{{.*}}+d{{.*}}"
+// IR: ![[#]] = !{i32 1, !"target-abi", !"lp64d"}
+// IR-NEXT: ![[#]] = !{i32 6, !"riscv-isa", ![[#ISA:]]}
+// IR-NEXT: ![[#ISA]] = !{!"{{.*}}_d2p2_{{.*}}"}
+
+// RUN: %clang --target=riscv64-linux-android -march=rv64gcv -flto -shared -nostdlib -fuse-ld=lld %s -### 2>&1 \
+// RUN: | FileCheck %s --check-prefix=DRIVER
+
+// DRIVER: "-cc1"{{.*}}"-target-cpu" "generic-rv64"{{.*}}"-target-feature" "+d"{{.*}}"-target-abi" "lp64d"
+// DRIVER: "{{[^"]*}}ld.lld{{(\.exe)?}}"
+// DRIVER-NOT: mattr
+// DRIVER-SAME: "-plugin-opt=mcpu=generic-rv64"
+// DRIVER-NOT: mattr
+
+__asm__("nop");
+
+void _start(void) {
+ __asm__ volatile("nop");
+}
diff --git a/cross-project-tests/.clang-format b/cross-project-tests/.clang-format
new file mode 100644
index 0000000000000..f5e3ec5b16d1f
--- /dev/null
+++ b/cross-project-tests/.clang-format
@@ -0,0 +1,2 @@
+BasedOnStyle: LLVM
+ReflowComments: false
diff --git a/cross-project-tests/riscv/lit.local.cfg b/cross-project-tests/riscv/lit.local.cfg
new file mode 100644
index 0000000000000..e966a6d4ed56e
--- /dev/null
+++ b/cross-project-tests/riscv/lit.local.cfg
@@ -0,0 +1,6 @@
+if (
+ "clang" not in config.available_features
+ or "ld.lld" not in config.available_features
+ or "RISCV" not in config.targets_to_build
+):
+ config.unsupported = True
diff --git a/cross-project-tests/riscv/lto-inline-asm-abi.c b/cross-project-tests/riscv/lto-inline-asm-abi.c
new file mode 100644
index 0000000000000..e89e1e4f0687d
--- /dev/null
+++ b/cross-project-tests/riscv/lto-inline-asm-abi.c
@@ -0,0 +1,89 @@
+// REQUIRES: ld.lld
+/// Regression test for https://github.com/llvm/llvm-project/pull/213410:
+/// Check that module-level inline assembly (including .symver imported by
+/// ThinLTO) and function-level inline assembly link cleanly under RegularLTO
+/// and ThinLTO when targeting riscv64 with lp64d ABI and -march=rv64gcv.
+// RUN: rm -rf %t && split-file %s %t
+// RUN: %clang --target=riscv64-linux-android -march=rv64gcv -O2 -flto -c %t/a.c -o %t1.o
+// RUN: %clang --target=riscv64-linux-android -march=rv64gcv -O2 -flto -c %t/b.c -o %t2.o
+// RUN: %clang --target=riscv64-linux-android -march=rv64gcv -O2 -flto -shared -nostdlib -fuse-ld=lld -Wl,-save-temps -Wl,--version-script=%t/ver.ver %t1.o %t2.o -o %t.so 2>&1 \
+// RUN: | FileCheck %s --allow-empty --implicit-check-not="error:" --implicit-check-not="warning:" --implicit-check-not="note:"
+// RUN: llvm-dis %t.so.0.5.precodegen.bc -o - | FileCheck %s --check-prefix=REGULAR-IR
+// RUN: llvm-readobj --file-headers %t.so | FileCheck %s --check-prefix=FLAGS
+// RUN: llvm-objdump -d --show-all-symbols --no-show-raw-insn %t.so | FileCheck %s --check-prefix=DISASM
+// RUN: llvm-objdump -t %t.so | FileCheck %s --check-prefix=SYMS --implicit-check-not='\$x'
+//
+// RUN: %clang --target=riscv64-linux-android -march=rv64gcv -O2 -flto=thin -c %t/a.c -o %t1.thin.o
+// RUN: %clang --target=riscv64-linux-android -march=rv64gcv -O2 -flto=thin -c %t/b.c -o %t2.thin.o
+// RUN: %clang --target=riscv64-linux-android -march=rv64gcv -O2 -flto=thin -shared -nostdlib -fuse-ld=lld -Wl,-save-temps -Wl,--version-script=%t/ver.ver %t1.thin.o %t2.thin.o -o %t.thin.so 2>&1 \
+// RUN: | FileCheck %s --allow-empty --implicit-check-not="error:" --implicit-check-not="warning:" --implicit-check-not="note:"
+// RUN: llvm-dis %t2.thin.o.5.precodegen.bc -o - | FileCheck %s --check-prefix=THIN-IR
+// RUN: llvm-readobj --file-headers %t.thin.so | FileCheck %s --check-prefix=FLAGS
+// RUN: llvm-objdump -d --show-all-symbols --no-show-raw-insn %t.thin.so | FileCheck %s --check-prefix=DISASM
+// RUN: llvm-objdump -t %t.thin.so | FileCheck %s --check-prefix=SYMS --implicit-check-not='\$x'
+//
+/// TODO: LTO::addRegularLTO and IRLinker::run drop target_features and
+/// target_cpu when synthesizing .lto_discard and imported .symver directives.
+// REGULAR-IR: module asm{{$}}
+// REGULAR-IR-NEXT: ".lto_discard "
+// REGULAR-IR-NEXT: module asm(target_features: "+64bit,{{.*}}", target_cpu: "generic-rv64")
+// REGULAR-IR-NEXT: "nop"
+// REGULAR-IR-NEXT: ".symver symver_fn, symver_fn at VER_1.0"
+//
+// THIN-IR: module asm{{$}}
+// THIN-IR-NEXT: ".symver symver_fn, symver_fn at VER_1.0"
+//
+// FLAGS: Flags [ (0x5)
+// FLAGS-NEXT: EF_RISCV_FLOAT_ABI_DOUBLE (0x4)
+// FLAGS-NEXT: EF_RISCV_RVC (0x1)
+// FLAGS-NEXT: ]
+//
+/// TODO: RISCVTargetELFStreamer::emitTextAttribute does not update the
+/// streamer's ArchString when emitting the module's RISCVAttrs::ARCH attribute
+/// ("rv64i2p1_m2p0_a2p1_f2p2_d2p2_c2p0_v1p0_...").
+// DISASM-LABEL: Disassembly of section .text:
+// DISASM-EMPTY:
+// DISASM-NEXT: [[#%x,]] <$xrv64i2p1>:
+// DISASM-NEXT: [[#%x,]]: nop
+// DISASM-EMPTY:
+// DISASM-NEXT: [[#%x,]] <$xrv64i2p1>:
+// DISASM-NEXT: [[#%x,]] <symver_fn>:
+// DISASM-NEXT: [[#%x,]]: ret
+// DISASM-EMPTY:
+// DISASM-NEXT: [[#%x,]] <$xrv64i2p1>:
+// DISASM-NEXT: [[#%x,]] <fn>:
+// DISASM-NEXT: [[#%x,]]: nop
+// DISASM-NEXT: [[#%x,]]: ret
+// DISASM-EMPTY:
+// DISASM-NEXT: [[#%x,]] <$xrv64i2p1>:
+// DISASM-NEXT: [[#%x,]] <caller>:
+// DISASM-NEXT: [[#%x,]]: nop
+// DISASM-NEXT: [[#%x,]]: ret
+// DISASM-NOT: {{.}}
+//
+// SYMS: [[#%x,]] l .text 0000000000000000 $xrv64i2p1{{$}}
+// SYMS: [[#%x,]] l .text 0000000000000000 $xrv64i2p1{{$}}
+// SYMS: [[#%x,]] l .text 0000000000000000 $xrv64i2p1{{$}}
+// SYMS: [[#%x,]] l .text 0000000000000000 $xrv64i2p1{{$}}
+// SYMS: [[#%x,]] g F .text 0000000000000004 fn{{$}}
+// SYMS: [[#%x,]] g F .text 0000000000000002 symver_fn{{$}}
+// SYMS: [[#%x,]] g F .text 0000000000000004 caller{{$}}
+
+//--- ver.ver
+VER_1.0 {};
+
+//--- a.c
+__asm__("nop");
+__asm__(".symver symver_fn, symver_fn at VER_1.0");
+
+void symver_fn(void) {}
+
+void fn(void) { __asm__ volatile("nop"); }
+
+//--- b.c
+extern void fn(void);
+extern void symver_fn(void);
+void caller(void) {
+ fn();
+ symver_fn();
+}
diff --git a/lld/test/ELF/lto/riscv-target-abi.ll b/lld/test/ELF/lto/riscv-target-abi.ll
index 23f722e91a6fb..3859377445a8d 100644
--- a/lld/test/ELF/lto/riscv-target-abi.ll
+++ b/lld/test/ELF/lto/riscv-target-abi.ll
@@ -1,40 +1,44 @@
; REQUIRES: riscv
+; RUN: rm -rf %t && split-file %s %t
-;; The module flag asks for lp64d, but without -mcpu we default to no D extension,
-;; so we print a warning and ignore the module flag.
-; RUN: llvm-as %s -o %t.bc
-; RUN: ld.lld -shared %t.bc -o %t.so 2>&1 | FileCheck %s --check-prefix=WARN \
-; RUN: --implicit-check-not="ignoring target-abi" --implicit-check-not="error:" --implicit-check-not="warning:"
; WARN: note: hard-float 'd' ABI can't be used for a target that doesn't support the D instruction set extension (ignoring target-abi)
+; NOWARN-NOT: ignoring target-abi
+
+; FLAGS-ABI-IGNORED: Flags [ (0x4)
+; FLAGS-ABI-IGNORED-NEXT: EF_RISCV_FLOAT_ABI_DOUBLE (0x4)
+; FLAGS-ABI-IGNORED-NEXT: ]
+
+; FLAGS-MCPU: Flags [ (0x5)
+; FLAGS-MCPU-NEXT: EF_RISCV_FLOAT_ABI_DOUBLE (0x4)
+; FLAGS-MCPU-NEXT: EF_RISCV_RVC (0x1)
+; FLAGS-MCPU-NEXT: ]
+
+;--- no-ext.ll
+;; The module flag asks for lp64d, and _start() has no target-features attribute.
+;; Without -mcpu we default to no D extension, so RISCVSubtarget prints a note
+;; and ignores the module flag.
+; RUN: llvm-as %t/no-ext.ll -o %t/no-ext.bc
+; RUN: ld.lld -shared %t/no-ext.bc -o %t/no-ext.so 2>&1 | FileCheck %s --check-prefix=WARN \
+; RUN: --implicit-check-not="ignoring target-abi" --implicit-check-not="error:" --implicit-check-not="warning:"
;; TODO: This is inconsistent: RISCVAsmPrinter::emitStartOfAsmFile sets e_flags
;; based on the raw module flag not the ABI actually used for codegen.
;; This means we are setting EF_RISCV_FLOAT_ABI_DOUBLE on a file built for soft float ABI
-; RUN: llvm-readobj --file-headers %t.so | FileCheck %s --check-prefix=FLAGS-ABI-IGNORED
-; FLAGS-ABI-IGNORED: Flags [ (0x4)
-; FLAGS-ABI-IGNORED-NEXT: EF_RISCV_FLOAT_ABI_DOUBLE (0x4)
-; FLAGS-ABI-IGNORED-NEXT: ]
+; RUN: llvm-readobj --file-headers %t/no-ext.so | FileCheck %s --check-prefix=FLAGS-ABI-IGNORED
-;; Passing -mcpu that has D makes the ABI valid again, so no warning.
-; RUN: ld.lld -mllvm -mcpu=sifive-u74 -shared %t.bc -o %t.so 2>&1 | FileCheck %s --check-prefix=NOWARN --allow-empty \
-; RUN: --implicit-check-not="error:" --implicit-check-not="warning:"
-; RUN: llvm-readobj --file-headers %t.so | FileCheck %s --check-prefix=FLAGS-MCPU
-; RUN: ld.lld -plugin-opt=mcpu=sifive-u74 -shared %t.bc -o %t.so 2>&1 | FileCheck %s --check-prefix=NOWARN --allow-empty \
-; RUN: --implicit-check-not="error:" --implicit-check-not="warning:"
-; RUN: llvm-readobj --file-headers %t.so | FileCheck %s --check-prefix=FLAGS-MCPU
-; NOWARN-NOT: ignoring target-abi
-; FLAGS-MCPU: Flags [ (0x5)
-; FLAGS-MCPU-NEXT: EF_RISCV_FLOAT_ABI_DOUBLE (0x4)
-; FLAGS-MCPU-NEXT: EF_RISCV_RVC (0x1)
-; FLAGS-MCPU-NEXT: ]
+;; Passing -mcpu that has D makes the ABI valid again, so no warning/note.
+; RUN: ld.lld -mllvm -mcpu=sifive-u74 -shared %t/no-ext.bc -o %t/no-ext.so 2>&1 | FileCheck %s --check-prefix=NOWARN --allow-empty \
+; RUN: --implicit-check-not="error:" --implicit-check-not="warning:" --implicit-check-not="note:"
+; RUN: llvm-readobj --file-headers %t/no-ext.so | FileCheck %s --check-prefix=FLAGS-MCPU
+; RUN: ld.lld -plugin-opt=mcpu=sifive-u74 -shared %t/no-ext.bc -o %t/no-ext.so 2>&1 | FileCheck %s --check-prefix=NOWARN --allow-empty \
+; RUN: --implicit-check-not="error:" --implicit-check-not="warning:" --implicit-check-not="note:"
+; RUN: llvm-readobj --file-headers %t/no-ext.so | FileCheck %s --check-prefix=FLAGS-MCPU
-target datalayout = "e-m:e-p:64:64-i64:64-i128:128-n64-S128"
+target datalayout = "e-m:e-p:64:64-i64:64-i128:128-n32:64-S128"
target triple = "riscv64"
module asm "nop"
-;; Module asm with target features not including 'd' (would fail before fix)
module asm(target_features: "+c") "c.nop"
-;; Module asm with target features enabling 'd'
module asm(target_features: "+d") "fld f0, 0(sp)"
define void @_start() {
@@ -44,3 +48,72 @@ define void @_start() {
!llvm.module.flags = !{!0}
!0 = !{i32 1, !"target-abi", !"lp64d"}
+
+;--- fn-inline-asm-no-ext.ll
+;; Function-level inline asm without +d on the function emits the missing D
+;; note once from RISCVSubtarget, without re-validating target-abi in RISCVAsmParser.
+; RUN: llvm-as %t/fn-inline-asm-no-ext.ll -o %t/fn-inline-asm-no-ext.bc
+; RUN: ld.lld -shared %t/fn-inline-asm-no-ext.bc -o %t/fn-inline-asm-no-ext.so 2>&1 | FileCheck %s --check-prefix=WARN \
+; RUN: --implicit-check-not="ignoring target-abi" --implicit-check-not="error:" --implicit-check-not="warning:"
+
+target datalayout = "e-m:e-p:64:64-i64:64-i128:128-n32:64-S128"
+target triple = "riscv64"
+
+define void @_start() {
+ call void asm sideeffect "nop", ""()
+ ret void
+}
+
+!llvm.module.flags = !{!0}
+!0 = !{i32 1, !"target-abi", !"lp64d"}
+
+;--- module-asm-no-ext.ll
+;; Module-level inline asm without target_features does not re-validate
+;; target-abi in RISCVAsmParser when functions in the module have +f,+d.
+; RUN: llvm-as %t/module-asm-no-ext.ll -o %t/module-asm-no-ext.bc
+; RUN: ld.lld -plugin-opt=mcpu=generic-rv64 -shared %t/module-asm-no-ext.bc -o %t/module-asm-no-ext.so 2>&1 \
+; RUN: | FileCheck %s --check-prefix=NOWARN --allow-empty \
+; RUN: --implicit-check-not="ignoring target-abi" --implicit-check-not="error:" --implicit-check-not="warning:" --implicit-check-not="note:"
+; RUN: ld.lld -plugin-opt=mcpu=sifive-u74 -shared %t/module-asm-no-ext.bc -o %t/module-asm-no-ext.so 2>&1 \
+; RUN: | FileCheck %s --check-prefix=NOWARN --allow-empty \
+; RUN: --implicit-check-not="error:" --implicit-check-not="warning:" --implicit-check-not="note:"
+
+target datalayout = "e-m:e-p:64:64-i64:64-i128:128-n32:64-S128"
+target triple = "riscv64"
+
+module asm "nop"
+
+define void @_start() #0 {
+ ret void
+}
+attributes #0 = { "target-features"="+f,+d" }
+
+!llvm.module.flags = !{!0}
+!0 = !{i32 1, !"target-abi", !"lp64d"}
+
+;--- module-asm-abi.ll
+;; Regression test for https://github.com/llvm/llvm-project/pull/213410:
+;; Module asm and function target-features specifying +c,+d should link cleanly
+;; even when the LTO backend is invoked with -plugin-opt=mcpu=generic-rv64.
+; RUN: llvm-as %t/module-asm-abi.ll -o %t/module-asm-abi.bc
+; RUN: ld.lld -plugin-opt=mcpu=generic-rv64 -shared %t/module-asm-abi.bc -o %t/module-asm-abi.so 2>&1 \
+; RUN: | FileCheck %s --check-prefix=NOWARN --allow-empty \
+; RUN: --implicit-check-not="error:" --implicit-check-not="warning:" --implicit-check-not="note:"
+; RUN: llvm-readobj --file-headers %t/module-asm-abi.so | FileCheck %s --check-prefix=FLAGS-MCPU
+
+target datalayout = "e-m:e-p:64:64-i64:64-i128:128-n32:64-S128"
+target triple = "riscv64"
+
+module asm(target_features: "+c,+d")
+ "nop"
+
+define void @_start() #0 {
+ call void asm sideeffect "nop", ""()
+ ret void
+}
+attributes #0 = { "target-features"="+c,+d" }
+
+!llvm.module.flags = !{!0, !1}
+!0 = !{i32 1, !"target-abi", !"lp64d"}
+!1 = !{i32 6, !"riscv-isa", !2}
+!2 = !{!"rv64i2p1_c2p0_d2p2"}
diff --git a/llvm/test/CodeGen/RISCV/module-asm-features.ll b/llvm/test/CodeGen/RISCV/module-asm-features.ll
index ab16acab48688..7fee7188e0ed8 100644
--- a/llvm/test/CodeGen/RISCV/module-asm-features.ll
+++ b/llvm/test/CodeGen/RISCV/module-asm-features.ll
@@ -1,14 +1,29 @@
; RUN: llc -mtriple=riscv64-unknown-linux-gnu < %s | FileCheck %s --check-prefixes=CHECK,EXTRA-FEATURES
; RUN: llc -mtriple=riscv64-unknown-linux-gnu -mattr=+d < %s | FileCheck %s --check-prefixes=CHECK,SAME-FEATURES
+; RUN: llc -mtriple=riscv64-unknown-linux-gnu -filetype=obj < %s | llvm-objdump -d --show-all-symbols --no-show-raw-insn - | FileCheck %s --check-prefix=OBJ
; This should work fine, because the module asm specifies the necessary
; target features
; SAME-FEATURES-NOT: .option arch
-; EXTRA-FEATURES: .option push
-; EXTRA-FEATURES: .option arch, +d
-; CHECK: fld ft0, 0(sp)
-; EXTRA-FEATURES: .option pop
+; EXTRA-FEATURES: .option push
+; EXTRA-FEATURES-NEXT: .option arch, +d, +f, +zicsr{{$}}
+; CHECK: .globl func
+; CHECK-NEXT: func:
+; CHECK-NEXT: fld ft0, 0(sp)
+; CHECK-NEXT: ret
+; EXTRA-FEATURES-NEXT: .option pop
+
+;; TODO: emitTargetFeaturePush does not call setArchString(), so the mapping
+;; symbol does not record +d/+f/+zicsr when assembling directly to an object
+;; file, causing llvm-objdump to fail to disassemble `fld`.
+; OBJ-LABEL: Disassembly of section .text:
+; OBJ-EMPTY:
+; OBJ-NEXT: 0000000000000000 <$xrv64i2p1>:
+; OBJ-NEXT: 0000000000000000 <func>:
+; OBJ-NEXT: 0: <unknown>
+; OBJ-NEXT: 4: ret
+; OBJ-NOT: {{.}}
module asm(target_features: "+d")
".globl func"
diff --git a/llvm/test/CodeGen/RISCV/riscv-func-target-feature.ll b/llvm/test/CodeGen/RISCV/riscv-func-target-feature.ll
index d627ae9c90394..de3de8c27df1d 100644
--- a/llvm/test/CodeGen/RISCV/riscv-func-target-feature.ll
+++ b/llvm/test/CodeGen/RISCV/riscv-func-target-feature.ll
@@ -1,44 +1,72 @@
; RUN: llc -mtriple=riscv64 -mcpu=sifive-u74 -verify-machineinstrs < %s | FileCheck %s
+; RUN: llc -mtriple=riscv64 -mcpu=sifive-u74 -filetype=obj < %s \
+; RUN: | llvm-objdump -d --show-all-symbols --no-show-raw-insn - | FileCheck %s --check-prefix=OBJ
-; CHECK: .option push
-; CHECK-NEXT: .option arch, +v, +zve32f, +zve32x, +zve64d, +zve64f, +zve64x, +zvl128b, +zvl32b, +zvl64b
+;; TODO: emitTargetFeaturePush does not call setArchString(), so per-function
+;; target-features are not reflected in the $x<arch> mapping symbols.
+; OBJ-LABEL: Disassembly of section .text:
+; OBJ-EMPTY:
+; OBJ-NEXT: 0000000000000000 <$xrv64i2p1_m2p0_a2p1_f2p2_d2p2_c2p0_zicsr2p0_zifencei2p0_zmmul1p0_zaamo1p0_zalrsc1p0_zca1p0_zcd1p0>:
+; OBJ-NEXT: 0000000000000000 <test1>:
+; OBJ-NEXT: 0: ret
+; OBJ-EMPTY:
+; OBJ-NEXT: 0000000000000002 <test2>:
+; OBJ-NEXT: 2: ret
+; OBJ-EMPTY:
+; OBJ-NEXT: 0000000000000004 <test3>:
+; OBJ-NEXT: 4: ret
+; OBJ-EMPTY:
+; OBJ-NEXT: 0000000000000006 <test4>:
+; OBJ-NEXT: 6: ret
+; OBJ-EMPTY:
+; OBJ-NEXT: 0000000000000008 <test5>:
+; OBJ-NEXT: 8: ret
+; OBJ-NOT: {{.}}
+
+; CHECK: .option push
+; CHECK-NEXT: .option arch, +v, +zve32f, +zve32x, +zve64d, +zve64f, +zve64x, +zvl128b, +zvl32b, +zvl64b{{$}}
define void @test1() "target-features"="+a,+d,+f,+m,+c,+v,+zifencei,+zve32f,+zve32x,+zve64d,+zve64f,+zve64x,+zvl128b,+zvl32b,+zvl64b" {
-; CHECK-LABEL: test1
-; CHECK: .option pop
+; CHECK-LABEL: test1:
+; CHECK: ret
+; CHECK: .option pop
entry:
ret void
}
-; CHECK: .option push
-; CHECK-NEXT: .option arch, +zihintntl
+; CHECK-NEXT: .option push
+; CHECK-NEXT: .option arch, +zihintntl{{$}}
define void @test2() "target-features"="+a,+d,+f,+m,+zihintntl,+zifencei" {
-; CHECK-LABEL: test2
-; CHECK: .option pop
+; CHECK-LABEL: test2:
+; CHECK: ret
+; CHECK: .option pop
entry:
ret void
}
-; CHECK: .option push
-; CHECK-NEXT: .option arch, -a, -d, -f, -m
+; CHECK-NEXT: .option push
+; CHECK-NEXT: .option arch, -a, -d, -f, -m, -zcd{{$}}
define void @test3() "target-features"="-a,-d,-f,-m" {
-; CHECK-LABEL: test3
-; CHECK: .option pop
+; CHECK-LABEL: test3:
+; CHECK: ret
+; CHECK: .option pop
entry:
ret void
}
; CHECK-NOT: .option push
define void @test4() {
-; CHECK-LABEL: test4
-; CHECK-NOT: .option pop
+; CHECK-LABEL: test4:
+; CHECK: ret
+; CHECK-NOT: .option pop
entry:
ret void
}
; CHECK-NOT: .option push
define void @test5() "target-features"="+unaligned-scalar-mem" {
-; CHECK-LABEL: test5
-; CHECK-NOT: .option pop
+; CHECK-LABEL: test5:
+; CHECK: ret
+; CHECK-NOT: .option pop
entry:
ret void
}
diff --git a/llvm/test/LTO/RISCV/module-asm.ll b/llvm/test/LTO/RISCV/module-asm.ll
index 73320185e778a..e213ec14a2a58 100644
--- a/llvm/test/LTO/RISCV/module-asm.ll
+++ b/llvm/test/LTO/RISCV/module-asm.ll
@@ -2,11 +2,24 @@
; RUN: llvm-lto2 run -save-temps -filetype=asm -o %t.s %t.o -r=%t.o,func,p
; RUN: llvm-nm %t.o | FileCheck %s --check-prefix NM
; RUN: llvm-nm %t.s.0.5.precodegen.bc | FileCheck %s --check-prefix NM
+; RUN: llvm-dis %t.s.0.5.precodegen.bc -o - | FileCheck %s --check-prefix=IR
; RUN: FileCheck %s --input-file %t.s.0
; NM: T func
-; CHECK: fld ft0, 0(sp)
+;; TODO: LTO::addRegularLTO prepends ".lto_discard" without preserving the
+;; existing module inline asm's TargetCPU and TargetFeatures.
+; IR: module asm
+; IR-NEXT: ".lto_discard"
+; IR-NEXT: module asm(target_features: "+d")
+; IR-NEXT: ".globl func"
+; IR-NEXT: "func:"
+; IR-NEXT: "fld f0, 0(sp)"
+; IR-NEXT: "ret"
+
+; CHECK-LABEL: func:
+; CHECK-NEXT: fld ft0, 0(sp)
+; CHECK-NEXT: ret
target datalayout = "e-m:e-p:64:64-i64:64-i128:128-n32:64-S128"
target triple = "riscv64-unknown-linux-gnu"
diff --git a/llvm/test/MC/RISCV/mapping-across-sections.s b/llvm/test/MC/RISCV/mapping-across-sections.s
index ecb8292dd6664..9a741e792a244 100644
--- a/llvm/test/MC/RISCV/mapping-across-sections.s
+++ b/llvm/test/MC/RISCV/mapping-across-sections.s
@@ -18,6 +18,13 @@
.text
nop
+# Pushing a data section and popping back to .text should also preserve .text's
+# mapping symbol state and not emit a redundant $x.
+ .pushsection .starts_data
+ .word 42
+ .popsection
+ nop
+
# With all those constraints, we want:
# + .text to have $x<ISA> at 0 and no others
# + .wibble to have $x<ISA> at 0 (each code section records the active ISA
@@ -28,9 +35,17 @@
# CHECK: [[#WIBBLE:]]] .wibble
# CHECK: [[#STARTS_DATA:]]] .starts_data
-# CHECK: Value Size Type Bind Vis Ndx Name
-# CHECK-RV32: 00000000 0 NOTYPE LOCAL DEFAULT [[#TEXT]] $xrv32i2p1{{$}}
-# CHECK-RV64: 00000000 0 NOTYPE LOCAL DEFAULT [[#TEXT]] $xrv64i2p1{{$}}
-# CHECK-RV32: 00000000 0 NOTYPE LOCAL DEFAULT [[#WIBBLE]] $xrv32i2p1{{$}}
-# CHECK-RV64: 00000000 0 NOTYPE LOCAL DEFAULT [[#WIBBLE]] $xrv64i2p1{{$}}
-# CHECK: 00000000 0 NOTYPE LOCAL DEFAULT [[#STARTS_DATA]] $d{{$}}
+## TODO: RISCVELFStreamer::changeSection saves mapping symbol state to
+## getPreviousSection() instead of getCurrentSection() on popSection(), causing
+## a duplicate $x mapping symbol at offset 8 in .text.
+# CHECK: Symbol table '.symtab' contains 5 entries:
+# CHECK-NEXT: Num: Value Size Type Bind Vis Ndx Name
+# CHECK-NEXT: 0: {{0+}} 0 NOTYPE LOCAL DEFAULT UND {{$}}
+# CHECK-RV32-NEXT: 1: 00000000 0 NOTYPE LOCAL DEFAULT [[#TEXT]] $xrv32i2p1{{$}}
+# CHECK-RV64-NEXT: 1: {{0+}} 0 NOTYPE LOCAL DEFAULT [[#TEXT]] $xrv64i2p1{{$}}
+# CHECK-RV32-NEXT: 2: 00000000 0 NOTYPE LOCAL DEFAULT [[#WIBBLE]] $xrv32i2p1{{$}}
+# CHECK-RV64-NEXT: 2: {{0+}} 0 NOTYPE LOCAL DEFAULT [[#WIBBLE]] $xrv64i2p1{{$}}
+# CHECK-NEXT: 3: {{0+}} 0 NOTYPE LOCAL DEFAULT [[#STARTS_DATA]] $d{{$}}
+# CHECK-RV32-NEXT: 4: 00000008 0 NOTYPE LOCAL DEFAULT [[#TEXT]] $xrv32i2p1{{$}}
+# CHECK-RV64-NEXT: 4: {{0+}}8 0 NOTYPE LOCAL DEFAULT [[#TEXT]] $xrv64i2p1{{$}}
+# CHECK-NOT: {{.}}
>From f7de023a59f9a9aa73d6207cbfd75a0b00352481 Mon Sep 17 00:00:00 2001
From: Alexander Richardson <alexrichardson at google.com>
Date: Sat, 26 Sep 2026 19:01:44 -0700
Subject: [PATCH 31/53] [LTO] Preserve module inline asm target properties for
.lto_discard and symvers (#225130)
Previously, LTO::addRegularLTO() and IRLinker::run() called
prependModuleInlineAsm() and appendModuleInlineAsm() with a plain string when
synthesizing `.lto_discard` and imported `.symver` directives, creating a new
GlobalAsmFragment with empty TargetCPU and TargetFeatures instead of preserving
the existing module inline asm's properties. Copy the front fragment's Props so
these synthesized directives are merged into the module's inline asm with the
same target features.
This commit was created with the help of AI tools
Pull-Request: https://github.com/llvm/llvm-project/pull/225130
---
cross-project-tests/riscv/lto-inline-asm-abi.c | 7 ++-----
llvm/lib/LTO/LTO.cpp | 2 +-
llvm/lib/Linker/IRMover.cpp | 3 ++-
llvm/test/LTO/RISCV/module-asm.ll | 5 +----
4 files changed, 6 insertions(+), 11 deletions(-)
diff --git a/cross-project-tests/riscv/lto-inline-asm-abi.c b/cross-project-tests/riscv/lto-inline-asm-abi.c
index e89e1e4f0687d..883244f1854f7 100644
--- a/cross-project-tests/riscv/lto-inline-asm-abi.c
+++ b/cross-project-tests/riscv/lto-inline-asm-abi.c
@@ -22,15 +22,12 @@
// RUN: llvm-objdump -d --show-all-symbols --no-show-raw-insn %t.thin.so | FileCheck %s --check-prefix=DISASM
// RUN: llvm-objdump -t %t.thin.so | FileCheck %s --check-prefix=SYMS --implicit-check-not='\$x'
//
-/// TODO: LTO::addRegularLTO and IRLinker::run drop target_features and
-/// target_cpu when synthesizing .lto_discard and imported .symver directives.
-// REGULAR-IR: module asm{{$}}
+// REGULAR-IR: module asm(target_features: "+64bit,{{.*}}", target_cpu: "generic-rv64")
// REGULAR-IR-NEXT: ".lto_discard "
-// REGULAR-IR-NEXT: module asm(target_features: "+64bit,{{.*}}", target_cpu: "generic-rv64")
// REGULAR-IR-NEXT: "nop"
// REGULAR-IR-NEXT: ".symver symver_fn, symver_fn at VER_1.0"
//
-// THIN-IR: module asm{{$}}
+// THIN-IR: module asm(target_features: "+64bit,{{.*}}", target_cpu: "generic-rv64")
// THIN-IR-NEXT: ".symver symver_fn, symver_fn at VER_1.0"
//
// FLAGS: Flags [ (0x5)
diff --git a/llvm/lib/LTO/LTO.cpp b/llvm/lib/LTO/LTO.cpp
index 4594c52fb5f6e..e307b7f1c16e8 100644
--- a/llvm/lib/LTO/LTO.cpp
+++ b/llvm/lib/LTO/LTO.cpp
@@ -1129,7 +1129,7 @@ LTO::addRegularLTO(InputFile &Input, ArrayRef<SymbolResolution> InputRes,
NewIA += " " + llvm::join(NonPrevailingAsmSymbols, ", ");
}
NewIA += "\n";
- M.prependModuleInlineAsm(NewIA);
+ M.prependModuleInlineAsm({NewIA, M.getModuleInlineAsm().front().Props});
}
assert(MsymI == MsymE);
diff --git a/llvm/lib/Linker/IRMover.cpp b/llvm/lib/Linker/IRMover.cpp
index 3b72b412d0b2e..d96c5d18a0ae3 100644
--- a/llvm/lib/Linker/IRMover.cpp
+++ b/llvm/lib/Linker/IRMover.cpp
@@ -1576,7 +1576,8 @@ Error IRLinker::run() {
S += Name;
S += ", ";
S += Alias;
- DstM.appendModuleInlineAsm(std::string(S));
+ DstM.appendModuleInlineAsm(
+ {std::string(S), SrcM->getModuleInlineAsm().back().Props});
}
});
}
diff --git a/llvm/test/LTO/RISCV/module-asm.ll b/llvm/test/LTO/RISCV/module-asm.ll
index e213ec14a2a58..f27943f5457ac 100644
--- a/llvm/test/LTO/RISCV/module-asm.ll
+++ b/llvm/test/LTO/RISCV/module-asm.ll
@@ -7,11 +7,8 @@
; NM: T func
-;; TODO: LTO::addRegularLTO prepends ".lto_discard" without preserving the
-;; existing module inline asm's TargetCPU and TargetFeatures.
-; IR: module asm
+; IR: module asm(target_features: "+d")
; IR-NEXT: ".lto_discard"
-; IR-NEXT: module asm(target_features: "+d")
; IR-NEXT: ".globl func"
; IR-NEXT: "func:"
; IR-NEXT: "fld f0, 0(sp)"
>From 9993a7e62ba4d371311f13b96dbe03176b83e849 Mon Sep 17 00:00:00 2001
From: Alexander Richardson <alexrichardson at google.com>
Date: Sat, 26 Sep 2026 19:01:58 -0700
Subject: [PATCH 32/53] [RISC-V][MC] Fix mapping symbol section tracking on
popSection() (#225131)
Previously, RISCVELFStreamer::changeSection() saved LastEMS and LastEmittedArch
under getPreviousSection().first instead of getCurrentSection().first. When
MCStreamer::popSection() switches back to a previous section,
getPreviousSection() already points to the destination section being restored
rather than the section being exited. This clobbered the destination section's
saved mapping symbol state and caused duplicate `$x<arch>` mapping symbols to
be emitted whenever returning to `.text`.
Use getCurrentSection().first instead, matching AArch64ELFStreamer and
ARMELFStreamer.
This commit was created with the help of AI tools
Pull-Request: https://github.com/llvm/llvm-project/pull/225131
---
llvm/lib/Target/RISCV/MCTargetDesc/RISCVELFStreamer.cpp | 6 +++---
llvm/test/MC/RISCV/mapping-across-sections.s | 7 +------
2 files changed, 4 insertions(+), 9 deletions(-)
diff --git a/llvm/lib/Target/RISCV/MCTargetDesc/RISCVELFStreamer.cpp b/llvm/lib/Target/RISCV/MCTargetDesc/RISCVELFStreamer.cpp
index 524d0bb79d5f8..10702a836de33 100644
--- a/llvm/lib/Target/RISCV/MCTargetDesc/RISCVELFStreamer.cpp
+++ b/llvm/lib/Target/RISCV/MCTargetDesc/RISCVELFStreamer.cpp
@@ -218,9 +218,9 @@ void RISCVELFStreamer::changeSection(MCSection *Section, uint32_t Subsection) {
// default constructor by DenseMap::lookup. The last ISA suffix emitted in
// each section is also preserved so that re-entering a section only emits a
// new "$x<ISA>" symbol when the active ISA has actually changed.
- const MCSection *Prev = getPreviousSection().first;
- LastMappingSymbols[Prev] = LastEMS;
- LastEmittedArchInSection[Prev] = LastEmittedArch;
+ const MCSection *Cur = getCurrentSection().first;
+ LastMappingSymbols[Cur] = LastEMS;
+ LastEmittedArchInSection[Cur] = LastEmittedArch;
LastEMS = LastMappingSymbols.lookup(Section);
auto It = LastEmittedArchInSection.find(Section);
LastEmittedArch = It != LastEmittedArchInSection.end() ? It->second : "";
diff --git a/llvm/test/MC/RISCV/mapping-across-sections.s b/llvm/test/MC/RISCV/mapping-across-sections.s
index 9a741e792a244..54c420085da12 100644
--- a/llvm/test/MC/RISCV/mapping-across-sections.s
+++ b/llvm/test/MC/RISCV/mapping-across-sections.s
@@ -35,10 +35,7 @@
# CHECK: [[#WIBBLE:]]] .wibble
# CHECK: [[#STARTS_DATA:]]] .starts_data
-## TODO: RISCVELFStreamer::changeSection saves mapping symbol state to
-## getPreviousSection() instead of getCurrentSection() on popSection(), causing
-## a duplicate $x mapping symbol at offset 8 in .text.
-# CHECK: Symbol table '.symtab' contains 5 entries:
+# CHECK: Symbol table '.symtab' contains 4 entries:
# CHECK-NEXT: Num: Value Size Type Bind Vis Ndx Name
# CHECK-NEXT: 0: {{0+}} 0 NOTYPE LOCAL DEFAULT UND {{$}}
# CHECK-RV32-NEXT: 1: 00000000 0 NOTYPE LOCAL DEFAULT [[#TEXT]] $xrv32i2p1{{$}}
@@ -46,6 +43,4 @@
# CHECK-RV32-NEXT: 2: 00000000 0 NOTYPE LOCAL DEFAULT [[#WIBBLE]] $xrv32i2p1{{$}}
# CHECK-RV64-NEXT: 2: {{0+}} 0 NOTYPE LOCAL DEFAULT [[#WIBBLE]] $xrv64i2p1{{$}}
# CHECK-NEXT: 3: {{0+}} 0 NOTYPE LOCAL DEFAULT [[#STARTS_DATA]] $d{{$}}
-# CHECK-RV32-NEXT: 4: 00000008 0 NOTYPE LOCAL DEFAULT [[#TEXT]] $xrv32i2p1{{$}}
-# CHECK-RV64-NEXT: 4: {{0+}}8 0 NOTYPE LOCAL DEFAULT [[#TEXT]] $xrv64i2p1{{$}}
# CHECK-NOT: {{.}}
>From 9a07a6b8d36a5e424dc2b8df8289f6d07e39097b Mon Sep 17 00:00:00 2001
From: Alexander Richardson <alexrichardson at google.com>
Date: Sat, 26 Sep 2026 19:02:12 -0700
Subject: [PATCH 33/53] [RISC-V][MC] Update ELF streamer ArchString in
setFlagsFromFeatures() (#225140)
During LTO, the TargetMachine subtarget is initialized with the linker's
default CPU (e.g. `generic-rv64`, `rv64i2p1`), while
RISCVAsmPrinter::emitStartOfAsmFile() reconstructs the module's actual ISA from
the `riscv-isa` module flag and calls RISCVTargetStreamer::setFlagsFromFeatures()
and RISCVTargetStreamer::emitTargetAttributes(). Because
RISCVTargetELFStreamer previously only initialized InitialArchString and
ArchString in its constructor rather than in setFlagsFromFeatures(), direct
object emission bypassed the update and tagged `.text` with `$xrv64i2p1`
instead of the module's full architecture string.
Move the InitialArchString and ArchString initialization into
RISCVTargetELFStreamer::setFlagsFromFeatures() and also call setArchString()
alongside emitTextAttribute(RISCVAttrs::ARCH, ...) in
RISCVTargetStreamer::emitTargetAttributes().
This commit was created with the help of AI tools
Pull-Request: https://github.com/llvm/llvm-project/pull/225140
---
.../riscv/lto-inline-asm-abi.c | 19 +++++-------
.../RISCV/MCTargetDesc/RISCVELFStreamer.cpp | 29 +++++++++++--------
.../RISCV/MCTargetDesc/RISCVELFStreamer.h | 5 ++--
.../MCTargetDesc/RISCVTargetStreamer.cpp | 4 ++-
.../RISCV/MCTargetDesc/RISCVTargetStreamer.h | 2 +-
5 files changed, 32 insertions(+), 27 deletions(-)
diff --git a/cross-project-tests/riscv/lto-inline-asm-abi.c b/cross-project-tests/riscv/lto-inline-asm-abi.c
index 883244f1854f7..897413428500f 100644
--- a/cross-project-tests/riscv/lto-inline-asm-abi.c
+++ b/cross-project-tests/riscv/lto-inline-asm-abi.c
@@ -35,33 +35,30 @@
// FLAGS-NEXT: EF_RISCV_RVC (0x1)
// FLAGS-NEXT: ]
//
-/// TODO: RISCVTargetELFStreamer::emitTextAttribute does not update the
-/// streamer's ArchString when emitting the module's RISCVAttrs::ARCH attribute
-/// ("rv64i2p1_m2p0_a2p1_f2p2_d2p2_c2p0_v1p0_...").
// DISASM-LABEL: Disassembly of section .text:
// DISASM-EMPTY:
-// DISASM-NEXT: [[#%x,]] <$xrv64i2p1>:
+// DISASM-NEXT: [[#%x,]] <$xrv64i2p1_m2p0_a2p1_f2p2_d2p2_c2p0_v1p0_zicsr2p0_zifencei2p0_zmmul1p0_zaamo1p0_zalrsc1p0_zca1p0_zcd1p0_zve32f1p0_zve32x1p0_zve64d1p0_zve64f1p0_zve64x1p0_zvl128b1p0_zvl32b1p0_zvl64b1p0>:
// DISASM-NEXT: [[#%x,]]: nop
// DISASM-EMPTY:
-// DISASM-NEXT: [[#%x,]] <$xrv64i2p1>:
+// DISASM-NEXT: [[#%x,]] <$xrv64i2p1_m2p0_a2p1_f2p2_d2p2_c2p0_v1p0_zicsr2p0_zifencei2p0_zmmul1p0_zaamo1p0_zalrsc1p0_zca1p0_zcd1p0_zve32f1p0_zve32x1p0_zve64d1p0_zve64f1p0_zve64x1p0_zvl128b1p0_zvl32b1p0_zvl64b1p0>:
// DISASM-NEXT: [[#%x,]] <symver_fn>:
// DISASM-NEXT: [[#%x,]]: ret
// DISASM-EMPTY:
-// DISASM-NEXT: [[#%x,]] <$xrv64i2p1>:
+// DISASM-NEXT: [[#%x,]] <$xrv64i2p1_m2p0_a2p1_f2p2_d2p2_c2p0_v1p0_zicsr2p0_zifencei2p0_zmmul1p0_zaamo1p0_zalrsc1p0_zca1p0_zcd1p0_zve32f1p0_zve32x1p0_zve64d1p0_zve64f1p0_zve64x1p0_zvl128b1p0_zvl32b1p0_zvl64b1p0>:
// DISASM-NEXT: [[#%x,]] <fn>:
// DISASM-NEXT: [[#%x,]]: nop
// DISASM-NEXT: [[#%x,]]: ret
// DISASM-EMPTY:
-// DISASM-NEXT: [[#%x,]] <$xrv64i2p1>:
+// DISASM-NEXT: [[#%x,]] <$xrv64i2p1_m2p0_a2p1_f2p2_d2p2_c2p0_v1p0_zicsr2p0_zifencei2p0_zmmul1p0_zaamo1p0_zalrsc1p0_zca1p0_zcd1p0_zve32f1p0_zve32x1p0_zve64d1p0_zve64f1p0_zve64x1p0_zvl128b1p0_zvl32b1p0_zvl64b1p0>:
// DISASM-NEXT: [[#%x,]] <caller>:
// DISASM-NEXT: [[#%x,]]: nop
// DISASM-NEXT: [[#%x,]]: ret
// DISASM-NOT: {{.}}
//
-// SYMS: [[#%x,]] l .text 0000000000000000 $xrv64i2p1{{$}}
-// SYMS: [[#%x,]] l .text 0000000000000000 $xrv64i2p1{{$}}
-// SYMS: [[#%x,]] l .text 0000000000000000 $xrv64i2p1{{$}}
-// SYMS: [[#%x,]] l .text 0000000000000000 $xrv64i2p1{{$}}
+// SYMS: [[#%x,]] l .text 0000000000000000 $xrv64i2p1_m2p0_a2p1_f2p2_d2p2_c2p0_v1p0_zicsr2p0_zifencei2p0_zmmul1p0_zaamo1p0_zalrsc1p0_zca1p0_zcd1p0_zve32f1p0_zve32x1p0_zve64d1p0_zve64f1p0_zve64x1p0_zvl128b1p0_zvl32b1p0_zvl64b1p0{{$}}
+// SYMS: [[#%x,]] l .text 0000000000000000 $xrv64i2p1_m2p0_a2p1_f2p2_d2p2_c2p0_v1p0_zicsr2p0_zifencei2p0_zmmul1p0_zaamo1p0_zalrsc1p0_zca1p0_zcd1p0_zve32f1p0_zve32x1p0_zve64d1p0_zve64f1p0_zve64x1p0_zvl128b1p0_zvl32b1p0_zvl64b1p0{{$}}
+// SYMS: [[#%x,]] l .text 0000000000000000 $xrv64i2p1_m2p0_a2p1_f2p2_d2p2_c2p0_v1p0_zicsr2p0_zifencei2p0_zmmul1p0_zaamo1p0_zalrsc1p0_zca1p0_zcd1p0_zve32f1p0_zve32x1p0_zve64d1p0_zve64f1p0_zve64x1p0_zvl128b1p0_zvl32b1p0_zvl64b1p0{{$}}
+// SYMS: [[#%x,]] l .text 0000000000000000 $xrv64i2p1_m2p0_a2p1_f2p2_d2p2_c2p0_v1p0_zicsr2p0_zifencei2p0_zmmul1p0_zaamo1p0_zalrsc1p0_zca1p0_zcd1p0_zve32f1p0_zve32x1p0_zve64d1p0_zve64f1p0_zve64x1p0_zvl128b1p0_zvl32b1p0_zvl64b1p0{{$}}
// SYMS: [[#%x,]] g F .text 0000000000000004 fn{{$}}
// SYMS: [[#%x,]] g F .text 0000000000000002 symver_fn{{$}}
// SYMS: [[#%x,]] g F .text 0000000000000004 caller{{$}}
diff --git a/llvm/lib/Target/RISCV/MCTargetDesc/RISCVELFStreamer.cpp b/llvm/lib/Target/RISCV/MCTargetDesc/RISCVELFStreamer.cpp
index 10702a836de33..11093a700ac27 100644
--- a/llvm/lib/Target/RISCV/MCTargetDesc/RISCVELFStreamer.cpp
+++ b/llvm/lib/Target/RISCV/MCTargetDesc/RISCVELFStreamer.cpp
@@ -28,18 +28,6 @@ RISCVTargetELFStreamer::RISCVTargetELFStreamer(MCStreamer &S,
const MCSubtargetInfo &STI)
: RISCVTargetStreamer(S), CurrentVendor("riscv") {
setFlagsFromFeatures(STI);
-
- // Compute the initial ISA string. This serves two purposes:
- // 1. Deduplication: subsequent .option arch/rvc/norvc directives compare
- // against ArchString to avoid propagating redundant ISA updates.
- // 2. Initial symbol: seed the streamer's active ISA so a "$x<ArchString>"
- // mapping symbol is emitted before the first instruction, recording
- // the full ISA in the object even when no .option directive is present.
- if (auto ParseResult = RISCVFeatures::parseFeatureBits(STI)) {
- InitialArchString = (*ParseResult)->toString();
- ArchString = InitialArchString;
- getStreamer().setMappingSymbolArch(ArchString);
- }
}
RISCVELFStreamer::RISCVELFStreamer(MCContext &C,
@@ -52,6 +40,23 @@ RISCVELFStreamer &RISCVTargetELFStreamer::getStreamer() {
return static_cast<RISCVELFStreamer &>(Streamer);
}
+void RISCVTargetELFStreamer::setFlagsFromFeatures(const MCSubtargetInfo &STI) {
+ RISCVTargetStreamer::setFlagsFromFeatures(STI);
+
+ // Compute the initial ISA string. This serves two purposes:
+ // 1. Deduplication: subsequent .option arch/rvc/norvc directives compare
+ // against ArchString to avoid propagating redundant ISA updates.
+ // 2. Initial symbol: seed the streamer's active ISA so a "$x<ArchString>"
+ // mapping symbol is emitted before the first instruction, recording
+ // the full ISA in the object even when no .option directive is present.
+ if (auto ParseResult = RISCVFeatures::parseFeatureBits(STI)) {
+ InitialArchString = (*ParseResult)->toString();
+ setArchString(InitialArchString);
+ } else {
+ consumeError(ParseResult.takeError());
+ }
+}
+
void RISCVTargetELFStreamer::setArchString(StringRef Arch) {
if (Arch == ArchString)
return;
diff --git a/llvm/lib/Target/RISCV/MCTargetDesc/RISCVELFStreamer.h b/llvm/lib/Target/RISCV/MCTargetDesc/RISCVELFStreamer.h
index ac738307922d7..e86144ac20c9b 100644
--- a/llvm/lib/Target/RISCV/MCTargetDesc/RISCVELFStreamer.h
+++ b/llvm/lib/Target/RISCV/MCTargetDesc/RISCVELFStreamer.h
@@ -56,8 +56,8 @@ class RISCVTargetELFStreamer : public RISCVTargetStreamer {
private:
StringRef CurrentVendor;
- // Initial ISA string derived from the subtarget features in the constructor.
- // Used to re-establish state on reset().
+ // Initial ISA string derived from the subtarget features in
+ // setFlagsFromFeatures(). Used to re-establish state on reset().
std::string InitialArchString;
// Current ISA string, kept in sync with each .option arch/rvc/norvc/pop
@@ -81,6 +81,7 @@ class RISCVTargetELFStreamer : public RISCVTargetStreamer {
RISCVELFStreamer &getStreamer();
RISCVTargetELFStreamer(MCStreamer &S, const MCSubtargetInfo &STI);
+ void setFlagsFromFeatures(const MCSubtargetInfo &STI) override;
// Update ArchString and propagate the change to the streamer so the next
// instruction-run emits an ISA-specific mapping symbol. A no-op when
// Arch == ArchString (deduplication).
diff --git a/llvm/lib/Target/RISCV/MCTargetDesc/RISCVTargetStreamer.cpp b/llvm/lib/Target/RISCV/MCTargetDesc/RISCVTargetStreamer.cpp
index d6523cb42c879..f5f62bc5200f1 100644
--- a/llvm/lib/Target/RISCV/MCTargetDesc/RISCVTargetStreamer.cpp
+++ b/llvm/lib/Target/RISCV/MCTargetDesc/RISCVTargetStreamer.cpp
@@ -86,7 +86,9 @@ void RISCVTargetStreamer::emitTargetAttributes(const MCSubtargetInfo &STI,
report_fatal_error(ParseResult.takeError());
} else {
auto &ISAInfo = *ParseResult;
- emitTextAttribute(RISCVAttrs::ARCH, ISAInfo->toString());
+ std::string Arch = ISAInfo->toString();
+ emitTextAttribute(RISCVAttrs::ARCH, Arch);
+ setArchString(Arch);
}
if (RiscvAbiAttr && STI.hasFeature(RISCV::FeatureStdExtA)) {
diff --git a/llvm/lib/Target/RISCV/MCTargetDesc/RISCVTargetStreamer.h b/llvm/lib/Target/RISCV/MCTargetDesc/RISCVTargetStreamer.h
index cd59089fae4c7..c7b4f752d7264 100644
--- a/llvm/lib/Target/RISCV/MCTargetDesc/RISCVTargetStreamer.h
+++ b/llvm/lib/Target/RISCV/MCTargetDesc/RISCVTargetStreamer.h
@@ -64,7 +64,7 @@ class RISCVTargetStreamer : public MCTargetStreamer {
void setTargetABI(RISCVABI::ABI ABI);
RISCVABI::ABI getTargetABI() const { return TargetABI; }
bool hasTargetABI() const { return TargetABI != RISCVABI::ABI_Unknown; }
- void setFlagsFromFeatures(const MCSubtargetInfo &STI);
+ virtual void setFlagsFromFeatures(const MCSubtargetInfo &STI);
bool hasRVC() const { return HasRVC; }
bool hasTSO() const { return HasTSO; }
};
>From f0c1697b06aece8a94aca2a134b594f1667ad27a Mon Sep 17 00:00:00 2001
From: Alexander Richardson <alexrichardson at google.com>
Date: Sat, 26 Sep 2026 19:02:23 -0700
Subject: [PATCH 34/53] [RISC-V] Update streamer ArchString in
emitTargetFeaturePush() (#225133)
Previously, RISCVAsmPrinter::emitTargetFeaturePush() only emitted `.option push`
and `.option arch` without updating the streamer's active ArchString. When
emitting an ELF object file directly (`-filetype=obj`),
RISCVTargetELFStreamer::emitDirectiveOptionArch() is a no-op while
emitTargetFeaturePop() resets ArchString back to the pushed ArchString, so
module-level inline assembly and functions with custom `target-features` failed
to emit updated `$x<arch>` mapping symbols.
Call RTS.setArchString() with the parsed ISA string in emitTargetFeaturePush()
so `-filetype=obj` records the active `$x<arch>` mapping symbol alongside
`.option arch`.
This commit was created with the help of AI tools
Pull-Request: https://github.com/llvm/llvm-project/pull/225133
---
llvm/lib/Target/RISCV/RISCVAsmPrinter.cpp | 2 ++
llvm/test/CodeGen/RISCV/module-asm-features.ll | 7 ++-----
llvm/test/CodeGen/RISCV/riscv-func-target-feature.ll | 7 ++++---
3 files changed, 8 insertions(+), 8 deletions(-)
diff --git a/llvm/lib/Target/RISCV/RISCVAsmPrinter.cpp b/llvm/lib/Target/RISCV/RISCVAsmPrinter.cpp
index c2cadf361a0cd..eccf36e9af184 100644
--- a/llvm/lib/Target/RISCV/RISCVAsmPrinter.cpp
+++ b/llvm/lib/Target/RISCV/RISCVAsmPrinter.cpp
@@ -562,6 +562,8 @@ bool RISCVAsmPrinter::emitTargetFeaturePush(const MCSubtargetInfo &STI) {
if (!NeedEmitStdOptionArgs.empty()) {
RTS.emitDirectiveOptionPush();
RTS.emitDirectiveOptionArch(NeedEmitStdOptionArgs);
+ RTS.setArchString(
+ cantFail(RISCVFeatures::parseFeatureBits(STI))->toString());
return true;
}
diff --git a/llvm/test/CodeGen/RISCV/module-asm-features.ll b/llvm/test/CodeGen/RISCV/module-asm-features.ll
index 7fee7188e0ed8..87c970956a520 100644
--- a/llvm/test/CodeGen/RISCV/module-asm-features.ll
+++ b/llvm/test/CodeGen/RISCV/module-asm-features.ll
@@ -14,14 +14,11 @@
; CHECK-NEXT: ret
; EXTRA-FEATURES-NEXT: .option pop
-;; TODO: emitTargetFeaturePush does not call setArchString(), so the mapping
-;; symbol does not record +d/+f/+zicsr when assembling directly to an object
-;; file, causing llvm-objdump to fail to disassemble `fld`.
; OBJ-LABEL: Disassembly of section .text:
; OBJ-EMPTY:
-; OBJ-NEXT: 0000000000000000 <$xrv64i2p1>:
+; OBJ-NEXT: 0000000000000000 <$xrv64i2p1_f2p2_d2p2_zicsr2p0>:
; OBJ-NEXT: 0000000000000000 <func>:
-; OBJ-NEXT: 0: <unknown>
+; OBJ-NEXT: 0: fld ft0, 0x0(sp)
; OBJ-NEXT: 4: ret
; OBJ-NOT: {{.}}
diff --git a/llvm/test/CodeGen/RISCV/riscv-func-target-feature.ll b/llvm/test/CodeGen/RISCV/riscv-func-target-feature.ll
index de3de8c27df1d..61b06c605ba2a 100644
--- a/llvm/test/CodeGen/RISCV/riscv-func-target-feature.ll
+++ b/llvm/test/CodeGen/RISCV/riscv-func-target-feature.ll
@@ -2,20 +2,21 @@
; RUN: llc -mtriple=riscv64 -mcpu=sifive-u74 -filetype=obj < %s \
; RUN: | llvm-objdump -d --show-all-symbols --no-show-raw-insn - | FileCheck %s --check-prefix=OBJ
-;; TODO: emitTargetFeaturePush does not call setArchString(), so per-function
-;; target-features are not reflected in the $x<arch> mapping symbols.
; OBJ-LABEL: Disassembly of section .text:
; OBJ-EMPTY:
-; OBJ-NEXT: 0000000000000000 <$xrv64i2p1_m2p0_a2p1_f2p2_d2p2_c2p0_zicsr2p0_zifencei2p0_zmmul1p0_zaamo1p0_zalrsc1p0_zca1p0_zcd1p0>:
+; OBJ-NEXT: 0000000000000000 <$xrv64i2p1_m2p0_a2p1_f2p2_d2p2_c2p0_v1p0_zicsr2p0_zifencei2p0_zmmul1p0_zaamo1p0_zalrsc1p0_zca1p0_zcd1p0_zve32f1p0_zve32x1p0_zve64d1p0_zve64f1p0_zve64x1p0_zvl128b1p0_zvl32b1p0_zvl64b1p0>:
; OBJ-NEXT: 0000000000000000 <test1>:
; OBJ-NEXT: 0: ret
; OBJ-EMPTY:
+; OBJ-NEXT: 0000000000000002 <$xrv64i2p1_m2p0_a2p1_f2p2_d2p2_c2p0_zicsr2p0_zifencei2p0_zihintntl1p0_zmmul1p0_zaamo1p0_zalrsc1p0_zca1p0_zcd1p0>:
; OBJ-NEXT: 0000000000000002 <test2>:
; OBJ-NEXT: 2: ret
; OBJ-EMPTY:
+; OBJ-NEXT: 0000000000000004 <$xrv64i2p1_a2p1_c2p0_zicsr2p0_zifencei2p0_zmmul1p0_zaamo1p0_zalrsc1p0_zca1p0>:
; OBJ-NEXT: 0000000000000004 <test3>:
; OBJ-NEXT: 4: ret
; OBJ-EMPTY:
+; OBJ-NEXT: 0000000000000006 <$xrv64i2p1_m2p0_a2p1_f2p2_d2p2_c2p0_zicsr2p0_zifencei2p0_zmmul1p0_zaamo1p0_zalrsc1p0_zca1p0_zcd1p0>:
; OBJ-NEXT: 0000000000000006 <test4>:
; OBJ-NEXT: 6: ret
; OBJ-EMPTY:
>From 025177c168d2ee92b329842c0a35a80ba8c12c99 Mon Sep 17 00:00:00 2001
From: Aiden Grossman <aidengrossman at google.com>
Date: Sat, 26 Sep 2026 20:21:46 -0700
Subject: [PATCH 35/53] [Github] Build BOLT docs on changes (#226747)
We were already set up to build the BOLT docs, but the workflow did not
actually trigger on changes to the BOLT docs specifically. This change
fixes that.
---
.github/workflows/docs.yml | 2 ++
1 file changed, 2 insertions(+)
diff --git a/.github/workflows/docs.yml b/.github/workflows/docs.yml
index 286dbfcc1cba9..dc9f69f8b3c13 100644
--- a/.github/workflows/docs.yml
+++ b/.github/workflows/docs.yml
@@ -13,6 +13,7 @@ on:
branches:
- 'main'
paths:
+ - 'bolt/docs/**'
- 'llvm/docs/**'
- 'clang/docs/**'
- 'clang/include/clang/Basic/AttrDocs.td'
@@ -31,6 +32,7 @@ on:
- '.github/workflows/docs.yml'
pull_request:
paths:
+ - 'bolts/docs/**'
- 'llvm/docs/**'
- 'clang/docs/**'
- 'clang/include/clang/Basic/AttrDocs.td'
>From 28687894e88460fb5e5dd352b413416583ae1c39 Mon Sep 17 00:00:00 2001
From: Letu Ren <fantasquex at gmail.com>
Date: Sun, 27 Sep 2026 11:22:48 +0800
Subject: [PATCH 36/53] [MLIR][LLVM] Produce canonical const GEP in
convertGEPOp (#226699)
The DataLayout overload of ConstantExpr::getGetElementPtr is the form
that replaces typed constant GEPs. Use it for inrange GEPs and fail
translation when the indices cannot be reduced to a byte offset.
This resolves
https://github.com/llvm/llvm-project/pull/220424#discussion_r4092337505
Assisted-by: grok-4.7
Signed-off-by: Letu Ren <fantasquex at gmail.com>
---
.../LLVMIR/Dialect/LLVMIR/LLVMToLLVMIRTranslation.cpp | 7 ++++++-
mlir/test/Target/LLVMIR/llvmir.mlir | 2 +-
2 files changed, 7 insertions(+), 2 deletions(-)
diff --git a/mlir/lib/Target/LLVMIR/Dialect/LLVMIR/LLVMToLLVMIRTranslation.cpp b/mlir/lib/Target/LLVMIR/Dialect/LLVMIR/LLVMToLLVMIRTranslation.cpp
index 98ddf0a7c51ea..3f83a5f7f22ce 100644
--- a/mlir/lib/Target/LLVMIR/Dialect/LLVMIR/LLVMToLLVMIRTranslation.cpp
+++ b/mlir/lib/Target/LLVMIR/Dialect/LLVMIR/LLVMToLLVMIRTranslation.cpp
@@ -267,10 +267,15 @@ static LogicalResult convertGEPOp(GEPOp op, llvm::IRBuilderBase &builder,
constIndices.reserve(indices.size());
for (llvm::Value *value : indices)
constIndices.push_back(cast<llvm::Constant>(value));
+ const llvm::DataLayout &dataLayout =
+ moduleTranslation.getLLVMModule()->getDataLayout();
res = llvm::ConstantExpr::getGetElementPtr(
- elementType, baseConst, constIndices, nwFlags,
+ dataLayout, elementType, baseConst, constIndices, nwFlags,
llvm::ConstantRange::getNonEmpty(inrangeAttr.getLower(),
inrangeAttr.getUpper()));
+ if (!res)
+ return op.emitError(
+ "failed to lower 'inrange' GEP to a constant byte offset");
} else {
res = builder.CreateGEP(elementType, base, indices, "", nwFlags);
}
diff --git a/mlir/test/Target/LLVMIR/llvmir.mlir b/mlir/test/Target/LLVMIR/llvmir.mlir
index fb0cf9b493af4..5557788b039f7 100644
--- a/mlir/test/Target/LLVMIR/llvmir.mlir
+++ b/mlir/test/Target/LLVMIR/llvmir.mlir
@@ -124,7 +124,7 @@ llvm.mlir.global internal constant @int_gep() : !llvm.ptr {
// CHECK: @vt = external constant { [3 x ptr] }
llvm.mlir.global external constant @vt() : !llvm.struct<(array<3 x ptr>)>
-// CHECK: @int_gep_inrange = internal constant ptr getelementptr inbounds inrange(-16, 8) ({ [3 x ptr] }, ptr @vt, i32 0, i32 0, i32 2)
+// CHECK: @int_gep_inrange = internal constant ptr getelementptr inbounds inrange(-16, 8) (i8, ptr @vt, i64 16)
llvm.mlir.global internal constant @int_gep_inrange() : !llvm.ptr {
%addr = llvm.mlir.addressof @vt : !llvm.ptr
%gepinit = llvm.getelementptr inbounds inrange <i64, -16, 8> %addr[0, 0, 2] : (!llvm.ptr) -> !llvm.ptr, !llvm.struct<(array<3 x ptr>)>
>From 80abf9a1a8dc87542b3a4b853e2927e8c7951a59 Mon Sep 17 00:00:00 2001
From: Alan Zhao <ayzhao at google.com>
Date: Sat, 26 Sep 2026 20:46:30 -0700
Subject: [PATCH 37/53] [cmake][Windows] Fix CMake variable (#226745)
As a follow-up fix to #226635, the variable is `WIN32` not `Win32`.
---
llvm/cmake/modules/GetLibraryName.cmake | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/llvm/cmake/modules/GetLibraryName.cmake b/llvm/cmake/modules/GetLibraryName.cmake
index 5aa42c343a3d3..06a4b8d49fe5e 100644
--- a/llvm/cmake/modules/GetLibraryName.cmake
+++ b/llvm/cmake/modules/GetLibraryName.cmake
@@ -7,7 +7,7 @@ function(get_library_name path name)
list(FILTER suffixes EXCLUDE REGEX "^\\s*$")
# Do not strip the "lib" prefix for Windows because MSVC-style linkers don't
# implicitly add the "lib" prefix.
- if(prefixes AND NOT Win32)
+ if(prefixes AND NOT WIN32)
string(REPLACE ";" "|" prefixes "${prefixes}")
string(REGEX REPLACE "^(${prefixes})" "" path ${path})
endif()
>From 5ae850a88b31c7aeeae5f88a4d340e0b5a73e738 Mon Sep 17 00:00:00 2001
From: Craig Topper <craig.topper at sifive.com>
Date: Sat, 26 Sep 2026 21:43:38 -0700
Subject: [PATCH 38/53] [SDPatternMatch] Simplify EffectiveOperands and drop
the template specialization. NFC (#226036)
The chain and glue operands are in fixed locations, we don't need a loop
to find them.
Use the template parameter to skip the constructor body instead of using
template specialization.
---
llvm/include/llvm/CodeGen/SDPatternMatch.h | 27 ++++++++--------------
1 file changed, 9 insertions(+), 18 deletions(-)
diff --git a/llvm/include/llvm/CodeGen/SDPatternMatch.h b/llvm/include/llvm/CodeGen/SDPatternMatch.h
index 2e82693708fca..80804a22ae4f0 100644
--- a/llvm/include/llvm/CodeGen/SDPatternMatch.h
+++ b/llvm/include/llvm/CodeGen/SDPatternMatch.h
@@ -404,29 +404,20 @@ template <bool ExcludeChain> struct EffectiveOperands {
unsigned Size = 0;
unsigned FirstIndex = 0;
- explicit EffectiveOperands(SDValue N) {
- const unsigned TotalNumOps = N->getNumOperands();
- FirstIndex = TotalNumOps;
- for (unsigned I = 0; I < TotalNumOps; ++I) {
- // Count the number of non-chain and non-glue nodes (we ignore chain
- // and glue by default) and retreive the operand index offset.
- EVT VT = N->getOperand(I).getValueType();
- if (VT != MVT::Glue && VT != MVT::Other) {
- ++Size;
- if (FirstIndex == TotalNumOps)
- FirstIndex = I;
+ explicit EffectiveOperands(SDValue N) : Size(N->getNumOperands()) {
+ if (ExcludeChain) {
+ // Glue if present, is the last operand.
+ if (Size != 0 && N->getOperand(Size - 1).getValueType() == MVT::Glue)
+ --Size;
+ // Chain if present, is the first operand.
+ if (Size != 0 && N->getOperand(0).getValueType() == MVT::Other) {
+ ++FirstIndex;
+ --Size;
}
}
}
};
-template <> struct EffectiveOperands<false> {
- unsigned Size = 0;
- unsigned FirstIndex = 0;
-
- explicit EffectiveOperands(SDValue N) : Size(N->getNumOperands()) {}
-};
-
// === Ternary operations ===
template <typename T0_P, typename T1_P, typename T2_P, bool Commutable = false,
bool ExcludeChain = false>
>From fe831e1db6cc80f3de92f718435036d972da4af1 Mon Sep 17 00:00:00 2001
From: Timm Baeder <tbaeder at redhat.com>
Date: Sun, 27 Sep 2026 07:18:51 +0200
Subject: [PATCH 39/53] [clang][AST] Mark APValue as LLVM_ATTRIBUTE_WARN_UNUSED
(#226632)
---
clang/include/clang/AST/APValue.h | 3 ++-
clang/lib/AST/ExprConstant.cpp | 3 +--
2 files changed, 3 insertions(+), 3 deletions(-)
diff --git a/clang/include/clang/AST/APValue.h b/clang/include/clang/AST/APValue.h
index c5c871ef953cb..ef25f99b48544 100644
--- a/clang/include/clang/AST/APValue.h
+++ b/clang/include/clang/AST/APValue.h
@@ -21,6 +21,7 @@
#include "llvm/ADT/PointerIntPair.h"
#include "llvm/ADT/PointerUnion.h"
#include "llvm/Support/AlignOf.h"
+#include "llvm/Support/Compiler.h"
namespace clang {
namespace serialization {
@@ -119,7 +120,7 @@ namespace clang {
/// APValue - This class implements a discriminated union of [uninitialized]
/// [APSInt] [APFloat], [Complex APSInt] [Complex APFloat], [Expr + Offset],
/// [Vector: N * APValue], [Array: N * APValue]
-class APValue {
+class LLVM_ATTRIBUTE_WARN_UNUSED APValue {
typedef llvm::APFixedPoint APFixedPoint;
typedef llvm::APSInt APSInt;
typedef llvm::APFloat APFloat;
diff --git a/clang/lib/AST/ExprConstant.cpp b/clang/lib/AST/ExprConstant.cpp
index 2df754dc9007f..72baa9406cdde 100644
--- a/clang/lib/AST/ExprConstant.cpp
+++ b/clang/lib/AST/ExprConstant.cpp
@@ -20828,9 +20828,8 @@ bool FloatExprEvaluator::VisitCastExpr(const CastExpr *E) {
if (!hlslElementwiseCastHelper(Info, SubExpr, E->getType(), SrcVals,
SrcTypes))
return false;
- APValue Val;
- // cast our single element
+ // Cast our single element.
const FPOptions FPO = E->getFPFeaturesInEffect(Info.Ctx.getLangOpts());
APValue ResultVal;
if (!handleScalarCast(Info, FPO, E, SrcTypes[0], E->getType(), SrcVals[0],
>From b387e2eb63e63320471fc28fb5e281b84c98c729 Mon Sep 17 00:00:00 2001
From: Reid Kleckner <rkleckner at nvidia.com>
Date: Sat, 26 Sep 2026 22:34:40 -0700
Subject: [PATCH 40/53] [docs] Replace clang.llvm.org/docs links with Sphinx
links (#222507)
Use Sphinx document and option roles or project-relative links for links
within the Clang documentation. Repair stale generated-document
fragments found while validating the replacements. This ensures that
standalone documentation builds are self-contained, although
cross-project links (Clang->LLVM) typically go via absolute llvm.org
hrefs.
Part of #214861
Assisted-by: Codex
---
clang/docs/AllocToken.md | 4 ++--
clang/docs/ClangLinkerWrapper.md | 2 +-
clang/docs/ClangOffloadBundler.md | 2 +-
clang/docs/ClangTransformerTutorial.md | 2 +-
clang/docs/ControlFlowIntegrity.md | 2 +-
clang/docs/ControlFlowIntegrityDesign.md | 2 +-
clang/docs/InternalsManual.md | 2 +-
clang/docs/IntroductionToTheClangAST.md | 4 +---
clang/docs/LanguageExtensions.md | 10 +++++-----
clang/docs/LibASTImporter.md | 2 +-
clang/docs/LifetimeSafety.md | 12 ++++++------
clang/docs/SafeBuffers.md | 7 +++----
.../user-docs/SummaryExtraction.md | 3 +--
clang/docs/WarningSuppressionMappings.md | 3 +--
clang/docs/analyzer/checkers.md | 8 ++++----
clang/docs/analyzer/user-docs/Annotations.md | 6 +++---
clang/docs/conf.py | 5 ++++-
clang/include/clang/Basic/AttrDocs.td | 2 ++
llvm/docs/SphinxQuickstartTemplate.md | 3 +--
utils/docs/llvm_sphinx/ext/absolute_links.py | 1 +
.../llvm_sphinx/ext/absolute_links_test/markdown.md | 5 +++++
21 files changed, 46 insertions(+), 41 deletions(-)
diff --git a/clang/docs/AllocToken.md b/clang/docs/AllocToken.md
index b3a72ed480097..136f76bfb3111 100644
--- a/clang/docs/AllocToken.md
+++ b/clang/docs/AllocToken.md
@@ -166,8 +166,8 @@ the allocation call the wrapper returns, which is then instrumented normally.
Wrappers that are not inlined still require
`-fsanitize-alloc-token-extended`.
-[malloc-attribute]: https://clang.llvm.org/docs/AttributeReference.html#malloc
-[alloc-size-attribute]: https://clang.llvm.org/docs/AttributeReference.html#alloc-size
+[malloc-attribute]: AttributeReference.md#malloc
+[alloc-size-attribute]: AttributeReference.md#alloc-size
### Disabling Instrumentation
diff --git a/clang/docs/ClangLinkerWrapper.md b/clang/docs/ClangLinkerWrapper.md
index 081bfce2ea359..7245d3cead6ad 100644
--- a/clang/docs/ClangLinkerWrapper.md
+++ b/clang/docs/ClangLinkerWrapper.md
@@ -87,7 +87,7 @@ cause it be linked with any other device code with the same target triple.
The linker wrapper performs a lot of steps internally, such as input matching,
symbol resolution, and image registration. This makes it difficult to debug in
some scenarios. The behavior of the linker-wrapper is controlled mostly through
-metadata, described in [clang documentation](https://clang.llvm.org/docs/OffloadingDesign.html).
+metadata, described in [Clang documentation](OffloadingDesign.md).
The individual tool invocations the wrapper performs can be printed with the
`--wrapper-verbose` flag, and the intermediate files they operate on can be
diff --git a/clang/docs/ClangOffloadBundler.md b/clang/docs/ClangOffloadBundler.md
index 0dd93849eb275..45f9fba0a50aa 100644
--- a/clang/docs/ClangOffloadBundler.md
+++ b/clang/docs/ClangOffloadBundler.md
@@ -279,7 +279,7 @@ without differentiation based on offload kind.
**target-triple**
-: The target triple of the code object. See [Target Triple](https://clang.llvm.org/docs/CrossCompilation.html#target-triple).
+: The target triple of the code object. See [Target Triple](CrossCompilation.md#target-triple).
LLVM target triples can be with or without the optional environment field:
diff --git a/clang/docs/ClangTransformerTutorial.md b/clang/docs/ClangTransformerTutorial.md
index eba7c5e6b6085..28a7a428c7bec 100644
--- a/clang/docs/ClangTransformerTutorial.md
+++ b/clang/docs/ClangTransformerTutorial.md
@@ -367,7 +367,7 @@ introductions on clang's site:
- {doc}`Introduction to the Clang AST <IntroductionToTheClangAST>`
- {doc}`Matching the Clang AST <LibASTMatchers>`
-- [AST Matcher Reference](https://clang.llvm.org/docs/LibASTMatchersReference.html)
+- [AST Matcher Reference](LibASTMatchersReference.html){.external}
:::{rubric} Footnotes
:::
diff --git a/clang/docs/ControlFlowIntegrity.md b/clang/docs/ControlFlowIntegrity.md
index 045b70a4e23b3..c25a27fd910c5 100644
--- a/clang/docs/ControlFlowIntegrity.md
+++ b/clang/docs/ControlFlowIntegrity.md
@@ -39,7 +39,7 @@ CFI checks for classes without visibility attributes. Most users will want
to specify `-fvisibility=hidden`, which enables CFI checks for such classes.
When using `-fsanitize=cfi*` with `-flto=thin`, it is recommended
-to reduce link times by passing [-funique-source-file-names](https://clang.llvm.org/docs/UsersManual.html#cmdoption-f-no-unique-source-file-names), provided
+to reduce link times by passing {option}`-funique-source-file-names <-f[no-]unique-source-file-names>`, provided
that your program is compatible with it.
Experimental support for {ref}`cross-DSO control flow integrity
diff --git a/clang/docs/ControlFlowIntegrityDesign.md b/clang/docs/ControlFlowIntegrityDesign.md
index 0885b7a7b5891..1b3348928c859 100644
--- a/clang/docs/ControlFlowIntegrityDesign.md
+++ b/clang/docs/ControlFlowIntegrityDesign.md
@@ -780,5 +780,5 @@ ability to protect against invalid casts between polymorphic types.
[globalsplit]: https://github.com/llvm/llvm-project/blob/main/llvm/lib/Transforms/IPO/GlobalSplit.cpp
[intel cet]: https://software.intel.com/en-us/blogs/2016/06/09/intel-release-new-technology-specifications-protect-rop-attacks
[rfg]: https://xlab.tencent.com/en/2016/11/02/return-flow-guard
-[safestack]: https://clang.llvm.org/docs/SafeStack.html
+[safestack]: SafeStack.md
[type metadata]: https://llvm.org/docs/TypeMetadata.html
diff --git a/clang/docs/InternalsManual.md b/clang/docs/InternalsManual.md
index c2925563aa388..8d5549cb168be 100644
--- a/clang/docs/InternalsManual.md
+++ b/clang/docs/InternalsManual.md
@@ -2930,7 +2930,7 @@ allowing the programmer to pass semantic information along to the compiler for
various uses. For example, attributes may be used to alter the code generation
for a program construct, or to provide extra semantic information for static
analysis. This document explains how to add a custom attribute to Clang.
-Documentation on existing attributes can be found [here](https://clang.llvm.org/docs/AttributeReference.html).
+Documentation on existing attributes can be found [here](AttributeReference.md).
#### Attribute Basics
diff --git a/clang/docs/IntroductionToTheClangAST.md b/clang/docs/IntroductionToTheClangAST.md
index 56431b9ee5fec..b429e0b770f8e 100644
--- a/clang/docs/IntroductionToTheClangAST.md
+++ b/clang/docs/IntroductionToTheClangAST.md
@@ -108,8 +108,7 @@ and then recursively traverses everything that can be reached from that
node - this information has to be encoded for each specific node type.
This algorithm is encoded in the
[RecursiveASTVisitor](https://clang.llvm.org/doxygen/classclang_1_1RecursiveASTVisitor.html).
-See the [RecursiveASTVisitor
-tutorial](https://clang.llvm.org/docs/RAVFrontendAction.html).
+See the [RecursiveASTVisitor tutorial](RAVFrontendAction.md).
The two most basic nodes in the Clang AST are statements
([Stmt](https://clang.llvm.org/doxygen/classclang_1_1Stmt.html)) and
@@ -118,4 +117,3 @@ declarations
that expressions
([Expr](https://clang.llvm.org/doxygen/classclang_1_1Expr.html)) are
also statements in Clang's AST.
-
diff --git a/clang/docs/LanguageExtensions.md b/clang/docs/LanguageExtensions.md
index bd47b18da2481..74d6042961f72 100644
--- a/clang/docs/LanguageExtensions.md
+++ b/clang/docs/LanguageExtensions.md
@@ -1735,7 +1735,7 @@ mode.
Use `__has_feature(modules)` to determine if Modules have been enabled.
For example, compiling code with `-fmodules` enables the use of Modules.
-More information can be found [here](https://clang.llvm.org/docs/Modules.html).
+More information can be found [here](Modules.md).
## Language Extensions Back-ported to Previous Standards
@@ -2582,7 +2582,7 @@ and `-respondsToSelector:` or `+instancesRespondToSelector:` for
Objective-C methods. If such a check was missed, the program would compile
fine, run fine on newer systems, but crash on older systems.
-As of LLVM 5.0, `-Wunguarded-availability` uses the [availability attributes](https://clang.llvm.org/docs/AttributeReference.html#availability) together
+As of LLVM 5.0, `-Wunguarded-availability` uses the [availability attributes](AttributeReference.md#availability) together
with the new `@available()` keyword to assist with this issue.
When a method that's introduced in the OS newer than the target OS is called, a
-Wunguarded-availability warning is emitted if that call is not guarded:
@@ -2624,7 +2624,7 @@ void my_fun(NSSomeClass* var) {
```
If the caller of `my_fun()` already checks that `my_fun()` is only called
-on 10.12, then add an [availability attribute](https://clang.llvm.org/docs/AttributeReference.html#availability) to it,
+on 10.12, then add an [availability attribute](AttributeReference.md#availability) to it,
which will also suppress the warning and require that calls to my_fun() are
checked:
@@ -4715,7 +4715,7 @@ The effect of passing some other value to `__builtin_flt_rounds` is
implementation-defined. `__builtin_set_flt_rounds` is currently only supported
to work on x86, x86_64, powerpc, powerpc64, Arm and AArch64 targets. These builtins
read and modify the floating-point environment, which is not always allowed and may
-have unexpected behavior. Please see the section on [Accessing the floating point environment](https://clang.llvm.org/docs/UsersManual.html#accessing-the-floating-point-environment) for more information.
+have unexpected behavior. Please see the section on [Accessing the floating point environment](UsersManual.md#accessing-the-floating-point-environment) for more information.
### String builtins
@@ -6623,7 +6623,7 @@ more information about subobjects to be determined, so the `type & 1 == 1`
case will often give imprecise results when used across a function call boundary
even when optimization is enabled.
-[The pass_object_size and pass_dynamic_object_size attributes](https://clang.llvm.org/docs/AttributeReference.html#pass-object-size-pass-dynamic-object-size)
+[The pass_object_size and pass_dynamic_object_size attributes](AttributeReference.md#pass-object-size-pass-dynamic-object-size)
can be used to invisibly pass the object size for a pointer parameter alongside
the pointer in a function call. This allows more precise object sizes to be
determined both when building without optimizations and in the `type & 1 == 1`
diff --git a/clang/docs/LibASTImporter.md b/clang/docs/LibASTImporter.md
index 2ccb05ba33b1c..9c5a22d099a0d 100644
--- a/clang/docs/LibASTImporter.md
+++ b/clang/docs/LibASTImporter.md
@@ -6,7 +6,7 @@ It imports nodes of an `ASTContext` into another `ASTContext`.
In this document, we assume basic knowledge about the Clang AST. See the {doc}`Introduction
to the Clang AST <IntroductionToTheClangAST>` if you want to learn more
about how the AST is structured.
-Knowledge about {doc}`matching the Clang AST <LibASTMatchers>` and the [reference for the matchers](https://clang.llvm.org/docs/LibASTMatchersReference.html) are also useful.
+Knowledge about [matching the Clang AST](LibASTMatchers.md) and the [reference for the matchers](LibASTMatchersReference.html){.external} are also useful.
## Introduction
diff --git a/clang/docs/LifetimeSafety.md b/clang/docs/LifetimeSafety.md
index dbf67cfb2ce77..2785d2c80ccb1 100644
--- a/clang/docs/LifetimeSafety.md
+++ b/clang/docs/LifetimeSafety.md
@@ -22,8 +22,8 @@ This is compile-time analysis; there is no run-time overhead.
It tracks pointer validity through intra-procedural data-flow analysis. While it does
not require lifetime annotations to get started, in their absence, the analysis
treats function calls optimistically, assuming no lifetime effects, thereby potentially missing dangling pointer issues. As more functions are annotated
-with attributes like [clang::lifetimebound](https://clang.llvm.org/docs/AttributeReference.html#lifetimebound), [gsl::Owner](https://clang.llvm.org/docs/AttributeReference.html#gsl-owner), and
-[gsl::Pointer](https://clang.llvm.org/docs/AttributeReference.html#gsl-pointer), the analysis can see through these lifetime contracts and enforce
+with attributes like [clang::lifetimebound](AttributeReference.md#lifetimebound), [gsl::Owner](AttributeReference.md#owner), and
+[gsl::Pointer](AttributeReference.md#pointer), the analysis can see through these lifetime contracts and enforce
lifetime safety at call sites with higher accuracy. This approach supports
gradual adoption in existing codebases.
@@ -142,8 +142,8 @@ void test() {
Without these annotations, the analysis may not be able to determine whether a
type is owning or borrowing, which can affect analysis precision. For more
details on these attributes, see the Clang attribute reference for
-[gsl::Owner](https://clang.llvm.org/docs/AttributeReference.html#gsl-owner) and
-[gsl::Pointer](https://clang.llvm.org/docs/AttributeReference.html#gsl-pointer).
+[gsl::Owner](AttributeReference.md#owner) and
+[gsl::Pointer](AttributeReference.md#pointer).
:::{note}
Types with mixed ownership semantics (owning some data while holding views to
@@ -211,7 +211,7 @@ void test() {
}
```
-For more details, see [lifetimebound](https://clang.llvm.org/docs/AttributeReference.html#lifetimebound).
+For more details, see [lifetimebound](AttributeReference.md#lifetimebound).
### NoEscape
@@ -223,7 +223,7 @@ parameter to escape its scope, for example, by returning it or assigning it to
a field or global variable. This is useful for parameters passed to callbacks
or visitors that are only used during the call and not stored.
-For more details, see [noescape](https://clang.llvm.org/docs/AttributeReference.html#noescape).
+For more details, see [noescape](AttributeReference.md#noescape).
## Checks Performed
diff --git a/clang/docs/SafeBuffers.md b/clang/docs/SafeBuffers.md
index cc7d74ed37efb..4cd98fcbda133 100644
--- a/clang/docs/SafeBuffers.md
+++ b/clang/docs/SafeBuffers.md
@@ -61,7 +61,7 @@ acting as "hardened custom containers" to replace raw pointers.
However, such approach would be very unergonomic in C, and safety guarantees
will be lower due to lack of good encapsulation technology. A better approach
to bounds safety for non-C++ programs,
-[-fbounds-safety](https://clang.llvm.org/docs/BoundsSafety.html),
+[-fbounds-safety](BoundsSafety.md),
is currently in development.
Technically, safety guarantees cannot be provided without hardening
@@ -321,7 +321,7 @@ int get_last_element(int *pointer, size_t size) {
}
```
-This behavior is analogous to `#pragma clang diagnostic` ([documentation](https://clang.llvm.org/docs/UsersManual.html#controlling-diagnostics-via-pragmas))
+This behavior is analogous to `#pragma clang diagnostic` ([documentation](UsersManual.md#controlling-diagnostics-via-pragmas))
However, `#pragma clang unsafe_buffer_usage` is specialized and recommended
over `#pragma clang diagnostic` for a number of technical and non-technical
reasons. Most importantly, `#pragma clang unsafe_buffer_usage` is more
@@ -391,7 +391,7 @@ passed into the wrapper is correct.**
### Flag bounds information discontinuities with `[[clang::unsafe_buffer_usage]]`
The clang attribute `[[clang::unsafe_buffer_usage]]`
-([attribute documentation](https://clang.llvm.org/docs/AttributeReference.html#unsafe-buffer-usage))
+([attribute documentation](AttributeReference.md#unsafe-buffer-usage))
allows the user to annotate various objects, such as functions or member
variables, as incompatible with the Safe Buffers programming model.
You are encouraged to do that for arbitrary reasons, but typically the main
@@ -587,4 +587,3 @@ significantly fewer warnings. It will also need to bypass
`#pragma clang unsafe_buffer_usage` suppressions and "see through"
unsafe wrappers such as `unsafe_forge_span` -- something that
the static analyzer is naturally capable of doing.
-
diff --git a/clang/docs/ScalableStaticAnalysis/user-docs/SummaryExtraction.md b/clang/docs/ScalableStaticAnalysis/user-docs/SummaryExtraction.md
index 50d5ccf510822..f493994d1554d 100644
--- a/clang/docs/ScalableStaticAnalysis/user-docs/SummaryExtraction.md
+++ b/clang/docs/ScalableStaticAnalysis/user-docs/SummaryExtraction.md
@@ -32,5 +32,4 @@ or just happens to have an error, then the error is forwarded as a `scalable-sta
These errors can be downgraded into warnings using `-Wno-error=scalable-static-analysis-framework`.
These errors can be completely suppressed using `-Wno-scalable-static-analysis-framework`.
-See the [diagnostic flags](https://clang.llvm.org/docs/DiagnosticsReference.html#wscalable-static-analysis-framework) for the full list of diagnostics controlled by `-Wscalable-static-analysis-framework`.
-
+See the [diagnostic flags](../../DiagnosticsReference.md#wscalable-static-analysis-framework) for the full list of diagnostics controlled by `-Wscalable-static-analysis-framework`.
diff --git a/clang/docs/WarningSuppressionMappings.md b/clang/docs/WarningSuppressionMappings.md
index 2c6ce42f7c668..350c4ca78dffc 100644
--- a/clang/docs/WarningSuppressionMappings.md
+++ b/clang/docs/WarningSuppressionMappings.md
@@ -26,7 +26,7 @@ flag.
Note that this mechanism won't enable any diagnostics on its own. Users should
still turn on warnings in their compilations with explicit `-Wfoo` flags.
-[Controlling diagnostics pragmas](https://clang.llvm.org/docs/UsersManual.html#controlling-diagnostics-via-pragmas)
+[Controlling diagnostics pragmas](UsersManual.md#controlling-diagnostics-via-pragmas)
take precedence over suppression mappings. Ensuring code author's explicit
intent is always preserved.
@@ -86,4 +86,3 @@ src:*foo/*=emit
# Only suppress for sources under bar/.
src:*bar/*
```
-
diff --git a/clang/docs/analyzer/checkers.md b/clang/docs/analyzer/checkers.md
index b3a64ccd564d4..e2bbf52943bb8 100644
--- a/clang/docs/analyzer/checkers.md
+++ b/clang/docs/analyzer/checkers.md
@@ -197,7 +197,7 @@ void test() {
Null pointer dereferences of pointers with address spaces are not always defined
as error. Specifically on x86/x86-64 target if the pointer address space is
256 (x86 GS Segment), 257 (x86 FS Segment), or 258 (x86 SS Segment), a null
-dereference is not defined as error. See [X86/X86-64 Language Extensions](https://clang.llvm.org/docs/LanguageExtensions.html#memory-references-to-specified-segments)
+dereference is not defined as error. See [X86/X86-64 Language Extensions](../LanguageExtensions.md#memory-references-to-specified-segments)
for reference.
If the analyzer option `suppress-dereferences-from-any-address-space` is set
@@ -808,7 +808,7 @@ This checker does not accept the coding pattern where an enum type is used to
store combinations of flag values.
Such enums should be annotated with the `__attribute__((flag_enum))` or by the
`[[clang::flag_enum]]` attribute to signal this intent. Refer to the
-[documentation](https://clang.llvm.org/docs/AttributeReference.html#flag-enum)
+[documentation](../AttributeReference.md#flag-enum)
of this Clang attribute.
```cpp
@@ -901,7 +901,7 @@ arguments -- even if there is no such call in the codebase.
This design rule is dictated by the SEI CERT rule [EXP47-C](https://wiki.sei.cmu.edu/confluence/display/c/EXP47-C.+Do+not+call+va_arg+with+an+argument+of+the+incorrect+type),
which describes several issues related to the use of `va_arg()`. (The problem
reported by this checker is shown in the second code example; the first,
-unrelated code example is covered by the clang diagnostic [-Wvarargs](https://clang.llvm.org/docs/DiagnosticsReference.html#wvarargs).)
+unrelated code example is covered by the clang diagnostic [-Wvarargs](../DiagnosticsReference.md#wvarargs).)
```cpp
// This function expects a list of variadic arguments terminated by a NULL pointer.
@@ -3357,7 +3357,7 @@ int *direct_return() {
The attribute states that the returned value is dangling after the lifetime
of the annotated parameter, or of the implicit object argument, has ended.
-Refer to the [documentation](https://clang.llvm.org/docs/AttributeReference.html#lifetimebound)
+Refer to the [documentation](../AttributeReference.md#lifetimebound)
of this Clang attribute.
```cpp
diff --git a/clang/docs/analyzer/user-docs/Annotations.md b/clang/docs/analyzer/user-docs/Annotations.md
index ce2f42fa1920f..d1551f5750f22 100644
--- a/clang/docs/analyzer/user-docs/Annotations.md
+++ b/clang/docs/analyzer/user-docs/Annotations.md
@@ -8,7 +8,7 @@ analyzer's ability to find bugs.
This page gives a practical overview of such annotations. For more technical
specifics regarding Clang-specific annotations please see the Clang's list of
-[language extensions](https://clang.llvm.org/docs/LanguageExtensions.html).
+[language extensions](../../LanguageExtensions.md).
Details of "standard" GCC attributes (that Clang also supports) can
be found in the [GCC manual](https://gcc.gnu.org/onlinedocs/gcc/), with the
majority of the relevant attributes being in the section on
@@ -212,7 +212,7 @@ conventions can cause the analyzer to miss bugs or flag false positives.
One can educate the analyzer (and others who read your code) about methods or
functions that deviate from the Cocoa and Core Foundation conventions using the
attributes described here. However, you should consider using proper naming
-conventions or the [objc_method_family](https://clang.llvm.org/docs/LanguageExtensions.html#the-objc-method-family-attribute)
+conventions or the [objc_method_family](../../AttributeReference.md#objc-method-family)
attribute, if applicable.
(ns_returns_retained)=
@@ -598,7 +598,7 @@ By default, the following summaries are assumed:
including the implicit `this` parameter.
These summaries can be overriden with the following
-[attributes](https://clang.llvm.org/docs/AttributeReference.html#os-returns-not-retained):
+{ref}`attributes <os-retained-attr-family>`:
#### Attribute 'os_returns_retained'
diff --git a/clang/docs/conf.py b/clang/docs/conf.py
index 1280d70b3be71..9b4991f3752bf 100644
--- a/clang/docs/conf.py
+++ b/clang/docs/conf.py
@@ -19,7 +19,7 @@
globals().update(common_conf(tags))
-myst_enable_extensions += ["deflist"]
+myst_enable_extensions += ["attrs_inline", "deflist"]
# -- General configuration -----------------------------------------------------
@@ -29,9 +29,12 @@
"sphinx.ext.todo",
"sphinx.ext.mathjax",
"sphinx.ext.graphviz",
+ "llvm_sphinx.ext.absolute_links",
"llvm_sphinx.ext.ghlinks",
]
+llvm_sphinx_doc_url_prefixes = ("https://clang.llvm.org/docs/",)
+
import sphinx
# General information about the project.
diff --git a/clang/include/clang/Basic/AttrDocs.td b/clang/include/clang/Basic/AttrDocs.td
index bac01d64ac811..0a8977d3c6b14 100644
--- a/clang/include/clang/Basic/AttrDocs.td
+++ b/clang/include/clang/Basic/AttrDocs.td
@@ -1922,6 +1922,8 @@ have the same respective semantics when applied to CoreFoundation objects.
These attributes affect code generation when interacting with ARC code, and
they are used by the Clang Static Analyzer.
+(os-retained-attr-family)=
+
Finally, in C++ interacting with XNU kernel (objects inheriting from OSObject),
the same attribute family is present:
`__attribute__((os_returns_not_retained))`,
diff --git a/llvm/docs/SphinxQuickstartTemplate.md b/llvm/docs/SphinxQuickstartTemplate.md
index 68fd7c60278cf..a7a69582579e4 100644
--- a/llvm/docs/SphinxQuickstartTemplate.md
+++ b/llvm/docs/SphinxQuickstartTemplate.md
@@ -109,8 +109,7 @@ GitHub. Likewise, do not link to generated paths such as `CMake.html`.
Use a Sphinx `{ref}` role when the target is an explicit label rather than a
generated heading. Explicit labels are useful when
-an anchor must remain stable after its heading or source file is renamed, or
-when a target must be exported to another Sphinx project through an inventory.
+an anchor must remain stable after its heading or source file is renamed.
Avoid adding explicit labels to ordinary headings when a checked Markdown link
is sufficient.
diff --git a/utils/docs/llvm_sphinx/ext/absolute_links.py b/utils/docs/llvm_sphinx/ext/absolute_links.py
index 43d9439e978bc..b7f42a1662aaf 100644
--- a/utils/docs/llvm_sphinx/ext/absolute_links.py
+++ b/utils/docs/llvm_sphinx/ext/absolute_links.py
@@ -259,6 +259,7 @@ def run_tests() -> None:
"target.html#target-document",
"project:rest.rst#rest-section",
"rest.html#rest-section",
+ "project:rest.rst",
)
for link in expected_nonportable_links:
if link not in warnings:
diff --git a/utils/docs/llvm_sphinx/ext/absolute_links_test/markdown.md b/utils/docs/llvm_sphinx/ext/absolute_links_test/markdown.md
index 61d8500fce212..f52817077d64c 100644
--- a/utils/docs/llvm_sphinx/ext/absolute_links_test/markdown.md
+++ b/utils/docs/llvm_sphinx/ext/absolute_links_test/markdown.md
@@ -17,12 +17,17 @@ These nonportable internal links should warn:
[project-target]: project:target.md#target-document
[html-target]: target.html#target-document
+:::{note}
+[A project link in Markdown directive content](project:rest.rst)
+:::
+
These links should not warn:
- [another project](https://other.example.test/docs/target.html)
- [a nonexistent document](https://example.test/docs/missing.html)
- [a non-document page](https://example.test/docs/downloads/package.tar.xz)
- [source document](target.md)
+- [reStructuredText source document](rest.rst)
- [source heading](target.md#target-section)
- [same-document heading](#markdown-absolute-link-tests)
- [an HTML file that is not a document](static.html)
>From 118097a73d9f40844f3af44769b5cc86f95ba8f1 Mon Sep 17 00:00:00 2001
From: Amilendra Kodithuwakku <amilendra.kodithuwakku at arm.com>
Date: Sun, 27 Sep 2026 11:06:45 +0530
Subject: [PATCH 41/53] [docs][bolt] Remove Markdown enum (#226411)
Remove use of Markdown enum from the bolt docs. Related to #223829.
This should fix our ATfL build
[failure](https://github.com/arm/arm-toolchain/actions/runs/36103090912/job/107969633028#step:7:22533)
```
Traceback (most recent call last):
File "/workspace/python/.venv/lib/python3.12/site-packages/sphinx/config.py", line 529, in eval_config_file
exec(code, namespace) # NoQA: S102
^^^^^^^^^^^^^^^^^^^^^
File "/workspace/src/bolt/docs/conf.py", line 18, in <module>
globals().update(common_conf(tags, markdown=Markdown.NEVER))
^^^^^^^^
NameError: name 'Markdown' is not defined
```
---
bolt/docs/conf.py | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/bolt/docs/conf.py b/bolt/docs/conf.py
index f641dc6f2b80e..6a3eba73ffd17 100644
--- a/bolt/docs/conf.py
+++ b/bolt/docs/conf.py
@@ -15,7 +15,7 @@
from llvm_sphinx import * # see llvm-project/utils/docs/README.md
-globals().update(common_conf(tags, markdown=Markdown.NEVER))
+globals().update(common_conf(tags))
# -- General configuration -----------------------------------------------------
>From 51b20b5a25dd18152c988b3019323cf17fbea648 Mon Sep 17 00:00:00 2001
From: Alexander Richardson <alexrichardson at google.com>
Date: Sat, 26 Sep 2026 22:47:33 -0700
Subject: [PATCH 42/53] [Clang][RISC-V] Fix lto-module-asm-abi.c on
non-asserts/no-lld builds (#226752)
Pass `-fno-discard-value-names` so the `entry:` label is preserved in
non-asserts builds, and drop `-fuse-ld=lld` so the driver check succeeds
on bots that do not have `ld.lld` installed.
Fixes: bf529329b5ce ("[RISC-V][LTO] Add baseline tests for LTO inline assembly and mapping symbols (#225129)")
Pull-Request: https://github.com/llvm/llvm-project/pull/226752
---
clang/test/CodeGen/RISCV/lto-module-asm-abi.c | 6 +++---
1 file changed, 3 insertions(+), 3 deletions(-)
diff --git a/clang/test/CodeGen/RISCV/lto-module-asm-abi.c b/clang/test/CodeGen/RISCV/lto-module-asm-abi.c
index cceafaac3ba0e..50242fdd3085d 100644
--- a/clang/test/CodeGen/RISCV/lto-module-asm-abi.c
+++ b/clang/test/CodeGen/RISCV/lto-module-asm-abi.c
@@ -4,7 +4,7 @@
/// Check that -march=rv64gcv -flto records +d in module asm and function
/// target-features even though the driver only passes mcpu=generic-rv64 to lld.
-// RUN: %clang --target=riscv64-linux-android -march=rv64gcv -flto %s -S -emit-llvm -o - \
+// RUN: %clang --target=riscv64-linux-android -march=rv64gcv -flto -fno-discard-value-names %s -S -emit-llvm -o - \
// RUN: | FileCheck %s --check-prefix=IR
// IR: module asm(target_features: "{{.*}}+d{{.*}}", target_cpu: "generic-rv64")
@@ -20,11 +20,11 @@
// IR-NEXT: ![[#]] = !{i32 6, !"riscv-isa", ![[#ISA:]]}
// IR-NEXT: ![[#ISA]] = !{!"{{.*}}_d2p2_{{.*}}"}
-// RUN: %clang --target=riscv64-linux-android -march=rv64gcv -flto -shared -nostdlib -fuse-ld=lld %s -### 2>&1 \
+// RUN: %clang --target=riscv64-linux-android -march=rv64gcv -flto -shared -nostdlib %s -### 2>&1 \
// RUN: | FileCheck %s --check-prefix=DRIVER
// DRIVER: "-cc1"{{.*}}"-target-cpu" "generic-rv64"{{.*}}"-target-feature" "+d"{{.*}}"-target-abi" "lp64d"
-// DRIVER: "{{[^"]*}}ld.lld{{(\.exe)?}}"
+// DRIVER: "-m" "elf64lriscv"
// DRIVER-NOT: mattr
// DRIVER-SAME: "-plugin-opt=mcpu=generic-rv64"
// DRIVER-NOT: mattr
>From 790ba52abb7f11c91d8f8b8323b5890c9e15242a Mon Sep 17 00:00:00 2001
From: Lang Hames <lhames at gmail.com>
Date: Sun, 27 Sep 2026 17:04:11 +1000
Subject: [PATCH 43/53] [orc-rt] Move ORC_RT_LOG_ENABLED out of config.h; catch
typos. (#226751)
Move ORC_RT_LOG_ENABLED to orc-rt-c/support/LogLevel.h so that config.h
is kept for build options only.
Switch to using function-like macros so that typos in level-names become
compilation errors.
---
orc-rt/include/CMakeLists.txt | 1 +
orc-rt/include/orc-rt-c/config.h.in | 32 -------------------
orc-rt/include/orc-rt-c/support/LogLevel.h | 37 ++++++++++++++++++++++
orc-rt/include/orc-rt-c/support/Logging.h | 3 +-
orc-rt/include/orc-rt/bedrock/Session.h | 1 +
5 files changed, 41 insertions(+), 33 deletions(-)
create mode 100644 orc-rt/include/orc-rt-c/support/LogLevel.h
diff --git a/orc-rt/include/CMakeLists.txt b/orc-rt/include/CMakeLists.txt
index 72069456c5e3b..016f9b9cfc932 100644
--- a/orc-rt/include/CMakeLists.txt
+++ b/orc-rt/include/CMakeLists.txt
@@ -8,6 +8,7 @@ set(ORC_RT_HEADERS
orc-rt-c/support/Compiler.h
orc-rt-c/support/CoreTypes.h
orc-rt-c/support/Error.h
+ orc-rt-c/support/LogLevel.h
orc-rt-c/support/Logging.h
orc-rt-c/support/WrapperFunction.h
orc-rt/bedrock/BootstrapInfo.h
diff --git a/orc-rt/include/orc-rt-c/config.h.in b/orc-rt/include/orc-rt-c/config.h.in
index 7af7b4b2bc565..c4ff9a0fb4cd2 100644
--- a/orc-rt/include/orc-rt-c/config.h.in
+++ b/orc-rt/include/orc-rt-c/config.h.in
@@ -38,36 +38,4 @@
#endif
#define ORC_RT_LOG_BACKEND @ORC_RT_LOG_BACKEND_VALUE@
-/*
- * ORC_RT_LOG_ENABLED(Level) is true (1) if log sites at the given level are
- * compiled in, and false (0) if they're compiled out, either because the
- * backend is none or because the level is below the ORC_RT_LOG_LEVEL floor. It
- * takes the same level token as ORC_RT_LOG (see orc-rt-c/support/Logging.h),
- * and can be used in preprocessor conditionals:
- *
- * #if ORC_RT_LOG_ENABLED(Error)
- * ...
- * #endif
- *
- * It's defined here, rather than in Logging.h, so that it can be used without
- * pulling in the logging backend's headers (e.g. <os/log.h>).
- *
- * Code that relies on logging to surface something important (e.g. an error
- * reporter that logs) can use this to pick an alternative when logging is
- * compiled out. Note that a compiled-in level may still be suppressed at
- * runtime (e.g. by the printf backend's runtime threshold).
- *
- * The ORC_RT_LOG_LEVEL_VALUE_<Level> aliases map ORC_RT_LOG's level tokens to
- * the ORC_RT_LOG_LEVEL_<LEVEL> values above. An unrecognized level token
- * pastes to an undefined identifier, which evaluates to 0 in #if.
- */
-#define ORC_RT_LOG_LEVEL_VALUE_Debug ORC_RT_LOG_LEVEL_DEBUG
-#define ORC_RT_LOG_LEVEL_VALUE_Info ORC_RT_LOG_LEVEL_INFO
-#define ORC_RT_LOG_LEVEL_VALUE_Warning ORC_RT_LOG_LEVEL_WARNING
-#define ORC_RT_LOG_LEVEL_VALUE_Error ORC_RT_LOG_LEVEL_ERROR
-
-#define ORC_RT_LOG_ENABLED(Level) \
- (ORC_RT_LOG_BACKEND != ORC_RT_LOG_BACKEND_NONE && \
- ORC_RT_LOG_LEVEL_VALUE_##Level >= ORC_RT_LOG_LEVEL)
-
#endif /* ORC_RT_C_CONFIG_H */
diff --git a/orc-rt/include/orc-rt-c/support/LogLevel.h b/orc-rt/include/orc-rt-c/support/LogLevel.h
new file mode 100644
index 0000000000000..e2df75557466f
--- /dev/null
+++ b/orc-rt/include/orc-rt-c/support/LogLevel.h
@@ -0,0 +1,37 @@
+/*===------ LogLevel.h - ORC Runtime compiled-in log levels -------*- C -*-===*\
+|* *|
+|* Part of the LLVM Project, under the Apache License v2.0 with LLVM *|
+|* Exceptions. *|
+|* See https://llvm.org/LICENSE.txt for license information. *|
+|* SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception *|
+|* *|
+|*===----------------------------------------------------------------------===*|
+|* *|
+|* ORC_RT_LOG_ENABLED, kept apart from Logging.h so that it can be used *|
+|* without pulling in the logging backend's headers (e.g. <os/log.h>). *|
+|* *|
+\*===----------------------------------------------------------------------===*/
+
+#ifndef ORC_RT_C_SUPPORT_LOGLEVEL_H
+#define ORC_RT_C_SUPPORT_LOGLEVEL_H
+
+#include "orc-rt-c/config.h"
+
+/**
+ * ORC_RT_LOG_ENABLED(Level) is 1 if log sites at Level (Error, Warning, Info or
+ * Debug, as for ORC_RT_LOG) are compiled in, and 0 otherwise. Usable in #if.
+ */
+#define ORC_RT_LOG_ENABLED(Level) \
+ (ORC_RT_LOG_BACKEND != ORC_RT_LOG_BACKEND_NONE && \
+ ORC_RT_LOG_ENABLED_##Level() >= ORC_RT_LOG_LEVEL)
+
+/*
+ * Per-level macros are function-like so that a mistyped level is a hard error
+ * rather than silently evaluating to 0.
+ */
+#define ORC_RT_LOG_ENABLED_Debug() ORC_RT_LOG_LEVEL_DEBUG
+#define ORC_RT_LOG_ENABLED_Info() ORC_RT_LOG_LEVEL_INFO
+#define ORC_RT_LOG_ENABLED_Warning() ORC_RT_LOG_LEVEL_WARNING
+#define ORC_RT_LOG_ENABLED_Error() ORC_RT_LOG_LEVEL_ERROR
+
+#endif /* ORC_RT_C_SUPPORT_LOGLEVEL_H */
diff --git a/orc-rt/include/orc-rt-c/support/Logging.h b/orc-rt/include/orc-rt-c/support/Logging.h
index 9b7413766b330..a10ad8449eb89 100644
--- a/orc-rt/include/orc-rt-c/support/Logging.h
+++ b/orc-rt/include/orc-rt-c/support/Logging.h
@@ -43,6 +43,7 @@
#include "orc-rt-c/config.h"
#include "orc-rt-c/support/Compiler.h"
+#include "orc-rt-c/support/LogLevel.h"
#if ORC_RT_LOG_BACKEND == ORC_RT_LOG_BACKEND_OS_LOG
#include <os/log.h>
@@ -136,7 +137,7 @@ int orc_rt_log_formatCheck(const char *Fmt, ...) ORC_RT_FORMAT_PRINTF(1, 2);
/*
* To check whether a level is compiled in, use ORC_RT_LOG_ENABLED(Level),
- * defined in orc-rt-c/config.h.
+ * defined in orc-rt-c/support/LogLevel.h.
*/
/**
diff --git a/orc-rt/include/orc-rt/bedrock/Session.h b/orc-rt/include/orc-rt/bedrock/Session.h
index 49542ca14df3c..fbcd2da13f90e 100644
--- a/orc-rt/include/orc-rt/bedrock/Session.h
+++ b/orc-rt/include/orc-rt/bedrock/Session.h
@@ -25,6 +25,7 @@
#include "orc-rt-c/config.h"
#include "orc-rt-c/support/CoreTypes.h"
+#include "orc-rt-c/support/LogLevel.h"
#include "orc-rt-c/support/WrapperFunction.h"
#include <cassert>
>From 45f10efca6826d00e23a8dfdff867efe87bd1970 Mon Sep 17 00:00:00 2001
From: Matt Arsenault <Matthew.Arsenault at amd.com>
Date: Sun, 27 Sep 2026 10:55:46 +0200
Subject: [PATCH 44/53] CodeGen: Remove shadowing TargetMachine members from
COFF and M68k TLOFs (#226769)
TargetLoweringObjectFile::Initialize already records the TargetMachine
in a base class field.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
---
llvm/include/llvm/CodeGen/TargetLoweringObjectFileImpl.h | 1 -
llvm/lib/CodeGen/TargetLoweringObjectFileImpl.cpp | 1 -
llvm/lib/Target/M68k/M68kTargetObjectFile.cpp | 3 ---
llvm/lib/Target/M68k/M68kTargetObjectFile.h | 2 --
4 files changed, 7 deletions(-)
diff --git a/llvm/include/llvm/CodeGen/TargetLoweringObjectFileImpl.h b/llvm/include/llvm/CodeGen/TargetLoweringObjectFileImpl.h
index e631d820eb637..e76a1ac172057 100644
--- a/llvm/include/llvm/CodeGen/TargetLoweringObjectFileImpl.h
+++ b/llvm/include/llvm/CodeGen/TargetLoweringObjectFileImpl.h
@@ -188,7 +188,6 @@ class LLVM_ABI TargetLoweringObjectFileMachO : public TargetLoweringObjectFile {
class LLVM_ABI TargetLoweringObjectFileCOFF : public TargetLoweringObjectFile {
mutable unsigned NextUniqueID = 0;
- const TargetMachine *TM = nullptr;
public:
~TargetLoweringObjectFileCOFF() override = default;
diff --git a/llvm/lib/CodeGen/TargetLoweringObjectFileImpl.cpp b/llvm/lib/CodeGen/TargetLoweringObjectFileImpl.cpp
index 8b81aaf9d12d4..68ff80ae1d376 100644
--- a/llvm/lib/CodeGen/TargetLoweringObjectFileImpl.cpp
+++ b/llvm/lib/CodeGen/TargetLoweringObjectFileImpl.cpp
@@ -2042,7 +2042,6 @@ void TargetLoweringObjectFileCOFF::emitLinkerDirectives(
void TargetLoweringObjectFileCOFF::Initialize(MCContext &Ctx,
const TargetMachine &TM) {
TargetLoweringObjectFile::Initialize(Ctx, TM);
- this->TM = &TM;
const Triple &T = TM.getTargetTriple();
if (T.isWindowsMSVCEnvironment() || T.isWindowsItaniumEnvironment()) {
StaticCtorSection =
diff --git a/llvm/lib/Target/M68k/M68kTargetObjectFile.cpp b/llvm/lib/Target/M68k/M68kTargetObjectFile.cpp
index 4986d5dbebb91..5aed6b8fe38e1 100644
--- a/llvm/lib/Target/M68k/M68kTargetObjectFile.cpp
+++ b/llvm/lib/Target/M68k/M68kTargetObjectFile.cpp
@@ -14,7 +14,6 @@
#include "M68kTargetObjectFile.h"
#include "M68kSubtarget.h"
-#include "M68kTargetMachine.h"
#include "llvm/BinaryFormat/ELF.h"
#include "llvm/IR/DataLayout.h"
@@ -37,8 +36,6 @@ void M68kELFTargetObjectFile::Initialize(MCContext &Ctx,
TargetLoweringObjectFileELF::Initialize(Ctx, TM);
InitializeELF(TM.Options.UseInitArray);
- this->TM = &static_cast<const M68kTargetMachine &>(TM);
-
// FIXME do we need `.sdata` and `.sbss` explicitly?
SmallDataSection = getContext().getELFSection(
".sdata", ELF::SHT_PROGBITS, ELF::SHF_WRITE | ELF::SHF_ALLOC);
diff --git a/llvm/lib/Target/M68k/M68kTargetObjectFile.h b/llvm/lib/Target/M68k/M68kTargetObjectFile.h
index 80a7d0d6e1206..3d145aa434750 100644
--- a/llvm/lib/Target/M68k/M68kTargetObjectFile.h
+++ b/llvm/lib/Target/M68k/M68kTargetObjectFile.h
@@ -17,9 +17,7 @@
#include "llvm/CodeGen/TargetLoweringObjectFileImpl.h"
namespace llvm {
-class M68kTargetMachine;
class M68kELFTargetObjectFile : public TargetLoweringObjectFileELF {
- const M68kTargetMachine *TM;
MCSection *SmallDataSection;
MCSection *SmallBSSSection;
>From ffdaf818710cdc927e225f0936ead38403d7ec67 Mon Sep 17 00:00:00 2001
From: Lang Hames <lhames at gmail.com>
Date: Sun, 27 Sep 2026 19:27:44 +1000
Subject: [PATCH 45/53] [orc-rt] Simplify ConnectorRegistry::connect and ogre
setup. (#226772)
ConnectorRegistry::connect and ConnectorFn now take the Session and
BootstrapInfo directly, replacing the GetAttachInfo callback and the
AttachInfo struct. The callback let the Session be built lazily, after
the connector had validated its spec, but a connection can fail after
validation anyway, so callers already had to handle discarding a Session
built for a failed connection. Build the Session first, then connect.
ogre's setup is restructured to match: makeSession detects the process
info, builds a thread-pool dispatcher and the Session, adds the host
services, then connects; runOgre waits for the detach. Errors are
reported via Session::logErrors when logging is enabled.
SocketConnectorTest now connects a real Session. Its ownership test
adopts a non-stream socket, which the connector accepts and
createSimpleRemoteCAOverSocket then rejects, rather than relying on a
failing GetAttachInfo.
---
.../orc-rt/bedrock/ConnectorRegistry.h | 18 +---
orc-rt/lib/bedrock/ConnectorRegistry.cpp | 10 +-
.../lib/bedrock/sys/posix/SocketConnector.cpp | 12 +--
.../bedrock/sys/posix/SocketConnectorTest.cpp | 42 ++++----
orc-rt/tools/ogre/ogre.cpp | 102 ++++++++++--------
5 files changed, 88 insertions(+), 96 deletions(-)
diff --git a/orc-rt/include/orc-rt/bedrock/ConnectorRegistry.h b/orc-rt/include/orc-rt/bedrock/ConnectorRegistry.h
index 72150f3b66df1..3faf757411673 100644
--- a/orc-rt/include/orc-rt/bedrock/ConnectorRegistry.h
+++ b/orc-rt/include/orc-rt/bedrock/ConnectorRegistry.h
@@ -34,18 +34,10 @@ class Session;
/// ensures that only requested transport mechanisms are available.
class ConnectorRegistry {
public:
- struct AttachInfo {
- Session &S;
- BootstrapInfo BI;
- };
-
- /// Supplies the Session and BootstrapInfo to connect.
- using GetAttachInfoFn = move_only_function<Expected<AttachInfo>() noexcept>;
-
- /// Establishes the connection CS describes and attaches it to the Session
- /// that GetSession returns.
+ /// Establishes the connection CS describes and attaches it to S, handing
+ /// over BI.
using ConnectorFn = move_only_function<Error(
- GetAttachInfoFn GetAttachInfo, const ConnectionSpec &) noexcept>;
+ const ConnectionSpec &CS, Session &S, BootstrapInfo BI) noexcept>;
/// Registers Connector as the handler for Transport.
///
@@ -58,8 +50,8 @@ class ConnectorRegistry {
///
/// Fails if no connector is registered for it, which is how a spec naming a
/// transport this process was not built with is reported.
- Error connect(GetAttachInfoFn GetAttachInfo,
- const ConnectionSpec &CS) noexcept;
+ Error connect(const ConnectionSpec &CS, Session &S,
+ BootstrapInfo BI) noexcept;
private:
std::mutex M;
diff --git a/orc-rt/lib/bedrock/ConnectorRegistry.cpp b/orc-rt/lib/bedrock/ConnectorRegistry.cpp
index c850cca7f4cbd..d50350b76b00e 100644
--- a/orc-rt/lib/bedrock/ConnectorRegistry.cpp
+++ b/orc-rt/lib/bedrock/ConnectorRegistry.cpp
@@ -34,8 +34,8 @@ Error ConnectorRegistry::registerConnector(std::string Transport,
return Error::success();
}
-Error ConnectorRegistry::connect(GetAttachInfoFn GetAttachInfo,
- const ConnectionSpec &CS) noexcept {
+Error ConnectorRegistry::connect(const ConnectionSpec &CS, Session &S,
+ BootstrapInfo BI) noexcept {
ConnectorFn *Connector = nullptr;
{
std::scoped_lock<std::mutex> Lock(M);
@@ -49,11 +49,7 @@ Error ConnectorRegistry::connect(GetAttachInfoFn GetAttachInfo,
Connector = &I->second;
}
- // Run the connector without the lock: it blocks on IO, and may register
- // further connectors or connect again. The pointer stays good because
- // unordered_map does not move its elements on insert, and nothing removes
- // them.
- return (*Connector)(std::move(GetAttachInfo), CS);
+ return (*Connector)(CS, S, std::move(BI));
}
} // namespace orc_rt
diff --git a/orc-rt/lib/bedrock/sys/posix/SocketConnector.cpp b/orc-rt/lib/bedrock/sys/posix/SocketConnector.cpp
index 1ddbd474959e0..dd1fff2719417 100644
--- a/orc-rt/lib/bedrock/sys/posix/SocketConnector.cpp
+++ b/orc-rt/lib/bedrock/sys/posix/SocketConnector.cpp
@@ -24,8 +24,8 @@ using namespace orc_rt;
namespace {
-Error socketConnector(ConnectorRegistry::GetAttachInfoFn GetAttachInfo,
- const ConnectionSpec &CS) noexcept {
+Error socketConnector(const ConnectionSpec &CS, Session &S,
+ BootstrapInfo BI) noexcept {
auto BadCS = [&](const std::string &Reason) noexcept {
return make_error<StringError>((StringOutputStream()
<< "Invalid connection spec \"" << CS.str()
@@ -59,16 +59,12 @@ Error socketConnector(ConnectorRegistry::GetAttachInfoFn GetAttachInfo,
return BadCS("file descriptor " + std::string(FDStr) +
" is not a socket (" + sys::strError(ErrNum) + ")");
}
- SocketHandle Sock(FD);
- auto AI = GetAttachInfo();
- if (!AI)
- return AI.takeError();
- auto CA = createSimpleRemoteCAOverSocket(AI->S, std::move(Sock));
+ auto CA = createSimpleRemoteCAOverSocket(S, SocketHandle(FD));
if (!CA)
return CA.takeError();
- AI->S.attach(std::move(*CA), std::move(AI->BI));
+ S.attach(std::move(*CA), std::move(BI));
return Error::success();
}
diff --git a/orc-rt/test/unit/bedrock/sys/posix/SocketConnectorTest.cpp b/orc-rt/test/unit/bedrock/sys/posix/SocketConnectorTest.cpp
index 43d5d7784612f..2bf7648b1608f 100644
--- a/orc-rt/test/unit/bedrock/sys/posix/SocketConnectorTest.cpp
+++ b/orc-rt/test/unit/bedrock/sys/posix/SocketConnectorTest.cpp
@@ -11,7 +11,10 @@
//===----------------------------------------------------------------------===//
#include "orc-rt/bedrock/SocketConnector.h"
+#include "orc-rt/bedrock/Session.h"
+#include "BedrockTestUtils.h"
+#include "CommonTestUtils.h"
#include "ErrorMatchers.h"
#include "bedrock/SocketTestUtils.h"
@@ -28,22 +31,18 @@ using ::testing::HasSubstr;
namespace {
-/// Runs "socket:adopt=<FD>" through the socket connector. The GetAttachInfo it
-/// supplies records that it was called and then fails, so a descriptor that
-/// passes validation stops there rather than being attached.
-Error adoptFD(int FD, bool &AttachInfoRequested) {
+/// Runs "socket:adopt=<FD>" through the socket connector, targeting a fresh
+/// Session. Every case below fails before attach, so the Session never
+/// connects (and noErrors would catch an unexpected attach failure).
+Error adoptFD(int FD) {
ConnectorRegistry R;
if (auto Err = registerSocketConnector(R))
return Err;
auto CS = ConnectionSpec::parse("socket:adopt=" + std::to_string(FD));
if (!CS)
return CS.takeError();
- return R.connect(
- [&]() noexcept -> Expected<ConnectorRegistry::AttachInfo> {
- AttachInfoRequested = true;
- return make_error<StringError>("attach info requested");
- },
- *CS);
+ Session S(mockExecutorProcessInfo(), noDispatch, noErrors);
+ return R.connect(*CS, S, BootstrapInfo(S));
}
bool isOpen(int FD) { return ::fcntl(FD, F_GETFD) != -1; }
@@ -52,10 +51,8 @@ TEST(SocketConnectorTest, RejectsPipe) {
int P[2];
ASSERT_EQ(::pipe(P), 0);
- bool AttachInfoRequested = false;
- EXPECT_THAT_ERROR(adoptFD(P[0], AttachInfoRequested),
+ EXPECT_THAT_ERROR(adoptFD(P[0]),
FailedWithMessage(HasSubstr("is not a socket")));
- EXPECT_FALSE(AttachInfoRequested);
EXPECT_TRUE(isOpen(P[0])) << "a rejected descriptor must be left open";
::close(P[0]);
@@ -68,20 +65,19 @@ TEST(SocketConnectorTest, RejectsClosedDescriptor) {
::close(P[0]);
::close(P[1]);
- bool AttachInfoRequested = false;
- EXPECT_THAT_ERROR(adoptFD(P[0], AttachInfoRequested),
+ EXPECT_THAT_ERROR(adoptFD(P[0]),
FailedWithMessage(HasSubstr("is not a socket")));
- EXPECT_FALSE(AttachInfoRequested);
}
TEST(SocketConnectorTest, TakesOwnershipOfASocketEvenOnFailure) {
- auto H = makeNativeSocket();
- ASSERT_TRUE(H.has_value());
-
- bool AttachInfoRequested = false;
- EXPECT_THAT_ERROR(adoptFD(*H, AttachInfoRequested),
- FailedWithMessage("attach info requested"));
- EXPECT_TRUE(AttachInfoRequested);
+ // A non-stream socket passes the connector's is-a-socket check, so the
+ // connector takes ownership of it, but is then rejected when creating the
+ // ControllerAccess, which requires a stream socket.
+ auto H = makeNativeNonStreamSocket();
+ ASSERT_TRUE(H.has_value()) << "could not create a socket for the test";
+
+ EXPECT_THAT_ERROR(adoptFD(*H),
+ FailedWithMessage(HasSubstr("requires a stream socket")));
EXPECT_FALSE(isNativeSocketOpen(*H))
<< "an adopted socket must be closed when the connection fails";
}
diff --git a/orc-rt/tools/ogre/ogre.cpp b/orc-rt/tools/ogre/ogre.cpp
index 398622275066a..1e9f1f3f12215 100644
--- a/orc-rt/tools/ogre/ogre.cpp
+++ b/orc-rt/tools/ogre/ogre.cpp
@@ -21,6 +21,8 @@
#include "orc-rt/bedrock/ThreadPoolRunner.h"
#include "orc-rt/bedrock/sps/AllSPSCI.h"
+#include "orc-rt-c/support/Logging.h"
+
#include <cstdio>
#include <cstring>
#include <future>
@@ -76,20 +78,24 @@ static std::variant<Options, int> parseArgs(int argc, char *argv[]) noexcept {
return O;
}
-void reportError(Error Err) noexcept {
- fprintf(stderr, "reported error: %s\n", toString(std::move(Err)).c_str());
+static void reportError(Session &S, Error Err) noexcept {
+#if ORC_RT_LOG_ENABLED(Error)
+ Session::logErrors(S, std::move(Err));
+#else
+ fprintf(stderr, "Session %p error: %s\n", &S,
+ toString(std::move(Err)).c_str());
+#endif // ORC_RT_LOG_ENABLED(Error)
}
-void printExecutorProcessInfo(const ExecutorProcessInfo &EPI) noexcept {
- fprintf(stderr,
- "executor info: triple = \"%s\", page-size = %zu, "
- "cpu-features = \"%s\"\n",
- EPI.targetTriple().c_str(), EPI.pageSize(),
- EPI.targetCPUFeatures().c_str());
+static Expected<Session::DispatchFn> makeDispatcher() noexcept {
+ return [R = std::make_unique<ThreadPoolRunner>(4)](Session::Task T) {
+ (*R)(std::move(T));
+ };
}
-Error setupSession(Session &S, const Options &Opts,
- BootstrapInfo &BI) noexcept {
+/// Adds the services a host executor provides, publishing their entry points
+/// in BI for the controller.
+static Error addHostServices(Session &S, BootstrapInfo &BI) noexcept {
if (auto Err = sps_ci::addAll(BI.symbols()))
return Err;
@@ -104,53 +110,59 @@ Error setupSession(Session &S, const Options &Opts,
return Error::success();
}
-Error trySetupAndConnect(Session &S, const Options &Opts) noexcept {
- ConnectorRegistry ConnRegistry;
- if (auto Err = registerSocketConnector(ConnRegistry))
- return Err;
- // registerTCPConnect(ConnRegistry);
+static Expected<std::unique_ptr<Session>>
+makeSession(const Options &Opts) noexcept {
+ auto EPI = ExecutorProcessInfo::Detect();
+ if (!EPI)
+ return EPI.takeError();
- auto BI = BootstrapInfo::CreateDefault(S);
+ if (Opts.Verbose) {
+ fprintf(stderr,
+ "executor info: triple = \"%s\", page-size = %zu, "
+ "cpu-features = \"%s\"\n",
+ EPI->targetTriple().c_str(), EPI->pageSize(),
+ EPI->targetCPUFeatures().c_str());
+ }
+
+ auto D = makeDispatcher();
+ if (!D)
+ return D.takeError();
+
+ auto S =
+ std::make_unique<Session>(std::move(*EPI), std::move(*D), reportError);
+
+ auto BI = BootstrapInfo::CreateDefault(*S);
if (!BI)
return BI.takeError();
- if (auto Err = setupSession(S, Opts, *BI))
+ if (auto Err = addHostServices(*S, *BI))
return Err;
- return ConnRegistry.connect(
- [&]() noexcept -> Expected<ConnectorRegistry::AttachInfo> {
- return ConnectorRegistry::AttachInfo{S, std::move(*BI)};
- },
- Opts.ConnSpec);
-}
+ ConnectorRegistry Connectors;
+ if (auto Err = registerSocketConnector(Connectors))
+ return Err;
-Expected<int> runOgre(const Options &Opts) noexcept {
- // Get the process info.
- auto EPI = ExecutorProcessInfo::Detect();
- if (!EPI)
- return EPI.takeError();
- if (Opts.Verbose)
- printExecutorProcessInfo(*EPI);
+ if (auto Err = Connectors.connect(Opts.ConnSpec, *S, std::move(*BI)))
+ return Err;
- // Build the session.
- ThreadPoolRunner Run(4);
- Session S(
- std::move(*EPI), [&Run](Session::Task T) { Run(std::move(T)); },
- [](Session &, Error Err) noexcept { reportError(std::move(Err)); });
+ return S;
+}
- std::promise<void> StopP;
- auto StopF = StopP.get_future();
- S.setOnDisconnect([StopP = std::move(StopP)](Error Err) mutable noexcept {
- if (Err)
- reportError(std::move(Err));
- StopP.set_value();
- });
+static Expected<int> runOgre(const Options &Opts) noexcept {
+ auto S = makeSession(Opts);
+ if (!S)
+ return S.takeError();
- if (auto Err = trySetupAndConnect(S, Opts))
- return Err;
+ // makeSession has already connected, so the Session may have detached by
+ // now. That's fine: addOnDetach runs the callback immediately if so, so the
+ // wait below cannot miss the detach.
+ std::promise<void> StopP;
+ std::future<void> StopF = StopP.get_future();
+ (*S)->addOnDetach(
+ [StopP = std::move(StopP)]() mutable noexcept { StopP.set_value(); });
+ // Wait for detach.
StopF.get();
-
return 0;
}
>From b54588918b8932eac1524b843e6ae59dbbe1a029 Mon Sep 17 00:00:00 2001
From: Letu Ren <fantasquex at gmail.com>
Date: Sun, 27 Sep 2026 17:45:15 +0800
Subject: [PATCH 46/53] [CIR][NFC] Add vector of boolean test to
__builtin_nondeterministic_value (#219550)
Assisted-by: grok-4.6
Signed-off-by: Letu Ren <fantasquex at gmail.com>
---
.../builtin-nondeterministic-value.c | 18 +++++++++++++-----
1 file changed, 13 insertions(+), 5 deletions(-)
diff --git a/clang/test/CIR/CodeGenBuiltins/builtin-nondeterministic-value.c b/clang/test/CIR/CodeGenBuiltins/builtin-nondeterministic-value.c
index d7d9224fc0440..4ef5a9308fa52 100644
--- a/clang/test/CIR/CodeGenBuiltins/builtin-nondeterministic-value.c
+++ b/clang/test/CIR/CodeGenBuiltins/builtin-nondeterministic-value.c
@@ -6,11 +6,7 @@
// RUN: FileCheck --input-file=%t.ll %s -check-prefix=LLVM
typedef float float4 __attribute__((ext_vector_type(4)));
-
-// TODO: the bool4 (ext_vector_type _Bool) case from the classic CodeGen test is
-// omitted here: CIR does not yet implement storing ext-vector-bool types
-// (emitStoreOfScalar ExtVectorBoolType is NYI). The cir.freeze lowering itself
-// works for <N x !cir.bool>; only the surrounding store is unsupported.
+typedef _Bool bool4 __attribute__((ext_vector_type(4)));
int clang_nondet_i(int x) {
return __builtin_nondeterministic_value(x);
@@ -66,3 +62,15 @@ void clang_nondet_fv(void) {
// LLVM-LABEL: @clang_nondet_fv
// LLVM: %[[RES:.*]] = freeze <4 x float> poison
+
+void clang_nondet_bv(void) {
+ bool4 x = __builtin_nondeterministic_value(x);
+}
+
+// CIR-LABEL: cir.func {{.*}}@clang_nondet_bv
+// CIR: %[[POISON:.*]] = cir.const #cir.poison : !cir.vector<4 x !cir.bool>
+// CIR: %[[RES:.*]] = cir.freeze %[[POISON]] : !cir.vector<4 x !cir.bool>
+// CIR: cir.store{{.*}} %[[RES]], {{.*}} : !cir.vector<4 x !cir.bool>, !cir.ptr<!cir.vector<4 x !cir.bool>>
+
+// LLVM-LABEL: @clang_nondet_bv
+// LLVM: %[[RES:.*]] = freeze <4 x i1> poison
>From d3a1d348a1e6dd2b33aebdca501edf4a4b0145bc Mon Sep 17 00:00:00 2001
From: Mamadou Wane <mamadouswane at gmail.com>
Date: Sun, 27 Sep 2026 06:18:14 -0400
Subject: [PATCH 47/53] [clang-tidy] Fix redundant-branch-condition false
positive in loops (#225827)
Fixes #205685.
`bugprone-redundant-branch-condition` decides whether the condition
variable changes between the outer and inner `if` by comparing source
positions. When a loop sits between the two, a mutation that comes after
the inner `if` in the source still runs before the inner condition is
evaluated again on the next iteration, so the check reports a redundant
condition that isn't, and the fix-it changes behavior.
The fix walks the parents of the inner `if` up to the outer `if` and
finds the outermost enclosing loop. If the variable is mutated anywhere
in that loop, the check does not warn. This follows NagyDonat's
suggestion in the issue: the existing check already covers mutations
between the outer condition and the loop, so only the loop itself needs
the extra query. The walk passes through declarations, so an inner `if`
inside a lambda stored in a variable is handled too, and it stops at the
enclosing function. I chose a syntactic walk instead of the CFG
(`utils::ExprSequence`), since it covers the range the issue asks for
without building a CFG on every match.
Tests cover `for`, `while`, `do`, and range-based `for`, a mutation in a
`do`/`while` condition, a lambda stored in a variable inside the loop,
and an inner loop whose enclosing loop does the mutation. Each of these
warns without the patch. Three positive cases confirm the warning is
kept when the loop never mutates the variable, when the mutation comes
after the loop, and when the loop encloses both `if` statements.
Limitations:
- Any mutation in the loop suppresses the warning, including one
followed by an exit from the loop (`break`, `return`, `throw`) or one
that cannot make the condition false. The inner condition is still
redundant in those cases; telling them apart needs the CFG. I added the
`break` case under Unhandled Cases. The change only removes warnings, so
it adds no new false positives.
- Loops built with `goto` and Objective-C `for ... in` loops are not
recognized, so the false positive remains there.
I used Claude to help write the code. I tested and reviewed all of it
myself, and worked through the approach and suggestions with Claude.
---
.../RedundantBranchConditionCheck.cpp | 34 ++++
clang-tools-extra/docs/ReleaseNotes.md | 5 +
.../bugprone/redundant-branch-condition.cpp | 164 ++++++++++++++++++
3 files changed, 203 insertions(+)
diff --git a/clang-tools-extra/clang-tidy/bugprone/RedundantBranchConditionCheck.cpp b/clang-tools-extra/clang-tidy/bugprone/RedundantBranchConditionCheck.cpp
index a6458d96055c3..8d608dd4b0e66 100644
--- a/clang-tools-extra/clang-tidy/bugprone/RedundantBranchConditionCheck.cpp
+++ b/clang-tools-extra/clang-tidy/bugprone/RedundantBranchConditionCheck.cpp
@@ -10,6 +10,8 @@
#include "../utils/Aliasing.h"
#include "../utils/LexerUtils.h"
#include "clang/AST/ASTContext.h"
+#include "clang/AST/ParentMapContext.h"
+#include "clang/AST/StmtCXX.h"
#include "clang/ASTMatchers/ASTMatchFinder.h"
#include "clang/Analysis/Analyses/ExprMutationAnalyzer.h"
#include "clang/Lex/Lexer.h"
@@ -40,6 +42,31 @@ static bool isChangedBefore(const Stmt *S, const Stmt *NextS, const Stmt *PrevS,
SM.isBeforeInTranslationUnit(MutS->getEndLoc(), NextS->getBeginLoc());
}
+/// Returns the outermost loop that encloses `S` and is itself enclosed by
+/// `Outer`, or null if there is no such loop. The walk passes through
+/// declarations, such as a variable initialized by a lambda, but stops at the
+/// enclosing function.
+static const Stmt *getOutermostLoopBetween(const Stmt *S, const Stmt *Outer,
+ ASTContext *Context) {
+ const Stmt *Loop = nullptr;
+ // getParents() returns only the direct parents of a node, usually exactly
+ // one, so the walk calls it once per level.
+ DynTypedNodeList Parents = Context->getParents(*S);
+ while (!Parents.empty()) {
+ const DynTypedNode Parent = Parents[0];
+ if (Parent.get<FunctionDecl>())
+ break;
+ if (const auto *ParentStmt = Parent.get<Stmt>()) {
+ if (ParentStmt == Outer)
+ break;
+ if (isa<ForStmt, WhileStmt, DoStmt, CXXForRangeStmt>(ParentStmt))
+ Loop = ParentStmt;
+ }
+ Parents = Context->getParents(Parent);
+ }
+ return Loop;
+}
+
void RedundantBranchConditionCheck::registerMatchers(MatchFinder *Finder) {
const auto ImmutableVar =
varDecl(anyOf(parmVarDecl(), hasLocalStorage()), hasType(isInteger()),
@@ -98,6 +125,13 @@ void RedundantBranchConditionCheck::check(
return;
}
+ // Inside a loop, a mutation anywhere in the loop runs before the inner
+ // condition is evaluated again, even if it comes later in the source.
+ const Stmt *Loop = getOutermostLoopBetween(InnerIf, OuterIf, Result.Context);
+ if (Loop &&
+ ExprMutationAnalyzer(*Loop, *Result.Context).findMutation(CondVar))
+ return;
+
// If the variable has an alias then it can be changed by that alias as well.
// FIXME: could potentially support tracking pointers and references in the
// future to improve catching true positives through aliases.
diff --git a/clang-tools-extra/docs/ReleaseNotes.md b/clang-tools-extra/docs/ReleaseNotes.md
index 833638a47abc6..3d9e34cc4d44e 100644
--- a/clang-tools-extra/docs/ReleaseNotes.md
+++ b/clang-tools-extra/docs/ReleaseNotes.md
@@ -185,6 +185,11 @@ infrastructure are described first, followed by tool-specific sections.
<clang-tidy/checks/bugprone/pointer-arithmetic-on-polymorphic-object>` when
the pointer points to an incomplete (forward-declared) type.
+- Improved {doc}`bugprone-redundant-branch-condition
+ <clang-tidy/checks/bugprone/redundant-branch-condition>` check by fixing
+ false positives when the condition variable is changed later in a loop that
+ encloses the inner `if`.
+
- Fixed a crash in {doc}`bugprone-std-namespace-modification
<clang-tidy/checks/bugprone/std-namespace-modification>` when checking
lambda closure types used as template arguments.
diff --git a/clang-tools-extra/test/clang-tidy/checkers/bugprone/redundant-branch-condition.cpp b/clang-tools-extra/test/clang-tidy/checkers/bugprone/redundant-branch-condition.cpp
index 40994b0ff884e..ad2238ff9984f 100644
--- a/clang-tools-extra/test/clang-tidy/checkers/bugprone/redundant-branch-condition.cpp
+++ b/clang-tools-extra/test/clang-tidy/checkers/bugprone/redundant-branch-condition.cpp
@@ -1127,6 +1127,52 @@ int positive_expr_with_cleanups() {
return 0;
}
+// Loops
+
+void positive_loop_not_mutated() {
+ bool onFire = isBurning();
+ if (onFire) {
+ while (someOtherCondition()) {
+ if (onFire) {
+ // CHECK-MESSAGES: :[[@LINE-1]]:7: warning: redundant condition 'onFire' [bugprone-redundant-branch-condition]
+ // CHECK-FIXES: {{^\ *$}}
+ scream();
+ }
+ // CHECK-FIXES: {{^\ *$}}
+ }
+ }
+}
+
+void positive_loop_mutated_after_loop() {
+ bool onFire = isBurning();
+ if (onFire) {
+ while (someOtherCondition()) {
+ if (onFire) {
+ // CHECK-MESSAGES: :[[@LINE-1]]:7: warning: redundant condition 'onFire' [bugprone-redundant-branch-condition]
+ // CHECK-FIXES: {{^\ *$}}
+ scream();
+ }
+ // CHECK-FIXES: {{^\ *$}}
+ }
+ tryToExtinguish(onFire);
+ }
+}
+
+void positive_loop_around_both_ifs() {
+ bool onFire = isBurning();
+ while (someOtherCondition()) {
+ if (onFire) {
+ if (onFire) {
+ // CHECK-MESSAGES: :[[@LINE-1]]:7: warning: redundant condition 'onFire' [bugprone-redundant-branch-condition]
+ // CHECK-FIXES: {{^\ *$}}
+ scream();
+ }
+ // CHECK-FIXES: {{^\ *$}}
+ }
+ tryToExtinguish(onFire);
+ }
+}
+
//===--- Special Negatives ------------------------------------------------===//
// Aliasing
@@ -1351,6 +1397,109 @@ void negative_comma_after_condition() {
}
}
+// Loops
+
+void negative_for_mutated_later_in_body(int n) {
+ bool onFire = isBurning();
+ if (onFire) {
+ for (int i = 0; i < n; ++i) {
+ switch (i) {
+ case 4:
+ case 5:
+ if (onFire) {
+ // NO-MESSAGE: fire may have been extinguished in a previous iteration
+ onFire = false;
+ scream();
+ }
+ break;
+ }
+ }
+ }
+}
+
+void negative_while_mutated_later_in_body() {
+ bool onFire = isBurning();
+ if (onFire) {
+ while (someOtherCondition()) {
+ if (onFire) {
+ // NO-MESSAGE: fire may have been extinguished in a previous iteration
+ scream();
+ }
+ tryToExtinguish(onFire);
+ }
+ }
+}
+
+void negative_do_mutated_later_in_body() {
+ bool onFire = isBurning();
+ if (onFire) {
+ do {
+ if (onFire) {
+ // NO-MESSAGE: fire may have been extinguished in a previous iteration
+ scream();
+ }
+ onFire = isBurning();
+ } while (someOtherCondition());
+ }
+}
+
+void negative_range_for_mutated_later_in_body() {
+ bool onFire = isBurning();
+ int floors[3] = {1, 2, 3};
+ if (onFire) {
+ for (int floor : floors) {
+ if (onFire) {
+ // NO-MESSAGE: fire may have been extinguished in a previous iteration
+ scream();
+ }
+ onFire = floor > 1;
+ }
+ }
+}
+
+void negative_loop_condition_mutates() {
+ bool onFire = isBurning();
+ if (onFire) {
+ do {
+ if (onFire) {
+ // NO-MESSAGE: fire may have been extinguished by the loop condition
+ scream();
+ }
+ } while (tryToExtinguish(onFire));
+ }
+}
+
+void negative_loop_mutated_after_lambda_variable() {
+ bool onFire = isBurning();
+ if (onFire) {
+ while (someOtherCondition()) {
+ auto check = [onFire] {
+ if (onFire) {
+ // NO-MESSAGE: fire may have been extinguished in a previous iteration
+ scream();
+ }
+ };
+ check();
+ tryToExtinguish(onFire);
+ }
+ }
+}
+
+void negative_mutated_in_outer_loop() {
+ bool onFire = isBurning();
+ if (onFire) {
+ while (someOtherCondition()) {
+ for (int i = 0; i < 3; ++i) {
+ if (onFire) {
+ // NO-MESSAGE: fire may have been extinguished in a previous iteration
+ scream();
+ }
+ }
+ tryToExtinguish(onFire);
+ }
+ }
+}
+
//===--- Unhandled Cases --------------------------------------------------===//
void negated_in_else() {
@@ -1397,3 +1546,18 @@ void volatile_concrete_address() {
}
}
}
+
+void loop_mutated_then_break() {
+ bool onFire = isBurning();
+ if (onFire) {
+ while (someOtherCondition()) {
+ if (onFire) {
+ // Redundant, but not diagnosed: the loop exits before onFire is checked
+ // again. Telling this apart from a later mutation needs the CFG.
+ onFire = false;
+ scream();
+ break;
+ }
+ }
+ }
+}
>From 6df49b28d7c6982fdebc4bfecde30e473a0c00bc Mon Sep 17 00:00:00 2001
From: Florian Hahn <flo at fhahn.com>
Date: Sun, 27 Sep 2026 11:21:19 +0100
Subject: [PATCH 48/53] [TTI] Provide conservative legality + costs for
@llvm.speculative.load. (#180036)
Add TTI support for @llvm.speculative.load, including cost and legaltiy
checking support.
The initial implementation for AArch64 checks if the loaded type is
<= 16 bytes, and only considers such cases legal due to MTE.
PR:
https://github.com/llvm/llvm-project/pull/180036
Depends on https://github.com/llvm/llvm-project/pull/179642
---
.../llvm/Analysis/TargetTransformInfo.h | 6 ++
.../llvm/Analysis/TargetTransformInfoImpl.h | 6 ++
llvm/include/llvm/CodeGen/BasicTTIImpl.h | 11 +++
llvm/lib/Analysis/TargetTransformInfo.cpp | 5 ++
.../AArch64/AArch64TargetTransformInfo.cpp | 12 ++++
.../AArch64/AArch64TargetTransformInfo.h | 3 +
.../CostModel/AArch64/speculative-load.ll | 70 +++++++++++--------
.../CostModel/X86/speculative-load.ll | 10 +--
8 files changed, 88 insertions(+), 35 deletions(-)
diff --git a/llvm/include/llvm/Analysis/TargetTransformInfo.h b/llvm/include/llvm/Analysis/TargetTransformInfo.h
index ba74de115ac3d..a33e6f62e941e 100644
--- a/llvm/include/llvm/Analysis/TargetTransformInfo.h
+++ b/llvm/include/llvm/Analysis/TargetTransformInfo.h
@@ -941,6 +941,12 @@ class TargetTransformInfo {
isLegalMaskedLoad(Type *DataType, Align Alignment, unsigned AddressSpace,
MaskKind MaskKind = VariableOrConstantMask) const;
+ /// Return true if the target supports speculatively loading \p DataType from
+ /// address space \p AddressSpace, i.e. @llvm.can.load.speculatively can
+ /// return true for the store size of \p DataType.
+ LLVM_ABI bool isLegalSpeculativeLoad(Type *DataType,
+ unsigned AddressSpace) const;
+
/// Return true if the target supports nontemporal store.
LLVM_ABI bool isLegalNTStore(Type *DataType, Align Alignment) const;
/// Return true if the target supports nontemporal load.
diff --git a/llvm/include/llvm/Analysis/TargetTransformInfoImpl.h b/llvm/include/llvm/Analysis/TargetTransformInfoImpl.h
index 3a37af79cf0e1..a19c122c16f20 100644
--- a/llvm/include/llvm/Analysis/TargetTransformInfoImpl.h
+++ b/llvm/include/llvm/Analysis/TargetTransformInfoImpl.h
@@ -368,6 +368,11 @@ class LLVM_ABI TargetTransformInfoImplBase {
return false;
}
+ virtual bool isLegalSpeculativeLoad(Type *DataType,
+ unsigned AddressSpace) const {
+ return false;
+ }
+
virtual bool isLegalNTStore(Type *DataType, Align Alignment) const {
// By default, assume nontemporal memory stores are available for stores
// that are aligned and have a size that is a power of 2.
@@ -990,6 +995,7 @@ class LLVM_ABI TargetTransformInfoImplBase {
case Intrinsic::vp_gather:
case Intrinsic::masked_compressstore:
case Intrinsic::masked_expandload:
+ case Intrinsic::speculative_load:
return 1;
}
return InstructionCost::getInvalid();
diff --git a/llvm/include/llvm/CodeGen/BasicTTIImpl.h b/llvm/include/llvm/CodeGen/BasicTTIImpl.h
index 57aa72c5a9a11..9fffe55325421 100644
--- a/llvm/include/llvm/CodeGen/BasicTTIImpl.h
+++ b/llvm/include/llvm/CodeGen/BasicTTIImpl.h
@@ -2523,6 +2523,13 @@ class BasicTTIImplBase : public TargetTransformInfoImplCRTPBase<T> {
return thisT()->getMemIntrinsicInstrCost(
MemIntrinsicCostAttributes(IID, Ty, TyAlign, 0), CostKind);
}
+ case Intrinsic::speculative_load: {
+ const IntrinsicInst *I = ICA.getInst();
+ Align Alignment = I ? I->getParamAlign(0).valueOrOne() : Align(1);
+ unsigned AS = Tys[0]->getPointerAddressSpace();
+ return thisT()->getMemIntrinsicInstrCost(
+ MemIntrinsicCostAttributes(IID, RetTy, Alignment, AS), CostKind);
+ }
case Intrinsic::experimental_vp_strided_store: {
auto *Ty = cast<VectorType>(ICA.getArgTypes()[0]);
Align Alignment = thisT()->DL.getABITypeAlign(Ty->getElementType());
@@ -3279,6 +3286,10 @@ class BasicTTIImplBase : public TargetTransformInfoImplCRTPBase<T> {
}
case Intrinsic::vp_load_ff:
return InstructionCost::getInvalid();
+ case Intrinsic::speculative_load:
+ // Speculative loads are lowered to regular loads of the full type.
+ return thisT()->getMemoryOpCost(Instruction::Load, DataTy, Alignment,
+ MICA.getAddressSpace(), CostKind);
default:
llvm_unreachable("unexpected intrinsic");
}
diff --git a/llvm/lib/Analysis/TargetTransformInfo.cpp b/llvm/lib/Analysis/TargetTransformInfo.cpp
index 1479246c1d0ef..af73a615f25fa 100644
--- a/llvm/lib/Analysis/TargetTransformInfo.cpp
+++ b/llvm/lib/Analysis/TargetTransformInfo.cpp
@@ -494,6 +494,11 @@ bool TargetTransformInfo::isLegalMaskedLoad(Type *DataType, Align Alignment,
MaskKind);
}
+bool TargetTransformInfo::isLegalSpeculativeLoad(Type *DataType,
+ unsigned AddressSpace) const {
+ return TTIImpl->isLegalSpeculativeLoad(DataType, AddressSpace);
+}
+
bool TargetTransformInfo::isLegalNTStore(Type *DataType,
Align Alignment) const {
return TTIImpl->isLegalNTStore(DataType, Alignment);
diff --git a/llvm/lib/Target/AArch64/AArch64TargetTransformInfo.cpp b/llvm/lib/Target/AArch64/AArch64TargetTransformInfo.cpp
index 5b690e5e8043c..ddea78fa4e808 100644
--- a/llvm/lib/Target/AArch64/AArch64TargetTransformInfo.cpp
+++ b/llvm/lib/Target/AArch64/AArch64TargetTransformInfo.cpp
@@ -5947,6 +5947,18 @@ bool AArch64TTIImpl::isLegalMaskedExpandLoad(Type *DataTy,
(ST->isSVEorStreamingSVEAvailable() && ST->hasSME2p2());
}
+bool AArch64TTIImpl::isLegalSpeculativeLoad(Type *DataType,
+ unsigned AddressSpace) const {
+ // Matches AArch64TargetLowering::emitCanLoadSpeculatively: only address
+ // space 0 and power-of-2 sizes up to the 16-byte MTE tag granule.
+ // TODO: Support scalable vectors.
+ if (AddressSpace != 0)
+ return false;
+ TypeSize Size = DL.getTypeStoreSize(DataType);
+ return !Size.isScalable() && isPowerOf2_64(Size.getFixedValue()) &&
+ Size.getFixedValue() <= 16;
+}
+
unsigned
AArch64TTIImpl::getMaxInterleaveFactor(ElementCount VF,
bool HasUnorderedReductions) const {
diff --git a/llvm/lib/Target/AArch64/AArch64TargetTransformInfo.h b/llvm/lib/Target/AArch64/AArch64TargetTransformInfo.h
index f086ba1844965..d090f69c1476a 100644
--- a/llvm/lib/Target/AArch64/AArch64TargetTransformInfo.h
+++ b/llvm/lib/Target/AArch64/AArch64TargetTransformInfo.h
@@ -279,6 +279,9 @@ class AArch64TTIImpl final : public BasicTTIImplBase<AArch64TTIImpl> {
bool isLegalMaskedExpandLoad(Type *DataTy, Align Alignment) const override;
+ bool isLegalSpeculativeLoad(Type *DataType,
+ unsigned AddressSpace) const override;
+
void getUnrollingPreferences(Loop *L, ScalarEvolution &SE,
TTI::UnrollingPreferences &UP,
OptimizationRemarkEmitter *ORE) const override;
diff --git a/llvm/test/Analysis/CostModel/AArch64/speculative-load.ll b/llvm/test/Analysis/CostModel/AArch64/speculative-load.ll
index 4b7b906ce3fee..343757e03d631 100644
--- a/llvm/test/Analysis/CostModel/AArch64/speculative-load.ll
+++ b/llvm/test/Analysis/CostModel/AArch64/speculative-load.ll
@@ -1,30 +1,30 @@
; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py
-; RUN: opt -passes="print<cost-model>" 2>&1 -disable-output -mtriple=aarch64 < %s | FileCheck %s --check-prefixes=COMMON
-; RUN: opt -passes="print<cost-model>" 2>&1 -disable-output -mtriple=aarch64 -mattr=+sve < %s | FileCheck %s --check-prefixes=COMMON
+; RUN: opt -passes="print<cost-model>" 2>&1 -disable-output -mtriple=aarch64 < %s | FileCheck %s --check-prefixes=COMMON,NOSVE
+; RUN: opt -passes="print<cost-model>" 2>&1 -disable-output -mtriple=aarch64 -mattr=+sve < %s | FileCheck %s --check-prefixes=COMMON,SVE
define void @speculative_load_cost_fixed(ptr %p) {
- ; Scalar types - all valid (<= 16 bytes)
+ ; Scalar types (<= 16 bytes)
; COMMON-LABEL: 'speculative_load_cost_fixed'
; COMMON-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %1 = call b8 (ptr, i1, ...) @llvm.speculative.load.b8.p0(ptr %p, i1 false, i64 0)
; COMMON-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %2 = call b16 (ptr, i1, ...) @llvm.speculative.load.b16.p0(ptr %p, i1 false, i64 0)
; COMMON-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %3 = call b32 (ptr, i1, ...) @llvm.speculative.load.b32.p0(ptr %p, i1 false, i64 0)
; COMMON-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %4 = call b64 (ptr, i1, ...) @llvm.speculative.load.b64.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %5 = call b128 (ptr, i1, ...) @llvm.speculative.load.b128.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 6 for instruction: %6 = call <2 x i32> (ptr, i1, ...) @llvm.speculative.load.v2i32.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 12 for instruction: %7 = call <4 x i32> (ptr, i1, ...) @llvm.speculative.load.v4i32.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 6 for instruction: %8 = call <2 x i64> (ptr, i1, ...) @llvm.speculative.load.v2i64.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 10 for instruction: %9 = call <4 x float> (ptr, i1, ...) @llvm.speculative.load.v4f32.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 4 for instruction: %10 = call <2 x double> (ptr, i1, ...) @llvm.speculative.load.v2f64.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 24 for instruction: %11 = call <8 x i8> (ptr, i1, ...) @llvm.speculative.load.v8i8.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 48 for instruction: %12 = call <16 x i8> (ptr, i1, ...) @llvm.speculative.load.v16i8.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 12 for instruction: %13 = call <4 x i16> (ptr, i1, ...) @llvm.speculative.load.v4i16.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 24 for instruction: %14 = call <8 x i16> (ptr, i1, ...) @llvm.speculative.load.v8i16.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 24 for instruction: %15 = call <8 x i32> (ptr, i1, ...) @llvm.speculative.load.v8i32.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 12 for instruction: %16 = call <4 x i64> (ptr, i1, ...) @llvm.speculative.load.v4i64.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 96 for instruction: %17 = call <32 x i8> (ptr, i1, ...) @llvm.speculative.load.v32i8.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 48 for instruction: %18 = call <16 x i16> (ptr, i1, ...) @llvm.speculative.load.v16i16.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 20 for instruction: %19 = call <8 x float> (ptr, i1, ...) @llvm.speculative.load.v8f32.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 8 for instruction: %20 = call <4 x double> (ptr, i1, ...) @llvm.speculative.load.v4f64.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 2 for instruction: %5 = call b128 (ptr, i1, ...) @llvm.speculative.load.b128.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %6 = call <2 x i32> (ptr, i1, ...) @llvm.speculative.load.v2i32.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %7 = call <4 x i32> (ptr, i1, ...) @llvm.speculative.load.v4i32.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %8 = call <2 x i64> (ptr, i1, ...) @llvm.speculative.load.v2i64.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %9 = call <4 x float> (ptr, i1, ...) @llvm.speculative.load.v4f32.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %10 = call <2 x double> (ptr, i1, ...) @llvm.speculative.load.v2f64.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %11 = call <8 x i8> (ptr, i1, ...) @llvm.speculative.load.v8i8.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %12 = call <16 x i8> (ptr, i1, ...) @llvm.speculative.load.v16i8.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %13 = call <4 x i16> (ptr, i1, ...) @llvm.speculative.load.v4i16.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %14 = call <8 x i16> (ptr, i1, ...) @llvm.speculative.load.v8i16.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 2 for instruction: %15 = call <8 x i32> (ptr, i1, ...) @llvm.speculative.load.v8i32.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 2 for instruction: %16 = call <4 x i64> (ptr, i1, ...) @llvm.speculative.load.v4i64.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 2 for instruction: %17 = call <32 x i8> (ptr, i1, ...) @llvm.speculative.load.v32i8.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 2 for instruction: %18 = call <16 x i16> (ptr, i1, ...) @llvm.speculative.load.v16i16.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 2 for instruction: %19 = call <8 x float> (ptr, i1, ...) @llvm.speculative.load.v8f32.p0(ptr %p, i1 false, i64 0)
+; COMMON-NEXT: Cost Model: Found an estimated cost of 2 for instruction: %20 = call <4 x double> (ptr, i1, ...) @llvm.speculative.load.v4f64.p0(ptr %p, i1 false, i64 0)
; COMMON-NEXT: Cost Model: Found an estimated cost of 0 for instruction: ret void
;
call b8 (ptr, i1, ...) @llvm.speculative.load.b8.p0(ptr %p, i1 false, i64 0)
@@ -33,7 +33,7 @@ define void @speculative_load_cost_fixed(ptr %p) {
call b64 (ptr, i1, ...) @llvm.speculative.load.b64.p0(ptr %p, i1 false, i64 0)
call b128 (ptr, i1, ...) @llvm.speculative.load.b128.p0(ptr %p, i1 false, i64 0)
- ; Vector types <= 16 bytes - valid
+ ; Vector types <= 16 bytes
call <2 x i32> (ptr, i1, ...) @llvm.speculative.load.v2i32.p0(ptr %p, i1 false, i64 0)
call <4 x i32> (ptr, i1, ...) @llvm.speculative.load.v4i32.p0(ptr %p, i1 false, i64 0)
call <2 x i64> (ptr, i1, ...) @llvm.speculative.load.v2i64.p0(ptr %p, i1 false, i64 0)
@@ -44,7 +44,7 @@ define void @speculative_load_cost_fixed(ptr %p) {
call <4 x i16> (ptr, i1, ...) @llvm.speculative.load.v4i16.p0(ptr %p, i1 false, i64 0)
call <8 x i16> (ptr, i1, ...) @llvm.speculative.load.v8i16.p0(ptr %p, i1 false, i64 0)
- ; Vector types > 16 bytes - invalid
+ ; Vector types > 16 bytes
call <8 x i32> (ptr, i1, ...) @llvm.speculative.load.v8i32.p0(ptr %p, i1 false, i64 0)
call <4 x i64> (ptr, i1, ...) @llvm.speculative.load.v4i64.p0(ptr %p, i1 false, i64 0)
call <32 x i8> (ptr, i1, ...) @llvm.speculative.load.v32i8.p0(ptr %p, i1 false, i64 0)
@@ -55,15 +55,25 @@ define void @speculative_load_cost_fixed(ptr %p) {
}
define void @speculative_load_cost_scalable(ptr %p) {
-; COMMON-LABEL: 'speculative_load_cost_scalable'
-; COMMON-NEXT: Cost Model: Invalid cost for instruction: %1 = call <vscale x 2 x i64> (ptr, i1, ...) @llvm.speculative.load.nxv2i64.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Invalid cost for instruction: %2 = call <vscale x 4 x i32> (ptr, i1, ...) @llvm.speculative.load.nxv4i32.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Invalid cost for instruction: %3 = call <vscale x 8 x i16> (ptr, i1, ...) @llvm.speculative.load.nxv8i16.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Invalid cost for instruction: %4 = call <vscale x 16 x i8> (ptr, i1, ...) @llvm.speculative.load.nxv16i8.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Invalid cost for instruction: %5 = call <vscale x 2 x double> (ptr, i1, ...) @llvm.speculative.load.nxv2f64.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Invalid cost for instruction: %6 = call <vscale x 4 x float> (ptr, i1, ...) @llvm.speculative.load.nxv4f32.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Invalid cost for instruction: %7 = call <vscale x 8 x float> (ptr, i1, ...) @llvm.speculative.load.nxv8f32.p0(ptr %p, i1 false, i64 0)
-; COMMON-NEXT: Cost Model: Found an estimated cost of 0 for instruction: ret void
+; NOSVE-LABEL: 'speculative_load_cost_scalable'
+; NOSVE-NEXT: Cost Model: Invalid cost for instruction: %1 = call <vscale x 2 x i64> (ptr, i1, ...) @llvm.speculative.load.nxv2i64.p0(ptr %p, i1 false, i64 0)
+; NOSVE-NEXT: Cost Model: Invalid cost for instruction: %2 = call <vscale x 4 x i32> (ptr, i1, ...) @llvm.speculative.load.nxv4i32.p0(ptr %p, i1 false, i64 0)
+; NOSVE-NEXT: Cost Model: Invalid cost for instruction: %3 = call <vscale x 8 x i16> (ptr, i1, ...) @llvm.speculative.load.nxv8i16.p0(ptr %p, i1 false, i64 0)
+; NOSVE-NEXT: Cost Model: Invalid cost for instruction: %4 = call <vscale x 16 x i8> (ptr, i1, ...) @llvm.speculative.load.nxv16i8.p0(ptr %p, i1 false, i64 0)
+; NOSVE-NEXT: Cost Model: Invalid cost for instruction: %5 = call <vscale x 2 x double> (ptr, i1, ...) @llvm.speculative.load.nxv2f64.p0(ptr %p, i1 false, i64 0)
+; NOSVE-NEXT: Cost Model: Invalid cost for instruction: %6 = call <vscale x 4 x float> (ptr, i1, ...) @llvm.speculative.load.nxv4f32.p0(ptr %p, i1 false, i64 0)
+; NOSVE-NEXT: Cost Model: Invalid cost for instruction: %7 = call <vscale x 8 x float> (ptr, i1, ...) @llvm.speculative.load.nxv8f32.p0(ptr %p, i1 false, i64 0)
+; NOSVE-NEXT: Cost Model: Found an estimated cost of 0 for instruction: ret void
+;
+; SVE-LABEL: 'speculative_load_cost_scalable'
+; SVE-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %1 = call <vscale x 2 x i64> (ptr, i1, ...) @llvm.speculative.load.nxv2i64.p0(ptr %p, i1 false, i64 0)
+; SVE-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %2 = call <vscale x 4 x i32> (ptr, i1, ...) @llvm.speculative.load.nxv4i32.p0(ptr %p, i1 false, i64 0)
+; SVE-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %3 = call <vscale x 8 x i16> (ptr, i1, ...) @llvm.speculative.load.nxv8i16.p0(ptr %p, i1 false, i64 0)
+; SVE-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %4 = call <vscale x 16 x i8> (ptr, i1, ...) @llvm.speculative.load.nxv16i8.p0(ptr %p, i1 false, i64 0)
+; SVE-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %5 = call <vscale x 2 x double> (ptr, i1, ...) @llvm.speculative.load.nxv2f64.p0(ptr %p, i1 false, i64 0)
+; SVE-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %6 = call <vscale x 4 x float> (ptr, i1, ...) @llvm.speculative.load.nxv4f32.p0(ptr %p, i1 false, i64 0)
+; SVE-NEXT: Cost Model: Found an estimated cost of 2 for instruction: %7 = call <vscale x 8 x float> (ptr, i1, ...) @llvm.speculative.load.nxv8f32.p0(ptr %p, i1 false, i64 0)
+; SVE-NEXT: Cost Model: Found an estimated cost of 0 for instruction: ret void
;
call <vscale x 2 x i64> (ptr, i1, ...) @llvm.speculative.load.nxv2i64.p0(ptr %p, i1 false, i64 0)
call <vscale x 4 x i32> (ptr, i1, ...) @llvm.speculative.load.nxv4i32.p0(ptr %p, i1 false, i64 0)
diff --git a/llvm/test/Analysis/CostModel/X86/speculative-load.ll b/llvm/test/Analysis/CostModel/X86/speculative-load.ll
index 81b3b56eeb10c..2dba86ceabff9 100644
--- a/llvm/test/Analysis/CostModel/X86/speculative-load.ll
+++ b/llvm/test/Analysis/CostModel/X86/speculative-load.ll
@@ -7,11 +7,11 @@ define void @speculative_load_cost(ptr %p) {
; CHECK-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %2 = call b16 (ptr, i1, ...) @llvm.speculative.load.b16.p0(ptr %p, i1 false, i64 0)
; CHECK-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %3 = call b32 (ptr, i1, ...) @llvm.speculative.load.b32.p0(ptr %p, i1 false, i64 0)
; CHECK-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %4 = call b64 (ptr, i1, ...) @llvm.speculative.load.b64.p0(ptr %p, i1 false, i64 0)
-; CHECK-NEXT: Cost Model: Found an estimated cost of 11 for instruction: %5 = call <4 x i32> (ptr, i1, ...) @llvm.speculative.load.v4i32.p0(ptr %p, i1 false, i64 0)
-; CHECK-NEXT: Cost Model: Found an estimated cost of 22 for instruction: %6 = call <8 x i32> (ptr, i1, ...) @llvm.speculative.load.v8i32.p0(ptr %p, i1 false, i64 0)
-; CHECK-NEXT: Cost Model: Found an estimated cost of 5 for instruction: %7 = call <2 x i64> (ptr, i1, ...) @llvm.speculative.load.v2i64.p0(ptr %p, i1 false, i64 0)
-; CHECK-NEXT: Cost Model: Found an estimated cost of 7 for instruction: %8 = call <4 x float> (ptr, i1, ...) @llvm.speculative.load.v4f32.p0(ptr %p, i1 false, i64 0)
-; CHECK-NEXT: Cost Model: Found an estimated cost of 3 for instruction: %9 = call <2 x double> (ptr, i1, ...) @llvm.speculative.load.v2f64.p0(ptr %p, i1 false, i64 0)
+; CHECK-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %5 = call <4 x i32> (ptr, i1, ...) @llvm.speculative.load.v4i32.p0(ptr %p, i1 false, i64 0)
+; CHECK-NEXT: Cost Model: Found an estimated cost of 2 for instruction: %6 = call <8 x i32> (ptr, i1, ...) @llvm.speculative.load.v8i32.p0(ptr %p, i1 false, i64 0)
+; CHECK-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %7 = call <2 x i64> (ptr, i1, ...) @llvm.speculative.load.v2i64.p0(ptr %p, i1 false, i64 0)
+; CHECK-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %8 = call <4 x float> (ptr, i1, ...) @llvm.speculative.load.v4f32.p0(ptr %p, i1 false, i64 0)
+; CHECK-NEXT: Cost Model: Found an estimated cost of 1 for instruction: %9 = call <2 x double> (ptr, i1, ...) @llvm.speculative.load.v2f64.p0(ptr %p, i1 false, i64 0)
; CHECK-NEXT: Cost Model: Found an estimated cost of 0 for instruction: ret void
;
call b8 (ptr, i1, ...) @llvm.speculative.load.b8.p0(ptr %p, i1 false, i64 0)
>From 8f8aebcf3fbb656f3423a56889e2711b1d8b3d0b Mon Sep 17 00:00:00 2001
From: Florian Hahn <flo at fhahn.com>
Date: Sun, 27 Sep 2026 11:21:50 +0100
Subject: [PATCH 49/53] [ConstraintElim] Don't request SCEV if there are no
loops. (NFC) (#226736)
ConstraintElimination only uses SCEV when there are any loops. Don't
request the analysis if there are no loops.
This reduces compile-time a bit (more for workloads with large number of
functions w/o loops):
stage1-O3: -0.03%
stage1-ReleaseThinLTO: -0.04%
stage1-ReleaseLTO-g: -0.02%
stage1-aarch64-O3: -0.03%
stage2-O3: -0.04%
stage2-clang: -0.09%
https://llvm-compile-time-tracker.com/compare.php?from=899d817c7950c997402cd229935cd822acf45b08&to=c4453240417d8705b0752c3feda8d02365f8175c&stat=instructions:u
PR: https://github.com/llvm/llvm-project/pull/226736
---
.../Scalar/ConstraintElimination.cpp | 24 ++++++++++---------
.../analysis-invalidation.ll | 9 +++----
2 files changed, 16 insertions(+), 17 deletions(-)
diff --git a/llvm/lib/Transforms/Scalar/ConstraintElimination.cpp b/llvm/lib/Transforms/Scalar/ConstraintElimination.cpp
index 0d1dae4e667c0..ae5f85374ad39 100644
--- a/llvm/lib/Transforms/Scalar/ConstraintElimination.cpp
+++ b/llvm/lib/Transforms/Scalar/ConstraintElimination.cpp
@@ -217,11 +217,12 @@ struct MonotonicInfo {
struct State {
DominatorTree &DT;
LoopInfo &LI;
- ScalarEvolution &SE;
+ /// Only available for functions with loops.
+ ScalarEvolution *SE;
TargetLibraryInfo &TLI;
SmallVector<FactOrCheck, 64> WorkList;
- State(DominatorTree &DT, LoopInfo &LI, ScalarEvolution &SE,
+ State(DominatorTree &DT, LoopInfo &LI, ScalarEvolution *SE,
TargetLibraryInfo &TLI)
: DT(DT), LI(LI), SE(SE), TLI(TLI) {}
@@ -1150,14 +1151,14 @@ MonotonicInfo State::getMonotonicityInfo(PHINode &PN, Value *Step) {
if (Info.Unsigned || Info.Signed || !StepOffset)
return Info;
- const auto *AR = dyn_cast<SCEVAddRecExpr>(SE.getSCEV(&PN));
+ const auto *AR = dyn_cast<SCEVAddRecExpr>(SE->getSCEV(&PN));
if (!AR)
return Info;
ScalarEvolution::MonotonicPredicateType Expected =
Info.Decreasing ? ScalarEvolution::MonotonicallyDecreasing
: ScalarEvolution::MonotonicallyIncreasing;
auto IsMonotonic = [&](CmpInst::Predicate Pred) {
- return SE.getMonotonicPredicateType(AR, Pred) == Expected;
+ return SE->getMonotonicPredicateType(AR, Pred) == Expected;
};
Info.Signed = IsMonotonic(CmpInst::ICMP_SGT);
Info.Unsigned = !Info.Decreasing && IsMonotonic(CmpInst::ICMP_UGT);
@@ -1238,7 +1239,7 @@ void State::addInfoForInductions(BasicBlock &BB) {
}
if (PN->getParent() != Header || PN->getNumIncomingValues() != 2 ||
- !SE.isSCEVable(PN->getType()))
+ !SE->isSCEVable(PN->getType()))
return;
// For latch conditions, we need to inject the condition that holds for the
@@ -1307,7 +1308,7 @@ void State::addInfoForInductions(BasicBlock &BB) {
if (StepOffset->isZero())
return;
} else {
- const SCEV *Expr = SE.getSCEV(PN);
+ const SCEV *Expr = SE->getSCEV(PN);
if (!match(Expr,
m_scev_AffineAddRec(m_SCEV(StartSCEV), m_scev_APInt(StepOffset),
m_SpecificLoop(L))))
@@ -1362,10 +1363,10 @@ void State::addInfoForInductions(BasicBlock &BB) {
if (!StepOffset->isOne()) {
// Check whether B-Start is known to be a multiple of StepOffset.
if (!StartSCEV)
- StartSCEV = SE.getSCEV(StartValue);
- const SCEV *BMinusStart = SE.getMinusSCEV(SE.getSCEV(B), StartSCEV);
+ StartSCEV = SE->getSCEV(StartValue);
+ const SCEV *BMinusStart = SE->getMinusSCEV(SE->getSCEV(B), StartSCEV);
if (isa<SCEVCouldNotCompute>(BMinusStart) ||
- !SE.getConstantMultiple(BMinusStart).urem(*StepOffset).isZero())
+ !SE->getConstantMultiple(BMinusStart).urem(*StepOffset).isZero())
return;
}
@@ -2359,7 +2360,7 @@ tryToSimplifyOverflowMath(WithOverflowInst *II, ConstraintInfo &Info,
}
static bool eliminateConstraints(Function &F, DominatorTree &DT, LoopInfo &LI,
- ScalarEvolution &SE,
+ ScalarEvolution *SE,
OptimizationRemarkEmitter &ORE,
TargetLibraryInfo &TLI) {
bool Changed = false;
@@ -2698,7 +2699,8 @@ PreservedAnalyses ConstraintEliminationPass::run(Function &F,
FunctionAnalysisManager &AM) {
auto &DT = AM.getResult<DominatorTreeAnalysis>(F);
auto &LI = AM.getResult<LoopAnalysis>(F);
- auto &SE = AM.getResult<ScalarEvolutionAnalysis>(F);
+ // SCEV is only used for loops, only construct it if there are some.
+ auto *SE = LI.empty() ? nullptr : &AM.getResult<ScalarEvolutionAnalysis>(F);
auto &ORE = AM.getResult<OptimizationRemarkEmitterAnalysis>(F);
auto &TLI = AM.getResult<TargetLibraryAnalysis>(F);
if (!eliminateConstraints(F, DT, LI, SE, ORE, TLI))
diff --git a/llvm/test/Transforms/ConstraintElimination/analysis-invalidation.ll b/llvm/test/Transforms/ConstraintElimination/analysis-invalidation.ll
index cb51f8c852503..85d310798fa52 100644
--- a/llvm/test/Transforms/ConstraintElimination/analysis-invalidation.ll
+++ b/llvm/test/Transforms/ConstraintElimination/analysis-invalidation.ll
@@ -12,9 +12,8 @@
; CHECK-NEXT: Running analysis: DominatorTreeAnalysis on ssub_no_overflow_due_to_or_conds
; CHECK-NEXT: Running pass: ConstraintEliminationPass on ssub_no_overflow_due_to_or_conds
; CHECK-NEXT: Running analysis: LoopAnalysis on ssub_no_overflow_due_to_or_conds
-; CHECK-NEXT: Running analysis: ScalarEvolutionAnalysis on ssub_no_overflow_due_to_or_conds
-; CHECK-NEXT: Running analysis: TargetLibraryAnalysis on ssub_no_overflow_due_to_or_conds
; CHECK-NEXT: Running analysis: OptimizationRemarkEmitterAnalysis on ssub_no_overflow_due_to_or_conds
+; CHECK-NEXT: Running analysis: TargetLibraryAnalysis on ssub_no_overflow_due_to_or_conds
; CHECK-NEXT: Invalidating analysis: DemandedBitsAnalysis on ssub_no_overflow_due_to_or_conds
; CHECK-NEXT: Running pass: RequireAnalysisPass
; CHECK-NEXT: Running analysis: DemandedBitsAnalysis on ssub_no_overflow_due_to_or_conds
@@ -26,9 +25,8 @@
; CHECK-NEXT: Running analysis: DominatorTreeAnalysis on uge_zext
; CHECK-NEXT: Running pass: ConstraintEliminationPass on uge_zext
; CHECK-NEXT: Running analysis: LoopAnalysis on uge_zext
-; CHECK-NEXT: Running analysis: ScalarEvolutionAnalysis on uge_zext
-; CHECK-NEXT: Running analysis: TargetLibraryAnalysis on uge_zext
; CHECK-NEXT: Running analysis: OptimizationRemarkEmitterAnalysis on uge_zext
+; CHECK-NEXT: Running analysis: TargetLibraryAnalysis on uge_zext
; CHECK-NEXT: Invalidating analysis: DemandedBitsAnalysis on uge_zext
; CHECK-NEXT: Running pass: RequireAnalysisPass
; CHECK-NEXT: Running analysis: DemandedBitsAnalysis on uge_zext
@@ -40,9 +38,8 @@
; CHECK-NEXT: Running analysis: DominatorTreeAnalysis on test_mul_const_nuw_unsigned_14
; CHECK-NEXT: Running pass: ConstraintEliminationPass on test_mul_const_nuw_unsigned_14
; CHECK-NEXT: Running analysis: LoopAnalysis on test_mul_const_nuw_unsigned_14
-; CHECK-NEXT: Running analysis: ScalarEvolutionAnalysis on test_mul_const_nuw_unsigned_14
-; CHECK-NEXT: Running analysis: TargetLibraryAnalysis on test_mul_const_nuw_unsigned_14
; CHECK-NEXT: Running analysis: OptimizationRemarkEmitterAnalysis on test_mul_const_nuw_unsigned_14
+; CHECK-NEXT: Running analysis: TargetLibraryAnalysis on test_mul_const_nuw_unsigned_14
declare { i8, i1 } @llvm.ssub.with.overflow.i8(i8, i8)
>From 30d6051bb5126fa2bf48717b3cd3fc2c90896608 Mon Sep 17 00:00:00 2001
From: "Chibuoyim (Wilson) Ogbonna" <chibuoyim.faith.ogbonna at huawei.com>
Date: Sun, 27 Sep 2026 12:56:40 +0100
Subject: [PATCH 50/53] [clang][CIR] Add tests for SVE ADDV intrinsics
(#225648)
This adds CIR tests for `svaddv` (plain SVE) intrinsics following the
task description in
[223963](https://github.com/llvm/llvm-project/issues/223963).
Moves and adapts `sve-intrinsics/acle_sve_addv.c` to `sve/addv.c`.
---
.../AArch64/sve-intrinsics/acle_sve_addv.c | 206 ------------------
clang/test/CodeGen/AArch64/sve/addv.c | 205 +++++++++++++++++
2 files changed, 205 insertions(+), 206 deletions(-)
delete mode 100644 clang/test/CodeGen/AArch64/sve-intrinsics/acle_sve_addv.c
create mode 100644 clang/test/CodeGen/AArch64/sve/addv.c
diff --git a/clang/test/CodeGen/AArch64/sve-intrinsics/acle_sve_addv.c b/clang/test/CodeGen/AArch64/sve-intrinsics/acle_sve_addv.c
deleted file mode 100644
index efd0046998957..0000000000000
--- a/clang/test/CodeGen/AArch64/sve-intrinsics/acle_sve_addv.c
+++ /dev/null
@@ -1,206 +0,0 @@
-// NOTE: Assertions have been autogenerated by utils/update_cc_test_checks.py
-// REQUIRES: aarch64-registered-target
-// RUN: %clang_cc1 -triple aarch64 -target-feature +sve -disable-O0-optnone -Werror -Wall -emit-llvm -o - %s | opt -S -passes=mem2reg,tailcallelim | FileCheck %s
-// RUN: %clang_cc1 -triple aarch64 -target-feature +sve -disable-O0-optnone -Werror -Wall -emit-llvm -o - -x c++ %s | opt -S -passes=mem2reg,tailcallelim | FileCheck %s -check-prefix=CPP-CHECK
-// RUN: %clang_cc1 -DSVE_OVERLOADED_FORMS -triple aarch64 -target-feature +sve -disable-O0-optnone -Werror -Wall -emit-llvm -o - %s | opt -S -passes=mem2reg,tailcallelim | FileCheck %s
-// RUN: %clang_cc1 -DSVE_OVERLOADED_FORMS -triple aarch64 -target-feature +sve -disable-O0-optnone -Werror -Wall -emit-llvm -o - -x c++ %s | opt -S -passes=mem2reg,tailcallelim | FileCheck %s -check-prefix=CPP-CHECK
-// RUN: %clang_cc1 -triple aarch64 -target-feature +sve -S -disable-O0-optnone -Werror -Wall -o /dev/null %s
-// RUN: %clang_cc1 -triple aarch64 -target-feature +sme -S -disable-O0-optnone -Werror -Wall -o /dev/null %s
-
-#include <arm_sve.h>
-
-#if defined __ARM_FEATURE_SME
-#define MODE_ATTR __arm_streaming
-#else
-#define MODE_ATTR
-#endif
-
-#ifdef SVE_OVERLOADED_FORMS
-// A simple used,unused... macro, long enough to represent any SVE builtin.
-#define SVE_ACLE_FUNC(A1,A2_UNUSED,A3,A4_UNUSED) A1##A3
-#else
-#define SVE_ACLE_FUNC(A1,A2,A3,A4) A1##A2##A3##A4
-#endif
-
-// CHECK-LABEL: @test_svaddv_s8(
-// CHECK-NEXT: entry:
-// CHECK-NEXT: [[TMP0:%.*]] = tail call i64 @llvm.aarch64.sve.saddv.nxv16i8(<vscale x 16 x i1> [[PG:%.*]], <vscale x 16 x i8> [[OP:%.*]])
-// CHECK-NEXT: ret i64 [[TMP0]]
-//
-// CPP-CHECK-LABEL: @_Z14test_svaddv_s8u10__SVBool_tu10__SVInt8_t(
-// CPP-CHECK-NEXT: entry:
-// CPP-CHECK-NEXT: [[TMP0:%.*]] = tail call i64 @llvm.aarch64.sve.saddv.nxv16i8(<vscale x 16 x i1> [[PG:%.*]], <vscale x 16 x i8> [[OP:%.*]])
-// CPP-CHECK-NEXT: ret i64 [[TMP0]]
-//
-int64_t test_svaddv_s8(svbool_t pg, svint8_t op) MODE_ATTR
-{
- return SVE_ACLE_FUNC(svaddv,_s8,,)(pg, op);
-}
-
-// CHECK-LABEL: @test_svaddv_s16(
-// CHECK-NEXT: entry:
-// CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 8 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv8i1(<vscale x 16 x i1> [[PG:%.*]])
-// CHECK-NEXT: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.saddv.nxv8i16(<vscale x 8 x i1> [[TMP0]], <vscale x 8 x i16> [[OP:%.*]])
-// CHECK-NEXT: ret i64 [[TMP1]]
-//
-// CPP-CHECK-LABEL: @_Z15test_svaddv_s16u10__SVBool_tu11__SVInt16_t(
-// CPP-CHECK-NEXT: entry:
-// CPP-CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 8 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv8i1(<vscale x 16 x i1> [[PG:%.*]])
-// CPP-CHECK-NEXT: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.saddv.nxv8i16(<vscale x 8 x i1> [[TMP0]], <vscale x 8 x i16> [[OP:%.*]])
-// CPP-CHECK-NEXT: ret i64 [[TMP1]]
-//
-int64_t test_svaddv_s16(svbool_t pg, svint16_t op) MODE_ATTR
-{
- return SVE_ACLE_FUNC(svaddv,_s16,,)(pg, op);
-}
-
-// CHECK-LABEL: @test_svaddv_s32(
-// CHECK-NEXT: entry:
-// CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 4 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv4i1(<vscale x 16 x i1> [[PG:%.*]])
-// CHECK-NEXT: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.saddv.nxv4i32(<vscale x 4 x i1> [[TMP0]], <vscale x 4 x i32> [[OP:%.*]])
-// CHECK-NEXT: ret i64 [[TMP1]]
-//
-// CPP-CHECK-LABEL: @_Z15test_svaddv_s32u10__SVBool_tu11__SVInt32_t(
-// CPP-CHECK-NEXT: entry:
-// CPP-CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 4 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv4i1(<vscale x 16 x i1> [[PG:%.*]])
-// CPP-CHECK-NEXT: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.saddv.nxv4i32(<vscale x 4 x i1> [[TMP0]], <vscale x 4 x i32> [[OP:%.*]])
-// CPP-CHECK-NEXT: ret i64 [[TMP1]]
-//
-int64_t test_svaddv_s32(svbool_t pg, svint32_t op) MODE_ATTR
-{
- return SVE_ACLE_FUNC(svaddv,_s32,,)(pg, op);
-}
-
-// CHECK-LABEL: @test_svaddv_s64(
-// CHECK-NEXT: entry:
-// CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 2 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv2i1(<vscale x 16 x i1> [[PG:%.*]])
-// CHECK-NEXT: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.saddv.nxv2i64(<vscale x 2 x i1> [[TMP0]], <vscale x 2 x i64> [[OP:%.*]])
-// CHECK-NEXT: ret i64 [[TMP1]]
-//
-// CPP-CHECK-LABEL: @_Z15test_svaddv_s64u10__SVBool_tu11__SVInt64_t(
-// CPP-CHECK-NEXT: entry:
-// CPP-CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 2 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv2i1(<vscale x 16 x i1> [[PG:%.*]])
-// CPP-CHECK-NEXT: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.saddv.nxv2i64(<vscale x 2 x i1> [[TMP0]], <vscale x 2 x i64> [[OP:%.*]])
-// CPP-CHECK-NEXT: ret i64 [[TMP1]]
-//
-int64_t test_svaddv_s64(svbool_t pg, svint64_t op) MODE_ATTR
-{
- return SVE_ACLE_FUNC(svaddv,_s64,,)(pg, op);
-}
-
-// CHECK-LABEL: @test_svaddv_u8(
-// CHECK-NEXT: entry:
-// CHECK-NEXT: [[TMP0:%.*]] = tail call i64 @llvm.aarch64.sve.uaddv.nxv16i8(<vscale x 16 x i1> [[PG:%.*]], <vscale x 16 x i8> [[OP:%.*]])
-// CHECK-NEXT: ret i64 [[TMP0]]
-//
-// CPP-CHECK-LABEL: @_Z14test_svaddv_u8u10__SVBool_tu11__SVUint8_t(
-// CPP-CHECK-NEXT: entry:
-// CPP-CHECK-NEXT: [[TMP0:%.*]] = tail call i64 @llvm.aarch64.sve.uaddv.nxv16i8(<vscale x 16 x i1> [[PG:%.*]], <vscale x 16 x i8> [[OP:%.*]])
-// CPP-CHECK-NEXT: ret i64 [[TMP0]]
-//
-uint64_t test_svaddv_u8(svbool_t pg, svuint8_t op) MODE_ATTR
-{
- return SVE_ACLE_FUNC(svaddv,_u8,,)(pg, op);
-}
-
-// CHECK-LABEL: @test_svaddv_u16(
-// CHECK-NEXT: entry:
-// CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 8 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv8i1(<vscale x 16 x i1> [[PG:%.*]])
-// CHECK-NEXT: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.uaddv.nxv8i16(<vscale x 8 x i1> [[TMP0]], <vscale x 8 x i16> [[OP:%.*]])
-// CHECK-NEXT: ret i64 [[TMP1]]
-//
-// CPP-CHECK-LABEL: @_Z15test_svaddv_u16u10__SVBool_tu12__SVUint16_t(
-// CPP-CHECK-NEXT: entry:
-// CPP-CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 8 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv8i1(<vscale x 16 x i1> [[PG:%.*]])
-// CPP-CHECK-NEXT: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.uaddv.nxv8i16(<vscale x 8 x i1> [[TMP0]], <vscale x 8 x i16> [[OP:%.*]])
-// CPP-CHECK-NEXT: ret i64 [[TMP1]]
-//
-uint64_t test_svaddv_u16(svbool_t pg, svuint16_t op) MODE_ATTR
-{
- return SVE_ACLE_FUNC(svaddv,_u16,,)(pg, op);
-}
-
-// CHECK-LABEL: @test_svaddv_u32(
-// CHECK-NEXT: entry:
-// CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 4 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv4i1(<vscale x 16 x i1> [[PG:%.*]])
-// CHECK-NEXT: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.uaddv.nxv4i32(<vscale x 4 x i1> [[TMP0]], <vscale x 4 x i32> [[OP:%.*]])
-// CHECK-NEXT: ret i64 [[TMP1]]
-//
-// CPP-CHECK-LABEL: @_Z15test_svaddv_u32u10__SVBool_tu12__SVUint32_t(
-// CPP-CHECK-NEXT: entry:
-// CPP-CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 4 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv4i1(<vscale x 16 x i1> [[PG:%.*]])
-// CPP-CHECK-NEXT: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.uaddv.nxv4i32(<vscale x 4 x i1> [[TMP0]], <vscale x 4 x i32> [[OP:%.*]])
-// CPP-CHECK-NEXT: ret i64 [[TMP1]]
-//
-uint64_t test_svaddv_u32(svbool_t pg, svuint32_t op) MODE_ATTR
-{
- return SVE_ACLE_FUNC(svaddv,_u32,,)(pg, op);
-}
-
-// CHECK-LABEL: @test_svaddv_u64(
-// CHECK-NEXT: entry:
-// CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 2 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv2i1(<vscale x 16 x i1> [[PG:%.*]])
-// CHECK-NEXT: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.uaddv.nxv2i64(<vscale x 2 x i1> [[TMP0]], <vscale x 2 x i64> [[OP:%.*]])
-// CHECK-NEXT: ret i64 [[TMP1]]
-//
-// CPP-CHECK-LABEL: @_Z15test_svaddv_u64u10__SVBool_tu12__SVUint64_t(
-// CPP-CHECK-NEXT: entry:
-// CPP-CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 2 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv2i1(<vscale x 16 x i1> [[PG:%.*]])
-// CPP-CHECK-NEXT: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.uaddv.nxv2i64(<vscale x 2 x i1> [[TMP0]], <vscale x 2 x i64> [[OP:%.*]])
-// CPP-CHECK-NEXT: ret i64 [[TMP1]]
-//
-uint64_t test_svaddv_u64(svbool_t pg, svuint64_t op) MODE_ATTR
-{
- return SVE_ACLE_FUNC(svaddv,_u64,,)(pg, op);
-}
-
-// CHECK-LABEL: @test_svaddv_f16(
-// CHECK-NEXT: entry:
-// CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 8 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv8i1(<vscale x 16 x i1> [[PG:%.*]])
-// CHECK-NEXT: [[TMP1:%.*]] = tail call half @llvm.aarch64.sve.faddv.nxv8f16(<vscale x 8 x i1> [[TMP0]], <vscale x 8 x half> [[OP:%.*]])
-// CHECK-NEXT: ret half [[TMP1]]
-//
-// CPP-CHECK-LABEL: @_Z15test_svaddv_f16u10__SVBool_tu13__SVFloat16_t(
-// CPP-CHECK-NEXT: entry:
-// CPP-CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 8 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv8i1(<vscale x 16 x i1> [[PG:%.*]])
-// CPP-CHECK-NEXT: [[TMP1:%.*]] = tail call half @llvm.aarch64.sve.faddv.nxv8f16(<vscale x 8 x i1> [[TMP0]], <vscale x 8 x half> [[OP:%.*]])
-// CPP-CHECK-NEXT: ret half [[TMP1]]
-//
-float16_t test_svaddv_f16(svbool_t pg, svfloat16_t op) MODE_ATTR
-{
- return SVE_ACLE_FUNC(svaddv,_f16,,)(pg, op);
-}
-
-// CHECK-LABEL: @test_svaddv_f32(
-// CHECK-NEXT: entry:
-// CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 4 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv4i1(<vscale x 16 x i1> [[PG:%.*]])
-// CHECK-NEXT: [[TMP1:%.*]] = tail call float @llvm.aarch64.sve.faddv.nxv4f32(<vscale x 4 x i1> [[TMP0]], <vscale x 4 x float> [[OP:%.*]])
-// CHECK-NEXT: ret float [[TMP1]]
-//
-// CPP-CHECK-LABEL: @_Z15test_svaddv_f32u10__SVBool_tu13__SVFloat32_t(
-// CPP-CHECK-NEXT: entry:
-// CPP-CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 4 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv4i1(<vscale x 16 x i1> [[PG:%.*]])
-// CPP-CHECK-NEXT: [[TMP1:%.*]] = tail call float @llvm.aarch64.sve.faddv.nxv4f32(<vscale x 4 x i1> [[TMP0]], <vscale x 4 x float> [[OP:%.*]])
-// CPP-CHECK-NEXT: ret float [[TMP1]]
-//
-float32_t test_svaddv_f32(svbool_t pg, svfloat32_t op) MODE_ATTR
-{
- return SVE_ACLE_FUNC(svaddv,_f32,,)(pg, op);
-}
-
-// CHECK-LABEL: @test_svaddv_f64(
-// CHECK-NEXT: entry:
-// CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 2 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv2i1(<vscale x 16 x i1> [[PG:%.*]])
-// CHECK-NEXT: [[TMP1:%.*]] = tail call double @llvm.aarch64.sve.faddv.nxv2f64(<vscale x 2 x i1> [[TMP0]], <vscale x 2 x double> [[OP:%.*]])
-// CHECK-NEXT: ret double [[TMP1]]
-//
-// CPP-CHECK-LABEL: @_Z15test_svaddv_f64u10__SVBool_tu13__SVFloat64_t(
-// CPP-CHECK-NEXT: entry:
-// CPP-CHECK-NEXT: [[TMP0:%.*]] = tail call <vscale x 2 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv2i1(<vscale x 16 x i1> [[PG:%.*]])
-// CPP-CHECK-NEXT: [[TMP1:%.*]] = tail call double @llvm.aarch64.sve.faddv.nxv2f64(<vscale x 2 x i1> [[TMP0]], <vscale x 2 x double> [[OP:%.*]])
-// CPP-CHECK-NEXT: ret double [[TMP1]]
-//
-float64_t test_svaddv_f64(svbool_t pg, svfloat64_t op) MODE_ATTR
-{
- return SVE_ACLE_FUNC(svaddv,_f64,,)(pg, op);
-}
diff --git a/clang/test/CodeGen/AArch64/sve/addv.c b/clang/test/CodeGen/AArch64/sve/addv.c
new file mode 100644
index 0000000000000..bf9e693173dcf
--- /dev/null
+++ b/clang/test/CodeGen/AArch64/sve/addv.c
@@ -0,0 +1,205 @@
+// REQUIRES: aarch64-registered-target
+
+// DEFINE: %{optimize} = opt -passes=mem2reg,instcombine,tailcallelim -S
+
+// RUN: %if cir-enabled %{%clang_cc1_cg_arm64_sve -fclangir -emit-cir -disable-O0-optnone -o - %s | FileCheck %s --check-prefixes=C,CIR %}
+// RUN: %if cir-enabled %{%clang_cc1_cg_arm64_sve -DSVE_OVERLOADED_FORMS -fclangir -emit-cir -disable-O0-optnone -o - %s | FileCheck %s --check-prefixes=C,CIR %}
+
+// RUN: %if cir-enabled %{%clang_cc1_cg_arm64_sve -fclangir -emit-llvm -disable-O0-optnone -o - %s | %{optimize} | FileCheck %s --check-prefixes=C,LLVM %}
+// RUN: %if cir-enabled %{%clang_cc1_cg_arm64_sve -DSVE_OVERLOADED_FORMS -fclangir -emit-llvm -disable-O0-optnone -o - %s | %{optimize} | FileCheck %s --check-prefixes=C,LLVM %}
+// RUN: %if cir-enabled %{%clang_cc1_cg_arm64_sve -DSVE_OVERLOADED_FORMS -fclangir -emit-llvm -disable-O0-optnone -o - -x c++ %s | %{optimize} | FileCheck %s --check-prefixes=CPP,LLVM %}
+
+// RUN: %clang_cc1_cg_arm64_sve -emit-llvm -disable-O0-optnone -o - %s | %{optimize} | FileCheck %s --check-prefixes=C,LLVM
+// RUN: %clang_cc1_cg_arm64_sve -DSVE_OVERLOADED_FORMS -emit-llvm -disable-O0-optnone -o - %s | %{optimize} | FileCheck %s --check-prefixes=C,LLVM
+// RUN: %clang_cc1_cg_arm64_sve -DSVE_OVERLOADED_FORMS -emit-llvm -disable-O0-optnone -o - -x c++ %s | %{optimize} | FileCheck %s --check-prefixes=CPP,LLVM
+
+//=============================================================================
+// NOTES
+//
+// Tests for SVE ADDV intrinsics
+//=============================================================================
+
+#include <arm_sve.h>
+
+#if defined __ARM_FEATURE_SME
+#define MODE_ATTR __arm_streaming
+#else
+#define MODE_ATTR
+#endif
+
+#ifdef SVE_OVERLOADED_FORMS
+// A simple used,unused... macro, long enough to represent any SVE builtin.
+#define SVE_ACLE_FUNC(A1,A2_UNUSED,A3,A4_UNUSED) A1##A3
+#else
+#define SVE_ACLE_FUNC(A1,A2,A3,A4) A1##A2##A3##A4
+#endif
+
+// C-LABEL: @test_svaddv_s8(
+// CPP-LABEL: @_Z14test_svaddv_s8u10__SVBool_tu10__SVInt8_t(
+int64_t test_svaddv_s8(svbool_t pg, svint8_t op) MODE_ATTR
+{
+// CIR: %[[RES:.*]] = cir.call_llvm_intrinsic "aarch64.sve.saddv" %{{.*}}, %{{.*}} :
+// CIR-SAME: -> !s64i
+
+// LLVM-SAME: <vscale x 16 x i1> [[PG:%.*]], <vscale x 16 x i8> [[OP:%.*]])
+// LLVM: [[TMP0:%.*]] = tail call i64 @llvm.aarch64.sve.saddv.nxv16i8(<vscale x 16 x i1> [[PG]], <vscale x 16 x i8> [[OP]])
+// LLVM: ret i64 [[TMP0]]
+ return SVE_ACLE_FUNC(svaddv,_s8,,)(pg, op);
+}
+
+// C-LABEL: @test_svaddv_s16(
+// CPP-LABEL: @_Z15test_svaddv_s16u10__SVBool_tu11__SVInt16_t(
+int64_t test_svaddv_s16(svbool_t pg, svint16_t op) MODE_ATTR
+{
+// CIR: %[[CONVERT_PG:.*]] = cir.call_llvm_intrinsic "aarch64.sve.convert.from.svbool" %{{.*}} :
+// CIR-SAME: -> !cir.vector<[8] x !cir.int<u, 1>>
+// CIR: %[[RES:.*]] = cir.call_llvm_intrinsic "aarch64.sve.saddv" %[[CONVERT_PG]], %{{.*}} :
+// CIR-SAME: -> !s64i
+
+// LLVM-SAME: <vscale x 16 x i1> [[PG:%.*]], <vscale x 8 x i16> [[OP:%.*]])
+// LLVM: [[TMP0:%.*]] = tail call <vscale x 8 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv8i1(<vscale x 16 x i1> [[PG]])
+// LLVM: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.saddv.nxv8i16(<vscale x 8 x i1> [[TMP0]], <vscale x 8 x i16> [[OP]])
+// LLVM: ret i64 [[TMP1]]
+ return SVE_ACLE_FUNC(svaddv,_s16,,)(pg, op);
+}
+
+// C-LABEL: @test_svaddv_s32(
+// CPP-LABEL: @_Z15test_svaddv_s32u10__SVBool_tu11__SVInt32_t(
+int64_t test_svaddv_s32(svbool_t pg, svint32_t op) MODE_ATTR
+{
+// CIR: %[[CONVERT_PG:.*]] = cir.call_llvm_intrinsic "aarch64.sve.convert.from.svbool" %{{.*}} :
+// CIR-SAME: -> !cir.vector<[4] x !cir.int<u, 1>>
+// CIR: %[[RES:.*]] = cir.call_llvm_intrinsic "aarch64.sve.saddv" %[[CONVERT_PG]], %{{.*}} :
+// CIR-SAME: -> !s64i
+
+// LLVM-SAME: <vscale x 16 x i1> [[PG:%.*]], <vscale x 4 x i32> [[OP:%.*]])
+// LLVM: [[TMP0:%.*]] = tail call <vscale x 4 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv4i1(<vscale x 16 x i1> [[PG]])
+// LLVM: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.saddv.nxv4i32(<vscale x 4 x i1> [[TMP0]], <vscale x 4 x i32> [[OP]])
+// LLVM: ret i64 [[TMP1]]
+ return SVE_ACLE_FUNC(svaddv,_s32,,)(pg, op);
+}
+
+// C-LABEL: @test_svaddv_s64(
+// CPP-LABEL: @_Z15test_svaddv_s64u10__SVBool_tu11__SVInt64_t(
+int64_t test_svaddv_s64(svbool_t pg, svint64_t op) MODE_ATTR
+{
+// CIR: %[[CONVERT_PG:.*]] = cir.call_llvm_intrinsic "aarch64.sve.convert.from.svbool" %{{.*}} :
+// CIR-SAME: -> !cir.vector<[2] x !cir.int<u, 1>>
+// CIR: %[[RES:.*]] = cir.call_llvm_intrinsic "aarch64.sve.saddv" %[[CONVERT_PG]], %{{.*}} :
+// CIR-SAME: -> !s64i
+
+// LLVM-SAME: <vscale x 16 x i1> [[PG:%.*]], <vscale x 2 x i64> [[OP:%.*]])
+// LLVM: [[TMP0:%.*]] = tail call <vscale x 2 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv2i1(<vscale x 16 x i1> [[PG]])
+// LLVM: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.saddv.nxv2i64(<vscale x 2 x i1> [[TMP0]], <vscale x 2 x i64> [[OP]])
+// LLVM: ret i64 [[TMP1]]
+ return SVE_ACLE_FUNC(svaddv,_s64,,)(pg, op);
+}
+
+// C-LABEL: @test_svaddv_u8(
+// CPP-LABEL: @_Z14test_svaddv_u8u10__SVBool_tu11__SVUint8_t(
+uint64_t test_svaddv_u8(svbool_t pg, svuint8_t op) MODE_ATTR
+{
+// CIR: %[[RES:.*]] = cir.call_llvm_intrinsic "aarch64.sve.uaddv" %{{.*}}, %{{.*}} :
+// CIR-SAME: -> !u64i
+
+// LLVM-SAME: <vscale x 16 x i1> [[PG:%.*]], <vscale x 16 x i8> [[OP:%.*]])
+// LLVM: [[TMP0:%.*]] = tail call i64 @llvm.aarch64.sve.uaddv.nxv16i8(<vscale x 16 x i1> [[PG]], <vscale x 16 x i8> [[OP]])
+// LLVM: ret i64 [[TMP0]]
+ return SVE_ACLE_FUNC(svaddv,_u8,,)(pg, op);
+}
+
+// C-LABEL: @test_svaddv_u16(
+// CPP-LABEL: @_Z15test_svaddv_u16u10__SVBool_tu12__SVUint16_t(
+uint64_t test_svaddv_u16(svbool_t pg, svuint16_t op) MODE_ATTR
+{
+// CIR: %[[CONVERT_PG:.*]] = cir.call_llvm_intrinsic "aarch64.sve.convert.from.svbool" %{{.*}} :
+// CIR-SAME: -> !cir.vector<[8] x !cir.int<u, 1>>
+// CIR: %[[RES:.*]] = cir.call_llvm_intrinsic "aarch64.sve.uaddv" %[[CONVERT_PG]], %{{.*}} :
+// CIR-SAME: -> !u64i
+
+// LLVM-SAME: <vscale x 16 x i1> [[PG:%.*]], <vscale x 8 x i16> [[OP:%.*]])
+// LLVM: [[TMP0:%.*]] = tail call <vscale x 8 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv8i1(<vscale x 16 x i1> [[PG]])
+// LLVM: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.uaddv.nxv8i16(<vscale x 8 x i1> [[TMP0]], <vscale x 8 x i16> [[OP]])
+// LLVM: ret i64 [[TMP1]]
+ return SVE_ACLE_FUNC(svaddv,_u16,,)(pg, op);
+}
+
+// C-LABEL: @test_svaddv_u32(
+// CPP-LABEL: @_Z15test_svaddv_u32u10__SVBool_tu12__SVUint32_t(
+uint64_t test_svaddv_u32(svbool_t pg, svuint32_t op) MODE_ATTR
+{
+// CIR: %[[CONVERT_PG:.*]] = cir.call_llvm_intrinsic "aarch64.sve.convert.from.svbool" %{{.*}} :
+// CIR-SAME: -> !cir.vector<[4] x !cir.int<u, 1>>
+// CIR: %[[RES:.*]] = cir.call_llvm_intrinsic "aarch64.sve.uaddv" %[[CONVERT_PG]], %{{.*}} :
+// CIR-SAME: -> !u64i
+
+// LLVM-SAME: <vscale x 16 x i1> [[PG:%.*]], <vscale x 4 x i32> [[OP:%.*]])
+// LLVM: [[TMP0:%.*]] = tail call <vscale x 4 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv4i1(<vscale x 16 x i1> [[PG]])
+// LLVM: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.uaddv.nxv4i32(<vscale x 4 x i1> [[TMP0]], <vscale x 4 x i32> [[OP]])
+// LLVM: ret i64 [[TMP1]]
+ return SVE_ACLE_FUNC(svaddv,_u32,,)(pg, op);
+}
+
+// C-LABEL: @test_svaddv_u64(
+// CPP-LABEL: @_Z15test_svaddv_u64u10__SVBool_tu12__SVUint64_t(
+uint64_t test_svaddv_u64(svbool_t pg, svuint64_t op) MODE_ATTR
+{
+// CIR: %[[CONVERT_PG:.*]] = cir.call_llvm_intrinsic "aarch64.sve.convert.from.svbool" %{{.*}} :
+// CIR-SAME: -> !cir.vector<[2] x !cir.int<u, 1>>
+// CIR: %[[RES:.*]] = cir.call_llvm_intrinsic "aarch64.sve.uaddv" %[[CONVERT_PG]], %{{.*}} :
+// CIR-SAME: -> !u64i
+
+// LLVM-SAME: <vscale x 16 x i1> [[PG:%.*]], <vscale x 2 x i64> [[OP:%.*]])
+// LLVM: [[TMP0:%.*]] = tail call <vscale x 2 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv2i1(<vscale x 16 x i1> [[PG]])
+// LLVM: [[TMP1:%.*]] = tail call i64 @llvm.aarch64.sve.uaddv.nxv2i64(<vscale x 2 x i1> [[TMP0]], <vscale x 2 x i64> [[OP]])
+// LLVM: ret i64 [[TMP1]]
+ return SVE_ACLE_FUNC(svaddv,_u64,,)(pg, op);
+}
+
+// C-LABEL: @test_svaddv_f16(
+// CPP-LABEL: @_Z15test_svaddv_f16u10__SVBool_tu13__SVFloat16_t(
+float16_t test_svaddv_f16(svbool_t pg, svfloat16_t op) MODE_ATTR
+{
+// CIR: %[[CONVERT_PG:.*]] = cir.call_llvm_intrinsic "aarch64.sve.convert.from.svbool" %{{.*}} :
+// CIR-SAME: -> !cir.vector<[8] x !cir.int<u, 1>>
+// CIR: %[[RES:.*]] = cir.call_llvm_intrinsic "aarch64.sve.faddv" %[[CONVERT_PG]], %{{.*}} :
+// CIR-SAME: -> !cir.f16
+
+// LLVM-SAME: <vscale x 16 x i1> [[PG:%.*]], <vscale x 8 x half> [[OP:%.*]])
+// LLVM: [[TMP0:%.*]] = tail call <vscale x 8 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv8i1(<vscale x 16 x i1> [[PG]])
+// LLVM: [[TMP1:%.*]] = tail call half @llvm.aarch64.sve.faddv.nxv8f16(<vscale x 8 x i1> [[TMP0]], <vscale x 8 x half> [[OP]])
+// LLVM: ret half [[TMP1]]
+ return SVE_ACLE_FUNC(svaddv,_f16,,)(pg, op);
+}
+
+// C-LABEL: @test_svaddv_f32(
+// CPP-LABEL: @_Z15test_svaddv_f32u10__SVBool_tu13__SVFloat32_t(
+float32_t test_svaddv_f32(svbool_t pg, svfloat32_t op) MODE_ATTR
+{
+// CIR: %[[CONVERT_PG:.*]] = cir.call_llvm_intrinsic "aarch64.sve.convert.from.svbool" %{{.*}} :
+// CIR-SAME: -> !cir.vector<[4] x !cir.int<u, 1>>
+// CIR: %[[RES:.*]] = cir.call_llvm_intrinsic "aarch64.sve.faddv" %[[CONVERT_PG]], %{{.*}} :
+// CIR-SAME: -> !cir.float
+
+// LLVM-SAME: <vscale x 16 x i1> [[PG:%.*]], <vscale x 4 x float> [[OP:%.*]])
+// LLVM: [[TMP0:%.*]] = tail call <vscale x 4 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv4i1(<vscale x 16 x i1> [[PG]])
+// LLVM: [[TMP1:%.*]] = tail call float @llvm.aarch64.sve.faddv.nxv4f32(<vscale x 4 x i1> [[TMP0]], <vscale x 4 x float> [[OP]])
+// LLVM: ret float [[TMP1]]
+ return SVE_ACLE_FUNC(svaddv,_f32,,)(pg, op);
+}
+
+// C-LABEL: @test_svaddv_f64(
+// CPP-LABEL: @_Z15test_svaddv_f64u10__SVBool_tu13__SVFloat64_t(
+float64_t test_svaddv_f64(svbool_t pg, svfloat64_t op) MODE_ATTR
+{
+// CIR: %[[CONVERT_PG:.*]] = cir.call_llvm_intrinsic "aarch64.sve.convert.from.svbool" %{{.*}} :
+// CIR-SAME: -> !cir.vector<[2] x !cir.int<u, 1>>
+// CIR: %[[RES:.*]] = cir.call_llvm_intrinsic "aarch64.sve.faddv" %[[CONVERT_PG]], %{{.*}} :
+// CIR-SAME: -> !cir.double
+
+// LLVM-SAME: <vscale x 16 x i1> [[PG:%.*]], <vscale x 2 x double> [[OP:%.*]])
+// LLVM: [[TMP0:%.*]] = tail call <vscale x 2 x i1> @llvm.aarch64.sve.convert.from.svbool.nxv2i1(<vscale x 16 x i1> [[PG]])
+// LLVM: [[TMP1:%.*]] = tail call double @llvm.aarch64.sve.faddv.nxv2f64(<vscale x 2 x i1> [[TMP0]], <vscale x 2 x double> [[OP]])
+// LLVM: ret double [[TMP1]]
+ return SVE_ACLE_FUNC(svaddv,_f64,,)(pg, op);
+}
>From d895e362f18f4ad647011ed22ef2d0cf37642a55 Mon Sep 17 00:00:00 2001
From: David Green <david.green at arm.com>
Date: Sun, 27 Sep 2026 13:02:24 +0100
Subject: [PATCH 51/53] [DAG] Scalarise trivial splat operations (#226257)
This fixes an issue reported on #224255, where a single active lane
masked store is converted to a v1f16 extract_subvector, which becomes a
v1f16 splat, which fails to scalarize. Add the necessary trivial
scalarisation of the splat by using the f16 input operand.
(I was originally going to fix this by generating a scalar f16 extract
directly (which would have the benefit of treating f16 as legal), but
that seems to cause a regression in 2 x86 tests).
---
llvm/lib/CodeGen/SelectionDAG/LegalizeTypes.h | 2 +-
llvm/lib/CodeGen/SelectionDAG/LegalizeVectorTypes.cpp | 9 ++++++---
llvm/test/CodeGen/AArch64/lowmaskedlanes.ll | 11 +++++++++++
3 files changed, 18 insertions(+), 4 deletions(-)
diff --git a/llvm/lib/CodeGen/SelectionDAG/LegalizeTypes.h b/llvm/lib/CodeGen/SelectionDAG/LegalizeTypes.h
index b638d00673490..0cf31089e08bd 100644
--- a/llvm/lib/CodeGen/SelectionDAG/LegalizeTypes.h
+++ b/llvm/lib/CodeGen/SelectionDAG/LegalizeTypes.h
@@ -837,7 +837,7 @@ class LLVM_LIBRARY_VISIBILITY DAGTypeLegalizer {
SDValue ScalarizeVecRes_ADDRSPACECAST(SDNode *N);
SDValue ScalarizeVecRes_BITCAST(SDNode *N);
- SDValue ScalarizeVecRes_BUILD_VECTOR(SDNode *N);
+ SDValue ScalarizeVecRes_BUILD_VECTOR_OR_SPLAT(SDNode *N);
SDValue ScalarizeVecRes_EXTRACT_SUBVECTOR(SDNode *N);
SDValue ScalarizeVecRes_FP_ROUND(SDNode *N);
SDValue ScalarizeVecRes_CONVERT_FROM_ARBITRARY_FP(SDNode *N);
diff --git a/llvm/lib/CodeGen/SelectionDAG/LegalizeVectorTypes.cpp b/llvm/lib/CodeGen/SelectionDAG/LegalizeVectorTypes.cpp
index 913aa4536bc81..ee1907b0c67bd 100644
--- a/llvm/lib/CodeGen/SelectionDAG/LegalizeVectorTypes.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/LegalizeVectorTypes.cpp
@@ -63,7 +63,10 @@ void DAGTypeLegalizer::ScalarizeVectorResult(SDNode *N, unsigned ResNo) {
break;
case ISD::MERGE_VALUES: R = ScalarizeVecRes_MERGE_VALUES(N, ResNo);break;
case ISD::BITCAST: R = ScalarizeVecRes_BITCAST(N); break;
- case ISD::BUILD_VECTOR: R = ScalarizeVecRes_BUILD_VECTOR(N); break;
+ case ISD::SPLAT_VECTOR:
+ case ISD::BUILD_VECTOR:
+ R = ScalarizeVecRes_BUILD_VECTOR_OR_SPLAT(N);
+ break;
case ISD::EXTRACT_SUBVECTOR: R = ScalarizeVecRes_EXTRACT_SUBVECTOR(N); break;
case ISD::FP_ROUND: R = ScalarizeVecRes_FP_ROUND(N); break;
case ISD::CONVERT_FROM_ARBITRARY_FP:
@@ -475,10 +478,10 @@ SDValue DAGTypeLegalizer::ScalarizeVecRes_BITCAST(SDNode *N) {
NewVT, Op);
}
-SDValue DAGTypeLegalizer::ScalarizeVecRes_BUILD_VECTOR(SDNode *N) {
+SDValue DAGTypeLegalizer::ScalarizeVecRes_BUILD_VECTOR_OR_SPLAT(SDNode *N) {
EVT EltVT = N->getValueType(0).getVectorElementType();
SDValue InOp = N->getOperand(0);
- // The BUILD_VECTOR operands may be of wider element types and
+ // The BUILD_VECTOR / SPLAT operands may be of wider element types and
// we may need to truncate them back to the requested return type.
if (EltVT.isInteger())
return DAG.getNode(ISD::TRUNCATE, SDLoc(N), EltVT, InOp);
diff --git a/llvm/test/CodeGen/AArch64/lowmaskedlanes.ll b/llvm/test/CodeGen/AArch64/lowmaskedlanes.ll
index 1a3f88ef11b33..e8de2bdcfb98d 100644
--- a/llvm/test/CodeGen/AArch64/lowmaskedlanes.ll
+++ b/llvm/test/CodeGen/AArch64/lowmaskedlanes.ll
@@ -245,3 +245,14 @@ define void @store_low4_nxv8i16(ptr %p, <vscale x 8 x i16> %a) {
tail call void @llvm.masked.store(<vscale x 8 x i16> %a, ptr align 1 %p, <vscale x 8 x i1> %m)
ret void
}
+
+define void @store_low1_nxv8f16_splat(ptr %0) {
+; CHECK-LABEL: store_low1_nxv8f16_splat:
+; CHECK: // %bb.0:
+; CHECK-NEXT: fmov h0, #1.00000000
+; CHECK-NEXT: str h0, [x0]
+; CHECK-NEXT: ret
+ %2 = tail call <vscale x 8 x i1> @llvm.get.active.lane.mask.nxv8i1.i32(i32 0, i32 1)
+ tail call void @llvm.masked.store.nxv8f16.p0(<vscale x 8 x half> splat (half 0xH3C00), ptr align 2 %0, <vscale x 8 x i1> %2)
+ ret void
+}
>From 007eaba80432d16a8b084d99038f4abb43b3cae4 Mon Sep 17 00:00:00 2001
From: Alexey Bataev <a.bataev at outlook.com>
Date: Sun, 27 Sep 2026 09:14:46 -0400
Subject: [PATCH 52/53] [SLP]Fix erasing gathered leftover narrowed reduction
leaves
A leftover narrowed reduction leaf, only gathered in the tree, loses all
users once the tree scalars are erased and is swept as a dead operand,
though the reduction epilogue still uses it. Keep such values alive.
Fixes https://github.com/llvm/llvm-project/pull/224919#issuecomment-5855368565
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/226782
---
.../Transforms/Vectorize/SLPVectorizer.cpp | 22 +++-
...rrowed-reduction-gathered-leftover-leaf.ll | 106 ++++++++++++++++++
2 files changed, 122 insertions(+), 6 deletions(-)
create mode 100644 llvm/test/Transforms/SLPVectorizer/X86/narrowed-reduction-gathered-leftover-leaf.ll
diff --git a/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp b/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
index f76a754418422..0f1fbea086abc 100644
--- a/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+++ b/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
@@ -3813,9 +3813,9 @@ class slpvectorizer::BoUpSLP {
/// uses.
SmallPtrSet<Value *, 4> ExternalUsesWithNonUsers;
- /// Replacements emitted for the external uses without users, consumed after
- /// the tree vectorization; must not be collected as dead operands of the
- /// erased scalars.
+ /// Externally used values and replacements emitted for the external uses
+ /// without users, consumed after the tree vectorization; must not be
+ /// collected as dead operands of the erased scalars.
SmallPtrSet<Value *, 4> ExternalUseReplacements;
/// Values used only by @llvm.assume calls.
@@ -26608,6 +26608,16 @@ BoUpSLP::vectorizeTree(const ExtraValueToDebugLocsMap &ExternallyUsedValues,
SI->setCondition(Constant::getNullValue(SI->getCondition()->getType()));
}
}
+ // Externally used values, which are not operands of the reduction ops (e.g.
+ // gathered narrowed reduction leaves), may lose all users once the tree
+ // scalars are erased; keep them alive for the reduction epilogue. Too many
+ // users - keep conservatively.
+ if (UserIgnoreList)
+ for (Instruction *I : make_isa_range<Instruction>(ExternallyUsedValues))
+ if (I->hasNUsesOrMore(UsesLimit) || none_of(I->users(), [&](User *U) {
+ return UserIgnoreList->contains(U);
+ }))
+ ExternalUseReplacements.insert(I);
// Retain to-be-deleted instructions for some debug-info bookkeeping and alias
// cache correctness.
// NOTE: removeInstructionAndOperands only marks the instruction for deletion
@@ -32688,9 +32698,9 @@ class HorizontalReduction {
RdxVal = It->second;
if (!Visited.insert(RdxVal).second)
continue;
- // Check if the scalar was vectorized as part of the vectorization
- // tree but not the top node.
- if (!VLScalars.contains(RdxVal) && V.isVectorized(RdxVal)) {
+ // Scalars not reduced by the top node may be used by the reduction
+ // epilogue, even if not vectorized as part of the tree.
+ if (!VLScalars.contains(RdxVal)) {
LocalExternallyUsedValues.insert(RdxVal);
continue;
}
diff --git a/llvm/test/Transforms/SLPVectorizer/X86/narrowed-reduction-gathered-leftover-leaf.ll b/llvm/test/Transforms/SLPVectorizer/X86/narrowed-reduction-gathered-leftover-leaf.ll
new file mode 100644
index 0000000000000..1aed39d118ac4
--- /dev/null
+++ b/llvm/test/Transforms/SLPVectorizer/X86/narrowed-reduction-gathered-leftover-leaf.ll
@@ -0,0 +1,106 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
+; RUN: opt -passes=slp-vectorizer -S -mtriple=x86_64-unknown-linux-gnu -mcpu=x86-64-v3 < %s | FileCheck %s
+
+define i32 @extracted_fields_leaf(ptr %p, i1 %c, i8 %x) {
+; CHECK-LABEL: define i32 @extracted_fields_leaf(
+; CHECK-SAME: ptr [[P:%.*]], i1 [[C:%.*]], i8 [[X:%.*]]) #[[ATTR0:[0-9]+]] {
+; CHECK-NEXT: [[ENTRY:.*]]:
+; CHECK-NEXT: br label %[[LOOP:.*]]
+; CHECK: [[LOOP]]:
+; CHECK-NEXT: [[TMP0:%.*]] = phi <4 x i8> [ zeroinitializer, %[[ENTRY]] ], [ [[TMP5:%.*]], %[[LOOP]] ]
+; CHECK-NEXT: [[LD:%.*]] = load i64, ptr [[P]], align 4
+; CHECK-NEXT: [[F0:%.*]] = trunc i64 [[LD]] to i8
+; CHECK-NEXT: [[TMP1:%.*]] = shufflevector <4 x i8> [[TMP0]], <4 x i8> <i8 poison, i8 poison, i8 1, i8 1>, <4 x i32> <i32 0, i32 poison, i32 6, i32 7>
+; CHECK-NEXT: [[TMP2:%.*]] = insertelement <4 x i8> [[TMP1]], i8 [[X]], i64 1
+; CHECK-NEXT: [[TMP3:%.*]] = bitcast i64 [[LD]] to <8 x i8>
+; CHECK-NEXT: [[TMP4:%.*]] = shufflevector <8 x i8> [[TMP3]], <8 x i8> poison, <4 x i32> <i32 0, i32 1, i32 0, i32 0>
+; CHECK-NEXT: [[TMP5]] = or <4 x i8> [[TMP2]], [[TMP4]]
+; CHECK-NEXT: br i1 [[C]], label %[[EXIT:.*]], label %[[LOOP]]
+; CHECK: [[EXIT]]:
+; CHECK-NEXT: [[TMP6:%.*]] = zext <4 x i8> [[TMP0]] to <4 x i32>
+; CHECK-NEXT: [[TMP7:%.*]] = shl <4 x i32> [[TMP6]], <i32 0, i32 1, i32 0, i32 0>
+; CHECK-NEXT: [[TMP8:%.*]] = call i32 @llvm.vector.reduce.or.v4i32(<4 x i32> [[TMP7]])
+; CHECK-NEXT: [[TMP9:%.*]] = zext i8 [[F0]] to i32
+; CHECK-NEXT: [[OP_RDX:%.*]] = or i32 [[TMP8]], [[TMP9]]
+; CHECK-NEXT: ret i32 [[OP_RDX]]
+;
+entry:
+ br label %loop
+
+loop:
+ %p0 = phi i8 [ 0, %entry ], [ %o0, %loop ]
+ %p1 = phi i8 [ 0, %entry ], [ %o1, %loop ]
+ %p2 = phi i8 [ 0, %entry ], [ %o2, %loop ]
+ %p3 = phi i8 [ 0, %entry ], [ %o3, %loop ]
+ %ld = load i64, ptr %p, align 4
+ %f0 = trunc i64 %ld to i8
+ %o0 = or i8 %p0, %f0
+ %sh = lshr i64 %ld, 8
+ %f1 = trunc i64 %sh to i8
+ %o1 = or i8 %x, %f1
+ %o2 = or i8 %f0, 1
+ %o3 = or i8 %f0, 1
+ br i1 %c, label %exit, label %loop
+
+exit:
+ %z1 = zext i8 %p1 to i32
+ %s1 = shl i32 %z1, 1
+ %z0 = zext i8 %o0 to i32
+ %r0 = or i32 %s1, %z0
+ %z2 = zext i8 %p2 to i32
+ %r1 = or i32 %r0, %z2
+ %z3 = zext i8 %p3 to i32
+ %r2 = or i32 %r1, %z3
+ ret i32 %r2
+}
+
+define i32 @extractelement_leaf(ptr %p, i1 %c, i8 %x) {
+; CHECK-LABEL: define i32 @extractelement_leaf(
+; CHECK-SAME: ptr [[P:%.*]], i1 [[C:%.*]], i8 [[X:%.*]]) #[[ATTR0]] {
+; CHECK-NEXT: [[ENTRY:.*]]:
+; CHECK-NEXT: br label %[[LOOP:.*]]
+; CHECK: [[LOOP]]:
+; CHECK-NEXT: [[TMP0:%.*]] = phi <4 x i8> [ zeroinitializer, %[[ENTRY]] ], [ [[TMP4:%.*]], %[[LOOP]] ]
+; CHECK-NEXT: [[LD:%.*]] = load <8 x i8>, ptr [[P]], align 4
+; CHECK-NEXT: [[F0:%.*]] = extractelement <8 x i8> [[LD]], i64 0
+; CHECK-NEXT: [[TMP1:%.*]] = shufflevector <4 x i8> [[TMP0]], <4 x i8> <i8 poison, i8 poison, i8 1, i8 1>, <4 x i32> <i32 0, i32 poison, i32 6, i32 7>
+; CHECK-NEXT: [[TMP2:%.*]] = insertelement <4 x i8> [[TMP1]], i8 [[X]], i64 1
+; CHECK-NEXT: [[TMP3:%.*]] = shufflevector <8 x i8> [[LD]], <8 x i8> poison, <4 x i32> <i32 0, i32 1, i32 0, i32 0>
+; CHECK-NEXT: [[TMP4]] = or <4 x i8> [[TMP2]], [[TMP3]]
+; CHECK-NEXT: br i1 [[C]], label %[[EXIT:.*]], label %[[LOOP]]
+; CHECK: [[EXIT]]:
+; CHECK-NEXT: [[TMP5:%.*]] = zext <4 x i8> [[TMP0]] to <4 x i32>
+; CHECK-NEXT: [[TMP6:%.*]] = shl <4 x i32> [[TMP5]], <i32 0, i32 1, i32 0, i32 0>
+; CHECK-NEXT: [[TMP7:%.*]] = call i32 @llvm.vector.reduce.or.v4i32(<4 x i32> [[TMP6]])
+; CHECK-NEXT: [[TMP8:%.*]] = zext i8 [[F0]] to i32
+; CHECK-NEXT: [[OP_RDX:%.*]] = or i32 [[TMP7]], [[TMP8]]
+; CHECK-NEXT: ret i32 [[OP_RDX]]
+;
+entry:
+ br label %loop
+
+loop:
+ %p0 = phi i8 [ 0, %entry ], [ %o0, %loop ]
+ %p1 = phi i8 [ 0, %entry ], [ %o1, %loop ]
+ %p2 = phi i8 [ 0, %entry ], [ %o2, %loop ]
+ %p3 = phi i8 [ 0, %entry ], [ %o3, %loop ]
+ %ld = load <8 x i8>, ptr %p, align 4
+ %f0 = extractelement <8 x i8> %ld, i64 0
+ %o0 = or i8 %p0, %f0
+ %f1 = extractelement <8 x i8> %ld, i64 1
+ %o1 = or i8 %x, %f1
+ %o2 = or i8 %f0, 1
+ %o3 = or i8 %f0, 1
+ br i1 %c, label %exit, label %loop
+
+exit:
+ %z1 = zext i8 %p1 to i32
+ %s1 = shl i32 %z1, 1
+ %z0 = zext i8 %o0 to i32
+ %r0 = or i32 %s1, %z0
+ %z2 = zext i8 %p2 to i32
+ %r1 = or i32 %r0, %z2
+ %z3 = zext i8 %p3 to i32
+ %r2 = or i32 %r1, %z3
+ ret i32 %r2
+}
>From e634a3f203d2973580e134262dd6997291b31c0d Mon Sep 17 00:00:00 2001
From: "Agaev G." <cpp.thread at gmail.com>
Date: Sun, 27 Sep 2026 17:14:55 +0400
Subject: [PATCH 53/53] [clang-tidy] Fix test and add a release note
---
clang-tools-extra/docs/ReleaseNotes.md | 4 ++++
.../checkers/performance/noexcept-move-constructor.cpp | 2 +-
2 files changed, 5 insertions(+), 1 deletion(-)
diff --git a/clang-tools-extra/docs/ReleaseNotes.md b/clang-tools-extra/docs/ReleaseNotes.md
index 3d9e34cc4d44e..033733762b98d 100644
--- a/clang-tools-extra/docs/ReleaseNotes.md
+++ b/clang-tools-extra/docs/ReleaseNotes.md
@@ -333,6 +333,10 @@ infrastructure are described first, followed by tool-specific sections.
<clang-tidy/checks/readability/use-std-min-max>` check by fixing spurious
trailing semicolons and lost comments when the `if` body has no braces.
+- Improved {doc}`noexcept-move-constructors
+ <clang-tidy/checks/performance/noexcept-move-constructor>` check by fixing
+ false positives for implicitly declared noexcept(false).
+
#### Removed checks
- Removed the deprecated `zircon-temporary-objects` check. Users should migrate to
diff --git a/clang-tools-extra/test/clang-tidy/checkers/performance/noexcept-move-constructor.cpp b/clang-tools-extra/test/clang-tidy/checkers/performance/noexcept-move-constructor.cpp
index 42a1287163d92..1b705aada6c2b 100644
--- a/clang-tools-extra/test/clang-tidy/checkers/performance/noexcept-move-constructor.cpp
+++ b/clang-tools-extra/test/clang-tidy/checkers/performance/noexcept-move-constructor.cpp
@@ -199,7 +199,7 @@ struct P {
void p() {
P P1{};
P P2{std::move(P1)};
- P P3 = std::move(P2);
+ P1 = std::move(P2);
}
class OK {};
More information about the flang-commits
mailing list