[llvm] [MC] Skip AsmToken::Comment at statement start in AsmParser (PR #218456)
Matt Turner via llvm-commits
llvm-commits at lists.llvm.org
Mon Aug 24 14:46:14 PDT 2026
https://github.com/mattst88 updated https://github.com/llvm/llvm-project/pull/218456
>From 1d3bf6e83b0f60b7ddfe54d3a7845b984d8415a7 Mon Sep 17 00:00:00 2001
From: Matt Turner <mattst88 at gmail.com>
Date: Thu, 9 Jul 2026 22:00:47 -0400
Subject: [PATCH] [MC] Skip AsmToken::Comment at statement start in AsmParser
eatToEndOfStatement() advances the lexer with Lexer.Lex() rather than
AsmParser::Lex(), and only the latter filters comment tokens. Its trailing
"eat EOL" step can therefore leave a block comment as the current token, which
parseStatement() then takes for the start of a statement and rejects:
.extern foo
/* comment */
nop
error: unexpected token at start of statement
GNU as accepts that input. Nor is this confined to error recovery -- .extern
above is an ordinary non-error path. On an error path the same thing adds a
second, spurious diagnostic after the real one.
Skip comment tokens alongside spaces in the start-of-statement loop. Hand
each one to the streamer on the way past, so -preserve-comments still emits
it, as it does for a comment the AsmParser::Lex() path sees.
Assisted-by: Claude Code
---
llvm/lib/MC/MCParser/AsmParser.cpp | 10 +++++--
.../MC/AsmParser/block-comment-after-error.s | 17 +++++++++++
.../block-comment-at-statement-start.s | 30 +++++++++++++++++++
.../MC/AsmParser/block-comment-preserve.s | 23 ++++++++++++++
4 files changed, 78 insertions(+), 2 deletions(-)
create mode 100644 llvm/test/MC/AsmParser/block-comment-after-error.s
create mode 100644 llvm/test/MC/AsmParser/block-comment-at-statement-start.s
create mode 100644 llvm/test/MC/AsmParser/block-comment-preserve.s
diff --git a/llvm/lib/MC/MCParser/AsmParser.cpp b/llvm/lib/MC/MCParser/AsmParser.cpp
index c4c5b1da9aa91..c6ab206141dee 100644
--- a/llvm/lib/MC/MCParser/AsmParser.cpp
+++ b/llvm/lib/MC/MCParser/AsmParser.cpp
@@ -1735,9 +1735,15 @@ bool AsmParser::parseBinOpRHS(unsigned Precedence, const MCExpr *&Res,
bool AsmParser::parseStatement(ParseStatementInfo &Info,
MCAsmParserSemaCallback *SI) {
assert(!hasPendingError() && "parseStatement started with pending error");
- // Eat initial spaces and comments
- while (Lexer.is(AsmToken::Space))
+ // Eat initial spaces and comments. eatToEndOfStatement() lexes raw, so a
+ // block comment can still be the current token here; taken as the start of a
+ // statement it rejects input GNU as accepts. Lex() drops the ones that
+ // follow, but it cannot see the current token, so preserve that one here.
+ while (Lexer.is(AsmToken::Space) || Lexer.is(AsmToken::Comment)) {
+ if (Lexer.is(AsmToken::Comment) && MAI.preserveAsmComments())
+ Out.addExplicitComment(Twine(Lexer.getTok().getString()));
Lex();
+ }
if (Lexer.is(AsmToken::EndOfStatement)) {
// if this is a line comment we can drop it safely
if (getTok().getString().empty() || getTok().getString().front() == '\r' ||
diff --git a/llvm/test/MC/AsmParser/block-comment-after-error.s b/llvm/test/MC/AsmParser/block-comment-after-error.s
new file mode 100644
index 0000000000000..02f6ea0a18a3f
--- /dev/null
+++ b/llvm/test/MC/AsmParser/block-comment-after-error.s
@@ -0,0 +1,17 @@
+# RUN: not llvm-mc -triple=x86_64-unknown-linux-gnu %s -o /dev/null 2>&1 \
+# RUN: | FileCheck %s
+
+## The error path into eatToEndOfStatement(). Without the block comment skip
+## in parseStatement() the comment-only line below is taken for the start of a
+## statement and draws a second, spurious error. See
+## block-comment-at-statement-start.s for the error-free cases.
+##
+## The directives sit above the input on purpose: a # line comment between the
+## two lines lexes as an EndOfStatement and resets the lexer, which hides the
+## very thing this is testing.
+
+# CHECK: [[#@LINE+2]]:{{[0-9]+}}: error: unexpected token
+# CHECK-NOT: error: unexpected token at start of statement
+ .byte 1 2 /* trailing */
+/* a line that is nothing but a comment */
+ nop
diff --git a/llvm/test/MC/AsmParser/block-comment-at-statement-start.s b/llvm/test/MC/AsmParser/block-comment-at-statement-start.s
new file mode 100644
index 0000000000000..6d1f603e2011c
--- /dev/null
+++ b/llvm/test/MC/AsmParser/block-comment-at-statement-start.s
@@ -0,0 +1,30 @@
+# RUN: llvm-mc -triple=x86_64-unknown-linux-gnu %s -o /dev/null 2>&1 \
+# RUN: | FileCheck %s --allow-empty --implicit-check-not={{.}}
+
+## eatToEndOfStatement() lexes raw, so it can leave a block comment as the
+## current token. parseStatement() has to skip it, or each shape below is
+## rejected with "unexpected token at start of statement" -- input GNU as
+## accepts. .extern reaches eatToEndOfStatement() with no error involved, so
+## none of this depends on error recovery.
+
+ .extern a
+/* comment-only line */
+ nop
+
+ .extern b
+/* spanning
+ several
+ lines */
+ nop
+
+ .extern c
+/* two */ /* on one line */
+ nop
+
+ .extern d
+/* followed by an instruction */ nop
+
+## A block comment as the last thing in the file exercises the Lex()-at-EOF
+## path in the skip loop.
+ .extern e
+/* at end of file */
diff --git a/llvm/test/MC/AsmParser/block-comment-preserve.s b/llvm/test/MC/AsmParser/block-comment-preserve.s
new file mode 100644
index 0000000000000..37eb623ef76b3
--- /dev/null
+++ b/llvm/test/MC/AsmParser/block-comment-preserve.s
@@ -0,0 +1,23 @@
+# RUN: llvm-mc -triple=x86_64-unknown-linux-gnu -preserve-comments %s \
+# RUN: | FileCheck %s --strict-whitespace
+
+## parseStatement() skips a block comment left as the current token by
+## eatToEndOfStatement(). It has to hand that comment to the streamer on the
+## way past, or -preserve-comments drops it -- unlike a block comment the
+## ordinary Lex() path sees. See block-comment-at-statement-start.s for why
+## the comment lands there in the first place.
+##
+## The {{^}} anchors keep these directives from matching their own echo, which
+## -preserve-comments emits as well.
+
+ .extern a
+/* preserved */
+ nop
+# CHECK: {{^}} # preserved
+
+ .extern b
+/* spanning
+ two lines */
+ nop
+# CHECK: {{^}} # spanning
+# CHECK: {{^}} # two lines
More information about the llvm-commits
mailing list