[Lldb-commits] [lldb] [lldb] Reject NULL key in PythonDictionary::GetItem (PR #205753)
Yao Qi via lldb-commits
lldb-commits at lists.llvm.org
Mon Jun 29 04:10:06 PDT 2026
https://github.com/qiyao updated https://github.com/llvm/llvm-project/pull/205753
>From 26be35369124d816832293606545f80d735c49d4 Mon Sep 17 00:00:00 2001
From: Yao Qi <yao_qi at apple.com>
Date: Wed, 24 Jun 2026 21:23:21 +0100
Subject: [PATCH 1/2] [lldb] Reject NULL key in PythonDictionary::GetItem
A `breakpoint set -P <class>` with an invalid-UTF-8 class name crashes
lldb with `EXC_BAD_ACCESS`. The class name is passed through
unmodified, so the byte sequence `\xd0p` (an incomplete two-byte UTF-8
sequence) reaches the Python layer:
```
./bin/lldb -b -o $'breakpoint set -P \xd0p -f main.cpp' ./bin/lldb
...
PLEASE submit a bug report to https://github.com/llvm/llvm-project/issues/
Stack dump:
#4 PythonDictionary::GetItem(PythonObject const&) const
#6 PythonDictionary::GetItemForKey(PythonObject const&) const
#7 PythonObject::ResolveNameWithDictionary(StringRef, PythonDictionary const&)
#9 ScriptedPythonInterface::CreatePluginObject<...>(...)
```
`ResolveNameWithDictionary` builds a `PythonString` from the class name
and looks it up in the module dictionary. When the name is not valid
UTF-8, `PyUnicode_FromStringAndSize` returns `NULL`, so the
`PythonString` is left with `m_py_obj == NULL`. `GetItemForKey` forwards
that object to `GetItem`, which passed the `NULL` `PyObject *` straight
to `PyDict_GetItemWithError`.
Debugging the crashing lldb under lldb shows the `NULL` key arriving in
`GetItem`:
```
(lldb) break set -n lldb_private::python::PythonDictionary::GetItem
(lldb) run --no-use-colors -b -s /tmp/bp_cmds.txt
* stop reason = breakpoint 1.1
(lldb) up
(lldb) frame variable key.m_py_obj
(PyObject *) key.m_py_obj = nullptr
```
Continuing lands on the dereference of that `NULL` inside CPython:
```
(lldb) continue
* stop reason = EXC_BAD_ACCESS (code=1, address=0x8)
frame #0: PyDict_GetItemWithError + 40
-> ldr x8, [x1, #0x8]
```
`x1` holds the key argument, which is `NULL`, so reading the type field
at offset `0x8` faults at address `0x8`.
`GetItem` already guards against an invalid dictionary (`!IsValid()`);
extend that guard to reject an invalid key as well and return
`nullDeref()`. `GetItemForKey` consumes the error and returns an empty
`PythonObject`, so existing callers such as `ResolveNameWithDictionary`
keep working and the breakpoint command now fails cleanly instead of
crashing.
This was found by `lldb-commandinterpreter-fuzzer`.
Adds `PythonDataObjectsTest.TestPythonDictionaryGetItemWithInvalidKey`,
which feeds an invalid-UTF-8 key to `GetItem`. Without the fix the test
dereferences the `NULL` key and crashes (`SIGSEGV`).
---
.../Python/PythonDataObjects.cpp | 2 +-
.../Python/PythonDataObjectsTests.cpp | 24 +++++++++++++++++++
2 files changed, 25 insertions(+), 1 deletion(-)
diff --git a/lldb/source/Plugins/ScriptInterpreter/Python/PythonDataObjects.cpp b/lldb/source/Plugins/ScriptInterpreter/Python/PythonDataObjects.cpp
index 800f991680c85..fba98abf9e83b 100644
--- a/lldb/source/Plugins/ScriptInterpreter/Python/PythonDataObjects.cpp
+++ b/lldb/source/Plugins/ScriptInterpreter/Python/PythonDataObjects.cpp
@@ -697,7 +697,7 @@ PythonObject PythonDictionary::GetItemForKey(const PythonObject &key) const {
Expected<PythonObject>
PythonDictionary::GetItem(const PythonObject &key) const {
- if (!IsValid())
+ if (!IsValid() || !key.IsValid())
return nullDeref();
PyObject *o = PyDict_GetItemWithError(m_py_obj, key.get());
if (PyErr_Occurred())
diff --git a/lldb/unittests/ScriptInterpreter/Python/PythonDataObjectsTests.cpp b/lldb/unittests/ScriptInterpreter/Python/PythonDataObjectsTests.cpp
index 0d4b04b7a1284..24ed721049e67 100644
--- a/lldb/unittests/ScriptInterpreter/Python/PythonDataObjectsTests.cpp
+++ b/lldb/unittests/ScriptInterpreter/Python/PythonDataObjectsTests.cpp
@@ -490,6 +490,30 @@ TEST_F(PythonDataObjectsTest, TestPythonDictionaryValueEquality) {
EXPECT_EQ(value_1, chk_str.GetString());
}
+TEST_F(PythonDataObjectsTest, TestPythonDictionaryGetItemWithInvalidKey) {
+ // A PythonString constructed from invalid UTF-8 silently fails (the
+ // PyUnicode_FromStringAndSize call returns NULL) and leaves m_py_obj == NULL.
+ // PythonDictionary::GetItem(const PythonObject&) must not pass that NULL key
+ // through to PyDict_GetItemWithError, which would then dereference it and
+ // crash.
+ llvm::StringRef invalid_utf8("\xd0p", 2);
+ PythonString invalid_key(invalid_utf8);
+ EXPECT_FALSE(invalid_key.IsValid());
+
+ PythonDictionary dict(PyInitialValue::Empty);
+ dict.SetItemForKey(PythonString("real_key"), PythonInteger(42));
+
+ // GetItem must return an error (not crash) when given an invalid key.
+ llvm::Expected<PythonObject> item = dict.GetItem(invalid_key);
+ EXPECT_FALSE(static_cast<bool>(item));
+ llvm::consumeError(item.takeError());
+
+ // GetItemForKey wraps GetItem, consumes any error, and returns an empty
+ // PythonObject. It must not crash either.
+ PythonObject result = dict.GetItemForKey(invalid_key);
+ EXPECT_FALSE(result.IsAllocated());
+}
+
TEST_F(PythonDataObjectsTest, TestPythonDictionaryManipulation) {
// Test that manipulation of a dictionary behaves correctly when wrapped
// by a PythonDictionary.
>From 25ecf86e9c8b3f76727cac021de92ab0ff426b2d Mon Sep 17 00:00:00 2001
From: Yao Qi <yao_qi at apple.com>
Date: Mon, 29 Jun 2026 11:21:43 +0100
Subject: [PATCH 2/2] [lldb] Add API test for invalid-UTF-8 scripted process
class name
Follow-up to the `PythonDictionary::GetItem` NULL-key fix, adding the
requested API integration test in `lldb/test/API` for the invalid-UTF-8
class name crash.
Why the test is written this way
--------------------------------
The crash is reached only when a class name that is not valid UTF-8
arrives at `PythonObject::ResolveNameWithDictionary`, which builds a
`PythonString` from it. `PyUnicode_FromStringAndSize` then returns
`NULL`,mthe `PythonString` is left invalid, and (before the fix)
`GetItem` dereferenced that `NULL` key.
To exercise this from an API test the malformed bytes must be delivered
to lldb verbatim. Every existing entry point that takes a scripted
class name, including the breakpoint resolver path
(`SBTarget::BreakpointCreateFromScript` / "breakpoint set -P") and
`SBLaunchInfo::SetScriptedProcessClassName(const char *)`, accepts the
name as a NUL-terminated C string. The Python bindings encode the
incoming Python str to UTF-8 at the SWIG boundary, so an invalid
sequence fails to encode and never reaches lldb. A normal breakpoint
(or any other) API test therefore cannot reproduce the crash: the bad
bytes are rejected by the bindings before they get anywhere near
GetItem. The only way they previously reached the C++ layer was through
the command line (e.g. "breakpoint set -P '\xd0p'"), which is why a
direct API test was not previously possible without spawning lldb as a
subprocess.
To make the path testable through the SB API, add an overload
```
SBLaunchInfo::SetScriptedProcessClassName(const char *src, size_t len)
```
that accepts a Python bytes object and preserves the exact bytes
(including invalid UTF-8 and embedded NULs). `ScriptedMetadata` already
stores the class name as a byte-preserving std::string, so no other
layer needs to change.
SWIG note: the existing (const char *src, size_t src_len) in-typemap
already accepts bytes, but SWIG selects an overload using the typecheck
typemap, not the in-typemap. Without a matching typecheck, the new
overload rejects a bytes argument. Add a typecheck for
(const char *src, size_t src_len) that accepts str/bytes/bytearray so
the overload dispatches correctly. The only overloaded user of that
signature is the new method; non-overloaded users (e.g.
SBProcess::PutSTDIN) are unaffected because they generate no dispatcher.
The test creates a scripted process whose class name is the invalid
sequence b"\xd0p" and asserts the launch fails cleanly. Against the
unfixed code the launch crashes in `PythonDictionary::GetItem;` with the
fix it returns a clean error.
---
lldb/bindings/python/python-typemaps.swig | 17 +++++++++
lldb/include/lldb/API/SBLaunchInfo.h | 9 +++++
lldb/source/API/SBLaunchInfo.cpp | 11 ++++++
.../TestScriptedProcessInvalidClassName.py | 37 +++++++++++++++++++
4 files changed, 74 insertions(+)
create mode 100644 lldb/test/API/functionalities/scripted_process/TestScriptedProcessInvalidClassName.py
diff --git a/lldb/bindings/python/python-typemaps.swig b/lldb/bindings/python/python-typemaps.swig
index 072e688c4bde1..19ab7a1822895 100644
--- a/lldb/bindings/python/python-typemaps.swig
+++ b/lldb/bindings/python/python-typemaps.swig
@@ -272,6 +272,23 @@ AND call SWIG_fail at the same time, because it will result in a double free.
SWIG_fail;
}
}
+
+// Overload dispatch typecheck for (const char *src, size_t src_len).
+//
+// The %typemap(in) above already accepts str/bytes/bytearray, but SWIG's
+// generated overload dispatcher selects a candidate using the typecheck
+// typemap, not the in typemap. Without this, an overloaded method taking
+// (const char *src, size_t src_len) (e.g. SBLaunchInfo::SetScriptedProcess
+// ClassName) would reject a Python bytes argument because the default char*
+// typecheck only matches str. Accept the same input kinds as the in typemap.
+%typemap(typecheck, precedence=SWIG_TYPECHECK_STRING) (const char *src,
+ size_t src_len) {
+ $1 = (PyUnicode_Check($input) || PyBytes_Check($input) ||
+ PyByteArray_Check($input))
+ ? 1
+ : 0;
+}
+
// For SBProcess::WriteMemory, SBTarget::GetInstructions and SBDebugger::DispatchInput.
%typemap(in) (const void *buf, size_t size),
(const void *data, size_t data_len),
diff --git a/lldb/include/lldb/API/SBLaunchInfo.h b/lldb/include/lldb/API/SBLaunchInfo.h
index 06e72efc30f9f..355cb119a6951 100644
--- a/lldb/include/lldb/API/SBLaunchInfo.h
+++ b/lldb/include/lldb/API/SBLaunchInfo.h
@@ -202,6 +202,15 @@ class LLDB_API SBLaunchInfo {
void SetScriptedProcessClassName(const char *class_name);
+ /// Set the scripted process class name from a buffer of \a src_len bytes.
+ ///
+ /// Unlike the NUL-terminated overload, this preserves the exact bytes of
+ /// \a src, including embedded NULs and byte sequences that are not valid
+ /// UTF-8. It exists so that tests can exercise the class-name resolution
+ /// path with arbitrary bytes; the NUL-terminated overload cannot carry such
+ /// input across the scripting-language bindings.
+ void SetScriptedProcessClassName(const char *src, size_t src_len);
+
lldb::SBStructuredData GetScriptedProcessDictionary() const;
void SetScriptedProcessDictionary(lldb::SBStructuredData dict);
diff --git a/lldb/source/API/SBLaunchInfo.cpp b/lldb/source/API/SBLaunchInfo.cpp
index 572590f4cae76..46858786054fd 100644
--- a/lldb/source/API/SBLaunchInfo.cpp
+++ b/lldb/source/API/SBLaunchInfo.cpp
@@ -355,6 +355,17 @@ void SBLaunchInfo::SetScriptedProcessClassName(const char *class_name) {
m_opaque_sp->SetScriptedMetadata(metadata_sp);
}
+void SBLaunchInfo::SetScriptedProcessClassName(const char *src,
+ size_t src_len) {
+ LLDB_INSTRUMENT_VA(this, src, src_len);
+ ScriptedMetadataSP metadata_sp = m_opaque_sp->GetScriptedMetadata();
+ StructuredData::DictionarySP dict_sp =
+ metadata_sp ? metadata_sp->GetArgsSP() : nullptr;
+ metadata_sp = std::make_shared<ScriptedMetadata>(
+ llvm::StringRef(src, src_len), dict_sp);
+ m_opaque_sp->SetScriptedMetadata(metadata_sp);
+}
+
lldb::SBStructuredData SBLaunchInfo::GetScriptedProcessDictionary() const {
LLDB_INSTRUMENT_VA(this);
diff --git a/lldb/test/API/functionalities/scripted_process/TestScriptedProcessInvalidClassName.py b/lldb/test/API/functionalities/scripted_process/TestScriptedProcessInvalidClassName.py
new file mode 100644
index 0000000000000..90cb198400643
--- /dev/null
+++ b/lldb/test/API/functionalities/scripted_process/TestScriptedProcessInvalidClassName.py
@@ -0,0 +1,37 @@
+"""
+Test that creating a scripted process with an invalid-UTF-8 class name fails
+cleanly instead of crashing lldb.
+"""
+
+import lldb
+from lldbsuite.test.decorators import *
+from lldbsuite.test.lldbtest import *
+from lldbsuite.test import lldbtest_config
+
+
+class ScriptedProcessInvalidClassNameTestCase(TestBase):
+ NO_DEBUG_INFO_TESTCASE = True
+
+ def test_invalid_utf8_class_name(self):
+ """A scripted process class name that is not valid UTF-8 must produce a
+ clean error, not crash lldb.
+
+ ``b"\xd0p"`` is an incomplete two-byte UTF-8 sequence. When used as a
+ scripted process class name it reaches PyUnicode_FromStringAndSize,
+ which returns NULL; PythonDictionary::GetItem then dereferenced that
+ NULL key and crashed with EXC_BAD_ACCESS.
+ """
+ target = self.dbg.CreateTarget(lldbtest_config.lldbExec)
+ self.assertTrue(target.IsValid(), VALID_TARGET)
+
+ launch_info = lldb.SBLaunchInfo(None)
+ launch_info.SetProcessPluginName("ScriptedProcess")
+ launch_info.SetScriptedProcessClassName(b"\xd0p")
+
+ error = lldb.SBError()
+ process = target.Launch(launch_info, error)
+
+ # The launch must fail cleanly rather than crash. Reaching this
+ # assertion at all means lldb survived the invalid class name.
+ self.assertTrue(error.Fail(), "launch should fail, not crash")
+ self.assertFalse(process.IsValid())
More information about the lldb-commits
mailing list