Repository navigation
SIGSEGV using journal.Reader, thread safety issue? #143
Description
Activity
- added 2 commits that reference this issue
on Mar 30, 2025 As the libsystemd' documentation says:
All functions listed here are thread-agnostic and only a single specific thread may operate on a given object during its entire lifetime. It's safe to allocate multiple independent objects and use each from a specific thread in parallel. However, it's not safe to allocate such an object in one thread, and operate or free it from any other, even if locking is used to ensure these threads don't operate on it at the very same time.
Basically this mean the pythonic
systemd.journal.Readeris not thread-safe currently and normally,Reader.close()or whatever other functions shall be prohibited across the threads.
Completely, regardless the GIL and co.- added a commit that references this issue
on Apr 1, 2025 You are right that the the C code is "thread agnostic", i.e. only one thread can ever access those objects at a single time. This means that those
Py_BEGIN_ALLOW_THREADS/Py_END_ALLOW_THREADSare generally broken. You observed a segfault if.close()was called at the wrong time, but in general, the problem is much wider. If the journal object is passed to multiple threads, and they make any kind of calls on the the object, that will create concurrent calls into the journal C API from multiple threads. So just adding reference counting to prevent the close from happening is not enough.One broad possibility would be to remove all those
Py_BEGIN_ALLOW_THREADS/Py_END_ALLOW_THREADS. Then the GIL will protect us, because only one Python thread will ever be active, and the journal object is accessible only through the Python object.The other option would be put all the access to journal object behind a separate mutex. I think that for Python 3.13+ we'll want to do something like this anyway, because of free-threading mode. With ff, different Python threads are allowed to access the same Python objects at the same time. We could either say that the objects provided this module cannot be shared between threads, or add a new mutex around each journal object. I think this is a nicer option, but I only started reading about the details of ff mode, so I'm not sure if I have the full picutre.
Reacted by Jörg BehrmannYou observed a segfault if .close() was called at the wrong time, but in general, the problem is much wider.
But only close() is dangerous in the way that it'd even segfault (because it destroying the handle, which can be accessed by another thread at the same time). The issues with interop cross-thread access (without to consider close) may surely cause some undefined behavior, but they are at least "safe" against segfault and doesn't cause a crash.
Normally a documentation entry would be enough, and actually I don't think someone really want to access journal multi-threaded.
Close() is rather an exception and used from another thread just in order to signal "stop" of monitoring if another thread is in waiting state.Then the GIL will protect us, because only one Python thread will ever be active
And globally lock for several seconds, if some waiting operation is involved? I don't think it is alright.
The other option would be put all the access to journal object behind a separate mutex.
As I wrote above it is not necessary in my opinion. Let alone I don't think it'd be good idea generally.
I think fixing close() is important, another stuff seems to be matter of UB either, at least as long as systemd API doesn't get completely "multi-threaded".
However I guess the set of functions called thread agnostic mostly refers to issue with the close().
For anything else I'm sure the entry in the documentation would be enough.- added 2 commits that reference this issue
on Oct 3, 2025 - added a commit that references this issue
on Oct 4, 2025 - added 4 commits that reference this issue
on Oct 10, 2025
Env:
Issue:
I got sporadically a segfault using journal.Reader from python-systemd module.
Thereby it seemed to happen by stop of thread, monitoring journal with Reader.
I could reproduce it also under dgb, here is the callstack:
Details
It can be either some missing lock or atomic handling somewhere in journal.Reader, because if I rewrote the code to avoid too earlier clean-up, closing the journal with
Reader.close()before join of thread (now it ensures that the thread working with Reader is really stopped, before clean-up happens), it doesn't happen anymore.So my assumption is if TH2 is somewhere within
reader.wait()and TH1 would invokereader.close(), it could segfault.Sure, it can be avoided on the user side, just SF is SF and shall not happen normally.
Naive attempts to reproduce it with simple wait/close in different threads did not work, but I could find a bit complex scenario (with a busy-wait & lock using to synchronize the timing a bit), where it is pretty reproducible, almost every time not later than on 10 iteration: