I have a function that I want to put into a static library:
run_me.h:
void run_me();
run_me.cpp:
namespace {
std::mutex python_mutex;
struct PythonScopedLock {
PythonScopedLock() : _lock(python_mutex) {
PyEval_AcquireLock();
}
~PythonScopedLock() {
PyEval_ReleaseLock();
}
private:
std::lock_guard<decltype(python_mutex)> _lock;
};
struct PythonInterpreter {
PythonInterpreter() {
Py_InitializeEx(0);
PyEval_InitThreads();
PyEval_ReleaseLock();
}
~PythonInterpreter() {
PyEval_AcquireLock();
Py_Finalize();
}
};
const PythonInterpreter python_interpreter;
} //namespace
void run_me()
{
PythonScopedLock lock;
// import python module and call some python code
}
Next, in the following executable, I get a crash almost instantly on PyObject_Malloc.
#include <tbb/tbb.h>
#include "run_me.h"
int main(){
using namespace tbb;
parallel_for(blocked_range<size_t>(0, 1234, 1), [&](const blocked_range<size_t>& r) {
for (auto it = r.begin(); it != r.end(); ++it) {
run_me();
}
});
}
I checked the threads with gdb and they are nicely waiting for lock to be unlocked, but one of them crashes:
1 Thread 0x7ffff7fdb900 (LWP 21474) "tests" 0x00007ffff69374ed in __lll_lock_wait () from /lib64/libpthread.so.0
2 Thread 0x7fffef349700 (LWP 21483) "tests" 0x00007ffff69374ed in __lll_lock_wait () from /lib64/libpthread.so.0
3 Thread 0x7fffeef48700 (LWP 21484) "tests" 0x00007ffff69374ed in __lll_lock_wait () from /lib64/libpthread.so.0
* 4 Thread 0x7fffeeb47700 (LWP 21485) "tests" PyObject_Malloc (nbytes=41) at Objects/obmalloc.c:831
5 Thread 0x7fffee746700 (LWP 21486) "tests" 0x00007ffff69374ed in __lll_lock_wait () from /lib64/libpthread.so.0
6 Thread 0x7fffedf44700 (LWP 21488) "tests" 0x00007ffff69374ed in __lll_lock_wait () from /lib64/libpthread.so.0
7 Thread 0x7fffee345700 (LWP 21487) "tests" 0x00007ffff69374ed in __lll_lock_wait () from /lib64/libpthread.so.0
Situation changes when I move out PythonScopedLock from the body of run_me and put it into lambda body. This runs just fine:
#include <tbb/tbb.h>
#include "run_me.h"
namespace {
std::mutex python_mutex;
struct PythonScopedLock {
PythonScopedLock() : _lock(python_mutex) {
PyEval_AcquireLock();
}
~PythonScopedLock() {
PyEval_ReleaseLock();
}
private:
std::lock_guard<decltype(python_mutex)> _lock;
};
} // namespace
int main(){
using namespace tbb;
parallel_for(blocked_range<size_t>(0, 1234, 1), [&](const blocked_range<size_t>& r) {
for (auto it = r.begin(); it != r.end(); ++it) {
PythonScopedLock lock;
run_me();
}
});
}
I have few questions:
What is the difference, in this context, between heaving a lock inside of the lambda or a body of the function? To me it's the same, I just want to hide that lock inside of body of the function.What are possible causes of the crash?- What are other problems besides static initialisation fiasco, that I can encounter during interpreter initialisation through a variable in anon namespace through
const PythonInterpreter python_interpreter;?
EDIT
After a bit of debugging I noticed when lock is inside of run_me definition, in the static library, PyGILState_GetThisThreadState returns NULL. When lock is outside, in the main thread then it returns same pointer as _PyThreadState_Current. That kind of answers question number 1. and 2.
Question 3. still remains open.