I built a simulator in C++ with a pybind11 interface to run deep learning in Python using PyTorch. At each time step, I draw certain things from the simulator's scene using the SFML library (wrapper around openGL). I draw that on a texture, then get the pixels from that texture as follows:
glBindTexture(GL_TEXTURE_2D, imageTexture.getTexture().getNativeHandle());
glGetTexImage(GL_TEXTURE_2D, 0, GL_RGBA, GL_UNSIGNED_BYTE, img.data());
I then move the pixels vector img from C++ to Python using the pybind11 interface. Problem is that this GPU-to-CPU operation is very slow. Since the vector in Python is then transferred back to the GPU for fast deep learning (CNN) operations, I was wondering how I could avoid that step.
My best guess so far is that I should at each step bind the texture in C++ (as in the code above), then right after that in the Python get the bound texture using CUDA, while keeping it on the GPU. However I couldn't figure out how to do that, I don't know much about GPUs and how CUDA/OpenGL work. Pointers to the right direction would be very appreciated!