Each time I load a transformer model into GPU it takes ~60 seconds.
So, I want to access the model in GPU across flask requests without initiating it each time.
So, I tried to save the model in BaseManager and then access it.
from multiprocessing.managers import BaseManager
manager = BaseManager(('', 37844), b'password')
manager.connect()
generator = pipeline('text-generation', model=MODEL_NAME, device=1)
manager.register('generator', generator)
but while trying to access the model
with
generator = manager.generator()
I get the following error
Cannot re-initialize CUDA in forked subprocess. To use CUDA with multiprocessing, you must use the 'spawn' start method
and while digging further into the error it asked to use multiprocessor from torch instead but that doesn't have a BaseManager.
from torch.multiprocessing import Pool, Process, set_start_method
So, how does one efficiently use Models across requests in Flask?