Multiprocessing in Python crashes without explanation

Viewed 86

I have a code for training Neural Networks in Python.

I have two GPUs and a list of different trainings to do.

I want to have the GPUs performing different trainings in parallel.

In each training I perform a 3-fold Validation (I called each fold as "splitset")

I am doing something like I write in the code below and sometime I print some message in order to check whether everything is working fine

training_list = [0,1,2]
NUM_OF_PROCESSES = 2
job_queue = Queue()
for training in training_list:
    job_queue.put(training)
start_processes(NUM_OF_PROCESSES,job_queue)



def start_processes(num_of_processes, job_queue):
    device0, device1 = '0', '1'
    m = Manager()
    lock = m.Lock()
    process_dev0 = Process(target=manage_gpus, args=(device0, job_queue,lock),
                           name=device0 + '_Manager')
    process_dev0.start()
    if num_of_processes > 1:
        process_dev1 = Process(target=manage_gpus, args=(device1, job_queue,lock),
                               name=device1 + '_Manager')
        process_dev1.start()
    process_dev0.join()
    if num_of_processes > 1:
        process_dev1.join()



def manage_gpus(device, job_queue, lock):
    while not job_queue.empty():
        trainUid = job_queue.get()
        worker = Process(target=train_and_export_wrapper,
                         args=(trainUid, device,lock),
                         name="Worker_tUid_" + str(trainUid) + "_" + device)
        worker.start()
        worker.join()



def train_and_export_wrapper(trainingUid, device, lock):
    if '0' in device:
        os.environ["CUDA_VISIBLE_DEVICES"] = '0'
    elif '1' in device:
        os.environ["CUDA_VISIBLE_DEVICES"] = '1'
    else:
        raise NotImplementedError(device)
    ...setting parameters...
    print('Device {}: start training {}'.format(device, trainingUid))
    training_function(device,parameters)
    with lock:
        print('Device {}: training {} completed'.format(device, trainingUid))



def training_function(device, parameters):
    splitset_list = [0,1,2]
    for splitsetidx in splitset_list:
        training_set = ...
        validation_set = ...
        test_set = ...
        model = ...
        print('Device {}: start training split {}'.format(device, splitsetidx))
        history = model.fit(...)
        print('Device {}: predictions'.format(device))
        predictions = model.predict(test_set)
        print('Device {}: end of split {}'.format(device, splitsetidx))  
        del training_set, validation_set, test_set, model, history, predictions 

When I run this code, these are the output I see:

Device 0: start training 0
Device 1: start training 1
Device 0: start training split 0
Device 1: start training split 0
Device 0: predictions
Device 1: predictions
Device 0: end of split 0
Device 1: end of split 0
Device 0: start training split 1
Device 1: start training split 1
Device 0: predictions
Device 1: predictions
Device 0: end of split 1
Device 1: end of split 1
Device 0: start training split 2
Device 1: start training split 2
Device 0: predictions
Device 1: predictions
---->HERE DEVICE 0 DOES NOT finish his job on training 0 and immediately starts training 2, while device 1 ends his task correctly.
Device 1: end of split 2
Device 0: start training 2
Device 1: training 1 completed
----> now Device 0 does all the necessary steps regarding training 2, correctly

So, what is happening here? I have two hypotesis: either there is a mistake in my code, either the data I am using is too huge (this is possible) and at a certain point one GPU crashes, it exits from training_function and starts another training. But, in order to prevent the crash, I am always deleting everything before starting another split/fold, so crashes sohould not happen.

0 Answers
Related