I have a code for training Neural Networks in Python.
I have two GPUs and a list of different trainings to do.
I want to have the GPUs performing different trainings in parallel.
In each training I perform a 3-fold Validation (I called each fold as "splitset")
I am doing something like I write in the code below and sometime I print some message in order to check whether everything is working fine
training_list = [0,1,2]
NUM_OF_PROCESSES = 2
job_queue = Queue()
for training in training_list:
job_queue.put(training)
start_processes(NUM_OF_PROCESSES,job_queue)
def start_processes(num_of_processes, job_queue):
device0, device1 = '0', '1'
m = Manager()
lock = m.Lock()
process_dev0 = Process(target=manage_gpus, args=(device0, job_queue,lock),
name=device0 + '_Manager')
process_dev0.start()
if num_of_processes > 1:
process_dev1 = Process(target=manage_gpus, args=(device1, job_queue,lock),
name=device1 + '_Manager')
process_dev1.start()
process_dev0.join()
if num_of_processes > 1:
process_dev1.join()
def manage_gpus(device, job_queue, lock):
while not job_queue.empty():
trainUid = job_queue.get()
worker = Process(target=train_and_export_wrapper,
args=(trainUid, device,lock),
name="Worker_tUid_" + str(trainUid) + "_" + device)
worker.start()
worker.join()
def train_and_export_wrapper(trainingUid, device, lock):
if '0' in device:
os.environ["CUDA_VISIBLE_DEVICES"] = '0'
elif '1' in device:
os.environ["CUDA_VISIBLE_DEVICES"] = '1'
else:
raise NotImplementedError(device)
...setting parameters...
print('Device {}: start training {}'.format(device, trainingUid))
training_function(device,parameters)
with lock:
print('Device {}: training {} completed'.format(device, trainingUid))
def training_function(device, parameters):
splitset_list = [0,1,2]
for splitsetidx in splitset_list:
training_set = ...
validation_set = ...
test_set = ...
model = ...
print('Device {}: start training split {}'.format(device, splitsetidx))
history = model.fit(...)
print('Device {}: predictions'.format(device))
predictions = model.predict(test_set)
print('Device {}: end of split {}'.format(device, splitsetidx))
del training_set, validation_set, test_set, model, history, predictions
When I run this code, these are the output I see:
Device 0: start training 0
Device 1: start training 1
Device 0: start training split 0
Device 1: start training split 0
Device 0: predictions
Device 1: predictions
Device 0: end of split 0
Device 1: end of split 0
Device 0: start training split 1
Device 1: start training split 1
Device 0: predictions
Device 1: predictions
Device 0: end of split 1
Device 1: end of split 1
Device 0: start training split 2
Device 1: start training split 2
Device 0: predictions
Device 1: predictions
---->HERE DEVICE 0 DOES NOT finish his job on training 0 and immediately starts training 2, while device 1 ends his task correctly.
Device 1: end of split 2
Device 0: start training 2
Device 1: training 1 completed
----> now Device 0 does all the necessary steps regarding training 2, correctly
So, what is happening here? I have two hypotesis: either there is a mistake in my code, either the data I am using is too huge (this is possible) and at a certain point one GPU crashes, it exits from training_function and starts another training. But, in order to prevent the crash, I am always deleting everything before starting another split/fold, so crashes sohould not happen.