I am writing an algorithm where I have a neural network written with tensorflow and I create different processes copying the aforementioned network and then I update the weights using gradients computed in the subprocesses. When I start the processes I get the following error from tensorflow, how can I fix this problem?
for _ in range(num_processes):
# creation of an agent
agent_process = agent_class(obs_space, act_space)
# policy and value_net are tensorflow Model object and clone_model is a tensorflow function
agent_process.actor = clone_model(policy)
agent_process.critic = clone_model(value_net)
tuple_process = (
agent_process,
env_class(),
gamma,
max_steps,
queue_actor,
queue_critic
)
processes.append(Process(target=train_a2c_single_agent, args=tuple_process))
# start the processes
for proc in processes:
proc.start()
I am using multiprocessing python library for the multiprocessing part and tensorflow for the neural network. I have executed the single process function without multiprocessing and it works well, therefore I think it is a problem of multiprocessing settings. This is the error I obtain:
Traceback (most recent call last): File "", line 1, in File "C:\Users\boezi\AppData\Local\Programs\Python\Python310\lib\multiprocessing\spawn.py", line 116, in spawn_main exitcode = _main(fd, parent_sentinel) File "C:\Users\boezi\AppData\Local\Programs\Python\Python310\lib\multiprocessing\spawn.py", line 126, in _main self = reduction.pickle.load(from_parent) File "C:\Users\boezi\PycharmProjects\AWSDeepRacerChallenge\env\lib\site-packages\keras\saving\pickle_utils.py", line 48, in deserialize_model_from_bytecode model = save_module.load_model(temp_dir) File "C:\Users\boezi\PycharmProjects\AWSDeepRacerChallenge\env\lib\site-packages\keras\utils\traceback_utils.py", line 67, in error_handler raise e.with_traceback(filtered_tb) from None File "C:\Users\boezi\PycharmProjects\AWSDeepRacerChallenge\env\lib\site-packages\tensorflow\python\saved_model\load.py", line 915, in load_partial raise FileNotFoundError( FileNotFoundError: Unsuccessful TensorSliceReader constructor: Failed to find any matching files for ram://0af1921e-ad76-4108-8fe1-525f2697738d/variables/variables You may be trying to load on a different device from the computational device. Consider setting the
experimental_io_deviceoption intf.saved_model.LoadOptionsto the io_device such as '/job:localhost'.
I tried the suggested solution setting tf.saved_model.LoadOptions(experimental_io_device='/job:localhost') but I get the same error