How can I run multiple processes avoiding this error?

Viewed 56

I am writing an algorithm where I have a neural network written with tensorflow and I create different processes copying the aforementioned network and then I update the weights using gradients computed in the subprocesses. When I start the processes I get the following error from tensorflow, how can I fix this problem?

        for _ in range(num_processes):
            # creation of an agent
            agent_process = agent_class(obs_space, act_space)
            # policy and value_net are tensorflow Model object and clone_model is a tensorflow function
            agent_process.actor = clone_model(policy)
            agent_process.critic = clone_model(value_net)
            tuple_process = (
                agent_process,
                env_class(),
                gamma,
                max_steps,
                queue_actor,
                queue_critic
            )
            processes.append(Process(target=train_a2c_single_agent, args=tuple_process))

        # start the processes
        for proc in processes:
            proc.start()

I am using multiprocessing python library for the multiprocessing part and tensorflow for the neural network. I have executed the single process function without multiprocessing and it works well, therefore I think it is a problem of multiprocessing settings. This is the error I obtain:

Traceback (most recent call last): File "", line 1, in File "C:\Users\boezi\AppData\Local\Programs\Python\Python310\lib\multiprocessing\spawn.py", line 116, in spawn_main exitcode = _main(fd, parent_sentinel) File "C:\Users\boezi\AppData\Local\Programs\Python\Python310\lib\multiprocessing\spawn.py", line 126, in _main self = reduction.pickle.load(from_parent) File "C:\Users\boezi\PycharmProjects\AWSDeepRacerChallenge\env\lib\site-packages\keras\saving\pickle_utils.py", line 48, in deserialize_model_from_bytecode model = save_module.load_model(temp_dir) File "C:\Users\boezi\PycharmProjects\AWSDeepRacerChallenge\env\lib\site-packages\keras\utils\traceback_utils.py", line 67, in error_handler raise e.with_traceback(filtered_tb) from None File "C:\Users\boezi\PycharmProjects\AWSDeepRacerChallenge\env\lib\site-packages\tensorflow\python\saved_model\load.py", line 915, in load_partial raise FileNotFoundError( FileNotFoundError: Unsuccessful TensorSliceReader constructor: Failed to find any matching files for ram://0af1921e-ad76-4108-8fe1-525f2697738d/variables/variables You may be trying to load on a different device from the computational device. Consider setting the experimental_io_device option in tf.saved_model.LoadOptions to the io_device such as '/job:localhost'.

I tried the suggested solution setting tf.saved_model.LoadOptions(experimental_io_device='/job:localhost') but I get the same error

0 Answers
Related