Not able to start Ray client server at the second and next executions

Viewed 107

I have a Ray cluster deployed on an AKS cluster. Currently I have only 1 worker node. The version of Ray I'm using is the 1.9.0.

Ray namespace pods

The pod called linear-model-5cd66b57d8-rn6ft contains the code to run the training of a Pytorch model. I got the code from the official documentation: Ray Pytorch code example.

I'm trying to connect to the Ray cluster using

ray.init(address="ray://10.0.210.159:10001", runtime_env=ray_env)

where 10.0.210.159 is the IP address of the Kubernetes service that creates Ray automatically installing it using the helm chart.

ray_env contains some needed Python modules:

ray_env = {"conda": {"dependencies": ["pytorch", "torchvision", "pip", {"pip": ["numpy"]}]}}

Now, the first time that the pod has been scheduled, all worked fine (Pytorch model trained), but in the second and in the next executions I got the following error:

enter image description here

Does someone know why I get this error?

Thanks in advance.

UPDATE

Now I'm getting this error:

Conda error

Something wrong with Conda or pip.

0 Answers
Related