I'm encountering some strange behaviour when running xarray with a dask client on our PBS Cluster.
When choosing multiple nodes, aka - cluster.scale(<int>), the process constantly fails prompting the error below.
The strange thing is that when running only one 1 machine 'cluster.scale(1)' the process runs smoothly
Code to summon workers - '''
cluster = PBSCluster(queue = 'some_q_name',
project = 'project1',
cores = 16,
memory = '100GB',
processes = 1,
walltime = '48:00:00')
cluster.scale(2)
client = Client(cluster)
'''
Error- '''
distributed.utils - ERROR - "('mean_combine-partial-8d870d263fee8efd863c386e1a41a643', 48, 0, 0, 0)"
Traceback (most recent call last):
File "/work/stavn/anaconda3/envs/Dask_v2/lib/python3.7/site-packages/distributed/utils.py", line 656, in log_errors
yield
File "/work/stavn/anaconda3/envs/Dask_v2/lib/python3.7/site-packages/distributed/scheduler.py", line 1736, in add_worker
typename=types[key],
KeyError: "('mean_combine-partial-8d870d263fee8efd863c386e1a41a643', 48, 0, 0, 0)"
distributed.core - ERROR - "('mean_combine-partial-8d870d263fee8efd863c386e1a41a643', 48, 0, 0, 0)"
Traceback (most recent call last):
File "/work/stavn/anaconda3/envs/Dask_v2/lib/python3.7/site-packages/distributed/core.py", line 459, in handle_comm
result = await result
File "/work/stavn/anaconda3/envs/Dask_v2/lib/python3.7/site-packages/distributed/scheduler.py", line 1736, in add_worker
typename=types[key],
KeyError: "('mean_combine-partial-8d870d263fee8efd863c386e1a41a643', 48, 0, 0, 0)"
'''