Error using PBSCluster with multiple nodes

Viewed 104

I'm encountering some strange behaviour when running xarray with a dask client on our PBS Cluster. When choosing multiple nodes, aka - cluster.scale(<int>), the process constantly fails prompting the error below. The strange thing is that when running only one 1 machine 'cluster.scale(1)' the process runs smoothly

Code to summon workers - '''

cluster = PBSCluster(queue      = 'some_q_name',
                             project    = 'project1',
                             cores      = 16,
                             memory     = '100GB',
                             processes  = 1,
                             walltime   = '48:00:00')
cluster.scale(2)
client = Client(cluster)

'''

Error- '''

distributed.utils - ERROR - "('mean_combine-partial-8d870d263fee8efd863c386e1a41a643', 48, 0, 0, 0)"

Traceback (most recent call last):

File "/work/stavn/anaconda3/envs/Dask_v2/lib/python3.7/site-packages/distributed/utils.py", line 656, in log_errors
    yield

File "/work/stavn/anaconda3/envs/Dask_v2/lib/python3.7/site-packages/distributed/scheduler.py", line 1736, in add_worker
    typename=types[key],

KeyError: "('mean_combine-partial-8d870d263fee8efd863c386e1a41a643', 48, 0, 0, 0)"
distributed.core - ERROR - "('mean_combine-partial-8d870d263fee8efd863c386e1a41a643', 48, 0, 0, 0)"

Traceback (most recent call last):
  File "/work/stavn/anaconda3/envs/Dask_v2/lib/python3.7/site-packages/distributed/core.py", line 459, in handle_comm
    result = await result
  File "/work/stavn/anaconda3/envs/Dask_v2/lib/python3.7/site-packages/distributed/scheduler.py", line 1736, in add_worker
    typename=types[key],

KeyError: "('mean_combine-partial-8d870d263fee8efd863c386e1a41a643', 48, 0, 0, 0)"

'''

0 Answers
Related