I have a single server with 4 gpu's. I'm trying to run some tensorflow operations in parallel, 4 at a time. I've oversimplified the usecase in this post by the idea is to run different model trainings in parallel on the 4 gpu's.
I'm running Spark in standalone mode.
From what I can see from the Spark UI and in the executor logs, gpu resources are correctly assigned to each executor.
But it seems that tensorflow is not aware of such scheduling in my case because model training endup on the same gpu.
Here is my Spark spark-env.sh configuration:
spark.worker.resource.gpu.discoveryScript=/opt/spark/conf/getGpusResources.sh
spark.worker.resource.gpu.amount=1
spark.executor.resource.gpu.discoveryScript=/opt/spark/conf/getGpusResources.sh
spark.executor.resource.gpu.amount=1
spark.task.resource.gpu.amount=0.25
I launch Spark master and spark worker the following way:
spark/sbin/start-master.sh
spark/sbin/start-worker.sh spark://0.0.0.0:7077
Here is my simple test.py python script:
import tensorflow as tf
from pyspark import SparkContext
from pyspark import SparkConf
sc = SparkContext()
def launch_on_executor():
tf.debugging.set_log_device_placement(True)
# Create some tensors
a = tf.constant([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]])
b = tf.constant([[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]])
c = tf.matmul(a, b)
return c
executors = [(1,),]*4
res = sc.parallelize(executors).map(lambda x: launch_on_executor()).collect()
print(res)
Once I launch the test.py via:
spark/bin/spark-submit --master spark://0.0.0.0:7077 --conf "spark.executor.cores=1" test.py
I have then 4 executor logs, in each I can see something similar expect on the gpu addresses that show different value for each executor (which seems good):
21/10/11 09:31:21 INFO ResourceUtils: ==============================================================
21/10/11 09:31:21 INFO ResourceUtils: Custom resources for spark.executor:
gpu -> [name: gpu, addresses: 0]
21/10/11 09:31:21 INFO ResourceUtils: ==============================================================
My thoughts were that each executor would only be able to use the gpu instance it was assigned by the resource scheduler but it is not the case.
If I look at executor 1 logs, I can see that gpu address 1 has been assigned by Spark to it but still all gpus present on the server are visible by tensorflow and it fails to run the op on device:GPU:0 since it is also used by another executor at the same time.
21/10/13 07:39:36 INFO ResourceUtils: ==============================================================
21/10/13 07:39:36 INFO ResourceUtils: Custom resources for spark.executor:
gpu -> [name: gpu, addresses: 1]
21/10/13 07:39:36 INFO ResourceUtils: ==============================================================
21/10/13 07:39:36 INFO CoarseGrainedExecutorBackend: Successfully registered with driver
21/10/13 07:39:36 INFO Executor: Starting executor ID 1 on host 172.17.0.3
21/10/13 07:39:36 INFO Utils: Successfully started service 'org.apache.spark.network.netty.NettyBlockTransferService' on port 33061.
21/10/13 07:39:36 INFO NettyBlockTransferService: Server created on 172.17.0.3:33061
21/10/13 07:39:36 INFO BlockManager: Using org.apache.spark.storage.RandomBlockReplicationPolicy for block replication policy
21/10/13 07:39:36 INFO BlockManagerMaster: Registering BlockManager BlockManagerId(1, 172.17.0.3, 33061, None)
21/10/13 07:39:36 INFO BlockManagerMaster: Registered BlockManager BlockManagerId(1, 172.17.0.3, 33061, None)
21/10/13 07:39:36 INFO BlockManager: Initialized BlockManager: BlockManagerId(1, 172.17.0.3, 33061, None)
21/10/13 07:39:44 INFO CoarseGrainedExecutorBackend: Got assigned task 2
21/10/13 07:39:44 INFO Executor: Running task 0.1 in stage 0.0 (TID 2)
21/10/13 07:39:44 INFO TorrentBroadcast: Started reading broadcast variable 0 with 1 pieces (estimated total size 4.0 MiB)
21/10/13 07:39:44 INFO TransportClientFactory: Successfully created connection to /172.17.0.3:33127 after 3 ms (0 ms spent in bootstraps)
21/10/13 07:39:44 INFO MemoryStore: Block broadcast_0_piece0 stored as bytes in memory (estimated size 3.3 KiB, free 366.3 MiB)
21/10/13 07:39:44 INFO TorrentBroadcast: Reading broadcast variable 0 took 125 ms
21/10/13 07:39:44 INFO MemoryStore: Block broadcast_0 stored as values in memory (estimated size 4.9 KiB, free 366.3 MiB)
2021-10-13 07:39:45.768689: I tensorflow/stream_executor/platform/default/dso_loader.cc:53] Successfully opened dynamic library libcudart.so.11.0
2021-10-13 07:39:46.827378: I tensorflow/stream_executor/platform/default/dso_loader.cc:53] Successfully opened dynamic library libcuda.so.1
2021-10-13 07:39:48.211112: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1734] Found device 0 with properties:
pciBusID: 0000:18:00.0 name: Tesla V100-PCIE-16GB computeCapability: 7.0
coreClock: 1.38GHz coreCount: 80 deviceMemorySize: 15.78GiB deviceMemoryBandwidth: 836.37GiB/s
2021-10-13 07:39:48.214266: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1734] Found device 1 with properties:
pciBusID: 0000:5e:00.0 name: Tesla V100-PCIE-16GB computeCapability: 7.0
coreClock: 1.38GHz coreCount: 80 deviceMemorySize: 15.78GiB deviceMemoryBandwidth: 836.37GiB/s
2021-10-13 07:39:48.215748: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1734] Found device 2 with properties:
pciBusID: 0000:86:00.0 name: Tesla V100-PCIE-16GB computeCapability: 7.0
coreClock: 1.38GHz coreCount: 80 deviceMemorySize: 15.78GiB deviceMemoryBandwidth: 836.37GiB/s
2021-10-13 07:39:48.217202: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1734] Found device 3 with properties:
pciBusID: 0000:af:00.0 name: Tesla V100-PCIE-16GB computeCapability: 7.0
coreClock: 1.38GHz coreCount: 80 deviceMemorySize: 15.78GiB deviceMemoryBandwidth: 836.37GiB/s
2021-10-13 07:39:48.217240: I tensorflow/stream_executor/platform/default/dso_loader.cc:53] Successfully opened dynamic library libcudart.so.11.0
2021-10-13 07:39:48.224854: I tensorflow/stream_executor/platform/default/dso_loader.cc:53] Successfully opened dynamic library libcublas.so.11
2021-10-13 07:39:48.224896: I tensorflow/stream_executor/platform/default/dso_loader.cc:53] Successfully opened dynamic library libcublasLt.so.11
2021-10-13 07:39:48.226907: I tensorflow/stream_executor/platform/default/dso_loader.cc:53] Successfully opened dynamic library libcufft.so.10
2021-10-13 07:39:48.227193: I tensorflow/stream_executor/platform/default/dso_loader.cc:53] Successfully opened dynamic library libcurand.so.10
2021-10-13 07:39:48.227719: I tensorflow/stream_executor/platform/default/dso_loader.cc:53] Successfully opened dynamic library libcusolver.so.11
2021-10-13 07:39:48.228677: I tensorflow/stream_executor/platform/default/dso_loader.cc:53] Successfully opened dynamic library libcusparse.so.11
2021-10-13 07:39:48.228813: I tensorflow/stream_executor/platform/default/dso_loader.cc:53] Successfully opened dynamic library libcudnn.so.8
2021-10-13 07:39:48.234907: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1872] Adding visible gpu devices: 0, 1, 2, 3
2021-10-13 07:39:49.218065: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1734] Found device 0 with properties:
pciBusID: 0000:18:00.0 name: Tesla V100-PCIE-16GB computeCapability: 7.0
coreClock: 1.38GHz coreCount: 80 deviceMemorySize: 15.78GiB deviceMemoryBandwidth: 836.37GiB/s
2021-10-13 07:39:49.219650: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1734] Found device 1 with properties:
pciBusID: 0000:5e:00.0 name: Tesla V100-PCIE-16GB computeCapability: 7.0
coreClock: 1.38GHz coreCount: 80 deviceMemorySize: 15.78GiB deviceMemoryBandwidth: 836.37GiB/s
2021-10-13 07:39:49.221142: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1734] Found device 2 with properties:
pciBusID: 0000:86:00.0 name: Tesla V100-PCIE-16GB computeCapability: 7.0
coreClock: 1.38GHz coreCount: 80 deviceMemorySize: 15.78GiB deviceMemoryBandwidth: 836.37GiB/s
2021-10-13 07:39:49.222625: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1734] Found device 3 with properties:
pciBusID: 0000:af:00.0 name: Tesla V100-PCIE-16GB computeCapability: 7.0
coreClock: 1.38GHz coreCount: 80 deviceMemorySize: 15.78GiB deviceMemoryBandwidth: 836.37GiB/s
2021-10-13 07:39:49.233616: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1872] Adding visible gpu devices: 0, 1, 2, 3
2021-10-13 07:39:49.233683: I tensorflow/stream_executor/platform/default/dso_loader.cc:53] Successfully opened dynamic library libcudart.so.11.0
2021-10-13 07:39:50.682630: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1258] Device interconnect StreamExecutor with strength 1 edge matrix:
2021-10-13 07:39:50.682681: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1264] 0 1 2 3
2021-10-13 07:39:50.682690: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1277] 0: N Y Y Y
2021-10-13 07:39:50.682695: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1277] 1: Y N Y Y
2021-10-13 07:39:50.682701: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1277] 2: Y Y N Y
2021-10-13 07:39:50.682706: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1277] 3: Y Y Y N
2021-10-13 07:39:50.714730: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1418] Created TensorFlow device (/job:localhost/replica:0/task:0/device:GPU:0 with 14210 MB memory) -> physical GPU (device: 0, name: Tesla V100-PCIE-16GB, pci bus id: 0000:18:00.0, compute capability: 7.0)
2021-10-13 07:39:50.716431: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1418] Created TensorFlow device (/job:localhost/replica:0/task:0/device:GPU:1 with 14210 MB memory) -> physical GPU (device: 1, name: Tesla V100-PCIE-16GB, pci bus id: 0000:5e:00.0, compute capability: 7.0)
2021-10-13 07:39:50.718775: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1418] Created TensorFlow device (/job:localhost/replica:0/task:0/device:GPU:2 with 14210 MB memory) -> physical GPU (device: 2, name: Tesla V100-PCIE-16GB, pci bus id: 0000:86:00.0, compute capability: 7.0)
2021-10-13 07:39:50.720430: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1418] Created TensorFlow device (/job:localhost/replica:0/task:0/device:GPU:3 with 14210 MB memory) -> physical GPU (device: 3, name: Tesla V100-PCIE-16GB, pci bus id: 0000:af:00.0, compute capability: 7.0)
2021-10-13 07:39:50.730822: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 13.88G (14900789248 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.731733: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 12.49G (13410709504 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.732625: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 11.24G (12069638144 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.733513: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 10.12G (10862673920 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.734405: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 9.10G (9776406528 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.735296: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 8.19G (8798766080 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.736239: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 7.38G (7918889472 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.737124: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 6.64G (7127000576 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.738008: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 5.97G (6414300160 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.738892: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 5.38G (5772870144 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.739779: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 4.84G (5195582976 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.740664: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 4.35G (4676024320 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.741507: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 3.92G (4208421888 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.742287: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 3.53G (3787579648 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.743067: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 3.17G (3408821504 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.743852: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 2.86G (3067939328 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.744647: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 2.57G (2761145344 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.745427: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 2.31G (2485030656 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.746224: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 2.08G (2236527616 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.747073: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 1.87G (2012874752 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.747883: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 1.69G (1811587328 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.748665: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 1.52G (1630428672 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.749453: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 1.37G (1467385856 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.750252: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 1.23G (1320647424 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.751021: I tensorflow/stream_executor/cuda/cuda_driver.cc:732] failed to allocate 1.11G (1188582656 bytes) from device: CUDA_ERROR_OUT_OF_MEMORY: out of memory
2021-10-13 07:39:50.769748: I tensorflow/core/common_runtime/eager/execute.cc:733] Executing op MatMul in device /job:localhost/replica:0/task:0/device:GPU:0
2021-10-13 07:39:50.771307: I tensorflow/stream_executor/platform/default/dso_loader.cc:53] Successfully opened dynamic library libcublas.so.11
2021-10-13 07:39:50.895375: E tensorflow/stream_executor/cuda/cuda_blas.cc:226] failed to create cublas handle: CUBLAS_STATUS_NOT_INITIALIZED
2021-10-13 07:39:50.895435: W tensorflow/stream_executor/stream.cc:1421] attempting to perform BLAS operation using StreamExecutor without BLAS support
21/10/13 07:39:50 ERROR Executor: Exception in task 0.1 in stage 0.0 (TID 2)
I've gone through many Spark and Nvidia docs but I cannot find what I'm doing wrong here. Any thoughts?
thanks