Simple PytorchBenchmark from transformers lib of Huggingface script gives CUDA initialisation error

Viewed 276

I am trying to run a simple benchmark script from the transformers lib from Huggingface, but it fails due to a CUDA error, which leads to another error:

1 / 1
Cannot re-initialize CUDA in forked subprocess. To use CUDA with multiprocessing, you must use the 'spawn' start method
Cannot re-initialize CUDA in forked subprocess. To use CUDA with multiprocessing, you must use the 'spawn' start method
Traceback (most recent call last):
  File "/home/cbarkhof/code-thesis/Experimentation/Benchmarking/benchmarking-models-claartje.py", line 12, in <module>
    benchmark.run()
  File "/home/cbarkhof/.local/lib/python3.6/site-packages/transformers/benchmark/benchmark_utils.py", line 674, in run
    memory, inference_summary = self.inference_memory(model_name, batch_size, sequence_length)
ValueError: too many values to unpack (expected 2)

The script simply follows the example like shown on this page:

from transformers import PyTorchBenchmark, PyTorchBenchmarkArguments

benchmark_args = PyTorchBenchmarkArguments(models=["bert-base-uncased"],
                                           batch_sizes=[8],
                                           sequence_lengths=[8, 32, 128, 512],
                                           save_to_csv=True,
                                           log_filename='log',
                                           env_info_csv_file='env_info')

benchmark = PyTorchBenchmark(benchmark_args)
benchmark.run()

If anyone can point me to why this might be happening. Please let me know :). Cheers!

1 Answers

This error is due to the CUDA runtime not supporting process forking which is the default method of multiprocessing in PyTorch. The suggested workaround is to call torch.multiprocessing.set_start_method('spawn') before importing transformers, but in my experiments this raises an AttributeError: Can't pickle local object error.

If you're happy with the script running as a single process when using GPUs you can pass the multi_process=False arg, or --no_multi_process flag if using the command line script. Alternatively, if running this on CPU only without CUDA, you should be able to make use of multiple cores and passing the cuda=False arg or --no_cuda flag should also work.

Related