Running the snippet below takes twice as long when run inside a docker-container (keeping all python versions the same) using the Nvidia-docker runtime, then locally executed.
I tried various different base images, but they have a similar (slower) runtime. I would have expected that running code inside a Docker container does not impact runtime performance? So I feel that I'm missing something.
import numpy as np
from time import time
import torch
def run():
print("doing run")
random_data = []
for _ in range(320):
random_data.append(np.random.randint(256, size=(1, 84, 84)))
tensor = torch.tensor(random_data, device="cuda")
print(tensor.shape)
n_runs = int(1e3)
runtimes = []
for i in range(n_runs):
start = time()
run()
end = time()
took = end-start
runtimes.append(took)
print(f"It took {took} second")
print("====")
print(f"avg_runtime: {np.average(runtimes)}")