Since I need to upload a large number of files over 100000 to azure blob storage, I wrote a program to upload by multi-thread processing like this.
from azure.storage.blob import BlobServiceClient, BlobClient
from itertools import repeat
from concurrent.futures import ThreadPoolExecutor
import os
def upload_single_blob(blob_service_client, blob_path):
# Create a blob client using the local file name as the name for the blob
blob_client = blob_service_client.get_blob_client(container='MyContainer',
blob=blob_path)
# Upload the file
with open(blob_path, "rb") as data:
blob_client.upload_blob(data)
# make blob service client from connect str
blob_service_client = BlobServiceClient.from_connection_string(connect_str)
# make file path list to upload
blob_path_list = os.listdir("./blob_files/")
blob_path_list = map(lambda x: "./blob_files/"+x, blob_path_list)
blob_path_list = list(blob_path_list)
# multi threading upload to blob
with ThreadPoolExecutor(max_workers=100) as executor:
executor.map(upload_single_blob, repeat(blob_service_client), blob_path_list)
However, when I ran this program at azure VM (OS is ubuntu18.04), I got the warning a lot.
urllib3.connectionpool WARNING --Connection pool is full, discarding connection: myblobaccount.blob.core.windows.net
I didn't measure it accurately, but it seemed that there were only about 10 connections at the same time, even though uploading in parallel with 100 threads.
How can I increase the number of connections any more?