The problem I'm trying to solve is the following: I have a desktop computer with large amounts of data (~5 TB), that I want to analyze. The data consist of 500k files, and each file can be analyzed individually. For the analysis I have a series of servers at the university available, however, the server does not have space for all this data, nor does it have space to store the output of the analysis.
So my idea is to copy the data over to the server in segments, run the analysis, transfer the results back to the desktop, delete both input and output data on the server, and repeat.
For the file transfer I installed paramiko yesterday and it seems to work great:
remote_get = 'test'
local_deliver = './test'
ssh = paramiko.SSHClient()
ssh.set_missing_host_key_policy(paramiko.AutoAddPolicy())
ssh.load_host_keys(os.path.expanduser(os.path.join("~", ".ssh", "known_hosts")))
ssh.connect(server, username=username, password=password)
sftp = ssh.open_sftp()
for root, dirs, files in os.walk(local_path):
for fname in files:
full_fname = os.path.join(root, fname)
full_remote = os.path.join(remote_path, fname)
sftp.put(full_fname, full_remote)
sftp.close()
ssh.close()
However my only problem is that the amount of data I will need to transfer will likely take days to get back and forth, and hence I would love to start the data transfer asynchronously if possible, such that I can do analysis on the current dataset while transferring the next dataset to be analyzed.
But I don't have any clue how to do such a thing, can anyone point me in the right direction?