How to parallelize a processor- and an I/O-intensive operations?

Viewed 46

I'm writing a helper utility to ease SFTP-uploads of large files to a remote server. In addition to the actual transfer (which command-line sftp or scp could accomplish), the utility is supposed to print the SHA256 of each file.

The files are large and the "local" storage of them is, actually, a slow NFS-mount -- so I don't want to re-read them again for the checksum. While the buffer is in memory, it can be both digested and pushed to the remote.

So, my code does:

async def upload(fName, SFTP):
    inp = open(fName, "rb")
    digest = hashlib.sha256()
    bsize = os.stat(inp.fileno()).st_blksize
    out = SFTP.open(os.path.split(fName)[-1], "w")

    while True:
        buf = inp.read(bsize)
        if not buf:
            break
        digest.update(buf)
        out.write(buf)

    inp.close()
    out.close()
    print('SHA256 (%s) = %s' % (fName, digest.hexdigest()))

This, actually, works -- and multiple instances of upload are running over multiple SFTP-connections.

But it annoys me, that the CPU-intensive digest.update() is not in parallel with the I/O-bound out.write(). Can the two be parallelized?

How about the two open calls -- the local input and the remote output?

1 Answers

On mainstreams platforms (eg. Windows, Linux and MacOS with standard hardware) file writes are buffered and asynchronous: data is put to a memory buffer contiguously written back to the disk. Whether this is the case for networking file system like a NFS is dependent of the actual software stack used but it would be very surprising to be synchronous (due to the network latency). In fact, the TCP stack for example behave similarly by storing data in a waiting buffer and weakly synchronizing packet acknowledgments (via a packet window). Thus, the low-level asynchronicity should already help to overlap the writes with the computation (assuming there is not to much jitter). That being said, if the computation is slower than the writes, then you can compute the digest of multiple files in parallel using a pipeline strategy. The idea is to create multiple worker processes, send them the work to do, asynchronously wait for the result, so the workers are kept buzzy as much as possible. The inter-process communication may be a problem though (it is slow). Alternatively, the whole operation can be done in workers (reads+write+digests). It may help to mitigate the latency of the NFS (assuming requests are not serialized on the server side which is quite frequent unfortunately). Functions like apply_async or map_async of the multiprocessing package should be helpful for that. If using multiple processes is a problem, then note that a ThreadPool can also be used instead to mitigate the latency of the IO operations, but not using use multiple cores (due to the global interpreter lock -- aka GIL).

Related