best way to migrate billions of files on a single partition in a data center to s3?

Viewed 678

We have a data center with a 10G direct connect circuit to AWS. In the data center, we have an IBM XIV storage infrastructure with GPFS filesystems containing 1.5 BILLION images (about 50k each) in the single top level directory. We could argue all day about how dumb this was, but I'd rather seek advice for my task which is moving all these files into an s3 bucket.

I can't use any physical transport solution, as the data center is physically locked down and obtaining on-premises physical clearance is a 6-month process.

What is the best way to do this file migration?

The best idea I have so far is building an EC2 linux server in AWS, mounting the s3 destination bucket using s3fs-fuse (https://github.com/s3fs-fuse/s3fs-fuse/wiki/Fuse-Over-Amazon) as a filesystem on the EC2 server, and then running some netcat + tar command between the data center server holding the GPFS mount and the EC2 server. I found this suggestion on another post: Destination box: nc -l -p 2342 | tar -C /target/dir -xzf - Source box: tar -cz /source/dir | nc Target_Box 2342

Before I embark on a task that could take a month, I wanted to see if anyone here had a better way to do this?

2 Answers
Related