I'm working on a project that incorporates file storage and sharing features and after months of researching the best method to leverage AWS I'm still a little concerned.
Basically my decision is between using EBS storage to house user files or S3. The system will incorporate on-the-fly zip archiving when the user wants to download a handful of files. Also, when users download any files I don't want the URL to the files exposed.
The two best options I've come up with are:
Have an EC2 instance which has a number of EBS volumes mounted to store user files.
- pros: It seems much faster than S3, and zipping files from the EBS volume is straight forward.
- cons: I believe Amazon caps how much EBS storage you can use and there is not as redundant as S3.
After files are uploaded and processed, the system pushes those files to an S3 bucket for long term storage. When files are requested I will retrieve the files from S3 and output back to the client.
- pros: Redundancy, no file storage limits
- cons: It seems very SLOW, no way to mount an S3 bucket as a volume in filesystem, serving zipped files would mean transferring each file to the EC2 instance, zipping, and then finally sending output (again, slow!)
Are any of my assumptions flawed? Can anyone think of a better way of managing massive amounts of file storage?