I am trying to copy a large number of files between buckets, and am only getting around 15 files per second. That is not usable, with 500k files...
So I was wondering if it actually makes any difference to use a wildcard in a cp statement, as opposed to sending individual cp statements? What is the "standard" to use here? Or are both resulting in the same client-side and server load?
as an example, I have now written code to group files based on their batch id and send them in groups. But I do not get the impression (from a very basic test) that it is faster?
e.g.,
aws s3 cp <path>/XY.15937610001 <path_to>
aws s3 cp <path>/XY.15937610002 <path_to>
aws s3 cp <path>/XY.15937610003 <path_to>
:
aws s3 cp <path>/XY.15937615999 <path_to>
versus:
cmd
aws s3 cp <path> <path_to> --recursive --exclude="*" --include="XY.159376*"
thank you
PS edit - is the only way to speed this up, using max_concurrent_sessions, or something like S3DistCp (s3-dist-cp) (whatever that may be)? Both options are not available to me currently...