I'm trying to use the du command but I'm not sure how to filter by file size. Trying to do this to delete big files that I simply don't need and are costing me money.
I'm trying to use the du command but I'm not sure how to filter by file size. Trying to do this to delete big files that I simply don't need and are costing me money.
GSutil offers some ways of sortering the objects inside a bucket, but not by file size; you can use a mix of linux/gsutil commands to help you out. For example, this:
List objects sorted by size descending with human redible sizes:
gsutil ls -lh gs://{bucket} | sort -n -k 1
Breaking a little the command:
gsutil ls: List providers, buckets, or objects
-l: Prints long listing (owner, length)
h: When used with -l, prints object sizes in human readable format
In case you need to do it recursively, add
-r-n: To sort a file numerically
-k 1: to sort on a certain column. For example, use
-k 2to sort on the second column
With python you can get a list of all blobs in your bucket and loop through that list to get those with a large size:
from google.cloud import storage
storage_client = storage.Client()
blobs_list = storage_client.list_blobs(bucket_or_name='name_of_your_bucket')
large_files = []
for blob in blobs_list:
if blob.size > 1_000_000: # size is in bytes
large_files.append(blob.id)
There isn't native command to achieve this. You need to parse all the file (only the metadata, you don't need to download the content) to get their size and act accordingly.
If your bucket contains lot of files, it could take hours. You can try to parallelize and partition the process by prefixes for example.