Best approach for batch processing objects located on a Google Cloud Storage Bucket

Viewed 449

I have a bunch on videos on a GCP Storage bucket that I need to concatenate using ffmpeg, then save the result to another bucket.

I'm quite new to GCP so what I would normally do is to spin up a VM and have it (through a script) download the videos using gsutil, process them and then upload the result with gsutil again, but as I understand this is would be very inefficient on network traffic, processing costs and scalability.

So, in a very general way, which will be the best GCP's inbuilt feature to run such a script: App Engine, Cloud Functions or Cloud Run, and what will it entail?

2 Answers

I would say that there are no wrong answers to your question, since App Engine, Cloud Functions and Cloud Run would all work for what you are trying to achieve and the costs associated with this would be similar.

While you make your decision you should consider:

  • What tool you are most familiar with;
  • What does Google Cloud recommends while choosing a serverless product for my app, you can find a good article about this on their documentation;
  • How you want this to scale.

Personally I would go with App Engine with the problem you described.

NOTE: I know that this answer was pretty generic but it really depends on specificities of your use case and what you want you solution to look like.

You can automate the process by creating an application and using the Google libraries to upload the result into a bucket. Here is an example of an upload in python:

namespace gcs = google::cloud::storage;
using ::google::cloud::StatusOr;
[](gcs::Client client, std::string const& file_name,
   std::string const& bucket_name, std::string const& object_name) {
  // Note that the client library automatically computes a hash on the
  // client-side to verify data integrity during transmission.
  StatusOr<gcs::ObjectMetadata> metadata = client.UploadFile(
      file_name, bucket_name, object_name, gcs::IfGenerationMatch(0));
  if (!metadata) throw std::runtime_error(metadata.status().message());

  std::cout << "Uploaded " << file_name << " to object " << metadata->name()
            << " in bucket " << metadata->bucket()
            << "\nFull metadata: " << *metadata << "\n";
}

I can see that ffmpeg also works with python so I believe that including all this into a single program will make it easier. Other possibilities for uploading/downloading an object in Storage buckets are the Cloud Console and REST APIs.

For more information you can also check the Google documentation where you can find sample codes which might be helpful to you.

Related