I am trying to unzip uploaded zip files to Cloud Storage which contains only image files without any other folders inside.
I was able to do that with cloud functions but seems like I get memory-related issues when files get bigger. I found Dataflow templates (Bulk Decompress Cloud Storage Files) for this specific case and tried to run some jobs with similar to below parameters.
{
"jobName": "unique_job_name",
"environment": {
"bypassTempDirValidation": false,
"numWorkers": 2,
"tempLocation": "gs://bucket_name/temp",
"ipConfiguration": "WORKER_IP_UNSPECIFIED",
"additionalExperiments": []
},
"parameters": {
"inputFilePattern": "gs://bucket_name/root_path/zip_to_extract.zip",
"outputDirectory": "gs://bucket_name/root_path/",
"outputFailureFile": "gs://bucket_name/root_path/failure.csv"
}
}
As an output, I only get 1 file with the same name of my zip file without a file extension and with the type of text/plain.
Is this an expected behaviour? If someone could help me to unzip the file with Dataflow, I would be glad.
Thanks