My goal is creating a mechanism that when a new file is uploaded into the Cloud Storage, it'll trigger a Cloud Function. Eventually, this Cloud function will trigger a Cloud Dataflow job.
I have a restriction that the Cloud Dataflow job should be written in Go, and the Cloud Function should be written in Python.
The problem I have been facing right now is, I cannot call Cloud Dataflow job from a Cloud Function.
The problem in Cloud Dataflow written in Go is there is no template-location variable defined in Apache Beam Go SDK. That's why I cannot create dataflow templates. And, since there is no dataflow templates, the only way that I can call Cloud Dataflow job from a cloud function is writing a Python job which calls a bash script which runs dataflow job.
The bash script looks like that:
go run wordcount.go \
--runner dataflow \
--input gs://dataflow-samples/shakespeare/kinglear.txt \
--output gs://${BUCKET?}/counts \
--project ${PROJECT?} \
--temp_location gs://${BUCKET?}/tmp/ \
--staging_location gs://${BUCKET?}/binaries/ \
--worker_harness_container_image=apache-docker-beam-snapshots-docker.bintray.io/beam/go:20180515
But above mechanism cannot create a new dataflow job and it seems it's cumbersome.
Is there a better way to achieve my goal? And what am I doing wrong on the above mechanism?