AWS GLUE transform for binary files in S3 from Protobuf (Google Protocol Buffers) for AWS Athena

Viewed 1054

firstly. I am a bit new to this so apologies if my terms are not correct.

What we are doing

We have files already in S3 in a Binary file format (e.g. Google Protocol Buffers) which we would like to run an ETL job to create our Data-Lake of transformed data that that will be accessed using either Amazon Redshift or Amazon Athena. In future we may stream via Kinesis.

Issue we face

We are looking at using AWS glue but its list of supported formats is limited (CSV, Json, Parquet, Orc, Avro, Grok) and doesn't provide a 'Custom/other' in the docs https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-format.html

Thoughts

  • Is the a cost-effective way to pre-transform the data in S3 wihtin the Glue job to a Parquet input?
  • Is there to extend AWS Glue with our custom binary format
  • Maybe we are using the wrong AWS tool?

Key Considerations

  • Costs i.e. If we had to duplicate all the data in S3 for Glue to work on it versus Streaming in-memory transform somehow!
  • Later we hope to stream the data using Kinesis

Any help or experience you may have is greatly appreciated, especially examples or existing use-cases as I don't feel what we are trying to do is out of the ordinary... or is it?

0 Answers
Related