firstly. I am a bit new to this so apologies if my terms are not correct.
What we are doing
We have files already in S3 in a Binary file format (e.g. Google Protocol Buffers) which we would like to run an ETL job to create our Data-Lake of transformed data that that will be accessed using either Amazon Redshift or Amazon Athena. In future we may stream via Kinesis.
Issue we face
We are looking at using AWS glue but its list of supported formats is limited (CSV, Json, Parquet, Orc, Avro, Grok) and doesn't provide a 'Custom/other' in the docs https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-format.html
Thoughts
- Is the a cost-effective way to pre-transform the data in S3 wihtin the Glue job to a Parquet input?
- Is there to extend AWS Glue with our custom binary format
- Maybe we are using the wrong AWS tool?
Key Considerations
- Costs i.e. If we had to duplicate all the data in S3 for Glue to work on it versus Streaming in-memory transform somehow!
- Later we hope to stream the data using Kinesis
Any help or experience you may have is greatly appreciated, especially examples or existing use-cases as I don't feel what we are trying to do is out of the ordinary... or is it?