I know that in Big Data areas the usage of data storage file formats such as Parquet, Avro and more is very wide. I know that these formats are meant to improve performance, compatibility, schema evolution, compression and more. I want to focus on compression and understand why exactly do these formats use behind the scenes compression formats like gzip, zlib and snappy?
And this leads me to the main question I have - what’s the difference between keeping my data in a gzip format, to keeping it in Parquet? Why are compression formats take place in a different category, and not just other options for data storage formats?