from Spark RDDs, I want to stage and archive JSON data to AWS S3. It only makes sense to compress it, and I have a process working using hadoop's GzipCodec, but there's things that make me nervous about this.
When I look at the type signature of org.apache.spark.rdd.RDD.saveAsTextFile here:
https://spark.apache.org/docs/2.3.0/api/scala/index.html#org.apache.spark.rdd.RDD
the type signature is:
def saveAsTextFile(path: String, codec: Class[_ <: CompressionCodec]): Unit
but when I check the available compression codecs here:
https://spark.apache.org/docs/2.3.0/api/scala/index.html#org.apache.spark.io.CompressionCodec
the parent trait CompressionCodec and subtypes all say:
The wire protocol for a codec is not guaranteed compatible across versions of Spark. This is intended for use as an internal compression utility within a single Spark application
That's no good... but it's fine, because gzip is probably easier to deal with across ecosystems anyway.
The type signature says the codec must be a subtype of CompressionCodec... but I tried the following to save as .gz, and it works fine, even though hadoop's GzipCodec is not <: CompressionCodec.
import org.apache.hadoop.io.compress.GzipCodec
rdd.saveAsTextFile(bucketName, classOf[GzipCodec])
my questions:
- this works, but are there any reasons to not do it this way... or is there a better way?
- is this going to be robust across Spark versions (and elsewhere) unlike the built in compression codecs?