How to write Scala tests for component "Write data to HDFS directory"

Viewed 195

I have a simple code which writes data to hdfs in csv and parquet format, How can I write scala tests here which can test the below component. I can't actually write the data to hdfs(in the tests) as code is running in jules pipeline. Any suggestions will be helpful

   df
   .write.format("com.databricks.spark.csv")
   .option("header", "true")
   .mode("append")
   .save(hdfspath)
1 Answers

You can write a sample data with your schema to a local path, read it using spark and compare the expected and actual output.

Here is an example with ScalaTest:

import org.scalatest.FunSuite
import org.scalatest.Matchers
import org.apache.spark.sql.functions.input_file_name

case class RecordSchema(id: Int, value: String) // define here your real schema

class WriteTest extends FunSuite with Matchers {

    test("test data was written properly") {
        import spark.implicits._
        val path = "local/path/dir"
        val expectedData = List(RecordSchema(1, "dummyValue1"), RecordSchema(2, "dummyValue2"))
        expectedData.toDF
            .write.format("com.databricks.spark.csv")
            .option("header", "true")
            .mode("append")
            .save(path)
        val actualData = spark.read.format("com.databricks.spark.csv")
.load(path)
        
        // test that the data was written as expected
        actualData.as[RecordSchema].collect should contain theSameElementsAs expectedData
    
    }
}

This is just an example, you can encapsulate the write component into a separate method (in order to test it as a component and not copy its code). Please pay attention to write the data in your test to a new path (or delete the content of the path beforehand in your test) because otherwise, since it writes with append mode, the logic of this test won't work.

Related