I am currently using XUbuntu 16.04, Apache Spark 2.1.1, IntelliJ, and Scala 2.11.8
I am trying to load some simple text data in CSV format into an Apache Spark Dataset but instead of using regular text files, I am dumping data into a named pipe and then I want to read that data directly into the DataSet. It works perfectly if the data is a regular file, but the exact same data doesn’t work if it is coming from a named pipe. My Scala code is real simple and follows:
import org.apache.spark.sql.SparkSession
object PipeTest {
def main(args: Array[String]): Unit = {
val spark = SparkSession
.builder()
.appName("PipeTest")
.master("local")
.getOrCreate()
// Read data in from a text file and input to a DataSet
var dataFromTxt = spark.read.csv("csvData.txt")
dataFromTxt.show()
// Read data in from a pipe and input to a DataSet
var dataFromPipe = spark.read.csv("csvData.pipe")
dataFromPipe.show()
}
}
The 1st code section loads the csv data from a regular file and works fine. The 2nd code section fails with the following error:
Exception in thread “main” java.io.IOException: Error accessing file: /home/andersonlab/test/csvData.pipe
Anyone know how you would go about using named pipes with Spark Datasets and getting something like the above to work?