Spark read files recursively from all sub folders with same name

Viewed 2827

I have a process that is pushing bunch of data to the Blob store every hour and creating the following folder structure inside my storage container as below: /year=16/Month=03/Day=17/Hour=16/mydata.csv /year=16/Month=03/Day=17/Hour=17/mydata.csv

and so on

form inside my Spark context I want to access all the mydata.csv and process them. I figured out that I needed to set the sc.hadoopConfiguration.set("mapreduce.input.fileinputformat.input.dir.recursive","true") so that we can use recursive search like below:

val csvFile2 = sc.textFile("wasb://mycontainer@mystorage.blob.core.windows.net/*/*/*/mydata.csv")

but when I execute the following command to see how many files I have received, it gives me some really large number like below

csvFile2.count res41: Long = 106715282 ideally it should be returning me 24*16=384, also i verified on the container, it only has 384 mydata.csv files, but for some reasons i see it returns 106715282.

can someone please help me understand where I went wrong?

Regards Kiran

1 Answers
Related