I have a text file as below
1234_4567_DigitalDoc_XRay-01.pdf
2345_5678_DigitalDoc_CTC-03.png
1234_5684_DigitalDoc_XRay-05.pdf
1234_3345_DigitalDoc_XRay-02.pdf
I am expecting the output as
| catg|sub_catg| doc_name |revision_label|extension|
|1234| 4567|DigitalDoc_XRay-01.pdf| 01 |pdf |
I have created a custom schema
val customSchema = StructType(
StructField("catg", StringType, true)
:: StructField("sub_catg", StringType, true)
:: StructField("doc_name", StringType, true)
:: StructField("revision_label", StringType, true)
:: StructField("extension", StringType, true)
:: Nil
)
I am trying to create a dataframe as
val df = sparkSession.read
.format("csv")
.schema(customSchema)
.option("delimiter", "_")
.load("src/main/resources/data/sample.txt")
df.show()
I am wondering how to break that each line by custom record
I could probably write a java code something of this kind, can someone please help me with the spark. I am new to spark.
String word[] = line.split("_");
String filenName[] = word[3].split("-");
String revision = filenName[1];
word[0]+","+word[1]+","+ word[2]+"_"+word[3]+","+revision.replace(".", " ");