I have a JSON file. Every line is an array of objects. For instance, this is the first line:
[{"id":"JklsdgNkl543", "field1":"value1", "field2":{"nestedField1":"1", "nestedField2":2}},
{"id":"rweiuTH2325d", "field1":"smthng", "field2":{"nestedField1":"6", "nestedField2":8}},
...,
]
So, the file contains hundreds of thousands of these lines, every line contains hundreds of objects in array. And these objects have a lot of columns.
The problem is that there might be a nested extra column in, for example, one or two objects in array.
[
...,
{"id":"QwerTy1", "field1":"smthng", "field2":{"nestedField1":"0", "EXTRA_FIELD": "EXTRA VALUE", "nestedField2":8}},
...,
]
And if I use spark.read.json(), the inferred schema is being generalized for the entire array. That means, if only 1 or 2 objects (out of 1_000) have extra field then the schema will infer this field for the whole file. And this is not what I want.
Is there a way I can infer a separate and exact schema for every object in array? Or, maybe if not for every object but at least for every line, not the whole file?
And a bit off-topic. Maybe there is no need in spark at all, if I just want to save a JSON-file in hdfs as a parquet-file?