I have events coming as json which is converted to data frame to do processing. These events are deeply nested, recently I am receiving events which have field names having decimal numbers e.g 0.5, 0.75. This is causing my application to break as spark perceives them as struct field. In below case 0.5 is perceived as, field having name as 0 and type to be struct which is wrong and causes my application to fail.
{
"id": "asbuibdfknsdifjofiohfisdj1212423",
"object_response": {
"response_timestamp": "2020-04-04T14:26:00",
"addon_pricing_packages": {
"1": 10,
"0.5": 5,
"0.25": 2.5,
"0.75": 7.25
},
"object_id": "id_12321ijisdansdiu"
}
}
I can deploy recursive functionality which would loop within the fields and drill down to the lowest level and replace the field name. But I believe that will not be optimized as the events are deeply nested and could probably slow down my application. Any efficient solution or thought process. Also the fields are dynamic in nature and one thought process was to create a schema and then cast but as the fields are dynamic so can't use that. Any thoughts/pointer on this will be great.