I have a large JSON file like the following:
[(3, (2, 'Child')), (2, (1, 'Parent')), (1, (None, 'Root'))]
where the key of each element is a unique index for that element and 1st element in the value pair signifies the index of its parent element.
Now, the ultimate goal is to convert this JSON file into the following:
[(3, (2, 'Child Parent Root')), (2, (1, 'Parent Root')), (1, (None, 'Root'))]
where the 2nd element in the value pair for each item will be modified such that it has the concatenation of all the values up to its root ancestor.
The no. of levels is not fixed and can be up to 256. I know I can solve this problem by creating a tree DS and traversing it but the problem is the JSON file is huge (almost 180M items in the list).
Any idea on how can I achieve this efficiently? Suggestions involving Apache Spark would be fine as well.