I'm using Java 11 with Spark 3.3.0. Let's say I have a DataFrame with two columns, id (a string identifier) and properties (a string containing a single JSON object, e.g. {"foo":123,"bar":"test"}).
How can I "explode" all the name-value pairs of the properties object into multiple columns, e.g. id, foo, bar.
I've seen solutions using select(). But then I would have to take care to select the existing columns, wouldn't I? Is there a simpler way to simply add additional columns as needed in one command?
I know I can add a single column at a time if I already know what properties to expect, like this:
df = df.withColumn("foo", get_json_object(col("properties"), "$.foo"))
df = df.withColumn("bar", get_json_object(col("properties"), "$.bar"))
Is there a way to add all the columns necessary for the properties JSON object in the properties column in one fell swoop, even without knowing ahead of time what properties will be in each JSON object?