Denormalize data in AWS Glue PySpark

Viewed 22

Good day.

I have json data which I have successfully loaded into a dataframe. It has the following format:

id: abc123
name: some_text
date: some_date
groups: aa,bb,cc,dd

id: def456
name: some_more_text
date: some_other_date
groups: dd,ee

...

I need to expand the data based on the groups field, so that instead of

abc123; some_text;      some_date;       aa,bb,cc,dd
def456; some_more_text; some_other_date; dd,ee

I get:

abc123; some_text;      some_date;       aa
abc123; some_text;      some_date;       bb
abc123; some_text;      some_date;       cc
abc123; some_text;      some_date;       dd
def456; some_more_text; some_other_date; dd
def456; some_more_text; some_other_date; ee

I was started by using the pyspark.sql function "split" which can split the groups into columns, but that is not quite what I need.

Any ideas would be greatly appreciated.

0 Answers
Related