PySpark DataFrame: Custom Explode Function

Viewed 2582

How to implement a custom explode function using udfs, so we can have extra information on items? For example, along with items, I want to have items' indices.

The part I do not know how to do is when a udf returns multiple values and we should place those values as separate rows.

2 Answers

In Spark v. 2.1+, there is pyspark.sql.functions.posexplode() which will explode the array and provide the index:

Using the same example as @Mariusz:

df.show()
#+---------+
#|    array|
#+---------+
#|[a, b, c]|
#|   [d, e]|
#+---------+

df.select(f.posexplode('array')).show()
#+---+---+
#|pos|col|
#+---+---+
#|  0|  a|
#|  1|  b|
#|  2|  c|
#|  0|  d|
#|  1|  e|
#+---+---+
Related