I have a PySpark dataframe having two array columns as seen below:
| col1 | col2 |
|---|---|
| [1, 2, 3] | [1, 4, 3] |
| [5, 4, 3] | [5] |
I want the result to be like this:
| col3 |
|---|
| [1, 0, 1] |
| [1, 0, 0] |
Also, number of items in col1 is always fixed and greater than or equal to number of items in col2.
I know we can do it using udf. But I am looking for an optimised way of doing it using PySpark SQL functions.