I have a PySpark df with this schema:
root
|-- name: string (nullable = true)
|-- products: struct (nullable = true)
| |-- product_hist: map (nullable = true)
| | |-- key: string
| | |-- value: integer (valueContainsNull = true)
| |-- tot_visits: long (nullable = true)
Example of a row:
Mary, {{A -> 2000, B -> 100, C -> 250}, 4}
Given a python dict
my_dict = {'A': 1, 'C': 2}
I'd like to change the keys in the MapType field using the Python dict and filter out any keys that are not in the dict. I'd then get:
Mary, {{1 -> 2000, 2 -> 250}, 4}
What's the best way to do that?