Let's say I have the following dataframe schema:
+-------+-------+
| body | rules |
+-------+-------+
I have a udf that takes in the body column and rule-list column for each row, and parses and evaluated the conditions for the rules based on the row (and returns a list of booleans whether it each rule matches or not). Right now, every single row in the DF has a copy of these rules because I don't know any other way to pass in these rules to the UDF. This feels very redundant and wasteful to me.
The rules are joined onto the row based on some join conditions, so each row doesn't have the exact same data but there is still a lot of redundancy (each rule is probably listed ~5000 redundant times across 1 million rows). I would prefer to join the ruleIds onto each row instead and pass a map(ruleId -> rule) into the udf. This map may be somewhat large though so however it is passed in would have to be able to handle that (ideally it would be some sort of shared variable stored at the partition level)