I have two dataframe, one is large, the other is small:
val small_df = sc.parallelize(List(("Alice", 15), ("Bob", 20)).toDF("name", "age")
val large_df = sc.parallelize(("Bob", 40), ("SomeOne", 50) , ... ).toDF("name", "age")
I want to add up these two dataframe but only those with the key in my small table, that is, I want my result to be like:
List(("Alice", 15), ("Bob", 60))
My first attempt is try to do union and reduceByKey, but I can't seem to find a way to union two tables and keep those rows with keys in the smaller one only.
Is there a way to do something like "left union" or other way to approach my answer?