What I have is two DataFrames, each representing a probability distribution, but each stored one entry per row. For example one is df1:
item_id | probability
--------|---------------
item1 | 0.1
item2 | 0.2
item3 | 0.7
and another, let's call it df2:
item_id | probability
--------|---------------
item2 | 0.3
item3 | 0.5
item4 | 0.2
and please notice that the item space for these two are different. But this is ok because what it means is that df1 has zero probability for item4 and df2 has zero probability for item1. What I'd like is code without heavy use of custom UDFs that basically produces a DataFrame that, given some alpha double value, blends those two distribution. I can write this with custom UDFs but I'm wondering if there's some pure Spark SQL based code that does this with only built-in functions.
item_id | probability
--------|---------------
item1 | 0.1 * alpha + 0.0 * (1 - alpha)
item2 | 0.2 * alpha + 0.3 * (1 - alpha)
item3 | 0.7 * alpha + 0.5 * (1 - alpha)
item4 | 0.0 * alpha + 0.2 * (1 - alpha)