I am struggling to result the similarity between a series of two rows into a new series of columns if and only if another column meets a specific criteria. For example, suppose I have a df with four people, their friend status, and their social preferences.
preference = {'person': ["Sara","Jordan","Amish","Kimmie"],'game_night':[30,10,50,30], 'movies': [10,10,20,10], 'dinner_out': [20,20,30,10] }
near = {'person': ["Sara","Jordan","Amish","Kimmie"], 'friendSara':[0,1,0,0], 'friendJordan': [1,0,1,1], 'friendAmish': [0,1,0,1], 'friendKimmie': [0,1,1,0]}
df = pd.DataFrame(data=preference)
near_df = pd.DataFrame(data=near)
Please challenge me if you feel there is a better way to organize the df or to approach the problem, but I'm looking to, in this example, create a series of new columns named 'simSara', 'simJordan', etc. that fill with the dot(person1_preferences, person2_preferences)/(norm(person1_preferences)*norm(person2_preferences)) between each person's 3 social preferences and the others. For example, the first column added named 'simSara' would have a second row populated by 0.873 (because Jordan and Sara are friends)