I have a larger dataset that is similarly structured to this dataframe (incl. the [ ]):
Day Worker_ID Skills Team_members
0 1 1 [1 3] [1 3]
1 1 2 [2 5] [4 2]
2 1 3 [4 2] [3 1]
3 1 4 [3 3] [2 4]
4 2 1 [2 4] [1 3]
5 2 2 [3 5] [4 2]
6 2 3 [4 3] [3 1]
7 2 4 [2 2] [2 4]
I would like to group my dataframe by the team of the workers so it looks like this (the [ ] are optional]:
Day Team_ID Team_Skills Team_members
0 1 1 [2.5 2.5] [1 3]
1 1 2 [2.5 4] [2 4]
2 2 1 [3 3.5] [1 3]
3 2 2 [2.5 3.5] [2 4]
I would assume the process looks like this:
- Create a .copy() of original dataframe
- Sort the vectors in the team_members-column
- Group by the team-members column & and the day-column
- Delete Worker_ID column
- Create a new Team_ID-column so that every time a new combination of team_members is introduced, a new team number is allocated
- Calculate the mean of skills for each team for that specific day and rename the column
Here is the code, if you want to try it out:
import pandas as pd
data = {'Day': [1, 1, 1, 1, 2, 2, 2, 2],
'Worker_ID': [1, 2, 3, 4, 1, 2, 3, 4],
'Skills': ['[1 3]', '[2 5]', '[4 2]', '[3 3]', '[2 4]', '[3 5]', '[4 3]', '[2 2]'],
'Team_members': ['[1 3]', '[4 2]', '[3 1]', '[2 4]', '[1 3]', '[4 2]', '[3 1]', '[2 4]']}
df = pd.DataFrame(data)