Good evening,
I would like to know, which is the best way to compare two dataframes and return a combination of them? Or if there's even a build-in function inside pandas?
For example, these are my two dataframes:
Dataframe 01:
first_name | age | id | value_a | value_b | value_c
peter | 37 | 19 | 4562 | 78 | 21.5
jane | 32 | 5 | 3832 | 85 | 17.0
michael | 43 | 41 | 2195 | 63 | 44.4
Dataframe 02:
first_name | age | id | value_a | value_b | value_c
sarah | 51 | 2 | 63 | 81 | 4.1
peter | 37 | 19 | 4562 | 81 | 21.5
tom | 22 | 89 | 107 | 14 | 0.0
michael | 43 | 41 | 1838 | 63 | 44.4
As you can see, there are some new entrys troughout the whole dataframe (Dataframe 02) and some of the already existing ones are also listed --> some changes were made in these rows! What I want to achieve is a new(?) dataframe that contains all the new rows, the already existing ones and those who got updated! In this case:
Dataframe New
first_name | age | id | value_a | value_b | value_c
peter | 37 | 19 | 4562 | 81 | 21.5
jane | 32 | 5 | 3832 | 85 | 17.0
michael | 43 | 41 | 1838 | 63 | 44.4
sarah | 51 | 2 | 63 | 81 | 4.1
tom | 22 | 89 | 107 | 14 | 0.0
Notes:
- there is always a column (here: 'id') that can be seen as a non changing key
- the amount of rows may differ
- the amount and names of the colums are always staying the same
- the order of the rows is not important
Thanks for all your help and a great evening!