I have a pandas DataFrame looking as follows:
| ID | x | y | z |
| -- | - | --- | --- |
| 1 | 0 | nan | 36 |
| 1 | 1 | 12 | nan |
| 1 | 2 | nan | 38 |
| 1 | 3 | 11 | 37 |
| 2 | 0 | nan | 37 |
| 2 | 1 | nan | 37 |
| 2 | 2 | nan | nan |
| 2 | 3 | nan | nan |
I now want to fill the nan values for each ID in the following way:
- if values for a given ID exist, interpolate between the subsequent values (i.e.: When looking at ID 1: The value of z (in row x1) is what I'm looking for. I have z values for x0, x2 and x3, but the z value corresponding to x1 is missing. I Thus want to find a value for z (in the row of x1) by interpolating between the z values in rows x0 and x2.
- if no values are given for an ID (i.e.: all y values for ID 2 are nan), I want to calculate the median across the entire column (i.e.: across all y values of all IDs) and fill the nan values with that median number.
The result should be a pandas DataFrame in which all nan values are filled by the scheme as explained above. However, I am a beginner with pandas and don't know how to go about this problem to get the full DataFrame.