I would like to analyze a dataset from a clinical study using pandas. Patients come at different visits to the clinic and some parameters are measured. I would like to normalize the bloodparameters to the values of the first visit (baseline values), i.e: Normalized = Parameter[Visit X] / Parameter[Visit 1]. The dataset looks roughly like the following example:
import pandas as pd
import numpy as np
rng = np.random.RandomState(0)
df = pd.DataFrame({'Patient': ['A','A','A','B','B','B','C','C','C'],
'Visit': [1,2,3,1,2,3,1,2,3],
'Parameter': rng.randint(0, 100, 9)},
columns = ['Patient', 'Visit', 'Parameter'])
df
Patient Visit Parameter
0 A 1 44
1 A 2 47
2 A 3 64
3 B 1 67
4 B 2 67
5 B 3 9
6 C 1 83
7 C 2 21
8 C 3 36
Now I would like to add a column that includes each parameter normalized to the baseline value, i.e. the value at Visit 1. The simplest thing would be to add a column, which contains only the Visit 1 value for each patient and then simply divide the parameter column by this added column. However I fail to create such a column, which would add the baseline value for each respective patient. But maybe there are also one-line solutions without adding another column.
The result should look like this:
Patient Visit Parameter Normalized
0 A 1 44 1.0
1 A 2 47 1.07
2 A 3 64 1.45
3 B 1 67 1.0
4 B 2 67 1.0
5 B 3 9 0.13
6 C 1 83 1.0
7 C 2 21 0.25
8 C 3 36 0.43