We have a dataframe with three different columns, like shown in the example above (df). The goal of this task is to replace the first element of the column 2 by a np.nan, everytime the letter in the column 1 changes. Since the database under study is very big, it cannot be used a for loop. Also every solution that involves a shift is excluded because it is too slow.
I believe the easiest way is to use the groupby and the head method, however I don't know how to replace in the original dataframe.
Examples:
df = pd.DataFrame([['A','Z',1.11],['B','Z',2.1],['C','Z',3.1],['D', 'X', 2.1], ['E','X',4.3],['E', 'X', 2.1], ['F','X',4.3]])
to select the elements that we want to change, we can do the following:
df.groupby(by=1).head(1)[2] = np.nan
However in the original dataframe nothing changes.
The goal is to obtain the following:
Edit:
Based on comments, we won't df[1] returning to a group already seen, e.g. ['Z', 'Z', 'X', 'Z'] is not possible.

