import pandas as pd
import numpy as np
df = pd.DataFrame([
['Quality Engineer','Financial Services'],
['Progammer',np.nan],
['Quality Engineer',np.nan],
['Progammer',"IT"],
['General manager',np.nan]],
columns=['job_title','job_industry'])
with pd.option_context('mode.use_inf_as_null', True):
df = df.sort_values('job_industry', ascending=False, na_position='last')
df["job_industry"].loc[(df['job_title'] == "General manager") & (df['job_industry'].isnull())] = "Manufacturing"
df['job_industry'] = df.groupby('job_title')['job_industry'].fillna(method="ffill")
df['job_industry'].isnull(), this will verify job_industry column is null or not.
The following code will sort by null value in descending order of by the column job_industry,because if nan value appears before, the initial value of nan will not replace.
with pd.option_context('mode.use_inf_as_null', True):
df = df.sort_values('job_industry', ascending=False, na_position='last')
if your prefer ordering to the output, you can try,df.sort_index()
O/P
+----+------------------+-------------------------------------------------------+
| | job_title | job_industry |
|----+------------------+-------------------------------------------------------|
| 0 | Quality Engineer | Financial Services |
| 1 | Progammer | IT |
| 2 | Quality Engineer | Financial Services |
| 3 | Progammer | IT |
| 4 | General manager | Manufacturing |
+----+------------------+-------------------------------------------------------+