I want to read a column where the first word in each row is the quarter, and year the survey was taken, as well as the name of the survey. Initially I was trying to rename the survey name where I was keeping the quarter and year constant throughout the column but if I ran this script against files from other quarters then the whole line would not be detected and my script would not work.
My example:
Survey Name
0 Q321 Your Voice - Information Tech
1 Q321 Your Voice - Information Tech
2 Q321 Your Voice - Information Tech
3 Q321 Your Voice - Information Tech
4 Q321 Your Voice - Information Tech
9630 Q321 Your Voice - Business Group
9631 Q321 Your Voice - Business Group
(Q321 = Quarter 3, 2021)
What my code converts it into:
Survey Name
0 Q321 YV - IT
1 Q321 YV - IT
2 Q321 YV - IT
3 Q321 YV - IT
4 Q321 YV - IT
9630 Q321 YV - BG
9631 Q321 YV - BG
The code I use:
print(df.loc[:, "Survey.Name"])
'isolate to column of interest and replace commonly incorrect string with the correct output'
df.loc[df['Survey.Name'].str.contains('Q321 Your Voice - Information Tech'), 'Survey.Name'] = \
'Q321 YV - IT'
df.loc[df['Survey.Name'].str.contains('Q321 Your Voice - Business Group'), 'Survey.Name'] = \
'Q321 YV - BG'
df.loc[df['Survey.Name'].str.contains('Q321 Your Voice - Study Group'), 'Survey.Name'] = \
'Q321 YV - SG'
print(df.loc[:, "Survey.Name"])
But say that I run this script against a file from a different quarter, say Quarter 4, 2021:
Survey Name
0 Q421 Your Voice - Information Tech
1 Q421 Your Voice - Information Tech
2 Q421 Your Voice - Information Tech
3 Q421 Your Voice - Information Tech
4 Q421 Your Voice - Information Tech
9630 Q421 Your Voice - Business Group
9631 Q421 Your Voice - Business Group
I would have to change my script every time a new quarter is used. Is there a way for me to "detect" the first word which luckily happens to be the quarter and year of the survey and include it in the transformed version, whilst also replacing the string that requires to be changed in that column?