I have a pandas data frame with the below structure
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 annotation 10237 non-null object
1 note_sentence 10237 non-null object
2 listofsentence 10237 non-null object
The last column 'listofsentence' is a list initialised with 'O's same as the number of words in the note_sentence column. Now I want to match the individual strings present inside the annotation column with the long string in the 'note_sentence' column and wherever it gets the matching word I want to update the 'listofsentence' column value from 'O' to 'I'. For example, in the below sample record the value of 'listofsentence' should be updated to [O, I, I, I, O, I, I, I, O, O, O] from its default state.
I used the below code which is able to return the starting and ending indexes of the match but I want it at the word level.
def find_index(string, sentence):
for match in re.finditer(string, sentence):
print (match.start(), match.end())
How can I do this?
