If I didn't get you wrong, based on your definition in the comments of the main post, I've figured out a way that will do the job.
First, according to you, the data will look like this:
data = {'text' : ['Barack Obama was president of the United States in 2008.'],
'annotation' : ['MWE_type 0 12 MWE_type 34 47']}
We will maintain a final_list which is basically a list of list, where is inner list will be the output for each row.
We can iterate over each row by df.iterrows() and extract the result for each row from row['text'] and using row['annotation'].
for index, row in df.iterrows():
We can extract the pair of indexes through the use of regular expression:
re.findall(r'\d+ \d+', row['annotation'])
We can iterate over this list of index pairs and append the corresponding substring to our row based result list.
for indexes in index_list:
start, end = map(int, indexes.split())
result.append(row['text'][start:end])
At the end of iterating a row, we can append the row based result list to the final_list:
final_list.append(result)
Finally, assign the final_list to df['result']:
df['result'] = final_list
The whole program is as below:
import pandas as pd
import re
data = {'text' : ['Barack Obama was president of the United States in 2008.'],
'annotation' : ['MWE_type 0 12 MWE_type 34 47']}
df = pd.DataFrame(data)
final_list = []
for index, row in df.iterrows():
result = []
index_list = re.findall(r'\d+ \d+', row['annotation'])
for indexes in index_list:
start, end = map(int, indexes.split())
result.append(row['text'][start:end])
final_list.append(result)
df['result'] = final_list
print(df)
And you'll get:
text ... result
0 Barack Obama was president of the United State... ... [Barack Obama, United States]