Assume, I have the list of stopwords:
STOP = ['under', 'its', 'agreement', 'financed']
For the given dataframe:
lst = ['Kan.-based National', 'Kan.-based National Pizza', 'stock market',
'Pittsburg Kan.-based National Pizza', 'the stock market', 'revolving credit',
'revolving credit agreement', 'its revolving credit agreement', 'under its revolving credit agreement',
'financed under its revolving credit agreement']
df = pd.DataFrame(lst)
which is:
0 Kan.-based National
1 Kan.-based National Pizza
2 stock market
3 Pittsburg Kan.-based National Pizza
4 the stock market
5 revolving credit
6 revolving credit agreement
7 its revolving credit agreement
8 under its revolving credit agreement
9 financed under its revolving credit agreement
I want to obtain:
out = ['Pittsburg Kan.-based National Pizza', 'the stock market', 'revolving credit',
'revolving credit agreement', 'its revolving credit agreement', 'under its revolving credit agreement',
'financed under its revolving credit agreement']
df_out = pd.DataFrame(out)
which is:
0 Pittsburg Kan.-based National Pizza
1 the stock market
2 revolving credit
3 revolving credit agreement
4 its revolving credit agreement
5 under its revolving credit agreement
6 financed under its revolving credit agreement
Note: The order of rows isn't important.
Explanation:
Since, 'Kan.-based National' and 'Kan.-based National Pizza' differ by just one word 'Pizza' and there are no words present from the STOP list, we want to choose the longest span, i.e. 'Kan.-based National Pizza'.
But, 'Pittsburg Kan.-based National Pizza' also differs from 'Kan.-based National Pizza' by just one word 'Pittsburg' and there are no words present from the STOP list, we want to choose the longest span, i.e. 'Pittsburg Kan.-based National Pizza'.
We CANNOT choose 'financed under its revolving credit agreement' as the longest span starting with 'revolving credit' because the words are present in the STOP list. Hence, we will NOT delete it's smaller spans.
Alternatively, on a side track, if the string starts with (a|an|the) and the difference between it's common spans is just one word. For eg - "stock market" and "the stock market", we want to choose the longest span, i.e, "the stock market".
I tried to do:
delete_from_best_constituents = []
for u in best_parse_constituents:
for v in best_parse_constituents:
if u.lower().startswith('the') or v.lower().startswith('the'):
u_part = u.lower().split('the')[-1].strip()
v_part = v.lower().split('the')[-1].strip()
cond1 = all([w.lower() not in STOP for w in u_part.split()])
cond2 = all([w.lower() not in STOP for w in v_part.split()])
if u_part == v.lower() or v_part == u.lower() and cond1 and cond2:
if not u.lower().startswith('the'):
delete_from_best_constituents.append(u)