Align Multiple Matches in SpaCy nlp into a Pandas Dataframe

Viewed 93

I have written a Code that will search a multiple terms in a Text file which are 'Capex' & Much more in my case

nlp = spacy.load('en_core_web_sm')
pm = PhraseMatcher(nlp.vocab)
tipe = PhraseMatcher(nlp.vocab)
doc = nlp(text)
sents = [sent for sent in doc.sents]


phrases = ['capex', 'capacity expansion', 'Capacity expansion', 'CAPEX', 'Capacity Expansion', 'Capex']
patterns = [nlp(text) for text in phrases]
pm.add('CAPEX ',None,*patterns)
matches = pm(doc)

Then after that when i get where these terms were is a text file , I try to get the sentence where these terms were used. After that i further search for Date , Value & Type of 'CAPEX' further in that Sentence now the issue that i am facing is that when i do so their will have multiple instances where Type of 'CAPEX' which are "Greenfield etc etc" are used multiple times. Although my code only runs the no.of times the matches of the word 'CAPEX' are. Any solution to align all these into one Dataframe

def findmatch(doc,phrases,name):
    p = phrases
    pa = [nlp(text) for text in p]
    name = PhraseMatcher(nlp.vocab)
    name.add('Type',None,*pa)
    results = name(doc)
    return results

def getext(matches):
    for match_id,start,end in matches:
        string_id = nlp.vocab.strings[match_id]
        span = doc[start:end]
        text = span.text
        return text

allcapex = pd.DataFrame( columns = ['Type', 'Value', 'Date','business segment','Location', 'source'])

for ind,match in enumerate(matches):
    for sent in sents:
        if matches[ind][1] < sent.end:
            typematches = findmatch(sent,['Greenfield','greenfield', 'brownfield','Brownfield', 'de-bottlenecking', 'De-bottlenecking'],'Type')
            valuematches = findmatch(sent,['Crore', 'Cr','crore', 'cr'],'Value')
            datematches = findmatch(sent,['2020', '2021','2022', '2023','2024', '2025', 'FY21', 'FY22', 'FY23', 'FY24', 'FY25','FY26'],'Date')
            capextype = getext(typematches)
            capexvalue = getext(valuematches)
            capexdate = getext(datematches)
            allcapex.loc[len(allcapex.index)] = [capextype,capexvalue,capexdate,'','',sent]
            break

print(allcapex)
0 Answers
Related