I am writing a Python program with the goal of intaking a word document and returning the same document with certain words/phrases highlighted. The program leverages the python-docx and spaCy APIs.
Currently, the program intakes the text from a .docx word document into a Python-docx Doc, converts it into a list of Spacy Docs (one Doc for each paragraph in the docx Doc), and uses the spaCy Matcher to output all of the word/phrase matches. The output is in the form of a list with the string word, it's starting character index in its respective paragraph, and it's ending character index in its respective paragraph. See below:
import spacy
import docx
from spacy.matcher import Matcher
#Create the nlp object
nlp = spacy.load("en_core_web_sm")
#Intake a word doc into a list of nlp docs
docx = docx.Document('file.docx')
docs = []
for paragraph in docx.paragraphs:
docs.append(nlp(paragraph.text))
# Initialize the matcher with the shared vocab
matcher = Matcher(nlp.vocab)
#matcher rule logic omitted
# Call the matcher on each nlp Doc that each represent a paragraph in the python docx Doc.
matches = []
docx_matches = []
#Calls the matcher on each nlp Doc and adds each doc's match outputs into the matches list
for doc in docs:
matches.append(matcher(doc))
#Iterate through the matches found in each nlp Doc, and return them in a format that python docx can understand (by character indices rather than token indices)
for index, match in enumerate(matches):
for match_id, start, end in match:
span = docs[index][start:end]
match_text = span.text
match_start_index = span.start_char
match_end_index = span.end_char
docx_matches.append([match_text, match_start_index, match_end_index])
I'm struggling to now take the output of the matches (the docx_matches list) and properly highlight only those spans of text in the original file.docx document using the start and end character indices. python-docx only has an add_run() function for adding runs to the end of an existing document. Ideally, we would be able to create a run at the start and end index of each match and change the highlight of that run.
I tried leveraging the code here: https://github.com/python-openxml/python-docx/issues/980. However, I receive an error when attempting to:
return Run(r, paragraph)
in the form of "NameError: name 'Run' is not defined". I also tried setting the run "r" in the code from the above link to
r.font.highlight_color = WD_COLOR_INDEX.YELLOW
but received the error AttributeError: 'CT_R' object has no attribute 'font', meaning that "r" isn't of type Run in the first place.
Can anyone help with a different solution or how to fix the ones above? Thank you for your time - I appreciate it a lot! Happy to answer questions/provide more code as necessary.