How can I insert and modify specific python-docx runs within an existing doxc Doc initially made up of only 1 run? Given desired run indices as inputs

Viewed 111

I am writing a Python program with the goal of intaking a word document and returning the same document with certain words/phrases highlighted. The program leverages the python-docx and spaCy APIs.

Currently, the program intakes the text from a .docx word document into a Python-docx Doc, converts it into a list of Spacy Docs (one Doc for each paragraph in the docx Doc), and uses the spaCy Matcher to output all of the word/phrase matches. The output is in the form of a list with the string word, it's starting character index in its respective paragraph, and it's ending character index in its respective paragraph. See below:

import spacy
import docx
from spacy.matcher import Matcher

#Create the nlp object
nlp = spacy.load("en_core_web_sm")

#Intake a word doc into a list of nlp docs
docx = docx.Document('file.docx')
docs = []

for paragraph in docx.paragraphs:
    docs.append(nlp(paragraph.text))

# Initialize the matcher with the shared vocab
matcher = Matcher(nlp.vocab)

#matcher rule logic omitted

# Call the matcher on each nlp Doc that each represent a paragraph in the python docx Doc.
matches = []
docx_matches = []

#Calls the matcher on each nlp Doc and adds each doc's match outputs into the matches list
for doc in docs:
    matches.append(matcher(doc))

#Iterate through the matches found in each nlp Doc, and return them in a format that python docx can understand (by character indices rather than token indices)
for index, match in enumerate(matches):
    for match_id, start, end in match:
        span = docs[index][start:end]
        
        match_text = span.text
        match_start_index = span.start_char
        match_end_index = span.end_char

        docx_matches.append([match_text, match_start_index, match_end_index])

I'm struggling to now take the output of the matches (the docx_matches list) and properly highlight only those spans of text in the original file.docx document using the start and end character indices. python-docx only has an add_run() function for adding runs to the end of an existing document. Ideally, we would be able to create a run at the start and end index of each match and change the highlight of that run.

I tried leveraging the code here: https://github.com/python-openxml/python-docx/issues/980. However, I receive an error when attempting to:

return Run(r, paragraph)

in the form of "NameError: name 'Run' is not defined". I also tried setting the run "r" in the code from the above link to

r.font.highlight_color = WD_COLOR_INDEX.YELLOW

but received the error AttributeError: 'CT_R' object has no attribute 'font', meaning that "r" isn't of type Run in the first place.

Can anyone help with a different solution or how to fix the ones above? Thank you for your time - I appreciate it a lot! Happy to answer questions/provide more code as necessary.

0 Answers
Related