I have been trying to solve this problem for a while in Python3.
I normally use to extract some information from DOCX documents, by using python-docx library.
from docx.document import Document
from docx import Document
document = Document("test.docx")
for paragraph in document.paragraphs:
for run in paragraph.runs:
print(run.font.name)
#returns None
So, as you can see from the above code, this is a very simple python-docx code to extract some information. I can access some properties such as; font name, size, outline levels, etc.
However, all of these properties are returning None. Because they haven't been explicitly defined.
I have checked StackOverflow for a similar problem and found these.
Extracting word document with styles associated to the content
How to get actual style of text in word document using python docx
In the documentation, it also says, if it returns None, then it's the Default style, that is inherited.
Also tried some XML parsing, but could not reach the desired parameters:
words = document._element.xpath('//w:r')
WORD_NAMESPACE = '{http://schemas.openxmlformats.org/wordprocessingml/2006/main}'
PARA = WORD_NAMESPACE + 'p'
for elem in document.element.getiterator():
if elem.tag == WORD_NAMESPACE + 'p':
for i, child in enumerate(elem.getchildren()):
if child.tag == WORD_NAMESPACE + 'pPr':
...
# No idea how to access, all the styles with which
# tags etc.
How do we extract these default styles too? I would want to extract, the indentation levels, bold, italic, font name, size, etc properties from a DOCX. What could be the alternative ways. I want to solve it in Python3.