I have a PDF document which I am currently parsing using Tika-Python. I would like to split the document into paragraphs.
My idea is to split the document into paragraphs and then create a list of paragraphs using the isspace() function
I also tried splitting using \n\n however nothing works.
This is my current code:
file_data = (parser.from_file('/Users/graziellademartino/Desktop/UNIBA/Research Project/UK cases/file1.pdf'))
file_data_content = file_data['content']
paragraph = ''
for line in file_data_content:
if line.isspace():
if paragraph:
yield paragraph
paragraph = ''
else:
continue
else:
paragraph += ' ' + line.strip()
yield paragraph