We're building an app and we need to extract data from many pdf e-books. They have different formats, font sizes, line spaces, etc...
We're currently configuring the pdfminer.layout.LAParams for each pdf and it's not feasible.
LAParams(line_overlap=0.5, char_margin=2.0, line_margin=0.3, word_margin=0.1, boxes_flow=0.5, detect_vertical=False, all_texts=False)
The biggest issue is that for some books, paragraphs aren't recognized so we end up having only one-line paragraphs.
Is there any known way to automate the configuration for each pdf?