Splitting PDF with PyPDF2 removes alt text from images

Viewed 79

I need to split a pdf into individual pdfs for each of the pages. Overall, the functionality is working (I get pdfs for each page of the original pdf). The problem is that a screen reader cannot find alt-text or a label for images in the pdfs even though: I've verified that the original pdf has alt text that can be found and for the individual pdfs the screen reader can find the unlabelled image.

Here is a snippet of the code I'm using for PyPDF2 1.26.0:

        with open(filename, "rb") as original_pdf:
            try:
                pdf_reader = PdfFileReader(original_pdf)
                max_digits = len(str(pdf_reader.numPages))
                for i in range(pdf_reader.numPages):
                    individual_pdf_writer = PdfFileWriter()
                    individual_pdf_writer.addPage(pdf_reader.getPage(i))
                    # pad leading zeros for sorting filenames
                    padded_count = str(i + 1).zfill(max_digits)
                    new_filename = f"{self.workspace}/slide-{padded_count}.pdf"
                    with open(new_filename, "wb") as new_pdf_out_stream:
                        individual_pdf_writer.write(new_pdf_out_stream)
            except Exception as e:  # pylint: disable=broad-except
                ...

Has anyone run into this problem? Maybe where PyPDF2 is removing other metadata when splitting? Are there any alternatives you'd recommend?

0 Answers
Related