Accessing alt-text for an image in a PDF using PDFBox

Viewed 435

I'm looking for a way to access alt text for images in a PDF automatically, using PDFBox. I know that you can manually extract the alt text using the other tools, but I'm looking for a way to do this at scale. I found this question addressing this issue, but it seems outdated. How do I extract alt-text for images in a PDF using PDFBox?

Edit: This is what I have so far, but it is still returning null on a pdf that I know has alt text on the figures.

PDFMarkedContentExtractor extractor = new PDFMarkedContentExtractor();
for (int p = 1; p <= document.getNumberOfPages(); ++p) {
        PDPage page = document.getDocumentCatalog().getPages().get(p);
        extractor.processPage(page);
        List<PDMarkedContent> annotations = extractor.getMarkedContents();
        // do some nice output with a header
        String pageStr = String.format("page %d:", p);
        System.out.println(pageStr);
        for (int i = 0; i < pageStr.length(); ++i) {
            System.out.print("-");
        }
        System.out.println();
        for (int i = 0; i < annotations.size(); i++) {
         System.out.println(annotations.get(i).getAlternateDescription());
            }
        System.out.println();
}
0 Answers
Related