I'm looking for a way to access alt text for images in a PDF automatically, using PDFBox. I know that you can manually extract the alt text using the other tools, but I'm looking for a way to do this at scale. I found this question addressing this issue, but it seems outdated. How do I extract alt-text for images in a PDF using PDFBox?
Edit: This is what I have so far, but it is still returning null on a pdf that I know has alt text on the figures.
PDFMarkedContentExtractor extractor = new PDFMarkedContentExtractor();
for (int p = 1; p <= document.getNumberOfPages(); ++p) {
PDPage page = document.getDocumentCatalog().getPages().get(p);
extractor.processPage(page);
List<PDMarkedContent> annotations = extractor.getMarkedContents();
// do some nice output with a header
String pageStr = String.format("page %d:", p);
System.out.println(pageStr);
for (int i = 0; i < pageStr.length(); ++i) {
System.out.print("-");
}
System.out.println();
for (int i = 0; i < annotations.size(); i++) {
System.out.println(annotations.get(i).getAlternateDescription());
}
System.out.println();
}