PDFBox text extraction ligatures "fi", "fl" problem on Android Studio

Viewed 619

I'm using this https://github.com/TomRoush/PdfBox-Android PDFBox on Android Studio library to extract text from a PDF document. Here's what I'm doing:

File pdf_file = new File(file_path);

to create the file, then

PDDocument document = null;
document = PDDocument.load(pdf_file);

to load the file into a PDDocument object, and then

PDFTextStripper pdfStripper = new PDFTextStripper();
pdfStripper.setStartPage(...);
pdfStripper.setEndPage(...);
String page_text = pdfStripper.getText(document);

to get the text content of the page. The issue is that when there's for example the word "firm" it displays it like "fi rm". It basically puts a space after fi (and I guess fls and other ligatures). I tried reading this Problems with extracting OpenTypeFont text using pdfBox but I don't understand how to fix it. There are no solution details.

Important: As it turns out, in my PDF file, I don't have any ligatures such as fi but I have regular fi and yet, there's space after it. A solution is unclear.

PDF file: https://wetransfer.com/downloads/09e9036dda4a7962ccad32b1cbcd8edc20200506050349/ab4752

2 Answers

The issue is that when there's for example the word "firm" it displays it like "fi rm".

The reason is simple: There is a space after the "fi"!

This is the text drawing instruction drawing the line with the first occurrence of "firm" in your sample file:

 [( )360.3(Mr Dursley was the director of a “)250( )110.3(rm called Grunnings, )]TJ

The byte “ (147) by means of the font encoding is mapped to the glyph name fi and by means of the ToUnicode map of the font to the Unicode character U+fb01, the Latin small ligature fi.

Thus, PDF viewers display the ligature glyph fi and text extractors extract either the Unicode ligature character fi or after expansion the characters f and i.

After that ligature the start point for drawing the next glyph is moved left by 250 units, then a space is drawn, then the next start point is moved left by 110.3 units, and then "rm" is drawn.

Thus, you don't see a gap between "fi" and "rm" in viewers (because the moves left counteract the drawing of the space glyph) but text extractors extract a space character (because it's there).

You can check that this is not a PDFBox quirk, e.g. Adobe Reader with copy&paste extracts that text line as

Mr Dursley was the director of a fi rm called Grunnings,

Just like PDFBox it expands the ligature and extracts the space character.

As mentioned in the comment I had a similar problem once with ligatures. I had to check PDF files for certain strings and was wondering why it didn't work for some. After analysis I found that those files contained ligatures and thus I could not find "Textfield" even though it visually contained it. My solutions was to not only search for textfield but also for textfield - so search two Strings one with and one without ligature.

You said you want to extract text from pdf files. So I would add a post processing step.

  1. Extract the text like you do now
  2. Search all ligatures e.g. "fi " and "fi" and replace it with "fi".

I had documents with no space following a ligature - so I would consider both cases. And cases of word endings (e.g. buffi) should also be considered (might be two spaces then?).

A general word: The topic is not easy as you already researched. This step is called NFKC normalization. In pdfbox 2.X this is done internally (cp. PDFBOX-2384) now but in pdfbox 1.X the TextNormalize.java was doing it.

Upate:

One other possibility you could try is to change the PDFTextStripper.java. There is a method called normalizeWord(...). It converts the single "fi" ligature to "f" and "i". There you could add

//line 1971...
//for PDFs where ligatures are followed by a space (e.g. "fi ve") 
if(word.substring(q+1,q+2).equals(" ")) {
  p = q + 2;
}
else {
  p = q + 1;
}

But I tried it only with pdfbox 2.0.19 (and it seems you are using 1.8.X). The good thing is it is only applied when a ligature was found. However it seems not to be a general solution due to problems with words which end with a ligature. But in your case you should be fine since there consistently seems to be a space after each ligature.

Related