I am debugging a strange problem of parsing PDF text using itext: the character "起" cannot be parsed correctly.
Finally I find that character (with CID code 2510) been mapped to Unicode <e44a>. After reading the document for CID-Keyed Font (https://www.adobe.com/content/dam/acom/en/devnet/font/pdfs/5092.CID_Overview.pdf), I am quite confused why one CID number can be mapped to two different UTF16 encodings?
First occurrence: https://github.com/adobe-type-tools/cmap-resources/blob/master/Adobe-CNS1-7/CMap/UniCNS-UTF16-H#L12636
Second occurrence: https://github.com/adobe-type-tools/cmap-resources/blob/master/Adobe-CNS1-7/CMap/UniCNS-UTF16-H#L12636
And... the second one (the wrong one) wins.
...
<8d77> 2510
...
<e44a> 2510
...
==== The mentioned PDF file ====
PDF URL: CK Hutchison 2018 Annual Report. The 13th character in line 3, page 128.
Result looks like this:
This file works fine with PDFBox.
Env: itext7-core: 7.1.10, macOS 10.13.6
Code:
InputStream inputStream = getClass().getResourceAsStream("/hk-annual-report/00001.pdf");
PdfDocument pdfDocument = new PdfDocument(new PdfReader(inputStream));
PdfPage page = pdfDocument.getPage(128);
String s = PdfTextExtractor.getTextFromPage(page);
System.out.println(s);


