Why one CID number maps to two UTF16 encodings in one CMap File?

Viewed 254

I am debugging a strange problem of parsing PDF text using itext: the character "起" cannot be parsed correctly.

Finally I find that character (with CID code 2510) been mapped to Unicode <e44a>. After reading the document for CID-Keyed Font (https://www.adobe.com/content/dam/acom/en/devnet/font/pdfs/5092.CID_Overview.pdf), I am quite confused why one CID number can be mapped to two different UTF16 encodings?

enter image description here

First occurrence: https://github.com/adobe-type-tools/cmap-resources/blob/master/Adobe-CNS1-7/CMap/UniCNS-UTF16-H#L12636

Second occurrence: https://github.com/adobe-type-tools/cmap-resources/blob/master/Adobe-CNS1-7/CMap/UniCNS-UTF16-H#L12636

And... the second one (the wrong one) wins.

...
<8d77> 2510
...
<e44a> 2510
...

==== The mentioned PDF file ====

PDF URL: CK Hutchison 2018 Annual Report. The 13th character in line 3, page 128.

enter image description here

Result looks like this:

enter image description here

This file works fine with PDFBox.

Env: itext7-core: 7.1.10, macOS 10.13.6

Code:

    InputStream inputStream = getClass().getResourceAsStream("/hk-annual-report/00001.pdf");
    PdfDocument pdfDocument = new PdfDocument(new PdfReader(inputStream));
    PdfPage page = pdfDocument.getPage(128);
    String s = PdfTextExtractor.getTextFromPage(page);
    System.out.println(s);
0 Answers
Related