Reading text from pdf with iText7 + C#, text not recognized

Viewed 10719

i want to read data from pdf document. I use iText7:

var src = "<file location>";
var pdfDocument = new PdfDocument(new PdfReader(src));
var strategy = new LocationTextExtractionStrategy();
for (int i = 1; i <= pdfDocument.GetNumberOfPages(); ++i)
{
     var page = pdfDocument.GetPage(i);
     string text = PdfTextExtractor.GetTextFromPage(page, strategy);
     string processed = Encoding.UTF8.GetString(ASCIIEncoding.Convert(Encoding.Default, Encoding.UTF8, Encoding.Default.GetBytes(text)));
}
pdfDocument.Close();

It works, but doesn't recognize letters. All text looks like

"����������\n�������������������������\n�����������������������������������\n

It is in English, so I don't expect any problems with encoding. What is the cause of this issue and how can I fix it?

1 Answers

You don't need the conversion you're doing. Change the code to:

StringBuilder processed = new StringBuilder();

    for (int i = 1; i <= pdfDocument.GetNumberOfPages(); ++i)
    {
         var page = pdfDocument.GetPage(i);
         string text = PdfTextExtractor.GetTextFromPage(page, strategy);
         processed.Append(text);
    }
Related