Read Text from Image (OCR) in C# with IronOCR Tesseract

Viewed 2173

According this Link I installed the IronOcr package and I try the follow code.

using IronOcr;
var Result = new IronTesseract().Read(path);
string currentSubText = Result.Text;
textBox1.Text += currentSubText + Environment.NewLine + Environment.NewLine;

I tested it with six pictures:

Picture 1

Picture 2

Picture 3

Picture 4

I could just upload four pictures.

Actually it looks good. There are just a few mistakes with some special German language characters (äöü)

Result 1:5

I google and found it is possible to use a language package in OCR. I try it with the follow code.

var Ocr = new IronTesseract();

//Ocr.Language = OcrLanguage.German;
Ocr.Language = OcrLanguage.GermanBest;

using (var Input = new OcrInput(path))
{
    var Result = Ocr.Read(Input);
    string currentSubText = Result.Text;
    textBox1.Text += currentSubText + Environment.NewLine + Environment.NewLine;
}

Unfortunately the result is very, very bad.

Result 2:6

Can someone help me here?

Thanks and best regards

1 Answers

Did you try using the built in inversion color filter?

All OCR tends to work best for me with black text on white. I use this code based on code found in the IronOCR documentation:

https://ironsoftware.com/csharp/ocr/examples/ocr-image-filters-for-net-tesseract/

Simplified source code:

using IronOcr;
 
var Ocr = new IronTesseract();
Ocr.Language = OcrLanguage.GermanBest;
using (var Input = new OcrInput(@"image.png"))
{
    
    //Input.EnhanceResolution(300);
    Input.Invert();
    
    
   
    /*
    // Optional: Export modified images so you can view them.
    foreach(var page in  Input.Pages){
          page.SaveAsImage("filtered.bmp")
    }
    */
   
 
    var Result = Ocr.Read(Input);
    Console.WriteLine(Result.Text);
}

MSDN style docs: https://ironsoftware.com/csharp/ocr/object-reference/api/IronOcr.OcrInput.html#IronOcr_OcrInput_Invert_System_Boolean_

Related