What is the correct way to use google vision for OCR

Viewed 540

Hope you're all fine.

For the past few days, I've been spending some time with google vision for a work project. I'm quiet happy with the results but there are few things I can't figure out. Here it is:

I'm trying to use Google Vision API to read information out of a Tyre picture, this one for instance: enter image description here

This is the list of features I'm using to call the API:

const features = [
  {
    "maxResults": 50,
    "type": "LOGO_DETECTION"
  },
  {
    "maxResults": 100,
    "type": "DOCUMENT_TEXT_DETECTION"
  }];

And my results are the following:

description: 'GOOD YEAR\n' +
        'POSTER\n' +
        'RADIAL\n' +
        'YUDELESS\n' +
        'EXTRA LOAD\n' +
        'CSFY\n' +
        'MADE IN GERMANY\n' +
        'ROTATION\n' +
        'II SGR\n' +
        '(ED\n' +
        'MINT\n' +
        'M66 Lage\n' +
        'VEU 900?\n'

I'm happy with this, but I'm lacking few information that I know the API can detect.

Case 1: When I crop a part of the picture and use the exact same API and parameters enter image description here I get the following results:

{
      locale: 'und',
      description: '225 55R16 99W\n',
      boundingPoly: [Object]

And, case 2, even when I'm using the online google vision try it service I'm getting some results for the digits enter image description here

So at the end, I'm looking for the maximum information out of a picture, even if I need to sort it out after.

Ideas, answers, tips, I take everything.

Cheers, Ivan

1 Answers

There is no general answer as to what is the best way to use Cloud Vision. It's powered by Machine Learning models and results depend on many factors like zoom, quality of the picture and method.

As you can see Cloud Vision API - How To Guides you have many specific functions.

  • OCR
  • Faces - detects multiple faces within an image along with the associated key facial attributes such as emotional
  • Image properties - detects general attributes of the image, such as dominant color.
  • Logos - popular product logos within an image.

and a few other features. Those features are using different algorithms to recognize specific things like text or logos, etc.

In your example you have a tire with the GoodYear logo, which has the name of the company. However if you would use Logo Detection on just a logo without anything it will return the name of the company (database of logos is maintained by google). For example logo of the Nike (Nike Logo URL) it will return name of the company.

Also quality of results depends on the zoom of the image. If text is too small it might not be recognized by algorithms. That's why you have differences when you have used the whole tire picture and zoomed part of it.

In general use TEXT_DETECTION is used for recognizing text in the picture and DOCUMENT_TEXT_DETECTION is used for extracting text from an image, but the response is optimized for dense text and documents.

Even TEXT_DETECTION and DOCUMENT_TEXT_DETECTION doing the same thing, they are using different algorithms for better results from picture recognizing text (TEXT_DETECTION) and from documents(DOCUMENT_TEXT_DETECTION).

To sum up, Cloud Vision has many features which are using different algorithms to fulfill specific needs like getting logos, detecting faces or text.

I hope it gives you a better understanding of Cloud Vision.

Related