Google Cloud Vision API DOCUMENT_TEXT_DETECTION returning incorrect bounding box

Viewed 1237

I'm using the "DOCUMENT_TEXT_DETECTION" option from the Google Cloud Vision API.

It seems that it's returning correct text value, but incorrect coordinates bounding box.

Why this problem occurred?

Thank you.

raw picture

enter image description here

draw bounding box picture

enter image description here

returning json


appendix

draw bounding box words and overall

enter image description here

2 Answers

The DOCUMENT_TEXT_DETECTION is meant for dense text, I recommend to use the TEXT_DETECTION for that image.

I use DOCUMENT_TEXT_DETECTION model and I have the same problem.

The symbol level bounding boxes are very offset, overlapping other symbols. And this even though the OCR did a good job and was able to find a matching character... See attached picture for illustration. (OCR result is perfect in this straightforward situation):

enter image description here

I notice this model have become legacy as per https://cloud.google.com/vision/docs/release-notes#May_15_2020, maybe the replacement does a better job at this.

Related