Get lines and bounding box coordinates from Google Vision API OCR in Python

Viewed 1094

My goal is to generate a table with a row for each line and columns for the following:

  • Content (textual content for the line)
  • X (X coordinate for the top left of the bounding-box)
  • Y (Y coordinate for the top left of the bounding-box)
  • H (height of the line’s bounding box)
  • W (width of the line’s bounding box)

The Google Vision API for OCR does not group content by lines, only by paragraphs. The API does return data at the symbol level that indicates when a symbol is followed by a space or a line break. Here’s an example of the symbol data for a letter 'E' that is followed by a space:

{
  'property': {
    'detectedLanguages': [{
      'languageCode': 'it'
    }],
    'detectedBreak': {
      'type': 'SPACE'
    }
  },
  'boundingBox': {
    'vertices': [{
      'x': 197,
      'y': 56
    }, {
      'x': 268,
      'y': 59
    }, {
      'x': 263,
      'y': 167
    }, {
      'x': 192,
      'y': 164
    }]
  },
  'text': 'E',
  'confidence': 0.9900000095367432
}

I think it’s possible to loop through a list of symbols and group them into lines based on whether or not they have particular values for the detectedBreak attribute. Using the API documentation for break types, it seems that values of 'EOL_SURE_SPACE', 'HYPHEN', 'LINE_BREAK' indicate line breaks and values of 'UNKNOWN', 'SPACE', 'SURE_SPACE' should indicate spaces. As for the bounding box coordinates, one could collect the coordinates for upper left of the first symbol in a line (vertices[0]) and the bottom right of the last symbol in a line (vertices[2]) to construct a rough bounding box for the line.

My question is this, can anybody suggest code for aggregating the symbol data into line data so that it could be output into a table like that one at the beginning of the question? The following is the beginning of my script but I’m just not sure how to aggregate the data.

from google.cloud import vision
from google.cloud.vision import types
import json
from google.protobuf.json_format import MessageToJson

# other code…

with io.open(file, 'rb') as image_file:
        content = image_file.read()
        image = types.Image(content=content)
        response = client.document_text_detection(image=image)
        document = response.full_text_annotation

        serialized = MessageToJson(document)

        data = json.loads(serialized)

        symbols = []
        for page in data['pages']:
            for block in page['blocks']:
                for paragraph in block['paragraphs']:
                    for word in paragraph['words']:
                        for symbol in word['symbols']:
                            symbols.append(symbol)

        break_values = ['EOL_SURE_SPACE', 'HYPHEN', 'LINE_BREAK']
        space_values = ['UNKNOWN', 'SPACE', 'SURE_SPACE']

        lines = [] # or should lines be a dictionary?
        for symbol in symbols:
            # need help here

Thanks in advance.

1 Answers

I came across this question while working with the same API. I had to solve the same problem, and here's the code I ended up with:

def get_lines(response):
    annotation = response.full_text_annotation

    def has_line_break(symbol):
        line_break = vision.types.TextAnnotation.DetectedBreak.BreakType.LINE_BREAK
        return symbol.property and symbol.property.detected_break.type_ == line_break

    lines = []
    line = []
    for page in annotation.pages:
        for block in page.blocks:
            for paragraph in block.paragraphs:
                for word in paragraph.words:
                    for symbol in word.symbols:
                        line.append(symbol.text)
                        if has_line_break(symbol):
                            lines.append(line)
                            line = []
    return lines

This returns each line as a list of one character strings. in my use case, I ended up needing to convert the symbols to my own datatype with a line number field, but it's the same basic idea

Probably you figured something out by now but, maybe this will be useful for someone else

Related