How can I use KNeighbors predictions to choose the right text position in Google Vision OCR?

Viewed 39

I'm using Google Vision OCR for blood results. Users can take picture of their blood results and my job is to extract meaningful information.

For example, in the picture below, I might want to extract the number related to the TSH (here: 0.722).

Example:

enter image description here

But the pictures I get are never that straight, they are often crooked. And I cannot get proper "line by line" OCR from Google. All I have is the coordinates of all words/numbers.

To get the right number associated to TSH (for example) I gathered over 300 blood results. I've collected 3 types of informations on each of these:

  • first, the straight distance between the position of the word TSH and the position of every number on the page.
  • second, the orthogonal distance from the position of the word TSH and the position of every number on the page.
  • and finally, a boolean : 0 if the number is NOT the TSH value, 1 if it is.

This constitutes my dataset. I might look like this:

22.17605594351686,22.156573116691284,0
40.390995176389076,4.948301329394384,0
42.816362012232894,5.169867060561296,0
39.66602473983306,2.4372230428360426,0
58.589347000389644,2.141802067946825,0
65.27678257473366,2.141802067946825,0
39.34331591288292,0.07385524372230544,1
58.42795056763215,0.4431314623338253,0
62.14532216492161,0.4431314623338253,0
39.35232078475228,2.3633677991137336,0
58.98873313300486,2.8064992614475623,0
61.65003480785281,2.8064992614475623,0

Next step: I train a model using scikit-learn and KNeighbors Classifier.

Here's my code so far:

import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score


def run_model(data):
    cols = [col for col in data.columns if col not in ['winner']]
    target = data['winner']
    data = data[cols]

    data_train, data_test, target_train, target_test = train_test_split(data, target, test_size=0.30, random_state=10)

    # create object of the Classifier
    neigh = KNeighborsClassifier(n_neighbors=3)
    # Train the algorithm
    neigh.fit(data_train, target_train)
    # predict the response
    pred = neigh.predict(data_test)

    # evaluate accuracy
    return accuracy_score(target_test, pred)


def main():
    items = ['TSH']
    for item in items:
        data = pd.read_csv(f'./data/{item}.csv')
        result = run_model(data)
        print(f"KNeighbors accuracy score for {item}: {result}")

The goal here was to predict if a certain number on the page (based on the distance and the orthogonal distance) was the right number (1) or not (0). And, according to the accuracy score, it works. I'm around 99.5, which is great.

How can I use these information to implement a formula/equation to choose the right number on new blood results ?

0 Answers
Related