How to highlight (only) word errors using difflib?

Viewed 246

I'm trying to compare the output of a speech-to-text API with a ground truth transcription. What I'd like to do is capitalize the words in the ground truth which the speech-to-text API either missed or misinterpreted.

For Example:

Truth: The quick brown fox jumps over the lazy dog.

Speech-to-text Output: the quick brown box jumps over the dog

Desired Result: The quick brown FOX jumps over the LAZY dog.

My initial instinct was to remove the capitalization and punctuation from the ground truth and use difflib. This gets me an accurate diff, but I'm having trouble mapping the output back to positions in the original text. I would like to keep the ground truth capitalization and punctuation to display the results, even if I'm only interested in word errors.

Is there any way to express difflib output as word-level changes on an original text?

3 Answers

I would also like to suggest a solution using difflib but I'd prefer using RegEx for word detection since it will be more precise and more tolerant to weird characters and other issues.

I've added some weird text to your original strings to show what I mean:

import re
import difflib

truth = 'The quick! brown - fox jumps, over the lazy dog.'
speech = 'the quick... brown box jumps. over the dog'

truth = re.findall(r"[\w']+", truth.lower())
speech = re.findall(r"[\w']+", speech.lower())

for d in difflib.ndiff(truth, speech):
    print(d)

Output

  the
  quick
  brown
- fox
+ box
  jumps
  over
  the
- lazy
  dog

Another possible output:

diff = difflib.unified_diff(truth, speech)
print(''.join(diff))

Output

---
+++
@@ -1,9 +1,8 @@
 the quick brown-fox+box jumps over the-lazy dog

Why not just split the sentence into words then use difflib on those?

import difflib

truth = 'The quick brown fox jumps over the lazy dog.'.lower().strip(
    '.').split()

speech = 'the quick brown box jumps over the dog'.lower().strip('.').split()

for d in difflib.ndiff(truth, speech):
    print(d)

So I think I've solved the problem. I realised that difflib's "contextdiff" provides indices of lines that have changes in them. To get the indices for the "ground truth" text, I remove the capitalization / punctuation, split the text into individual words, and then do the following:


altered_word_indices = []
diff = difflib.context_diff(transformed_ground_truth, transformed_hypothesis, n=0)
for line in diff:
  if line.startswith('*** ') and line.endswith(' ****\n'):
    line = line.replace(' ', '').replace('\n', '').replace('*', '')
    if ',' in line:
      split_line = line.split(',')
      for i in range(0, (int(split_line[1]) - int(split_line[0])) + 1):
        altered_word_indices.append((int(split_line[0]) + i) - 1)
    else:
      altered_word_indices.append(int(line) - 1)

Following this, I print it out with the changed words capitalized:

split_ground_truth = ground_truth.split(' ')
for i in range(0, len(split_ground_truth)):
    if i in altered_word_indices:
        print(split_ground_truth[i].upper(), end=' ')
    else:
        print(split_ground_truth[i], end=' ')

This allows me to print out "The quick brown FOX jumps over the LAZY dog." (capitalization / punctuation included) instead of "the quick brown FOX jumps over the LAZY dog".

This is...not a super elegant solution, and it's subject to testing, cleanup, error handling, etc. But it seems like a decent start and is potentially useful for someone else running into the same problem. I'll leave this question open for a few days in case someone comes up with a less gross way of getting the same result.

Related