I have a very large dictionary which stores large numbers of English sentences and their Spanish translations. When given a random English sentence, I intend to use Python's fuzzywuzzy library to find its closest match in the dictionary. My code:
from fuzzywuzzy import process
sentencePairs = {'How are you?':'¿Cómo estás?', 'Good morning!':'¡Buenos días!'}
query= 'How old are you?'
match = process.extractOne(query, sentencePairs.keys())[0]
print(match, sentencePairs[match], sep='\n')
In real life scenario, the sentencePairs dictionary would be very large, with at least one million items stored. So it will take a long time to get the result with fuzzywuzzy, even if python-Levenshtein is installed to provide speedup.
So is there a better way to achieve better performance? My goal is to get the result in less than a few seconds, or even in real time.