first question ever here so bear with me if my etiquette is poor. I'm currently working on a project where the goal is to implement a voice assistant using python. We were recommended to use natural language processing to help the assistant parse problems more effectively and I've successfully installed nltk for this. I'm totally new to natural language processing so I've run into some confusion.
Right now my code will take a verbal input to the mic such as:
"what is the weather in Chicago?"
and sucessfully tokenize it, remove the stopwords and tag it as follows:
import nltk # importing the natural language toolkit
from nltk import word_tokenize # this allows us to tokenize a sentence
from nltk.corpus import stopwords # this allows us to filter out stopwords
# Tokenizes the sentence
tokens = word_tokenize(text)
print(tokens)
# Removes stopwords from the sentence
sWords = set(stopwords.words('english'))
cleanTokens = [w for w in tokens if not w in sWords]
print(cleanTokens)
# Tags the sentence
tagged = nltk.pos_tag(cleanTokens)
print(tagged)
# Prints fully processed sentence with tags attatched
print(nltk.ne_chunk(tagged))
Output:
['what', 'is', 'the', 'weather', 'in', 'Chicago']
['weather', 'Chicago']
[('weather', 'NN'), ('Chicago', 'NNP')]
(S weather/NN (GPE Chicago/NNP))
Essentially my problem is that I'm not sure where to go from here. I haven't really found any good examples of how text like this should be used with API's to actually return the weather in Chicago.
Would I be right to simply use if/else statements like in this pseudocode?:
if tagged.contains("weather")
city = searchForCities(tagged)
return city.weatherReport
elif tagged.contains("time) ect...
To summarize, when you have tokenized/tagged nltk text, what's the best way for your code to determine what to do next so that the relevant information is used by the correct library?