This is a pretty straightforward question, but I couldn't find any related posts. I hope I'm not generating duplicates, but here's the issue, I'm building a text classifier, which is derived from a table looking like this:
id text cat1 cat2 target
111 'A mí, no me gusta...' A 1 1
888 'Bueníssimo...' A 2 0
999 'Yo pienso que...' B 2 1
132 'Lo que no me...' C 1 0
555 'Terrible...' C 2 0
.
.
.
One can think of this as the public opinion of your product. I'm trying to put text (text column) and categorical variables (cat1 and cat2 columns) together. Target is 1 if the customer recommends the product, 0 if doesn't.
And then I've just applied these methods:
tfv = TfidfVectorizer(stop_words=stopwords.words("spanish"))
tfv_train = tfv.fit_transform(X_train.text).toarray()
tfv_test = tfv.transform(X_test.text).toarray()
# dimensionality reduction to improve sparse matrix
svd = TruncatedSVD()
trainsvd = svd.fit_transform(tfv_train)
testsvd = svd.transform(tfv_test)
# putting categorical and text-transformed variables together
features_train = np.hstack([train.drop(columns="text").values, trainsvd])
features_test = np.hstack([test.drop(columns="text").values, testsvd])
Since these packages are usually better for English transformations, I wanted to see which words are still in the text after the removal of the stopwords. So I was trying to plot the wordcloud, but I just found out that I don't know how to do this after applying the tfidf transformation. This is what I would normally do:
%matplotlib inline
from wordcloud import WordCloud
words = " ".join([text for text in df.request_text])
word_cloud = WordCloud().generate(words)
plt.figure()
plt.imshow(word_cloud)
Any clues?
Ps.: Sorry about the oversimplification of the data, but I can't post the real one. I hope someone is able to understand and help.