Assign twitter to a cluster from hashtag k-Means

Viewed 67

Hi dears I have a problem: i have to do a tweets clustering, but i don't know how i could do it in the right way. My scope is to do a sentiment analysis of the tweets on the back of the hashtag used (if you have a suggestion to improve my work i'll be very happy you will share it with me, thank you). I did all preprocessing work on test, in order to have a clean text to work with. So i extracted the text of my tweets and i trained a word2vec model on them in order to have a vector representation of each word because I would to do a K-Means clustering. Then I asked for the normed vectors obtained from the model of the 1000 most common hashtags, which could give me a label for the tweets. Here is my problem: if I want to give a label to each tweet from the labelled hashtag, how can i do it? Or what I could to to label every tweet from a given hashtag? I post below the code used starting from the word2vec model:

text= df["token_text"]

model = gensim.models.Word2Vec(
    window=3, # they are tweet, so I choose a window of 3 elements
    min_count=5, #min number of token word in each tweet
    workers = 5, #cpu 
)

model.build_vocab(text)
model.train(text, total_examples= model.corpus_count, epochs=30)

Then in an other page:

model= gensim.models.Word2Vec.load("../../2/Modelli/model.bin")#model trained by me
corpus= df.hashtag
model.build_vocab(corpus, update=True)
model.train(corpus, total_examples= model.corpus_count, epochs=30)
word_counts = Counter(itertools.chain.from_iterable(corpus))
words_by_freq = (k for k, v in word_counts.most_common()[0:1000])
word_vectors = model.wv.vectors_for_all(words_by_freq) 
normed_vectors=word_vectors.get_normed_vectors()

k=5
km_model = KMeans(
    n_clusters=k,
    init='k-means++',
    max_iter=35,
    n_init=30,
    verbose=True
)
groups = km_model.fit_predict(normed_vectors)
ordered_centroids = km_model.cluster_centers_.argsort()[:, ::-1]
tags = word_vectors.index_to_key

TAGS_IN_CATEGORY= 5
for idx, centroids in enumerate(ordered_centroids):
    print("Centroid %s:" % idx)
    for centroid_tag in centroids[:TAGS_IN_CATEGORY]:
        print("#%s" % tags[centroid_tag])

Finally I try do a PCA with the clusters. I don't know if all is the right way and I don't how now I can give a label to each tweet based on the labeled hashtag. I'll appreciate every suggestions. Thank you for the attention and patience.

1 Answers

after a lot of thinking and a lot of work, I want to post my solution: it could be useful for someone. Let you free to give me any suggestion. On the back of the idea that the word2vec model trained knows the context of each tweet(I trained my model on the all tweet corpus, then I tuned this model updating it with the hashtag corpus), and because i have the division in cluster of each hashtag, so the perspective of the rate of the hashtag in each cluster, I choose to count every hashtag in each line for each tweet, keep of this hashtag the label given to it through KMeans, calculate the max number of iteration for each label in each row for each tweet and finally give to each tweet the respective label. I expected that there was some tweets without label, and it happened, luckly! So I discard the tweets without labels and finally gave to each tweet the cluster with the maximum sum of the labeled counted for the hashtag of each tweet. I'll show some lines of code. I hope you appreciate.

label=[]    #here take the label for the hashtag in each tweet

for i in range(0, len(df)):
    word_list=[]
    for word in df['hashtag'][i]:
        if word_labeled.get(word) is not None:
            word_list.append(word_labeled.get(word))
    label.append(Counter(word_list).most_common())

df['cluster']= label #assign the labels to my dataset

df= df[~df.cluster.astype(str).str.contains('\[]')]  #delete the twitters without label

df.reset_index(inplace=True) #after some manipulations i reset the index

label_unpack= [] #extract the max in labels

for i in range(0, len(df)):
    label_unpack.append(df['cluster'][i][0][0]) 

df.drop(columns=['cluster'], inplace=True)
df['cluster']= label_unpack #reassign the right labels

Finaly I print the division beetween the cluster for the tweets. Thank you for the attention!

Related