I m trying to create clusters from data based on the string value of each row. I m using the R langage. What I m calling a "cluster" is a big thematic (= family) that can define each keywords. I imagine something autogenearated based on the keyword, maybe by using lemmatization or ngram.
For example both keywords "cloud services" and "the cloud service" should be in the "service" cluster.
Here is my input vector:
keywords_df <- c("cloud storage", "cloud computing", "google cloud storage", "the cloud service",
"free cloud storage", "what is cloud computing", "best cloud storage","cloud computing definition",
"amazon cloud services", "cloud service providers", "cloud services", "google cloud computing", "cloud computing services", "benefits of cloud computing")
Here is the expected output dataframe:
| Keyword | Thematic |
|---------------------------|:---------:|
|cloud storage |storage |
|cloud computing |computing|
|google cloud storage |storage |
|the cloud service |service |
|free cloud storage |storage |
|what is cloud computing |computing|
|best cloud storage |storage |
|cloud computing definition |computing|
|amazon cloud service |service |
|cloud service providers |services |
|cloud service |service |
|google cloud computing |computing|
|cloud computing services |service |
|benefits of cloud computing|computing|
The goal is to clean up the data in the "keyword" column and auto extract a kind of lemm or ngram.
Here is what I have done for now :
Create the "Thematic" column based on keyword column:
keywords_df <- mutate(keywords_df,Thematic=Keyword) keywords_df$Thematic <- as.character(keywords_df$Thematic)Remove Stopwords:
stopwords_list<-(c("cloud")) #Remove the main word stopwords <- stopwords(kind = "en") stopwords <- append(stopwords,stopwords_list) x = keywords_df$Thematic x = removeWords(x,stopwords) keywords_df$Thematic <- x