I have large questionaire dataset where some of the features need to be stemmed, with the goal being to assign a topic to each response. However, I'm having trouble stemming some words using the package tm.
Here is a reproducible (simplified) example:
library(tm)
# Words that need to be stemmed
test_vec <- c("easier","easy","easiest","closest","close","closer","near","nearest")
# Preprocessing function to clean corpus
# Note that, this is my full pipeline, but only the last command will be used in this case example
clean_corpus<- function(corpus){
corpus <- tm_map(corpus, stripWhitespace)
corpus <- tm_map(corpus, removePunctuation)
corpus <- tm_map(corpus, removeNumbers)
corpus <- tm_map(corpus, content_transformer(tolower))
corpus <- tm_map(corpus, removeWords, stopwords("en"))
corpus <- tm_map(corpus,stemDocument)
return(corpus)
}
# Create corpus with test_vec
test_corpus <- VCorpus(VectorSource(test_vec))
# Apply cleaning
test_corpus <- clean_corpus(test_corpus)
# Print out stemmed values
for(i in 1:length(test_corpus)){
print(test_corpus[[i]]$content)
}
[1] "easier"
[2] "easi"
[3] "easiest"
[4] "closest"
[5] "close"
[6] "closer"
[7] "near"
[8] "nearest"
Question 1
Why isn't [1] "easier" and [3] "easiest" stemmed to be "easi" (like "easy" has been). Similarly, why isn't "close" or "near" stemmed. Am I missing something?
Question 2
This is a side question, but is there a way to relate words like "close" and "near" from a dictionary that would be able to verify that these are synonyms. If they are synonyms, all instances of "near" will then be changed to "close", for example.