Some words won't be stemmed using tm ("easier" or "easiest")

Viewed 41

I have large questionaire dataset where some of the features need to be stemmed, with the goal being to assign a topic to each response. However, I'm having trouble stemming some words using the package tm.

Here is a reproducible (simplified) example:

library(tm)

# Words that need to be stemmed
test_vec <- c("easier","easy","easiest","closest","close","closer","near","nearest")

# Preprocessing function to clean corpus
# Note that, this is my full pipeline, but only the last command will be used in this case example
clean_corpus<- function(corpus){
  corpus <- tm_map(corpus, stripWhitespace)
  corpus <- tm_map(corpus, removePunctuation)
  corpus <- tm_map(corpus, removeNumbers)
  corpus <- tm_map(corpus, content_transformer(tolower))
  corpus <- tm_map(corpus, removeWords, stopwords("en"))
  corpus <- tm_map(corpus,stemDocument)
    return(corpus)
}

# Create corpus with test_vec
test_corpus <- VCorpus(VectorSource(test_vec))
# Apply cleaning
test_corpus <- clean_corpus(test_corpus)

# Print out stemmed values
for(i in 1:length(test_corpus)){
  print(test_corpus[[i]]$content)
}

[1] "easier"
[2] "easi"
[3] "easiest"
[4] "closest"
[5] "close"
[6] "closer"
[7] "near"
[8] "nearest"

Question 1 Why isn't [1] "easier" and [3] "easiest" stemmed to be "easi" (like "easy" has been). Similarly, why isn't "close" or "near" stemmed. Am I missing something?

Question 2 This is a side question, but is there a way to relate words like "close" and "near" from a dictionary that would be able to verify that these are synonyms. If they are synonyms, all instances of "near" will then be changed to "close", for example.

1 Answers

There are multiple stemmers (quick overview here), but porter is used the most. Python also has Lancaster stemming which would return the following based on your test_vec:

easier easy
easy easy
easiest easiest
closest closest
close clos
closer clos
near near
nearest nearest

But still there are issues as iest, is not shortened.

But you could also use lemmatization which would return the following:

library(textstem)
lemmatize_words(test_vec)
"easy"  "easy"  "easy"  "close" "close" "close" "near"  "near" 

For topic assignment lemmatising might be preferred to stemming because it groups better. But you need to be aware of the differences between both.

Lemmatisation (or lemmatization) in linguistics is the process of grouping together the inflected forms of a word so they can be analysed as a single item, identified by the word's lemma, or dictionary form. Wikipedia: lemmatisation

Stemming is the process of reducing inflected (or sometimes derived) words to their word stem, base or root form—generally a written word form. The stem need not be identical to the morphological root of the word; it is usually sufficient that related words map to the same stem, even if this stem is not in itself a valid root. Wikipedia: stemming

As for your second question, there is a package called syn (only for English), which contains all the synonyms, but it will create a list of all of them and for "close" or "near" that is a very long list. Or package qdap, that also has a synonym function.

Related