I am using the following function (based on https://rpubs.com/sprishi/twitterIBM) to extract bigrams from text. However, I want to keep the hash symbol for analysis purposes. The function to clean text works fine, but the unnest tokens function removes special characters. Is there any way to run unnest tokens without removing special characters?
x <- (c("I went to afternoon tea with her majesty and #queen @Victoria in the palace.", "Does tea have extra caffeine?"))
clean_Twitter_Corpus <- function(x) {
x = tolower(x) # convert to lower case characters
x = stripWhitespace(x) # removing white space
x = gsub("^\\s+|\\s+$", "", x) # remove leading and trailing white space
x = removeWords(x,stopwords("english")) # remove stopwords
return(x)
}
# clean the twitter texts. call the clean_Twitter_Corpus function
tweets <- clean_Twitter_Corpus(x)
tweets
text <- as.character(tweets)
text <- as.data.frame(text)
tidy_descr_ngrams <- text %>%
unnest_tokens(bigram, text, token = "ngrams", n = 2) %>%
separate(bigram, c("word1", "word2"), sep = " ")
tidy_descr_ngrams
bigram_counts <- tidy_descr_ngrams %>%
count(word1, word2, sort = TRUE)
bigram_counts