I have corpus and I am trying to find the frequencies of multiple phrases totalled by year and plot this. For example, if the phrase "American economy" and "Canadian economy" is mentioned 2 times each in 2004, I would want this to give a frequency of 4 in 2004.
I have managed to do this for single tokens, but am having trouble trying it for phrases. This is the code I used to do for single tokens.
a_corpus <- corpus(df, text = "text")
my_dict <- dictionary(list(america = c("America", "President")))
freq_grouped_creators <- textstat_frequency(dfm(tokens(a_corpus)),
groups = a_corpus$Year)
freq_word_creators <- subset(freq_grouped_creators, freq_grouped_creators$feature %in% my_dict$america)
# collapsing rows by year to total frequencies for tokens
freq_word_creators_2 <- freq_word_creators %>%
group_by(group) %>%
summarize(Sum_frequency = sum(frequency))
# plotting
ggplot(freq_word_creators_2, aes(x = group, y =
Sum_frequency)) +
geom_point() +
scale_y_continuous(limits = c(0, 300), breaks = c(seq(0, 300, 30))) +
xlab(NULL) +
ylab("Frequency") +
theme(axis.text.x = element_text(angle = 90, hjust = 1))
