I'm using the excellent tidytext package to tokenize sentences in several paragraphs. For instance, I want to take the following paragraph:
"I am perfectly convinced by it that Mr. Darcy has no defect. He owns it himself without disguise."
and tokenize it into the two sentences
- "I am perfectly convinced by it that Mr. Darcy has no defect."
- "He owns it himself without disguise."
However, when I use the default sentence tokenizer of tidytext I get three sentences.
Code
df <- data_frame(Example_Text = c("I am perfectly convinced by it that Mr. Darcy has no defect. He owns it himself without disguise."))
unnest_tokens(df, input = "Example_Text", output = "Sentence", token = "sentences")
Result
# A tibble: 3 x 1
Sentence
<chr>
1 i am perfectly convinced by it that mr.
2 darcy has no defect.
3 he owns it himself without disguise.
What is a simple way to use tidytext to tokenize sentences but without running into issues with common abbreviations such as "Mr." or "Dr." being interpreted as sentence endings?