It's certainly a good solution to fix this in the character inputs before creating quanteda objects (corpus, tokens, etc.). An alternative in quanteda is to tokenise the texts with the ending hyphens, then:
- compound the hyphenated tokens with the token that follows
- remove the new tokens with their internal hyphens
Example:
library("quanteda")
## Package version: 3.0.0
## Unicode version: 10.0
## ICU version: 61.1
## Parallel computing: 12 of 12 threads used.
## See https://quanteda.io for tutorials and examples.
txt <- c(
"The sun is shin- ing.",
"Hyphen- ation is fun",
"text an- alysis"
)
toks <- tokens(txt)
toks
## Tokens consisting of 3 documents.
## text1 :
## [1] "The" "sun" "is" "shin-" "ing" "."
##
## text2 :
## [1] "Hyphen-" "ation" "is" "fun"
##
## text3 :
## [1] "text" "an-" "alysis"
The compounding step:
toksc <- tokens_compound(toks, phrase("*- *"), concatenator = "")
toksc
## Tokens consisting of 3 documents.
## text1 :
## [1] "The" "sun" "is" "shin-ing" "."
##
## text2 :
## [1] "Hyphen-ation" "is" "fun"
##
## text3 :
## [1] "text" "an-alysis"
And finally the replacement without hyphens step:
toks_hyphenated <- grep("\\w+-\\w+", types(toksc), value = TRUE)
tokens_replace(toksc, toks_hyphenated, gsub("-", "", toks_hyphenated))
## Tokens consisting of 3 documents.
## text1 :
## [1] "The" "sun" "is" "shining" "."
##
## text2 :
## [1] "Hyphenation" "is" "fun"
##
## text3 :
## [1] "text" "analysis"
Edit: Added to question
If you really want to recombine these to make a corpus from the processed tokens, you can apply this step:
> toks_rejoined <- tokens_replace(toksc, toks_hyphenated, gsub("-", "",
> corpus(sapply(toks_rejoined, paste, collapse = " "))
Corpus consisting of 3 documents.
text1 :
"The sun is shining ."
text2 :
"Hyphenation is fun"
text3 :
"text analysis"