Replacing punctuation in string in different ways by word length in R

Viewed 429

I have a data.frame with a large number of (lengthy) strings. I'm trying to clean them up a little bit before processing them, but I've run into a problem when dealing with periods. I'd like to be able to differentiate between when a period is used to end a sentence and when it's used as part of an abbreviation. I'd like to do this by length of word, but haven't figured out the right regex for it.

Say I have a string like this: mystring <- "hello.world from the u.s.a.". I'd like to replace this with something like "hello world from the usa".

I could try splitting the data.frame by spaces using split_string <- unlist(strsplit(mystring, split=" ")) , and then running something like

split_string <- ifelse(nchar(split_string) < 7, gsub(".", "", split_string), gsub(".", " ", split_string))

But as the body of text is rather large, this is a very slow (and rather ugly) process. How could I do this in a more efficient and cleaner manner?

2 Answers

You can test this to see if this is any faster. It looks for a delimiter, up to 6 non-space characters and a delimiter and for any such match it runs the anonymous function specified in formula notation in the second argument of gsubfn. That anonymous function removes any periods in the match. In what is left the gsub replaces each period with a space.

library(gsubfn)
pat <- "(?<=^| )(\\S{1,6})(?=$| )"
gsub("[.]", " ", gsubfn(pat, ~ gsub("[.]", "", ..1), mystring, perl = TRUE))
## [1] "hello world from the usa"

How about the following...

mystring2 <- gsub("(\\w)\\.(\\w)","\\1 \\2",gsub("\\.(\\w+)\\.","\\1",mystring))

mystring2
[1] "hello world from the usa."

For dots either side of letters, it deletes them first, then for the remaining dots with letters either side, it replaces them with a space.

It even keeps the last dot in your example as the end of a sentence!

Related