I have a data.frame with a large number of (lengthy) strings. I'm trying to clean them up a little bit before processing them, but I've run into a problem when dealing with periods. I'd like to be able to differentiate between when a period is used to end a sentence and when it's used as part of an abbreviation. I'd like to do this by length of word, but haven't figured out the right regex for it.
Say I have a string like this: mystring <- "hello.world from the u.s.a.". I'd like to replace this with something like "hello world from the usa".
I could try splitting the data.frame by spaces using split_string <- unlist(strsplit(mystring, split=" ")) , and then running something like
split_string <- ifelse(nchar(split_string) < 7, gsub(".", "", split_string), gsub(".", " ", split_string))
But as the body of text is rather large, this is a very slow (and rather ugly) process. How could I do this in a more efficient and cleaner manner?