I'm trying to analyze a list of emails stored inside a dataframe (data$Email.Address) and I want to start by splitting the emails into parts, so that example1@gmail.com, example2@outlook.org, and example3@comcast.net end up like this:
email firstpart secondpart thirdpart
1 example1@gmail.com example1 gmail com
2 example2@outlook.org example2 outlook org
3 example3@comcast.net example3 comcast net
With my current code, however, I can't all match all strings — since some include domains like (some-url.com) or (us.army.mil). This means that example4@us.army.mil shows up as:
email firstpart secondpart thirdpart
4 example4@us.army.mil example4 us army
My goal is to read "some-url" or "us.army" as the second part, and "com" and "mil" as the third parts, so that is shows up like this:
email firstpart secondpart thirdpart
4 example4@us.army.mil example4 us.army mil
Here's the code I have:
library(tidyverse)
library(dplyr)
library(stringr)
library(rebus)
email_pattern <- capture(one_or_more(WRD)) %R%
"@" %R% capture(one_or_more(x = WRD)) %R%
DOT %R% capture(one_or_more(WRD))
#Split the emails into parts based on the pattern
email_parts <- str_match(data$Email.Address, pattern = email_pattern)
How can I change the code so that all the domains can be read? Thank you!