I am trying to write a function to detect capitalized words that are all capitalised
currently, code:
df <- data.frame(title = character(), id = numeric())%>%
add_row(title= "THIS is an EXAMPLE where I DONT get the output i WAS hoping for", id = 6)
df <- df %>%
mutate(sec_code_1 = unlist(str_extract_all(title," [A-Z]{3,5} ")[[1]][1])
, sec_code_2 = unlist(str_extract_all(title," [A-Z]{3,5} ")[[1]][2])
, sec_code_3 = unlist(str_extract_all(title," [A-Z]{3,5} ")[[1]][3]))
df
Where output is:
| title | id | sec_code_1 | sec_code_2 | sec_code_3 |
|---|---|---|---|---|
| THIS is an EXAMPLE where I DONT get the output i WAS hoping for | 6 | DONT | WAS |
The first 3-5 letter capitalized word is "THIS", second should skip example (>5) and be "DONT", third example should be "WAS". ie:
| title | id | sec_code_1 | sec_code_2 | sec_code_3 |
|---|---|---|---|---|
| THIS is an EXAMPLE where I DONT get the output i WAS hoping for | 6 | THIS | DONT | WANT |
does anyone know where Im going wrong? specifically how I can denote "space or beginning of string" or "space or end of string" logically using stringr.