TLDR: Generally, you can get the second occurrence of a PATTERN using one of the following
sub('.*?PATTERN.*?(PATTERN).*', '\\1', x)
stringr::str_match(x, 'PATTERN.*?(PATTERN)')[,2]
regmatches(x, regexpr('PATTERN.*?\\KPATTERN', x, perl=TRUE))
Details
You can use
x <- c('SEF001DT45','BV004MF')
sub('.*?[A-Z]{2}.*?([A-Z]{2}).*', '\\1', x)
## => [1] "DT" "MF"
See the R demo online and the regex demo. The point here is to match up to the second occurrence of the pattern, capture it, and then match the rest, and replace with the backreference to the capturing group value.
Note that sub will perform a single search and replace operation, and this is fine since the regex here requires the whole string match.
Details:
.*? - any zero or more chars as few as possible
[A-Z]{2} - two uppercase ASCII letters
.*? - any zero or more chars as few as possible
([A-Z]{2}) - Group 1 (\1 refers to this group value): two uppercase ASCII letters
.* - any zero or more chars as many as possible.
You can achieve this with a simpler regex using stringr::str_match:
x <- c('SEF001DT45','BV004MF')
library(stringr)
results <- stringr::str_match(x, '[A-Z]{2}.*?([A-Z]{2})')
results[,2] ## Get Group 1 values
See this R demo.
Or, with regmatches/regexpr in base R:
x <- c('SEF001DT45','BV004MF')
results <- regmatches(x, regexpr('[A-Z]{2}.*?\\K[A-Z]{2}', x, perl=TRUE))
results
See this R demo.
Here, [A-Z]{2}.*?\\K[A-Z]{2} finds the first two uppercase ASCII letters, then matches any zero or more chars (other than line break chars since the PCRE engine is used) as few as possible, and then \K discards the matched text and the [A-Z]{2} at the end of the pattern matches the second occurrence of the two-letter chunk. regexpr only finds the first match.