Finding and matching reversed strings efficiently in R

Viewed 149

I have a large number of strings (~280,000) that all have the following format "ABC12D/XYZ34A". In my data, each of those strings has a duplicate entry that is identical but in reverse, e.g. "XYZ34A/ABC12D" for the example above. So, my data looks something like this:

1    "ABC12D/XYZ34A"
2    "TUR44F/SWP29R"
3    "PLL93S/WQQ22F"
4    "YNV77C/AAZ05S"
5    "SWP29R/TUR44F"
6    "AAZ05S/YNV77C"
7    "CLK86G/ERF74Q"
8    "XYZ34A/ABC12D"
9    "ERF74Q/CLK86G"
10   "WQQ22F/PLL93S"

Row 1 matches row 8, row 2 matches row 5, etc.

My aims are: 1) for a given string, find where its reversed entry is and keep this index and then 2) replace the reverse entry with the non-reverse entry:

1   "ABC12D/XYZ34A" 8
2   "TUR44F/SWP29R" 5
3   "PLL93S/WQQ22F" 10
4   "YNV77C/AAZ05S" 6
5   "TUR44F/SWP29R" 0
6   "YNV77C/AAZ05S" 0
7   "CLK86G/ERF74Q" 9
8   "ABC12D/XYZ34A" 0
9    "CLK86G/ERF74Q" 0
10   "PLL93S/WQQ22F" 0

Currently, I do this in the following way using a loop:

df <- data.frame(c("ABC12D/XYZ34A", "TUR44F/SWP29R", "PLL93S/WQQ22F", 
"YNV77C/AAZ05S", "SWP29R/TUR44F", "AAZ05S/YNV77C", "CLK86G/ERF74Q", 
"XYZ34A/ABC12D", "ERF74Q/CLK86G", "WQQ22F/PLL93S"), stringsAsFactors = 
FALSE)

colnames(df) <- "entries"
df

# Reverse function
reverse.entry <- function(string) {
  string.reversed <- paste(rev(strsplit(string, "/")[[1]]), collapse = '/')
  string.reversed
}

duplicate.flag <- list() 
duplicate.idx <- list() 

# Find and replace reversed entries
for (i in 1:dim(df)[[1]]) {
  # current entry
  string = df[i,]

  # reverse the current entry
  string.reversed <- reverse.entry(string)

  # if any other entry matches the reversed string get match index 
  if (grepl(string.reversed, df)) {

    print(sprintf("%d found a reversal", i))
    idx <- which(df == string.reversed)
    duplicate.flag[i] <- 1;
    duplicate.idx[i] <- idx;
    # replace reversed strings with original strings
    df[idx,] <- string
  } else {
    duplicate.flag[i] <- 0;
    duplicate.idx[i] <- 0;
  }

}

data.frame(df, unlist(duplicate.idx), unlist(duplicate.flag))

However, this is quite slow and is taking several hours. Is there a better way of programming this? I'm fairly new to R and programming so am not terribly good at vectorization etc. Since each entry has one reverse entry, I could also just have the loop for 1:dim(df)[[1]] / 2. Would that already save a lot of time?

Many thanks!

3 Answers
Related