I have several data frames which has the following format:
| sequence | range | sequence_ID |
|---|---|---|
| K142442 | 283-423 | 58654 |
| K142442 | 283-414 | 58322 |
| K142442 | 192-342 | 36762 |
| K123123 | 771-250 | 21456 |
| K123123 | 771-250 | 76846 |
| K123123 | 771-250 | 41234 |
| K343232 | 320-642 | 82657 |
| K343232 | 320-642 | 36245 |
| K343232 | 1670-1521 | 25264 |
(showing the relevant columns for time issues)
As you notice, there are some ranges (range column) that are repeated along the same sequences (sequence column) as shown in the table. So I'm trying to remove those redundant rows by looking the same range starting number from left to right.
In brief, if for example the sequence K142442 has 3 representatives as shown in the table, but two of these representatives have the same initial range (283 in the table) I only want to keep the sequence with the longer range of those sequences in the table.
This is the example data frame:
df = data_frame(sequence = c(rep("K142442",3),rep("K123123",3),
rep("343232",3)),
range = c("283-423","283-414","192-342","771-250","771-250",
"771-250","320-642","320-642","1670-1521"),
sequence_ID = c(58645,58322,36762,21456,76846,41232,82657,36246,25264))
Desired output:
| sequence | range | sequence_ID |
|---|---|---|
| K142442 | 283-423 | 58654 |
| K142442 | 192-342 | 36762 |
| K123123 | 771-250 | 21456 |
| K343232 | 320-642 | 82657 |
| K343232 | 1670-1521 | 25264 |