I have a dataframe of old ensembl stable transcript IDs that I am in the process of mapping to current genes. Ensembl stable transcript IDs are a combination of characters and integers. As new information comes out they may choose to update the version number, which is the number after the decimal. I want to subset this dataframe further so that it only includes the most recent (largest) version number.
Here is some example data:
| stableID | release |
|---|---|
| ENSMUST00000080572.10 | 80 |
| ENSMUST00000080572.11 | 81 |
| ENSMUST00000080572.12 | 82 |
| ENSMUST00000071062.6 | 80 |
| ENSMUST00000071062.7 | 81 |
| ENSMUST00000071062.8 | 82 |
| ENSMUST00000124232.1 | 62 |
| ENSMUST00000124232.1 | 64 |
Here is what I would like to return through some code:
| stableID | release |
|---|---|
| ENSMUST00000080572.12 | 82 |
| ENSMUST00000071062.8 | 82 |
| ENSMUST00000124232.1 | 62 |
| ENSMUST00000124232.1 | 64 |
I haven't the slightest clue of what to write for this. Current thought is to split the decimal into a third column and then some dplyr::group_by() %>% dplyr::filter() shenanigans perhaps?