UPDATED QUESTION
I have this character vector
str_ <- "H3K9me0S10ph1K14ac1me0"
I would like to break it into pieces such that I get an output like:
"H3K9: me0 | S10: ph1 | K14: ac1,me0"
Preferably this is done in a manner that utilizes {dplyr}, such that I can perform this operation on a tibble and get a new column with the desired character string output. Any ideas?
As the below section suggests, I'm struggling with getting a table that denotes which modifications are paired with what, e.g. that the me0 goes with H3K9 and BOTH the ac1,me0 go with K14
Any assistance would be so helpful!
Pieces of attempts
Using a slightly different example,
str_ <- "H3K9ac1K14ac1K18ac1me0"
So I've tried breaking the character vector into pieces by extracting all "me[0-9]*" or "ac[0-9]*" etc, then giving them an id which corresponds to their index in the character vector.
# A tibble: 4 x 2
i m
<int> <chr>
1 12 ac1
2 17 ac1
3 23 ac1
4 26 me0
I need a way to create a column together that tells whether two modifications belong to the same protein, i.e. in this example K14 has ac1 and me0, so their 'together' values should be 'TRUE'. I've tried using the distance between their indices as a surrogate for togetherness, but I don't think this is the best way to do it:
# A tibble: 4 x 2
i m unit_diff together
<int> <chr> <int> <lgl>
1 12 ac1 0 FALSE
2 17 ac1 5 FALSE
3 23 ac1 6 TRUE
4 26 me0 3 TRUE
Any ideas? I've tried using modulo 3, but this doesn't seem to generalize. Is this even the correct way to be doing this? I'm open to suggestions