I spent the last hour trying to reformat a 2-column format into something more usable.
I have the following input (a 2 column data frame / tibble) :
Input
TGGGAAGGTTATGTGC-1 CMO305|CMO306|CMO312 3698|3806|12182
TGTTCTACATGACAGG-1 CMO305|CMO306|CMO312 3027|1449|4184
ACTGATGCAGAGTGAC-1 CMO305|CMO307 6802|4715
ATCGTCCGTTACCCAA-1 CMO305|CMO307 5599|7019
ATGCATGTCATGACAC-1 CMO305|CMO307 10872|16729
GTGAGTTAGTCCGCCA-1 CMO305|CMO307 10096|3434
Desired output (A - wide)
| CMO305 | CMO306 | CMO307 | CMO312 | |
|---|---|---|---|---|
| TGGGAAGGTTATGTGC-1 | 3698 | 3806 | 0 | 12182 |
| TGTTCTACATGACAGG-1 | 3027 | 1449 | 0 | 4184 |
| ACTGATGCAGAGTGAC-1 | 6802 | 0 | 4715 | 0 |
| ATCGTCCGTTACCCAA-1 | 5599 | 0 | 7019 | 0 |
| ATGCATGTCATGACAC-1 | 10872 | 0 | 16729 | 0 |
| GTGAGTTAGTCCGCCA-1 | 10096 | 0 | 3434 | 0 |
Desired output (B - long format)
> CMO.umis.long
feature_call num_umis
<chr> <dbl>
1 CMO304 2168
2 CMO304 14210
3 CMO304 7009
4 CMO304 5931
5 CMO304 7147
6 CMO304 1683
I am pretty sure this has been answered already, but I can't seem to find the right search terms.
separate_rows() may be the way but I cannot get it to split correclty...
Thank you, I appreciate your help!