I have a small follow-up question to a previous one answered here. With a dummy dataset as included below, I would like to filter off rows in which one or more of the samples (denoted S1, S2, S3, S4 in the dataset) has only "1" and "0", as well as all rows where the samples have only 1", "0" and "."
| ID | Pos | S1 | S2 | S3 | S4 |
|---|---|---|---|---|---|
| A | 22 | . | 1 | 0 | . |
| B | 21 | 1 | 0 | . | 1 |
| C | 50 | 0 | . | . | . |
| D | 11 | . | 1 | . | . |
| E | 13 | 0 | 0 | 0 | 0 |
| F | 14 | 1 | 1 | 1 | 1 |
| G | 10 | 1 | 0 | 0 | 0 |
In other words, I want to keep only those rows that either has "1" in all the samples, or "0" in all samples, or "1" and "." in all samples, or lastly those with "0" and "." such that at the end, I have the final dataset looking as below
| ID | Pos | S1 | S2 | S3 | S4 |
|---|---|---|---|---|---|
| C | 50 | 0 | . | . | . |
| D | 11 | . | 1 | . | . |
| E | 13 | 0 | 0 | 0 | 0 |
| F | 14 | 1 | 1 | 1 | 1 |
I am trying to do this in R and I request any suggestions from you.
Thanks and regards!