If I have a data frame like the following
| group1 | group2 | col1 | col2 |
|---|---|---|---|
| A | 1 | ABC | 5 |
| A | 1 | DEF | 2 |
| B | 1 | AB | 1 |
| C | 1 | ABC | 5 |
| C | 1 | DEF | 2 |
| A | 2 | BC | 8 |
| B | 2 | AB | 1 |
We can see that the the (A, 1) and (C, 1) groups have the same rows (since col1 and col2 are the same within this group). The same is true for (B,1) and (B, 2).
So really we are left with 3 distinct "larger groups" (call them categories) in this data frame, namely:
| category | group1 | group2 |
|---|---|---|
| 1 | A | 1 |
| 1 | C | 1 |
| 2 | B | 1 |
| 2 | B | 2 |
| 3 | A | 2 |
And I am wondering how can I return the above data frame in R given a data frame like the first? The order of the "category" column doesn't matter here, for example (A,2) could be group 1 instead of {(A,1), (C,1)}, as long as these have a distinct category index.
I have tried a few very long/inefficient ways of doing this in Dplyr but I'm sure there must be a more efficient way to do this. Thanks