I have a dataset with thousands of rows and almost a hundred columns. Each row only contains unique elements, however, these elements may also be found in other rows.
Basically, I want to create two new columns in my data frame, one to store how many Unique and another to store how many Ambiguous elements there are in a given row but compared to the whole dataset.
Note there are NAs in the dataframe that should not be considered when counting unique and ambiguous elements.
df <- data.frame(
col1 = c('Ab', 'Cd', 'Ef', 'Gh', 'Ij'),
col2 = c('Ac', 'Ce', 'Eg', 'Gi', 'Ik'),
col3 = c('Acc', NA, 'Ab', 'Gef', 'Il'),
col4 = c(NA, NA, NA, 'Ce', 'Im')
)
In the dataframe created above, Ab is not unique, so in row 1 there are 2 unique and 1 ambiguous elements when compared to the whole dataset.
In my expected output, Unique in row 1 would be equal to 2, and Ambiguous = 1. In row five, it would be 4 and 0, respectively.
I've searched for possible solutions, but most only deals with unique or repeated elements in a particular row, or across multiple rows for a particular column. Anyway, any help would be greatly appreciated.