I'm am doing some coding in R. I would want to show the rows that have duplicates for columns ID and NAME but have different values for AGE.
For example I have this table:
ID | NAME | AGE
111| Mark| 22
222| Anne| 21
333| Chery| 30
444| Megan| 16
555| Charles| 37
111| Mark| 23
222| Anne| 22
333| Chery| 30
111| Mark| 22
As of now I have this code:
readfile <- read.csv(file='/home/user/shane/names.csv')
dat <- data.frame(ID=c(readfile$ID),NAME=c(readfile$NAME),AGE=c(readfile$AGE))
nam <- duplicated(dat[,c('ID','NAME)]) | duplicated(dat[,c('ID','NAME], fromLast = TRUE)
readfile[nam,]
The output looks like this:
ID | NAME | AGE
111| Mark| 22
222| Anne| 21
333| Chery| 30
111| Mark| 23
222| Anne| 22
333| Chery| 30
111| Mark| 22
I would want the output to be:
ID | NAME | AGE
111| Mark| 22
222| Anne| 21
111| Mark| 23
222| Anne| 22
111| Mark| 22
I would want to remove the columns with the ID = 333 as they have the same value in Age. would anyone have a suggestion?