My data looks something like this
zz <- 'wb_iso3c country year wbclass
1: YUG "Serbia and Montenegro (former)" 1990 NA
2: YUG "Yugoslavia (former)" 1990 UM
3: YUG "Yugoslavia (former)" 1991 NA
4: YUG "Serbia and Montenegro (former)" 1991 UM
5: YUG "Serbia and Montenegro (former)" 1992 NA
6: YUG "Yugoslavia (former)" 1992 NA'
Data <- read.table(text=zz, header = TRUE)
I would like to find a way to systematically and efficiently drop observations when:
My observation is duplicated only considering wb_iso3c and year. (hence I do not care if another variable such as country has different values). Among the "duplicated" observations, I would like to keep
the observation where wbclass is not NA. If
wbclass is NA for both observations, it is indifferent which row to keep.
The final dataset should look something like this
wb_iso3c country year wbclass
1: YUG Yugoslavia (former) 1990 UM
2: YUG Serbia and Montenegro (former) 1991 UM
3: YUG Serbia and Montenegro (former) 1992 <NA>
Thanks a lot in advance for your help. If you can use data.table of dyplr it would be great.