I'm having a dataset containing firms involving in a certain category of products. Dataset looks like this:
df <- data.table(year=c(1979,1979,1980,1980,1980,1981,1981,1982,1982,1982,1982),
category = c("A","A","B","C","A","D","C","F","F","A","B"))
I want to create a new variable as follows: If a firm enters into a new category that it has not been previously engaged in previous years (not the same year), then that entry is labeld as "NEW", otherwise it will be labeld as "OLD".
As such, the desired outcome will be:
year category Newness
1: 1979 A NEW
2: 1979 A NEW
3: 1980 B NEW
4: 1980 C NEW
5: 1980 A OLD
6: 1981 D NEW
7: 1981 C OLD
8: 1982 F NEW
9: 1982 F NEW
10: 1982 A OLD
11: 1982 B OLD
I'm inclined to use data.table as I have over 1.5 million observations, and want to be able to replicate the solution by grouping by firm IDs.
Any help would be greatly appreciated, and thank you in advance.