I have a data frame with Date, city, sales -
Date City Sales
2008-01-01 C1. 10000
2008-01-01 C2 2000
2008-01-02 C1. 13000
2008-01-02 C2 5000
and so on...
I have a function for outliers -
def detect_discrete_outliers(data):
outliers=[]
threshold=3
mean = np.mean(data)
std =np.std(data)
for i in data:
z_score= (i - mean)/std
if np.abs(z_score) > threshold:
outliers.append(i)
return outliers
Now, I want to use this outlier function to remove outliers from df
detect_discrete_outliers(df['sales'])
df = df[~ df['sales'].isin(detect_discrete_outliers(df['sales']))]
This removes the outliers in sales
However, I think its not accurate to do that, I need to remove outliers for each city and not overall outliers.
Can anyone suggest what's the best way to do this?