I'm currently helping a client identify fraud activity using login activity. I've boiled the data down to a specific week and aggregated over day. for example:
| Date | IP Address | Total Login Attempts | Login Failure |
|---|---|---|---|
| 7/11/2022 | 1.256.xxx.xxx | 1000 | .5 |
| 7/12/2022 | 1.256.xxx.xxx | 50000 | .2 |
| 7/11/2022 | 1.126.xxx.xxx | 1000 | .5 |
IP address and Date are the primary keys in this table. Would it be appropriate to randomly sample the Total Login Attempts or Login Failure per day to create a normal distribution to determine outliers in the dataset? Is there a rule in statistics that would stop this form being effective? i.e. the timeseries components not being independent