How to use pd.interpolate fill the gap with only one missing data

Viewed 125

I have a time series data for air pollution with several missing gaps like this:

Date             AMB_TEMP      CO         PM10       PM2.5
2010-01-01 0         8         10         ...          15
2010-01-01 1         10        15         ...          20
...
2010-01-02 0                   5          ...           
2010-01-02 1                              ...          20
...
2010-01-03 1         4         13         ...          34     

To specify, here's the data link: shorturl.at/blBN1

The gaps were composed of several consecutive or inconsecutive NAs, and there are some helpful statistics done by R like:

  1. Length of time series: 87648
  2. Number of Missing Values:746
  3. Percentage of Missing Values: 0.85 %
  4. Number of Gaps: 136
  5. Average Gap Size: 5.485294
  6. Longest NA gap (series of consecutive NAs): 32
  7. Most frequent gap size (series of consecutive NA series): 1(occurring 50 times)

Generally if I use the df.interpolate(limit=1), gaps with more than one missing will be interpolated as well.

So I guess a better way to interpolate the gap with only one missing is to get the gap id.

To do so, I grouped the different size of gap and used the following function:

    cum = df.notna().cumsum()
    cum[cum.duplicated()]

and got the result:

                       PM2.5
2019-01-09 13:00:00     205
2019-01-10 15:00:00     230
2019-01-10 16:00:00     230
2019-01-16 11:00:00     368
2019-01-23 14:00:00     538
                       ... 
2019-12-02 10:00:00    7971
2019-12-10 09:00:00    8161
2019-12-16 15:00:00    8310
2019-12-24 12:00:00    8498
2019-12-31 10:00:00    8663

How to get the index of each first missing value in each gap like this?

                     PM2.5 gap size
2019-01-09 13:00:00     1
2019-01-10 15:00:00     2
2019-01-16 11:00:00     1
2019-01-23 14:00:00     1
                       ... 
2019-12-02 10:00:00     1
2019-12-10 09:00:00     1
2019-12-16 15:00:00     1
2019-12-24 12:00:00     1
2019-12-31 10:00:00     1

but when I used cum[cum.duplicated()].groupby(cum[cum.duplicated()]).count() the index would miss.

Are there better solutions to do these?

OR How to interpolate case by case?

Anyone can help me?

0 Answers
Related