how many rows does interpolation consider?

Viewed 40

How does pandas' DataFrame.interpolation() work in relation to the amount of rows it considers:

  1. is it just the row before the NaNs and the row right after?
  2. Or is it the whole DataFrame (how does that work at 1 million rows?)
  3. Or another way (please explain)

Edit: (with method=='polynomial' ideally)

1 Answers

With method='polynomial', DataFrame.interpolate() starts with the first non-NaN values, and stops at the last non-NaN value, leaving leading and trailing NaNs unchanged:

>>> n = np.nan
>>> df = pd.DataFrame([n,n,4,2,4,n,n,2,n,n,n,n,n,])
>>> df
      0
0   NaN
1   NaN
2   4.0
3   2.0
4   4.0
5   NaN
6   NaN
7   2.0
8   NaN
9   NaN
10  NaN
11  NaN
12  NaN

>>> df.interpolate(method='polynomial', order=2)
           0
0        NaN  <--- 
1        NaN  <--- first 2 NaN values are unchanged
2   4.000000
3   2.000000
4   4.000000
5   5.076923
6   4.410256
7   2.000000  <--- last non-NaN value
8        NaN  <--- this and subsequent NaN values unchanged
9        NaN
10       NaN
11       NaN
12       NaN

If you'd like the leading and trailing NaN values filled, just use bfill and ffill:

>>> df.interpolate(method='polynomial', order=2).bfill().ffill()
           0
0   4.000000
1   4.000000
2   4.000000
3   2.000000
4   4.000000
5   5.076923
6   4.410256
7   2.000000
8   2.000000
9   2.000000
10  2.000000
11  2.000000
12  2.000000
Related