How to resample value in pandas column?

Viewed 54

I know about the resample function on time-series data. I want something similar on a normal column with 3000 examples. I want to keep the length. I want every row to have the value of the last occurrence in a n- long window.

I know about group by as well and the last function, but here I am groupping based on length not on some value.

I want non-overlapping windows, so rolling does not help either.

Example of window of size three:

0         sakshijoshii
1         medpagetoday
2            nickmmark
3      mukeshm07384110
4         DipakBiswas_       
5      jaysanchezdorta
6         Terry6969696
7            LizShelby
8            wlharper1
9       BruhOriginalMe

What I want:

0            nickmmark
1            nickmmark
2            nickmmark
3      jaysanchezdorta
4      jaysanchezdorta      
5      jaysanchezdorta
6            wlharper1
7            wlharper1
8            wlharper1
9       BruhOriginalMe
1 Answers

You can go for

df.groupby(np.arange(len(df)) // n)[col_name].transform("last")

Grouping by every n'th element of the frame can be done by looking at 0...N-1 values' dividents after dividing by n. e.g., for 0..7 values with n = 3, we get 0, 0, 0, 1, 1, 1, 2. Then transform with last gets the last entry of each group and produces a like-indexed series via repeating it for each group member.

For the sample given:

>>> df

             names
0     sakshijoshii
1     medpagetoday
2        nickmmark
3  mukeshm07384110
4     DipakBiswas_
5  jaysanchezdorta
6     Terry6969696
7        LizShelby
8        wlharper1
9   BruhOriginalMe

>>> n = 3
>>> col_name = "names"
>>> df.groupby(np.arange(len(df)) // n)[col_name].transform("last")

0          nickmmark
1          nickmmark
2          nickmmark
3    jaysanchezdorta
4    jaysanchezdorta
5    jaysanchezdorta
6          wlharper1
7          wlharper1
8          wlharper1
9     BruhOriginalMe
Related