How could I add the content of the skipped row to the previous row in a dataframe

Viewed 42

I have a datafrme like this:

test = pd.DataFrame({'label':['a','C','D','E','b','b','c','c','c'], 'text':['a','c','d','e','b','b','c','c','c'],'title':['a','c','d','e','b','b','c','c','c']})

the original dataframe

How could I add the content of the skipped row to the previous row when 'C','D','E' appear as a sequence. The ideal output would be:

test = pd.DataFrame({'label':['a','C','D','E','b','b','c','c','c'], 'text':['a','c','d','e','b','b','c','c','c'],'title':['a','c(e)','d','e','b','b','c','c','c']})

the ideal output

1 Answers

You can use shift

s = test['label']
idx = s.eq('C')&s.shift(-1).eq('D')&s.shift(-2).eq('E')
idx = idx[idx].index
test.loc[idx, 'title']+='('+test.shift(-2).loc[idx,'title']+')'

input:

test = pd.DataFrame({'label':['a','C','D','E','b','b','C','D','E'],
                     'text': ['a','c','d','e','b','b','c','c','c'],
                     'title':['a','c','d','e','b','b','c','c','c']})

output:

  label text title
0     a    a     a
1     C    c  c(e)
2     D    d     d
3     E    e     e
4     b    b     b
5     b    b     b
6     C    c  c(c)
7     D    c     c
8     E    c     c
Related