Get the first n/2 of n words in a column in a pandas data frame

Viewed 1682

I would like to get the first n/2 of n words in a column in a pandas data frame. Each row can have a different number of words, but every row has an even number of words. This column contains the name of an item, but every name is duplicated. For example, One became One One and One Two became One Two One Two.

I thought the following would work.

  1. count the number of words
  2. split the column on spaces
  3. get the first n/2 words in this split

But it doesn't work (I only casually use Python and pandas). Here is an MWE.

import pandas as pd
df = pd.DataFrame(['One One', 'One Two One Two'])
df[1] = df[0].str.count('\w+')
df[2] = df[0].str.split()
df[3] = df[0].get(df[2])

P.S. Please let me know if you have a good reference on pandas for the R user.

3 Answers
df[column_name].apply(lambda x: ' '.join(x.split()[:2]))

this takes the first n (2 in this above case) from the column names listed in the dataframe.

Related