Efficient Way to Slice Strings in Pandas

Viewed 327

I have a dataset that has over 100 million rows that I am trying to manipulate in pandas. I am trying to slice the string in a based on the values in b and c as the start and end points respectively.

enter image description here

I can do this with list comprehension like so:

df['d'] = [a[1]['a'][a[1]['b']:a[1]['c']] for a in df.iterrows()]

This is really slow. I can do the same thing with an apply like this:

df['d'] = df.apply(lambda x: x['a'][x['b']:x['c']],axis=1)

This is also quite slow. My question is, what is the most efficient way to slice the strings in a using the values in b and c as the start and end for the slice?

1 Answers

Iterating over df.iterrows() is really slow because for each row it creates a separate pd.Series object. For 100 million rows this means 100 million such objects are being created (and discarded). It's better to zip the columns and use this in a comprehension like so:

df.assign(d=[a[b:c] for a, b, c in zip(df['a'], df['b'], df['c'])])

This will only create three Series objects and then iterate over them which saves a lot of overhead.

You can also take a look at Numba to write your own function that loops over the data frame.

Related