Performance difference when doing string indexing using [0] and [0:1] on list, numpy and series

Viewed 87

System Information

Python version : 3.7.12
Pandas version : 1.3.4
OS : Debian 10

Background

I'm doing some experiment on the fastest way to extract first several letter of list/series of string and I notice some very weird behaviour of [ ]. Here's 3 different experiment I tried :


1. Analyzing performance of [0] compared to [0:1] on list comprehension

Let's say X is an array-like object. Then X[0] and X[0:1] must be doing the same thing.

But from this experiment:

N = 1000000
list_str = ['ABC']*N
%timeit [x[0] for x in list_str]   # 37.8 ms ± 141 µs per loop 
%timeit [x[0:1] for x in list_str] # 69.7 ms ± 113 µs per loop

N = 10
list_str = ['ABC']*N
%timeit [x[0] for x in list_str]   # 592 ns ± 3.72 ns per loop
%timeit [x[0:1] for x in list_str] # 956 ns ± 6.29 ns per loop

The x[0] is almost 60%-90% faster than x[0:1], which is quite weird. This indicating that [0] and [0:1] is not doing the same thing behind the scenes even if it produces the same output.

2. Analyzing performance of [0] compared to [0:1] on numpy array comprehension

I curious if using numpy array will yield different performance.

N = 1000000
list_str = np.array(['ABC']*N)
%timeit [x[0] for x in list_str]  # 262 ms ± 9.95 ms per loop
%timeit [x[0:1] for x in list_str]# 293 ms ± 8.34 ms per loop

N = 10
list_str = np.array(['ABC']*N)
%timeit [x[0] for x in list_str]  # 3.61 µs ± 22.3 ns per loop
%timeit [x[0:1] for x in list_str]# 4.04 µs ± 58.7 ns per loop

The x[0] is faster only by around 10% than x[0:1]

3. Analyzing performance of [0] compared to [0:1] on string indexing with .str accessor on pandas' series.

On pandas, we can use .str[n] to get the n-th element of each member on the series. The experiment goes like this:

N = 1000000
series = pd.Series(['ABC']*N)
%timeit series.str[0]   # 388 ms ± 10.5 ms per loop
%timeit series.str[0:1] # 172 ms ± 712 µs per loop

N = 10
series = pd.Series(['ABC']*N)
%timeit series.str[0]   # 131 µs ± 2.02 µs per loop 
%timeit series.str[0:1] # 133 µs ± 4.25 µs per loop

On this case, [0:1] is over 50% than [0] approach on long series.

Question

I thought that [0] and [0:1] doing the same thing, but from this experiment it is already clear that [0] and [0:1] can give significant difference in speed. So, what do I miss in here? Is because of the data structure? Or [0] and [0:1] just simply not same?

0 Answers
Related