Slicing of a memoryview should be faster than bytes, should'nt it?

Viewed 110

I receive a long, structured byte string (up to several mbytes) over internet and need to unpack it to python types.

In order to do so, i'm using struct and wanted to try using memoryviews to take slices efficiently, since taking a slice of a string creates a copy. Have to mention that struct.iter_unpack is unavailable, since it was added in 3.4.

And while benchmarking implementation using memoryview, in comparison with current implementation (regular slices), I stumbled upon interesting (to me) results:

Creating fake data:

import struct
import random

# Create fake data
data = {}
for i in range(5001):
    data[i] = [random.randint(1, 100), random.randint(1, 100), random.randint(1, 100)]
    
st = struct.Struct('!I3B')
stsize = st.size

data = b"".join(st.pack(k, *v) for k, v in data.items())
data_buffer = memoryview(data)

Test functions:

def unpack_mview(mview):
    end = len(mview)
    for i in range(0, end, stsize):
        st.unpack(mview[i:i+stsize])
            
def unpack_slice(byte_array):
    end = len(byte_array)
    for i in range(0, end, stsize):
        st.unpack(byte_array[i:i+stsize])

which yields me following results:

timeit.timeit('unpack_mview(data_buffer)', setup='from __main__ import unpack_mview, data_buffer', number=1000)
0.97220089999999
timeit.timeit('unpack_slice(data)', setup='from __main__ import unpack_slice, data', number=1000)
0.7721589999999878

If I use while loop instead of for loop, memoryview indeed works better:

def unpack_mview_while(mview):
    while mview:
        st.unpack(mview[:stsize])
        mview = mview[stsize:]

def unpack_slice_while(byte_array):
    while byte_array:
        st.unpack(byte_array[:stsize])
        byte_array = byte_array[stsize:]
>>> timeit.timeit('unpack_mview_while(data_buffer)', setup='from __main__ import unpack_mview_while, data_buffer', number=1000)
1.2701496999999904
>>> timeit.timeit('unpack_slice_while(data)', setup='from __main__ import unpack_slice_while, data', number=1000)
3.474370599999986

But why would I.

I expected for slicing of a memoryview to be faster than taking a slice of bytestring. But it seems that initializing a subview takes more time than creating a copy of a substring with length of 10 bytes. Is this correct and speed up will only be visible as soon as slice is bigger than 10 bytes or I misunderstood the purpose or usecases of memoryviews? Or I plainly using memoryview in the wrong way?

0 Answers
Related