Python fastest way to merge and average series / list data

Viewed 41

I'm trying to combine and average multiple (10-100 per call) data series, with each data series of approx shape=(1,100). I want to average the values of each result and output a series of equal length, ie output[i] = mean(series0[i], series1[i], series2[i].... This will be called ~10k times a day in early production, hopefully much more later, so I'm interested in wider tips or references if possible.

Currently the extant code in development is heavy on pandas for it's easy readability but is easily amended to output pandas.Series, python3 lists, or numpy.arrays, so anything goes. At a guess, I imagine some or all pandas will eventually be cut in favour of numpy.arrays and lists/dicts for speed/memory/cost reasons. I know enough to write the below code and know just about enough that a list comprehension may be a good contender, but I'm very much learning as I go so please be gentle.

I could find posts on merge/concat speeds, but rarely is this combined with further functions. So... suggestions on faster ways to produce an average series?

import numpy as np

series_length = 100
repeats=10

def foo(length):
    return np.random.randint(0,500,length,int)

results = []
for i in range(repeats):
    results.append(foo(series_length)) # produce a list n long, each containing a len=100 series/array/list (format optional) of integers

def some_code_here(data):
    avg_results = [np.mean([series[i] for series in data]) for i in range(series_length)]
    return avg_results

# Output length = series_length 
final_solution = some_code_here(results) 
1 Answers

You could use np.stack to create a single array from your data and then take the mean along axis = 0, which results in a speed improvement of ~45x on my machine:

import numpy as np
avg_results = np.stack(results).mean(axis=0)

Timings:

series_length = 100
repeats = 10
rng = np.random.default_rng()
results = [rng.integers(0, 500, series_length) for _ in range(repeats)]

#Check they're the same
assert ([np.mean([series[i] for series in results]) for i in range(series_length)] == np.stack(results).mean(axis=0)).all()
#OP
%timeit [np.mean([series[i] for series in results]) for i in range(series_length)]
#Me
%timeit np.stack(results).mean(axis=0)

Output:

787 µs ± 6.71 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)
17.6 µs ± 38.9 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each)
Related