The problem with zip(*iter) is that it will iterate over the entire iterable and pass the resulting sequence in as args to zip.
So these are functionally the same:
Using *: xs, ys = zip(*[(p.x, p.y) for p in ((0,1),(0,2),(0,3))])
Using positionals: xz, ys = zip((0,1),(0,2),(0,3)).
Obviously, if there are millions of positional arguments this will be slow.
The iterator approach is the only work around.
I did a web search for python itertools unzip. Sadly the closest itertools gets is tee. In the link to the gist above, a tuple of iterators from itertools.tee is returned from this implementation of iunzip: https://gist.github.com/andrix/106334.
I had to convert it to python3:
from random import randint
import itertools
import time
from operator import itemgetter
def iunzip(iterable):
"""Iunzip is the same as zip(*iter) but returns iterators, instead of
expand the iterator. Mostly used for large sequence"""
_tmp, iterable = itertools.tee(iterable, 2)
iters = itertools.tee(iterable, len(next(_tmp)))
return (map(itemgetter(i), it) for i, it in enumerate(iters))
class Point:
def __init__(self, x, y):
self.x = x
self.y = y
points = [Point(randint(1, 10), randint(1, 10)) for _ in range(1000000)]
itime = time.time()
xs = [point.x for point in points]
ys = [point.y for point in points]
otime = time.time() - itime
itime += otime
print(f"original: {otime}")
xs, ys = zip(*[(p.x, p.y) for p in points])
otime = time.time() - itime
itime += otime
print(f"unpacking into zip: {otime}")
xs, ys = iunzip(((p.x, p.y) for p in points))
for _ in zip(xs, ys): pass
otime = time.time() - itime
itime += otime
print(f"iunzip: {otime}")
Output:
original: 0.1282501220703125
unpacking into zip: 1.286362886428833
iunzip: 0.3046858310699463
So the iterator is definitely better than unpacking into positional args. Not to mention the fact my 4GB of memory got eaten up when I went to 10 million points... However, I'm not convinced the iunzip iterator above is as optimal as it could be if it was a python builtin given that iterating twice to do the unzip as in the "original" method is still by far the fastest (~4x faster trying with various length of points).
Seems like iunzip should be a thing. I'm surprised it's not a python builtin or part of itertools...