python memory usage: txt file much smaller than python list containing file text

Viewed 422

I have a 543 MB txt file containing a single line of space separated, utf-8 tokens:

aaa algeria americansamoa appliedethics accessiblecomputing ada anarchism ...

But, when I load this text data into a python list, it uses ~8 GB of memory (~900 MB for the list and ~8 GB for the tokens):

with open('tokens.txt', 'r') as f:
    tokens = f.read().decode('utf-8').split()

import sys

print sys.getsizeof(tokens)
# 917450944 bytes for the list
print sum(sys.getsizeof(t) for t in tokens)
# 7067732908 bytes for the actual tokens

I expected the memory usage to be approximately file size + list overhead = 1.5 GB. Why do the tokens consume so much more memory when loaded into a list?

1 Answers
Related