How to mmap a 2d array from a text file

Viewed 786

I have very large files containing 2d arrays of positive integers

  • Each file contains a matrix

I would like to process them without reading the files into memory. Luckily I only need to look at the values from left to right in the input file. I was hoping to be able to mmap each file so I can process them as if they were in memory but without actually reading in the files into memory.

Example of smaller version:

[[2, 2, 6, 10, 2, 6, 7, 15, 14, 10, 17, 14, 7, 14, 15, 7, 17], 
 [3, 3, 7, 11, 3, 7, 0, 11, 7, 16, 0, 17, 17, 7, 16, 0, 0], 
 [4, 4, 8, 7, 4, 13, 0, 0, 15, 7, 8, 7, 0, 7, 0, 15, 13], 
 [5, 5, 9, 12, 5, 14, 7, 13, 9, 14, 16, 12, 13, 14, 7, 16, 7]]

Is it possible to mmap such a file so I can then process the np.int64 values with

for i in range(rownumber):
    for j in range(rowlength):
        process(M[i, j])

To be clear, I don't want ever to have all my input file in memory as it won't fit.

5 Answers

Updated Answer

On the basis of your comments and clarifications, it appears you actually have a text file with a bunch of square brackets in it that is around 4 lines long with 1,000,000,000 ASCII integers per line separated by commas. Not a very efficient format! I would suggest you simply pre-process the file to remove all square brackets, linefeeds, and spaces and convert the commas to newlines so that you get one value per line which you can easily deal with.

Using the tr command to transliterate, that would be this:

# Delete all square brackets, newlines and spaces, change commas into newlines
tr -d '[] \n' < YourFile.txt | tr , '\n' > preprocessed.txt

Your file then looks like this and you can readily process one value at a time in Python.

2
2
6
10
2
6
...
...

In case you are on Windows, the tr tool is available for Windows in GNUWin32 and in the Windows Subsystem for Linux thing (git bash?).

You can go still further and make a file that you can memmap() like in the second part of my answer, you could then randomly find any byte in the file. So, taking the preprocessed.txt created above, you can make a binary version like this:

import struct

# Make binary memmapable version
with open('preprocessed.txt', 'r') as ifile, open('preprocessed.bin', 'wb') as ofile:
    for line in ifile:
        ofile.write(struct.pack('q',int(line)))

Original Answer

You can do that like this. The first part is just setup:

#!/usr/bin/env python3

import numpy as np

# Create 2,4 Numpy array of int64
a = np.arange(8, dtype=np.int64).reshape(2,4)

# Write to file as binary
a.tofile('a.dat')

Now check the file by hex-dumping it in the shell:

xxd a.dat

00000000: 0000 0000 0000 0000 0100 0000 0000 0000  ................
00000010: 0200 0000 0000 0000 0300 0000 0000 0000  ................
00000020: 0400 0000 0000 0000 0500 0000 0000 0000  ................
00000030: 0600 0000 0000 0000 0700 0000 0000 0000  ................

Now that we are all set up, let's memmap() the file:

# Memmap file and access values via 'mm'
mm = np.memmap('a.dat', dtype=np.int64, mode='r', shape=(2,4))

print(mm[1,2])      # prints 6

The primary problem is that the file is too large, and it doesn't seem to be split in lines either. (For reference, array.txt is the example you provided and arr_map.dat is an empty file)

import re
import numpy as np 

N = [str(i) for i in range(10)]
arrayfile = 'array.txt'
mmapfile = 'arr_map.dat'
R = 4
C = 17
CHUNK = 20

def read_by_chunk(file, chunk_size=CHUNK):
    return file.read(chunk_size)

fp = np.memmap(mmapfile, dtype=np.uint8, mode='w+', shape=(R,C)) 

with open(arrayfile,'r') as f:
    curr_row = curr_col = 0
    while True:
        data = read_by_chunk(f)
        if not data:
            break

        # Make sure that chunk reading does not break a number
        while data[-1] in N:
            data += read_by_chunk(f,1)

        # Convert chunk into numpy array
        nums = np.array(re.findall(r'[0-9]+', data)).astype(np.uint8)
        num_len = len(nums)

        if num_len == 0:
            break

        # CASE 1: Number chunk can fit into current row
        if curr_col + num_len <= C: 
            fp[curr_row, curr_col : curr_col + num_len] = nums

            curr_col = curr_col + num_len

        # CASE 2: Number chunk has to be split into current and next row
        else: 
            col_remaining = C-curr_col
            fp[curr_row, curr_col : C] = nums[:col_remaining] # Fill in row i

            curr_row, curr_col = curr_row+1, 0                # Move to row i+1 and fill the rest
            fp[curr_row, :num_len-col_remaining] = nums[col_remaining:]

            curr_col = num_len-col_remaining

        if curr_col>=C:
            curr_col = curr_col%C
            curr_row += 1

        #print('\n--debug--\n',fp,'\n--debug--\n')

Basically, read small parts of the array file at a time (making sure not to break the numbers), finding the numbers from the junk characters like commas, brackets etc. using regex, and then inserting the numbers into the memory map.

The situation you describe seems to be more suitable for a generator that fetches the next integer, or the next row from the file and allows you to process that.

def sanify(s):
    while s.startswith('['):
        s = s[1:]
    while s.endswith(']'):
        s = s[:-1]
    return int(s)


def get_numbers(file_obj):
    file_obj.seek(0)
    i = j = 0
    for line in file_obj:
        for item in line.split(', '):
            if item and not item.isspace():
                yield sanify(item), i, j
                j += 1
        i += 1
        j = 0

This ensures only one line at a time ever resides in memory.

This can be used like:

import io


s = '''[[2, 2, 6, 10, 2, 6, 7, 15, 14, 10, 17, 14, 7, 14, 15, 7, 17], 
[3, 3, 7, 11, 3, 7, 0, 11, 7, 16, 0, 17, 17, 7, 16, 0, 0], 
[4, 4, 8, 7, 4, 13, 0, 0, 15, 7, 8, 7, 0, 7, 0, 15, 13], 
[5, 5, 9, 12, 5, 14, 7, 13, 9, 14, 16, 12, 13, 14, 7, 16, 7]]'''


items = get_numbers(io.StringIO(s))
for item, i, j in items:
    print(item, i, j)

If you really want to be able to access an arbitrary element of the matrix, you could adapt the above logic into a class implementing __getitem__ and you would only need to keep track of the position of the beginning of each line. In code, this would look like:

class MatrixData(object):
    def __init__(self, file_obj):
        self._file_obj = file_obj
        self._line_offsets = list(self._get_line_offsets(file_obj))[:-1]
        file_obj.seek(0)
        row = list(self._read_row(file_obj.readline()))
        self.shape = len(self._line_offsets), len(row)
        self.length = self.shape[0] * self.shape[1]


    def __len__(self):
        return self.length


    def __iter__(self):
        self._file_obj.seek(0)
        i = j = 0
        for line in self._file_obj:
            for item in _read_row(line):
                    yield item, i, j
                    j += 1
            i += 1
            j = 0


    def __getitem__(self, indices):
        i, j = indices
        self._file_obj.seek(self._line_offsets[i])
        line = self._file_obj.readline()
        row = self._read_row(line)
        return row[j]


    @staticmethod
    def _get_line_offsets(file_obj):
        file_obj.seek(0)
        yield file_obj.tell()
        for line in file_obj:
            yield file_obj.tell()


    @staticmethod
    def _read_row(line):
        for item in line.split(', '):
            if item and not item.isspace():
                yield MatrixData._sanify(item)


    @staticmethod
    def _sanify(item, dtype=int):
        while item.startswith('['):
            item = item[1:]
        while item.endswith(']'):
            item = item[:-1]
        return dtype(item)


class MatrixData(object):
    def __init__(self, file_obj):
        self._file_obj = file_obj
        self._line_offsets = list(self._get_line_offsets(file_obj))[:-1]
        file_obj.seek(0)
        row = list(self._read_row(file_obj.readline()))
        self.shape = len(self._line_offsets), len(row)
        self.length = self.shape[0] * self.shape[1]


    def __len__(self):
        return self.length


    def __iter__(self):
        self._file_obj.seek(0)
        i = j = 0
        for line in self._file_obj:
            for item in self._read_row(line):
                    yield item, i, j
                    j += 1
            i += 1
            j = 0


    def __getitem__(self, indices):
        i, j = indices
        self._file_obj.seek(self._line_offsets[i])
        line = self._file_obj.readline()
        row = list(self._read_row(line))
        return row[j]


    @staticmethod
    def _get_line_offsets(file_obj):
        file_obj.seek(0)
        yield file_obj.tell()
        for line in file_obj:
            yield file_obj.tell()


    @staticmethod
    def _read_row(line):
        for item in line.split(', '):
            if item and not item.isspace():
                yield MatrixData._sanify(item)


    @staticmethod
    def _sanify(item, dtype=int):
        while item.startswith('['):
            item = item[1:]
        while item.endswith(']'):
            item = item[:-1]
        return dtype(item)

To be used as:

m = MatrixData(io.StringIO(s))

# get total number of elements
len(m)

# get number of row and col
m.shape

# access a specific element
m[3, 12]

# iterate through
for x, i, j in m:
    ...

This seems to be exactly what the mmap module does in python. See: https://docs.python.org/3/library/mmap.html

Example from documentation

import mmap

# write a simple example file
with open("hello.txt", "wb") as f:
    f.write(b"Hello Python!\n")

with open("hello.txt", "r+b") as f:
    # memory-map the file, size 0 means whole file
    mm = mmap.mmap(f.fileno(), 0)
    # read content via standard file methods
    print(mm.readline())  # prints b"Hello Python!\n"
    # read content via slice notation
    print(mm[:5])  # prints b"Hello"
    # update content using slice notation;
    # note that new content must have same size
    mm[6:] = b" world!\n"
    # ... and read again using standard file methods
    mm.seek(0)
    print(mm.readline())  # prints b"Hello  world!\n"
    # close the map
    mm.close()

It's depends on the operation you would like to perform on your input matrix, if it was a matrix operation, then you can use a partial matrix, Most of the time you are able to partially process small batches of your input file as partial matrix, by this way you can process the file very efficient, you just have to develop your algorithm to read and partially process the input and caching the result, for some operations you may just need to decide what is the best representation of your input matrix (i.e. row major or column major).

The main advantage of using the partial matrix approach, that you can take the advantage of applying a parallel processing techniques to process n partial matrix at each iteration using CUDA GPU for example, if you are familiar with C or C++, Then using the Python C API might improve the time complexity a lot for partial matrix operations, but even using Python not too much worse, because you just need to process your partial matrix using Numpy.

Related