numpy broadcasting on pandas dataframe gives memory error

Viewed 125

I have two data frames. Dataframe A is of shape (1269345,5) and dataframe B is of shape (18583586, 3).

Dataframe A looks like:

Name.   gender     start_coordinate    end_coordinate    ID      
Peter     M             30                  150           1      
Hugo      M            4500                6000           2      
Jennie    F             300                 700           3   

Dataframe (B) looks like

ID_sim.  position      string      
  1         89            aa      
  4         568            bb     
  5        938437         cc

I want to make extract rows and make two data frames for which position column in dataframe B falls in the interval (specified by start_coordinate and end_coordinate column) in dataframe A.So resulting dataframe would look like:

###Final dataframe A
Name.   gender     start_coordinate    end_coordinate    ID      
Peter     M             30                  150           1 
Jennie    F             300                 700           3  


###Final dataframe B

ID_sim.    position     string          
   1          89           aa 
   4          568           bb 

I tried using numpy broadcasting like this:

s, e = dfA[['start_coordinate', 'end_coordinate']].to_numpy().T
p = dfB['position'].to_numpy()[:, None]

dfB[((p >= s) & (p <= e)).any(1)]

But this gave me the following error:

MemoryError: Unable to allocate 2.72 TiB for an array with shape (18583586, 160711) and data type bool

I think its because my numpy becomes quite large when I try broadcasting. How can I achieve my task without numpy broadcasting considering that my dataframes are very large. Insights will be appreciated.

1 Answers

This is likely due to your system overcommit mode.

It will be 0 by default,

Heuristic overcommit handling. Obvious overcommits of address space are refused. Used for a typical system. It ensures a seriously wild allocation fails while allowing overcommit to reduce swap usage. The root is allowed to allocate slightly more memory in this mode. This is the default.

By Running below command to check your current overcommit mode

$ cat /proc/sys/vm/overcommit_memory
0

In this case, you're allocating

> 156816 * 36 * 53806 / 1024.0**3
282.8939827680588

~282 GB and the kernel is saying well obviously there's no way I'm going to be able to commit that many physical pages to this, and it refuses the allocation.

If (as root) you run:

$ echo 1 > /proc/sys/vm/overcommit_memory 

This will enable the "always overcommit" mode, and you'll find that indeed the system will allow you to make the allocation no matter how large it is (within 64-bit memory addressing at least).

I tested this myself on a machine with 32 GB of RAM. With overcommit mode 0 I also got a MemoryError, but after changing it back to 1 it works:

>>> import numpy as np 
>>> a = np.zeros((156816, 36, 53806), dtype='uint8')
>>> a.nbytes

303755101056

You can then go ahead and write to any location within the array, and the system will only allocate physical pages when you explicitly write to that page. So you can use this, with care, for sparse arrays.

Related