Python numpy array add / update / delete row on value from other array

Viewed 544

I have a numpy array and depending on the value from another array, I would like to either update the value of the row, or delete it, or add one.

Example:

I have arr, the one with all values and to keep updated with value from new_arr. If a value in the first column of new_arr exists in arr, then the second column of arr is updated. If the value does no exist, then add a new row. If the second column in new_arr == 0, then delete the row in arr with the matching first column.

arr = np.array([[1, 10],
                [2, 15],
                [3,  5],
                [4, 10]])

new_arr = np.array([[2, 20], # 2 exists in arr and 20 > 0 --> update in arr
                    [5, 20], # 5 does not exists in arr --> add row in arr
                    [1, 0]]) # 1 exists in arr but col 2 == 0--> delete row in arr

Then I would like to obtain:

arr = np.array([[2, 20],
                [3,  5],
                [4, 10],
                [5, 20]])

Observe that arr is ordered by the first column. Also arr has a maximum lenght of 1000 rows.

Any simple and fast method please?

Initially arr and new_arr are lists. I've turned them into numpy arrays. However, as I do not do any strong calculation with arr, most likely it would be faster to keep it as a list.

2 Answers

Keeping the input arrays as numpy constructs, here's how I would do it.

def process_arrays(np1, np2):
    np1d = dict((np1[x][0], np1[x][1]) for x in range(len(np1)))
    np2d = dict((np2[x][0], np2[x][1]) for x in range(len(np2)))
    for ky2 in np2d.keys():
        if ky2 in np1d.keys():
            if np2d[ky2] == 0:
                del np1d[ky2]
            else:
                np1d[ky2] = np2d[ky2]
        else:
            np1d[ky2] = np2d[ky2]
    return np.array(np1d)   

Given you input executing:

process_arrays(arr, newArr)  

Yields:

array({2: 20, 3: 5, 4: 10, 5: 20}, dtype=object)

You really want three different operations, each of which is easy to implement. Adding and deleting allocates new arrays, so you want to do those in bulk one time. Your goal is therefore mostly to split the new data into three portions.

First identify and split the delete portion:

mask = new_arr[:, -1] == 0
to_del = new_arr[mask, :]
to_add_update = new_arr[~mask, :]

Now you can find the insertion indices of the add and update portions:

insert_index = np.searchsorted(arr[:, 0], to_add_update[:, 0])

The elements whose insertion indices match between the arrays are places where you want to update vs the ones you want to insert. Let's define a function for this since we can use it twice:

 def get_insert_update_index(a, v):
     """
     Get new and existing insertion indices.

     Parameters
     ----------
     a :
         The array to insert into.
     v :
         The values to insert

     Returns
     -------
     mask :
         True indicates new elements of `v`.
     insert :
         Insertion indices of new elements in `a`
     update :
         Indices if existing elements in `a`
     """
     index = np.searchsorted(a, v)
     mask = index >= a.size
     mask2 = a[index[~mask]] != v[~mask]
     mask[~mask] = mask2
     return mask, index[mask], index[~mask]

mask, insert_index, update_index = get_insert_update_index(arr[:, 0], to_add_update[:, 0])

If you process the updates first, your indices won't change:

arr[update_index, -1] = to_add_update[~mask, -1]

Processing an update involves making a new array:

arr = np.insert(arr, insert_index, to_add_update[mask, :], axis=0)

Now that you've done the insertion, you will need to recompute the indices of the deletions you found in the beginning. You probably don't want to remove the non-existent stuff, so get_insert_update_index is going to come in handy again:

_, _, delete_index = get_insert_update_index(arr[:, 0], to_del[:, 0])
arr = np.delete(are, delete_index, axis=0)
Related