My system crashes when trying to apply a function to a pandas dataframe via vectorizing a numpy array

Viewed 320

I have created a process in which I am matching records in the pandas dataframe via names. These names are not necessarily standardized by format, spelling, or anything like that, so this matching is accomplished through a fuzzy matching approach.

Essentially I am trying to take every record's name field, and match those to all other associated records. After some research, it was apparent that passing the matching function though a vectorized numpy array would be the faster approach. However, running my code on even 100 records crashes my system. It works on 1 however. I am not yet knowledgeable enough to know about memory constraints, but my fear is that the memory requirements on this function are being overloaded.

The name field values are a string. I am trying to write a list of matching records to a new column. The "match_name" call is to a separate function that works by using a tuple of the name's elements and creating a phonetic code for the combinations of those elemental names. It then accesses a dictionary that has already been populated with every records phonetic code combinations and returns the unique IDs for where those codes match. This list of record IDs is what I am trying to write into the new field.

There are a few hundred thousand records I am trying to test this on right now.

NOTE: the initial few lines are simply creating the standardized name tuple. They refence some keyword dictionaries to drop unique characters and expand abbreviated names.

def test(in_name):
    name=str(table["name_field"]).upper()
    for clean in NameCleaner:
        name=name.replace(clean ,'')            
    for expand in NamesExpander:
        name = name.replace(expand, NamesExpander[expand])
    name =  re.sub(r"\b[a-zA-Z]\b", "", name)
    name = (normalize_unicode_to_ascii(name)).strip()
    tp = tuple(name.split(" "))
    match_list = match_name(tp)
    
    return match_list

def main(df):
    df["Matches"] = [test(df["name_field"].values)]```
0 Answers
Related