Pandas: Getting indices (numeric position) from external array for each value in Column

Viewed 30

I have an fixed value with arrays: ['string1', 'string2', 'string3'] and a Pandas Datafrae:

>>> pd.DataFrame({'column': ['string1', 'string1', 'string2']})
    column
0  string1
1  string1
2  string2

And I want to add a new column with the indices position from the previous array, so it becomes:

>>> pd.DataFrame({'column': ['string1', 'string1', 'string2', pd.NA], 'indices': [0,0,1, pd.NA]})
    column indices
0  string1       0
1  string1       0
2  string2       1
3     <NA>    <NA>

I.e the position of the value in the main array. This will be later fed into pyarrow's DictionaryArray[1]. The Dataframe can have null values as well

Is there any fast way to do this? Been trying to figure out how to vectorize it. Naive implementation:

def create_dictionary_array_indices(column_name, arrow_array):
    global dictionary_values
    values = arrow_array.to_pylist()
    indices = []
    for i, value in enumerate(values):
        if not value or value != value:
            indices.append(None)
        else:
            indices.append(
                dictionary_values[column_name].index(value)
            )
    indices = pd.array(indices, dtype=pd.Int32Dtype())
    return pa.DictionaryArray.from_arrays(indices, dictionary_values[column_name])

[1] https://lists.apache.org/thread/xkpyb3zboksbhmyqzzkj983y6l0t9bjs

1 Answers

Given your two dataframes:

import pandas as pd

df1 = pd.DataFrame({"column": ["string1", "string1", "string2"]})
df2 = pd.DataFrame({"column": ["string1", "string1", "string2", pd.NA]})

Here is one way to do it:

df1 = df1.drop_duplicates(keep="first").reset_index(drop=True)
indices = {value: key for key, value in df1["column"].items()}

df2["indices"] = df2["column"].apply(lambda x: indices.get(x, pd.NA))
print(df2)
# Output
    column indices
0  string1       0
1  string1       0
2  string2       1
3     <NA>    <NA>
Related