How to dynamically assign column names with Pandas?

Viewed 278

I am facing the following problem:

I need to assign column to a dataframe, based on applying a function for each row. However, the column names are hard coded and I would like to dynamically create them, in order to improve maintenance for the code.

Little background what the apply function does:

It takes values around a certain point ('location') from another dataframe and stores them in a numpy array (which are later on be used for cross correlation, but that's out of the scope of this question).

The current code looks like this:

def get_yearly_values(df,
                      df_2018, 
                      df_2017, 
                      df_2016):

    return df.assign(
        values_1_df_2018=df.apply(lambda d: get_value_range(d, 'location', df_2018, 'Value_1'), axis=1),
        values_2_df_2018=df.apply(lambda d: get_value_range(d, 'location', df_2018, 'Value_2'), axis=1),
        values_1_df_2017=df.apply(lambda d: get_value_range(d, 'location', df_2017, 'Value_1'), axis=1),
        values_2_df_2017=df.apply(lambda d: get_value_range(d, 'location', df_2017, 'Value_2'), axis=1),
        values_1_df_2016=df.apply(lambda d: get_value_range(d, 'location', df_2016, 'Value_1'), axis=1),
        values_2_df_2016=df.apply(lambda d: get_value_range(d, 'location', df_2016, 'Value_2'), axis=1),
    )

What I would like is to maintain one list of the years (and when a new yearly data is available, I just add the new year to the list in chronological order) and dynamically create the column names.

years = [df_2018, df_2017, df_2016]

Now, I know Pandas' assign doesn't work with string formatting, but I would like to do something like this:

def get_yearly_values(df,
                      years[0], 
                      years[1], 
                      years[2]):

    return df.assign(
        f'values_1_df_{years[0]}'=df.apply(lambda d: get_value_range(d, 'location', years[0], 'Value_1'), axis=1),
        f'values_2_df_{years[0]}'=df.apply(lambda d: get_value_range(d, 'location', years[0], 'Value_2'), axis=1),
        f'values_1_df_{years[1]}'=df.apply(lambda d: get_value_range(d, 'location', years[1], 'Value_1'), axis=1),
        f'values_2_df_{years[1]}'=df.apply(lambda d: get_value_range(d, 'location', years[1], 'Value_2'), axis=1),
        f'values_1_df_{years[2]}'=df.apply(lambda d: get_value_range(d, 'location', years[2], 'Value_1'), axis=1),
        f'values_2_df_{years[2]}'=df.apply(lambda d: get_value_range(d, 'location', years[2], 'Value_2'), axis=1),
    )

How can I make this work?

1 Answers

Not sure if I understand your "get_value_range" function correctly but this would be what I suggest.

    def get_yearly_values(df,
                      years[0], 
                      years[1], 
                      years[2]):
   
    kwards = [_ for _ in range(len(years)*2)]
    
    for i in range(len(years)):
        kwards[i] = { f'values_1_df_{years[i]}' : lambda x : get_value_range(d, 'location', years[i], 'Value_1')
        kwards[i+1] = { f'values_2_df_{years[i]}' : lambda x : get_value_range(d, 'location', years[0], 'Value_2')

    for i in range(len(kwards)):
        df.assign(**kwards[i])

    return df
    )
Related