Getting "Performance Warning" when trying to add multiple columns in pandas DataFrame

Viewed 76

Please find below a dataframe:

enter image description here

Logic: For every new entry, first I need to check time if it exists. If it exists, I want to add new column suppose 'vlan3' with some value at the same index ('time') row. If there is no such time present, a new row with another time needs to be added.

I have written a code trying to add multiple columns.

enter image description here

In this code, I am getting an error as below:

enter image description here

Please advise how to add multiple columns satisfying the above mentioned logic.

1 Answers

Here is how you can avoid this warning while adding values one at a time:

from datetime import datetime

import pandas as pd

df = pd.DataFrame(columns=["time"]).set_index("time")

start = pd.to_datetime(datetime.now())

for i in range(288):
    temp_df = pd.DataFrame()
    for j in range(1_000):
        temp_df = pd.concat(
            [
                temp_df,
                pd.DataFrame(
                    data={"vlan" + str(j): 5_465},
                    index=[start + pd.Timedelta(minutes=5 * i)],
                ),
            ],
            axis=1,
        )
    df = pd.concat([df, temp_df], axis=0)

After a couple of minutes:

print(df)
# Output
                            vlan0  vlan1  vlan2  vlan3  vlan4  ...  vlan995  vlan996  vlan997  vlan998  vlan999
2022-06-19 16:21:46.494248   5465   5465   5465   5465   5465  ...     5465     5465     5465     5465     5465
2022-06-19 16:26:46.494248   5465   5465   5465   5465   5465  ...     5465     5465     5465     5465     5465
2022-06-19 16:31:46.494248   5465   5465   5465   5465   5465  ...     5465     5465     5465     5465     5465
2022-06-19 16:36:46.494248   5465   5465   5465   5465   5465  ...     5465     5465     5465     5465     5465
2022-06-19 16:41:46.494248   5465   5465   5465   5465   5465  ...     5465     5465     5465     5465     5465
...                           ...    ...    ...    ...    ...  ...      ...      ...      ...      ...      ...
2022-06-20 15:56:46.494248   5465   5465   5465   5465   5465  ...     5465     5465     5465     5465     5465
2022-06-20 16:01:46.494248   5465   5465   5465   5465   5465  ...     5465     5465     5465     5465     5465
2022-06-20 16:06:46.494248   5465   5465   5465   5465   5465  ...     5465     5465     5465     5465     5465
2022-06-20 16:11:46.494248   5465   5465   5465   5465   5465  ...     5465     5465     5465     5465     5465
2022-06-20 16:16:46.494248   5465   5465   5465   5465   5465  ...     5465     5465     5465     5465     5465

[288 rows x 1000 columns]

Beyond not getting the warning, you can confirm that this is more efficient memory wise by profiling the code:

Line #    Mem usage Occurrences   Line Contents
===============================================
    11                           def func(df):
    12     75.1 MiB          1       start = pd.to_datetime(datetime.now())
    13    104.9 MiB        289       for i in range(288):
    14    104.9 MiB        288           temp_df = pd.DataFrame()
    15    104.9 MiB     288288           for j in range(1_000):
    16    104.9 MiB     576000               temp_df = pd.concat(
    17    104.9 MiB     288000                   [
    18    104.9 MiB     288000                       temp_df,
    19    104.9 MiB     576000                       pd.DataFrame(
    20    104.9 MiB     288000                           data={"vlan" + str(j): 5_465},
    21    104.9 MiB     288000                           index=[start + pd.Timedelta(minutes=5 * i)],
    22                                               ),
    23                                           ],
    24    104.9 MiB     288000                   axis=1,
    25                                       )
    26    104.9 MiB        288           df = pd.concat([df, temp_df], axis=0)
    27    104.8 MiB          1       return df
Related