Simulate electric vehicle (EV) charging demand in Python

Viewed 100

I'm trying to simulate electric vehicle charging demand in python, given the vehicle's energy consumption per second, a charging rate for the charger, and a starting state of charge. The true end goal is to be able to create a graph of hourly charging profile over the course of the dataset (which stretches a couple of months).

I have a per-second dataset that is very large (~6 million rows), each row currently has a timestamp (date time object), energy consumption (kWh), speed (m/s), and indicator variable for whether the vehicle is Stopped or Moving. The dataset looks something like this:

DateTime Average speed (m/s) Status Energy Consumption (kWh)
2022-01-01-01:00:00 0.0 Stopped 0.0
2022-01-01-01:00:01 0.0 Stopped 0.0
2022-01-01-01:00:02 0.0 Stopped 0.0
2022-01-01-01:00:03 5.0 Moving 0.0050
2022-01-01-01:00:04 6.2 Moving 0.0062

My goal is to add columns that include vehicle battery state of charge (kWh) and vehicle charging demand (kWh) (assume for this example that the battery starts full at 100 kWh). I want to have it so that the vehicle charges if it is stopped and the battery is below 100%.

DateTime Average speed (m/s) Status Energy Consumption (kWh) State of Charge (%) Charging Demand (kWh)
2022-01-01-01:00:00 0.0 Stopped 0.0 100 0.0
2022-01-01-01:00:01 0.0 Stopped 0.0 100 0.0
2022-01-01-01:00:02 0.0 Stopped 0.0 100 0.0
2022-01-01-01:00:03 5.0 Moving 0.0050 99.995 0.0
2022-01-01-01:00:04 6.2 Moving 0.0062 99.988 0.0
2022-01-01-01:00:05 3.8 Moving 0.0038 99.950 0.0
2022-01-01-01:00:06 1.5 Moving 0.0015 99.935 0.0
2022-01-01-01:00:07 0.0 Stopped 0.0 99.960 0.0061
2022-01-01-01:00:08 0.0 Stopped 0.0 99.9661 0.0061
2022-01-01-01:00:09 0.0 Stopped 0.0 99.9722 0.0061
2022-01-01-01:00:10 0.0 Stopped 0.0 99.9783 0.0061

etc...

I have tried a normal for-loop and for loop with iterrows to run this as a simulation, but the dataset is ~6 million rows, and I have multiple datasets to run this on, so it would take far too long. Here are the attempts:

starting_state_of_charge = 100 # %
charger_rating = 22 #kW

df['State of Charge (%)'] = starting_state_of_charge #initialize state of charge column
df['Charging Demand (kWh)'] = 0 #initialize charging demand column

for i in range(len(df)):
# update state of charge every time energy is consumed by the vehicle
    df['State of Charge (%)'][i] = df['State of Charge (%)'][i - 1] - df['Energy Consumption (kWh)'][i]

# if the vehicle is not moving and the state of charge is less than 100%, then charge the vehicle
    if df['Status'][i] == 'Stopped' and df['State of Charge (%)'][i] < 100:
        charging_rate = charger_rating/3.6e3 #unit conversion to per second charging rate
        df['Charging Demand (kWh)'][i] = charging_rate
        df['State of Charge (%)'][i] = min(100,df['State of Charge (%)'][i] + charging_rate)
starting_state_of_charge = 100 # %
charger_rating = 22 #kW

df['State of Charge (%)'] = starting_state_of_charge #initialize state of charge column
df['Charging Demand (kWh)'] = 0 #initialize charging demand column

for idx, row in df.iterrows():
    if idx != 0:
        df.loc[idx,'State of Charge (%)'] = df.loc[idx - 1, 'State of Charge (%)'] - df.loc[idx,'Energy Consumption (kWh)']
        if df.loc[idx, 'Status'] == 'Stopped' and df.loc[idx,'State of Charge (%)'] < 100:
            charging_rate = charger_rating/3.6e3
            df.loc[idx,'Charging Demand (kWh)'] = charging_rate
            df.loc[idx,'State of Charge (%)']+= min(100, df['State of Charge (%)'][i] + charging_rate)

It seems the normal solutions for a large data frame are to use the 'apply' or 'itertools' functions, but I haven't figured out how to do so. The problem is that the State of Charge column and Charging Demand columns that I want to create would depend on each other, even though neither exists yet. Specifically, the State of Charge and would increase every time there is positive Charging Demand, but the vehicle can only charge into the battery if the State of Charge is less than 100%.

How can I run this simulation with python code?

2 Answers

Have you tried slicing the dataframe into subsets of 1000 rows each and then running your code on each subset?

Here's an approach that might work. I haven't tested it with a dataframe as large as yours, so am not sure about the performance. Assuming your 'DateTime' column is already formatted as a pd Timestamp, do the following:

Step 1. Add a new column to the frame which determines the time difference from the start using:

df['Time_diff'] = [int((x - pd.to_datetime('2022-01-01-01:00:00')).total_seconds()//60) for x in df['DateTime'].to_list()] 

This adds a column 'Time_diff' which contains the total number of minutes since start of simulation as an integer.

Step 2. Now group the data frame by 'Time_diff' and 'Status' using:

df.groupby(['Time_diff', 'Status']).sum()  

This produces a view of the data frame which is organized by each minute and has a sum of the energy consumption by 'Status'. Similar to the following:

                  Average speed (m/s)   Energy Consumption (kWh)
Time_diff   Status      
0           Moving  39.2                0.0392
            Stopped 0.0                 0.0000
60          Moving  5.0                 0.0050
180         Moving  3.8                 0.0038
300         Moving  1.5                 0.0015  

This reduces the size of your data frame to two lines/minute and maintains the consumption by minute for each minute of the simulation. Which is a net improvement of at least 96%

Related