I use stable baseline 3. The model.learn function in SB3 is to generate one action according to the state and then obtain the reward, then the model is trained. However, if I have multiple state-action-reward pairs generated by an (old) model, can I feed these pairs to RL and let it learn from them and get a new trained model? It seems I can't do it using the current model.learn function because the current workflow of model.learn is: (old model) state->action->reward (get a new model) state->action->reward (get a new model)... But what I want is: (old model) multiple state-action-reward pairs (get a new model). I can't use the current model.learn function because after it learned from one state-action pair, the model will be updated, and the generated action may be different from the action generated by the old model. Anyone can help me? Thanks in advance.
I have read the [documentation](https://stable-baselines3.readthedocs.io/en/master/modules/dqn.html?highlight=model.learn(#stable_baselines3.dqn.DQN.learn)