is actor-critic agent learning?

karim bio gassi

21 Ag. 2021

1 Respuesta

Actualizado a las 4 Oct. 2022

3 Visualizaciones (30 días)

Iniciar sesión para responder a esta pregunta.

Iniciar sesión para seguir la actividad

Iniciar sesión para responder a esta pregunta.

Iniciar sesión para seguir la actividad

Mostrar comentarios más antiguos

0 votos

I built a actor critic agent for microgrid energy management. it has to decide the discharging/charging energy among a set of action

in total 9 action can be taken for 7008 time steps. I am training the agent over 2000 episodes. But I notice when the agent start learning at a cetain episodes,

at the next episodes it fall completely down. I tattached the training for the first 250 episodes.

I wonder if there something wrong in my code.

0 comentarios
Mostrar -2 comentarios más antiguos Ocultar -2 comentarios más antiguos

Iniciar sesión para comentar.

Iniciar sesión para responder a esta pregunta.

Iniciar sesión para seguir la actividad

Respuestas (1)

Ahmed R. Sayed el 4 de Oct. de 2022

0 votos

Hi, karim bio gassi,

From your figure, the discounted reward value is very large. try to rescale it to a certain value [-10, 10] in the environment. For example, r(t) = 10 * Microgrid operational cost (t) / MaxCost , where MaxCost is the maximum possible cost per time step.

Another point is you can use another agent.

I hope these suggestions can solve your concerns.