Uploaded December 2024 | Updated September 2026, 2 weeks ago
Not my proudest work from a visual point of view - especially not since a GPT helped me a lot with the coding and bug fixing. That is, a machine helped me to learn a machine.
Left: A very simple maze that an reinforcement, RL, learning agent improves its ability to find the fastest way between start (upper left) and goal (lower right) utilizing a grid-based Q-learning simulation.
Right: The number of steps the agent takes in each episode.
The excellent song is a reuse from a previous video, made by @gpcbass. It's called Swing Zero.
Not my proudest work from a visual point of view - especially not since a GPT helped me a lot with the coding and bug fixing. That is, a machine helped me to learn a machine.
Left: A very simple maze that an reinforcement, RL, learning agent improves its ability to find the fastest way between start (upper left) and goal (lower right) utilizing a grid-based Q-learning simulation.
Right: The number of steps the agent takes in each episode.
The excellent song is a reuse from a previous video, made by @gpcbass. It's called Swing Zero.










