innovator · thinks

Reinforcement Learning Basics

Discover how agents learn from rewards and penalties to master environments. Build your own RL toy worlds and watch intelligent behavior emerge.

20 modules·Difficulty: ★★★★★· 5 Free
Start module 1

Modules

  • 1
    Free8 min
    What Is Reinforcement Learning?
    Meet agents that learn by trial and error, just like you learn to ride a bike 🚴
  • 2
    Free9 min
    Agents, Environments, and States
    Discover the building blocks: who acts, where they act, and what they observe.
  • 3
    Free10 min
    Actions and Rewards
    Learn how agents choose moves and receive feedback signals to improve over time.
  • 4
    Free11 min
    Reward Shaping Fundamentals
    Design reward signals that guide your agent toward smart behavior without confusing it.
  • 5
    Free10 min
    Exploration vs. Exploitation
    Balance trying new things with sticking to what works—a core dilemma in decision-making.
  • 6
    Paid12 min
    Policies and Action Selection
    Policies map states to actions, acting as the brain of your agent.
  • 7
    Paid11 min
    Episodes and Trajectories
    Understand how sequences of states and actions form complete learning episodes.
  • 8
    Paid10 min
    Discounting Future Rewards
    Learn why agents value immediate rewards more than distant ones using discount factors.
  • 9
    Paid12 min
    Returns and Value Functions
    Calculate the total expected reward from any state using value functions.
  • 10
    Paid11 min
    Building a Simple Grid World
    Code your first environment where an agent navigates a grid to reach a goal.
  • 11
    Paid10 min
    Q-Values and Action Quality
    Assign scores to state-action pairs to identify the best moves.
  • 12
    Paid12 min
    Introduction to Q-Learning
    Implement the classic algorithm that updates Q-values from experience.
  • 13
    Paid11 min
    Temporal Difference Learning
    Learn step-by-step from immediate feedback without waiting for episode ends.
  • 14
    Paid10 min
    Epsilon-Greedy Strategy
    Balance random exploration with greedy exploitation using a simple parameter.
  • 15
    Paid12 min
    Training a Grid-World Agent
    Watch your agent learn optimal paths through thousands of practice episodes.
  • 16
    Paid11 min
    Debugging Reward Signals
    Identify and fix reward structures that cause unintended agent behaviors.
  • 17
    Paid10 min
    Adding Obstacles and Penalties
    Enrich your environment with hazards that teach smarter navigation strategies.
  • 18
    Paid12 min
    Multi-Goal Environments
    Design scenarios where agents must choose between competing objectives dynamically.
  • 19
    Paid11 min
    Capstone: Your RL Toy World
    Combine all concepts to build a custom environment and train an agent from scratch.
  • 20
    Final exam25 min
    Final Assessment: Reinforcement Learning Mastery
    Prove your understanding of agents, policies, rewards, and Q-learning fundamentals.