Reinforcement Learning Fundamentals
This course introduces the fundamental concepts, algorithms, and real-world applications of reinforcement learning (RL). Topics include Multi-Armed Bandits (value functions, UCB, Thompson Sampling), Markov Decision Processes (Bellman equations, policy and value iteration), Monte Carlo Methods (on-policy and off-policy learning), and Temporal Difference Learning (TD(0), SARSA, Q-Learning). The course also covers Policy Gradient Methods (REINFORCE, Policy Gradient Theorem), Deep Reinforcement Learning (DQN, Actor-Critic), and advanced topics including Multiagent RL and Inverse RL. Through coding assignments, projects, and real-world case studies, students develop a strong theoretical foundation and hands-on experience applying RL techniques to AI, robotics, and autonomous systems.
Course Overview
RO3002: Reinforcement Learning Fundamentals provides a comprehensive introduction to RL, covering key concepts, algorithms, and real-world applications. Students will explore Multi-Armed Bandits (value functions, UCB, Thompson Sampling), Markov Decision Processes (MDP) (Bellman equations, policy and value iteration), Monte Carlo Methods (on-policy and off-policy learning), and Temporal Difference (TD) Learning (TD(0), SARSA, Q-Learning). The course also covers Policy Gradient Methods (REINFORCE, Policy Gradient Theorem), Deep Reinforcement Learning (DQN, Actor-Critic), and advanced topics like Multiagent RL and Inverse RL. Through coding assignments, projects, and real-world case studies, students will develop a strong theoretical foundation and hands-on experience in applying RL techniques to AI, robotics, and autonomous systems.
Learning Objectives
By the end of this course, each student will have had the opportunity to:
- Understand fundamental RL principles and algorithms.
- Implement RL models using Python and deep learning frameworks.
- Design and apply RL techniques to real-world decision-making problems.
- Explore advanced topics like Deep RL, Multiagent RL, and Inverse RL.
This course is ideal for students interested in AI, robotics, autonomous systems, and data-driven decision-making.
Learning Outcomes
By the end of this course, each student should have done the following,
- Implemented Core RL Algorithms – Developed and coded Multi-Armed Bandits, Q-Learning, SARSA, Monte Carlo Methods, and Policy Gradient algorithms.
- Built and Trained Deep RL Models – Designed Deep Q-Networks (DQN) and Actor-Critic models for complex decision-making tasks.
- Modeled Real-World Problems Using RL – Applied RL techniques to robotics, autonomous systems, finance, and healthcare through hands-on projects.
- Learnt about Exploration and Exploitation – Designed strategies to optimize decision-making in uncertain environments.
- Learnt about Multiagent and Inverse RL – Explored multiagent reinforcement learning and learning from expert demonstrations (Inverse RL).
- Completed a Capstone RL Project – Designed, implemented, and presented a comprehensive RL-based solution addressing a practical challenge.
- Analyzed and Optimized RL Models – Assessed the performance, stability, and scalability of different RL approaches.
- Collaborated on Research and Development – Engaged in team-based projects and discussions on cutting-edge RL advancements.
Recommended Textbooks
- An Introduction to Reinforcement Learning, by Sutton and Barto, MIT Press, Second edition
Available free online!
https://www.andrew.cmu.edu/course/10-703/textbook/BartoSutton.pdf
Additional Readings
- Algorithms for Reinforcement Learning, by Szepesvari, Morgan and Claypool Publishers, 2010
Available free online!
https://sites.ualberta.ca/~szepesva/papers/RLAlgsInMDPs.pdf
