Sequential Decision

ESIEE Paris, 5th year (work-study program), 2026-2027.

A system learns sequentially when its decisions determine the data it observes next: recommending an article, setting a price, choosing the answer of an assistant. The course starts with multi-armed bandits, where the whole problem is the trade-off between exploration and exploitation, then moves to reinforcement learning, and ends with the alignment of language models from human preferences. It consists of 10 sessions of 3 hours, each combining a lecture, a tutorial and a lab. All materials are in English.

Materials

SessionTopicMaterials
1The K-armed bandit: regret, Hoeffding’s inequality, greedy, epsilon-greedy, explore-then-commitslides · handout · tutorial sheet
2Optimism: UCB
3Thompson sampling and linear bandits
4Pure exploration, experimental design and Gaussian processes
5Exponential weights and off-policy evaluation
6Learning from preferences
7Markov decision processes and Q-learning
8Policy gradient
9Preferences and RLHF
10Projects

The materials of each session are added after the session.

Labs

The labs are Jupyter notebooks in Python. They run on a laptop CPU, or online with Google Colab.

Download the lab folder (zip archive). It contains:

  • SETUP.pdf: installation page, to read first;
  • lab01-02.ipynb: Labs 1 and 2 (naive strategies and UCB), and lab00.ipynb, an optional warm-up on Python, numpy and matplotlib;
  • a -colab version of each notebook, which runs on Google Colab without installation;
  • seqlearn/, the Python package used in the labs, environment.yml and check_install.py.

References

  • T. Lattimore, C. Szepesvári, Bandit Algorithms, Cambridge University Press, 2020.
  • A. Slivkins, Introduction to Multi-Armed Bandits, Foundations and Trends in Machine Learning, 2019.
  • R. Sutton, A. Barto, Reinforcement Learning: An Introduction, 2nd edition, MIT Press, 2018.
  • N. Lambert, Reinforcement Learning from Human Feedback, 2025.