ESIEE Paris, 5th year (work-study program), 2026-2027.
A system learns sequentially when its decisions determine the data it observes next: recommending an article, setting a price, choosing the answer of an assistant. The course starts with multi-armed bandits, where the whole problem is the trade-off between exploration and exploitation, then moves to reinforcement learning, and ends with the alignment of language models from human preferences. It consists of 10 sessions of 3 hours, each combining a lecture, a tutorial and a lab. All materials are in English.
| Session | Topic | Materials |
|---|---|---|
| 1 | The K-armed bandit: regret, Hoeffding’s inequality, greedy, epsilon-greedy, explore-then-commit | slides · handout · tutorial sheet |
| 2 | Optimism: UCB | |
| 3 | Thompson sampling and linear bandits | |
| 4 | Pure exploration, experimental design and Gaussian processes | |
| 5 | Exponential weights and off-policy evaluation | |
| 6 | Learning from preferences | |
| 7 | Markov decision processes and Q-learning | |
| 8 | Policy gradient | |
| 9 | Preferences and RLHF | |
| 10 | Projects |
The materials of each session are added after the session.
The labs are Jupyter notebooks in Python. They run on a laptop CPU, or online with Google Colab.
Download the lab folder (zip archive). It contains:
SETUP.pdf: installation page, to read first;lab01-02.ipynb: Labs 1 and 2 (naive strategies and UCB), and lab00.ipynb,
an optional warm-up on Python, numpy and matplotlib;-colab version of each notebook, which runs on Google Colab without
installation;seqlearn/, the Python package used in the labs, environment.yml and
check_install.py.