PhD defence by Chloé Alice Florence Rouyer
Summary
Originally, this framework was studied by making the strong assumption that the environment provides easy data. The opposite scenario where the learner forgoes any assumption on the data and runs as if the environment will provide worst-case data has also been studied and led to the development of robust but slower learning algorithms. Recently the goal has been to derive algorithms that adapt to both worst-case and easy data simultaneously. In each case, the learner is faced with different constraints on her behavior and on the feedback she is allowed to observe, which affect the trade-off between exploration and exploitation in different ways. We propose algorithms for each of those problems and analyze them to provide theoretical guarantees against both worst-case and easy data.