This talk has moved to Fall 2016 — Contextual bandit and off-policy evaluation: minimax bounds and new algorithm.