The Price of Hidden Curvature: An Ω˜(d5/4T−−√) Lower Bound for Bandit Convex Optimization

  • Nived Rajaraman

arXiv

We establish a widetildeOmega(d^{5/4}sqrt T)widetildeOmega(d^{5/4}sqrt T) lower bound on the minimax expected regret of stochastic bandit convex optimization of 1-Lipschitz functions on the Euclidean ball. This presents the first nontrivial regret lower bound that grows faster than dT for this problem, establishing that stochastic bandit convex optimization is fundamentally harder than linear bandits. The hard class of convex functions we construct takes the following form in dimension 2d: for an action a=(a1,a2)𝔹22d, each function is the scaled soft maximum of a”tube”, r^{-1} \| W^star a^1 – frac{r}{8varepsilon} a^2 \|_2r^{-1} \| W^star a^1 – frac{r}{8varepsilon} a^2 \|_2 (hyperparameterized by ε,r), and a squared distance function, frac12 \| a^1 – u^star \|_2^2 – frac12 \| u^star \|_2^2frac12 \| a^1 – u^star \|_2^2 – frac12 \| u^star \|_2^2. Here, Wd×d is an unknown linear transformation, and ud is an unknown vector which must be learned to minimize the function. Observations are informative about u only when the learner’s action lies near the tube determined by W, satisfying a28εrWa1: thus the learner must either find this tube without knowing W, or spend observations learning useful directions of W. Formally, our regret analysis exploits this tradeoff by bounding the posterior spread of Fisher information matrices obtained under an adaptive sequence of actions. Together, these ingredients give a sample complexity lower bound of widetilde{Omega}(d^{5/2}/varepsilon^2)widetilde{Omega}(d^{5/2}/varepsilon^2) to find an ε-optimal action, which translates to an widetilde{Omega} (d^{5/4} sqrt{T})widetilde{Omega} (d^{5/4} sqrt{T}) regret lower bound. We also extend this lower bound to the unconstrained setting where the action space is d.