Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates. Mei, J., Dai, B., Agarwal, A., Vaswani, S., Raj, A., Szepesvári, C., & Schuurmans, D. In NeurIPS, 2024.
Url
Pdf abstract bibtex 4 downloads We provide a new understanding of the stochastic gradient bandit algorithm by showing that it converges to a globally optimal policy almost surely using any constant learning rate. This result demonstrates that the stochastic gradient algorithm continues to balance exploration and exploitation appropriately even in scenarios where standard smoothness and noise control assumptions break down. The proofs are based on novel findings about action sampling rates and the relationship between cumulative progress and noise, and extend the current understanding of how simple stochastic gradient methods behave in bandit settings.
@inproceedings{MeDaAgVaRaSzeSch24,
author = {Jincheng Mei and Bo Dai and Alekh Agarwal and Sharan Vaswani and Anant Raj and Csaba Szepesv\'ari and Dale Schuurmans},
title = {Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates},
crossref = {NeurIPS2024poster},
booktitle = {NeurIPS},
year = {2024},
url_url = {http://papers.nips.cc/paper_files/paper/2024/hash/8815c983a1f46b477fe4fbf7042a3ba3-Abstract-Conference.html},
url_pdf = {https://proceedings.neurips.cc/paper_files/paper/2024/file/8815c983a1f46b477fe4fbf7042a3ba3-Paper-Conference.pdf},
abstract = {We provide a new understanding of the stochastic gradient bandit algorithm by showing that it converges to a globally optimal policy almost surely using any constant learning rate. This result demonstrates that the stochastic gradient algorithm continues to balance exploration and exploitation appropriately even in scenarios where standard smoothness and noise control assumptions break down. The proofs are based on novel findings about action sampling rates and the relationship between cumulative progress and noise, and extend the current understanding of how simple stochastic gradient methods behave in bandit settings.},
keywords = {stochastic multi-armed bandits, policy gradient, stochastic gradient bandit algorithm, large stepsize, convergence},
}
Downloads: 4
{"_id":"FpDyfgLRca29bQrAX","bibbaseid":"mei-dai-agarwal-vaswani-raj-szepesvri-schuurmans-smallstepsnomoreglobalconvergenceofstochasticgradientbanditsforarbitrarylearningrates-2024","author_short":["Mei, J.","Dai, B.","Agarwal, A.","Vaswani, S.","Raj, A.","Szepesvári, C.","Schuurmans, D."],"bibdata":{"bibtype":"inproceedings","type":"inproceedings","author":[{"firstnames":["Jincheng"],"propositions":[],"lastnames":["Mei"],"suffixes":[]},{"firstnames":["Bo"],"propositions":[],"lastnames":["Dai"],"suffixes":[]},{"firstnames":["Alekh"],"propositions":[],"lastnames":["Agarwal"],"suffixes":[]},{"firstnames":["Sharan"],"propositions":[],"lastnames":["Vaswani"],"suffixes":[]},{"firstnames":["Anant"],"propositions":[],"lastnames":["Raj"],"suffixes":[]},{"firstnames":["Csaba"],"propositions":[],"lastnames":["Szepesvári"],"suffixes":[]},{"firstnames":["Dale"],"propositions":[],"lastnames":["Schuurmans"],"suffixes":[]}],"title":"Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates","crossref":"NeurIPS2024poster","booktitle":"NeurIPS","year":"2024","url_url":"http://papers.nips.cc/paper_files/paper/2024/hash/8815c983a1f46b477fe4fbf7042a3ba3-Abstract-Conference.html","url_pdf":"https://proceedings.neurips.cc/paper_files/paper/2024/file/8815c983a1f46b477fe4fbf7042a3ba3-Paper-Conference.pdf","abstract":"We provide a new understanding of the stochastic gradient bandit algorithm by showing that it converges to a globally optimal policy almost surely using any constant learning rate. This result demonstrates that the stochastic gradient algorithm continues to balance exploration and exploitation appropriately even in scenarios where standard smoothness and noise control assumptions break down. The proofs are based on novel findings about action sampling rates and the relationship between cumulative progress and noise, and extend the current understanding of how simple stochastic gradient methods behave in bandit settings.","keywords":"stochastic multi-armed bandits, policy gradient, stochastic gradient bandit algorithm, large stepsize, convergence","bibtex":"@inproceedings{MeDaAgVaRaSzeSch24,\n author = {Jincheng Mei and Bo Dai and Alekh Agarwal and Sharan Vaswani and Anant Raj and Csaba Szepesv\\'ari and Dale Schuurmans},\n title = {Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates},\n crossref = {NeurIPS2024poster},\n booktitle = {NeurIPS},\n year = {2024},\n url_url = {http://papers.nips.cc/paper_files/paper/2024/hash/8815c983a1f46b477fe4fbf7042a3ba3-Abstract-Conference.html},\n url_pdf = {https://proceedings.neurips.cc/paper_files/paper/2024/file/8815c983a1f46b477fe4fbf7042a3ba3-Paper-Conference.pdf},\n abstract = {We provide a new understanding of the stochastic gradient bandit algorithm by showing that it converges to a globally optimal policy almost surely using any constant learning rate. This result demonstrates that the stochastic gradient algorithm continues to balance exploration and exploitation appropriately even in scenarios where standard smoothness and noise control assumptions break down. The proofs are based on novel findings about action sampling rates and the relationship between cumulative progress and noise, and extend the current understanding of how simple stochastic gradient methods behave in bandit settings.},\n\tkeywords = {stochastic multi-armed bandits, policy gradient, stochastic gradient bandit algorithm, large stepsize, convergence},\n}\n\n","author_short":["Mei, J.","Dai, B.","Agarwal, A.","Vaswani, S.","Raj, A.","Szepesvári, C.","Schuurmans, D."],"key":"MeDaAgVaRaSzeSch24","id":"MeDaAgVaRaSzeSch24","bibbaseid":"mei-dai-agarwal-vaswani-raj-szepesvri-schuurmans-smallstepsnomoreglobalconvergenceofstochasticgradientbanditsforarbitrarylearningrates-2024","role":"author","urls":{" url":"http://papers.nips.cc/paper_files/paper/2024/hash/8815c983a1f46b477fe4fbf7042a3ba3-Abstract-Conference.html"," pdf":"https://proceedings.neurips.cc/paper_files/paper/2024/file/8815c983a1f46b477fe4fbf7042a3ba3-Paper-Conference.pdf"},"keyword":["stochastic multi-armed bandits","policy gradient","stochastic gradient bandit algorithm","large stepsize","convergence"],"metadata":{"authorlinks":{}},"downloads":4},"bibtype":"inproceedings","biburl":"https://sites.ualberta.ca/~szepesva/papers/p2.bib","dataSources":["JAZPSdjiP95Ah92D9"],"keywords":["stochastic multi-armed bandits","policy gradient","stochastic gradient bandit algorithm","large stepsize","convergence"],"search_terms":["small","steps","more","global","convergence","stochastic","gradient","bandits","arbitrary","learning","rates","mei","dai","agarwal","vaswani","raj","szepesvári","schuurmans"],"title":"Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates","year":2024,"downloads":5}