Research

I work on the theoretical foundations of learning, with a current focus on language models and the algorithms used to train, adapt, and interact with them. My current interests include:

  1. Foundations of Reinforcement Learning. Post-training improves language models through reinforcement learning from feedback and interaction, yet we still lack a clear understanding of when standard algorithms work with rich function approximation, how they generalize, and when they are computationally efficient. I develop theory to address these questions and guide more principled and efficient post-training algorithms, with particular focus on policy-gradient methods such as PPO, generalization under rich function approximation, and the computational limits of learning from feedback.
  2. Learning Language Models Through API Access. The inner workings of modern language models are often confidential even when the models themselves are available through APIs. I investigate whether this access can provably make it easier to learn a new model, with particular focus on low-rank structure and the connections between classical sequence models, such as hidden Markov models, and modern language models.

Selected Publications

  1. Learning Hidden Markov Models Using Conditional Samples
    with Sham Kakade, Akshay Krishnamurthy and Cyril Zhang
    COLT 2023
    arxiv

  2. Computational-Statistical Gaps in Reinforcement Learning
    with Daniel Kane, Sihan Liu and Shachar Lovett
    COLT 2022
    talk | arxiv | tweet

  3. Realizable Learning is All You Need
    with Max Hopkins, Daniel Kane and Shachar Lovett
    COLT 2022
    arxiv | tweet

  4. Optimality and Approximation with Policy Gradient Methods in Markov Decision Processes
    with Alekh Agarwal, Sham Kakade and Jason Lee
    COLT 2020
    colt | arxiv

All Publications

Google Scholar
  1. Cloning is as Hard as Learning for Stabilizer States
    with Nikhil Bansal and Matthias Caro
    COLT 2026, QCTiP 2026, TQC 2026
    arxiv

  2. Generalized de-Finiti Theorem: Lifting pure state algorithms to mixed states
    with Adam Bene Watts, John Bostanci, Matthias Caro, Daniel Grier and Sihan Liu
    Under submission, 2026

  3. Improved classical shadows from local symmetries in the Schur basis
    with Daniel Grier and Sihan Liu
    Under submission, 2025
    arxiv

  4. Do PAC-Learners Learn the Marginal Distribution?
    with Max Hopkins, Daniel Kane and Shachar Lovett
    ALT 2025
    arxiv

  5. Learning Hidden Markov Models Using Conditional Samples
    with Sham Kakade, Akshay Krishnamurthy and Cyril Zhang
    COLT 2023
    arxiv

  6. Exponential Hardness of Reinforcement Learning with Linear Function Approximation
    with Daniel Kane, Sihan Liu, Shachar Lovett, Csaba Szepesvari and Gillert Weisz
    COLT 2023
    arxiv

  7. Computational-Statistical Gaps in Reinforcement Learning
    with Daniel Kane, Sihan Liu and Shachar Lovett
    COLT 2022
    arxiv | tweet

  8. Realizable Learning is All You Need
    with Max Hopkins, Daniel Kane and Shachar Lovett
    COLT 2022
    arxiv | tweet

  9. Learning What To Remember
    with Robi Bhattacharjee
    ALT 2022
    arxiv

  10. Convergence of online k-means
    with Sanjoy Dasgupta and Geelon So
    AISTATS 2022
    arxiv

  11. Bilinear Classes: A Structural Framework for Provable Generalization in RL
    with Simon Du, Sham Kakade, Jason Lee, Shachar Lovett, Wen Sun and Ruosong Wang
    ICML 2021 (Long Talk)
    icml | arxiv | tweet

  12. On the Theory of Policy Gradient Methods: Optimality, Approximation, and Distribution Shift
    with Alekh Agarwal, Sham Kakade and Jason Lee
    JMLR 2021
    jmlr | arxiv

  13. Point Location and Active Learning: Learning Halfspaces Almost Optimally
    with Max Hopkins, Daniel Kane and Shachar Lovett
    FOCS 2020
    focs | arxiv

  14. Optimality and Approximation with Policy Gradient Methods in Markov Decision Processes
    with Alekh Agarwal, Sham Kakade and Jason Lee
    COLT 2020
    colt | arxiv

  15. Noise-tolerant, Reliable Active Classification with Comparison Queries
    with Max Hopkins, Daniel Kane and Shachar Lovett
    COLT 2020
    arxiv

  16. Q-learning with Function Approximation in Deterministic Systems
    with Simon Du Jason Lee and Ruosong Wang
    NeurIPS 2020
    arxiv

Lecture Notes

  1. Reinforcement Learning Theory
    ICTS Reinforcement Learning Bootcamp, 2025
    lecture notes | course page

  2. Quantum Codes and Applications to Complexity
    CPSC 648, Yale University, Fall 2024
    lecture notes | course page

Talks

  1. Learning Low-Rank Language Models with Conditional Queries (slides)
    Flatiron Institute, New York
    Toulouse School of Economics
    Cal Poly San Luis Obispo
    University of Victoria
    Barnard College
    Oberlin College
    University of Birmingham

  2. Computational-Statistical Gaps in Reinforcement Learning: (talk) (slides)
    Microsoft Research New York Seminar
    Yale Foundations of Data Science Seminar
    EnCORE Fall Retreat
    TTIC Machine Learning Seminar Series
    UCLA Big Data and Machine Learning weekly seminar
    Brown Robotics Group (George Konidaris's group)
    Berkeley RL Reading Group (Jiantao Jiao's group)

  3. Equivalence between Realizable and Agnostic Learning: (slides)
    Cornell Theory Seminar
    MSR New York ML Reading Group

  4. Theory of Generalization in Reinforcement Learning: (slides)
    RL Theory Seminar
    UCSD AI Seminar
    UCSD Theory Seminar