Zaiwei Chen
Computer Science · Purdue University West Lafayette
Publications
59
Citations
242
Est. group size
—
Recurring co-author estimate
Active years
8
Publishing since 2019
Zaiwei Chen studies the mathematical foundations of reinforcement learning, a branch of machine learning where an agent learns to make decisions through trial-and-error feedback. Much of the work develops rigorous theory explaining why and how fast common learning algorithms (such as Q-learning, actor-critic methods, and policy gradient methods) converge to good solutions, often using tools from stochastic approximation and control theory. The research aims to provide provable performance guarantees for these algorithms rather than only empirical demonstrations.
Publication output has grown substantially over the last decade, rising from no recorded outputs in 2017-2018 to a peak of 10 per year in 2022-2023 and a sharp increase to 13 in 2026, indicating an active and increasing research pace.
Generated by claude-sonnet-5 from public bibliographic data · Jul 20, 2026
- A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies
2026
- Achieving $\varepsilon^{-2}$ Dependence for Average-Reward Q-Learning with a New Contraction Principle
Open MIND · 2026
- Achieving $\varepsilon^{-2}$ Dependence for Average-Reward Q-Learning with a New Contraction Principle
arXiv (Cornell University) · 2026
- Natural Hypergradient Descent: Algorithm Design, Convergence Analysis, and Parallel Implementation
Open MIND · 2026
- Natural Hypergradient Descent: Algorithm Design, Convergence Analysis, and Parallel Implementation
arXiv (Cornell University) · 2026
- Bridging the Gap Between Average and Discounted TD Learning
arXiv (Cornell University) · 2026
- Bridging the Gap Between Average and Discounted TD Learning
arXiv (Cornell University) · 2026
- Natural Policy Gradient as Doubly Smoothed Policy Iteration: A Bellman-Operator Framework
arXiv (Cornell University) · 2026
- Natural Policy Gradient as Doubly Smoothed Policy Iteration: A Bellman-Operator Framework
arXiv (Cornell University) · 2026
- Achieving $ε^{-2}$ Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions
arXiv (Cornell University) · 2026
- Achieving $ε^{-2}$ Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions
arXiv (Cornell University) · 2026
- Non-Asymptotic Convergence of Stochastic Iterative Algorithms: A Lyapunov Framework
arXiv (Cornell University) · 2026
- Non-Asymptotic Convergence of Stochastic Iterative Algorithms: A Lyapunov Framework
arXiv (Cornell University) · 2026
- Concentration of contractive stochastic approximation: Additive and multiplicative noise
The Annals of Applied Probability · 2025
- An approximate policy iteration viewpoint of actor–critic algorithms
Automatica · 2025
- arXiv (Cornell University)×36
- ACM SIGMETRICS Performance Evaluation Review×3
- Automatica×2
- Proceedings of the ACM on Measurement and Analysis of Computing Systems×2
- Open MIND×2
- Tengyu Xu
Computer Science · The Ohio State University
- Ziwei Guan
Computer Science · The Ohio State University
- Qinbo Bai
Computer Science · Purdue University West Lafayette
- Andrew Perrault
Computer Science · The Ohio State University
- Nan Jiang
Computer Science · Purdue University West Lafayette
This profile was generated automatically from public scholarly data (OpenAlex). Group size and activity levels are estimates derived from co-authorship patterns.
Last updated Jul 20, 2026.
Claim or correct this profile