Xupeng Miao
Computer Science · Purdue University West Lafayette
Publications
100
Citations
1,169
Est. group size
—
Recurring co-author estimate
Active years
8
Publishing since 2019
Xupeng Miao works on computer systems for efficiently running and training large machine learning models, especially large language models (LLMs). His work covers topics such as speeding up LLM inference and serving, managing GPU memory and caches, distributing training across many machines, and building infrastructure for reinforcement learning and agentic AI systems. This research is aimed at making large-scale AI systems faster, cheaper, and more scalable to run.
Publication output has grown substantially over the last decade, rising from near zero in 2017-2018 to a peak of 24 in 2024, with continued high activity (12-15 per year) in 2025-2026.
Generated by claude-sonnet-5 from public bibliographic data · Jul 20, 2026
- Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
arXiv (Cornell University) · 2026
- Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
arXiv (Cornell University) · 2026
- AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
2026
- Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs
arXiv (Cornell University) · 2026
- Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs
arXiv (Cornell University) · 2026
- DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning
arXiv (Cornell University) · 2026
- AutoSci: A Memory-Centric Agentic System for the Full Scientific Research Lifecycle
arXiv (Cornell University) · 2026
- DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning
arXiv (Cornell University) · 2026
- AutoSci: A Memory-Centric Agentic System for the Full Scientific Research Lifecycle
arXiv (Cornell University) · 2026
- Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster Large Language Model Serving
arXiv (Cornell University) · 2026
- Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster Large Language Model Serving
arXiv (Cornell University) · 2026
- Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training
arXiv (Cornell University) · 2026
- Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training
arXiv (Cornell University) · 2026
- Adaptive Resource Management and Quality Control for Streaming Video Generation
arXiv (Cornell University) · 2026
- Adaptive Resource Management and Quality Control for Streaming Video Generation
arXiv (Cornell University) · 2026
- arXiv (Cornell University)×47
- Proceedings of the VLDB Endowment×6
- 2022 IEEE 38th International Conference on Data Engineering (ICDE)×4
- IEEE Transactions on Knowledge and Data Engineering×3
- The VLDB Journal×3
- Quentin Anthony
Computer Science · The Ohio State University
- Arghadip Das
Computer Science · Purdue University West Lafayette
- Lang Xu
Computer Science · The Ohio State University
- Qi Guo
Computer Science · Purdue University West Lafayette
- Xuwei Tan
Computer Science · The Ohio State University
This profile was generated automatically from public scholarly data (OpenAlex). Group size and activity levels are estimates derived from co-authorship patterns.
Last updated Jul 20, 2026.
Claim or correct this profile