Shutian Luo
University of Virginia. Postdoctoral Researcher
Email: ksy8xs@virginia.edu
Charlottesville, VA 22903
Shutian Luo is currently a Postdoctoral Researcher in the Department of Computer Science at the University of Virginia, working with Prof. Haiying Shen. He received his Ph.D. in Computer Application Technology from the University of Chinese Academy of Sciences in 2023, advised by Prof. Chengzhong Xu and working closedly with Prof. Huanle Xu. Previously, he was a Postdoctoral Associate at Yale University with Prof. Lin Zhong and Prof. Anurag Khandelwal.
His research lies at the intersection of AI systems and infrastructure and cloud-native AI platforms. His goal is to build hardware-aware systems that make modern AI workloads—especially LLMs and MoE models—efficient, scalable, and practical to deploy in real data centers.
Current interests include:
- LLM & MoE infrastructure: Superchip- and MIG-based serving systems, KV-cache offloading, expert streaming.
- Memory-centric AI systems: RDMA- and DSM-based hierarchical memory for cross-node / cross-GPU workloads.
- Cloud-native runtimes for AI workloads: Resource management for microservices, serverless platforms, and latency-sensitive AI pipelines.
news
| Mar 26, 2026 | DirectKV was accepted to OSDI ’26! The first work on LLM KV-cache offloading using NVIDIA Superchips. Code coming soon. |
|---|---|
| Oct 15, 2025 | Alibaba trace analysis about diffusion model was accepted by SoCC’25! Congrats to Yanyin! |
| Feb 15, 2025 | Grad was accepted by HPCA’25! Congrats to Chenliao! |
| Jan 15, 2025 | Imbres was accepted by ASPLOS’25! |
selected publications
- OSDI’26No Buffer, No Bottleneck: Efficient Zero-Copy KV Cache Offloading for Long-Context LLMsIn Proceedings of the 20th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’26), 2026
- ASPLOS’25Embracing Imbalance: Dynamic Load Shifting among Microservice Containers in Shared ClustersIn Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’25), 2025
- SoCC’25Understanding Diffusion Model Serving in Production: A Top-Down Analysis of Workload, Scheduling, and Resource EfficiencyIn Proceedings of the 16th ACM Symposium on Cloud Computing (SoCC ’25), 2025
- ISCA’24Derm: SLA-aware Resource Management for Highly Dynamic MicroservicesIn Proceedings of the 51st Annual IEEE/ACM International Symposium on Computer Architecture (ISCA ’24), 2024
- ASPLOS’23Erms: Efficient Resource Management for Shared Microservices with SLA GuaranteesIn Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’23), 2023
- SoCC’22The Power of Prediction: Microservice Auto Scaling via Workload LearningIn Proceedings of the 13th ACM Symposium on Cloud Computing (SoCC ’22), 2022
- SoCC’21Characterizing Microservice Dependency and Performance: Alibaba Trace AnalysisIn Proceedings of the 12th ACM Symposium on Cloud Computing (SoCC ’21), 2021
- TOCSOptimizing Resource Management for Shared Microservices: A Scalable System DesignACM Transactions on Computer Systems, 2024