Shutian Luo

University of Virginia. Postdoctoral Researcher

shutianLuo.jpg

Email: ksy8xs@virginia.edu

Charlottesville, VA 22903

Shutian Luo is currently a Postdoctoral Researcher in the Department of Computer Science at the University of Virginia, working with Prof. Haiying Shen. He received his Ph.D. in Computer Application Technology from the University of Chinese Academy of Sciences in 2023, advised by Prof. Chengzhong Xu and working closedly with Prof. Huanle Xu. Previously, he was a Postdoctoral Associate at Yale University with Prof. Lin Zhong and Prof. Anurag Khandelwal.

His research lies at the intersection of AI systems and infrastructure and cloud-native AI platforms. His goal is to build hardware-aware systems that make modern AI workloads—especially LLMs and MoE models—efficient, scalable, and practical to deploy in real data centers.

Current interests include:

  • LLM & MoE infrastructure: Superchip- and MIG-based serving systems, KV-cache offloading, expert streaming.
  • Memory-centric AI systems: RDMA- and DSM-based hierarchical memory for cross-node / cross-GPU workloads.
  • Cloud-native runtimes for AI workloads: Resource management for microservices, serverless platforms, and latency-sensitive AI pipelines.

news

Mar 26, 2026 DirectKV was accepted to OSDI ’26! The first work on LLM KV-cache offloading using NVIDIA Superchips. Code coming soon.
Oct 15, 2025 Alibaba trace analysis about diffusion model was accepted by SoCC’25! Congrats to Yanyin!
Feb 15, 2025 Grad was accepted by HPCA’25! Congrats to Chenliao!
Jan 15, 2025 Imbres was accepted by ASPLOS’25!

selected publications

  1. OSDI’26
    No Buffer, No Bottleneck: Efficient Zero-Copy KV Cache Offloading for Long-Context LLMs
    Shutian Luo and Haiying Shen
    In Proceedings of the 20th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’26), 2026
  2. ASPLOS’25
    Embracing Imbalance: Dynamic Load Shifting among Microservice Containers in Shared Clusters
    Shutian Luo, Jianxiong Liao, Huanle Xu, and 2 more authors
    In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’25), 2025
  3. SoCC’25
    Understanding Diffusion Model Serving in Production: A Top-Down Analysis of Workload, Scheduling, and Resource Efficiency
    Yanying Lin, Shuaipeng Wu, Shutian Luo, and 8 more authors
    In Proceedings of the 16th ACM Symposium on Cloud Computing (SoCC ’25), 2025
  4. ISCA’24
    Derm: SLA-aware Resource Management for Highly Dynamic Microservices
    Liao Chen, Shutian Luo, Chenyu Lin, and 4 more authors
    In Proceedings of the 51st Annual IEEE/ACM International Symposium on Computer Architecture (ISCA ’24), 2024
  5. ASPLOS’23
    Erms: Efficient Resource Management for Shared Microservices with SLA Guarantees
    Shutian Luo, Huanle Xu, Kejiang Ye, and 5 more authors
    In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’23), 2023
  6. SoCC’22
    The Power of Prediction: Microservice Auto Scaling via Workload Learning
    Shutian Luo, Huanle Xu, Kejiang Ye, and 4 more authors
    In Proceedings of the 13th ACM Symposium on Cloud Computing (SoCC ’22), 2022
  7. SoCC’21
    Characterizing Microservice Dependency and Performance: Alibaba Trace Analysis
    Shutian Luo, Huanle Xu, Chengzhi Lu, and 6 more authors
    In Proceedings of the 12th ACM Symposium on Cloud Computing (SoCC ’21), 2021
  8. TOCS
    Optimizing Resource Management for Shared Microservices: A Scalable System Design
    Shutian Luo, Chenyu Lin, Kejiang Ye, and 5 more authors
    ACM Transactions on Computer Systems, 2024