I am an associate professor in the Department of Computer Science at ETH Zurich and a part of the Systems group. Prior to this I was a faculty member at UW-Madison and a part of the madSystems group. I completed my PhD from UC Berkeley where I was advised by Ion Stoica and Mike Franklin. I also have a Masters from University of Illinois at Urbana-Champaign.

News: I am looking for PhD students and postdocs to join my group — please get in touch with your CV if you are interested!

Current research areas

My group works at the intersection of computer systems and machine learning, in both directions. On one side, we work across the stack to make modern ML workloads faster, cheaper, and more reliable: from efficient ML model architecture design to GPU energy modeling. On the other, we use machine learning inside core systems, in areas such as memory tiering and cloud configuration tuning.

Group

  • Rutwik Jain — co-advised w/ Matt Sinclair
  • Brandon Tran — co-advised w/ Matt Sinclair
  • Minghao Yan
  • Johannes Freischuetz
  • Tzu-Tao Chang
  • Fanchao Chen
  • Tareq Mahmood
  • Seth Ockerman

Alumni

+

PhD

  • Song Bian → NVIDIA Research Labs
  • Konstantinos Kanellis → AWS Learned Systems Group
  • Jason Mohoney → Post-doc at MIT
  • Saurabh Agarwal → Post-doc at UT-Austin

Post-doctoral researchers

  • Pengfei Zheng (co-advised with Aditya Akella) → Huawei Technologies

MS

  • Devesh Sarda → Databricks
  • Aditi Singh → Nutanix
  • Mohil Patel → Oracle
  • Rachit Tibrewal
  • Olesia Elfimova → Dropbox
  • Adarsh Kumar → Amazon Alexa AI
  • Arjun Balasubramanian → Amazon AWS

Selected recent publications

Fanchao Chen, Ziheng Jiang, Ziyun Wei, Zheng Zhong, Du Li, Chi Zhang, Haibin Lin, Shivaram Venkataraman. Towards Full Pipeline FP8 Reinforcement Learning for LLMs. COLM 2026.
Minghao Yan, Zhuang Wang, Zhen Jia, Shivaram Venkataraman, Yida Wang. PLoRA: Efficient Concurrent LoRA Training for Large Language Models. ICML 2026.
Rutwik Jain, Yiwei Jiang, Matt Sinclair, Shivaram Venkataraman. Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters. ACM SIGMETRICS 2026.
Brandon Tran, Matthias Maiterth, Woong Shin, Matt Sinclair, Shivaram Venkataraman. Wattchmen: Watching the Wattchers — High Fidelity, Flexible GPU Energy Modeling. ACM ICS 2026.
Konstantinos Kanellis, Sujay Yadalam, Hayden Coffey, Shivaram Venkataraman, Michael Swift. From Good to Great: Parameter Tuning in Memory Tiering Systems. IEEE Transactions on Computers 2026.
Song Bian, Tao Yu, Shivaram Venkataraman, Youngsuk Park. Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs. ICLR 2026.
Saurabh Agarwal, Bodun Hu, Anyong Mao, Aditya Akella, Shivaram Venkataraman. SYMPHONY: Enabling Compute-Memory Disaggregation in LLM Serving Systems. NSDI 2026.
Jason Mohoney, Devesh Sarda, Mengze Tang, et al., Shivaram Venkataraman. Quake: Adaptive Indexing for Vector Search. OSDI 2025.
Minghao Yan, Saurabh Agarwal, Shivaram Venkataraman. Decoding Speculative Decoding. NAACL 2025 · SAC Award for Generation.
Johannes Freischuetz, Konstantinos Kanellis, Brian Kroth, Shivaram Venkataraman. TUNA: Tuning Unstable and Noisy Cloud Applications. EuroSys 2025.
Seth Ockerman, Amal Gueroudji, Tanwi Mallick, Yixuan He, Line Pouchard, Rob Ross, Shivaram Venkataraman. PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training. Supercomputing 2025.

Please see Google Scholar for a complete list.