Scholar
Zili Zhang
Google Scholar ID: 310QUvQAAAAJ
Peking University
Distributed system
Deep learning
Follow
Homepage
↗
Google Scholar
↗
Citations & Impact
All-time
Citations
742
H-index
12
i10-index
13
Publications
16
Co-authors
15
list available
Contact
Email
zhangzili1201@gmail.com
CV
Open ↗
GitHub
Open ↗
Publications
15 items
Resolving Mixed Single-Photon LiDAR Returns for Foreground-View and Hidden Scene Reconstruction
2026
Cited
0
Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models
2026
Cited
0
ExpertPlex: A High-Goodput Disaggregated Serving System for MoE LLMs with Adaptive Persistent Kernels
2026
Cited
0
SP-TransientBench: A Real-Captured Single Photon Perception Benchmark
2026
Cited
0
UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing
2026
Cited
0
BigMac: Breaking the Pareto Frontier of Compute and Memory in Multimodal LLM Training
2026
Cited
0
ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning
2026
Cited
0
Heddle: A Distributed Orchestration System for Agentic RL Rollout
2026
Cited
0
Load more
Resume
Academic Achievements
Published multiple papers as first or co-author in top-tier venues including NSDI, OSDI, SIGCOMM, and TOCS, such as:
“TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving” (Preprint)
“StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation” (Preprint)
“RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation” (TOCS'25, To appear)
“Fast Distributed Inference Serving for Large Language Models” (NSDI'26, To appear, Equal contribution)
“DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models” (SIGCOMM 2025)
“RLHFuse: Efficient RLHF Training for Large Language Models with Inter- and Intra-Stage Fusion” (NSDI 2025)
“dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM Serving” (OSDI 2024)
“Jolteon: Unleashing the Promise of Serverless for Serverless Workflows” (NSDI 2024)
“Fast Vector Query Processing for Large Datasets Beyond GPU Memory with Reordered Pipelining” (NSDI 2024)
“Ditto: Efficient Serverless Analytics with Elastic Parallelism” (SIGCOMM 2023)
“Fast, Approximate Vector Queries on Very Large Unstructured Datasets” (NSDI 2023)
“Transparent GPU Sharing in Container Clouds for Deep Learning Workloads” (NSDI)
Co-authors
13 total
Xin Jin
Peking University
Bingyang Wu
Peking University
Yinmin Zhong
Peking University
Chao Jin
Peking University
Yibo Zhu
StepFun
Ruidong Zhu
Peking University
Sun Peng
MiroMind.ai
Zheng Ge
Senior Researcher, StepFun