ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving

2026-07-01Distributed, Parallel, and Cluster Computing

Distributed, Parallel, and Cluster Computing
AI summary

The authors address a problem in how requests are assigned to workers in systems that run large language models with mixture-of-experts (MoE) architectures. They note that simply balancing the number of tasks per worker does not account for differences in processing time caused by different expert usage. To fix this, they designed ELDR, which predicts which experts a request will use and assigns the task to the best matching worker with the least load. Their approach, tested on multiple GPUs and models, shortened waiting times without changing the output quality.

prefill-decodedisaggregated servingmixture-of-experts (MoE)decode routerexpert localityK-means clusteringKV cachevLLMGPU deployment
Authors
Sangjin Choi, Sukmin Cho, Yifan Xiong, Ziyue Yang, Youngjin Kwon, Peng Cheng
Abstract
In prefill-decode (PD) disaggregated LLM serving, each request is assigned to a decode worker after prefill. Existing decode routers balance only load; for mixture-of-experts (MoE) models this is incomplete: equally loaded workers can differ in latency, since each decode step loads the weights of every distinct expert its batch activates. We present ELDR, an expert-locality-aware decode router for PD-disaggregated MoE serving. From a request's prefill expert activations, ELDR builds an expert signature predicting the experts it will activate during generation. Offline, balanced K-means partitions signature space across decode workers; online, locality-band routing sends each request to the least-loaded worker among those best matching its signature. A signature cache, co-indexed with the KV cache at KV-block granularity, keeps signatures exact under prefix caching. Implemented in vLLM and evaluated on deployments of up to 40 GPUs, ELDR reduces median TPOT by 5.9-13.9% over the strongest of four load-balancing baselines across three MoE models and two workloads, with model outputs unchanged.