Distributed LLM Serving on Consumer-Grade GPUs by Reconciling Computation and Communication

Published in EMNLP Findings, 2025