RESEARCHInvestigateNEXT 12 MONTHS
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Researchers propose Cascade, an LLM serving framework that uses SLO-aware latency budgets to optimize inference cost and throughput.
Open source