Topics · OR for Generative AI

LLM Serving and Inference Scheduling

Batching, routing, admission control, prefill/decode scheduling, SLO-aware serving, and throughput optimization

Track this topic →
Living review: OR for Generative AI
35 recent papers · updated 2026-04-26
Read the review →