Queueing-theoretic models, stability guarantees, request admission, latency control, and traffic shaping for LLM workloads