| vLLM and KServe can serve the same OpenAI-compatible model path. | Serving layers compared |
LiteLLM owns tenant keys, budgets, spend, and the public /v1 facade. | LiteLLM guide |
| Gateway API plus GIE routes by inference signals instead of plain HTTP load balancing. | Inference gateway guide |
| Kueue and KEDA keep GPU admission, quota, and queue-depth autoscaling explicit. | GPU debugging guide |
| Benchmarks tie latency, throughput, GPU use, and KV-cache pressure to the same run. | Benchmark results |