vLLM on one RTX 5090

Serving open models locally on a single 32 GiB Blackwell card. Every number here was measured on that machine — single-stream decode, 300-token completions, thinking disabled, cold sample discarded.

Loading results.json